File query method and device, equipment and storage medium

By using metadata units and preset execution engines in the big data query system to directly obtain target files from the file system, the problems of lengthy query process and resource consumption in traditional methods are solved, and efficient and low-cost file query is achieved.

CN120596437APending Publication Date: 2025-09-05BAIDU (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510963323.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional file query methods in the big data field have lengthy query processes and complex data processing, leading to increased storage resource costs and excessive resource consumption.

Method used

The target address of the query request is obtained through the metadata unit, and the target file is directly obtained from the preset file system using the preset execution engine, avoiding data format conversion and external system dependence, and using the same service node to complete the query task.

Benefits of technology

It reduces query costs, improves file query efficiency, reduces data transmission risks and resource consumption, and improves query stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596437A_ABST
    Figure CN120596437A_ABST
Patent Text Reader

Abstract

The invention provides a file query method and device, equipment and a storage medium, and relates to the technical field of data processing, in particular to the fields of artificial intelligence, big data and the like. According to the specific implementation scheme, a query request is obtained, and the query request is used for requesting to query a target file from a preset file system; obtaining a target address corresponding to the query request through a metadata unit; loading the target address to a preset execution engine through an analysis unit; and obtaining a target file corresponding to the query request from a preset file system through an execution unit by utilizing the preset execution engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence, big data, etc. Background Art

[0002] In the big data space, with the growing demand for complex query tasks, traditional file query solutions have lengthy query processes and complex data processing. Furthermore, traditional methods rely on multiple storage mechanisms, increasing storage resource costs. Therefore, an efficient and cost-effective file query method is urgently needed. Summary of the Invention

[0003] The present disclosure provides a file query method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, a file query method is provided, comprising:

[0005] Obtaining a query request, wherein the query request is used to request to query a target file from a preset file system;

[0006] Obtaining the target address corresponding to the query request through the metadata unit;

[0007] Loading the target address into a preset execution engine through a parsing unit;

[0008] The target file targeted by the query request is obtained from the preset file system through the execution unit and by utilizing the preset execution engine.

[0009] According to another aspect of the present disclosure, there is provided a file query device, comprising:

[0010] A request obtaining unit, configured to obtain a query request, wherein the query request is used to request to query a target file from a preset file system;

[0011] A metadata unit, used to obtain a target address corresponding to the query request;

[0012] A parsing unit, configured to load the target address into a preset execution engine;

[0013] The execution unit is configured to obtain the target file targeted by the query request from a preset file system using the preset execution engine.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0020] In this way, after obtaining a query request, the disclosed solution can obtain the target address corresponding to the query request through the metadata unit. Furthermore, the target address is loaded into the preset execution engine through the parsing unit. Finally, the target file targeted by the query request is obtained from the preset file system through the execution unit and by using the preset execution engine and the loaded target address. Since the execution step of the disclosed solution does not require operations such as data format conversion, compared with the file query method based on the OLAP system, the disclosed solution avoids the problem of excessive resource consumption, reduces query costs, and at the same time improves file query efficiency.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0023] Figure 1 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 1 ;

[0024] Figure 2 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 2 ;

[0025] Figure 3 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 3 ;

[0026] Figure 4 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 4 ;

[0027] Figure 5 This is a schematic diagram of the overall process of file query according to an embodiment of the present application;

[0028] Figure 6 This is a schematic diagram of the structure of a file query device 600 according to an embodiment of the present disclosure. Figure 1 ;

[0029] Figure 7 This is a schematic diagram of the structure of a file query device 600 according to an embodiment of the present disclosure. Figure 2 ;

[0030] Figure 8 A schematic block diagram of an example electronic device 800 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0031] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0033] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0034] The following describes the related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any way, and all of them fall within the protection scope of the embodiments of the present disclosure.

[0035] In today's big data landscape, Online Analytical Processing (OLAP) systems have become indispensable tools for executing complex queries in big data scenarios. Specifically, a typical OLAP architecture consists of a data ingestion layer, a data transformation layer, a storage engine layer, and a query engine layer. The storage engine layer uses columnar storage technology, which improves query performance when processing large amounts of data. The specific query process is as follows:

[0036] First, the data ingestion layer extracts files from distributed file systems (such as the Hadoop Distributed File System (HDFS)) or data warehouses as the basis for subsequent processing. Second, the data transformation layer cleans, transforms, and reorganizes the data through an extract-transform-load (ETL) process, storing the processed data in the storage engine layer. Finally, the query engine layer, based on a massively parallel processing (MPP) architecture, calls the required files from the storage engine layer to implement distributed queries.

[0037] However, file query methods based on OLAP systems require a tight coupling between data storage and computation. This necessitates reorganizing and storing data in HDFS into a proprietary format, increasing the complexity and time cost of data processing. Furthermore, due to the need to maintain both the original data and the OLAP system's data, file query methods also suffer from data redundancy. Furthermore, these specialized storage engines often require independent computing resources to operate, resulting in excessive resource consumption and significant cost increases.

[0038] Based on this, the disclosed solution proposes a file query method, which can directly perform file queries on a distributed file system through a service node. This process does not require the use of an additional system. In other words, this process does not need to rely on an OLAP system, effectively reducing query costs. At the same time, it avoids unnecessary file processing steps, thereby effectively improving query efficiency.

[0039] Specifically, Figure 1 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0040] Furthermore, the method includes at least part of the following contents. Figure 1 Shown, including:

[0041] Step S101: Obtain a query request.

[0042] Here, the query request is used to request to query a target file from a preset file system.

[0043] It is understandable that the query request may be an information request for a preset file system, for example, a query request expressed using Structured Query Language (SQL).

[0044] Furthermore, in one example, the preset file system described above may include a distributed file system, such as HDFS, Andrew File System (AFS), etc. The target file includes a data set stored on a computer storage device, and the data type in the target file may include an integer type, a string type, etc. The present disclosure does not impose any restrictions on the data type.

[0045] Step S102: Obtain the target address corresponding to the query request through the metadata unit.

[0046] It can be understood that the target address is used to represent the address of the target file to be queried in the preset file system.

[0047] Step S103: Load the target address into a preset execution engine through a parsing unit.

[0048] Here, the preset execution engine may be a component that has been configured and is used to perform a specific task, and the disclosed solution does not impose any specific restrictions on the engine.

[0049] Step S104: The target file targeted by the query request is obtained from the preset file system through the execution unit and the preset execution engine.

[0050] In this way, after obtaining a query request, the disclosed solution can obtain the target address corresponding to the query request through the metadata unit. Furthermore, the target address is loaded into the preset execution engine through the parsing unit. Finally, the target file targeted by the query request is obtained from the preset file system through the execution unit and by using the preset execution engine and the loaded target address. Since the execution step of the disclosed solution does not require operations such as data format conversion, compared with the file query method based on the OLAP system, the disclosed solution avoids the problem of excessive resource consumption, reduces query costs, and at the same time improves file query efficiency.

[0051] Furthermore, in a specific example, at least two of the metadata unit, the parsing unit, and the execution unit for processing the query request are located in the same service node.

[0052] Here, it should be noted that the service node can be a physical server, cloud server, containerized application (such as Docker container), etc. with certain computing power and storage resources. The specific implementation form of the service node of the disclosed solution is not limited.

[0053] For example, in one example, the metadata unit and the parsing unit are located in the first service node, and the execution unit is located in the second service node. At this time, based on the first service node, the target address corresponding to the query request can be obtained, and the target address is loaded into the preset execution engine. Further, based on the second service node, and using the preset execution engine and the loaded target address, the target file targeted by the query request is obtained from the preset file system.

[0054] Alternatively, in another example, the parsing unit and the execution unit are located in the first service node, and the metadata unit is located in the second service node. At this time, the target address corresponding to the query request is obtained based on the second service node, and the target address is loaded into the preset execution engine based on the first service node, and the preset execution engine is used to obtain the target file targeted by the query request from the preset file system.

[0055] Furthermore, in one example, the metadata unit, parsing unit, and execution unit for processing the query request are all located in the same service node. In other words, for a query request, the query request is served by the units in the same service node, so that the required target file is finally queried through the units in the same service node.

[0056] In this way, since the disclosed solution can utilize the processing units deployed on the service node to complete the query task of the target file, it effectively avoids the additional query cost caused by relying on the external storage system, reduces the query cost, and at the same time, improves the query efficiency of the file.

[0057] Furthermore, since the disclosed solution can complete the query task of the target file at the same service node, it can also significantly reduce the need for data transmission between different nodes, thereby reducing potential risks during data transmission (such as data loss, leakage or damage).

[0058] Furthermore, in a specific example, the target file is a big data report, and the service node is one of multiple service nodes included in a big data query system (e.g., a distributed query system). In other words, the disclosed solution can be used in big data report query scenarios. Thus, compared to existing big data report query methods, the disclosed solution can effectively reduce the query cost of big data reports and effectively improve the query efficiency of big data reports.

[0059] Figure 2 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0060] Furthermore, the method includes at least part of the following contents. Figure 2 Shown, including:

[0061] Step S201: Obtain a query request.

[0062] Here, the query request is used to request to query a target file from a preset file system.

[0063] It should be noted that, for relevant examples of query requests, preset file systems and target files, please refer to the above description and will not be repeated here.

[0064] Step S202: Obtain target metadata corresponding to the query request through the metadata unit.

[0065] Here, in one example, the target metadata includes but is not limited to at least one of the following: basic information, attributes, dependency, association or combination relationship of the target file targeted by the query request.

[0066] Furthermore, in one example, the metadata corresponding to each file can be pre-stored in a preset metadata table, for example, in a preset metadata table pre-stored in a distributed database. Then, after receiving a query request, a corresponding metadata search algorithm can be used, and based on functions such as database sub-tables or views, the target metadata corresponding to the query request can be queried from the preset metadata table.

[0067] It should be noted that, in one example, the metadata search algorithm may include a binary search algorithm, a block search algorithm, an interpolation search algorithm, a hash search, an index search or a binary sorted tree search algorithm, etc. The present disclosure does not limit the specific metadata search algorithm.

[0068] Step S203: Determine the target address corresponding to the target metadata based on a preset mapping table through the metadata unit.

[0069] Here, the mapping relationship between metadata and file addresses is stored in the preset mapping table. That is, the metadata in the preset mapping table may correspond to file addresses, and these file addresses may point to the physical locations of the files in the preset file system.

[0070] For example, in one example, the preset mapping table can be represented by a predefined or constructed data structure. Furthermore, based on the determined target metadata and using the preset mapping table, the target address of the target file corresponding to the target metadata can be determined.

[0071] Step S204: Load the target address into a preset execution engine through a parsing unit.

[0072] Here, it should be noted that relevant examples of the parsing unit and the execution engine can be found in the above description and will not be repeated here.

[0073] Step S205: The target file targeted by the query request is obtained from the preset file system through the execution unit and the preset execution engine.

[0074] In this way, the disclosed solution can quickly parse the target metadata corresponding to the query request through the metadata unit, avoiding the tedious traversal and matching process in traditional data retrieval and improving the efficiency of querying the target metadata.

[0075] Furthermore, after obtaining the target metadata, the disclosed solution can also directly determine the target address of the target file corresponding to the target metadata based on a pre-set preset mapping table. By directly mapping the query, the redundant links in the information query process are effectively reduced, making the information query faster and more accurate, and laying the foundation for efficient file query.

[0076] Furthermore, in a specific example, the target file can be obtained by querying in the following manner. Specifically, the above-mentioned loading of the target address into the preset execution engine (such as step S103 or step S204) can specifically include:

[0077] The target address is loaded into a preset execution engine, and a target view table is created in the preset execution engine.

[0078] That is, the target address is loaded into the preset execution engine, and a target view table for calling the target file is created in the preset execution engine based on the target address, so that the target file can be quickly queried using the target view during the execution phase.

[0079] Here, it should be noted that the target view table can be a virtual table used to store query statements. For example, in this example, it is used to store query statements that can carry the target address. In this way, it is convenient to perform file queries through the stored query statements to obtain query results.

[0080] Furthermore, in a specific example, after the target view table is created, the target file targeted by the query request may be obtained in the following manner. Specifically, the above-described obtaining the target file targeted by the query request from the preset file system by the execution unit and the preset execution engine (such as step S104 or step S205) may specifically include:

[0081] The target file targeted by the query request is obtained from the preset file system through the execution unit and by using the target view table in the preset execution engine.

[0082] That is, in this example, the target file targeted by the query request can be found in the preset file system by the execution unit and by using the target view table.

[0083] In this way, the disclosed solution creates a target view table corresponding to the target address in a preset execution engine, and uses the target view table to abstract and encapsulate query statements, etc., thereby providing a more efficient and convenient data access interface.

[0084] Moreover, the target view table in the disclosed solution is a logical view containing query data and can carry relevant information of the physical storage location. Therefore, the execution unit can quickly utilize the relevant information in the target view table to quickly and directly locate the target file in the preset file system, thereby improving the speed and efficiency of data access.

[0085] Figure 3 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 3 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figures 1 to 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0086] Furthermore, the method includes at least part of the following contents. Figure 3 Shown, including:

[0087] Step S301: Obtain a query request.

[0088] Here, the query request is used to request to query a target file from a preset file system.

[0089] It should be noted that, for relevant examples of query requests, preset file systems and target files, please refer to the above description and will not be repeated here.

[0090] Step S302: Obtain target metadata corresponding to the query request through the metadata unit.

[0091] Step S303: Determine the target address corresponding to the target metadata based on a preset mapping table through the metadata unit.

[0092] Here, the preset mapping table stores a mapping relationship between metadata and file addresses.

[0093] It should be noted that, for relevant examples of metadata units and preset mapping tables, please refer to the above description and will not be repeated here.

[0094] Step S304: Load the target address into a preset execution engine through a parsing unit.

[0095] Here, it should be noted that relevant examples of the parsing unit and the execution engine can be found in the above description and will not be repeated here.

[0096] Step S305: The target file targeted by the query request is obtained from the preset file system through the execution unit and the preset execution engine.

[0097] Here, it should be noted that, for relevant examples of the execution unit, please refer to the above description, which will not be repeated here.

[0098] Step S306: The acquired target file is saved to the cache unit through the execution unit.

[0099] Step S307: Send the target file via the cache unit. For example, the target file is sent to the requester of the query request via the cache unit.

[0100] It should be noted that a cache unit can include the following features:

[0101] (1) Optimize the storage and computing efficiency of cache units through intelligent partitioning strategies.

[0102] For example, the intelligent partitioning strategy may include two partitioning strategies: partitioning by time and partitioning by hash.

[0103] Here, time partitioning is mainly achieved through pre-partitioning. Pre-partitioning is to divide the cache unit according to the time range in advance when creating the cache unit. Each time period corresponds to a partition, so that the data is distributed in different partitions in an orderly manner according to time when stored, thereby reducing the amount of data that needs to be scanned during query and improving query efficiency.

[0104] Furthermore, hash partitioning can utilize a hash algorithm. For example, all data is first converted into digital form, i.e., a hash value, via a hash function. Then, based on the hash value, the data is appropriately allocated to each partition. For example, in one example, the partition to which the data is ultimately assigned can be determined by dividing the hash value by the total number of partitions and taking the remainder. For example, if the cache unit has four partitions, the remainder obtained by dividing the hash value calculated by the hash function by 4 indicates the target partition where the data should be stored.

[0105] (2) Ability to achieve mixed storage of hot data and cold data.

[0106] Here, hot data refers to data that is frequently accessed and critical to the business or application. This data requires fast and efficient access and processing, so it is typically stored on high-performance, low-latency storage devices, such as solid-state drives (SSDs). Cold data refers to data that is less frequently accessed and less critical to the business and application. Cold data may not be frequently accessed or used within a specific time period and can usually be stored for a long time. Therefore, it is suitable for storage on storage devices with lower storage costs and higher capacity, such as hard disk drives (HDDs).

[0107] By configuring SSD and HDD for the cache unit, mixed storage of hot data and cold data in the cache unit can be achieved.

[0108] (3) Use column-based metadata indexing technology to store and retrieve data in the cache unit.

[0109] Specifically, column-based storage is a method of storing data in columns. In column-based storage, data is organized by column, meaning each column contains the values ​​of the corresponding fields in all data. Metadata indexes refer to descriptive information about data (such as data type and attributes). In this disclosed solution, by creating an index for each column in the cache unit, data retrieval can be accelerated and data location efficiency improved.

[0110] Furthermore, column-based metadata indexes combine the advantages of column-based storage and metadata indexes. For example, consider an e-commerce database that contains a product table containing order amounts and order dates. In this case, the column-based metadata in the product table includes the data type of the order amount column (e.g., integer) and the data type of the order date column (e.g., timestamp). Column-based metadata indexes can quickly determine the data type of each column, thereby optimizing data query and processing efficiency.

[0111] In this way, the disclosed solution can save the acquired target file into the cache unit. Thus, when the same target file is subsequently queried, the time required to read data from the preset file system can be reduced. In other words, when the target file is queried again, it can be read directly from the cache unit, thereby speeding up the response speed and saving query resources.

[0112] Figure 4 This is a schematic flow chart of a file query method according to an embodiment of the present application. Figure 4 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figures 1 to 3 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0113] Furthermore, the method includes at least part of the following contents. Figure 4 Shown, including:

[0114] Step S401: Obtain a query request.

[0115] Here, the query request is used to request to query a target file from a preset file system.

[0116] It should be noted that, for relevant examples of query requests, preset file systems and target files, please refer to the above description and will not be repeated here.

[0117] For example, in a specific example, it may be determined whether the target file is stored in the cache unit. Specifically, the method further includes:

[0118] The cache unit determines whether the target file required by the query request exists in the currently cached files.

[0119] That is, in one example, the cache unit first responds to the query request and determines whether the target file required by the query request exists among the currently cached files. Furthermore, if the target file required by the query request is determined to exist, the cache unit directly returns the stored target file to the requesting party. Otherwise, if the target file required by the query request is determined to not exist, the cache unit transmits the query request to the metadata unit, which then executes the subsequent file query process based on the query request.

[0120] In this way, when the target file required by the query request is stored in the cache unit, the disclosed solution can directly read the data from the cache without accessing the preset file system. When the target file required by the query request is not stored in the cache unit, the subsequent file query process is performed based on the query request. This mechanism can effectively reduce the response time of the query request, improve the overall file query speed, and at the same time, save query resources.

[0121] Step S402: When it is determined through the cache unit that the target file required by the query request does not exist in the currently cached files, the query request is sent to the metadata unit.

[0122] It should be noted that, for relevant examples of the cache unit, please refer to the above description, which will not be repeated here.

[0123] Step S403: Obtain target metadata corresponding to the query request through the metadata unit.

[0124] Step S404: Determine the target address corresponding to the target metadata based on a preset mapping table through the metadata unit.

[0125] Here, the preset mapping table stores a mapping relationship between metadata and file addresses.

[0126] Here, it should be noted that, for relevant examples of metadata units and preset mapping tables, please refer to the above description, which will not be repeated here.

[0127] Step S405: Load the target address into a preset execution engine through a parsing unit.

[0128] Here, it should be noted that relevant examples of the parsing unit and the execution engine can be found in the above description and will not be repeated here.

[0129] Step S406: The target file targeted by the query request is obtained from the preset file system through the execution unit and the preset execution engine.

[0130] Here, it should be noted that, for relevant examples of the execution unit, please refer to the above description, which will not be repeated here.

[0131] Step S407: The acquired target file is saved into the cache unit through the execution unit.

[0132] Step S408: Send the target file through the cache unit.

[0133] In this way, when the cache unit does not hit the target file, the disclosed solution sends the query request to the metadata unit to implement subsequent file query tasks, and then can locate the target address of the target file, reducing the waiting time of the file query and improving the efficiency of the file query.

[0134] The following is a further detailed description of the disclosed solution with reference to specific examples. Specifically, Figure 5 2 is a schematic diagram of the overall process of file query according to an embodiment of the present application.

[0135] like Figure 5 As shown, the file request end sends a query request to the cache unit. If the target file required by the query request exists in the files cached by the cache unit, the target file is directly fed back to the file request end; if the target file required by the query request does not exist in the files cached by the cache unit, the query request is sent to the metadata unit.

[0136] Furthermore, based on the query request, the metadata unit is used to obtain target metadata corresponding to the query request. Based on the target metadata, the metadata unit is used to obtain a target address corresponding to the target metadata from a preset mapping table. Here, the preset mapping table stores a mapping relationship between metadata and file addresses.

[0137] Furthermore, the parsing unit loads the target address into the preset execution engine and creates a target view table. Based on the target view table, the execution unit drives the preset execution engine to retrieve the target file targeted by the query request from the preset file system and feeds the target file back to the execution unit.

[0138] Furthermore, the target file is stored in the cache unit through the execution unit, and finally, the target file is sent to the file request end through the cache unit.

[0139] Here, the metadata unit, parsing unit and execution unit that execute the current query request are in the same service node; in other words, the disclosed solution adopts a back-end service model to provide a single-node query function, successfully eliminating the OLAP system and achieving a significant cost reduction.

[0140] In addition, since the disclosed solution eliminates the limitations of the OLAP system, it ensures the stability and reliability of file query while taking cost-effectiveness into consideration.

[0141] The disclosed solution also provides a file query device 600, such as Figure 6 Shown, including:

[0142] The request acquisition unit 601 is configured to acquire a query request, wherein the query request is configured to request a target file to be searched from a preset file system;

[0143] Metadata unit 602, used to obtain the target address corresponding to the query request;

[0144] The parsing unit 603 is used to load the target address into a preset execution engine;

[0145] The execution unit 604 is configured to use the preset execution engine to obtain the target file targeted by the query request from a preset file system.

[0146] In a specific example of the disclosed solution, at least two of the metadata unit, the parsing unit, and the execution unit for processing the query request are located in the same service node.

[0147] In a specific example of the disclosed solution, the target file is a big data report; and the service node is one of a plurality of service nodes included in the big data query system.

[0148] In a specific example of the disclosed solution, the metadata unit 602 is specifically used to:

[0149] Obtaining target metadata corresponding to the query request;

[0150] Based on a preset mapping table, a target address corresponding to the target metadata is determined; wherein the preset mapping table stores a mapping relationship between metadata and file addresses.

[0151] In a specific example of the disclosed solution, the parsing unit 603 is specifically configured to load the target address into a preset execution engine and create a target view table in the preset execution engine;

[0152] The execution unit 604 is specifically configured to obtain the target file targeted by the query request from the preset file system using the target view table in the preset execution engine.

[0153] In a specific example of the present disclosure, Figure 7 As shown, it also includes a cache unit 605; wherein,

[0154] The execution unit 604 is further configured to save the acquired target file into a cache unit;

[0155] The cache unit 605 is configured to send the target file.

[0156] In a specific example of the disclosed solution, the cache unit 605 is further configured to:

[0157] Determine whether the target file targeted by the query request exists in the currently cached files.

[0158] In a specific example of the disclosed solution, the cache unit 605 is further configured to:

[0159] When it is determined that the target file targeted by the query request does not exist in the currently cached files, the query request is sent to the metadata unit.

[0160] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0161] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0162] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0164] like Figure 8As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0165] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0166] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the file query method. For example, in some embodiments, the file query method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the file query method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the file query method by any other suitable means (e.g., by means of firmware).

[0167] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0168] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0169] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0171] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0172] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0173] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0174] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A file query method, comprising: Obtaining a query request, wherein the query request is used to request to query a target file from a preset file system; Obtaining the target address corresponding to the query request through the metadata unit; Loading the target address into a preset execution engine through a parsing unit; The target file targeted by the query request is obtained from the preset file system through the execution unit and by utilizing the preset execution engine.

2. The method according to claim 1, wherein At least two of the metadata unit, the parsing unit, and the execution unit for processing the query request are located in the same service node.

3. The method according to claim 1 or 2, wherein The target file is a big data report; the service node is one of multiple service nodes included in the big data query system.

4. The method according to any one of claims 1 to 3, wherein: Obtaining the target address corresponding to the query request through the metadata unit includes: Obtaining target metadata corresponding to the query request through the metadata unit; The target address corresponding to the target metadata is determined by the metadata unit based on a preset mapping table; wherein the preset mapping table stores a mapping relationship between metadata and file addresses.

5. The method according to claim 4, wherein The step of loading the target address into a preset execution engine includes: Loading the target address into a preset execution engine, and creating a target view table in the preset execution engine; The acquiring of the target file targeted by the query request from a preset file system by the execution unit and the preset execution engine includes: The target file targeted by the query request is obtained from the preset file system through the execution unit and by using the target view table in the preset execution engine.

6. The method according to any one of claims 1 to 4, further comprising: The target file is saved in the cache unit by the execution unit; The target file is sent through the cache unit.

7. The method according to claim 6, further comprising: The cache unit determines whether the target file targeted by the query request exists in the currently cached files.

8. The method according to claim 7, wherein: Also includes: When it is determined through the cache unit that the target file targeted by the query request does not exist in the currently cached files, the query request is sent to the metadata unit.

9. A file query device comprising: A request obtaining unit, configured to obtain a query request, wherein the query request is used to request to query a target file from a preset file system; A metadata unit, used to obtain a target address corresponding to the query request; A parsing unit, configured to load the target address into a preset execution engine; The execution unit is configured to obtain the target file targeted by the query request from a preset file system using the preset execution engine.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.