Data processing method and device
By separating the original metadata of the cold storage node into the hot storage node and updating it, the problem of slow access speed of cold storage nodes is solved, achieving more efficient data query and reducing storage costs.
Patent Information
- Application Number
- CN202410078265.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the data processing efficiency of the cold storage node is low, resulting in low access speed and query efficiency, and the inability to effectively reduce storage costs.
The original metadata in the cold storage node is separated, the metadata file is constructed and stored in the hot storage node, and the storage location information is updated based on the storage location information in the hot storage node, and the access information is recorded so that the hot storage node can access it instead of the cold storage node.
On the premise of ensuring data availability, access speed and data query efficiency are improved, and storage costs are reduced.
Smart Images

Figure CN120335708A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a data processing method and apparatus. Background Art
[0002] With the advent of the big data era, data storage and query also face huge challenges. In the scenario of massive big data, a large amount of historical data, such as order data or monitoring data, is often stored in a table. As time goes by, the frequency of accessing these data will gradually decrease and eventually be shelved. Reducing the storage cost of this part of the data has become a new problem. To solve this problem and reduce the storage cost at the same time, cold-hot separation storage has emerged. Cold-hot separation storage supports storing cold data and hot data on different storage media respectively.
[0003] In the prior art, cold data is usually stored in cold storage nodes and hot data is stored in hot storage nodes. When querying data, since the cold storage nodes have lower data processing efficiency compared with the hot storage nodes, it will also lead to a series of problems such as low access speed to the cold storage nodes and low data query efficiency. Therefore, there is an urgent need for a more effective data processing method to solve the above problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification simultaneously relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a data processing method is provided, including:
[0006] Determine the original metadata of the original data file in the cold storage node;
[0007] Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node;
[0008] Update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node;
[0009] Wherein, the original metadata updated and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
[0010] According to the second aspect of the embodiments of this specification, another data processing method is provided, including:
[0011] Determine the original metadata of the original data file in the cold storage node;
[0012] Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node;
[0013] Update the original metadata included in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node;
[0014] In response to receiving a data processing request, perform a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node.
[0015] According to the third aspect of the embodiments of the present specification, there is provided a data processing apparatus, including:
[0016] A determination module, configured to determine the original metadata of the original data file in the cold storage node;
[0017] A storage module, configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node;
[0018] An update module, configured to update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata after update and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
[0019] According to the fourth aspect of the embodiments of the present specification, there is provided another data processing apparatus, including:
[0020] A determination module, configured to determine the original metadata of the original data file in the cold storage node;
[0021] A storage module, configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node;
[0022] An update module, configured to update the original metadata included in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node;
[0023] An execution module, configured to perform a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node in response to receiving a data processing request.
[0024] According to the fifth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0025] Memory and processor;
[0026] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above data processing method are implemented.
[0027] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above data processing method are implemented.
[0028] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above data processing method are implemented.
[0029] In an embodiment of the present specification, the original metadata of the original data file in the cold storage node is determined; a metadata file corresponding to the original data file is constructed based on the original metadata and stored in the hot storage node; based on the storage location information of the metadata file in the hot storage node, the original metadata included in the metadata file is updated; wherein, the original metadata updated and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node in the form of a metadata file, so that when the cold storage node is accessed, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access. On the premise of ensuring data availability, the access speed and data query efficiency are improved while improving the access speed. Description of the Drawings
[0030] Figure 1 is a schematic diagram of the processing process of a data processing method provided by an embodiment of the present specification;
[0031] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present specification;
[0032] Figure 3 is a flowchart of the processing process of a data processing method provided by an embodiment of the present specification;
[0033] Figure 4 is a schematic diagram of data separation of a data processing method provided by an embodiment of the present specification;
[0034] Figure 5 is a schematic diagram of data update of a data processing method provided by an embodiment of the present specification;
[0035] Figure 6 It is a schematic structural diagram of a data processing device provided by an embodiment of this specification;
[0036] Figure 7 It is a flowchart of another data processing method provided by an embodiment of this specification;
[0037] Figure 8 It is a schematic structural diagram of another data processing device provided by an embodiment of this specification;
[0038] Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific implementation manners
[0039] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0040] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0041] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0042] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0043] First, the noun terms involved in one or more embodiments of this specification are explained.
[0044] SSD (Solid State Drive): It is a solid-state drive that uses flash memory chips to store data.
[0045] HDD (Hard Disk Drive): It is a traditional hard disk storage device that uses rotating disks and read / write heads to store and retrieve data.
[0046] Bloom Filter: It is used to determine whether an element exists in a set and has spatial and temporal efficiency.
[0047] Figure 1 It is a schematic diagram of the processing process of a data processing method provided by an embodiment of this specification; as Figure 1 shown, determine the original metadata of the original data file in the cold storage node, construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node. Based on the storage location information of the metadata file in the hot storage node, update the original metadata included in the metadata file. The original metadata updated and stored in the hot storage node records the access information of the original data file stored in the cold storage node. Thus, when a data read request is received, perform the data read task corresponding to the data read request based on the updated original metadata in the hot storage node.
[0048] Separate the original metadata of the original data file in the cold storage node from the cold storage node and store it in the hot storage node in the form of a metadata file, so that when the cold storage node is accessed, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access. On the premise of ensuring data availability, the access speed and data query efficiency are improved simultaneously.
[0049] In this specification, a data processing method is provided. This specification also relates to a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0050] See Figure 2 , Figure 2 which shows a flowchart of a data processing method provided according to an embodiment of this specification, specifically including the following steps.
[0051] Step 202: Determine the original metadata of the original data file in the cold storage node.
[0052] Specifically, the cold storage node is used to store cold data, that is, data with a low data access frequency; the storage type of the cold storage node is capacity-based storage; the cold storage node can be a storage medium in a database, and the original data file is the data file stored in the cold storage node; the original metadata is the meta-information in the original data file, and the original metadata is used to represent the data attribute information of the cold data stored in the cold storage node, which is a structured description of the data.
[0053] Based on this, determine the cold storage node corresponding to the database, and determine the original data file in the cold storage node. Furthermore, determine the original metadata in the original data file that is used to describe the data attribute information of the cold data, which is convenient for subsequent separation of the original metadata.
[0054] In practical applications, the cold storage node contains multiple original data files. When determining the original metadata, it is necessary to determine the original metadata corresponding to each original data file, so as to achieve a comprehensive separation of the metadata in the cold storage node.
[0055] Furthermore, considering that the data reading requirements are different when separating data for the original data file, the original metadata can be determined based on a data separation event. The specific implementation is as follows:
[0056] Based on the data separation event, determine the original data file in the cold storage node, where the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; determine the original metadata based on the original data file.
[0057] Specifically, the data separation event refers to an event for separating the original metadata in the original data file in the cold storage node; the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; the data separation event can be executed when the original data file in the cold storage node is updated, and then separate the original metadata in the updated original data file; the data separation event can also be executed based on the data separation policy corresponding to the database, which can be to execute the data separation event at fixed intervals, or to execute the data separation event for the cold storage node when there is a data separation requirement.
[0058] Based on this, determine the data separation event for the original data files in the cold storage node. Based on the data separation event, determine at least one original data file in the cold storage node, and then determine the original metadata of each original data file in the at least one original data file. Implement the data separation of the original metadata of multiple original data files in the cold storage node.
[0059] For example, in a big data scenario, cold data is stored in cold storage nodes and hot data is stored in hot storage nodes. When performing data reading, it is necessary to judge the data readability of the cold data in the cold storage node based on the original metadata in the original data files in the cold storage node. If the original metadata in the original data files in the cold storage node is separated and stored in the hot storage node, the hot storage node can replace the cold storage node to complete the data readability judgment. Determine the data separation event for the cold storage node, determine the original data file in the cold storage node based on the data separation event, and then read the original metadata in the original data file.
[0060] To sum up, read the original metadata in the original data files in the cold storage node based on the data separation event, so as to realize the separation of the original metadata in the cold storage node on demand.
[0061] Furthermore, considering that different data reading methods can be realized when performing data reading, such as reading by row or reading by range, different original metadata needs to be determined when determining the original metadata. The specific implementation is as follows:
[0062] In the case of the row data reading operation corresponding to the data separation event, determine the filtering metadata corresponding to the row data reading operation in the original data file, and use the filtering metadata as the original metadata; in the case of the range data reading operation corresponding to the data separation event, determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata; in the case of the row data reading operation and the range data reading operation corresponding to the data separation event, determine the filtering metadata corresponding to the row data reading operation in the original data file, and determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
[0063] Specifically, the row data reading operation refers to the operation of reading a single row of data in the storage node; the filtering metadata refers to the metadata associated with the Bloom filter, including but not limited to metadata such as Bloom blocks; correspondingly, the range reading operation refers to the operation of reading a range of data in the storage node, that is, after positioning the starting position corresponding to the data to be read, reading the data within the range; the index metadata contains data block index information, which is used to index the position of the data block storing the data in the original data file.
[0064] Based on this, the data separation event corresponds to the row data reading operation and / or the range data reading operation. In the case where the data separation event corresponds to the row data reading operation, it means that the single-row reading method is adopted when reading data from the cold storage node. Determine the filtering metadata associated with the Bloom filter corresponding to the row data reading operation in the original data file, and use the filtering metadata as the original metadata. In the case where the data separation event corresponds to the range data reading operation, it means that the range reading method is adopted when reading data from the cold storage node. Determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata. Correspondingly, in the case where the data separation event corresponds to both the row data reading operation and the range data reading operation, it means that both the single-row reading method and the range reading method are adopted when reading data from the cold storage node. Then determine the filtering metadata corresponding to the row data reading operation in the original data file, and determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
[0065] Continuing with the above example, for a single-row read request for a storage node, each file in the database saves a Bloom Filter to quickly determine whether the data to be read is in this file. Separating the Bloom filter from the data file of the cold storage node and saving it in the hot storage node with higher read and write speeds can accelerate the data reading judgment of the Bloom filter and speed up the reading of cold data and hot data. For a range read request, an index block (Index Block: used to index the position of the data block storing the data in the data file) is required to locate the starting position of the range read in the data file of the cold storage node. Therefore, separating the index block from the data file of the cold storage node and saving it in the hot storage node with higher read and write speeds can improve the data reading efficiency. The separation of the Bloom filter and the separation of the index block can be determined according to the actual data reading requirements.
[0066] In summary, the data separation event corresponds to the row data reading operation and / or the range data reading operation, and thus the original metadata can be flexibly determined to meet different data reading requirements.
[0067] Step 204: Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node.
[0068] Specifically, after determining the original metadata of the original data file in the cold storage node as described above, a metadata file corresponding to the original data file can be constructed based on the original metadata, and the metadata file is stored in the hot storage node. Among them, the data read and write efficiency of the cold storage node is lower than that of the hot storage node. The hot storage node is used to store hot data, that is, data with a high access frequency; the storage type of the hot storage node is standard storage, performance storage, local SSD or local HDD. The metadata file is a data file constructed based on the original metadata determined in the original data file, and is used to replace the original metadata in the original data file to perform subsequent data detection tasks. The data detection task is used to detect whether the cold storage node contains cold data that needs to be read.
[0069] Based on this, after determining the original metadata of the original data file in the cold storage node as described above, the original metadata is combined to construct a metadata file corresponding to the original data file in the cold storage node, and the metadata file is stored in the hot storage node, so that the metadata file can replace the original metadata in the original data file to perform subsequent data detection tasks. Utilize the characteristic of high data processing efficiency in the hot storage node to improve the task execution efficiency of the data detection task for the cold storage node.
[0070] Furthermore, when the original metadata is stored in the cold storage node, it has a certain storage structure. Considering that the storage structure will affect the data read efficiency and order during data reading, in order to ensure the consistency of the data processing results after separating the original metadata to the hot storage node, it is necessary to construct a metadata file based on the structure information of the original metadata. The specific implementation is as follows:
[0071] Determine the structure information of the original metadata in the original data file; determine the distribution information of the sub-original metadata included in the original metadata in the original data file based on the structure information; combine the sub-original metadata included in the original metadata according to the distribution information of the sub-original metadata to obtain the metadata file corresponding to the original data file.
[0072] Specifically, the structure information refers to the organization form and relationship of the original metadata, including but not limited to information such as the data type, data relationship, and storage structure of the original metadata; the distribution information refers to the arrangement information of the sub-original metadata in the original data file, including the arrangement position information and the relationship between the sub-original metadata.
[0073] Based on this, the structure information such as the data type, data relationship, and storage structure of the original metadata in the original data file is determined. Based on the structure information, the distribution information of each sub-original metadata included in the original metadata in the original data file is determined. The sub-original metadata included in the original metadata is combined according to the distribution information of each sub-original metadata to obtain the metadata file corresponding to the original data file.
[0074] Continuing with the above example, as Figure 3 shown in (a) above, in the case of data separation for the data files in the cold storage node, each of the N data files from data file 1 to data file N in the cold storage node is subjected to data separation. Each data file corresponds to a metadata file and is stored in the hot storage node. The metadata file contains metadata such as Bloom blocks, index blocks, meta-information indexes, and Bloom indexes that were originally stored in the data files in the cold storage node. As Figure 3 shown in (b) above, the original data file stored in the cold storage node contains data of multiple modules such as scanned blocks, non-scanned blocks, and the load of the open part. The scanned blocks correspond to Bloom blocks, the non-scanned blocks correspond to meta-information blocks, etc., and the load of the open part corresponds to data such as root data indexes, middle key fields, meta-indexes, file information, and Bloom filter metadata. When extracting data from the original data file, the data corresponding to the two modules of the scanned blocks and the load of the open part is determined as the original metadata to be extracted, and a metadata file is constructed according to the data structure when the original metadata is stored in the cold storage node and the distribution of each sub-original metadata.
[0075] In summary, combining the sub-original metadata included in the original metadata according to the distribution information of each sub-original metadata to obtain the metadata file corresponding to the original data file ensures the consistency of the structure information of the original metadata before and after data separation.
[0076] Step 206: Update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata updated and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
[0077] Specifically, after constructing the metadata file corresponding to the original data file based on the original metadata and storing the metadata file in the hot storage node, the original metadata included in the metadata file can be updated based on the storage location information of the metadata file in the hot storage node. Among them, the original metadata that has been updated and stored in the hot storage node records the access information of the original data file stored in the cold storage node. The storage location information indicates the location information of the metadata file in the hot storage node after the metadata file is stored in the hot storage node. The access information of the original data file refers to the metadata relied on when performing data detection tasks based on the original data file. The metadata includes, but is not limited to, Bloom blocks and Bloom indexes.
[0078] Based on this, after constructing the metadata file corresponding to the original data file based on the original metadata and storing the metadata file in the hot storage node, determine the storage location information of the metadata file in the hot storage node, update the original metadata included in the metadata file based on the storage location information, and the original metadata that has been updated and stored in the hot storage node records the access information of the original data file stored in the cold storage node. The access information is used to perform data detection tasks.
[0079] Furthermore, since the storage location of the original metadata in the cold storage node is different from that in the hot storage node, and when the original metadata is stored in the cold storage node, data block search is performed through the internal index information of the original data file, and the index information is also related to the storage location of the original metadata in the cold storage node. Therefore, when the original metadata is stored in the hot storage node, it is also necessary to update the index information according to the storage location information of the storage location target sub-metadata in the hot storage node in the original metadata. The specific implementation is as follows:
[0080] Determine the target sub-metadata in the original metadata of the metadata file, and read the initial data index information recorded by the target sub-metadata; based on the storage location information of the target sub-metadata in the hot storage node, update the initial data index information to the target data index information.
[0081] Specifically, the target sub-metadata refers to the metadata associated with the initial data index information corresponding to the original data file; the initial data index information is used to determine the data block containing the metadata before determining whether the cold storage node contains the data to be read; the storage location information is the storage address of the target sub-metadata in the hot storage node, and the target index information is the index information obtained after updating the initial data index information based on the storage address of the target sub-metadata in the hot storage node.
[0082] Based on this, determine the target sub - metadata in the original metadata of the metadata file, and read the initial data index information of the target sub - metadata record. In the case of row - data reading, the initial data index information associates the block offset and the Bloom block. In the case of range - data reading, the initial data index information associates the block offset and the index block. Based on the storage location information of the target sub - metadata in the hot - storage node, update the initial data index information to the target data index information.
[0083] Continuing with the above example, as Figure 4 shown in (a) below, in the case of row - data reading, separate the Bloom filter of the original metadata in the cold - storage node. At this time, it is necessary to update the initial data index information corresponding to the Bloom index block. The Bloom index block contains data such as version and Bloom index. In the cold - storage node, the Bloom index contains data such as block key and block offset. The block offset corresponds to the cold - storage index, that is, the block offset and the Bloom block in the cold - storage node constitute the initial data index information. After storing the original metadata in the form of a metadata file to the hot - storage node, it is necessary to modify the pointer of the block offset to the hot - storage index, pointing to the Bloom block stored in the hot - storage node, so as to achieve the purpose of index modification and update the initial data index information to the target data index information. As Figure 4 shown in (b) below, in the case of range - data reading, it is necessary to modify the pointer of the Bloom index in the root index block. The root index block contains data such as index items and Bloom index. Before separating the original metadata in the cold - storage node, the block offset corresponds to the index block, constituting the cold - storage index. After storing the original metadata to the hot - storage node, update the cold - storage index to the hot - storage index and modify the pointer of the block offset to point to the index block in the hot - storage node.
[0084] To sum up, based on the storage location information of the target sub - metadata in the hot - storage node, update the initial data index information to the target data index information, so as to ensure that when performing subsequent data - reading tasks based on the original metadata in the hot - storage node, data detection and data judgment can be correctly carried out, improving the accuracy of subsequent data reading.
[0085] Furthermore, considering that in an actual data - reading scenario, the data to be read may be stored in both the cold - storage node and the hot - storage node. Therefore, it is necessary to perform data - reading tasks based on both the cold - storage node and the hot - storage node simultaneously. The specific implementation is as follows:
[0086] Receive a data reading request; read the hot storage metadata stored in the hot storage node according to the data reading request, and the target metadata in the metadata file, where the target metadata is the metadata after updating the original metadata in the metadata file. Based on the hot storage metadata and the original metadata, perform a data reading task that associates the cold storage node and the hot storage node.
[0087] Specifically, the data reading request is used to read data in the cold storage node and / or the hot storage node; the hot storage metadata is used to determine whether the hot storage node contains the hot data to be read; the target metadata is used to determine whether the cold storage node contains the cold data to be read; and the data reading task is the task corresponding to the data reading request.
[0088] Based on this, receive a data reading request, read the hot storage metadata stored in the hot storage node according to the data reading request, and the target metadata in the metadata file stored in the hot storage node, where the target metadata is the metadata after updating the original metadata in the metadata file stored in the hot storage node. Based on the hot storage metadata and the original metadata, perform a data reading task that associates the cold storage node and the hot storage node, and read the data to be read corresponding to the data reading request in the cold storage node and / or the hot storage node.
[0089] In practical applications, both the hot storage node and the cold data storage node are storage nodes in the database. The data reading task includes a data detection task and a data determination task. During the execution of the data detection task in the data reading task, it includes the data storage detection of the cold storage node and the data storage detection of the hot storage node. During the execution of the data determination task in the data reading task, based on the hot detection result corresponding to the hot storage node, determine the cold data to be read in the cold storage node, and based on the cold detection result corresponding to the cold storage node, determine the hot data to be read in the hot storage node.
[0090] Following the above example, when receiving a data reading request for the database, determine the cold storage node and the hot storage node included in the database. Determine the target metadata for data storage detection of the cold storage node in the hot storage node, and determine the hot storage metadata for data storage detection of the hot storage node in the hot storage node. Perform a data reading task that associates the cold storage node and the hot storage node.
[0091] In summary, perform the data reading task simultaneously based on the cold storage node and the hot storage node, thereby improving the comprehensiveness of data reading.
[0092] Furthermore, considering that when performing data reading tasks on both cold storage nodes and hot storage nodes, a cold storage node may or may not contain the data to be read, and correspondingly, a hot storage node may or may not contain the data to be read. Therefore, it is necessary to perform data storage detection on both cold storage nodes and hot storage nodes simultaneously. The specific implementation is as follows:
[0093] Perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata; execute a data determination task according to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node, as the execution of the data reading task.
[0094] Specifically, data storage detection refers to the operation of determining whether the hot storage node contains the data to be read corresponding to the data reading request for the hot storage node, and the operation of determining whether the cold storage node contains the data to be read corresponding to the data reading request for the cold storage node; correspondingly, after performing data storage detection on the hot storage node, a hot detection result corresponding to the hot storage node can be obtained; after performing data storage detection on the cold storage node, a cold detection result corresponding to the cold storage node can be obtained. The hot detection result indicates whether the hot storage node contains the hot data to be read, and the cold detection result indicates whether the cold storage node contains the cold data to be read.
[0095] Based on this, perform data storage detection on the hot storage node based on the hot storage metadata to determine whether the hot storage node contains the data to be read corresponding to the data reading request, and perform data storage detection on the cold storage node based on the original metadata to determine whether the cold storage node contains the data to be read corresponding to the data reading request. Execute a data determination task for associating the hot storage node and the cold storage node according to the hot detection result obtained from the data storage detection of the hot storage node, or execute a data determination task for associating the hot storage node and the cold storage node according to the cold detection result corresponding to the cold storage node obtained from the data storage detection of the cold storage node. The above execution process is the execution process of the data reading task.
[0096] Following the above example, determine the target metadata for performing data storage detection on the cold storage node in the hot storage node, and determine the hot storage metadata for performing data storage detection on the hot storage node in the hot storage node. Perform data storage detection on whether the cold storage node contains the data to be read based on the target metadata to obtain a cold detection result; and perform data storage detection on whether the hot storage node contains the data to be read based on the hot storage metadata to obtain a hot detection result. Then, perform subsequent data determination tasks for associating the hot storage node and the cold storage node based on the obtained hot detection result and cold detection result.
[0097] In summary, data storage detection is performed on both cold storage nodes and hot storage nodes simultaneously to determine whether the cold storage nodes and hot storage nodes contain the data to be read, thereby improving the accuracy of subsequent data reading. Data storage detection is performed based on the target metadata stored in the hot storage nodes, without the need to perform data storage detection based on the original metadata in the cold storage nodes, thereby improving the efficiency of data storage detection.
[0098] Further, when performing data storage detection on cold storage nodes and hot storage nodes simultaneously, the detection results corresponding to the cold storage nodes and the hot storage nodes each correspond to two data storage situations. The data stored in the cold storage node may or may not need to be read. Correspondingly, the data stored in the hot storage node may or may not need to be read. Therefore, the data to be read can be determined according to the hot detection result and the cold detection result. The specific implementation is as follows:
[0099] In the case where it is determined according to the hot detection result corresponding to the hot storage node that the hot storage node contains the hot data to be read corresponding to the data reading request, the hot data to be read corresponding to the data reading request is read in the hot storage node as the execution of the data determination task; in the case where it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains the cold data to be read corresponding to the data reading request, the cold data to be read corresponding to the data reading request is read in the cold storage node as the execution of the data determination task.
[0100] Specifically, the hot data to be read is the hot data that needs to be read corresponding to the data reading request and is stored in the hot storage node; correspondingly, the cold data to be read is the cold data that needs to be read corresponding to the data reading request and is stored in the cold storage node.
[0101] Based on this, when it is determined according to the heat detection result corresponding to the heat storage node that the heat storage node contains the to-be-read heat data corresponding to the data reading request, it indicates that the heat data in the heat storage node needs to be read. Read the to-be-read heat data corresponding to the data reading request in the heat storage node, and take the determination of the to-be-read heat data as the execution of the data determination task. When it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains the to-be-read cold data corresponding to the data reading request, it indicates that the cold data in the cold storage node needs to be read. Read the to-be-read cold data corresponding to the data reading request in the cold storage node, and take the determination of the to-be-read cold data as the execution of the data determination task. Correspondingly, when it is determined according to the heat detection result corresponding to the heat storage node that the heat storage node does not contain the to-be-read heat data corresponding to the data reading request, it indicates that the heat data in the heat storage node does not need to be read; when it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node does not contain the to-be-read cold data corresponding to the data reading request, it indicates that the cold data in the cold storage node also does not need to be read.
[0102] Continuing with the above example, after determining the heat detection result, it is possible to determine based on the heat detection result whether the heat storage node contains the to-be-read heat data corresponding to the data reading request. When it is determined based on the heat detection result that the heat storage node contains the to-be-read heat data corresponding to the data reading request, the to-be-read heat data can be read out as the response to the data reading request. Correspondingly, after determining the cold detection result, it is possible to determine based on the cold detection result whether the cold storage node contains the to-be-read cold data corresponding to the data reading request. When it is determined based on the cold detection result that the cold storage node contains the to-be-read cold data corresponding to the data reading request, the to-be-read cold data can be read out as the response to the data reading request.
[0103] In summary, determine whether the heat storage node contains the to-be-read heat data according to the heat detection result, and determine whether the cold storage node contains the to-be-read cold data according to the cold detection result, so as to achieve accurate data reading for the heat storage node and the cold storage node.
[0104] Furthermore, considering that the metadata file in the heat storage node may be inaccessible, in this case, when performing the data processing task, it is impossible to determine the data readability through the metadata file in the heat storage node. Therefore, the original metadata still retained in the cold storage node can be used to determine the data readability of the cold storage node. The specific implementation is as follows:
[0105] In the case where the metadata file in the heat storage node is inaccessible, perform the data processing task based on the original metadata of the original data file in the cold storage node.
[0106] Based on this, in the case where the metadata file in the hot storage node is not updated and the metadata file cannot be accessed, determine the original metadata of the original data file in the cold storage node, perform a data processing task based on the original metadata in the cold storage node, and determine whether the cold data in the cold storage node needs to be read.
[0107] In summary, the original metadata stored in the cold storage node is used to assist in performing a data processing task in the case where the metadata file in the hot storage node cannot be accessed, complete the detection of the data readability in the cold storage node, so as to achieve the disaster recovery effect.
[0108] An embodiment of this specification determines the original metadata of the original data file in the cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata, and stores the metadata file in the hot storage node; updates the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata after update and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node. Separate the original metadata of the original data file in the cold storage node from the cold storage node and store it in the hot storage node in the form of a metadata file, so that when the cold storage node is accessed, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access, while ensuring data availability, improving the access speed and at the same time improving the data query efficiency.
[0109] The following combines the attached Figure 5 , taking the application of the data processing method provided in this specification in data reading as an example, further illustrates the data processing method. Among them, Figure 5 shows a process flow chart of a data processing method provided by an embodiment of this specification, specifically including the following steps.
[0110] Step 502: Determine the original data file in the cold storage node based on the data separation event.
[0111] In the scenario of massive big data, a large amount of historical data, such as order data or monitoring data, is often stored in a single table. As time goes by, the frequency of accessing these data will gradually decrease and eventually be shelved. The data with a lower access frequency is stored as cold data in cold storage nodes, and the data with a higher access frequency is stored as hot data in hot storage nodes, thereby reducing the data storage cost. The storage type of cold storage is capacity-based storage, and the storage types of hot storage are standard storage, performance storage, local SSD disks or local HDD disks. The data read and write speed of cold storage nodes is much lower than that of hot storage nodes. Even when only querying hot data, at least the metadata of the cold file storing the cold data will be read to determine whether to read the cold file. Therefore, the metadata of the cold file can be extracted from the original file structure, and the file structure can be reorganized to form a metadata file, which is stored in the hot storage node. The original data file in the cold storage node is the metadata of the cold file.
[0112] The data separation event is a data separation task determined according to a pre-established data separation rule or data separation strategy. When the data separation task is executed, the process of extracting the metadata of the cold file from the original file structure, reorganizing the file structure to form a metadata file, and storing it in the hot storage node is completed.
[0113] Step 504: In the case where the data separation event corresponds to a single-row data reading operation, determine the filtering metadata corresponding to the single-row data reading operation in the original data file, and use the filtering metadata as the original metadata.
[0114] Considering that there are two data reading methods, single-row reading and range reading, during data reading, the filtering metadata corresponding to the reading Bloom filter and / or the index metadata corresponding to the index block can be determined according to requirements.
[0115] Step 506: In the case where the data separation event corresponds to a range data reading operation, determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata.
[0116] Step 508: In the case where the data separation event corresponds to both a single-row data reading operation and a range data reading operation, determine the filtering metadata corresponding to the single-row data reading operation in the original data file, and determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
[0117] Step 510: Determine the structure information of the original metadata in the original data file, and based on the structure information, determine the distribution information of the metadata included in the original metadata in the original data file.
[0118] Step 512: Combine the metadata included in the original metadata according to the distribution information of the metadata included in the original metadata to obtain a metadata file corresponding to the original data file, and store the metadata file in the hot storage node.
[0119] When performing data separation, it is necessary to ensure the consistency of the original metadata in the cold storage node and in the hot storage node. Therefore, the distribution information of the metadata included in the original metadata in the original data file can be determined according to the structure information of the original metadata in the original data file, and then a metadata file can be constructed according to the distribution information of each metadata and stored in the hot storage node. The file structure of the original data file in the cold storage node is not damaged, achieving the effect of storing the original metadata in both the cold storage node and the hot storage node simultaneously.
[0120] Step 514: Determine the target metadata in the original metadata of the metadata file, and the initial data index information used to read the target metadata.
[0121] Step 516: Update the initial data index information to the target data index information based on the storage location information of the target metadata in the hot storage node.
[0122] In practical applications, since the storage location of the original metadata in the cold storage node is different from that in the hot storage node, it is necessary to modify the index information inside the metadata file to ensure that the data block can be correctly found, that is, modify the Bloom block pointed to by the block offset in the Bloom index block, and / or modify the index block pointed to by the block offset in the root index block.
[0123] Step 518: In response to a data read request, read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file.
[0124] When a data read request is received, it is possible to determine whether the data to be read is included in the hot storage node based on the hot storage metadata originally stored in the hot storage node, and determine whether the data to be read is included in the cold storage node based on the target metadata corresponding to the original metadata copied from the cold storage node in the hot storage node.
[0125] Step 520: Based on the hot storage metadata and the original metadata, execute a data read task that associates the cold storage node and the hot storage node.
[0126] In practical applications, data storage detection is performed on the hot storage nodes based on the hot storage metadata to determine whether the hot storage nodes contain the data to be read, and data storage detection is performed on the cold storage nodes based on the original metadata to determine whether the cold storage nodes contain the data to be read. In the case where it is determined based on the hot detection result corresponding to the hot storage node that the hot storage node contains the to-be-read hot data corresponding to the data read request, the to-be-read hot data corresponding to the data read request is read from the hot storage node; in the case where it is determined based on the cold detection result corresponding to the cold storage node that the cold storage node contains the to-be-read cold data corresponding to the data read request, the to-be-read cold data corresponding to the data read request is read from the cold storage node, and the data read task is completed.
[0127] In summary, the original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node in the form of a metadata file. Thus, when the cold storage node is accessed, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access. On the premise of ensuring data availability, the access speed is increased while the data query efficiency is improved. It can not only improve the cold data query efficiency, but also speed up the reading of hot data. It can not only improve the execution speed of a single-line read request, but also reduce the execution time of a range read request. In addition, according to the usage scenario, the Bloom filter and / or index block in the cold storage node can be freely selected to be separated, improving the flexibility of data separation.
[0128] In addition, since the Bloom filter has a false positive rate, that is, for the data judged to exist by the Bloom filter, there is a certain probability that it does not actually exist. The data separation can separately adjust the false positive rate of the Bloom filter of the metadata file stored in the hot storage node, which can avoid affecting the structure and space size of the original data file.
[0129] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing device. Figure 6 The structural schematic diagram of a data processing device provided by an embodiment of this specification is shown. As Figure 6 shown, the device includes:
[0130] A determination module 602, configured to determine the original metadata of the original data file in the cold storage node;
[0131] A storage module 604, configured to construct a metadata file corresponding to the original data file based on the original metadata and store the metadata file in the hot storage node;
[0132] An update module 606, configured to update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata stored in the hot storage node after the update records access information for accessing the original data file stored in the cold storage node.
[0133] In an optional embodiment, the determination module 602 is further configured to:
[0134] Determine an original data file in the cold storage node based on a data separation event, where the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node;
[0135] Determine the original metadata based on the original data file.
[0136] In an optional embodiment, the determination module 602 is further configured to:
[0137] In the case where the data separation event corresponds to a row data reading operation, determine filtering metadata corresponding to the row data reading operation in the original data file, and use the filtering metadata as the original metadata;
[0138] In the case where the data separation event corresponds to a range data reading operation, determine index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata;
[0139] In the case where the data separation event corresponds to a row data reading operation and a range data reading operation, determine filtering metadata corresponding to the row data reading operation in the original data file, and determine index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
[0140] In an optional embodiment, the update module 606 is further configured to:
[0141] Determine target sub-metadata in the original metadata of the metadata file, and read initial data index information recorded by the target sub-metadata;
[0142] Update the initial data index information to target data index information based on the storage location information of the target sub-metadata in the hot storage node.
[0143] In an optional embodiment, the storage module 604 is further configured to:
[0144] Determine the structure information of the original metadata in the original data file;
[0145] Based on the structure information, determine the distribution information of the sub-original metadata included in the original metadata in the original data file;
[0146] Combine the sub-original metadata included in the original metadata according to the distribution information of the sub-original metadata to obtain the metadata file corresponding to the original data file.
[0147] In an optional embodiment, the update module 606 is further configured to:
[0148] Receive a data reading request;
[0149] According to the data reading request, read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file, where the target metadata is the updated metadata of the original metadata in the metadata file.
[0150] Based on the hot storage metadata and the original metadata, perform a data reading task for associating the cold storage node and the hot storage node.
[0151] In an optional embodiment, the update module 606 is further configured to:
[0152] Perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata;
[0153] According to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node, perform a data determination task as the execution of the data reading task.
[0154] In an optional embodiment, the update module 606 is further configured to:
[0155] In the case where it is determined according to the hot detection result corresponding to the hot storage node that the hot storage node contains the to-be-read hot data corresponding to the data reading request, read the to-be-read hot data corresponding to the data reading request in the hot storage node as the execution of the data determination task;
[0156] In the case where it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains the to-be-read cold data corresponding to the data reading request, read the to-be-read cold data corresponding to the data reading request in the cold storage node as the execution of the data determination task.
[0157] In an optional embodiment, the update module 606 is further configured to:
[0158] In the case where the metadata file in the hot storage node is inaccessible, perform a data processing task based on the original metadata of the original data file in the cold storage node.
[0159] In summary, an embodiment of this specification determines the original metadata of the original data file in the cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata and stores the metadata file in the hot storage node; updates the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the updated original metadata stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node in the form of a metadata file, so that when the cold storage node is accessed, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access, improving the data query efficiency while increasing the access speed on the premise of ensuring data availability.
[0160] The above is a schematic solution of a data processing device according to an embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0161] See Figure 7 , Figure 7 shows a flowchart of another data processing method according to an embodiment of this specification, specifically including the following steps.
[0162] Step 702: Determine the original metadata of the original data file in the cold storage node;
[0163] Step 704: Construct a metadata file corresponding to the original data file based on the original metadata and store the metadata file in the hot storage node;
[0164] Step 706: Update the original metadata included in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node;
[0165] Step 708: In the case of receiving a data processing request, perform a data processing task related to the cold storage node based on the target metadata stored in the hot storage node.
[0166] In practical applications, in a massive big data scenario, data with a lower access frequency will be stored in cold storage nodes, while data with a higher access frequency will be stored in hot storage nodes. When reading data, a Bloom filter can be used to determine whether the data to be read exists in the cold storage nodes. Since the data read and write speed of cold storage nodes is relatively low, the efficiency of reading the original metadata of the original data files in cold storage nodes is also low. To solve this problem, by determining the original metadata of the original data files in cold storage nodes, constructing metadata files corresponding to the original data files based on the original metadata, and storing the metadata files in hot storage nodes, when a data processing request is received, the data processing task associated with the cold storage nodes is executed based on the target metadata stored in the hot storage nodes. In practical applications, since the storage location information of the metadata files in the hot storage nodes is different from the storage location information of the original data files in the cold storage nodes, the original metadata included in the metadata files can be updated to target metadata based on the storage location information of the metadata files in the hot storage nodes, and then the subsequent data processing tasks can be completed through the target metadata stored in the hot storage nodes.
[0167] In summary, an embodiment of this specification determines the original metadata of the original data files in cold storage nodes; constructs metadata files corresponding to the original data files based on the original metadata, and stores the metadata files in hot storage nodes; updates the original metadata included in the metadata files based on the storage location information of the metadata files in the hot storage nodes; wherein, the original metadata updated and stored in the hot storage nodes records the access information for accessing the original data files stored in the cold storage nodes. The original metadata of the original data files in the cold storage nodes is separated from the cold storage nodes and stored in the hot storage nodes in the form of metadata files. Thus, when the cold storage nodes are accessed during the execution of data processing tasks, the metadata files in the hot storage nodes can be used to replace the original data files in the cold storage nodes for access, improving the data query efficiency while enhancing the access speed on the premise of ensuring data availability.
[0168] Corresponding to the above method embodiment, this specification also provides an embodiment of a data processing device. Figure 8 The structural schematic diagram of another data processing device provided by an embodiment of this specification is shown. As Figure 8 shown, the device includes:
[0169] A determination module 802, configured to determine the original metadata of the original data files in cold storage nodes;
[0170] A storage module 804, configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in a hot storage node;
[0171] An update module 806, configured to update the original metadata included in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node;
[0172] An execution module 808, configured to execute a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node when a data processing request is received.
[0173] In an optional embodiment, the determination module 802 is further configured to:
[0174] Determine an original data file in the cold storage node based on a data separation event, where the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node;
[0175] Determine the original metadata based on the original data file.
[0176] In an optional embodiment, the determination module 802 is further configured to:
[0177] When the data separation event corresponds to a row data reading operation, determine filtering metadata corresponding to the row data reading operation in the original data file, and use the filtering metadata as the original metadata;
[0178] When the data separation event corresponds to a range data reading operation, determine index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata;
[0179] When the data separation event corresponds to both a row data reading operation and a range data reading operation, determine filtering metadata corresponding to the row data reading operation in the original data file, and determine index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
[0180] In an optional embodiment, the update module 806 is further configured to:
[0181] Determine target sub - metadata in the original metadata of the metadata file, and read initial data index information recorded by the target sub - metadata;
[0182] Update the initial data index information to target data index information based on the storage location information of the target sub - metadata in the hot storage node.
[0183] In an optional embodiment, the storage module 804 is further configured to:
[0184] Determine the structure information of the original metadata in the original data file;
[0185] Based on the structure information, determine the distribution information of the sub-original metadata included in the original metadata in the original data file;
[0186] Combine the sub-original metadata included in the original metadata according to the distribution information of the sub-original metadata to obtain the metadata file corresponding to the original data file.
[0187] In an optional embodiment, the update module 806 is further configured to:
[0188] Receive a data reading request;
[0189] According to the data reading request, read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file, where the target metadata is the metadata after the update of the original metadata in the metadata file.
[0190] Based on the hot storage metadata and the original metadata, perform a data reading task of associating the cold storage node and the hot storage node.
[0191] In an optional embodiment, the update module 806 is further configured to:
[0192] Perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata;
[0193] According to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node, perform a data determination task as the execution of the data reading task.
[0194] In an optional embodiment, the update module 806 is further configured to:
[0195] In the case where it is determined according to the hot detection result corresponding to the hot storage node that the hot storage node includes the to-be-read hot data corresponding to the data reading request, read the to-be-read hot data corresponding to the data reading request in the hot storage node as the execution of the data determination task;
[0196] In the case that it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains cold data to be read corresponding to the data reading request, read the cold data to be read corresponding to the data reading request in the cold storage node as the execution of the data determination task.
[0197] In an optional embodiment, the update module 806 is further configured to:
[0198] In the case that the metadata file in the hot storage node cannot be accessed, perform a data processing task based on the original metadata of the original data file in the cold storage node.
[0199] In summary, an embodiment of this specification determines the original metadata of the original data file in the cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata and stores the metadata file in the hot storage node; updates the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata after update and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node in the form of a metadata file, so that when the cold storage node is accessed during the execution of the data processing task, the metadata file in the hot storage node can be used to replace the original data file in the cold storage node for access, which improves the data query efficiency while improving the access speed on the premise of ensuring data availability.
[0200] The above is a schematic solution of a data processing device according to an embodiment of this specification. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0201] Figure 9 The structural block diagram of a computing device 900 provided according to an embodiment of this specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0202] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (e.g., network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (WiMAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0203] In one embodiment of the present specification, the above components of the computing device 900, as well as Figure 9 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.
[0204] The computing device 900 may be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smart phones), wearable computing devices (e.g., smart watches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 may also be a mobile or stationary server.
[0205] Among them, the processor 920 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above method are implemented.
[0206] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above method.
[0207] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above method are implemented.
[0208] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above method.
[0209] An embodiment of this specification also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0210] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above method.
[0211] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0212] The computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0213] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0214] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0215] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of this specification, many modifications and changes can be made. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Determining the original metadata of the original data file in the cold storage node; Constructing a metadata file corresponding to the original data file based on the original metadata, and storing the metadata file in the hot storage node; Updating the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; Wherein, the original metadata updated and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
2. The data processing method according to claim 1, wherein the determining the original metadata of the original data file in the cold storage node comprises: Determining the original data file in the cold storage node based on a data separation event, wherein the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; Determining the original metadata based on the original data file.
3. The data processing method according to claim 2, wherein the determining the original metadata based on the original data file comprises: In the case of a row data reading operation corresponding to the data separation event, determining filtering metadata corresponding to the row data reading operation in the original data file, and using the filtering metadata as the original metadata; In the case of a range data reading operation corresponding to the data separation event, determining index metadata associated with the original data file and corresponding to the range data reading operation, and using the index metadata as the original metadata; In the case of a row data reading operation and a range data reading operation corresponding to the data separation event, determining filtering metadata corresponding to the row data reading operation in the original data file, and determining index metadata associated with the original data file and corresponding to the range data reading operation, and using the filtering metadata and the index metadata as the original metadata.
4. The data processing method according to claim 1, wherein the updating the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node comprises: Determining target sub-metadata in the original metadata of the metadata file, and reading the initial data index information recorded by the target sub-metadata; Updating the initial data index information to target data index information based on the storage location information of the target sub-metadata in the hot storage node.
5. The data processing method according to claim 1, wherein the constructing a metadata file corresponding to the original data file based on the original metadata comprises: Determining the structure information of the original metadata in the original data file; Determining the distribution information of the sub-original metadata included in the original metadata in the original data file based on the structure information; Combining the sub-original metadata included in the original metadata according to the distribution information of the sub-original metadata to obtain the metadata file corresponding to the original data file.
6. The data processing method according to claim 1, the method further comprises: Receive a data reading request; Read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file according to the data reading request, where the target metadata is the updated metadata of the original metadata in the metadata file; Based on the hot storage metadata and the original metadata, perform a data reading task that associates the cold storage node and the hot storage node.
7. The data processing method according to claim 1, wherein the performing a data reading task that associates the cold storage node and the hot storage node based on the hot storage metadata and the original metadata includes: Perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata; According to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node, perform a data determination task as the execution of the data reading task.
8. The data processing method according to claim 7, wherein the performing a data determination task as the execution of the data reading task according to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node includes: In the case where it is determined according to the hot detection result corresponding to the hot storage node that the hot storage node contains the to-be-read hot data corresponding to the data reading request, read the to-be-read hot data corresponding to the data reading request in the hot storage node as the execution of the data determination task; In the case where it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains the to-be-read cold data corresponding to the data reading request, read the to-be-read cold data corresponding to the data reading request in the cold storage node as the execution of the data determination task.
9. The data processing method according to claim 1, the method further includes: In the case where the metadata file in the hot storage node cannot be accessed, perform a data processing task based on the original metadata of the original data file in the cold storage node.
10. A data processing method, including: Determine the original metadata of the original data file in the cold storage node; Based on the original metadata, construct a metadata file corresponding to the original data file, and store the metadata file in the hot storage node; Based on the storage location information of the metadata file in the hot storage node, update the original metadata included in the metadata file to target metadata; In the case of receiving a data processing request, perform a data processing task that associates the cold storage node based on the target metadata stored in the hot storage node.
11. A data processing device, including: A determination module configured to determine the original metadata of the original data file in the cold storage node; A storage module configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node; An update module, configured to update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; wherein, the original metadata after the update and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
12. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.