Data processing method and apparatus
By constructing the original metadata of the cold storage node into a metadata file and storing it to a hot storage node, the problem of cold storage node access speed and query efficiency is solved, and more efficient data access and query is achieved.
Patent Information
- Application Number
- PCT/IB2025/050175
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-24
AI Technical Summary
In the prior art, the data access speed and query efficiency of the cold storage node are low, resulting in an increase in the cost of cold data storage.
The original metadata in the cold storage node is separated, constructed into a metadata file and stored in a hot storage node, updated based on the storage location information in the hot storage node, and access information is recorded so that when a data processing request is received, data processing tasks are performed through the hot storage node.
It improves the access speed and data query efficiency of cold storage nodes, while reducing storage costs.
Smart Images

Figure IB2025050175_24072025_PF_FP_ABST
Abstract
Description
[0001] Data Processing Method and Apparatus This disclosure claims priority to Chinese patent application number 202410078265.7, filed with the China Patent Office on January 18, 2024, entitled "Data Processing Method and Apparatus," the entire contents of which are incorporated herein by reference. Technical Field: Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a data processing method and apparatus. Background: With the advent of the big data era, data storage and querying face significant challenges. In massive big data scenarios, a table often stores a large amount of historical data, such as order data or monitoring data. Over time, the frequency of access to this data decreases, and the data is eventually shelved. Reducing the storage cost of this data has become a new issue. To address this issue and reduce storage costs, hot-cold storage separation has emerged. Hot-cold storage separation supports storing cold and hot data on separate storage media. In the prior art, cold data is typically stored in cold storage nodes, and hot data is stored in hot storage nodes. During data queries, cold storage nodes have lower data processing efficiency than hot storage nodes, leading to a series of issues such as slow access speeds and low data query efficiency for cold storage nodes. Therefore, a more effective data processing method is urgently needed to address these issues. SUMMARY OF THE INVENTION In view of this, embodiments of the present disclosure provide a data processing method. One or more embodiments of the present disclosure relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art. According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: determining original metadata of an original data file in a cold storage node; constructing a metadata file corresponding to the original data file based on the original metadata, and storing the metadata file in a hot storage node; and updating the original metadata contained in the metadata file based on storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored in the cold storage node.According to a second aspect of an embodiment of the present disclosure, another data processing method is provided, comprising: determining original metadata of an original data file in a cold storage node; constructing a metadata file corresponding to the original data file based on the original metadata, and storing the metadata file in a hot storage node; updating the original metadata contained in the metadata file to target metadata based on storage location information of the metadata file in the hot storage node; and upon receiving a data processing request, executing a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node. According to a third aspect of an embodiment of the present disclosure, a data processing device is provided, comprising: a determination module configured to determine original metadata of an original data file in a cold storage node; a storage module configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in a hot storage node; and an update module configured to update the original metadata contained in the metadata file based on storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored in the cold storage node. According to a fourth aspect of an embodiment of the present disclosure, another data processing apparatus is provided, comprising: a determination module configured to determine original metadata of an original data file in a cold storage node; a storage module configured to construct a metadata file corresponding to the original data file based on the original metadata and store the metadata file in a hot storage node; an update module configured to update the original metadata contained in the metadata file to target metadata based on storage location information of the metadata file in the hot storage node; and an execution module configured to, upon receiving a data processing request, execute a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node. According to a fifth aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer-executable instructions; the processor is configured to execute the computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the steps of the aforementioned data processing method. According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-executable instructions are stored; wherein the computer-executable instructions, when executed by the processor, implement the steps of the aforementioned data processing method. According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.One embodiment of the present disclosure determines the original metadata of an original data file in a cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata, and stores the metadata file in a hot storage node; and updates the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for the original data file stored in the cold storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node as a metadata file. This allows the metadata file in the hot storage node to be accessed instead of the original data file in the cold storage node when the cold storage node is accessed. This improves access speed and data query efficiency while ensuring data availability. BRIEF DESCRIPTION OF THE DRAWINGS FIG1 is a schematic diagram of the processing process of a data processing method provided by one embodiment of the present disclosure; FIG2 is a flow chart of a data processing method provided by one embodiment of the present disclosure; FIG3 is a flow chart of the processing process of a data processing method provided by one embodiment of the present disclosure; FIG4 is a schematic diagram of data separation in a data processing method provided by one embodiment of the present disclosure; FIG5 is a schematic diagram of data updating in a data processing method provided by one embodiment of the present disclosure; FIG6 is a schematic diagram of the structure of a data processing device provided by one embodiment of the present disclosure; FIG7 is a flow chart of another data processing method provided by one embodiment of the present disclosure; FIG8 is a schematic diagram of the structure of another data processing device provided by one embodiment of the present disclosure; FIG9 is a block diagram of the structure of a computing device provided by one embodiment of the present disclosure. DETAILED DESCRIPTION The following description sets forth numerous specific details to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below. The terminology used in one or more embodiments of the present disclosure is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. As used in one or more embodiments of the present disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be understood that while the terms "first," "second," and so on may be employed to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another.For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the term "if" as used herein may be interpreted as "at the time of," "when," or "in response to a determination." Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data, etc.) involved in one or more embodiments of the present disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or deny. First, the terms involved in one or more embodiments of the present disclosure are explained.
[0002] SSD (Solid State Drive): is a solid state drive that uses flash memory chips to store data.
[0003] An HDD (Hard Disk Drive) is a traditional hard disk storage device that uses rotating disks and read / write heads to store and retrieve data. A Bloom filter is used to determine whether an element exists in a set and is space- and time-efficient. Figure 1 is a schematic diagram of a data processing method provided by one embodiment of the present disclosure. As shown in Figure 1, the original metadata of an original data file in a cold storage node is determined, a metadata file corresponding to the original data file is constructed based on the original metadata, and the metadata file is stored in a hot storage node. Based on the storage location information of the metadata file in the hot storage node, the original metadata contained in the metadata file is updated. The updated original metadata stored in the hot storage node records access information for the original data file stored in the cold storage node. Therefore, when a data read request is received, the data read task corresponding to the data read request is executed based on the updated original metadata in the hot storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node as a metadata file. This allows the metadata file in the hot storage node to replace the original data file in the cold storage node when the cold storage node is accessed. This improves access speed and data query efficiency while ensuring data availability. The present disclosure provides a data processing method, which also involves a data processing apparatus, a computing device, and a computer-readable storage medium. Each of these methods will be described in detail in the following embodiments. Referring to FIG. 2 , FIG. 2 shows a flowchart of a data processing method according to one embodiment of the present disclosure, specifically including the following steps: Step 202: Determine the original metadata of the original data file in the cold storage node. Specifically, cold storage nodes are used to store cold data, which is data with low access frequency. Cold storage nodes use capacity-based storage as their storage type. Cold storage nodes can be storage media in a database, and raw data files are data files stored in cold storage nodes. Raw metadata is metadata within raw data files, representing the data attributes of the cold data stored in the cold storage nodes and providing a structured description of the data. Based on this information, the corresponding cold storage node is determined for the database, and the raw data file is located within the cold storage node. Raw metadata describing the data attributes of the cold data is then located within the raw data file, facilitating subsequent separation of the raw metadata.In practical applications, a cold storage node contains multiple raw data files. When determining raw metadata, it is necessary to determine the raw metadata corresponding to each raw data file, thereby achieving comprehensive metadata separation in the cold storage node. Furthermore, considering that data reading requirements vary when performing data separation on raw data files, raw metadata can be determined based on a data separation event. Specifically, this is achieved as follows: Based on the data separation event, raw data files are determined in the cold storage node. The data separation event is used to copy the raw metadata in the raw data files in the cold storage node to the hot storage node. The raw metadata is determined based on the raw data files. Specifically, a data split event refers to an event that splits the original metadata in the original data files in the cold storage node. The data split event is used to copy the original metadata in the original data files in the cold storage node to the hot storage node. The data split event can be executed when the original data files in the cold storage node are updated, thereby splitting the original metadata in the updated original data files. The data split event can also be executed based on the data split policy corresponding to the database. The data split event can be executed at fixed intervals or for the cold storage node when data split is required. Based on this, a data split event is determined for the original data files in the cold storage node. Based on the data split event, at least one original data file is determined in the cold storage node, and then the original metadata of each original data file in the at least one original data file is determined. This achieves data splitting of the original metadata of multiple original data files in the cold storage node. For example, in a big data scenario, cold data is stored in the cold storage node and hot data is stored in the hot storage node. When reading data, it is necessary to determine the readability of the cold data in the cold storage node based on the original metadata in the original data files in the cold storage node. The original metadata from the original data files on the cold storage nodes is separated and stored on the hot storage nodes. The hot storage nodes can then perform data readability checks on behalf of the cold storage nodes. A data separation event is determined for the cold storage nodes, and based on the data separation event, the original data files are identified on the cold storage nodes. The original metadata is then read from the original data files. In summary, the original metadata is read from the original data files on the cold storage nodes based on the data separation event, thereby enabling on-demand separation of the original metadata on the cold storage nodes.Furthermore, considering that data reading may be performed in different ways, such as row-by-row reading or range-by-range reading, different raw metadata need to be determined when determining the original metadata, which is specifically implemented as follows: when the data separation event corresponds to a row data reading operation, filter metadata corresponding to the row data reading operation is determined in the original data file, and the filter metadata is used as the original metadata; when the data separation event corresponds to a range data reading operation, index metadata associated with the original data file and corresponding to the range data reading operation is determined, and the index metadata is used as the original metadata; when the data separation event corresponds to both a row data reading operation and a range data reading operation, filter metadata corresponding to the row data reading operation is determined in the original data file, and index metadata associated with the original data file and corresponding to the range data reading operation is determined, and the filter metadata and index metadata are used as the original metadata. Specifically, a row data read operation refers to a single row read of data in a storage node; filter metadata refers to metadata associated with a Bloom filter, including but not limited to Bloom block metadata; correspondingly, a range read operation refers to a range read of data in a storage node, i.e., after locating the starting position corresponding to the data to be read, the data within the range is read; index metadata includes data block index information, which is used to index the position of the data block storing the data in the original data file. Based on this, a data splitting event corresponds to a row data read operation and / or a range data read operation. If a data splitting event corresponds to a row data read operation, this indicates that a single row read is used when reading data from the cold storage node. The filter metadata associated with the Bloom filter corresponding to the row data read operation is determined in the original data file and used as the original metadata. If a data splitting event corresponds to a range data read operation, this indicates that a range read is used when reading data from the cold storage node. The index metadata associated with the original data file and corresponding to the range data read operation is determined and used as the original metadata. Accordingly, when the data splitting event corresponds to a row data read operation and a range data read operation, this indicates that both a single row read and a range read are used when reading data from the cold storage node. Filter metadata corresponding to the row data read operation is determined in the original data file, as well as index metadata associated with the original data file and corresponding to the range data read operation. The filter metadata and index metadata are used as original metadata.Continuing with the above example, for single-row read requests targeting storage nodes, each file in the database stores a Bloom filter to quickly determine whether the data being read is in the file. Separating the Bloom filter from the data files on cold storage nodes and storing it in hot storage nodes with higher read / write speeds can accelerate Bloom filter data read determination, speeding up the reading of both cold and hot data. For range read requests, index blocks (used to index the location of data blocks storing data in the data file) are required to locate the starting position of the range read in the data files on cold storage nodes. Therefore, separating the index blocks from the data files on cold storage nodes and storing them in hot storage nodes with higher read / write speeds can improve data reading efficiency. The separation of Bloom filters and index blocks can be determined based on actual data reading requirements. In summary, data separation events correspond to row data read operations and / or range data read operations, allowing for flexible determination of raw metadata to meet different data reading requirements. Step 204: Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node. Specifically, after determining the original metadata of the original data file in the cold storage node, construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node. The data read and write efficiency of the cold storage node is lower than that of the hot storage node. The hot storage node is used to store hot data, i.e., data with high access frequency. The storage type of the hot storage node can be standard storage, performance storage, local SSD, or local HDD. The metadata file is a data file constructed based on the original metadata determined in the original data file and is used to replace the original metadata in the original data file to perform subsequent data detection tasks. The data detection task is used to detect whether the cold storage node contains cold data to be read. Based on this, after determining the original metadata of the original data files in the cold storage nodes, the original metadata is combined to construct a metadata file corresponding to the original data files in the cold storage nodes. This metadata file is then stored in the hot storage nodes, allowing the metadata file to replace the original metadata in the original data files in subsequent data detection tasks. By leveraging the high data processing efficiency of the hot storage nodes, the efficiency of executing data detection tasks on the cold storage nodes is improved.Furthermore, when raw metadata is stored in a cold storage node, it has a certain storage structure. Considering that the storage structure can affect the efficiency and order of data reading, to ensure consistency in data processing results after separating the raw metadata to the hot storage node, it is necessary to construct a metadata file based on the structural information of the raw metadata. This is specifically implemented as follows: determining the structural information of the raw metadata in the raw data file; determining the distribution information of the sub-raw metadata contained in the raw metadata within the raw data file based on the structural information; and combining the sub-raw metadata contained in the raw metadata according to the distribution information of the sub-raw metadata to obtain a metadata file corresponding to the raw data file. Specifically, structural information refers to the organizational form and relationships of the raw metadata, including but not limited to information such as the data type, data relationship, and storage structure of the raw metadata; distribution information refers to the arrangement information of the sub-raw metadata within the raw data file, including arrangement position information and the relationships between the sub-raw metadata. Based on this, structural information such as the data type, data relationship, and storage structure of the raw metadata in the raw data file is determined, and distribution information of each sub-raw metadata contained in the raw metadata within the raw data file is determined based on the structural information. The sub-raw metadata contained in the original metadata is combined according to the distribution information of each sub-raw metadata to obtain the metadata file corresponding to the original data file. Continuing with the above example, as shown in Figure 3(a), when performing data separation for the data files in the cold storage node, data separation is performed on each of the N data files (data file 1 to data file N) in the cold storage node. Each data file corresponds to a metadata file, which is stored in the hot storage node. The metadata file contains metadata such as the Bloom blocks, index blocks, meta-information index, and Bloom index originally stored in the data file in the cold storage node. As shown in Figure 3(b), the original data file stored in the cold storage node contains data from multiple modules, including scanned blocks, non-scanned blocks, and the payload of the open portion. Scanned blocks correspond to Bloom blocks, non-scanned blocks correspond to meta-information blocks, and the payload of the open portion corresponds to data such as the root data index, middle key fields, meta-index, file information, and Bloom filter metadata. When extracting data from the original data file, the data corresponding to the scan block and the open portion of the load module are determined to be the original metadata to be extracted. The metadata file is constructed based on the data structure of the original metadata when stored in the cold storage node and the distribution of each sub-original metadata.In summary, the sub-primitive metadata contained in the original metadata are combined according to the distribution information of each sub-primitive metadata to obtain a metadata file corresponding to the original data file, thereby ensuring the consistency of the structural information of the original metadata before and after data separation. Step 206: Based on the storage location information of the metadata file in the hot storage node, the original metadata contained in the metadata file is updated. The updated original metadata stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node. Specifically, after constructing a metadata file corresponding to the original data file based on the original metadata and storing the metadata file on the hot storage node, the original metadata contained in the metadata file can be updated based on the storage location information of the metadata file on the hot storage node. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored on the cold storage node. The storage location information represents the location information of the metadata file on the hot storage node after the metadata file is stored on the hot storage node. The access information of the original data file refers to metadata relied upon when performing a data detection task based on the original data file, including but not limited to Bloom blocks and Bloom indexes. Based on this, after constructing a metadata file corresponding to the original data file based on the original metadata and storing the metadata file on the hot storage node, the storage location information of the metadata file on the hot storage node is determined. Based on the storage location information, the original metadata contained in the metadata file is updated. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored on the cold storage node. This access information is used to perform the data detection task. Furthermore, since the storage location of the original metadata in the cold storage node is different from that in the hot storage node, and when the original metadata is stored in the cold storage node, data blocks are searched through the internal index information of the original data file, and the index information is related to the storage location of the original metadata in the cold storage node, when the original metadata is stored in the hot storage node, the index information also needs to be updated according to the storage location information of the target sub-metadata in the original metadata in the hot storage node. The specific implementation is as follows: determining the target sub-metadata in the original metadata of the metadata file, and reading the initial data index information recorded by the target sub-metadata; and updating the initial data index information to the target data index information based on the storage location information of the target sub-metadata in the hot storage node.Specifically, target sub-metadata refers to metadata associated with the initial data index information corresponding to the original data file. Initial data index information is used to identify the data block containing the metadata before determining whether the cold storage node contains the data to be read. Storage location information is the storage address of the target sub-metadata in the hot storage node, and target index information is the index information obtained by updating the initial data index information based on the storage address of the target sub-metadata in the hot storage node. Based on this, the target sub-metadata is determined from the original metadata of the metadata file, and the initial data index information of the target sub-metadata record is read. When reading row data, the initial data index information is associated with the block offset and Bloom block; when reading range data, the initial data index information is associated with the block offset and index block. Based on the storage location information of the target sub-metadata in the hot storage node, the initial data index information is updated to the target data index information. Continuing with the above example, as shown in Figure 4(a), when reading row data, the Bloom filter of the original metadata in the cold storage node is separated. At this point, the initial data index information corresponding to the Bloom index block needs to be updated. A Bloom index block contains data such as the version and Bloom index. In the cold storage node, the Bloom index contains data such as the block key and block offset. The block offset corresponds to the cold storage index. In other words, the block offset and the Bloom block in the cold storage node constitute the initial data index information. After the original metadata is stored in the hot storage node as a metadata file, the block offset must be modified to point to the hot storage index, pointing to the Bloom block stored in the hot storage node. This achieves the index modification and updates the initial data index information to the target data index information. As shown in Figure 4(b), when reading range data, the Bloom index in the root index block must be modified. The root index block contains data such as index entries and Bloom indexes. Before the original metadata in the cold storage node is separated, the block offset corresponds to the index block, forming the cold storage index. After the original metadata is stored in the hot storage node, the cold storage index is updated to the hot storage index, and the block offset is modified to point to the index block in the hot storage node. In summary, based on the storage location information of the target sub-metadata in the hot storage node, the initial data index information is updated to the target data index information, thereby ensuring that when subsequent data reading tasks are executed based on the original metadata in the hot storage node, data detection and data judgment can be correctly performed, thereby improving the accuracy of subsequent data reading.Furthermore, considering that in actual data read scenarios, the data to be read may be stored in both cold storage nodes and hot storage nodes, it is necessary to simultaneously execute data read tasks based on both cold storage nodes and hot storage nodes. This is specifically implemented as follows: a data read request is received; hot storage metadata stored in the hot storage node and target metadata in the metadata file are read according to the data read request, where the target metadata is metadata updated from the original metadata in the metadata file. Based on the hot storage metadata and the original metadata, a data read task associated with the cold storage node and the hot storage node is executed. Specifically, the data read request is used to read data from the cold storage node and / or the hot storage node; the hot storage metadata is used to determine whether the hot storage node contains the hot data to be read; and the target metadata is used to determine whether the cold storage node contains the cold data to be read. The data read task is the task corresponding to the data read request. Based on this, a data read request is received. Hot storage metadata stored in the hot storage node and target metadata in the metadata file stored in the hot storage node are read according to the data read request. The target metadata is metadata that is updated from the original metadata in the metadata file stored in the hot storage node. Based on the hot storage metadata and the original metadata, a data read task is executed for the associated cold storage node and hot storage node, reading the data to be read corresponding to the data read request from the cold storage node and / or the hot storage node. In practical applications, both the hot storage node and the cold storage node are storage nodes in a database. The data read task includes a data detection task and a data determination task. The data detection task in the data read task includes data storage detection on the cold storage node and data storage detection on the hot storage node. The data determination task in the data read task determines the cold data to be read in the cold storage node based on the hot detection result corresponding to the hot storage node, and determines the hot data to be read in the hot storage node based on the cold detection result corresponding to the cold storage node. Continuing with the above example, when a data read request is received for a database, the cold storage nodes and hot storage nodes in the database are determined. Target metadata for performing data storage detection on the cold storage node is determined in the hot storage node, and hot storage metadata for performing data storage detection on the hot storage node is determined in the hot storage node. A data read task is executed for the associated cold storage node and the hot storage node. In summary, the data read task is executed simultaneously on the cold storage node and the hot storage node, thereby improving the comprehensiveness of data reading.Furthermore, considering that when a data read task is executed simultaneously on a cold storage node and a hot storage node, the cold storage node may or may not contain the data to be read, and accordingly, the hot storage node may or may not contain the data to be read. Therefore, data storage detection is required for both cold and hot storage nodes. This is specifically implemented as follows: a data storage detection is performed on the hot storage node based on the hot storage metadata, and a data storage detection is performed on the cold storage node based on the original metadata. A data determination task is executed as part of the data read task based on the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node. Specifically, the data storage detection refers to an operation performed on the hot storage node to determine whether the hot storage node contains the data to be read corresponding to the data read request, and an operation performed on the cold storage node to determine whether the cold storage node contains the data to be read corresponding to the data read request. Accordingly, performing the data storage detection on the hot storage node obtains the hot detection result corresponding to the hot storage node; and performing the data storage detection on the cold storage node obtains the cold detection result corresponding to the cold storage node. The hot detection result indicates whether the hot storage node contains hot data to be read, while the cold detection result indicates whether the cold storage node contains cold data to be read. Based on this, a data storage detection is performed on the hot storage node based on the hot storage metadata to determine whether the hot storage node contains the data to be read corresponding to the data read request. Furthermore, a data storage detection is performed on the cold storage node based on the original metadata to determine whether the cold storage node contains the data to be read corresponding to the data read request. The data determination task of associating the hot and cold storage nodes is executed based on the hot detection result obtained during the data storage detection of the hot storage node, or the data determination task of associating the hot and cold storage nodes is executed based on the cold detection result corresponding to the cold storage node obtained during the data storage detection of the cold storage node. This execution process is the execution process of the data read task. Continuing with the above example, target metadata for performing data storage detection on the cold storage node is determined in the hot storage node, and hot storage metadata for performing data storage detection on the hot storage node is determined in the hot storage node. A data storage check is performed based on the target metadata to determine whether the cold storage node contains the data to be read, obtaining a cold detection result. A data storage check is also performed based on the hot storage metadata to determine whether the hot storage node contains the data to be read, obtaining a hot detection result. Subsequently, a subsequent data determination task of associating the hot and cold storage nodes is performed based on the obtained hot and cold detection results.In summary, data storage detection is performed simultaneously on both cold and hot storage nodes to determine whether they contain the data to be read, thereby improving the accuracy of subsequent data reading. Data storage detection is performed based on the target metadata stored in the hot storage node, eliminating the need to perform data storage detection based on the original metadata in the cold storage node, thereby improving data storage detection efficiency. Furthermore, when data storage detection is performed on the cold storage node and the hot storage node simultaneously, the detection results corresponding to the cold storage node and the hot storage node respectively correspond to two data storage situations. The data stored in the cold storage node may or may not need to be read. Correspondingly, the data stored in the hot storage node may or may not need to be read. Therefore, the data to be read can be determined based on the hot detection result and the cold detection result. The specific implementation is as follows: when it is determined according to the hot detection result corresponding to the hot storage node that the hot storage node contains hot data to be read corresponding to the data read request, the hot data to be read corresponding to the data read request is read from the hot storage node as the execution of the data determination task; when it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains cold data to be read corresponding to the data read request, the cold data to be read corresponding to the data read request is read from the cold storage node as the execution of the data determination task. Specifically, the hot data to be read is the hot data stored in the hot storage node and corresponding to the data read request. Correspondingly, the cold data to be read is the cold data stored in the cold storage node and corresponding to the data read request. Based on this, if the hot detection result corresponding to the hot storage node determines that the hot storage node contains the hot data to be read corresponding to the data read request, it indicates that the hot data in the hot storage node needs to be read. The hot data to be read corresponding to the data read request is read from the hot storage node, and the determination of the hot data to be read is performed as the data determination task. If the cold detection result corresponding to the cold storage node determines that the cold storage node contains the cold data to be read corresponding to the data read request, it indicates that the cold data in the cold storage node needs to be read. The cold data to be read corresponding to the data read request is read from the cold storage node, and the determination of the cold data to be read is performed as the data determination task. Correspondingly, when it is determined based on the hot detection result corresponding to the hot storage node that the hot storage node does not contain the hot data to be read corresponding to the data read request, it means that the hot data in the hot storage node does not need to be read; when it is determined based on the cold detection result corresponding to the cold storage node that the cold storage node does not contain the cold data to be read corresponding to the data read request, it means that the cold data in the cold storage node also does not need to be read.Continuing with the above example, after determining the hot detection result, the hot storage node can be determined based on the hot detection result to determine whether it contains the hot data to be read corresponding to the data read request. If the hot detection result determines that the hot storage node contains the hot data to be read corresponding to the data read request, the hot data to be read can be read as a response to the data read request. Correspondingly, after determining the cold detection result, the cold detection result can be determined based on the cold detection result to determine whether the cold storage node contains the cold data to be read corresponding to the data read request. If the cold detection result determines that the cold storage node contains the cold data to be read corresponding to the data read request, the cold data to be read can be read as a response to the data read request. In summary, by determining whether a hot storage node contains the hot data to be read based on the hot detection result and determining whether a cold storage node contains the cold data to be read based on the cold detection result, accurate data reading can be achieved for both hot and cold storage nodes. Furthermore, considering that the metadata file in the hot storage node may be inaccessible, in this case, the data readability cannot be determined using the metadata file in the hot storage node when executing a data processing task. Therefore, the data readability determination for the cold storage node can be achieved using the original metadata still retained in the cold storage node. Specifically, the following implementation is performed: If the metadata file in the hot storage node is inaccessible, the data processing task is performed based on the original metadata of the original data file in the cold storage node. Based on this, if the metadata file in the hot storage node has not been updated, resulting in the metadata file being inaccessible, the original metadata of the original data file in the cold storage node is determined, and the data processing task is performed based on the original metadata in the cold storage node to determine whether the cold data in the cold storage node needs to be read. In summary, the original metadata stored in the cold storage node is used to assist in executing the data processing task and complete the data readability detection in the cold storage node when the metadata file in the hot storage node is inaccessible, thereby achieving disaster recovery. One embodiment of the present disclosure determines original metadata of an original data file in a cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata, and stores the metadata file in a hot storage node; and updates the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored in the cold storage node.The original metadata of the original data files in the cold storage node is separated from the cold storage node and stored in the hot storage node as metadata files. This allows the metadata files in the hot storage node to be accessed instead of the original data files in the cold storage node when the cold storage node is accessed. This improves access speed and data query efficiency while ensuring data availability. The data processing method provided by the present disclosure is further described below, using the application of the data processing method in data reading as an example, in conjunction with FIG5 . FIG5 shows a flowchart of a data processing method provided by one embodiment of the present disclosure, which specifically includes the following steps: Step 502: Determine the original data file in the cold storage node based on the data separation event. In massive big data scenarios, a table often stores a large amount of historical data, such as order data or monitoring data. Over time, the access frequency of this data gradually decreases and is eventually shelved. Data with lower access frequencies is stored as cold data in the cold storage node, while data with higher access frequencies is stored as hot data in the hot storage node, thereby reducing data storage costs. Cold storage is capacity-based storage, while hot storage is standard storage, performance-based storage, local SSDs, or local HDDs. The data read and write speeds of cold storage nodes are significantly lower than those of hot storage nodes. Even when querying only hot data, metadata of the cold files containing the cold data must be read to determine whether to read the cold files. Therefore, the metadata of the cold files can be extracted from the original file structure, reorganized into a metadata file, and stored in the hot storage node. The original data files in the cold storage node serve as the metadata of the cold files. A data separation event is a data separation task determined according to pre-defined data separation rules or policies. When the data separation task is executed, the cold file metadata is extracted from the original file structure, reorganized into a metadata file, and stored in the hot storage node. Step 504: If the data separation event corresponds to a single-row data read operation, filtered metadata corresponding to the single-row data read operation is determined in the original data file, and the filtered metadata is used as the original metadata. Considering that data reading includes single-row reading and range reading, the filter metadata corresponding to the Bloom filter and / or the index metadata corresponding to the index block can be determined based on requirements. Step 506: If the data split event corresponds to a range data read operation, index metadata associated with the original data file and corresponding to the range data read operation is determined, and the index metadata is used as the original metadata.Step 508: If the data separation event corresponds to a single-row data read operation and a range data read operation, filter metadata corresponding to the single-row data read operation and index metadata associated with the original data file and corresponding to the range data read operation are determined in the original data file. The filter metadata and index metadata are then used as original metadata. Step 510: The structural information of the original metadata in the original data file is determined, and based on the structural information, the distribution information of the metadata contained in the original metadata within the original data file is determined. Step 512: The metadata contained in the original metadata is combined according to the distribution information to obtain a metadata file corresponding to the original data file, and the metadata file is stored in the hot storage node. During data separation, it is necessary to ensure the consistency of the original metadata in the cold storage node and the hot storage node. Therefore, the distribution information of the metadata contained in the original metadata within the original data file can be determined based on the structural information of the original metadata in the original data file. Then, a metadata file is constructed based on the distribution information of each metadata and stored in the hot storage node. The file structure of the original data file in the cold storage node is not destroyed, achieving the effect of storing the original metadata in both the cold storage node and the hot storage node. Step 514: Determine the target metadata from the original metadata in the metadata file and read the initial data index information used by the target metadata. Step 516: Update the initial data index information to the target data index information based on the storage location information of the target metadata in the hot storage node. In actual applications, because the storage location of the original metadata in the cold storage node is different from that in the hot storage node, it is necessary to modify the index information within the metadata file to ensure that the data block can be correctly found. This means modifying the Bloom block pointed to by the block offset in the Bloom index block and / or modifying the index block pointed to by the block offset in the root index block. Step 518: Read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file in response to a data read request. Upon receiving the data read request, the hot storage metadata originally stored in the hot storage node can be used to determine whether the hot storage node contains the data to be read. Furthermore, the cold storage node can be used to determine whether the data to be read contains the target metadata corresponding to the original metadata copied from the cold storage node in the hot storage node. Step 520: Execute a data read task for the associated cold storage node and hot storage node based on the hot storage metadata and the original metadata. In actual applications, a data storage check is performed on the hot storage node based on the hot storage metadata to determine whether the hot storage node contains the data to be read. Furthermore, a data storage check is performed on the cold storage node based on the original metadata to determine whether the cold storage node contains the data to be read.If the hot detection result corresponding to the hot storage node determines that the hot storage node contains hot data to be read corresponding to the data read request, the hot data to be read corresponding to the data read request is read from the hot storage node. If the cold detection result corresponding to the cold storage node determines that the cold storage node contains cold data to be read corresponding to the data read request, the cold data to be read corresponding to the data read request is read from the cold storage node, completing the data read task. In summary, the original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node as a metadata file. This allows the metadata file in the hot storage node to replace the original data file in the cold storage node when the cold storage node is accessed. While ensuring data availability, this improves access speed and data query efficiency. This not only improves cold data query efficiency but also accelerates hot data reading, speeding up single-row read requests and reducing the execution time of range read requests. Furthermore, the Bloom filter and / or index block separation in the cold storage node can be freely selected based on the usage scenario, increasing the flexibility of data separation. Furthermore, since a Bloom filter has a false positive rate, meaning that data determined by the Bloom filter to exist has a certain probability of not actually existing, data separation can independently adjust the false positive rate of the Bloom filter for the metadata file stored in the hot storage node, thereby avoiding impacting the structure and space size of the original data file. Corresponding to the above-described method embodiments, the present disclosure also provides an embodiment of a data processing device. FIG6 shows a schematic structural diagram of a data processing device provided by one embodiment of the present disclosure. As shown in FIG6 , the device includes: a determination module 602 configured to determine original metadata of an original data file in a cold storage node; a storage module 604 configured to construct a metadata file corresponding to the original data file based on the original metadata and store the metadata file in the hot storage node; and an update module 606 configured to update the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for accessing the original data file stored in the cold storage node. In an optional embodiment, the determining module 602 is further configured to: determine an original data file in the cold storage node based on a data separation event, wherein the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; and determine the original metadata based on the original data file.In an optional embodiment, the determination module 602 is further configured to: if the data separation event corresponds to a row data read operation, determine, in the original data file, filter metadata corresponding to the row data read operation, and use the filter metadata as the original metadata; if the data separation event corresponds to a range data read operation, determine index metadata associated with the original data file and corresponding to the range data read operation, and use the index metadata as the original metadata; if the data separation event corresponds to both a row data read operation and a range data read operation, determine, in the original data file, filter metadata corresponding to the row data read operation, and determine index metadata associated with the original data file and corresponding to the range data read operation, and use the filter metadata and index metadata as the original metadata. In an optional embodiment, the update module 606 is further configured to: determine target sub-metadata from the original metadata of the metadata file, read initial data index information recorded by the target sub-metadata; and update the initial data index information to the target data index information based on the storage location information of the target sub-metadata in the hot storage node. In an optional embodiment, the storage module 604 is further configured to: determine structural information of raw metadata in the raw data file; determine distribution information of sub-raw metadata contained in the raw metadata in the raw data file based on the structural information; and combine the sub-raw metadata contained in the raw metadata according to the distribution information of the sub-raw metadata to obtain a metadata file corresponding to the raw data file. In an optional embodiment, the update module 606 is further configured to: receive a data read request; read the hot storage metadata stored by the hot storage node and the target metadata in the metadata file according to the data read request, wherein the target metadata is metadata updated from the raw metadata in the metadata file; and execute a data read task associated with the cold storage node and the hot storage node based on the hot storage metadata and the raw metadata. In an optional embodiment, the update module 606 is further configured to: perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata; and execute a data determination task as execution of the data reading task according to the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node.In an optional embodiment, the update module 606 is further configured to: if it is determined based on the hot detection result corresponding to the hot storage node that the hot storage node contains hot data to be read corresponding to the data read request, read the hot data to be read corresponding to the data read request from the hot storage node as execution of the data determination task; and if it is determined based on the cold detection result corresponding to the cold storage node that the cold storage node contains cold data to be read corresponding to the data read request, read the cold data to be read corresponding to the data read request from the cold storage node as execution of the data determination task. In an optional embodiment, the update module 606 is further configured to: if the metadata file in the hot storage node is inaccessible, perform a data processing task based on the original metadata of the original data file in the cold storage node. In summary, one embodiment of the present disclosure determines the original metadata of an original data file in a cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata, and stores the metadata file in a hot storage node; and updates the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for the original data file stored in the cold storage node. The original metadata of the original data file in the cold storage node is separated from the cold storage node and stored in the hot storage node as a metadata file. This allows the metadata file in the hot storage node to replace the original data file in the cold storage node when the cold storage node is accessed. This improves access speed and data query efficiency while ensuring data availability. The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the aforementioned data processing method share the same concept. For details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the aforementioned data processing method. 7 , which shows a flow chart of another data processing method provided according to an embodiment of the present disclosure, specifically including the following steps.Step 702: Determine the original metadata of the original data file in the cold storage node; Step 704: Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node; Step 706: Update the original metadata contained in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node; Step 708: Upon receiving a data processing request, execute the data processing task associated with the cold storage node based on the target metadata stored in the hot storage node. In practical applications, in massive data scenarios, less frequently accessed data is stored in the cold storage node, while more frequently accessed data is stored in the hot storage node. When reading data, a Bloom filter can be used to determine whether the data to be read exists in the cold storage node. Due to the low data read and write speed of cold storage nodes, the efficiency of reading the original metadata in the cold storage node is also low. To address this issue, the original metadata of the original data file in the cold storage node is determined, a metadata file corresponding to the original data file is constructed based on the original metadata, and the metadata file is stored in the hot storage node. Upon receiving a data processing request, the data processing task associated with the cold storage node is executed based on the target metadata stored in the hot storage node. In practical applications, because the storage location information of the metadata file in the hot storage node is different from the storage location information of the original data file in the cold storage node, the original metadata contained in the metadata file can be updated to the target metadata based on the storage location information of the metadata file in the hot storage node. Subsequent data processing tasks can then be completed using the target metadata stored in the hot storage node. In summary, one embodiment of the present disclosure determines the original metadata of the original data file in the cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata; and stores the metadata file in the hot storage node; and updates the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for the original data file stored in the cold storage node. The original metadata of the original data files in the cold storage nodes are separated from the cold storage nodes and stored in the hot storage nodes in the form of metadata files. Therefore, when the cold storage nodes are accessed during the execution of data processing tasks, the metadata files in the hot storage nodes can be accessed instead of the original data files in the cold storage nodes. Under the premise of ensuring data availability, the access speed and data query efficiency are improved.Corresponding to the above-mentioned method embodiments, the present disclosure also provides an embodiment of a data processing device. FIG8 shows a schematic structural diagram of another data processing device provided by one embodiment of the present disclosure. As shown in FIG8 , the device includes: a determination module 802 configured to determine original metadata of an original data file in a cold storage node; a storage module 804 configured to construct a metadata file corresponding to the original data file based on the original metadata and store the metadata file in a hot storage node; an update module 806 configured to update the original metadata contained in the metadata file to target metadata based on the storage location information of the metadata file in the hot storage node; and an execution module 808 configured to, upon receiving a data processing request, execute a data processing task associated with the cold storage node based on the target metadata stored in the hot storage node. In an optional embodiment, the determination module 802 is further configured to: determine an original data file in the cold storage node based on a data split event, wherein the data split event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; and determine the original metadata based on the original data file. In an optional embodiment, the determination module 802 is further configured to: if the data separation event corresponds to a row data read operation, determine, in the original data file, filter metadata corresponding to the row data read operation, and use the filter metadata as the original metadata; if the data separation event corresponds to a range data read operation, determine index metadata associated with the original data file and corresponding to the range data read operation, and use the index metadata as the original metadata; if the data separation event corresponds to both a row data read operation and a range data read operation, determine, in the original data file, filter metadata corresponding to the row data read operation, and determine index metadata associated with the original data file and corresponding to the range data read operation, and use the filter metadata and index metadata as the original metadata. In an optional embodiment, the update module 806 is further configured to: determine target sub-metadata from the original metadata of the metadata file, read initial data index information recorded by the target sub-metadata; and update the initial data index information to the target data index information based on the storage location information of the target sub-metadata in the hot storage node.In an optional embodiment, the storage module 804 is further configured to: determine structural information of raw metadata in the raw data file; determine distribution information of sub-raw metadata contained in the raw metadata in the raw data file based on the structural information; and combine the sub-raw metadata contained in the raw metadata according to the distribution information of the sub-raw metadata to obtain a metadata file corresponding to the raw data file. In an optional embodiment, the update module 806 is further configured to: receive a data read request; read the hot storage metadata stored by the hot storage node and the target metadata in the metadata file according to the data read request, wherein the target metadata is metadata updated from the raw metadata in the metadata file; and execute a data read task associated with the cold storage node and the hot storage node based on the hot storage metadata and the raw metadata. In an optional embodiment, the update module 806 is further configured to: perform a data storage check on the hot storage node based on the hot storage metadata, and perform a data storage check on the cold storage node based on the original metadata; and execute a data determination task based on the hot detection result corresponding to the hot storage node or the cold detection result corresponding to the cold storage node, as the execution of the data read task. In an optional embodiment, the update module 806 is further configured to: if it is determined based on the hot detection result corresponding to the hot storage node that the hot storage node contains hot data to be read corresponding to the data read request, read the hot data to be read corresponding to the data read request from the hot storage node as the execution of the data determination task; and if it is determined based on the cold detection result corresponding to the cold storage node that the cold storage node contains cold data to be read corresponding to the data read request, read the cold data to be read corresponding to the data read request from the cold storage node as the execution of the data determination task. In an optional embodiment, the update module 806 is further configured to: if the metadata file in the hot storage node is inaccessible, perform a data processing task based on the original metadata of the original data file in the cold storage node. In summary, one embodiment of the present disclosure determines the original metadata of the original data file in the cold storage node; constructs a metadata file corresponding to the original data file based on the original metadata, and stores the metadata file in the hot storage node; and updates the original metadata contained in the metadata file based on the storage location information of the metadata file in the hot storage node. The updated original metadata stored in the hot storage node records access information for the original data file stored in the cold storage node.The original metadata of the original data files in the cold storage nodes is separated from the cold storage nodes and stored in the hot storage nodes in the form of metadata files. Therefore, when the cold storage nodes are accessed during data processing tasks, the metadata files in the hot storage nodes can be accessed instead of the original data files in the cold storage nodes. This improves access speed and data query efficiency while ensuring data availability. The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solutions of this data processing device and the technical solutions of the aforementioned data processing method are based on the same concept. For details not described in detail in the technical solutions of the data processing device, please refer to the description of the technical solutions of the aforementioned data processing method. Figure 9 shows a block diagram of a computing device 900 according to one embodiment of the present disclosure. Components of computing device 900 include, but are not limited to, a memory 910 and a processor 920. oThe processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data. The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface. In one embodiment of the present disclosure, the aforementioned components of the computing device 900 and other components not shown in FIG. 9 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 9 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed. Computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 900 may also be a mobile or stationary server.The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned method. The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the aforementioned method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned method. An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned method. The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the aforementioned method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned method. An embodiment of the present disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the aforementioned method. The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above-described method share the same concept. For details not described in detail in the technical solution of the computer program, reference should be made to the description of the technical solution of the above-described method. The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. The computer instructions include computer program code, which may be in source code form, object code form, executable files, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.It should be noted that, for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present disclosure are not limited by the order of the actions described, as certain steps may be performed in a different order or simultaneously, depending on the embodiments of the present disclosure. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules described are not necessarily required for the embodiments of the present disclosure. In the above embodiments, the description of each embodiment has its own emphasis. For portions not described in detail in a particular embodiment, reference should be made to the relevant descriptions of other embodiments. The preferred embodiments disclosed above are merely intended to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific implementations described. Obviously, many modifications and variations are possible based on the content of the embodiments of the present disclosure. The present disclosure selects and describes these embodiments in detail to better explain the principles and practical applications of the embodiments of the present disclosure, thereby enabling those skilled in the art to better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
Claims 1. A data processing method, comprising: Determine the original metadata of the original data file in the cold storage node; Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the hot storage node; Update the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node; Among them, the original metadata that has been updated and stored in the hot storage node records the access information for accessing the original data file stored in the cold storage node.
2. The data processing method according to claim 1, wherein the determining the original metadata of the original data file in the cold storage node comprises: Determine the original data file in the cold storage node based on the data separation event, where the data separation event is used to copy the original metadata in the original data file in the cold storage node to the hot storage node; determine the original metadata based on the original data file.
3. The data processing method according to claim 2, wherein determining the original metadata based on the original data file comprises: In the case of a row data reading operation corresponding to the data separation event, determine the filtering metadata corresponding to the row data reading operation in the original data file, and use the filtering metadata as the original metadata; In the case of a range data reading operation corresponding to the data separation event, determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the index metadata as the original metadata; In the case of a row data reading operation and a range data reading operation corresponding to the data separation event, determine the filtering metadata corresponding to the row data reading operation in the original data file, and determine the index metadata associated with the original data file and corresponding to the range data reading operation, and use the filtering metadata and the index metadata as the original metadata.
4. The data processing method according to any one of claims 1 to 3, wherein updating the original metadata included in the metadata file based on the storage location information of the metadata file in the hot storage node comprises: Determine the target sub-metadata in the original metadata of the metadata file, and read the initial data index information recorded by the target sub-metadata; Update the initial data index information to target data index information based on the storage location information of the target sub-metadata in the hot storage node.
5. The data processing method according to any one of claims 1 to 4, wherein constructing the metadata file corresponding to the original data file based on the original metadata comprises: Determine the structure information of the original metadata in the original data file; Determine the distribution information of the sub-original metadata included in the original metadata in the original data file based on the structure information; combine the sub-original metadata included in the original metadata according to the distribution information of the sub-original metadata to obtain the metadata file corresponding to the original data file.
6. The data processing method according to any one of claims 1 to 5, the method further comprising: Receive a data reading request; Read the hot storage metadata stored in the hot storage node and the target metadata in the metadata file according to the data reading request, where the target metadata is the updated metadata of the original metadata in the metadata file metadata; Execute a data reading task that associates the cold storage node and the hot storage node based on the hot storage metadata and the original metadata.
7. The data processing method according to claim 6, wherein the step of performing a data reading task of associating the cold storage node and the hot storage node based on the thermal storage metadata and the original metadata includes: Perform data storage detection on the hot storage node based on the hot storage metadata, and perform data storage detection on the cold storage node based on the original metadata; Execute a data determination task according to the heat detection result corresponding to the heat storage node or the cold detection result corresponding to the cold storage node, as the execution of the data reading task.
8. The data processing method according to claim 7, wherein performing a data determination task according to the thermal detection result corresponding to the thermal storage node or the cold detection result corresponding to the cold storage node as the execution of the data reading task includes: In the case that it is determined according to the heat detection result corresponding to the heat storage node that the heat storage node contains the to-be-read heat data corresponding to the data reading request, read the to-be-read heat data corresponding to the data reading request in the heat storage node, as the execution of the data determination task; In the case that it is determined according to the cold detection result corresponding to the cold storage node that the cold storage node contains the to-be-read cold data corresponding to the data reading request, read the to-be-read cold data corresponding to the data reading request in the cold storage node, as the execution of the data determination task.
9. The data processing method according to any one of claims 1 to 8, the method further comprising: In the case that the metadata file in the heat storage node cannot be accessed, perform a data processing task based on the original metadata of the original data file in the cold storage node.
10. A data processing method, comprising: Determine the original metadata of the original data file in the cold storage node; Construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the heat storage node; Update the original metadata included in the metadata file to target metadata based on the storage location information of the metadata file in the heat storage node; In the case of receiving a data processing request, perform a data processing task related to the cold storage node based on the target metadata stored in the heat storage node.
11. A data processing device, comprising: A determination module, configured to determine the original metadata of the original data file in the cold storage node; A storage module, configured to construct a metadata file corresponding to the original data file based on the original metadata, and store the metadata file in the heat storage node; An update module, configured to update the original metadata included in the metadata file based on the storage location information of the metadata file in the heat storage node; wherein, the original metadata after update and stored in the heat storage node records the access information for accessing the original data file stored in the cold storage node.
12. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data processing method described in any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the data processing method described in any one of claims 1 to 10 are implemented.
14. A computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the data processing method described in any one of claims 1 to 10 are implemented. 19
Citation Information
Patent Citations
Cloud storage client and high-efficiency data access method thereof
CN104580437A
Reading method for distributed file system, client device and distributed file system
CN105468660A
Cold and hot data storage method and device, equipment and medium
CN117193652A