Data processing methods, apparatus, electronic equipment, media and program products
By building data access records in the Kafka system, hot and cold data on SSDs and HDDs are separated, which solves the problem of low data retrieval efficiency under hybrid storage, improves data retrieval efficiency and reduces storage costs.
Patent Information
- Application Number
- CN202410220809.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-02-28
AI Technical Summary
In the Kafka system, data retrieval efficiency is low when using a hybrid SSD and HDD storage solution because it requires multiple disk accesses.
By constructing data access records to record node information of data stored on SSDs and HDDs, and utilizing the principle of disk locality and data access characteristics to separate hot and cold data, hierarchical data storage is achieved, avoiding multiple hard drive accesses.
It improved data acquisition efficiency, reduced Kafka read/write time and disk I/O load, and lowered storage costs.
Smart Images

Figure CN118819395B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and in particular relates to a data processing method, apparatus, electronic device, medium and program product. Background Technology
[0002] Kafka's excellent I / O optimization and multiple asynchronous designs give it higher throughput compared to other message queue systems, while ensuring good latency, making it very suitable for use in the entire big data ecosystem. As a caching middleware, it can handle various data structures, decouple producers and consumers, and provide high throughput and low latency. However, Kafka's high throughput and low latency mainly rely on I / O optimization. Ordinary disk I / O, even with zero-copy and system caching, is still limited by the characteristics of the disk itself.
[0003] To provide high throughput and low latency, faster storage devices such as solid-state drives (SSDs) can be used. However, SSDs are expensive. To balance cost, a hybrid storage system of SSDs and hard disk drives (HDDs) can be used for Kafka data storage.
[0004] Currently, when Kafka uses SSDs and HDDs for data storage, when data needs to be retrieved, it first searches the SSD to determine if the data is stored there. If not, it then searches the HDD. This entire process requires multiple accesses to the hard drive, resulting in low data retrieval efficiency. Summary of the Invention
[0005] This application provides a data processing method, apparatus, electronic device, medium, and program product that can avoid the problem of low data acquisition efficiency caused by multiple accesses to the hard disk.
[0006] In a first aspect, embodiments of this application provide a data processing method, the method comprising:
[0007] Get the data access address;
[0008] The node information corresponding to the data access address is searched from the pre-acquired data access records to obtain the search result. The data access records are used to record the node information of the data stored in the first hard disk.
[0009] If the search result includes target node information, then the target data corresponding to the target node information is obtained from the first hard disk, wherein the target node information includes the description information of the target data;
[0010] If the search result does not include the target node information, the target data is obtained from the second hard disk, wherein the access frequency of the data stored on the first hard disk is higher than the access frequency of the data stored on the second hard disk.
[0011] Secondly, embodiments of this application provide a data processing apparatus, the apparatus comprising:
[0012] The first acquisition module is used to acquire the data access address;
[0013] The lookup module is used to search for the node information corresponding to the data access address from the pre-acquired data access records and obtain the search result. The data access records are used to record the node information of the data stored in the first hard disk.
[0014] The second acquisition module is used to acquire target data corresponding to the target node information from the first hard disk if the search result includes target node information, wherein the target node information includes descriptive information of the target data;
[0015] The third acquisition module is used to acquire the target data from the second hard disk if the search result does not include the target node information, wherein the access frequency of the data stored on the first hard disk is higher than the access frequency of the data stored on the second hard disk.
[0016] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions;
[0017] When the processor executes the computer program instructions, it implements the data processing method as described in the first aspect.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the data processing method as described in the first aspect.
[0019] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the data processing method as described in the first aspect.
[0020] The data processing method, apparatus, electronic device, medium, and program product of this application embodiment obtain a data access address; search for node information corresponding to the data access address from a pre-acquired data access record to obtain a search result, wherein the data access record is used to record node information of data stored in a first hard disk; if the search result includes target node information, then target data corresponding to the target node information is obtained from the first hard disk, wherein the target node information includes descriptive information of the target data; if the search result does not include the target node information, then the target data is obtained from a second hard disk, wherein the access frequency of data stored on the first hard disk is higher than the access frequency of data stored on the second hard disk. In the above, when obtaining target data, the storage location of the target data can be directly determined by querying the data access record, avoiding multiple hard disk accesses and improving data acquisition efficiency. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0023] Figure 2a This is another schematic flowchart of the data processing method provided in the embodiments of this application;
[0024] Figure 2b This is a schematic diagram of node information provided in an embodiment of this application;
[0025] Figure 2c This is a chain representation of the data processing method provided in the embodiments of this application;
[0026] Figure 3 This is a schematic diagram of the data processing method provided in the embodiments of this application;
[0027] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0029] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0030] To address the problems of the prior art, embodiments of this application provide a data processing method, apparatus, electronic device, medium, and program product. The data processing method provided in the embodiments of this application will be described first.
[0031] Figure 1 A flowchart illustrating a data processing method provided in one embodiment of this application is shown. Figure 1 As shown, the data processing method provided in this application embodiment is applied to an electronic device and includes the following steps 101-104, wherein:
[0032] Step 101: Obtain the data access address.
[0033] The data processing method provided in this application embodiment can be applied to electronic devices, such as the Kafka system. The Kafka system adopts a hierarchical storage structure that combines a first hard disk and a second hard disk. The access frequency of data stored on the first hard disk is higher than that of data stored on the second hard disk. The first hard disk can be an SSD and the second hard disk can be an HDD.
[0034] Step 102: Search for the node information corresponding to the data access address from the pre-acquired data access records to obtain the search results. The data access records are used to record the node information of the data stored in the first hard disk.
[0035] Data access records can be stored in the memory of electronic devices. Since data access to memory is very fast, storing data access records in memory can improve the efficiency of finding node information.
[0036] Based on the principle of locality of reference and data access characteristics, data that is frequently read and written within a certain period of time is more likely to be accessed later and should be stored on an SSD. On the other hand, data with low read and write frequency has a lower probability of being accessed and should be stored on an HDD. This allows for hierarchical storage of data and separation of hot and cold data.
[0037] Step 103: If the search result includes target node information, then obtain the target data corresponding to the target node information from the first hard disk. The target node information includes the description information of the target data.
[0038] For example, the target node information may include the storage address of the target data on the first hard disk, the number of times the target data has been accessed, and the time of the most recent modification of the target data.
[0039] The data access record includes multiple linked lists, each linked list corresponds to a number of accesses, and node information with the same number of accesses is located in the same linked list.
[0040] If the search results include target node information, then the target data is retrieved from the first hard drive according to the storage address recorded in the target node information.
[0041] Step 104: If the search result does not include the target node information, then obtain the target data from the second hard disk.
[0042] If the search results do not include the target node information, it means that the target data is stored on the second hard drive. The target data is retrieved from the second hard drive according to the data access address.
[0043] In this embodiment, a data access address is obtained; the node information corresponding to the data access address is searched from the pre-acquired data access records to obtain a search result. The data access records are used to record the node information of data stored on the first hard drive; if the search result includes target node information, the target data corresponding to the target node information is obtained from the first hard drive, and the target node information includes the description information of the target data; if the search result does not include the target node information, the target data is obtained from the second hard drive, wherein the access frequency of data stored on the first hard drive is higher than the access frequency of data stored on the second hard drive. In the above embodiment, when obtaining target data, the data access records can be queried to determine whether the target data is stored on the first hard drive or the second hard drive, thereby directly determining the storage location of the target data, avoiding multiple hard drive accesses, and improving data acquisition efficiency.
[0044] In one embodiment of this application, before searching for the node information corresponding to the data access address from the pre-acquired data access records, the method further includes:
[0045] Acquire the data to be stored and store the data to be stored in the first hard disk;
[0046] Construct node information for the data to be stored, including the storage address of the data to be stored on the first hard disk, the number of times the data to be stored has been accessed, and the time of the most recent modification of the data to be stored.
[0047] The node information of the data to be stored is recorded in the data access record. The data access record includes multiple linked lists, each linked list corresponds to a number of accesses, and node information with the same number of accesses is located in the same linked list.
[0048] In the above, the data to be stored can be messages reported by the smart gateway. Gateway attribute and running status data are collected through the OSGI API, and message data is sent to Kafka through Logstash. The data is processed by Kafka's algorithm eviction mechanism, and a hierarchical data storage architecture mechanism is constructed. Finally, the big data consumer reads the Kafka data for analysis and processing.
[0049] When Logstash sends data to Kafka, it's a write operation. The data to be stored is first written to the PageCache, and then flushed to the SSD after a certain period. This means the data is stored on the first hard drive, and the data access record is updated simultaneously. For example, when storing data for the first time, the access count in the node information of the data to be stored is set to 1, and the node information is placed in the linked list corresponding to record 1 in the data access record. The data access record can include multiple linked lists, and each linked list can include multiple node information. Node information with the same access count is located in the same linked list, and node information with different access counts is located in different linked lists.
[0050] The linked list can be a doubly linked list; there is no restriction on this.
[0051] In the above, recording node information of data stored on the first hard drive through data access records makes it easier to determine whether the data is stored on the first or second hard drive when retrieving data later, thereby reducing the number of hard drive accesses and improving data retrieval efficiency.
[0052] In one embodiment of this application, before searching for the node information corresponding to the data access address from the pre-acquired data access records, the method further includes:
[0053] Based on the data to be stored, storage nodes in the storage structure are constructed. Each storage node includes a first parameter value and a second parameter value. The first parameter value is determined based on the partition information where the data to be stored is located. The second parameter value is an identifier of the node information of the data to be stored. The storage structure includes multiple storage nodes, and each storage node has different first parameter values and second parameter values.
[0054] The step of searching for the node information corresponding to the data access address from the pre-acquired data access records to obtain the search result includes:
[0055] The value of the third parameter is determined based on the data access address;
[0056] The target storage node is obtained by searching the storage structure according to the third parameter value, and the first parameter value of the target storage node is the same as the third parameter value.
[0057] Obtain the second parameter value of the target storage node;
[0058] The search is performed in the data access record based on the second parameter value to obtain the search result.
[0059] In the above, the partition information may include the Kafka topic name, partition number, and file offset. The first parameter value can be a value composed of the Kafka topic name, partition number, and file offset, and the second parameter value can be an identifier of the node information of the data to be stored.
[0060] The data access address includes the subject name, partition number, and file sequence number of the target data storage. Based on the data access address, the third parameter value can be obtained. The third parameter value is compared with the first parameter values of each storage node in the storage structure to identify the target storage node, whose first parameter value is the same as its third parameter value.
[0061] The search is performed in the data access record based on the second parameter value of the target storage node. If there is a linked list in the data access record that includes the target node information, and the identifier of the target node information is the same as the second parameter value, then the search result includes the target node information, the target data is stored on the first hard disk, and the target data can be obtained by accessing the first hard disk; otherwise, the search result does not include the target node information, the target data is stored on the second hard disk, and the target data can be obtained by accessing the second hard disk.
[0062] In one embodiment of this application, the method further includes:
[0063] The access count of the data to be stored is updated based on one of the following:
[0064] If the number of accesses to the data to be stored is greater than or equal to a preset threshold, then the number of accesses to the data to be stored in the node information of the data to be stored is incremented by 1;
[0065] Alternatively, the decay value can be calculated based on the time of this reading of the data to be stored and the time of the most recent modification of the data to be stored.
[0066] If the attenuation value is greater than the number of times the data to be stored has been accessed, then the number of times the data to be stored has been accessed in the node information of the data to be stored is modified to 1.
[0067] If the attenuation value is less than or equal to the number of accesses to the data to be stored, a new number of accesses is calculated based on the attenuation value and the number of accesses to the data to be stored, and the new number of accesses is used to update the number of accesses to the data to be stored in the node information of the data to be stored.
[0068] The preset threshold mentioned above can be set according to the actual situation, such as 500, and is not limited here. When the number of accesses is greater than or equal to 500, it means that the data to be stored is data with a very high access frequency, and the access count Freq in the node information of the data to be stored is directly incremented by 1;
[0069] Alternatively, calculate the decay value (decay, rounded down to 0 if less than 0) = (time - lastFileTime) / lastFileTime based on the current file read time and the last time stored in the node information of the data to be stored.
[0070] Each time Freq is updated, if decay is greater than Freq, Freq is directly assigned 1 and awaits replacement and migration; otherwise, Freq = Freq + 1 - decay, where Freq + 1 - decay is the new number of visits.
[0071] In one embodiment of this application, if the search result includes target node information, then retrieving the target data corresponding to the target node information from the first hard disk includes:
[0072] If the search result includes target node information, then the first storage address of the target data on the first hard disk is obtained from the target node information;
[0073] The target data is obtained from the first hard disk according to the first storage address.
[0074] In the above, a Kafka partition consists of several log segments. Each log segment contains two index files and a log message file, and the log segments are arranged in order of relative time (offset). The first storage address can refer to the sequence number of the log segment.
[0075] Using the methods described above, the target data can be quickly obtained from the first hard drive.
[0076] In one embodiment of this application, the step of acquiring the data to be stored and storing the data to be stored in the first hard disk includes:
[0077] Obtain the data to be stored;
[0078] If the storage capacity of the first hard disk is greater than or equal to the preset capacity, then the first linked list with the fewest accesses is obtained from the data access records;
[0079] Obtain the first node information in the first linked list, and obtain the first data stored on the first hard disk based on the first node information;
[0080] The first node information is deleted from the first linked list, and the first data is migrated to the second hard disk for storage;
[0081] The data to be stored is stored in the first hard disk, and the node information of the data to be stored is inserted into the second linked list of the data access records.
[0082] In the above, the first node information can be any node information in the first linked list. The first data can be obtained based on the first node information and migrated to the second hard disk for storage, so as to free up storage space on the first hard disk. This storage space can be used to store the data to be stored.
[0083] In the above, the data stored in the Kafka system is classified into hot and cold data through the first hard drive and the second hard drive, and the migration management of hot and cold data is carried out through the algorithm to ensure that the data with high access frequency is stored in the first hard drive, thereby improving the real-time processing performance of Kafka and reducing the Kafka read and write time and disk I / O load.
[0084] Figure 2a The diagram shown is a schematic flowchart of the data processing method provided in an embodiment of this application. Figure 2a As shown, the data processing method includes the following steps:
[0085] Step 1: Data Acquisition
[0086] The raw data originates from messages reported by the smart gateway. For example... Figure 2a As shown, gateway attributes and running status data are collected through the OSGi API, and message data is sent to Kafka through Logstash. The data is processed by Kafka's algorithm eviction mechanism and a hierarchical data storage architecture is constructed. Finally, the big data consumer reads the Kafka data for analysis and processing.
[0087] Step 2: Data Preprocessing
[0088] In Kafka, a partition consists of several log segments. Each log segment contains two index files and a log message file, and the log segments are arranged in order of offset (relative time).
[0089] Based on the principle of disk locality and data access characteristics, data that is frequently read and written within a certain period has a significantly higher probability of being accessed subsequently and should be stored on SSDs. Conversely, data with low read and write frequency has a lower probability of being accessed and should be stored on HDDs, thus achieving tiered data storage and separating hot and cold data. Therefore, the number of times data is accessed (i.e., read and write counts) is calculated; a higher count indicates more frequent read and write operations. Therefore, based on Kafka's file storage characteristics, statistics are collected on Kafka's logSegment file information and read / write times. Based on the statistical results, entity node (Node) information is constructed, such as... Figure 2b As shown, Offset represents the sequence number of the file content, Freq represents the number of accesses, and lastFileTime represents the last time the file was last modified.
[0090] Step 3: Constructing the data algorithm structure
[0091] First, create a HashMap collection named logFeq to store the node information. The key corresponds to the access count, and the value corresponds to a Node type data in a linked list structure (LinkList). <node>It can also refer to the identifier of node information, where the data in the linked list is sorted according to the number of times it is read and written, forming an ordered linked list data structure.
[0092] Then, a HashMap collection storage structure (storing nodes) named `logDevice` is constructed. The key (i.e., the first parameter value) corresponds to a key value (`topic_partition_offset`) composed of the Kafka topic name, partition number, and file offset. The value (i.e., the second parameter value) corresponds to Node type data, such as... Figure 2c As shown.
[0093] Step 4: Data processing algorithm for Kafka data processing flow
[0094] For the algorithm data structure established in step 3, the storage contents of logFeq and logDevice need to be initialized. The number of times logFeq is initialized according to the read and write time, and it is currently initialized to 1. The contents of the linked list structure are sorted according to the number of times they are the same but the time is different.
[0095] Tiered data storage is based on access frequency and file read / write time. If certain data has not been accessed for a long time, even if it was accessed frequently in the past, the access count (Freq) needs to be decayed when the file read / write interval becomes long, and the data is then evicted and migrated to the HDD. This ensures that hot and real-time data are stored on the SSD. Furthermore, Freq cannot increase indefinitely; it needs to be set with boundary conditions of 0-500. The calculation steps are as follows:
[0096] 1. When the access count is greater than or equal to 500, it indicates that the data is accessed with extremely high frequency, so directly increment Freq by 1;
[0097] 2. Calculate the decay value of Freq (decay, rounded down to 0 if less than 0) based on the current file read time (time) and the last time (lastFileTime) stored in Node: (time - lastFileTime) / lastFileTime;
[0098] 3. Each time Freq is updated, if decay is greater than Freq, Freq is directly assigned 1 and awaits replacement migration; otherwise, Freq = Freq + 1 - decay;
[0099] The specific algorithm execution flow is as follows:
[0100] 1. When an external consumer accesses Kafka data, it checks if the data exists using logDevice. If the data does not exist, it means there is no data on the SSD, so it is returned from the HDD. If the data exists, it returns the value of the corresponding node, retrieves the offset value from it, and then consumes the data. It also needs to update the access count Freq of the node in logFeq, and then remove the node from the original doubly linked list and add it to the doubly linked list corresponding to the new access count Freq.
[0101] 2. When Logstash sends data to Kafka, it is a write operation. The data is first written to PageCache, and then flushed to SSD after a certain period of time. At the same time, the logDevice and logFeq are updated with new data. The SSD storage capacity is checked. If the maximum capacity has not been reached, a new data node is inserted, and the access count Freq of the node is 1. If the maximum capacity has been exceeded, the node with the lowest access count needs to be deleted first, and then a new node is inserted, with the access count Freq of the node being 1. The data of these deleted nodes will be migrated to HDD for storage.
[0102] 3. Kafka data is stored in a tiered manner using SSDs and HDDs for hot and cold data, and migration management of hot and cold data is performed through algorithms to ensure that data with high access frequency is stored in SSDs, thereby improving Kafka's real-time processing performance and reducing Kafka read and write time and disk I / O load.
[0103] The data processing method provided in this application embodiment utilizes two map structures, logDevice and logFeq, and a linked list structure to achieve data eviction and migration, enabling hierarchical data storage and data access diversion. This ensures that real-time hot data is stored in SSD, while HDD data is not migrated back to SSD, avoiding data pollution and high-load IO, and improving throughput.
[0104] Using both SSDs and HDDs simultaneously enables tiered data storage, reducing storage costs. Furthermore, by leveraging the device relationships stored in logDevice, data storage locations can be quickly located, improving data access efficiency without incurring additional performance overhead.
[0105] By using SSDs and HDDs as the medium for tiered data storage, and with SSDs serving as a cache layer that connects to the pageCache layer, the real-time processing capability of Kafka is improved and storage costs are reduced. Real-time hot data is stored in SSDs, while HDD data is not migrated back to SSDs, thus avoiding data pollution and high-load I / O.
[0106] Figure 3 A structural diagram of the data processing apparatus provided in an embodiment of this application is shown. Figure 3 As shown, the data processing device 300 includes:
[0107] The first acquisition module 301 is used to acquire the data access address;
[0108] The lookup module 302 is used to look up the node information corresponding to the data access address from the pre-acquired data access record and obtain the lookup result. The data access record is used to record the node information of the data stored in the first hard disk.
[0109] The second acquisition module 303 is used to acquire target data corresponding to the target node information from the first hard disk if the search result includes target node information, wherein the target node information includes description information of the target data;
[0110] The third acquisition module 304 is used to acquire the target data from the second hard disk if the search result does not include the target node information, wherein the access frequency of the data stored on the first hard disk is higher than the access frequency of the data stored on the second hard disk.
[0111] In one embodiment of this application, the apparatus further includes:
[0112] The third acquisition module is used to acquire the data to be stored and store the data to be stored in the first hard disk.
[0113] The first construction module is used to construct the node information of the data to be stored. The node information includes the storage address of the data to be stored on the first hard disk, the number of times the data to be stored has been accessed, and the last modification time of the data to be stored.
[0114] The recording module is used to record the node information of the data to be stored in the data access record. The data access record includes multiple linked lists, each linked list corresponds to a number of accesses, and node information with the same number of accesses is located in the same linked list.
[0115] In one embodiment of this application, the apparatus further includes:
[0116] The second construction module is used to construct storage nodes in the storage structure based on the data to be stored. The storage node includes a first parameter value and a second parameter value. The first parameter value is determined based on the partition information where the data to be stored is located. The second parameter value is an identifier of the node information of the data to be stored. The storage structure includes multiple storage nodes, and each storage node has a different first parameter value and a different second parameter value.
[0117] The search module 302 includes:
[0118] The first acquisition submodule is used to determine the value of the third parameter based on the data access address;
[0119] The second acquisition submodule is used to search in the storage structure according to the third parameter value to obtain the target storage node, wherein the first parameter value of the target storage node is the same as the third parameter value;
[0120] The third acquisition submodule is used to acquire the second parameter value of the target storage node;
[0121] The search submodule is used to search the data access record according to the second parameter value and obtain the search result.
[0122] In one embodiment of this application, the search submodule includes:
[0123] The search unit is used to search in the data access record according to the second parameter value. If there is a linked list in the data access record that includes target node information and the identifier of the target node information is the same as the second parameter value, then the search result includes the target node information; otherwise, the search result does not include the target node information.
[0124] In one embodiment of this application, the apparatus further includes:
[0125] The update module is used to update the access count of the data to be stored based on one of the following:
[0126] If the number of accesses to the data to be stored is greater than or equal to a preset threshold, then the number of accesses to the data to be stored in the node information of the data to be stored is incremented by 1;
[0127] Alternatively, the decay value can be calculated based on the time of this reading of the data to be stored and the time of the most recent modification of the data to be stored.
[0128] If the attenuation value is greater than the number of times the data to be stored has been accessed, then the number of times the data to be stored has been accessed in the node information of the data to be stored is modified to 1.
[0129] If the attenuation value is less than or equal to the number of accesses to the data to be stored, a new number of accesses is calculated based on the attenuation value and the number of accesses to the data to be stored, and the new number of accesses is used to update the number of accesses to the data to be stored in the node information of the data to be stored.
[0130] In one embodiment of this application, the second acquisition module 303 includes:
[0131] The fourth acquisition submodule is used to obtain the first storage address of the target data on the first hard disk from the target node information if the search result includes target node information;
[0132] The fifth acquisition submodule is used to acquire the target data from the first hard disk according to the first storage address.
[0133] In one embodiment of this application, the third acquisition module includes:
[0134] The first processing submodule is used to acquire the data to be stored;
[0135] The second processing submodule is used to obtain the first linked list with the fewest accesses from the data access records if the storage capacity of the first hard disk is greater than or equal to the preset capacity.
[0136] The third processing submodule is used to obtain the first node information in the first linked list and obtain the first data stored on the first hard disk based on the first node information.
[0137] The fourth processing submodule is used to delete the first node information from the first linked list and migrate the first data to the second hard disk storage;
[0138] The fifth processing submodule is used to store the data to be stored in the first hard disk and insert the node information of the data to be stored into the second linked list of the data access record.
[0139] The data processing apparatus 300 provided in this application embodiment can implement the various processes implemented in the aforementioned data processing method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0140] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0141] The electronic device may include a processor 401 and a memory 402 storing computer program instructions.
[0142] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0143] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.
[0144] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to the first or second aspect of this disclosure.
[0145] The processor 401 implements any of the information auditing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 402.
[0146] In one example, the electronic device may also include a communication interface 403 and a bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0147] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0148] Bus 410 includes hardware, software, or both, that couples components of an information auditing method or verification device together. For example, and not as a limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0149] Furthermore, in conjunction with the data processing methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the data processing methods in the above embodiments.
[0150] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0151] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0152] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0153] Alternatively, this application embodiment can be implemented using a computer program product, wherein the instructions in the computer program product, when executed by the processor of an electronic device, cause the electronic device to implement any of the data processing methods in the above embodiments.
[0154] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0155] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0156] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0157] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0158] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.< / node>
Claims
1. A data processing method, characterized in that, The method includes: Get the data access address; The node information corresponding to the data access address is searched from the pre-acquired data access records to obtain the search result. The data access records are used to record the node information of the data stored in the first hard disk. If the search result includes target node information, then the target data corresponding to the target node information is obtained from the first hard disk, wherein the target node information includes the description information of the target data; If the search result does not include the target node information, the target data is obtained from the second hard disk, wherein the access frequency of the data stored on the first hard disk is higher than the access frequency of the data stored on the second hard disk.
2. The method according to claim 1, characterized in that, Before searching for the node information corresponding to the data access address from the pre-acquired data access records, the method further includes: Acquire the data to be stored and store the data to be stored in the first hard disk; Construct node information for the data to be stored, including the storage address of the data to be stored on the first hard disk, the number of times the data to be stored has been accessed, and the time of the most recent modification of the data to be stored. The node information of the data to be stored is recorded in the data access record. The data access record includes multiple linked lists, each linked list corresponds to a number of accesses, and node information with the same number of accesses is located in the same linked list.
3. The method according to claim 2, characterized in that, Before searching for the node information corresponding to the data access address from the pre-acquired data access records, the method further includes: Based on the data to be stored, storage nodes in the storage structure are constructed. Each storage node includes a first parameter value and a second parameter value. The first parameter value is determined based on the partition information where the data to be stored is located. The second parameter value is an identifier of the node information of the data to be stored. The storage structure includes multiple storage nodes, and each storage node has different first parameter values and second parameter values. The step of searching for the node information corresponding to the data access address from the pre-acquired data access records to obtain the search result includes: The value of the third parameter is determined based on the data access address; The target storage node is obtained by searching the storage structure according to the third parameter value, and the first parameter value of the target storage node is the same as the third parameter value. Obtain the second parameter value of the target storage node; The search is performed in the data access record based on the second parameter value to obtain the search result.
4. The method according to claim 3, characterized in that, The step of searching the data access record according to the second parameter value to obtain the search result includes: The search is performed in the data access record according to the second parameter value. If there is a linked list in the data access record that includes target node information and the identifier of the target node information is the same as the second parameter value, then the search result includes the target node information; otherwise, the search result does not include the target node information.
5. The method according to claim 2, characterized in that, The method further includes: The access count of the data to be stored is updated based on one of the following: If the number of accesses to the data to be stored is greater than or equal to a preset threshold, then the number of accesses to the data to be stored in the node information of the data to be stored is incremented by 1; Alternatively, the decay value can be calculated based on the time of this reading of the data to be stored and the time of the most recent modification of the data to be stored. If the attenuation value is greater than the number of times the data to be stored has been accessed, then the number of times the data to be stored has been accessed in the node information of the data to be stored is modified to 1. If the attenuation value is less than or equal to the number of accesses to the data to be stored, a new number of accesses is calculated based on the attenuation value and the number of accesses to the data to be stored, and the new number of accesses is used to update the number of accesses to the data to be stored in the node information of the data to be stored.
6. The method according to claim 2, characterized in that, If the search result includes target node information, then the target data corresponding to the target node information is obtained from the first hard disk, including: If the search result includes target node information, then the first storage address of the target data on the first hard disk is obtained from the target node information; The target data is obtained from the first hard disk according to the first storage address.
7. The method according to claim 2, characterized in that, The step of acquiring the data to be stored and storing the data to be stored in the first hard disk includes: Obtain the data to be stored; If the storage capacity of the first hard disk is greater than or equal to the preset capacity, then the first linked list with the fewest accesses is obtained from the data access records; Obtain the first node information in the first linked list, and obtain the first data stored on the first hard disk based on the first node information; The first node information is deleted from the first linked list, and the first data is migrated to the second hard disk for storage; The data to be stored is stored in the first hard disk, and the node information of the data to be stored is inserted into the second linked list of the data access records.
8. A data processing apparatus, characterized in that, The device includes: The first acquisition module is used to acquire the data access address; The lookup module is used to search for the node information corresponding to the data access address from the pre-acquired data access records and obtain the search result. The data access records are used to record the node information of the data stored in the first hard disk. The second acquisition module is used to acquire target data corresponding to the target node information from the first hard disk if the search result includes target node information, wherein the target node information includes descriptive information of the target data; The third acquisition module is used to acquire the target data from the second hard disk if the search result does not include the target node information, wherein the access frequency of the data stored on the first hard disk is higher than the access frequency of the data stored on the second hard disk.
9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.
11. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method and system for reducing garbage collection and write amplification of key-value separation storage system
CN112395212A
Data synchronization method and device
CN114860826A