Data Processing Method and Electronic Device

By selecting offline or online merging methods based on the service type of data in distributed storage, the problems of low efficiency of merging and write performance impacts are solved, and efficient data merging and performance guarantees are achieved.

CN113946552BActive Publication Date: 2025-07-18BEIJING XSKY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111222136.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2025-07-18
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

In distributed storage scenarios, the merger of small files is low, and the offline merge method affects the writing performance of business data, while the online merge method affects the writing performance of business data.

Method used

By determining the service type of the target data during the writing process, the offline merge method is used to merge data with periodic fluctuations in the data volume, and the online merge method is used to merge data with balanced data volume.

Benefits of technology

It improves the efficiency of merging small files, ensures write performance, and improves the performance and space utilization of storage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113946552B_ABST
    Figure CN113946552B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method and an electronic device. The method includes: during the process of writing target data into a storage system, determining the service type corresponding to the target data, where the target data is data that occupies less storage space than a preset storage space, and the service types include a first type and a second type. The data writing fluctuation situation corresponding to the first type meets a preset fluctuation condition, and the data writing fluctuation situation corresponding to the first type does not meet the preset fluctuation condition; when the service type corresponding to the target data is the first type, performing data merging on the target data by using an offline merging method; when the service type corresponding to the target data is the second type, performing data merging on the target data by using an online merging method. The problem that in the related art, when merging small files in a distributed storage scenario, the merging efficiency of the offline merging method is relatively low, while the online merging method affects the writing performance of service data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular, to a data processing method and an electronic device. Background Art

[0002] With the advent of the big data era, the explosive growth of data poses a severe challenge to traditional storage. Distributed storage emerges in response to the "cloud" and can provide storage services with large capacity, high reliability, high scalability, and decentralization. However, in the scenario of massive storage, distributed storage still faces many problems. For example, in the process of unstructured data storage, the waste of storage space for small files is an important problem faced.

[0003] To solve the problem of waste of storage space for small files, in related technologies, the offline merging method is used to merge small objects in a single cluster. Specifically, a log file is recorded during the process of uploading small objects, and then the background task matches the objects by scanning the characteristic logs to perform file merging. However, this method has the following problems: 1. During the merging process, it is necessary to first scan the corresponding log objects and then read the data, resulting in a read penalty. At the same time, since the merging is performed offline, in the scenario of frequent user writes, the merging tasks will accumulate, resulting in a decrease in the write performance of subsequent business data or even business interruption; 2. During the merging process, small files are merged in units of clusters. If the small objects in the merged large object are deleted, it will cause waste of storage space for the merged large object; 3. There are certain differences in the access degrees of different files within the cluster, and the read and write operations of small files are inconsistent.

[0004] To solve the problems existing in offline merging, online merging technology has also emerged in related technologies. Although online merging technology can avoid write penalties through real-time merging, in the scenario where the user's business scenario is periodic writing and there is a certain peak writing scenario (a scenario where the user requires maximized performance writing), real-time merging will occupy the write bandwidth and affect the business performance.

[0005] In view of the problems that in the related technologies, when merging small files in the distributed storage scenario, the merging efficiency of the offline merging method is relatively low, and the online merging method affects the write performance of business data, no effective solution has been proposed yet. Summary of the Invention

[0006] The present application provides a data processing method and an electronic device to solve the problems that in the related technologies, when merging small files in the distributed storage scenario, the merging efficiency of the offline merging method is relatively low, and the online merging method affects the write performance of business data.

[0007] According to one aspect of the present application, a data processing method is provided. The method includes: during the process of writing target data into a storage system, determining the service type corresponding to the target data, where the target data is data that occupies less storage space than a preset storage space, the service type includes a first type and a second type, the data writing fluctuation condition corresponding to the first type meets a preset fluctuation condition, and the data writing fluctuation condition corresponding to the second type does not meet the preset fluctuation condition; when the service type corresponding to the target data is the first type, performing data merging on the target data in an offline merging manner; when the service type corresponding to the target data is the second type, performing data merging on the target data in an online merging manner.

[0008] Optionally, when the service type corresponding to the target data is the second type, performing data merging on the target data in an online merging manner includes: sequentially writing each target data into a first storage medium in the storage system, recording log data, and generating metadata corresponding to the target data; when each target data is written into the first storage medium in the storage system each time, adding the target data to a merging module, and sending a merging task through the merging module to store the target data in a second storage medium in the storage system until the set condition of the merging task is met to obtain the merged data, where the set condition is used to set the quantity or the occupied storage space size of the target data corresponding to the merged data; updating the metadata corresponding to each target data in the merged data.

[0009] Optionally, sending a merging task through the merging module to store the target data in a second storage medium in the storage system includes: generating a plurality of merging tasks, and adding the target data to one of the plurality of merging tasks according to a preset rule; controlling the plurality of merging tasks to concurrently write data into the second storage medium.

[0010] Optionally, controlling the plurality of merging tasks to concurrently write data into the second storage medium includes: when each merging task writes data into the second storage medium for the first time, adjusting the state corresponding to the merging task from an initial state to an execution state; when each merging task writes all the data into the second storage medium, adjusting the state corresponding to the merging task from the execution state to a completion state, and updating the metadata of the data corresponding to the merging task.

[0011] Optionally, the method further includes: when an exception occurs in the process corresponding to the merging task, re-executing the uncompleted merging task.

[0012] Optionally, when the business type corresponding to the target data is the first type, data merging of the target data by using the offline merging method includes: sequentially writing each target data into a first storage medium in the storage system, recording log data, and generating metadata corresponding to the target data; after a preset merging time is reached, obtaining storage locations of multiple target data from the log data, and generating a merging task, where the number of the multiple target data is a preset number, the storage location is a storage address in the first storage medium, and the merging task is used to merge the data stored at the corresponding storage location; obtaining multiple target data from the first storage medium according to the storage location corresponding to the merging task, and merging the multiple target data to obtain a set of merged data; storing the merged data into a second storage medium in the storage system, deleting the multiple target data from the first storage medium, and updating the metadata corresponding to the multiple target data.

[0013] Optionally, obtaining multiple target data from the first storage medium includes: determining whether the writing rate of the target data into the first storage medium is greater than a preset rate; when the writing rate of the target data into the first storage medium is greater than the preset rate, obtaining multiple target data from the first storage medium at a first speed; when the writing rate of the target data into the first storage medium is less than or equal to the preset rate, obtaining multiple target data from the first storage medium at a second speed, where the second speed is greater than the first speed.

[0014] Optionally, after storing the merged data into the second storage medium in the storage system, the method further includes: calculating a data missing ratio when there is missing data in a set of merged data; when the data missing ratio is greater than a preset ratio, reading the non-missing data in the merged data from the second storage medium, and reading a target number of target data from the first storage medium, where the target number is the number of the missing data; merging the non-missing data and the target number of target data to obtain re-merged data; deleting the non-missing data from the second storage medium, storing the re-merged data into the second storage medium, deleting the target number of target data from the first storage medium, and updating the metadata corresponding to each target data in the re-merged data.

[0015] Optionally, during the process of writing target data to a storage system, determining the service type corresponding to the target data includes: within a preset time period, obtaining the number of write operations of the target data at preset time intervals to obtain multiple write operation count values; determining the distribution of the multiple write operation count values within the preset time period, and determining whether the data write fluctuation condition of the target data meets a preset fluctuation condition according to the distribution; in the case where the data write fluctuation condition of the target data meets the preset fluctuation condition, determining that the service type corresponding to the target data is the first type; in the case where the data write fluctuation condition of the target data does not meet the preset fluctuation condition, determining that the service type corresponding to the target data is the second type.

[0016] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including a processor and a memory; computer-readable instructions are stored in the memory, and the processor is used to run the computer-readable instructions, wherein when the computer-readable instructions run, a data processing method is executed.

[0017] Through the present application, the following steps are adopted: during the process of writing target data to a storage system, determining the service type corresponding to the target data, wherein the target data is data whose occupied storage space is less than a preset storage space, and the service types include a first type and a second type, the data write fluctuation condition corresponding to the first type meets the preset fluctuation condition, and the data write fluctuation condition corresponding to the second type does not meet the preset fluctuation condition; in the case where the service type corresponding to the target data is the first type, performing data merging on the target data in an offline merging manner; in the case where the service type corresponding to the target data is the second type, performing data merging on the target data in an online merging manner. The problem that in the related art, when merging small files in a distributed storage scenario, the merging efficiency of the offline merging method is relatively low, and the online merging method affects the write performance of service data is solved. By determining the service type corresponding to the small file and selecting a corresponding method for data merging according to the service type, the effect of improving the merging efficiency of small files while ensuring the write performance of small files is achieved, and the storage scenario performance and space utilization rate are improved simultaneously. Description of the Drawings

[0018] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0019] Figure 1 is a flowchart of the data processing method provided by the embodiments of this application.

[0020] Figure 2 is a flowchart of an optional service type determination provided by the embodiments of this application.

[0021] Figure 3 It is a flowchart of an optional online merging method provided according to an embodiment of the present application.

[0022] Figure 4 It is a process of an optional offline merging method provided according to an embodiment of the present application.

[0023] Figure 5 It is a schematic diagram of a data processing device provided according to an embodiment of the present application. Detailed implementation manners

[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] In order to solve the problem that in the related art, when merging small files in a distributed storage scenario, the merging efficiency of the offline merging method is relatively low, while the online merging method affects the write performance of service data. Based on this, the present application hopes to provide a solution that can solve the above technical problems, and its detailed content will be elaborated in the subsequent embodiments.

[0028] According to an embodiment of the present application, a data processing method is provided.

[0029] Figure 1 It is a flowchart of the data processing method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0030] Step S101, during the process of writing target data into a storage system, determine the service type corresponding to the target data, where the target data is data that occupies less storage space than a preset storage space, and the service type includes a first type and a second type. The data writing fluctuation condition corresponding to the first type meets a preset fluctuation condition, and the data writing fluctuation condition corresponding to the first type does not meet the preset fluctuation condition.

[0031] Specifically, the preset storage space can be 1 kb, and the target data can be small files with a file size below 1 kb.

[0032] It should be noted that during the operation of the service, target data to be stored is continuously generated. The data writing fluctuation condition can be characterized by the difference in the number of multiple target data to be stored obtained at the same time interval within a cycle. The preset fluctuation condition can be the difference in the number of multiple target data to be stored corresponding to the same time interval within a cycle that has been set, and is used to compare with the actually obtained difference in quantity.

[0033] Furthermore, when the actual fluctuation condition of the target data to be stored and the preset fluctuation condition do not differ much, it can be considered that the data writing fluctuation condition meets the preset fluctuation condition, and the service type corresponding to the target data to be stored is a service type with balanced data volume, that is, the first type. When the actual fluctuation condition of the target data to be stored and the preset fluctuation condition differ too much, it can be considered that the data writing fluctuation condition does not meet the preset fluctuation condition, and the service type corresponding to the target data to be stored is a service type with periodic data volume fluctuations, that is, the second type.

[0034] Step S102, when the service type corresponding to the target data is the first type, perform data merging on the target data in an offline merging manner.

[0035] It should be noted that when the service type corresponding to the target data is the first type, it indicates that the data volume of the target data fluctuates periodically. In order to avoid the impact of the merging of the target data on the writing performance of the target data, the target data can be merged in an offline merging manner.

[0036] Specifically, the specific steps of the offline merging method are to record a log file during the process of uploading small target data, and then the background merging task matches the target data by scanning the log file and performs file merging. It should be noted that when writing too much target data, the user can customize the merging time, so as to stagger the execution of the writing and merging of the target data.

[0037] Step S103: When the service type corresponding to the target data is the second type, perform data merging on the target data in an online merging manner.

[0038] It should be noted that when the service type corresponding to the target data is the second type, it indicates that the data volume of the target data is balanced. To improve the data merging efficiency while not affecting the write performance of the target data, the target data can be merged in an online merging manner.

[0039] Specifically, the online merging method can achieve real-time merging, that is, merging while writing the target data. On the one hand, it avoids the read penalty during offline merging. On the other hand, it writes in an append-write manner, avoiding the write penalty of the method of writing after merging, thereby improving the merging efficiency.

[0040] The data processing method provided by the embodiments of the present application includes the following steps: During the process of writing target data into the storage system, determine the service type corresponding to the target data, where the target data is data whose occupied storage space is less than the preset storage space, and the service types include the first type and the second type. The data write fluctuation situation corresponding to the first type satisfies the preset fluctuation condition, and the data write fluctuation situation corresponding to the first type does not satisfy the preset fluctuation condition; when the service type corresponding to the target data is the first type, perform data merging on the target data in an offline merging manner; when the service type corresponding to the target data is the second type, perform data merging on the target data in an online merging manner. This solves the problem that in the related art, when merging small files in a distributed storage scenario, the merging efficiency of the offline merging method is relatively low, while the online merging method affects the write performance of service data. By determining the service type corresponding to the small file and selecting the corresponding method for data merging according to the service type, the effect of improving the merging efficiency of the small file while ensuring the write performance of the small file is achieved.

[0041] Optionally, in the data processing method provided by the embodiments of the present application, determining the service type corresponding to the target data during the process of writing the target data into the storage system includes: within a preset time period, obtain the number of write operations of the target data at preset time intervals to obtain multiple write operation count values; determine the distribution of the multiple write operation count values within the preset time period, and determine whether the data write fluctuation situation of the target data satisfies the preset fluctuation condition according to the distribution; when the data write fluctuation situation of the target data satisfies the preset fluctuation condition, determine that the service type corresponding to the target data is the first type; when the data write fluctuation situation of the target data does not satisfy the preset fluctuation condition, determine that the service type corresponding to the target data is the second type.

[0042] Specifically, the preset time period can be the total time for obtaining the number of write operations of the set target data. First, the distribution of multiple write operation count values obtained at preset time intervals within the preset time period can be determined, and this distribution can be used as the data write fluctuation condition of the target data within the preset time period. Then, the data write fluctuation condition of the target data is compared with the preset fluctuation condition, and the business type is determined based on the comparison result.

[0043] The average value and peak value of the write operation count values of the target data can characterize the distribution of the write operation count values. Optionally, in the data processing method provided in the embodiments of the present application, determining the distribution of multiple write operation count values within the preset time period and determining whether the data write fluctuation condition of the target data meets the preset fluctuation condition includes: calculating the average value and peak value of the write operation count values of the target data, and determining the target value based on the average value and the preset weight; in the case where the peak value is greater than the target value, determining that the data write fluctuation condition of the target data meets the preset fluctuation condition; in the case where the peak value is less than or equal to the target value, determining that the data write fluctuation condition of the target data does not meet the preset fluctuation condition.

[0044] Specifically, the peak value of the write operation count values of the target data can be the value at the 80% of the distribution statistics of the read and write count values obtained within the preset time period. The target value is determined by the average value of the write operation count values of the target data and the preset weight. Among them, the average value can be the value at the 50% of the distribution statistics of the read and write count values obtained within the preset time period. The preset weight can be set flexibly. For example, it can be set to 2, and whether the data write fluctuation condition of the target data meets the preset fluctuation condition is determined based on the comparison result between the peak value and twice the average value.

[0045] The following is an optional embodiment for determining the business type corresponding to the target data. Figure 2 It is a flowchart for business type determination, as Figure 2As shown, the preset time period T1 can be = 1 day, the preset time interval t1 can be = 5 minutes, and there are 288 time intervals in 1 day. The write operation count value of small files is statistically counted every 5 minutes. After 288 time intervals, the statistics for the entire cycle are completed. At this time, the average value m1 and the peak value m2 of the statistical distribution within this cycle are calculated. The average value m1 can be the read / write promotion corresponding to the 50% position of the distribution statistics, and the peak value m2 is the read / write count corresponding to the 80% position of the distribution statistics. The preset condition can be m2 > P * m, where P is the preset weight and can be set to 2. At this time, it is judged whether the size relationship between m2 and m1 * P satisfies m2 > P * m. When m2 > P * m1 is satisfied, it can be considered that the preset fluctuation condition is met, and the service type is determined to be the first type. When m2 > P * m1 is not satisfied, for example, when m2 < P * m1 or m2 = P * m1, it can be considered that the preset fluctuation condition is not met, and the service type is determined to be the second type.

[0046] Optionally, in the data processing method provided in the embodiments of the present application, when the service type corresponding to the target data is the second type, performing data merging on the target data by using an online merging method includes: sequentially writing each target data into a first storage medium in the storage system, recording log data, and generating metadata corresponding to the target data; when each target data is written into the first storage medium in the storage system each time, adding the target data to a merging module, and storing the target data into a second storage medium in the storage system through the merging module issuing a merging task until the set condition of the merging task is met, obtaining the merged data, where the set condition is used to set the number of target data corresponding to the merged data or the size of the storage space occupied; updating the metadata corresponding to each target data in the merged data.

[0047] Specifically, different from offline merging, in online merging, after the target data is written into the first storage medium in the storage system and the log data is recorded, the target data can be directly added to the merging task through the merging module, so that the target data is stored into the second storage medium in the storage system according to the merging task. The merging tasks can include multiple ones, and multiple merging tasks form a merging task list. The merging of the target data is controlled through the merging tasks, and at the same time, it is judged whether the merging process meets the set condition. After the set condition is met, the merging is completed, and the metadata corresponding to each target data in the merged data is updated. The set condition can be the condition set in a single merging task. For example, merging 16,000 target data or the size of the merged file reaches 64M. The metadata is updated from the storage address in the first storage medium to the storage address in the second storage medium.

[0048] Optionally, in the data processing method provided in the embodiments of the present application, storing the target data in the second storage medium in the storage system by the merging module through issuing a merging task includes: generating a plurality of merging tasks, and adding the target data to one of the plurality of merging tasks according to a preset rule; controlling the plurality of merging tasks to concurrently write data into the second storage medium.

[0049] Specifically, there may be a plurality of merging tasks, and the plurality of merging tasks form a merging task list. Each merging task in the merging task list is executed concurrently. According to the hash algorithm, the currently obtained target data is added to one of the plurality of merging tasks, achieving the effect of multiple merging tasks being carried out simultaneously and improving the efficiency of data merging.

[0050] Optionally, in the data processing method provided in the embodiments of the present application, controlling the plurality of merging tasks to concurrently write data into the second storage medium includes: when each merging task first writes data into the second storage medium, adjusting the state corresponding to the merging task from the initial state to the execution state; when each merging task writes all data into the second storage medium, adjusting the state corresponding to the merging task from the execution state to the completion state, and updating the metadata of the data corresponding to the merging task.

[0051] Specifically, when the previous merging task is completed or the merging module is initialized, a plurality of concurrent control tasks are created. It is in the initialization state when just created. After the first small object in each merging task is appended and written, it is updated to the execution state. Finally, when the data threshold set for the merging task is met, the metadata information of the object corresponding to the merging task is updated. It should be noted that during the merging process, if a small object changes, such as being deleted or the object metadata changes, the metadata of the merged large object will mark the small object as the deleted state, and a new small task will be generated after the update is completed.

[0052] Optionally, in the data processing method provided in the embodiments of the present application, the method further includes: when an exception occurs in the process corresponding to the merging task, re-executing the uncompleted merging task.

[0053] Specifically, when an exception occurs in one of the merging tasks, the target data that has been stored in the second storage medium is deleted, the target data that has been stored in the second storage medium is read from the first cache pool, stored in the first cache pool, and the target data in the merging task that has not completed the merging is continued to be merged.

[0054] The following is an optional online merging method provided according to the embodiments of the present application. Figure 3 It is a flowchart of an optional online merging method provided according to the embodiments of the present application, as Figure 3 shown:

[0055] The first storage medium is a cache pool, the second storage medium is a data pool, and the preset rule is to merge 16,000 target data into a large file.

[0056] Write the target data into the cache pool and generate log data. The log data includes information such as the time and location when the target data is written. At the same time, directly issue a merge task and add the target data to a merge task in the merge task list, so as to store the target data in the data pool.

[0057] For example, if there are 2 concurrent merge tasks in the current merge list, the merge control module can put the currently obtained target data into merge task 1. At this time, write the target data to the first target storage address in the data pool. When merge task 1 receives the next target data, append and write the target data to the storage address after the previous target data until 160,000 target data corresponding to merge task 1 are all written into the data pool, and a merged large file 1 is obtained. At this time, merge task 1 is completed, and the metadata of the 16,000 target data is updated. In addition, it should be noted that if an exception occurs when merging a target data in the merge task, the merge task is re-executed. For example, if the 10th target data in merge task 1 has an exception during the merge, delete the 10 target data in the data pool, read the 10 target data from the cache pool, and perform the merge task on the remaining 15,990 target data in merge task 1.

[0058] Optionally, in the data processing method provided in the embodiments of the present application, when the service type corresponding to the target data is the first type, using the offline merge method to perform data merge on the target data includes: sequentially writing each target data into the first storage medium in the storage system, recording log data, and generating metadata corresponding to the target data; after the preset merge time arrives, obtain the storage locations of multiple target data from the log data and generate a merge task, where the number of multiple target data is a preset number, the storage location is the storage address in the first storage medium, and the merge task is used to merge the data stored at the corresponding storage location; obtain multiple target data from the first storage medium according to the storage location corresponding to the merge task, and perform merge on the multiple target data to obtain a set of merged data; store the merged data in the second storage medium in the storage system, delete the multiple target data from the first storage medium, and update the metadata corresponding to the multiple target data.

[0059] Specifically, the first storage medium can be a hardware storage medium with better storage performance, such as a solid-state drive. The target data is stored sequentially according to the smallest storage unit of the first storage medium. Specifically, when the storage space occupied by the target data is less than one smallest storage unit, it also occupies one smallest storage unit.

[0060] While writing target data to the first storage medium, log data is recorded. The log data includes information such as the time when the target data is written to the first storage medium and the location where it is written to the first storage medium. Metadata of the target data can be generated based on the log data. The metadata corresponding to the target data can include the current storage location of the target data. After the target data is written to the first storage medium, the current storage location of the target data is the storage location of the target data in the first storage medium.

[0061] Among them, the preset merging time can be the merging execution time defined by the user. For example, when the business type corresponding to the target data has fluctuating data volumes during day and night, a large amount of target data to be stored is generated during the day, and a small amount of target data to be stored is generated at night. The preset merging time can be set to night; the preset quantity can be the number of target data to be merged that is set in advance. For example, it can be 100; the second storage medium can be a hardware storage medium with a large storage space, such as a mechanical hard disk, which is convenient for storing the merged large file.

[0062] Specifically, after the preset merging time arrives, scan the storage locations of 100 target data in the first storage medium from the log, read the 100 target data from the first storage medium into the memory, and generate a merging task. Scan and execute the merging task in the memory to achieve the merging of the 100 target data. Further, store the merged data in the second storage medium, update the metadata corresponding to the target data, that is, update the current storage location of the target data to the storage location in the second storage medium, and delete the 100 target data from the first storage medium to provide write space for the subsequent writing of target data.

[0063] To avoid occupying too many read and write resources of the first storage medium during data merging and blocking the process of writing target data to the first storage medium, optionally, in the data processing method provided in the embodiments of the present application, obtaining multiple target data from the first storage medium includes: judging whether the writing rate of the target data to the first storage medium is greater than a preset rate; when the writing rate of the target data to the first storage medium is greater than the preset rate, obtaining multiple target data from the first storage medium at a first speed; when the writing rate of the target data to the first storage medium is less than or equal to the preset rate, obtaining multiple target data from the first storage medium at a second speed, where the second speed is greater than the first speed.

[0064] Specifically, the preset rate is the rate at which data is normally written to the first storage medium as pre-set. The preset rate can be determined by counting the quantity and time of the target data being normally written to the first storage medium through QOS. Further, after obtaining the preset rate, the speed of obtaining multiple target data from the first storage medium is controlled according to the preset rate, achieving the normal writing of the target data while performing data reduction.

[0065] To prevent waste of storage space in the second storage medium caused by deletion of target data in the merged data, optionally, in the data processing method provided in the embodiments of the present application, after storing the merged data in the second storage medium in the storage system, the method further includes: calculating the data loss ratio in the case of data loss in a set of merged data; in the case where the data loss ratio is greater than the preset ratio, reading the non-lost data in the merged data from the second storage medium, and reading a target quantity of target data from the first storage medium, where the target quantity is the quantity of the lost data; merging the non-lost data and the target quantity of target data to obtain re-merged data; deleting the non-lost data from the second storage medium, storing the re-merged data in the second storage medium, deleting the target quantity of target data from the first storage medium, and updating the metadata corresponding to each target data in the re-merged data.

[0066] Specifically, the preset ratio is the data loss ratio in the merged data as pre-set. When the actual data loss ratio is greater than the preset ratio, the non-lost data in the merged data is read into the memory, and a quantity of target data equal to the lost data is read from the first storage medium. The non-lost data and the re-read target data are re-merged, the re-merged data is stored in the second storage medium, and at the same time, the relevant target data in the first storage medium and the second storage medium are deleted, and the metadata is updated.

[0067] The following is an optional offline merging method provided according to the embodiments of the present application. Figure 4 The flowchart of an optional offline merging method provided according to the embodiments of the present application is as Figure 4 shown:

[0068] The first storage medium is a cache pool, the second storage medium is a data pool, and the preset quantity is 100 target data. When data is written, the data is stored in the cache pool, and information such as the storage time and location is recorded in the log. After the preset merging time is reached, the QOS is used to statistically calculate the rate of writing the target data into the cache pool. When the rate of writing the target data into the first storage medium is less than or equal to the preset rate, the storage locations of the 100 target data in the cache pool are obtained from the log data, and the 100 target data are obtained from the cache pool at a relatively fast speed. The obtained target data is merged to obtain a set of merged data, and the merged data is stored in the data pool in the storage system. Further, the 100 target data are deleted from the cache pool, and the metadata corresponding to the 100 target data is updated, that is, the storage locations of the 100 target data are updated from the storage addresses in the first storage medium to the storage addresses in the second storage medium.

[0069] When there is data missing in a set of merged data, calculate the data missing ratio. It can be set that the preset ratio is 50%. When the first to 51st target data are all missing, the data missing ratio > 50%, then the remaining target data are triggered to join the merging task. Specifically, the remaining 49 target data of the merged data are read from the data pool into the memory, and 51 target data are read from the cache pool. The remaining 49 target data of the merged data are merged with the obtained 51 target data to obtain the secondarily merged data, and the secondarily merged data is stored in the data pool, and the corresponding metadata is updated at the same time.

[0070] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0071] The embodiment of the present application also provides a data processing device. It should be noted that the data processing device in the embodiment of the present application can be used to execute the data processing method provided by the embodiment of the present application. The data processing device provided by the embodiment of the present application is introduced below.

[0072] Figure 5 is a schematic diagram of the data processing device according to the embodiment of the present application. As Figure 5 shown, the device includes: a first determination unit 10, a first execution unit 20, and a second execution unit 30.

[0073] The first determination unit 10 is configured to determine the service type corresponding to the target data during the process of writing the target data into the storage system, where the target data is data that occupies a storage space smaller than a preset storage space, and the service types include a first type and a second type. The data writing fluctuation condition corresponding to the first type satisfies a preset fluctuation condition, and the data writing fluctuation condition corresponding to the first type does not satisfy the preset fluctuation condition.

[0074] The first execution module 20 is configured to perform data merging on the target data in an offline merging manner when the service type corresponding to the target data is the first type.

[0075] The second execution module 30 is configured to perform data merging on the target data in an online merging manner when the service type corresponding to the target data is the second type.

[0076] In the data processing device provided by the embodiment of the present application, the first determination unit 10 determines the service type corresponding to the target data during the process of writing the target data into the storage system, where the target data is data that occupies a storage space smaller than a preset storage space, and the service types include a first type and a second type. The data writing fluctuation condition corresponding to the first type satisfies a preset fluctuation condition, and the data writing fluctuation condition corresponding to the first type does not satisfy the preset fluctuation condition; the first execution module 20 performs data merging on the target data in an offline merging manner when the service type corresponding to the target data is the first type; the second execution module 30 performs data merging on the target data in an online merging manner when the service type corresponding to the target data is the second type. This solves the problem that in the related art, when merging small files in a distributed storage scenario, the merging efficiency of the offline merging method is relatively low, while the online merging method affects the writing performance of service data. By determining the service type corresponding to the small file and selecting the corresponding method for data merging according to the service type, the effect of improving the merging efficiency of small files while ensuring the writing performance of small files is achieved, and the performance and space utilization rate of the storage scenario are improved simultaneously.

[0077] Optionally, in the data processing apparatus provided in the embodiments of the present application, the first determination unit 10 includes: a first acquisition module, configured to obtain the number of write operations of target data at preset time intervals within a preset time period, so as to obtain a plurality of write operation count values; a first determination module, configured to determine the distribution of the plurality of write operation count values within the preset time period, and determine whether the data write fluctuation condition of the target data meets a preset fluctuation condition according to the distribution; a second determination module, configured to determine that the service type corresponding to the target data is a first type when the data write fluctuation condition of the target data meets the preset fluctuation condition; and a third determination module, configured to determine that the service type corresponding to the target data is a second type when the data write fluctuation condition of the target data does not meet the preset fluctuation condition.

[0078] Optionally, in the data processing apparatus provided in the embodiments of the present application, the second execution module 30 includes: a first execution sub-module, configured to sequentially write each target data into a first storage medium in the storage system, record log data, and generate metadata corresponding to the target data; a second execution sub-module, configured to add the target data to a merging module each time the target data is written into the first storage medium in the storage system, and issue a merging task through the merging module to store the target data in a second storage medium in the storage system until the set condition of the merging task is met, so as to obtain merged data, where the set condition is used to set the number or storage space size of the target data corresponding to the merged data; and a third execution sub-module, configured to update the metadata corresponding to each target data in the merged data.

[0079] Optionally, in the data processing apparatus provided in the embodiments of the present application, the second execution module 30 further includes: a generation sub-module, configured to generate a plurality of merging tasks, and add the target data to one of the plurality of merging tasks according to a preset rule; and a writing sub-module, configured to control the plurality of merging tasks to concurrently write data into the second storage medium.

[0080] Optionally, in the data processing apparatus provided in the embodiments of the present application, the writing sub-module includes: a first adjustment sub-module, configured to adjust the state of the merging task corresponding to the initial state to an execution state when each merging task first writes data into the second storage medium; and a second adjustment sub-module, configured to adjust the state of the merging task corresponding to the execution state to a completion state when each merging task writes all the data into the second storage medium, and update the metadata of the data corresponding to the merging task.

[0081] Optionally, in the data processing apparatus provided in the embodiments of the present application, the second execution module 30 further includes: a fourth execution sub-module, configured to re-execute the uncompleted merging task when an exception occurs in the process corresponding to the merging task.

[0082] Optionally, in the data processing device provided in the embodiments of the present application, the first execution module 20 includes: a fifth execution sub-module, configured to sequentially write each target data into a first storage medium in the storage system, record log data, and generate metadata corresponding to the target data; a sixth execution sub-module, configured to, after a preset merging time is reached, obtain storage locations of multiple target data from the log data, and generate a merging task, where the number of the multiple target data is a preset number, the storage location is a storage address in the first storage medium, and the merging task is used to merge the data stored at the corresponding storage locations; a seventh execution sub-module, configured to obtain multiple target data from the first storage medium according to the storage locations corresponding to the merging task, and merge the multiple target data to obtain a set of merged data; a storage sub-module, configured to store the merged data in a second storage medium in the storage system, delete the multiple target data from the first storage medium, and update the metadata corresponding to the multiple target data.

[0083] Optionally, in the data processing device provided in the embodiments of the present application, the first execution module 20 further includes: a judgment sub-module, configured to judge whether the rate of writing the target data into the first storage medium is greater than a preset rate; an eighth execution sub-module, configured to, when the rate of writing the target data into the first storage medium is greater than the preset rate, obtain multiple target data from the first storage medium at a first speed; a ninth execution sub-module, configured to, when the rate of writing the target data into the first storage medium is less than or equal to the preset rate, obtain multiple target data from the first storage medium at a second speed, where the second speed is greater than the first speed.

[0084] Optionally, in the data processing device provided in the embodiments of the present application, the first execution module 20 further includes: a calculation sub-module, configured to calculate a data missing ratio when there is data missing in a set of merged data; a reading sub-module, configured to, when the data missing ratio is greater than a preset ratio, read the non-missing data in the merged data from the second storage medium, and read a target number of target data from the first storage medium, where the target number is the number of missing data; a re-merging sub-module, configured to merge the non-missing data and the target number of target data to obtain re-merged data; a tenth execution sub-module, configured to delete the non-missing data from the second storage medium, store the re-merged data in the second storage medium, delete the target number of target data from the first storage medium, and update the metadata corresponding to each target data in the re-merged data.

[0085] The data processing device includes a processor and a memory. The above-mentioned first determination unit 10, first execution unit 20, second execution unit 30, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.

[0086] The processor contains cores, and the cores retrieve corresponding program units from the memory. One or more cores can be set, and by adjusting the core parameters, the performance and space utilization rate of the small file storage scenario can be improved simultaneously.

[0087] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0088] The embodiment of the present application also provides a non-volatile storage medium, and the non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute a data processing method.

[0089] The embodiment of the present application also provides an electronic device, which includes a processor and a memory; computer-readable instructions are stored in the memory, and the processor is used to run the computer-readable instructions, wherein when the computer-readable instructions run, they execute a data processing method. The electronic device herein can be a server, a PC, a PAD, a mobile phone, etc.

[0090] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the functions specified in one or more of the blocks and / or processes. Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the blocks and / or processes. Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0094] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0095] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory. The memory is an example of a computer-readable medium.

[0096] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0097] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0098] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A data processing method, characterized in that, Including: During the process of writing target data into a storage system, determine the service type corresponding to the target data. Herein, the target data is data that occupies less storage space than a preset storage space. The service type includes a first type and a second type. The data writing fluctuation condition corresponding to the first type meets a preset fluctuation condition, and the data writing fluctuation condition corresponding to the second type does not meet the preset fluctuation condition. Herein, the preset fluctuation condition is the difference situation of the number of the target data to be stored corresponding to the same time interval in a preset period. Determine whether the data writing fluctuation condition meets the preset fluctuation condition according to the difference between the actual fluctuation condition of the target data and the preset fluctuation condition; When the service type corresponding to the target data is the first type, perform data merging on the target data in an offline merging manner; When the service type corresponding to the target data is the second type, perform data merging on the target data in an online merging manner.

2. The method according to claim 1, wherein When the service type corresponding to the target data is the second type, performing data merging on the target data in an online merging manner includes: Sequentially write each target data into a first storage medium in the storage system, record log data, and generate metadata corresponding to the target data; When each target data is written into the first storage medium in the storage system each time, add the target data to a merging module, and issue a merging task through the merging module to store the target data in a second storage medium in the storage system until the set condition of the merging task is met, and obtain the merged data. Herein, the set condition is used to set the number or the occupied storage space size of the target data corresponding to the merged data; Update the metadata corresponding to each target data in the merged data.

3. The method according to claim 2, wherein Issuing a merging task through the merging module to store the target data in a second storage medium in the storage system includes: Generate a plurality of the merging tasks, and add the target data to one of the plurality of merging tasks according to a preset rule; Control the plurality of merging tasks to concurrently write data into the second storage medium.

4. The method according to claim 3, wherein Controlling the plurality of merging tasks to concurrently write data into the second storage medium includes: When each merging task writes data into the second storage medium for the first time, adjust the state corresponding to the merging task from an initial state to an execution state; When each merging task writes all the data into the second storage medium, adjust the state corresponding to the merging task from the execution state to a completion state, and update the metadata of the data corresponding to the merging task.

5. The method according to claim 2, characterized in that, The method further includes: When an exception occurs in the process corresponding to the merging task, re-execute the uncompleted merging task.

6. The method according to claim 1, characterized in that, When the service type corresponding to the target data is the first type, performing data merging on the target data in an offline merging manner includes: Write each of the target data to a first storage medium in the storage system in sequence, record log data, and generate metadata corresponding to the target data; After a preset merging time is reached, obtain storage locations of multiple pieces of the target data from the log data, and generate a merging task, where the number of the multiple pieces of the target data is a preset number, the storage location is a storage address in the first storage medium, and the merging task is used to merge data stored corresponding to the storage location; Obtain multiple pieces of the target data from the first storage medium according to the storage location corresponding to the merging task, and merge the multiple pieces of the target data to obtain a set of merged data; Store the merged data in a second storage medium in the storage system, delete the multiple pieces of the target data from the first storage medium, and update the metadata corresponding to the multiple pieces of the target data.

7. The method according to claim 6, characterized in that, Obtaining multiple pieces of the target data from the first storage medium includes: Determine whether a writing rate of the target data to the first storage medium is greater than a preset rate; When the writing rate of the target data to the first storage medium is greater than the preset rate, obtain multiple pieces of the target data from the first storage medium at a first speed; When the writing rate of the target data to the first storage medium is less than or equal to the preset rate, obtain multiple pieces of the target data from the first storage medium at a second speed, where the second speed is greater than the first speed.

8. The method according to claim 6, wherein After storing the merged data in a second storage medium in the storage system, the method further includes: When there is data missing in a set of the merged data, calculate a data missing ratio; When the data missing ratio is greater than a preset ratio, read non-missing data in the merged data from the second storage medium, and read a target number of the target data from the first storage medium, where the target number is the number of missing data; Merge the non-missing data and the target number of the target data to obtain re-merged data; Delete the non-missing data from the second storage medium, store the re-merged data in the second storage medium, delete the target number of the target data from the first storage medium, and update the metadata corresponding to each of the target data in the re-merged data.

9. The method according to claim 1, wherein During the process of writing target data to a storage system, determining a service type corresponding to the target data includes: Within a preset time period, obtain writing operation times of the target data at preset time intervals to obtain multiple writing operation times values; Determine a distribution condition of the multiple writing operation times values within the preset time period, and determine whether a data writing fluctuation condition of the target data meets the preset fluctuation condition according to the distribution condition; When the data writing fluctuation condition of the target data meets the preset fluctuation condition, determine that the service type corresponding to the target data is the first type; In the case where the data writing fluctuation condition of the target data does not meet the preset fluctuation condition, determine that the service type corresponding to the target data is the second type.

10. An electronic device, characterized in that, Comprising a processor and a memory, wherein computer-readable instructions are stored in the memory, and the processor is configured to run the computer-readable instructions. Among them, when the computer-readable instructions run, they execute the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data storage method, scheduling device, system and device and storage medium

    CN109375868A

  • Mass small file storage performance optimization method and device based on real-time merging

    CN112416880A