Data processing method and device, electronic equipment, storage medium and program product
By determining the IO data amount in the target queue and writing to the free storage blocks first, the problem of low data writing efficiency of hard disks is solved, and low write latency and high read performance under different load conditions are achieved.
Patent Information
- Application Number
- CN202510342566.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has low efficiency when writing data in hard disks, especially in scenarios where the IO dispatch frequency is low, or the IO write delay is large when the load is high, resulting in low data writing efficiency in hard disks.
By determining the amount of data for each IO in the target queue, the largest free memory block is obtained, IO is written in the memory block, and IO is written in the memory block one by one on multiple disks of the hard disk, reducing the number of zero compensation, and improving the aggregation degree and writing efficiency of the IO.
Reduces the number of zero-compensation, reduces the write delay, improves the efficiency of hard disk writing data, and adapts to low write delay and high read performance under different load conditions.
Smart Images

Figure CN120295565A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a data processing method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] As a storage device of a computer, a hard disk is used to store data after the memory runs. Currently, when writing data to the hard disk, multiple I / Os are aggregated into one I / O, and the data corresponding to this I / O is written to the hard disk at one time. This method has a low I / O issuing frequency. In scenarios where data needs to be quickly written to the hard disk, it will result in excessive zero padding. If waiting to write to the hard disk after accumulating a certain amount of data, it will result in a large write latency. In summary, there is a problem of low efficiency in current hard disk data writing. Summary of the Invention
[0003] This application provides a data processing method, apparatus, electronic device, storage medium, and program product to solve the problem of low efficiency in current hard disk data writing.
[0004] In a first aspect, this application provides a data processing method, including:
[0005] Determine the data volume corresponding to each I / O in the target queue, where the target queue stores at least one I / O;
[0006] Among multiple storage blocks in a stripe, obtain the storage block with the largest free space as the first storage block, and the stripe is deployed in the memory;
[0007] According to the data volume corresponding to each I / O, write at least one target I / O in the first storage block, and at least one target I / O is the I / O in the target queue;
[0008] Determine whether all at least one I / O has been written into the stripe;
[0009] If not, execute the step of determining the data volume corresponding to each I / O in the target queue. If so, write the I / O in the storage block to the multiple disks in the hard disk one by one.
[0010] In a feasible implementation, the hard disk includes: stripes, the stripes include multiple shards, the multiple shards correspond to the multiple disks one by one, and each shard is a partial storage space of the corresponding disk. Writing the I / O in the storage block to the multiple disks in the hard disk one by one includes:
[0011] Determine a second data volume, where the second data volume is the data volume of the I / O stored in the second storage block, and the second storage block is the storage block with the smallest data volume of the I / O among the multiple storage blocks;
[0012] For each shard, write the data of the second data volume in the corresponding storage block into the shard.
[0013] In one implementable manner, writing at least one target IO into the first storage block according to the data volume corresponding to each IO includes:
[0014] Determining a data volume threshold of the first storage block according to the data volume corresponding to each IO, where the data volume threshold is used to represent the data volume that the first storage block can currently store;
[0015] Determining at least one target IO from at least one IO in the target queue according to the data volume threshold. After writing the at least one target IO into the first storage block, the data volume of the first storage block exceeds the data volume threshold;
[0016] Writing at least one target IO into the first storage block.
[0017] In one implementable manner, determining a data volume threshold of the first storage block according to the data volume corresponding to each IO includes:
[0018] Determining the total data volume of at least one IO stored in the target queue as the first total data volume according to the data volume corresponding to each IO;
[0019] Determining the second total data volume of the data stored in each storage block;
[0020] Determining an average data volume according to the sum of the first total data volume and the second total data volume, and the number of each storage block;
[0021] Determining whether there is a storage block in each storage block whose data volume is greater than or equal to the average data volume;
[0022] If not, determining the data volume threshold of the first storage block according to the average data volume;
[0023] If so, performing the step of determining the second total data volume of the data stored in each storage block according to the storage block whose data volume is less than the average data volume.
[0024] In one implementable manner, determining at least one target IO from at least one IO in the target queue according to the data volume threshold includes:
[0025] Determining the first data volume of the first storage block;
[0026] Determining at least one target IO from at least one IO according to the first data volume and the data volume threshold, where the sum of the total data volume of the at least one target IO and the first data volume is greater than or equal to the data volume threshold. The at least one target IO includes at least one first target IO and a second target IO. The second target IO is the last target IO in the sorted at least one target IO, and the sum of the total data volume of the at least one first target IO and the first data volume is less than the data volume threshold.
[0027] In one implementable manner, after writing at least one target IO into the first storage block, where the at least one target IO is an IO of a target queue, the method further includes:
[0028] Determine whether the amount of data in the stripe reaches the storage space of the stripe;
[0029] If so, perform the step of writing the IO in the storage block corresponding one by one to multiple disks of the hard disk;
[0030] If not, perform the step of determining whether at least one IO has been written into the stripe.
[0031] In a second aspect, the present application provides a data processing apparatus, including:
[0032] A data amount determination module, configured to determine the data amount corresponding to each IO in the target queue, where at least one IO is stored in the target queue;
[0033] An acquisition module, configured to acquire, from multiple storage blocks in the stripe, the storage block with the largest free space as the first storage block, where the stripe is deployed in the memory;
[0034] An IO writing module, configured to write at least one target IO into the first storage block according to the data amount corresponding to each IO, where the at least one target IO is an IO of the target queue;
[0035] A determination module, configured to determine whether at least one IO has been written into the stripe;
[0036] A processing module, configured to, if not, perform the step of determining the data amount corresponding to each IO in the target queue, and if so, write the IO in the storage block corresponding one by one to multiple disks of the hard disk.
[0037] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0038] The memory stores computer-executable instructions;
[0039] The processor executes the computer-executable instructions stored in the memory to implement the method according to the first aspect.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to the first aspect.
[0041] In a fifth aspect, the present application provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the method according to the first aspect.
[0042] The data processing method, apparatus, electronic device, storage medium, and program product provided by this application determine the data volume corresponding to each IO in the target queue, where at least one IO is stored in the target queue; in multiple storage blocks in a stripe, obtain the storage block with the largest free space as the first storage block, and the stripe is deployed in memory; according to the data volume corresponding to each IO, write at least one target IO in the first storage block, where at least one target IO is an IO in the target queue; determine whether all of the at least one IO have been written into the stripe; if not, execute the step of determining the data volume corresponding to each IO in the target queue, and if so, write the IO in the storage block to the corresponding one of multiple disks in the hard disk, so as to realize preferentially storing the IO in the first storage block with the largest free space, which can reduce the number of zero-padding. In addition, after all the at least one IO in the target queue have been written into multiple storage blocks, the at least one IO in the stripe is written to the hard disk, without waiting to accumulate a certain amount of data before writing to the hard disk, thereby reducing the latency. Therefore, this application can improve the efficiency of writing data to the hard disk. Description of the Drawings
[0043] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0044] Figure 1 It is a diagram of an implementation scenario shown in an exemplary embodiment;
[0045] Figure 2 It is a flowchart of a data processing method shown in an exemplary embodiment;
[0046] Figure 3 It is a schematic diagram of a target queue shown in an exemplary embodiment;
[0047] Figure 4 It is a schematic diagram of a stripe shown in an exemplary embodiment;
[0048] Figure 5 It is a schematic diagram of a hard disk shown in an exemplary embodiment;
[0049] Figure 6 It is a flowchart of a data processing method shown in another exemplary embodiment;
[0050] Figure 7 It is a flowchart of a method for determining a data volume threshold shown in an exemplary embodiment;
[0051] Figure 8 It is a schematic diagram of the change of IO in a stripe shown in an exemplary embodiment;
[0052] Figure 9 It is a schematic diagram of a method for reading IO shown in an exemplary embodiment;
[0053] Figure 10 It is a structural diagram of a data processing device shown in an exemplary embodiment;
[0054] Figure 11 It is a block diagram of an electronic device shown in an exemplary embodiment.
[0055] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0057] RAID (Redundant Arrays of Independent Disks) technology mainly uses data striping, mirroring, and data verification technologies to obtain high performance, reliability, fault tolerance, and scalability. According to the strategies and architectures of applying or combining these three technologies, RAID can be divided into different levels to meet the needs of different data applications.
[0058] In the related art, a way of filling stripes is that the data length written to the hard disk from each stripe is fixed. For example, if it is 32k, then after a storage block is filled with 32k, the next column is immediately switched to fill the stripe. After all storage blocks are filled with 32k, they are written to the hard disk together. The data written to the hard disk in this way is fixed. This method has the problem that if the set data length of the stripe is too large, in a low-load scenario, the IO issuance frequency is low. If it is required to quickly write to the disk, it will result in too much zero-padding for the stripe. If waiting to fill the stripe before writing to the disk, it will result in a large IO write latency. If the data length of the stripe is too small, when the load is high, the read and write times of the hard disk will be high, and the IO aggregation degree on the hard disk will be low.
[0059] The data processing method, device, electronic device, storage medium, and program product provided by the present application aim to solve the above-mentioned technical problems in the prior art.
[0060] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0061] The data processing method provided by the present application can be implemented in any electronic device with data processing capabilities or a data processing system. It should be noted that the data processing system can be deployed alone on an electronic device in any environment (for example, alone on an edge server in the edge environment), can be fully deployed in the cloud environment, or can be distributed in different environments.
[0062] For example, the data processing system can be logically divided into multiple parts, each part having different functions. Each part in the data processing system can be respectively deployed in any two or three of the electronic devices (located on the user side, such as clients), the edge environment, and the cloud environment. The edge environment is an environment including a set of edge electronic devices close to the electronic device, and the edge electronic devices include: edge servers, edge small stations with computing power, etc. Each part of the container startup system deployed in different environments or devices collaborates to implement the functions of the data processing platform.
[0063] It should be understood that the present application does not restrictively divide which parts of the data processing system are specifically deployed in what environment. In actual applications, it can be adaptively deployed according to the computing power of the electronic device, the resource occupancy of the edge environment and the cloud environment, or specific application requirements.
[0064] Such as Figure 1 As shown in the implementation scenario diagram of an exemplary embodiment, the data processing system corresponding to this implementation scenario includes a target queue 11, a memory 12, and a hard disk 13.
[0065] In the embodiment of the present application, the IO sent from the front end is first stored in the target queue 11, then the IO in the target queue 11 is stored in the memory 12, and finally the data in the memory 12 is written into the hard disk 13.
[0066] It can be understood that the data processing device can be set in Figure 1 the data processing system 10 in, but the implementation environment shown in this embodiment such as Figure 1 is only exemplary. In other embodiments, the data processing method can also be applied to other implementation environments, and the data processing device can also be set in other structures in other implementation environments, which are not specifically limited here.
[0067] Figure 2 is the flowchart of a data processing method shown in an exemplary embodiment, such asFigure 2 As shown in Figure 2 , the method includes steps S201 to S205, which are introduced in detail as follows:
[0068] S201. Determine the data volume corresponding to each IO in the target queue.
[0069] Among them, at least one IO (Input / Output) is stored in the target queue.
[0070] In the embodiment of the present application, for the IO sent from the front end, it can be first stored in the target queue, and the number of IOs in the target queue and the total data volume corresponding to all IOs can be monitored in real time. It can be understood that the data volume of each IO stored in the target queue, the number of at least one IO, and the total data volume of at least one IO
[0071] For example, referring to Figure 3 , 6 IOs are stored in the target queue, namely IO a1, IO a2, IO a3, IO a4, IO a5, and IO a6. Among them, the data volume corresponding to IO a1 is 24k, the data volume corresponding to IO a2 is 8k, the data volume corresponding to IO a3 is 80k, the data volume corresponding to IO a4 is 32k, the data volume corresponding to IO a5 is 48k, and the data volume corresponding to IO a6 is 16k. The number of IOs from IO a1 to IO a6 is 6, and the total data volume of IOs in the target queue is 208k.
[0072] In the embodiment of the present application, every time an IO is enqueued in the target queue, the number of IOs increases by 1, and the total data volume of IOs increases by the data volume of this IO. When performing subsequent steps, the total number and total data volume of IOs in the current target queue are saved.
[0073] S202. Among the multiple storage blocks in the stripe, obtain the storage block with the largest free space as the first storage block.
[0074] Among them, the stripe is deployed in the memory. It can be understood that the memory includes stripes, and the stripes include multiple storage blocks. The first storage block is the storage block with the largest free storage space. In the embodiment of the present application, the storage spaces of all storage blocks are the same.
[0075] For example, referring to Figure 4 , it is a schematic diagram of the stripe. Among them, the stripe includes storage blocks X1, X2, X3, and X4. The storage spaces of storage blocks X1 to X4 are all 256k. The current storage situation of the stripe is: IOb1 (such as the data volume is 8k) and IOb2 (such as the data volume is 8k) are stored in storage block X1, IOb3 (such as the data volume is 8k) is stored in storage block X3, and IOb4 (such as the data volume is 80k) is stored in storage block X4. In Figure 4The first storage block is storage block X2.
[0076] Further, if there are multiple storage blocks with the largest free space, one of them can be randomly determined as the first storage block.
[0077] S203. Write at least one target IO into the first storage block according to the data volume corresponding to each IO.
[0078] Among them, at least one target IO is an IO in the target queue. It can be understood that the target IO is determined according to the storage order of the IO in the target queue. If there is one target IO, the first IO is read from the head of the queue as the target IO. If there are multiple target IOs, multiple IOs are sequentially read from the head of the queue as the target IOs.
[0079] For example, referring to Figure 3 , if the target IO is 1, then IO a1 can be read as the target IO. If the target IO is 3, then IO a1, IO a2, and IO a3 can all be read as the target IOs.
[0080] S204. Determine whether all at least one IO has been written into the stripe.
[0081] If not, execute S201. If so, execute S205.
[0082] In the embodiments of the present application, all the IOs in the target queue need to be written into the stripe to prepare for subsequent writing into the hard disk to reduce the latency.
[0083] Further, during the process of repeatedly executing S201 to S203, the first storage block with the largest free space is re-determined each time, so that each IO can be stored in the first storage block with the largest free space each time, reducing the number of zero-padding.
[0084] S205. Write the IOs in the storage block into the corresponding disks among the multiple disks of the hard disk one by one.
[0085] Among them, the hard disk includes multiple disks, and the disks correspond to the storage blocks in the stripe one by one, and the IOs in the storage block are written into the corresponding disks.
[0086] In one embodiment, the hard disk can be a RAID. Among them, the number of multiple disks included in the hard disk is the same as the number of storage blocks in the stripe.
[0087] For example, referring to Figure 5 , the hard disk includes: 4 disks, namely disk Y1, disk Y2, disk Y3, and disk Y4. Each disk includes multiple shards.
[0088] In the embodiments of the present application, a shard can be selected from each disk, and then multiple shards can form a stripe. It can be understood that at least one IO is written into the shard corresponding to the stripe.
[0089] For example, referring to Figure 5 , the stripe T is composed of the shard U1 of disk Y1, the shard U2 of disk Y2, the shard U3 of disk Y3, and U4 of disk Y4. The shards correspond to storage blocks one by one. Then, the IO of storage block X1 is written into shard U1, the IO of storage block X2 is written into shard U2, the IO of storage block X3 is written into shard U3, and the IO of storage block X4 is written into shard U4.
[0090] In the embodiments of the present application, the IO in the storage block is written into the corresponding disk, so as to implement writing each IO into the hard disk.
[0091] In summary, the present application repeatedly executes S201 to S204 to implement preferentially storing the IO into the first storage block with the largest free space, which can reduce the number of zero-padding. In addition, after at least one IO in the target queue is written into multiple storage blocks, at least one IO in the stripe is written into the hard disk, without waiting to accumulate a certain amount of data before writing into the hard disk, thereby reducing the latency. Therefore, the present application can improve the efficiency of writing data to the hard disk.
[0092] Figure 6 is a flowchart of another data processing method shown in an exemplary embodiment. The method includes step S601 and step S607, which are introduced in detail as follows:
[0093] S601. Determine the data volume corresponding to each IO in the target queue.
[0094] For the specific implementation process of this step, refer to S201, which will not be elaborated here.
[0095] S602. Among the multiple storage blocks in the stripe, obtain the storage block with the largest free space as the first storage block.
[0096] For the specific implementation process of this step, refer to S202, which will not be elaborated here.
[0097] S603. According to the data volume corresponding to each IO, determine the data volume threshold of the first storage block.
[0098] Among them, the data volume threshold is used to represent the data volume that the first storage block can currently store.
[0099] In some embodiments, referring to Figure 7 , the method for determining the data volume threshold of the disk corresponding to the first storage block includes the following steps:
[0100] S6031. Determine that the total data volume of at least one IO stored in the target queue is the first data volume according to the data volume corresponding to each IO.
[0101] It can be understood that the first data volume is the total data volume of each IO stored in the target queue. The first data volume can be determined in real time, such as obtained by summing the data volumes of each IO. The first data volume can also be read from the target queue, such as Figure 3 the total data volume stored therein is the first data volume.
[0102] For example, referring to Figure 3 , the first data volume is 208k.
[0103] S6032. Determine the second data volume of the data stored in each storage block.
[0104] It can be understood that the second data volume is the total data volume of the IOs currently stored in each storage block.
[0105] For example, referring to Figure 4 , if the current storage blocks include all the storage blocks in the striping, the data volume of IO b1 is 8k, the data volume of IO b2 is 8k, the data volume of IO b3 is 8k, and the data volume of IO b4 is 80k, then the third data volume is 104k.
[0106] S6033. Determine the average data volume according to the sum of the first data volume and the second data volume, and the number of each storage block.
[0107] It can be understood that the sum of the first data volume and the second data volume is m. The number of each storage block is n, then the average data volume is V1 = m / n.
[0108] For example, in summary, m = 312k. n = 4, then the average data volume is V1 = 78k.
[0109] S6034. Determine whether there is a storage block in each storage block whose data volume is greater than or equal to the average data volume.
[0110] Among them, if there is no storage block whose data volume is greater than or equal to the average data volume, execute S6035. If there is a storage block whose data volume is greater than or equal to the average data volume, execute S6032 according to the storage block whose data volume is less than the average data volume.
[0111] For example, referring to the above, the average data volume is 78k, and the data volume of storage block X4 is 80K, then there is a storage block whose data volume is greater than or equal to the average data volume. The storage blocks whose data volume is less than the average data volume are storage blocks X1 to storage blocks X3, and continue to execute S6032 to S6034. Among them, executing S6032 can obtain the second total data volume of 24k. Executing S6033 can obtain the sum m of the first total data volume (208k) and the second total data volume (24k) as 232k. If n is 3, the average data volume V1 = 232 / 3k.
[0112] Among them, the data amounts of storage blocks X1 to X3 are all smaller than the average data amount, and then S6035 is executed.
[0113] S6035: Determine a data volume threshold of the first storage block according to the average data volume.
[0114] One way is to determine the average data volume as the data volume threshold. For example, the data volume threshold is 232 / 3k
[0115] Another way is to determine the data volume threshold by making the data volume threshold an integer multiple of the first preset value (such as 8k) and the absolute value of the difference between the average data volume and the data volume threshold is less than 8k. For example, the data volume threshold is 72k or 80k.
[0116] In the embodiment of the present application, the data volume threshold is less than the stripe depth and the stripe depth. For example, if the stripe size is 1024k and the stripe includes 4 storage blocks, the stripe depth is 256k, and if the stripe size is 4M and the stripe includes 4 shards, the stripe depth is 1M. It can be understood that if the average data volume is greater than the stripe depth or the stripe depth, the data volume threshold takes the minimum value of the stripe depth and the stripe depth. For example, if the average data volume is greater than 256k, the data volume threshold is determined to be 256k.
[0117] In an embodiment of the present application, the data volume threshold of the target storage block is determined in real time according to the number of IOs in the target queue and the total data volume of the IOs, thereby achieving the effects of balancing IO aggregation and reducing zero padding.
[0118] S604: Determine at least one target IO in at least one IO in the target queue according to the data volume threshold.
[0119] In one embodiment, determining at least one target IO from at least one IO according to a data volume threshold includes: determining a first data volume of a first storage block; determining at least one target IO from at least one IO according to the first data volume and the data volume threshold, where the sum of the total data volume of at least one target IO and the first data volume is greater than or equal to the data volume threshold, at least one target IO includes at least one first target IO and a second target IO, the second target IO is the last target IO in the order of at least one target IO, and the sum of the total data volume of at least one first target IO and the first data volume is less than the data volume threshold.
[0120] For example, referring to the above, the first storage block is Figure 4 storage block X2, where the data volume of storage block X2 is 0k. Then in Figure 3 , at the head of the target queue, read IOa1 in order first. The sum of the data volume of storage block X2 (0k) and the data volume of IOa1 (24k) is 24k, 24k is less than the data volume threshold (72k), continue to read IOa2. The sum of the data volume of storage block X2 (0k) and the data volumes of IOa1 (24k) and IOa2 (8k) is 32k. Continue to read IOa3. Then the sum of the data volume of storage block X2 and the data volumes of IOa1 to a3 (80k) is 112k, 112k is greater than the data volume threshold (72k), so IOa1 and IOa3 are determined as target IOs.
[0121] It can be seen that in this application, at least one target IO is determined through the data volume threshold and the first data volume of the first storage block, and the aggregation of IOs can be achieved.
[0122] In addition, this application can also determine at least one target IO in other ways. For example, the difference between the sum of the total data volume of at least one target IO and the first data volume and the data volume threshold is less than a preset value, and at least one target IO is determined using this limiting condition.
[0123] S605. Write at least one target IO into the first storage block.
[0124] In the embodiments of the present disclosure, write the first storage block (such as IOa1 to IOa3) into storage block X2. Referring to Figure 8 , after IO a1 to IO a3 in storage block X2, the striping state is as Figure 8 in (P1).
[0125] S606. Determine whether all at least one IO has been written into the stripe.
[0126] If not, execute S607; if so, execute S601. For example, in summary, there are still IO a4 to IO a6 not written into the stripe, and S601 to S605 can be looped.
[0127] Further, after executing S601, there are still IOs a4 to a6 in the target queue. Execute S602 to determine Figure 8 that the first storage block corresponding to (P1) is storage block X3. Execute S6031 to determine that the total amount of the first data is 176k. Execute S6032 to determine that the total amount of the second data is 144k. Execute S6033 to determine that the average data amount is 78k. Execute S6034. Determine that the data amounts of storage blocks X2 and X4 are greater than the average data amount, and re-execute S6032 using storage blocks X1 and X2 to determine that the total amount of the second data is 24k. Execute S6033 to determine that the average data amount is 100k, and then execute S6035 to determine that the data amount threshold is 100k. Then execute S604 and S605 to write IOs a4 and a5 in storage block X3. At this time, the state in the stripe refers to Figure 8 (P2) in. Further, there is still IOa6 to be stored in the stripe. For IO a6, execute the above steps, and IO a6 can be written into storage block X1. At this time, the state parameter in the stripe Figure 8 (P3) in.
[0128] In one embodiment, in the present disclosure, if there is only one IO in the target queue, the IO can be directly written into the target storage block with the largest free space.
[0129] In one embodiment, after writing at least one target IO into the first storage block, where the at least one target IO is an IO in the target queue, it further includes: determining whether the data amount in the stripe reaches the storage space of the stripe; if so, executing the step of writing the IOs in the storage blocks corresponding to the multiple disks of the hard disk one by one; if not, executing the step of determining whether all the at least one IO is written into the stripe.
[0130] It can be understood that if the stripe is full and not all the at least one IO is written into the stripe, S607 is also executed. After executing S607, then execute S601 to S607 for the remaining IOs in the target queue.
[0131] This application writes to the hard disk first after the stripe is full, and executing the above steps can ensure the continuous execution of the above steps.
[0132] In the embodiment of the present application, aggregating each write IO in sequence can create conditions for read io aggregation. After the front end issues continuous IOs, if these IOs come for reading, they are more likely to be aggregated, reducing the number of read operations on the hard disk and improving the read performance.
[0133] S607, determine the second data amount.
[0134] Wherein, the second data volume is the data volume of the IOs stored in the second storage block, and the second storage block is the storage block with the smallest data volume of the IOs among the multiple storage blocks.
[0135] S608. For each shard, write data with the second data volume in the corresponding storage block into the shard.
[0136] In the present disclosure, the second data volume is flexibly determined according to the data volume in the target queue and is not fixed. Therefore, when the data volume in the target queue is large, the number of times of writing data to the hard disk can be reduced, and the aggregation degree of the IOs can be improved. When the data volume in the target queue is small, without waiting for the data in the stripe to reach a certain data volume, the data in the stripe can be written to the hard disk according to the second data volume, avoiding latency.
[0137] For example, referring to Figure 8 of (P3), it can be determined that the second storage block is storage block X1, and the second data volume is 32k. Then, write 32k of data volume of each storage block into the corresponding shard, and the remaining data volume of each storage block refers to (P4). It can be understood that all the data in storage block X1 is written into shard U1, part of the data (IOa1 and IOa2) of storage block X2 is written into shard U2, part of the data (part of IOb3 and IOa4) of storage block X3 is written into shard U3, and part of the data (part of IOb4) of storage block X4 is written into shard U4. Among them, the data written into each shard is the same, all being 32k.
[0138] Furthermore, after writing all of IO a1 to IO a6 to the hard disk, when reading data from the hard disk, the data in the hard disk can be read in the order of IO a1 to IO a6.
[0139] For example, referring to Figure 9 , sort IO a1 to IO a6 in the hard disk, then merge IO a1 to IO a3 to form a new IO c1, merge IO a3 to IO a6 to form a new IO c2, and write read operation data O1 and read operation data O2 into the stripe.
[0140] In the embodiments of the present application, the above-mentioned aggregation of each write IO in sequence can create conditions for read io aggregation. After the front-end issues continuous IOs, if these IOs come to be read, they are more likely to be aggregated, reducing the number of read operations on the hard disk and improving the read performance.
[0141] In one embodiment, after part of the data in the stripe is written into the shard, continue to execute S601 to S607 to supplement new IOs in the target queue into the stripe, as Figure 8 in (P4), to realize continuous writing of the data in the stripe to the hard disk, reduce latency, and reduce the number of zero-fillings.
[0142] In addition, if it is determined that there is no new IO in the target queue, padding with zeros is performed after the current striping. After padding with zeros, the amount of data in each storage block is the same. For example, in (P4) in Figure 8 , storage block X1 is padded with 80k zeros, storage block X3 is padded with 24k zeros, and storage block X4 is padded with 12k zeros. Then, the data in each storage block is written to the hard disk, further reducing the latency.
[0143] In summary, the present application determines the number and total amount of data of the IOs in the target queue, and determines the data volume threshold of the target storage block in real time according to the amount of data already stored in the striping, improves the aggregation degree of IOs, reduces the amount of zero padding, and at the same time solves the problems of reducing latency and low IO aggregation degree, so that a lower write latency can be achieved under different load conditions. In addition, highly aggregated IOs can also be achieved under high load conditions, reducing the number of writes on the hard disk and improving the read bandwidth performance. When the load is low, the write latency can be reduced and the amount of zero padding can be reduced.
[0144] Figure 10 FIG. is a structural diagram of a data processing device shown in an exemplary embodiment. The data processing device 100 may include:
[0145] A data volume determination module 110, configured to determine the data volume corresponding to each IO in the target queue, where at least one IO is stored in the target queue;
[0146] An acquisition module 120, configured to obtain, among multiple storage blocks in a striping, the storage block with the largest free space as the first storage block, and the striping is deployed in the memory;
[0147] An IO writing module 130, configured to write at least one target IO in the first storage block according to the data volume corresponding to each IO, where the at least one target IO is an IO in the target queue;
[0148] A determination module 140, configured to determine whether all at least one IOs are written into the striping;
[0149] A processing module 150, configured to, if not, execute the step of determining the data volume corresponding to each IO in the target queue, and if so, write the IOs in the storage block to the multiple disks of the hard disk one by one.
[0150] In an implementable manner, the hard disk includes: stripes, each stripe includes a plurality of shards, the plurality of shards correspond to the multiple disks one by one, and each shard is a partial storage space of the corresponding disk. The processing module 150 includes:
[0151] A first determination unit, configured to determine a second data volume, where the second data volume is the data volume of the IOs stored in a second storage block, and the second storage block is the storage block with the smallest data volume of the IOs among the multiple storage blocks;
[0152] The first writing unit is configured to write data with a second data volume in a corresponding storage block for each shard.
[0153] In an implementable manner, the IO writing module 130 includes:
[0154] The second determining unit is configured to determine a data volume threshold of the first storage block according to the data volume corresponding to each IO, and the data volume threshold is used to represent the data volume that the first storage block can currently store;
[0155] The third determining unit is configured to determine at least one target IO from at least one IO in the target queue according to the data volume threshold. After the at least one target IO is written into the first storage block, the data volume of the first storage block exceeds the data volume threshold;
[0156] The second writing unit is configured to write at least one target IO into the first storage block.
[0157] In an implementable manner, the second determining unit is specifically configured to:
[0158] Determine the total amount of data of at least one IO stored in the target queue as the first total amount of data according to the data volume corresponding to each IO;
[0159] Determine the second total amount of data of the data stored in each storage block;
[0160] Determine the average data volume according to the sum of the first total amount of data and the second total amount of data, and the number of each storage block;
[0161] Determine whether there is a storage block in each storage block whose data volume is greater than or equal to the average data volume;
[0162] If not, determine the data volume threshold of the first storage block according to the average data volume;
[0163] If so, execute the step of determining the second total amount of data of the data stored in each storage block according to the storage block whose data volume is less than the average data volume.
[0164] In an implementable manner, the third determining unit is specifically configured to:
[0165] Determine the first data volume of the first storage block;
[0166] Determine at least one target I / O among at least one I / O according to the first data volume and the data volume threshold, where the sum of the total data volume of the at least one target I / O and the first data volume is greater than or equal to the data volume threshold. The at least one target I / O includes at least one first target I / O and a second target I / O. The second target I / O is the last sorted target I / O among the at least one target I / O, and the sum of the total data volume of the at least one first target I / O and the first data volume is less than the data volume threshold.
[0167] In an implementable manner, the data volume determination module 110 is further configured to: after writing at least one target I / O, which is the I / O of the target queue, into the first storage block, determine whether the data volume in the stripe reaches the storage space of the stripe; if so, perform the step of writing the I / O in the storage block into the corresponding one of the multiple disks of the hard disk; if not, perform the step of determining whether at least one I / O has been written into the stripe.
[0168] The data processing device provided in this embodiment can be used to execute the above data processing method. The implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0169] Figure 11 It is a block diagram of an electronic device shown in an exemplary embodiment. Please refer to Figure 11 The electronic device 110 may include: a processor 111 and a memory 112, where the processor 111 and the memory 112 can communicate; exemplarily, the processor 111 and the memory 112 communicate through a communication bus 113. The memory 112 is used to store computer execution instructions, and the processor 111 is used to call the computer execution instructions in the memory to execute the data processing method shown in any of the above method embodiments.
[0170] The above processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.
[0171] This application provides a computer-readable storage medium, on which computer execution instructions are stored; when the computer execution instructions are executed by a processor, they are used to implement the data processing method as described in any of the above embodiments.
[0172] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the data processing method of the above-mentioned distributed system is implemented.
[0173] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0174] Furthermore, it should be noted that although the steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0175] It should be understood that the above device embodiments are only illustrative, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0176] In addition, unless otherwise specified, in each embodiment of the present application, the functional units / modules can be integrated into one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.
[0177] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0178] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned memory includes: USB flash drive, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disc, etc., all kinds of media that can store program codes.
[0179] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0180] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0181] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A data processing method, characterized in that, Including: Determine the data volume corresponding to each IO in the target queue, where at least one IO is stored in the target queue; Among multiple storage blocks in a stripe, obtain the storage block with the largest free space as the first storage block, and the stripe is deployed in the memory; According to the data volume corresponding to each IO, write at least one target IO in the first storage block, and the at least one target IO is the IO in the target queue; Determine whether all of the at least one IO are written into the stripe; If not, execute the step of determining the data volume corresponding to each IO in the target queue. If so, write the IO in the storage block into the corresponding ones of multiple disks of the hard disk.
2. The method according to claim 1, wherein The hard disk Including: a stripe, the stripe includes multiple shards, the multiple shards correspond to the multiple disks one by one, and each shard is a partial storage space of the corresponding disk. The writing the IO in the storage block into the corresponding ones of multiple disks of the hard disk includes: Determine a second data volume, where the second data volume is the data volume of the IO stored in the second storage block, and the second storage block is the storage block with the smallest data volume of the IO among the multiple storage blocks; For each shard, write the data of the second data volume in the corresponding storage block into the shard.
3. The method according to claim 2, wherein The writing at least one target IO in the first storage block according to the data volume corresponding to each IO includes: According to the data volume corresponding to each IO, determine a data volume threshold of the first storage block, and the data volume threshold is used to represent the data volume that the first storage block can currently store; According to the data volume threshold, determine at least one target IO among the at least one IO in the target queue, and after the at least one target IO is written into the first storage block, the data volume of the first storage block exceeds the data volume threshold; Write the at least one target IO in the first storage block.
4. The method according to claim 3, wherein The determining the data volume threshold of the first storage block according to the data volume corresponding to each IO includes: According to the data volume corresponding to each IO, determine the total data volume of at least one IO stored in the target queue as the first total data volume; Determine the second total data volume of the data stored in each storage block; According to the sum of the first total data volume and the second total data volume, and the number of each storage block, determine the average data volume; Determine whether there is a storage block in each storage block whose data volume is greater than or equal to the average data volume; If not, determine the data volume threshold of the first storage block according to the average data volume; If so, according to the storage block whose data volume is less than the average data volume, execute the step of determining the second total data volume of the data stored in each storage block.
5. The method according to claim 4, wherein The determining at least one target IO among the at least one IO in the target queue according to the data volume threshold includes: Determine the first data volume of the first storage block; Determine at least one target I / O among the at least one I / O according to the first data volume and the data volume threshold, where the sum of the total data volume of the at least one target I / O and the first data volume is greater than or equal to the data volume threshold. The at least one target I / O includes at least one first target I / O and a second target I / O, and the second target I / O is the last sorted target I / O among the at least one target I / O. The sum of the total data volume of the at least one first target I / O and the first data volume is less than the data volume threshold.
6. The method according to any one of claims 1 to 5, characterized in that After writing at least one target I / O into the first storage block, where the at least one target I / O is the I / O of the target queue, it further includes: Determine whether the data volume in the stripe reaches the storage space of the stripe; If so, execute the step of writing the I / O in the storage block corresponding to each of the multiple disks of the hard disk one by one; If not, execute the step of determining whether all the at least one I / O are written into the stripe.
7. A data processing device, characterized in that, It includes: A data volume determination module, configured to determine the data volume corresponding to each I / O in the target queue, where at least one I / O is stored in the target queue; An acquisition module, configured to acquire, among the multiple storage blocks in the stripe, the storage block with the largest free space as the first storage block, where the stripe is deployed in the memory; An I / O writing module, configured to write at least one target I / O into the first storage block according to the data volume corresponding to each I / O, where the at least one target I / O is the I / O of the target queue; A determination module, configured to determine whether all the at least one I / O are written into the stripe; A processing module, configured to, if not, execute the step of determining the data volume corresponding to each I / O in the target queue, and if so, write the I / O in the storage block corresponding to each of the multiple disks of the hard disk one by one.
8. An electronic device, characterized in that, It includes: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-6.