Data aggregation method and device, electronic equipment and storage medium

By calculating the data pulling overhead in the erasure code storage system and selecting the appropriate aggregation mode, the problem of excessive data pulling overhead in the replica write solution in the shard failure scenario is solved, and more efficient data aggregation is achieved.

CN120803348APending Publication Date: 2025-10-17CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510879939.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In erasure coded storage systems, the background aggregation process of the replica write solution severely amplifies the data pulling overhead in shard failure scenarios, resulting in excessive overhead in the data aggregation process.

Method used

By calculating the initial and target data pulling costs, the most suitable full or incremental aggregation mode is selected for data aggregation to reduce data pulling costs.

Benefits of technology

It effectively reduces the data pulling overhead during the data aggregation process and improves data aggregation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803348A_ABST
    Figure CN120803348A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a data aggregation method and apparatus, an electronic device and a storage medium, and can solve the problem of how to effectively reduce overhead for data pulling in a data aggregation process. The method comprises the steps of obtaining to-be-aggregated data and at least one to-be-aggregated data fragment corresponding to the to-be-aggregated data; according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data fragment and the data stored in the at least one verification fragment, initial data pulling overhead is calculated; when it is detected that the to-be-aggregated data meets the aggregation mode conversion condition, target data pulling overhead corresponding to a target aggregation mode after aggregation mode conversion is calculated, and the target aggregation mode is full-amount aggregation or incremental aggregation; and when the target data pulling overhead is smaller than the initial data pulling overhead, performing data aggregation according to the target aggregation mode through the target verification fragment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of data aggregation, and in particular to a data aggregation method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the increasing informatization, the amount of data generated by the whole society every day presents an explosive growth, therefore, people's demand for the reliability and availability of data storage becomes more and more urgent. For the erasure code redundancy strategy, in order to reduce the performance impact on the write IO, the replica write scheme emerges as the times require, however, although the replica write scheme solves the write performance problem under the condition that the strip is not full, and skillfully utilizes the background aggregation task to periodically aggregate the strip data, due to the characteristics of erasure code storage, whether it is incremental aggregation or full aggregation, as long as in the scene with shard failure, the overhead of pulling data in the background aggregation process of the replica write scheme may be seriously amplified, therefore, how to effectively reduce the overhead of data pulling in the data aggregation process has become a problem to be solved at present. SUMMARY

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present application provide a data aggregation method, device, electronic equipment and storage medium, to solve the problem of how to effectively reduce the overhead of data pulling in the data aggregation process.

[0004] In order to achieve the above purpose, the technical scheme provided by the embodiments of the present application is as follows:

[0005] In a first aspect, the embodiments of the present application provide a data aggregation method, which comprises: obtaining to-be-aggregated data and at least one to-be-aggregated data shard corresponding to the to-be-aggregated data;

[0006] According to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard and the data stored in at least one check shard, an initial data pulling overhead is calculated;

[0007] When it is detected that the to-be-aggregated data meets the aggregation mode conversion condition, a target data pulling overhead corresponding to a target aggregation mode after aggregation mode conversion is calculated, the target aggregation mode being full aggregation or incremental aggregation;

[0008] When the target data pulling overhead is less than the initial data pulling overhead, data aggregation is performed according to the target aggregation mode through a target check shard.

[0009] As an optional implementation, in the first aspect of the embodiment of the present application, the initial data pulling overhead is calculated according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard, and the initial aggregation mode is determined according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards.

[0010] The initial aggregation mode is determined according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards, and the initial aggregation mode is full-quantity aggregation or incremental aggregation.

[0011] The initial data pulling overhead corresponding to the initial aggregation mode is obtained by performing pulling overhead calculation according to the initial aggregation mode, according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard.

[0012] As an optional implementation, in the first aspect of the embodiment of the present application, the initial aggregation mode is determined according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards, and the initial aggregation mode is determined according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards.

[0013] The first quantity of the at least one to-be-aggregated data shard and the second quantity of the pre-stored plurality of data shards are determined.

[0014] The initial aggregation mode is determined according to the first quantity and the second quantity.

[0015] As an optional implementation, in the first aspect of the embodiment of the present application, the initial data pulling overhead corresponding to the initial aggregation mode is obtained by performing pulling overhead calculation according to the initial aggregation mode, according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard.

[0016] A plurality of original data pulling steps corresponding to the initial aggregation mode are determined according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard.

[0017] The data pulled by the plurality of original data pulling steps is de-duplicated and optimized to obtain a plurality of initial data pulling steps.

[0018] The initial data pulling overhead is calculated according to the plurality of initial data pulling steps.

[0019] As an optional implementation, in the first aspect of the embodiment of the present application, before the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion is calculated when it is detected that the to-be-aggregated data satisfies the aggregation mode conversion condition, the method further comprises:

[0020] When the initial aggregation mode is full-quantity aggregation, if it is detected that the current pullable data includes the to-be-aggregated data, the original data corresponding to the at least one to-be-aggregated data shard, and the check data of the target check shard, it is determined that the to-be-aggregated data satisfies the aggregation mode conversion condition, and the target aggregation mode after the aggregation mode conversion is incremental aggregation.

[0021] When the initial aggregation mode is incremental aggregation, if it is detected that the current pullable data includes the to-be-aggregated data and the original data corresponding to the data shards other than the at least one to-be-aggregated data shard, it is determined that the to-be-aggregated data satisfies the aggregation mode conversion condition, and the target aggregation mode after the aggregation mode conversion is full-quantity aggregation.

[0022] As an optional implementation, in the first aspect of the embodiment of the present application, when the target aggregation mode is full-quantity aggregation, the data aggregation by the target check shard according to the target aggregation mode when the target data pulling overhead is less than the initial data pulling overhead includes:

[0023] pulling the to-be-aggregated data from the target check shard when the target data pulling overhead is less than the initial data pulling overhead;

[0024] pulling at least one original data from at least one target data shard, the at least one target data shard being all data shards other than the at least one to-be-aggregated data shard in the pre-stored plurality of data shards;

[0025] performing full-quantity aggregation on the to-be-aggregated data and the at least one original data.

[0026] As an optional implementation, in the first aspect of the embodiment of the present application, when the target aggregation mode is incremental aggregation, the data aggregation by the target check shard according to the target aggregation mode when the target data pulling overhead is less than the initial data pulling overhead includes:

[0027] pulling the to-be-aggregated data from the target check shard when the target data pulling overhead is less than the initial data pulling overhead;

[0028] pulling at least one original data from the at least one to-be-aggregated data shard;

[0029] pulling at least one check data from the at least one check shard;

[0030] performing incremental aggregation on the to-be-aggregated data, the at least one original data, and the at least one check data.

[0031] In a second aspect, an embodiment of the present application provides a data aggregation apparatus, comprising: an obtaining module, configured to obtain to-be-aggregated data and at least one to-be-aggregated data shard corresponding to the to-be-aggregated data;

[0032] a processing module, configured to calculate an initial data pulling cost according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in at least one check shard;

[0033] The processing module is further configured to, when it is detected that the to-be-aggregated data satisfies an aggregation mode conversion condition, calculate a target data pulling cost corresponding to a target aggregation mode after aggregation mode conversion, the target aggregation mode being full-amount aggregation or incremental aggregation;

[0034] The processing module is further configured to, when the target data pulling cost is less than the initial data pulling cost, perform data aggregation according to the target aggregation mode through a target check shard.

[0035] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to determine an initial aggregation mode according to the at least one to-be-aggregated data shard and a plurality of pre-stored data shards, the initial aggregation mode being full-amount aggregation or incremental aggregation;

[0036] The processing module is specifically configured to perform pulling cost calculation according to the initial aggregation mode, to obtain the initial data pulling cost corresponding to the initial aggregation mode, according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in the at least one check shard.

[0037] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to determine a first quantity of the at least one to-be-aggregated data shard and a second quantity of the plurality of pre-stored data shards;

[0038] The processing module is specifically configured to determine the initial aggregation mode according to the first quantity and the second quantity.

[0039] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is specifically configured to determine a plurality of original data pulling steps corresponding to the initial aggregation mode, according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in the at least one check shard.

[0040] The processing module is specifically configured to perform deduplication optimization on data pulled by the plurality of original data pulling steps, to obtain a plurality of initial data pulling steps.

[0041] The processing module is specifically configured to calculate the initial data pulling overhead according to the plurality of initial data pulling steps.

[0042] As an optional implementation, in the second aspect of the embodiment of the present application, the processing module is further configured to, when the initial aggregation mode is full aggregation, determine that the to-be-aggregated data satisfies the aggregation mode conversion condition and the target aggregation mode after the aggregation mode conversion is incremental aggregation, if it is detected that the current pullable data includes the to-be-aggregated data, the original data corresponding to the at least one to-be-aggregated data shard and the check data of the target check shard.

[0043] The processing module is further configured to, when the initial aggregation mode is incremental aggregation, determine that the to-be-aggregated data satisfies the aggregation mode conversion condition and the target aggregation mode after the aggregation mode conversion is full aggregation, if it is detected that the current pullable data includes the to-be-aggregated data and the original data corresponding to the data shards other than the at least one to-be-aggregated data shard.

[0044] As an optional implementation, in the second aspect of the embodiment of the present application, when the target aggregation mode is full aggregation, the processing module is specifically configured to pull the to-be-aggregated data from the target check shard when the target data pulling overhead is less than the initial data pulling overhead.

[0045] The processing module is specifically configured to pull at least one original data from at least one target data shard, the at least one target data shard being all data shards other than the at least one to-be-aggregated data shard in the plurality of data shards pre-stored.

[0046] The processing module is specifically configured to perform full aggregation on the to-be-aggregated data and the at least one original data.

[0047] As an optional implementation, in the second aspect of the embodiment of the present application, when the target aggregation mode is incremental aggregation, the processing module is specifically configured to pull the to-be-aggregated data from the target check shard when the target data pulling overhead is less than the initial data pulling overhead.

[0048] The processing module is specifically configured to pull at least one original data from the at least one to-be-aggregated data shard.

[0049] The processing module is specifically configured to pull at least one check data from the at least one check shard.

[0050] The processing module is specifically configured to perform data incremental aggregation on the data to be aggregated, the at least one original data, and the at least one verification data.

[0051] In a third aspect, an embodiment of the present application provides an electronic device, comprising:

[0052] a memory storing executable program code;

[0053] a processor coupled to the memory;

[0054] The processor calls the executable program code stored in the memory to execute the data aggregation method in the first aspect of the embodiment of the present application.

[0055] In a fourth aspect, embodiments of the present application provide a computer-readable storage medium storing a computer program that causes a computer to execute the data aggregation method of the first aspect of the embodiments of the present application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, or an optical disk.

[0056] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute part or all of the steps of any one of the methods of the first aspect.

[0057] In a sixth aspect, an embodiment of the present application provides an application publishing platform, which is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps of any one of the methods of the first aspect.

[0058] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0059] The embodiments of the present application provide a data aggregation method, device, electronic device and storage medium, which obtain the data to be aggregated and at least one data shard to be aggregated corresponding to the data to be aggregated; calculate the initial data pulling overhead based on the data to be aggregated, the original data stored in at least one data shard to be aggregated and the data stored in at least one verification shard; when it is detected that the data to be aggregated meets the aggregation mode conversion conditions, calculate the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion, and the target aggregation mode is full aggregation or incremental aggregation; when the target data pulling overhead is less than the initial data pulling overhead, perform data aggregation according to the target aggregation mode through the target verification shard. In this solution, the aggregation mode that best suits the current data can be selected by converting the aggregation mode, which effectively reduces the overhead for data pulling during the data aggregation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings, which are incorporated herein and constitute part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application.

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings needed to be used in the embodiments will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0062] Figure 1 is a scene schematic diagram of a data aggregation method provided by an embodiment of the present application;

[0063] Figure 2 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 1 ;

[0064] Figure 3 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 2 ;

[0065] Figure 4 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 3 ;

[0066] Figure 4 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 5 ;

[0067] Figure 1 is a flow schematic diagram of a data aggregation method provided by an embodiment of the present application Figure 7 ;

[0068] Figure 2 is a flow schematic diagram of a data aggregation method provided by an embodiment of the present application Figure 8 ;

[0069] Figure 5 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 9 ;

[0070] Figure 6 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 10 ;

[0071] Figure 7 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 8 ;

[0072] Figure 12 is a data pulling step schematic diagram provided by an embodiment of the present application Figure 13 ;

[0073] Figure 1 is a structural schematic diagram of a data aggregation device provided by an embodiment of the present application.

[0074] Figure 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to more clearly understand the above objectives, features and advantages of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be explained that, in the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without any creative work, are within the protection scope of the present application.

[0076] The terms "first" and "second" and the like in the specification and claims of the present application are used to distinguish different objects, but are not used to describe a specific order of the objects.

[0077] The terms "include" and "have" and any variations thereof in the embodiments of the present application are intended to cover the non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units need not be limited to those clearly listed steps or units, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0078] It should be explained that, in the embodiments of the present application, the words "exemplary" or "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. In fact, the words "exemplary" or "for example" are intended to present the related concept in a specific way.

[0079] With the increasing of informatization, the amount of data generated by the whole society every day presents an explosive growth, therefore, people's demand for the reliability and availability of data storage becomes more and more urgent; replica and erasure code are two common redundancy strategies in the neighborhood of distributed storage; for some key business scenarios with high performance requirements, such as database scenarios, users often expect to adopt replica redundancy strategy to ensure the safety of data, and at the same time, the read-write performance can also be guaranteed; and for other business scenarios with high storage capacity utilization rate, such as video storage scenarios, users often expect to adopt erasure code redundancy strategy, which ensures the safety of data and also improves the storage capacity utilization rate.

[0080] Distributed storage system is a data storage system that stores data in multiple independent devices. The traditional network storage system uses a centralized storage server to store all data, and the storage server becomes the bottleneck of system performance and the focus of reliability and security, which cannot meet the needs of large-scale storage applications. Distributed network storage system uses a scalable system structure, uses multiple storage servers to share the storage load, and uses a location server to locate the storage information. It not only improves the reliability, availability and access efficiency of the system, but also is easy to expand.

[0081] Replica data storage is also a data protection method, which takes the original data as a whole, stores n copies of the same data in n different locations according to the redundancy level, to achieve the purpose of data protection.

[0082] Erasure code data storage is a data protection method, which cuts the original data into k original data blocks according to the specified size, then encodes them into p parity data blocks according to the redundancy level, and finally stores the original data blocks and parity data blocks in different locations to achieve the purpose of data protection.

[0083] For erasure code redundancy strategy, replica write scheme is born to reduce the performance impact on write IO; the general process of replica write scheme is: for the full stripe, directly calculate the parity data corresponding to the stripe data, then store the original data block together with the parity data block; for the not full stripe, directly write it as a whole to the corresponding data shard and parity shard in the form of replica; start a background aggregation task on the parity shard, which is responsible for periodically detecting the data of the written stripe, and aggregating the replica data on the stripe into parity data, and then deleting the replica data.

[0084] As shown in Figure 2 , based on erasure code 8+2, chunk size is equal to 4K, the stripe has been written full, and the parity data is generated on the parity shard P and Q; based on Figure 3The data states shown are respectively explained and described for incremental aggregation and full aggregation.

[0085] Incremental aggregation: for erasure code k+p, when the number of data coverage chunks to be aggregated is less than or equal to k / 2, incremental aggregation is performed.

[0086] When there is no shard failure, as shown in Figure 4 Data shards 6 and 7 are overwritten, that is, new data G' and H' are written, and the number of data coverage chunks to be aggregated is 2, satisfying the incremental aggregation condition. The check shard Q performs the aggregation task, and the data to be pulled includes:

[0087] (1) Pull data G' and H' from the local;

[0088] (2) Pull data Q from the local;

[0089] (3) Pull data P from the check shard P;

[0090] (4) Pull historical data G from data shard 6;

[0091] (5) Pull historical data H from data shard 7;

[0092] In summary, the cost of pulling data is: (G'+H')+P+Q+G+H; finally, the new check data can be encoded by diff((G'H'), (GH))+P+Q, that is, the aggregation is completed.

[0093] When there is a shard failure, as shown in Figure 5 Data shards 6 and 7 are overwritten, that is, new data G' and H' are written, and the number of data coverage chunks to be aggregated is 2, satisfying the incremental aggregation condition. The check shard Q performs the aggregation task, and the data to be pulled includes:

[0094] (1) Pull data G' and H' from the local;

[0095] (2) Pull data Q from the local;

[0096] (3) Pull data P from the check shard P;

[0097] (4) Pull data Q from the local, pull data A, B, C, D, E, and F from data shards 0, 1, 2, 3, 4, and 5, respectively, pull data P from the check shard P, and decode the historical data G by A+B+C+D+E+F+P+Q;

[0098] (5) Pull data Q from the local, pull data A, B, C, D, E, and F from data shards 0, 1, 2, 3, 4, and 5, respectively, pull data P from the check shard P, and decode the historical data H by A+B+C+D+E+F+P+Q;

[0099] In summary, the overhead of pulling data is: (G'+H')+P+Q+(A+B+C+D+E+F+P+Q)+(A+B+C+D+E+F+P+Q); finally, the new check data can be encoded by diff((G'H'),(GH))+P+Q, that is, the aggregation is completed.

[0100] As can be seen, in the scenario of shard failure, the overhead of pulling data will be greatly increased.

[0101] Full aggregation: for erasure code k+p, when the number of chunks covered by the data to be aggregated is greater than k / 2, full aggregation is performed.

[0102] When there is no shard failure, as shown in Figure 6 , data shards 3, 4, 5, 6, and 7 are overwritten, that is, new data D', E', F', G', and H' is written, and the number of chunks covered by the data to be aggregated is 5, which meets the full aggregation condition. Check shard Q performs the aggregation task, and the data to be pulled includes:

[0103] (1) pull data D', E', F', G', and H' from the local;

[0104] (2) pull data A from data shard 0;

[0105] (3) pull data B from data shard 1;

[0106] (4) pull data C from data shard 2;

[0107] In summary, the overhead of pulling data is: (D'+E'+F'+G'+H')+A+B+C; finally, the new check data can be encoded by D'+E'+F'+G'+H'+A+B+C, that is, the aggregation is completed.

[0108] When there is a shard failure, as shown in Figure 6 , data shards 3, 4, 5, 6, and 7 are overwritten, that is, new data D', E', F', G', and H' is written, and the number of chunks covered by the data to be aggregated is 5, which meets the full aggregation condition. Check shard Q performs the aggregation task, and the data to be pulled includes:

[0109] (1) pull data D', E', F', G', and H' from the local;

[0110] (2) pull data A from data shard 0;

[0111] (3) Pull data Q from the local computer, pull data A from data shard 0, pull historical data D, E, F, G, and H from data shards 3, 4, 5, 6, and 7 respectively, pull data P from the verification shard P, and decode data B through A+D+E+F+G+H+P+Q;

[0112] (4) Pull data Q from the local computer, pull data A from data shard 0, pull historical data D, E, F, G, and H from data shards 3, 4, 5, 6, and 7 respectively, pull data P from the verification shard P, and decode data C through A+D+E+F+G+H+P+Q;

[0113] In summary, the overhead of pulling data is: (D'+E'+F'+G'+H')+A+(A+D+E+F+G+H+P+Q)+(A+D+E+F+G+H+P+Q); finally, the new verification data can be encoded using D'+E'+F'+G'+H'+A+B+C, thus completing the aggregation.

[0114] It can be seen that in the scenario of shard failure, the cost of pulling data will increase significantly.

[0115] In summary, although the replica write solution solves the write performance issue to a certain extent when the stripe is not fully written, and cleverly uses background aggregation tasks to regularly aggregate stripe data, due to the characteristics of erasure code storage, whether it is incremental aggregation or full aggregation, as long as there is a shard failure scenario, the overhead of pulling data in the background aggregation process of the replica write solution may be severely amplified.

[0116] In order to solve some or all of the above technical problems, the embodiments of the present application provide a data aggregation method, device, electronic device and storage medium, which obtains data to be aggregated and at least one data shard to be aggregated corresponding to the data to be aggregated; calculates the initial data pulling overhead based on the data to be aggregated, the original data stored in at least one data shard to be aggregated and the data stored in at least one verification shard; when it is detected that the data to be aggregated meets the aggregation mode conversion conditions, calculates the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion, and the target aggregation mode is full aggregation or incremental aggregation; when the target data pulling overhead is less than the initial data pulling overhead, the data is aggregated according to the target aggregation mode through the target verification shard. In this solution, the aggregation mode that best suits the current data can be selected by converting the aggregation mode, which effectively reduces the overhead for data pulling during the data aggregation process.

[0117] like Figure 2 As shown, Figure 3 A flowchart of a data aggregation method is provided for an embodiment of the present application. The method may include the following steps:

[0118] 601、obtain the to-be-aggregated data and at least one to-be-aggregated data shard corresponding to the to-be-aggregated data.

[0119] In the embodiment of the present application, the to-be-aggregated data is the data in the data shard that has been written to be aggregated, which covers the original data in the data shard. It can be understood as Figure 4 and Figure 5 G' and H' cover the original data G and H. It can also be understood as Figure 1 and Figure 1 D', E', F', G' and H' cover the original data D, E, F, G and H. The to-be-aggregated data can correspond to at least one to-be-aggregated data shard. Since the data shard is divided according to the size of the stripe, it can be seen from Figure 2 that each data shard stores 4K size data in turn, so at least one to-be-aggregated data shard can be determined according to the position of the to-be-aggregated data in the stripe.

[0120] For example, as shown in Figure 3 , assuming that the to-be-aggregated data is 13k-24k data, then the corresponding data shards are data shard 3, data shard 4 and data shard 5.

[0121] It should be noted that the obtaining of the to-be-aggregated data described herein can be understood as only determining which data is to-be-aggregated data and the data partition newly written by the to-be-aggregated data, rather than referring to obtaining the specific data content. The specific data content needs to be obtained by subsequent data pulling.

[0122] 602、According to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard and the data stored in the at least one check shard, calculate the initial data pulling overhead.

[0123] In the embodiment of the present application, the initial data pulling overhead can be understood as the overhead generated for data pulling in the process of data aggregation according to the initial data aggregation method. The initial data aggregation method can be considered as the data aggregation method described in the above prior art such as Figure 4 , Figure 5 , Figures 1 to 5 and Figure 7 . Since data aggregation needs to ensure that the original data and newly written data can be obtained, if a data shard fails, the original data may also need to be obtained by reverse derivation through the data stored in the check shard. Therefore, when calculating the initial data pulling overhead, it needs to be determined according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard and the data stored in the at least one check shard.

[0124] It should be noted that the check shard includes replica space and check space. The replica space stores replica data, which can be understood as the replica data obtained by directly writing the data in the data shard into the replica space of the check shard. The check space stores check data, which can be understood as the check data obtained by performing a certain matrix deduction when the replica space stripe data is full. Similarly, the check data can also be reversely deduced into replica data through the matrix, that is, the original data in the data shard.

[0125] 603. When it is detected that the data to be aggregated meets the aggregation mode conversion condition, the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion is calculated.

[0126] It should be noted that data aggregation methods can include full aggregation and incremental aggregation, and full aggregation and incremental aggregation can be converted into each other. The calculation method of incremental aggregation is that each data immediately participates in the aggregation calculation when it arrives, and only maintains the intermediate state (such as cumulative sum, maximum value, etc.), without saving the original data. For example: calculate the sum of the numbers in the window, and update the cumulative value each time new data arrives. Incremental aggregation has the characteristics of low latency and high throughput, and is suitable for simple aggregation (such as sum and maximum value); the calculation method of full aggregation is that when the window is triggered, all the original data in the window are traversed, and the aggregation calculation (such as sorting and median) is completed at one time. For example: to calculate the median of all data in the window, all data must be saved and then sorted. Full aggregation is highly flexible and can access window metadata (such as time range, data volume), but it consumes a lot of resources.

[0127] In the embodiment of the present application, since full aggregation and incremental aggregation require different data for aggregation, incremental aggregation only focuses on the differences between the changed data. It can be understood that when there is newly written data, only the difference between the newly written data and the original data in the corresponding data shard is focused on, and then the difference is updated to the verification data of the verification shard to achieve incremental aggregation; while full aggregation focuses on all data. It can be understood that when there is newly written data, only the data in the latest data shards is focused on, and the original data before the new data is written is not focused on, and then the verification data of the verification shard is obtained again based on the latest data in each data shard.

[0128] If it is detected that after the data to be aggregated is written, all the data that can be pulled currently meets the data requirements corresponding to another aggregation mode, then it means that the aggregation mode can be converted at present. However, since data aggregating through different aggregation modes may require pulling different data, they all have certain overheads. If the overhead does not decrease after the aggregation mode conversion, then there is no need to convert the aggregation mode in the current scenario; if the overhead decreases after the aggregation mode conversion, it means that there is an overhead benefit in the current aggregation mode conversion. Therefore, when judging whether the aggregation mode conversion can be performed, the target data pulling overhead after the aggregation mode conversion can be calculated, and then the target data pulling overhead can be compared with the initial data pulling overhead before the aggregation mode conversion.

[0129] 604. When the target data pulling overhead is less than the initial data pulling overhead, data aggregation is performed according to the target aggregation mode through the target verification sharding.

[0130] In the embodiment of the present application, if the target data pulling overhead is less than the initial data pulling overhead, it means that after the aggregation mode conversion, the data pulling overhead is reduced, then it means that the aggregation mode conversion is profitable, so data aggregation can be performed according to the target aggregation mode, wherein data aggregation can be performed through the target verification shard, and the target verification shard can be any one of the p verification shards, which can be understood as Figure 7 The parity fragments P and / or Q in .

[0131] The embodiment of the present application provides a data aggregation method, which obtains the data to be aggregated and at least one data shard to be aggregated corresponding to the data to be aggregated; calculates the initial data pulling overhead based on the data to be aggregated, the original data stored in at least one data shard to be aggregated, and the data stored in at least one verification shard; when it is detected that the data to be aggregated meets the aggregation mode conversion conditions, calculates the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion, and the target aggregation mode is full aggregation or incremental aggregation; when the target data pulling overhead is less than the initial data pulling overhead, the data is aggregated according to the target aggregation mode through the target verification shard. In this solution, the aggregation mode that best suits the current data can be selected by converting the aggregation mode, which effectively reduces the overhead of data pulling during the data aggregation process.

[0132] like Figure 2 As shown, Figure 3 A flowchart of a data aggregation method is provided for an embodiment of the present application. The method may include the following steps:

[0133] 701. Obtain data to be aggregated and at least one data shard to be aggregated corresponding to the data to be aggregated.

[0134] In the embodiments of the present application, for the description of step 701, refer to the detailed description of step 601 in the above embodiments, and the embodiments of the present application will not be repeated.

[0135] 702. Determine an initial aggregation mode according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards.

[0136] In the embodiments of the present application, after determining the at least one to-be-aggregated data shard corresponding to the to-be-aggregated data, since the aggregation mode conversion may be involved, it is necessary to determine the initial aggregation mode, which includes full-amount aggregation or incremental aggregation.

[0137] In some embodiments, determining the initial aggregation mode according to the at least one to-be-aggregated data shard and the pre-stored plurality of data shards can specifically include: determining a first quantity of the at least one to-be-aggregated data shard and a second quantity of the pre-stored plurality of data shards; and determining the initial aggregation mode according to the first quantity and the second quantity.

[0138] It should be noted that the initial aggregation mode can be determined according to the proportion of the overwritten shard, that is, the proportion between the first quantity of the overwritten shard and the second quantity of the total data shard, the overwritten shard being the at least one to-be-aggregated data shard corresponding to the to-be-aggregated data, and the total data shard being all data shards corresponding to the strip data.

[0139] Further, for the first quantity and the second quantity, a proportion threshold value can be set, for example: 0.5, when the ratio of the first quantity and the second quantity is greater than the proportion threshold value, it can be determined that the initial aggregation mode is full-amount aggregation; when the ratio of the first quantity and the second quantity is less than or equal to the proportion threshold value, it can be determined that the initial aggregation mode is incremental aggregation.

[0140] For example, as shown in Figure 2 and Figure 3 , the second quantity of the total data shard is 8, the overwritten shard is data shard 6 and 7, the first quantity is 2, and 2 / 8=0.25<0.5, it can be determined that Figure 4 and Figure 5 the corresponding initial aggregation mode is incremental aggregation; as shown in Figure 4 and Figure 5 , the second quantity of the total data shard is 8, the overwritten shard is data shard 3, 4, 5, 6 and 7, the first quantity is 5, and 5 / 8=0.625>0.5, it can be determined that Figures 2-5 and Figures 2-5 the corresponding initial aggregation mode is full-amount aggregation.

[0141] 703、According to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard, perform pull overhead calculation according to the initial aggregation mode, to obtain initial data pull overhead corresponding to the initial aggregation mode.

[0142] In the embodiments of the present application, after the initial aggregation mode is determined, the overhead corresponding to the initial aggregation mode needs to be calculated. Since the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard may be used during aggregation, the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard can be used to perform data aggregation according to the initial aggregation mode, so as to calculate the initial data pull overhead corresponding to the initial aggregation mode. The process can be found in Figure 3 and the corresponding description.

[0143] In some embodiments, by Figure 3 As can be seen, in the scenario where a shard fails, since the original data of the failed shard needs to be derived reversely, a lot of data will be repeatedly pulled. In order to optimize the above steps, de-duplication optimization can be performed.

[0144] Specifically, according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard, performing pull overhead calculation according to the initial aggregation mode, to obtain initial data pull overhead corresponding to the initial aggregation mode, can specifically include: according to the to-be-aggregated data, the original data stored in the at least one to-be-aggregated data shard, and the data stored in the at least one check shard, determining a plurality of original data pull steps corresponding to the initial aggregation mode; performing de-duplication optimization on the data pulled by the plurality of original data pull steps, to obtain a plurality of initial data pull steps; and according to the plurality of initial data pull steps, calculating the initial data pull overhead.

[0145] The following will be described respectively for incremental aggregation and full aggregation.

[0146] It should be noted that when performing incremental aggregation, the plurality of original data pull steps before de-duplication optimization can be found in Figure 8 and the description. As can be seen, in the scenario where a shard fails, Figure 5In the example, since step (2) has already pulled data Q from the local, and step (3) has already pulled data P from the check shard P, then in steps (4) and (5), there is no need to pull data P from the check shard P again, nor is there any need to pull data Q locally again; and, in step (4), data A, B, C, D, E, F have been pulled from data shards 0, 1, 2, 3, 4, 5 respectively, then in step (5), there is no need to pull data A, B, C, D, E, F from data shards 0, 1, 2, 3, 4, 5 respectively; therefore, the above repeated steps can be omitted, and after deduplication optimization, the following can be obtained: Figure 5 Multiple initial data pulling steps shown:

[0147] (1) Pull data G', H' from the local;

[0148] (2) Pull data Q from the local machine;

[0149] (3) Pull data P from the verification shard P;

[0150] (4) Pull data A, B, C, D, E, and F from data shards 0, 1, 2, 3, 4, and 5 respectively.

[0151] In summary, the initial data fetching cost is: (G'+H')+P+Q+(A+B+C+D+E+F).

[0152] It should be noted that when performing full aggregation, the multiple steps of pulling raw data before deduplication optimization can be found in detail. Figure 9 And description. It can be seen that Figure 8 In the example, step (2) and step (3) are repeated to pull data A from data shard 0; and step (3) and step (4) are repeated to reversely deduce the original data B and C, respectively, to repeatedly pull data Q from the local, repeatedly pull data A from data shard 0, repeatedly pull historical data D, E, F, G, H from data shards 3, 4, 5, 6, 7, respectively, and repeatedly pull data P from the verification shard P; therefore, the above repeated steps can be omitted. After deduplication optimization, the following can be obtained: Figure 10 Multiple initial data pulling steps shown:

[0153] (1) Pull data D', E', F', G', H' from the local;

[0154] (2) Pull data Q from the local machine, pull data A from data shard 0, pull historical data D, E, F, G, and H from data shards 3, 4, 5, 6, and 7 respectively, and pull data P from the verification shard P.

[0155] In summary, the initial data fetching cost is: (D'+E'+F'+G'+H')+(A+D+E+F+G+H+P+Q).

[0156] 704. When the initial aggregation mode is full aggregation, if it is detected that the currently pullable data includes the data to be aggregated, the original data corresponding to at least one shard of the data to be aggregated, and the verification data of the target verification shard, it is determined that the data to be aggregated meets the aggregation mode conversion conditions, and the target aggregation mode after the aggregation mode conversion is incremental aggregation.

[0157] In an embodiment of the present application, when judging whether the aggregation mode can be converted, the judgment can be made based on the characteristics of the aggregation mode. When the initial aggregation mode is full aggregation, it is necessary to judge whether the current data can be converted into incremental aggregation. Since incremental aggregation can be simply understood as updating the verification data in the verification shard through the difference between the original data and the newly covered data to achieve aggregation, it is necessary to pay attention to whether the data to be aggregated can be obtained at present, as well as the original data of at least one data shard to be aggregated corresponding to the data to be aggregated.

[0158] It should be noted that when the initial aggregation mode is full aggregation, it is necessary to pull the data to be aggregated and the original data of all data shards except at least one data shard to be aggregated. If it can be converted to incremental aggregation, it is only necessary to pull the data to be aggregated, the original data corresponding to at least one data shard to be aggregated, and the verification data of the target verification shard. Therefore, the currently pullable data can be tested. If it is detected that the above-mentioned data to be aggregated, the original data corresponding to at least one data shard to be aggregated, and the verification data of the target verification shard can all be pulled, then it is considered that the current data can be converted to incremental aggregation, and then the overhead after conversion to incremental aggregation is calculated and compared with the overhead of the original full aggregation.

[0159] 705. When the initial aggregation mode is incremental aggregation, if it is detected that the currently pullable data includes the data to be aggregated and the original data corresponding to multiple data shards except at least one data shard to be aggregated, it is determined that the data to be aggregated meets the aggregation mode conversion conditions, and the target aggregation mode after the aggregation mode conversion is full aggregation.

[0160] In an embodiment of the present application, when judging whether the aggregation mode can be converted, the judgment can be made based on the characteristics of the aggregation mode. When the initial aggregation mode is incremental aggregation, it is necessary to judge whether the current data can be converted to full aggregation. Since full aggregation can be simply understood as aggregating all current data after overwriting, it is necessary to pay attention to whether the data to be aggregated can be obtained at present, as well as the original data of the data shards other than at least one data shard to be aggregated.

[0161] It should be noted that when the initial aggregation mode is incremental aggregation, the original data of the to-be-aggregated data, at least one to-be-aggregated data shard, and the check data of the target check shard need to be pulled, and if it can be converted into full aggregation, only the original data of the to-be-aggregated data and all data shards except at least one to-be-aggregated data shard need to be pulled, so the current pullable data can be detected, and if it is detected that the to-be-aggregated data and the original data of all data shards except at least one to-be-aggregated data shard can be pulled, it is considered that it can be converted into full aggregation, and then the overhead after conversion into full aggregation is calculated and compared with the overhead of original incremental aggregation.

[0162] 706、When it is detected that the to-be-aggregated data satisfies the aggregation mode conversion condition, the target data pulling overhead corresponding to the target aggregation mode after aggregation mode conversion is calculated.

[0163] In the embodiments of the present application, when it is determined that the current data supports aggregation mode conversion, it is also necessary to determine whether there is overhead benefit, that is, whether the data pulling overhead is reduced after aggregation mode conversion, and then the target data pulling overhead corresponding to the target aggregation mode after aggregation mode conversion needs to be calculated.

[0164] In some embodiments, when the target aggregation mode is full aggregation, in the multiple initial data pulling steps as shown in Figure 9 , step (2) of pulling data Q from the local does not need to be performed; in addition, step (3) of pulling data P from the check shard P does not need to be performed; therefore, after the aggregation mode is converted into full aggregation, the steps as shown in Figure 11 can be obtained.

[0165] (1) pulling data G' and H' from the local;

[0166] (2) pulling data A, B, C, D, E, and F from data shards 0, 1, 2, 3, 4, and 5, respectively.

[0167] In summary, the target data pulling overhead is (G'+H')+(A+B+C+D+E+F).

[0168] In some embodiments, when the target aggregation mode is incremental aggregation, in the multiple initial data pulling steps as shown in Figure 8 , step (2) of pulling data A from data shard 0 does not need to be performed; therefore, after the aggregation mode is converted into incremental aggregation, the steps as shown in Figure 10 can be obtained.

[0169] (1) pulling data D', E', F', G', and H' from the local;

[0170] (2) pull data Q from local, pull historical data D, E, F, G, H from data shards 3, 4, 5, 6, 7 respectively, and pull data P from check shard P.

[0171] In summary, the target data pulling overhead is (D'+E'+F'+G'+H')+(D+E+F+G+H+P+Q).

[0172] 707、When the target data pulling overhead is less than the initial data pulling overhead, perform data aggregation according to the target aggregation mode through the target check shard.

[0173] In the embodiments of the present application, in combination with Figure 9 and Figure 11 It can be seen that the initial data pulling overhead is (G'+H')+P+Q+(A+B+C+D+E+F), and the target data pulling overhead is (G'+H')+(A+B+C+D+E+F). The target data pulling overhead is less than the initial data pulling overhead. In combination with Figure 10 and Figure 11 It can be seen that the initial data pulling overhead is (D'+E'+F'+G'+H')+(A+D+E+F+G+H+P+Q), and the target data pulling overhead is (D'+E'+F'+G'+H')+(D+E+F+G+H+P+Q). The target data pulling overhead is also less than the initial data pulling overhead.

[0174] In some embodiments, when the target data pulling overhead is less than the initial data pulling overhead, data aggregation is performed according to the target aggregation mode through the target check shard. Specifically, when the target aggregation mode is full-amount aggregation, when the target data pulling overhead is less than the initial data pulling overhead, pull the data to be aggregated from the target check shard; pull at least one original data from at least one target data shard; and perform full-amount aggregation on the data to be aggregated and the at least one original data.

[0175] It should be noted that, in combination with Figure 12 , the data to be aggregated is pulled from the target check shard, that is, data G' and H' are pulled from local; at least one original data is pulled from at least one target data shard, that is, data A, B, C, D, E, and F are pulled from data shards 0, 1, 2, 3, 4, and 5 respectively. The at least one target data shard is all data shards in the pre-stored multiple data shards except the at least one data shard to be aggregated. Finally, new check data can be encoded through A+B+C+D+E+F+G'+H', that is, full-amount aggregation is completed.

[0176] In some embodiments, when the target data pulling overhead is less than the initial data pulling overhead, the data is aggregated according to the target aggregation mode by the target check shard, which can specifically include: when the target aggregation mode is incremental aggregation, when the target data pulling overhead is less than the initial data pulling overhead, pulling the to-be-aggregated data from the target check shard; pulling at least one original data from at least one to-be-aggregated data shard; pulling at least one check data from at least one check shard; and performing data incremental aggregation on the to-be-aggregated data, the at least one original data, and the at least one check data.

[0177] It should be noted that, for details Figure 13 , the to-be-aggregated data is pulled from the target check shard, that is, the data D', E', F', G', and H' is pulled locally; at least one original data is pulled from at least one to-be-aggregated data shard, that is, the historical data D, E, F, G, and H is pulled from the data shards 3, 4, 5, 6, and 7 respectively; at least one check data is pulled from at least one check shard, that is, the data Q is pulled locally and the data P is pulled from the check shard P; finally, the new check data can be encoded by diff((D'E'F'G'H'), (DEFGH))+P+Q, that is, the data incremental aggregation is completed.

[0178] As shown in ​ , the embodiment of the present application provides a data aggregation apparatus, which can include: an acquisition module 1201 configured to acquire to-be-aggregated data and at least one to-be-aggregated data shard corresponding to the to-be-aggregated data;

[0179] a processing module 1202 configured to calculate an initial data pulling overhead according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in at least one check shard;

[0180] The processing module 1202 is further configured to, when it is detected that the to-be-aggregated data satisfies an aggregation mode conversion condition, calculate a target data pulling overhead corresponding to a target aggregation mode after aggregation mode conversion, the target aggregation mode being full aggregation or incremental aggregation;

[0181] The processing module 1202 is further configured to, when the target data pulling overhead is less than the initial data pulling overhead, aggregate data according to the target aggregation mode by the target check shard.

[0182] In some embodiments, the processing module 1202 is specifically configured to determine an initial aggregation mode according to the at least one to-be-aggregated data shard and a plurality of pre-stored data shards, the initial aggregation mode being full aggregation or incremental aggregation;

[0183] The processing module 1202 is specifically configured to perform initial aggregation mode pull overhead calculation according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in the at least one check shard, to obtain initial data pull overhead corresponding to the initial aggregation mode.

[0184] In some embodiments, the processing module 1202 is specifically configured to determine a first quantity of the at least one to-be-aggregated data shard, and a second quantity of the plurality of pre-stored data shards.

[0185] The processing module 1202 is specifically configured to determine the initial aggregation mode according to the first quantity and the second quantity.

[0186] In some embodiments, the processing module 1202 is specifically configured to determine a plurality of original data pull steps corresponding to the initial aggregation mode according to the to-be-aggregated data, original data stored in the at least one to-be-aggregated data shard, and data stored in the at least one check shard.

[0187] The processing module 1202 is specifically configured to perform deduplication optimization on data pulled by the plurality of original data pull steps, to obtain a plurality of initial data pull steps.

[0188] The processing module 1202 is specifically configured to calculate the initial data pull overhead according to the plurality of initial data pull steps.

[0189] In some embodiments, the processing module 1202 is further configured to, when the initial aggregation mode is full aggregation, if it is detected that the current pullable data includes the to-be-aggregated data, original data corresponding to the at least one to-be-aggregated data shard, and check data of the target check shard, determine that the to-be-aggregated data satisfies the aggregation mode conversion condition, and the target aggregation mode after the aggregation mode conversion is incremental aggregation.

[0190] The processing module 1202 is further configured to, when the initial aggregation mode is incremental aggregation, if it is detected that the current pullable data includes the to-be-aggregated data and original data corresponding to data shards other than the at least one to-be-aggregated data shard among the plurality of data shards, determine that the to-be-aggregated data satisfies the aggregation mode conversion condition, and the target aggregation mode after the aggregation mode conversion is full aggregation.

[0191] In some embodiments, when the target aggregation mode is full aggregation, the processing module 1202 is specifically configured to, when the target data pull overhead is less than the initial data pull overhead, pull the to-be-aggregated data from the target check shard.

[0192] The processing module 1202 is specifically configured to pull at least one original data from at least one target data shard, the at least one target data shard being all data shards other than the at least one to-be-aggregated data shard among the plurality of pre-stored data shards.

[0193] The processing module 1202 is specifically configured to perform full data aggregation on the data to be aggregated and at least one original data.

[0194] In some embodiments, when the target aggregation mode is incremental aggregation, the processing module 1202 is specifically configured to pull the data to be aggregated from the target check shard when the target data pulling overhead is less than the initial data pulling overhead;

[0195] The processing module 1202 is specifically configured to pull at least one original data from at least one data shard to be aggregated;

[0196] The processing module 1202 is specifically configured to pull at least one verification data from at least one verification shard;

[0197] The processing module 1202 is specifically configured to perform data incremental aggregation on the data to be aggregated, at least one original data, and at least one verification data.

[0198] In the embodiment of the present application, each module can implement the data aggregation method provided in the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0199] like ​ As shown, an embodiment of the present application further provides an electronic device, which may include:

[0200] A memory 1301 storing executable program code;

[0201] a processor 1302 coupled to the memory 1301;

[0202] The processor 1302 calls the executable program code stored in the memory 1301 to execute the data aggregation method executed by the electronic device in the above-mentioned method embodiments.

[0203] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the data aggregation method in the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0204] An embodiment of the present application also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the data aggregation method in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0205] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0206] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a portion of code, and the module, program segment, or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.

[0207] In this application, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0208] In this application, memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0209] In the present application, those of ordinary skill in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by relevant hardware instructed by a program, and the program can be stored in a computer readable storage medium, including permanent and non-permanent, removable and non-removable storage medium. The storage medium can realize information storage by any method or technology, and the information can be computer readable instructions, data structure, program module or other data. Examples of computer storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), other types of random access memory (RAM), read-only memory (ROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, disk storage or other magnetic storage device or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, the computer readable medium does not include transitory computer readable media such as modulated data signals and carriers.

[0210] It should be noted that, in the specification, the terms "one embodiment" or "an embodiment" merely indicate a possible embodiment and do not necessarily specify a single or particular embodiment. Furthermore, the above-described features and characteristics can be combined in any suitable manner in various embodiments. One of ordinary skill in the art will recognize that the embodiments described herein can be practiced with various modifications and alterations, and equivalents thereof. Accordingly, the application should not be considered as limited to the specific recitations. Therefore, the above disclosure is not to be considered as limiting the scope of the application, but rather, is provided for illustrative purposes.

[0211] It should be understood that, throughout the specification, "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout the specification are not necessarily referring to the same embodiment. Further, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It will be appreciated by those skilled in the art that the embodiments described herein that can be implemented with or without employing any hardware and that the software that can be implemented operates on any platform capable of running a messaging application.

[0212] In various embodiments of the present application, it should be understood that the magnitude of the serial number of the above processes does not mean the inevitable sequence of execution, and the execution sequence of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0213] The units described above as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place, or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0214] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0215] The integrated units described above, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer accessible memory. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the present application or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of parts or all steps of the above-mentioned method for enabling a computer device (which can be a personal computer, a server or a network device, etc., and specifically can be a processor in the computer device) to execute the embodiments of the present application.

[0216] The above is only a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data aggregation method, characterized in that: The method comprises: Acquire data to be aggregated and at least one data shard to be aggregated corresponding to the data to be aggregated; Calculating an initial data pull overhead based on the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one check shard; When it is detected that the data to be aggregated meets the aggregation mode conversion condition, the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion is calculated, and the target aggregation mode is full aggregation or incremental aggregation; When the target data pulling overhead is less than the initial data pulling overhead, data aggregation is performed according to the target aggregation mode through the target verification sharding.

2. The method according to claim 1, characterized in that The calculating the initial data pulling overhead according to the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one check shard includes: Determining an initial aggregation mode according to the at least one data shard to be aggregated and the plurality of pre-stored data shards, wherein the initial aggregation mode is full aggregation or incremental aggregation; Based on the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one verification shard, the pulling overhead is calculated according to the initial aggregation mode to obtain the initial data pulling overhead corresponding to the initial aggregation mode.

3. The method according to claim 2, characterized in that The determining the initial aggregation mode according to the at least one data shard to be aggregated and the plurality of pre-stored data shards includes: Determining a first number of the at least one data shard to be aggregated and a second number of the plurality of pre-stored data shards; The initial aggregation mode is determined according to the first quantity and the second quantity.

4. The method according to claim 2, characterized in that The calculating the pull overhead according to the initial aggregation mode based on the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one check shard to obtain the initial data pull overhead corresponding to the initial aggregation mode includes: Determining a plurality of original data pulling steps corresponding to the initial aggregation mode according to the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one verification shard; performing deduplication optimization on the data pulled by the multiple original data pulling steps to obtain multiple initial data pulling steps; The initial data pulling overhead is calculated based on the multiple initial data pulling steps.

5. The method according to claim 1, wherein When it is detected that the data to be aggregated meets the aggregation mode conversion condition, before calculating the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion, the method further includes: When the initial aggregation mode is full aggregation, if it is detected that the currently pullable data includes the data to be aggregated, the original data corresponding to the at least one data shard to be aggregated, and the verification data of the target verification shard, it is determined that the data to be aggregated meets the aggregation mode conversion condition, and the target aggregation mode after the aggregation mode conversion is incremental aggregation; When the initial aggregation mode is incremental aggregation, if it is detected that the currently pullable data includes the data to be aggregated and the original data corresponding to multiple data shards other than the at least one data shard to be aggregated, it is determined that the data to be aggregated meets the aggregation mode conversion conditions, and the target aggregation mode after the aggregation mode conversion is full aggregation.

6. The method according to claim 1, characterized in that When the target aggregation mode is full aggregation, when the target data pulling overhead is less than the initial data pulling overhead, performing data aggregation according to the target aggregation mode through target verification shards, including: When the target data pulling overhead is less than the initial data pulling overhead, pulling the to-be-aggregated data from the target verification shard; Pulling at least one original data from at least one target data shard, where the at least one target data shard is all data shards from the pre-stored plurality of data shards except the at least one data shard to be aggregated; Full data aggregation is performed on the data to be aggregated and the at least one original data.

7. The method according to claim 1, characterized in that When the target aggregation mode is incremental aggregation, when the target data pulling overhead is less than the initial data pulling overhead, performing data aggregation according to the target aggregation mode through the target verification shards, including: When the target data pulling overhead is less than the initial data pulling overhead, pulling the to-be-aggregated data from the target verification shard; Pull at least one original data from the at least one data shard to be aggregated; Pulling at least one verification data from the at least one verification shard; Incremental data aggregation is performed on the data to be aggregated, the at least one original data, and the at least one verification data.

8. A data aggregation device, characterized in that: include: An acquisition module, configured to acquire data to be aggregated and at least one data fragment to be aggregated corresponding to the data to be aggregated; a processing module, configured to calculate an initial data pulling overhead based on the data to be aggregated, the original data stored in the at least one data shard to be aggregated, and the data stored in the at least one check shard; The processing module is further configured to calculate the target data pulling overhead corresponding to the target aggregation mode after the aggregation mode conversion when it is detected that the data to be aggregated meets the aggregation mode conversion condition, where the target aggregation mode is full aggregation or incremental aggregation; The processing module is further configured to perform data aggregation according to the target aggregation mode through target verification shards when the target data pulling overhead is less than the initial data pulling overhead.

9. An electronic device, characterized in that: include: a memory storing executable program code; and a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the data aggregation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that include: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the data aggregation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Multi-dimensional data cube increment aggregation and query optimization method

    CN102360379A

  • Data aggregation method, device and equipment and computer readable storage medium

    CN111881135A

  • Data additional aggregation method and device, electronic equipment and readable storage medium

    CN113031871A

  • Transaction processing method and device, computing equipment and storage medium

    CN115114344A

  • Method, device and equipment for reducing overhead of erasure code storage network and storage medium

    CN117931075A