Method and system for automatic capacity management of a distributed data analysis engine

By using management modules to divide data groups and formulate storage schedules in the distributed data analysis engine, using hidden storage methods and adjusting the number of node modules, the problem of low capacity management efficiency in the existing technology is solved, and the effective utilization and management of capacity is achieved.

CN119884254BActive Publication Date: 2025-06-27BEIJING YONGHONG SHANGZHI TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510368912.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing distributed data analysis engine is not efficient in capacity management and fails to effectively utilize capacity.

Method used

The management module obtains preset values, divides data groups and formulates storage plans, uses hidden storage methods to store data into node modules, and adjusts the number of node modules as needed to achieve automatic capacity management.

Benefits of technology

Efficient use of capacity, improve capacity management efficiency, and avoid data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884254B_ABST
    Figure CN119884254B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of distributed data analysis engines, and particularly to a method and system for automatic capacity management of a distributed data analysis engine. The method includes: S1, dividing different data to be stored into data groups with the number of characteristic values; S2, storing all the data in several data groups in corresponding node modules in a hidden storage manner; S3, determining whether it is necessary to adjust the number of currently used node modules; S4, if reducing the number of currently used node modules, transferring all the hidden data in the data groups with the difference in the number between the first ratio and the second ratio to the corresponding other node modules, otherwise, transferring all the hidden data in the data groups with the difference in the number between the third ratio and the fourth ratio to the corresponding node modules to be added; S5, performing recovery processing on the hidden data and using the data after the recovery processing for analysis processing. This application can effectively utilize the capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of distributed data analysis engines, and particularly to a method and system for automatically managing the capacity of a distributed data analysis engine. Background Art

[0002] A distributed data analysis engine refers to a system that can process and analyze large amounts of data in a distributed storage environment, and is widely used in scenarios such as big data processing, real-time data analysis, and machine learning.

[0003] The Chinese patent application with the publication number CN112099977A discloses a real-time data analysis engine for a distributed tracking system, including a data processing module, a data acquisition module connected to the data processing module, and a data analysis module. The data acquisition module and the data processing module are connected through a data receiving module. A Kafka cluster is used as an intermediate layer between the data receiving module and the data processing module. Asynchronous transmission is used between Kafka and the data receiving nodes in the data receiving module. Kafka counts the response time of each URL access within a specific time interval. The data processing module performs data pre-aggregation based on a time window, aggregates data by comparing the fields corresponding to the response time, and adds new data to the aggregated result. The result of the data pre-aggregation is stored in a Redis cache, and the data summary result is extracted and stored, while the data in the Redis cache is deleted. In addition, the Chinese patent application with the publication number CN111625696A provides a distributed scheduling method, a computing node device, a readable storage medium, a computing device, and a distributed scheduling system for a multi-source data analysis engine, which saves a large amount of communication costs and improves the distributed processing efficiency of the multi-source data analysis engine. The method includes: the computing node with the highest scheduling index of the multi-source data analysis engine receives a query task; the computing node determines a sub-query task including an intermediate result set of the query task and determines the storage node of the intermediate result set; the computing node calculates the first time cost of migrating the intermediate result set to the local, and calculates the second time cost of the storage node executing the sub-query task; the computing node selects whether to migrate the intermediate result set to the local by the computing node and execute the sub-query task or let the storage node execute the sub-query task according to the comparison result of the first time cost and the second time cost. However, neither of the above two patent applications considers how to effectively utilize the capacity, and the capacity management efficiency is not high. Summary of the Invention

[0004] The management module of the present application divides different data to be stored into data groups with the number of characteristic values, enabling each node module of the present application to evenly store the confidential data. And in the case where the number of currently used node modules needs to be adjusted, the management module of the present application performs capacity management in units of data groups. The present application aims to effectively utilize the capacity and improve the capacity management efficiency.

[0005] The present application provides a method for automatic capacity management of a distributed data analysis engine, including the following steps:

[0006] S1. The management module obtains a number of preset values, determines characteristic values based on the obtained number of preset values, and divides different data to be stored into data groups with the number of characteristic values;

[0007] S2. The management module formulates a storage schedule for different data to be stored based on the number of currently used node modules and the number of all data groups. The number of currently used node modules corresponds to one of the number of preset values. According to the storage schedule, all the data in several data groups are respectively stored in the corresponding node modules in a confidential storage manner;

[0008] S3. The management module determines whether it is necessary to adjust the number of currently used node modules. In the case of no, this step is repeated. In the case of yes, the next step is continued;

[0009] S4. In the case of reducing the number of currently used node modules, the management module calculates a first ratio of the characteristic value to the number of adjusted node modules, calculates a second ratio of the characteristic value to the number of node modules before adjustment, and transfers all the confidential data in the data groups with the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules respectively. In the case of increasing the number of currently used node modules, the management module calculates a third ratio of the characteristic value to the number of node modules before adjustment, calculates a fourth ratio of the characteristic value to the number of adjusted node modules, and transfers all the confidential data in the data groups with the difference between the third ratio and the fourth ratio from the currently used node modules to the corresponding node modules to be increased respectively;

[0010] S5. The extraction module extracts the stored confidential data from the corresponding node modules, performs recovery processing on the confidential data, and the analysis module uses the data after the recovery processing for analysis processing.

[0011] As a preferred technical solution of the present application, in S4, in the case of reducing the number of currently used node modules, the number of reduced node modules corresponds to one of the number of preset values.

[0012] As a preferred technical solution of the present application, in S4, when increasing the number of node modules currently in use, the increased number of node modules corresponds to one of several preset values.

[0013] As a preferred technical solution of the present application, the management module stores all the data in several data groups into the corresponding node modules in a hidden storage manner, including the following steps:

[0014] S21. For each data in several data groups, the management module performs hidden processing on the data to obtain hidden data;

[0015] S22. For the hidden data to be stored corresponding to each data in several data groups, the management module compares the hidden data to be stored with the different hidden data already stored in the corresponding node modules respectively;

[0016] S23. For the hidden data to be stored corresponding to each data in several data groups, the management module determines whether there is any stored hidden data that has the same part as the hidden data to be stored. If so, it performs storage processing on the different part of the hidden data to be stored. If not, it performs storage processing on the hidden data to be stored.

[0017] As a preferred technical solution of the present application, in S21, the management module performs hidden processing on the data to obtain hidden data, including the following steps:

[0018] S211. The management module divides the data into several partial data, and the management module sequentially generates a sequence number for each partial data according to the order of obtaining the several partial data;

[0019] S212. The management module uses all the other partial data except the partial data corresponding to the largest sequence number among the several partial data to form intermediate data, and performs a first arithmetic operation on the intermediate data to obtain a first arithmetic result data;

[0020] S213. The management module performs a second arithmetic operation on the first arithmetic result data and the partial data corresponding to the largest sequence number among the several partial data to obtain a second arithmetic result data, and generates scrambled data according to the second arithmetic result data;

[0021] S214. The management module performs a second arithmetic process on the obfuscated data and the intermediate data to obtain first hidden data, and encrypts the partial data corresponding to the largest sequence number among several partial data to obtain second hidden data, and connects the first hidden data and the second hidden data to obtain hidden data.

[0022] As a preferred technical solution of this application, in S22, the management module compares the hidden data to be stored with the different hidden data already stored in the corresponding node modules respectively, including the following steps:

[0023] S221. For each hidden data already stored in the corresponding node module, the management module performs a splitting process on the hidden data to obtain first partial data and second partial data, and performs the same splitting process on the hidden data to be stored to obtain third partial data and fourth partial data;

[0024] S222. The management module determines whether the first partial data is the same as the third partial data. If so, the third partial data is regarded as the same part of the hidden data to be stored, and the fourth partial data is regarded as the different part of the hidden data to be stored.

[0025] As a preferred technical solution of this application, the extraction module performs a recovery process on the hidden data, including the following steps:

[0026] S51. The extraction module performs a splitting process on the hidden data to obtain partial hidden data one and partial hidden data two, and continues to perform a decryption process on the partial hidden data two to obtain decryption result data;

[0027] S52. The extraction module generates mixed data according to the decryption result data, and performs a second arithmetic process on the mixed data and the partial hidden data one to obtain first recovery data;

[0028] S53. The extraction module performs a first arithmetic process on the first recovery data to obtain intermediate recovery data, and continues to perform a second arithmetic process on the intermediate recovery data and the decryption result data to obtain second recovery data, and connects the first recovery data and the second recovery data to obtain the data after the recovery process.

[0029] As a preferred technical solution of the present application, when the extraction module extracts the stored hidden data from the corresponding node module, it also extracts the metadata of the hidden data. Extracting the metadata of the hidden data includes: determining whether the number of metadata in the memory exceeds a preset number threshold. If it does not exceed, the metadata of the hidden data is extracted. If it exceeds, several metadata that have not been accessed in the last 7 days are unloaded from the memory, and then continue to determine whether the number of metadata in the memory exceeds the preset number threshold. If it does not exceed, the metadata of the hidden data is extracted. If it exceeds, several metadata are determined using the LRU algorithm, the determined several metadata are unloaded from the memory, and then the metadata of the hidden data is extracted.

[0030] The present application also provides a capacity automatic management system for a distributed data analysis engine, including the following modules:

[0031] A management module, configured to obtain several preset values, determine characteristic values based on the obtained several preset values, and divide different data to be stored into data groups with the number of characteristic values; and at the same time, based on the number of currently used node modules and the number of all data groups, formulate a storage schedule for different data to be stored. The number of currently used node modules corresponds to one of the several preset values, and all data in several data groups are respectively stored in the corresponding node modules in a hidden storage manner according to the storage schedule; and it is also configured to determine whether it is necessary to adjust the number of currently used node modules; and it is further configured to, in the case of reducing the number of currently used node modules, calculate a first ratio of the characteristic value to the number of adjusted node modules, calculate a second ratio of the characteristic value to the number of node modules before adjustment, and transfer all hidden data in the data groups with the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules respectively. In the case of increasing the number of currently used node modules, calculate a third ratio of the characteristic value to the number of node modules before adjustment, calculate a fourth ratio of the characteristic value to the number of adjusted node modules, and transfer all hidden data in the data groups with the difference between the third ratio and the fourth ratio from the currently used node modules to the corresponding node modules to be increased respectively;

[0032] A node module, configured to store the hidden data corresponding to the data to be stored;

[0033] An extraction module, configured to extract the stored hidden data from the corresponding node module and perform recovery processing on the hidden data;

[0034] An analysis module, configured to perform analysis processing using the data after the recovery processing by the extraction module.

[0035] Compared with the prior art, the beneficial effects of the present application are at least as follows:

[0036] In the technical solution provided by the present application, first, a number of preset values are obtained, and characteristic values are determined based on the obtained preset values. Different data to be stored are divided into data groups with the number of characteristic values. Secondly, based on the number of currently used node modules and the number of all data groups, a storage plan for different data to be stored is formulated. According to the storage plan, all data in several data groups are stored in the corresponding node modules in a hidden storage manner. Thirdly, in the case of needing to reduce the number of currently used node modules, the first ratio of the characteristic value to the adjusted number of node modules is calculated, and the second ratio of the characteristic value to the number of node modules before adjustment is calculated. All hidden data in the data groups with the difference in the number of the first ratio and the second ratio are transferred from the node modules to be reduced to the corresponding other node modules. In the case of needing to increase the number of currently used node modules, the third ratio of the characteristic value to the number of node modules before adjustment is calculated, and the fourth ratio of the characteristic value to the adjusted number of node modules is calculated. All hidden data in the data groups with the difference in the number of the third ratio and the fourth ratio are transferred from the currently used node modules to the corresponding node modules to be increased. Finally, the stored hidden data is extracted from the corresponding node modules, the extracted hidden data is subjected to recovery processing, and the data after the recovery processing is used for analysis processing. Through the present application, not only can the capacity be effectively utilized and the capacity management efficiency be improved, but also the problem of data leakage can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a flowchart of a method for automatic capacity management of a distributed data analysis engine in the present application;

[0039] Figure 2 It is a schematic diagram of a system for automatic capacity management of a distributed data analysis engine in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The embodiments of the present application provide a method and system for automatic capacity management of a distributed data analysis engine. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0041] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 , a method for automatic capacity management of a distributed data analysis engine in the embodiments of the present application mainly includes the following steps:

[0042] S1. The management module obtains a number of preset values, determines a characteristic value based on the obtained number of preset values, and divides different data to be stored into data groups with the number of characteristic values;

[0043] S2. The management module formulates a storage plan for different data to be stored based on the number of currently used node modules and the number of all data groups. The number of currently used node modules corresponds to one of the number of preset values. According to the storage plan, all data in several data groups are respectively stored in the corresponding node modules in a hidden storage manner;

[0044] S3. The management module determines whether it is necessary to adjust the number of currently used node modules. If not, this step is repeated. If so, continue to the next step;

[0045] S4. In the case of reducing the number of currently used node modules, the management module calculates a first ratio of the characteristic value to the number of adjusted node modules, calculates a second ratio of the characteristic value to the number of node modules before adjustment, and transfers all hidden data in the data groups with the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules respectively. In the case of increasing the number of currently used node modules, the management module calculates a third ratio of the characteristic value to the number of node modules before adjustment, calculates a fourth ratio of the characteristic value to the number of adjusted node modules, and transfers all hidden data in the data groups with the difference between the third ratio and the fourth ratio from the currently used node modules to the corresponding node modules to be increased respectively;

[0046] S5. The extraction module extracts the stored hidden data from the corresponding node module, performs recovery processing on the hidden data, and the analysis module uses the data after the recovery processing for analysis processing.

[0047] Specifically, the distributed data analysis engine usually works on a distributed storage environment. Since the capacity of the distributed storage environment is limited, the capacity should be utilized effectively. Therefore, S1 to S5 are proposed.

[0048] In S1, the management module obtains several preset values, and determines a characteristic value based on the obtained preset values. It should be noted that the preset values ​​correspond to the number of node modules. For ease of understanding, for example, if the preset values ​​are 3, 6, and 9, then a common multiple of 3, 6, and 9, 18, can be used as a characteristic value. The management module also divides the different data to be stored into data groups with the same number of characteristic values. There is no restriction on the specific division method, as long as the different data to be stored can be evenly divided into data groups with the same number of characteristic values. In S2, the management module formulates a storage plan for different data to be stored based on the number of node modules currently in use and the number of all data groups, wherein the number of node modules currently in use corresponds to one of several preset values. Following the above example, for example, the number of node modules currently in use is 6, and the storage plan records which different data groups each node module should store. Following the above example, for example, the node module with ID 1 to the node module with ID 6 respectively stores any 3 different data groups. The management module stores all data in several data groups in the corresponding node modules in a secret storage manner according to the storage plan. The specific process will be described below. So far, the different data to be stored are stored evenly in each node module, and the capacity can be effectively utilized. In S3, the management module determines whether it is necessary to adjust the number of node modules currently in use. For example, when the total amount of data to be stored increases, in order to effectively utilize the capacity and avoid the problem of centralized storage, the number of node modules must be increased. When the total amount of data to be stored decreases, in order to effectively utilize the capacity and avoid the problem of overly dispersed storage, the number of node modules must be reduced. If no adjustment is required, repeat this step. If adjustment is required, continue to S4. In S4, if the number of node modules currently in use is reduced, the management module calculates a first ratio of the characteristic value to the number of node modules after adjustment, calculates a second ratio of the characteristic value to the number of node modules before adjustment, and transfers all secret data in the data group of the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules. If the number of node modules currently in use is increased, the management module calculates a third ratio of the characteristic value to the number of node modules before adjustment, calculates a fourth ratio of the characteristic value to the number of node modules after adjustment, and transfers all secret data in the data group of the difference between the third ratio and the fourth ratio from the node modules currently in use to the corresponding node modules to be added. Through this step, regardless of whether the number of node modules needs to be increased or decreased, different node modules can always store data in a balanced manner, thereby effectively utilizing capacity. In addition, when the number of node modules currently in use needs to be adjusted, capacity management is performed in units of data groups, which can improve capacity management efficiency.In S5, the extraction module extracts the required hidden data from the corresponding node modules, performs recovery processing on the extracted hidden data, and the specific process will be described below. After that, the analysis module uses the data after recovery processing for analysis processing, which can be implemented using algorithms in the prior art, so it will not be elaborated here.

[0049] Further, in S4, when reducing the number of currently used node modules, the reduced number of node modules corresponds to one of several preset values.

[0050] Specifically, if it is necessary to reduce the number of currently used node modules, the reduced number of node modules corresponds to one of several preset values. Continuing with the above example, the reduced number of node modules can be 3. This ensures that when it is necessary to adjust the number of currently used node modules, capacity management can be carried out in units of data groups.

[0051] Further, in S4, when increasing the number of currently used node modules, the increased number of node modules corresponds to one of several preset values.

[0052] Specifically, if it is necessary to increase the number of currently used node modules, the increased number of node modules corresponds to one of several preset values. Continuing with the above example, the increased number of node modules can be 9. This ensures that when it is necessary to adjust the number of currently used node modules, capacity management can be carried out in units of data groups.

[0053] Further, the management module stores all the data in several data groups into the corresponding node modules in a hidden storage manner, including the following steps:

[0054] S21. For each data in several data groups, the management module performs hidden processing on the data to obtain hidden data;

[0055] S22. For the hidden data to be stored corresponding to each data in several data groups, the management module compares the hidden data to be stored with the different hidden data already stored in the corresponding node modules respectively;

[0056] S23. For the hidden data to be stored corresponding to each data in several data groups, the management module determines whether there is hidden data already stored that has the same part as the hidden data to be stored. In the case of yes, storage processing is performed on the different part of the hidden data to be stored. In the case of no, storage processing is performed on the hidden data to be stored.

[0057] Specifically, the process of the management module storing all the data in several data groups into the corresponding node modules in a hidden storage manner is introduced. In S21, for each data in several data groups, the management module performs hidden processing on the data to obtain hidden data. The process of hidden processing will be described below. Instead of directly storing the data, storing the corresponding hidden data of the data can avoid the problem of data leakage. In S22, for the hidden data to be stored corresponding to each data in several data groups, the management module respectively compares the hidden data to be stored with the different hidden data already stored in the corresponding node modules. The corresponding node module refers to the node module that will store the hidden data to be stored. The process of comparison processing will be described below. In S23, for the hidden data to be stored corresponding to each data in several data groups, the management module determines whether there is hidden data already stored that has the same part as the hidden data to be stored. If so, the management module performs storage processing on the different part of the hidden data to be stored. If not, the management module directly performs storage processing on the hidden data to be stored. Through the above method, in the case where there is hidden data already stored that has the same part as the hidden data to be stored, the amount of data to be stored can be reduced, thereby effectively utilizing the capacity. It should be noted that when storing, the different part of the hidden data to be stored can be stored corresponding to the corresponding already stored hidden data. The purpose of doing this is that when subsequent comparison processing is required, the complete hidden data to be stored can be obtained based on the corresponding already stored hidden data for comparison processing. The corresponding already stored hidden data can be complete or partial, that is, stored together corresponding to the complete other hidden data.

[0058] Further, the management module performs hidden processing on the data to obtain hidden data, including the following steps:

[0059] S211. The management module divides the data into several partial data, and the management module sequentially generates a sequence number for each partial data according to the order of obtaining the several partial data;

[0060] S212. The management module uses all the other partial data except the partial data corresponding to the largest sequence number in the several partial data to form intermediate data, and performs first operation processing on the intermediate data to obtain first operation result data;

[0061] S213. The management module performs second operation processing on the first operation result data and the partial data corresponding to the largest sequence number in the several partial data to obtain second operation result data, and generates confusion data according to the second operation result data;

[0062] S214. The management module performs a second operation process on the obfuscated data and the intermediate data to obtain the first hidden data, and performs an encryption process on the partial data corresponding to the largest sequence number among several partial data to obtain the second hidden data, and connects the first hidden data and the second hidden data to obtain the hidden data.

[0063] Specifically, the process of the management module performing hidden processing on the data to obtain the hidden data is introduced. In S211, the management module performs a division process, dividing the data into several partial data. It should be noted that the data scales of the several partial data are the same. According to the order of obtaining the several partial data, a sequence number is generated for each partial data in turn. For the sake of easy understanding, for example, if the data A is divided in the order from left to right, partial data A1, partial data A2, and partial data A3 are obtained respectively, then the sequence numbers of partial data A1 to partial data A3 are 1 to 3 in turn. In S212, the management module uses all the other partial data except the partial data corresponding to the largest sequence number among the several partial data to form the intermediate data. Continuing with the above example, that is, using partial data A1 to partial data A2 to form the intermediate data, and performing a first operation process on the intermediate data to obtain the first operation result data. The data scale of the first operation result data is the same as that of the partial data. The first operation process can be one of the sha series algorithms, mainly selected according to the length of the output data of each algorithm in the sha series algorithms. In S213, the management module performs a second operation process on the first operation result data and the partial data corresponding to the largest sequence number among the several partial data to obtain the second operation result data. The data scale of the second operation result data is the same as that of the partial data. For the sake of easy understanding of the second operation process, for example, performing a second operation process on "0101" and "1111" can obtain "1010". The management module generates obfuscated data according to the second operation result data. The data scale of the obfuscated data is the same as that of the intermediate data. Similarly, one of the sha series algorithms can also be used to generate the obfuscated data. The input data is the second operation result data, and the output data is the obfuscated data, also mainly selected according to the length of the output data of each algorithm in the sha series algorithms. In S214, the management module performs a second operation process on the obfuscated data and the intermediate data to obtain the first hidden data. The management module performs an encryption process on the partial data corresponding to the largest sequence number among the several partial data to obtain the second hidden data. The data scale of the second hidden data is the same as that of the partial data. For example, the aes algorithm can be used for encryption processing. Connecting the first hidden data and the second hidden data can obtain the hidden data. Through the above method, the hidden data corresponding to the data can be generated to protect the security of the data.

[0064] Further, the management module performs a comparison process on the confidential data to be stored with the different confidential data already stored in the corresponding node modules, including the following steps:

[0065] S221. For each piece of confidential data already stored in the corresponding node module, the management module performs a splitting process on the confidential data to obtain a first part of data and a second part of data, and performs the same splitting process on the confidential data to be stored to obtain a third part of data and a fourth part of data;

[0066] S222. The management module determines whether the first part of data is the same as the third part of data. If so, the third part of data is regarded as the same part of the confidential data to be stored, and the fourth part of data is regarded as the different part of the confidential data to be stored.

[0067] Specifically, the process of the management module performing a comparison process on the confidential data to be stored with the different confidential data already stored in the corresponding node modules is introduced. In S221, for each piece of confidential data already stored in the corresponding node module, the management module performs a splitting process on the confidential data. It should be noted that the splitting process is the same as the above-mentioned partitioning process, so it will not be elaborated here. Subsequently, the management module obtains a first part of data and a second part of data. The position of the first part of data in the already stored confidential data is equivalent to the position of the above-mentioned intermediate data in the above-mentioned data, and the second part of data is the part other than the first part of data. The management module also performs the same splitting process on the confidential data to be stored to obtain a third part of data and a fourth part of data. Similarly, the position of the third part of data in the confidential data to be stored is equivalent to the position of the above-mentioned intermediate data in the above-mentioned data, and the fourth part of data is the part other than the third part of data. In S222, the management module determines whether the first part of data is the same as the third part of data. If they are the same, the third part of data is regarded as the same part of the confidential data to be stored, and the fourth part of data is regarded as the different part of the confidential data to be stored. If they are not the same, it is determined that there is no same part between the confidential data to be stored and the already stored confidential data. It should be noted that only the first part of data and the third part of data are compared, and the second part of data and the fourth part of data are not compared because even if the original part of data corresponding to the second part of data is the same as the original part of data corresponding to the fourth part of data, if the key used to generate the second part of data is different from the key used to generate the fourth part of data, then the second part of data and the fourth part of data are not the same.

[0068] Further, the extraction module performs a restoration process on the confidential data, including the following steps:

[0069] S51. The extraction module performs segmentation processing on the hidden data to obtain a part of the hidden data one and a part of the hidden data two, and continues to perform decryption processing on the part of the hidden data two to obtain the decrypted result data;

[0070] S52. The extraction module generates mixed data according to the decrypted result data, and performs a second operation process on the mixed data and the part of the hidden data one to obtain the first restored data;

[0071] S53. The extraction module performs a first operation process on the first restored data to obtain the intermediate restored data, and continues to perform a second operation process on the intermediate restored data and the decrypted result data to obtain the second restored data, and connects the first restored data and the second restored data to obtain the data after restoration processing.

[0072] Specifically, when introducing the process of the extraction module for restoring the extracted hidden data, it should be noted that if the extracted hidden data is partial, then before the restoration process, it needs to be supplemented. In S51, the extraction module performs segmentation processing on the hidden data to obtain a part of the hidden data one and a part of the hidden data two. Similarly, the position of the part of the hidden data one in the hidden data is equivalent to the position of the above intermediate data in the above data, and the part of the hidden data two is the part except the part of the hidden data one. The extraction module continues to perform decryption processing on the part of the hidden data two to obtain the decrypted result data. For example, the aes algorithm can be used for decryption processing. In S52, the extraction module generates mixed data according to the decrypted result data. It should be noted that the method of generating mixed data from the decrypted result data is exactly the same as the method of generating obfuscated data from the second operation result data, that is, the same one in the sha series algorithms can be used. The extraction module performs a second operation process on the mixed data and the part of the hidden data one to obtain the first restored data. In S53, the extraction module performs a first operation process on the first restored data to obtain the intermediate restored data. The specific process is exactly the same as the process of performing the first operation process on the above intermediate data, that is, the same one in the sha series algorithms can be used. The extraction module continues to perform a second operation process on the intermediate restored data and the decrypted result data to obtain the second restored data. Thus, by connecting the first restored data and the second restored data, the data after restoration processing can be obtained. Through the above method, the data for generating the hidden data can be obtained.

[0073] Furthermore, while extracting the stored secret data from the corresponding node module, the extraction module also extracts the metadata of the secret data. Extracting the metadata of the secret data includes: determining whether the number of metadata in the memory exceeds a preset quantity threshold. If it does not exceed, the metadata of the secret data is extracted. If it exceeds, a number of metadata that have not been accessed in the last 7 days are unloaded from the memory, and then it is continued to determine whether the number of metadata in the memory exceeds the preset quantity threshold. If it does not exceed, the metadata of the secret data is extracted. If it exceeds, a number of metadata are determined using the LRU algorithm, the determined number of metadata are unloaded from the memory, and then the metadata of the secret data is extracted.

[0074] Specifically, as the amount of stored data continuously increases, the metadata will also continue to grow. Generally, the data analysis engine needs to load the metadata to support the analysis and calculation of the data. However, loading too much metadata will cause problems such as slow loading and excessive occupation of memory resources, and will also affect the subsequent data extraction performance. When the extraction module extracts the stored secret data from the corresponding node module, it also extracts the metadata of the secret data. The metadata helps with analysis and processing. When extracting the metadata of the secret data, a method of dynamically loading and unloading metadata can be considered to solve the above problems. The specific approach is as follows: First step, determine whether the number of metadata in the memory exceeds a preset quantity threshold. The quantity threshold is set according to the actual application scenario. If it does not exceed, the metadata of the secret data is extracted. If it exceeds, a number of metadata that have not been accessed in the last 7 days are unloaded from the memory. Second step, continue to determine whether the number of metadata in the memory exceeds the preset quantity threshold. If it does not exceed, the metadata of the secret data is extracted. If it exceeds, a number of metadata are determined using the LRU algorithm. The LRU algorithm determines a number of metadata that have been recently accessed but with fewer access times. The LRU algorithm is a prior art and will not be elaborated here. The determined number of metadata are unloaded from the memory, and then the metadata of the secret data is extracted.

[0075] In addition to the above, the distributed storage system in this application is also introduced. The system adopts a three-level capacity control strategy of "system - tenant - data set". System level: A global upper limit is set based on the total amount of physical resources. When the remaining capacity can only meet the usage for 2 - 3 days, a limit is triggered. This design takes into account the impact of the redundancy strategy on the available capacity and avoids storage overload through dynamic adjustment. Tenant level: The capacity of a single tenant does not exceed 2 / 3 of the total capacity, and a key - value pair configuration file is used to achieve flexible adjustment. This decentralized management mode is consistent with the tenant isolation feature of the distributed storage engine and also draws on the idea of dividing the storage space according to business scenarios in the sharding technology. Model level: The number of data rows is limited to 5 million to control the resource occupancy of a single analysis model. This fine - grained control is similar to the optimization strategy of columnar storage in a distributed database and also conforms to the practice of optimizing storage efficiency through wide - table design in the Hadoop ecosystem.

[0076] The system implements a three - level monitoring system. Storage space warning: The remaining space of the root directory or the cluster is detected every 15 minutes, and the threshold is configurable. This mechanism is similar to the capacity monitoring module of HDFS, but adds fine - grained detection at the tenant level, similar to the storage - layer monitoring design of Kudu. Data row warning: A row - number threshold is set for the data set. This design refers to the way of controlling the data scale in a columnar storage engine and combines the row - level compression technology of a distributed database. Alarm classification: Alarms are distinguished at the system / tenant / data set three levels, forming a three - level capacity limit to form a three - dimensional control network. This hierarchical alarm system has something in common with the cluster health status monitoring mechanism of Elasticsearch, but innovatively includes business scenarios (tenant directory, single data set) in the trigger conditions.

[0077] In addition to the above, the distributed data analysis engine in this application is also introduced. To improve the concurrent processing ability and analysis and calculation performance, a distributed data analysis engine generally needs to cache data and partial results in memory. However, caching will inevitably consume memory resources. Therefore, it is very necessary to limit the capacity of the cache.

[0078] Separate cache pools, namely DataContainer and ResultContainer, are set for data cache (DataCache) and result cache (ResultCache). Interface methods add, remove, and touch are defined to add, delete, and access cache objects. The size of the cache pool is automatically calculated based on the maximum available memory of the application. The specific calculation method is as follows: The data cache is defaulted to 1 / 3 of the maximum available memory of the application, and the result cache is defaulted to 1 / 6 of the maximum available memory of the application. To flexibly adapt to various scenarios, the size of the cache pool also supports user - defined configuration through system properties. The cache key is designed to include the dimensions and metrics of the analysis and calculation, and different analysis and calculations generate different cache keys.

[0079] The cache pool internally uses a linked list structure to store cache objects. This linked list supports the LRU algorithm and is thus defined as LRULink. In this linked list, the most recently accessed object is placed at the tail end of the linked list, and the least frequently accessed object is placed at the head end of the linked list. The objects in the linked list all implement a unified interface ISerializable. The main methods in the interface are complete, dispose, empty, and fill. Among them: The complete method is used to add the prepared object to LRULink, while the dispose method deletes the object from LRULink and releases it from memory.

[0080] To ensure that the size of the cache pool does not exceed the allocated capacity limit, when the free memory in the cache pool is less than the ISerializable object to be added, an operation to swap to disk will be triggered. The specific method is as follows: Traverse the objects starting from the head end of LRULink, call the empty method of the object to write the data to disk, and then remove the object from LRULink to release the occupied memory. When the released memory is greater than the object to be added, stop traversing and perform the addition operation. When the object swapped to disk is accessed again, the fill method needs to be called to read the data from disk. After the reading is completed, it needs to be added back to LRULink. Before adding, it is also necessary to check whether the free capacity of the cache pool is greater than the size of the object to be added.

[0081] According to another aspect of the embodiments of the present application, as shown in Figure 2 the present application also provides a capacity automatic management system for a distributed data analysis engine, including a management module, a node module, an extraction module, and an analysis module, to implement a capacity automatic management method for a distributed data analysis engine described above.

[0082] The functions of each module in the system are as follows:

[0083] A management module, which is used to obtain a number of preset values, determine characteristic values based on the obtained preset values, and divide different data to be stored into data groups with the number of characteristic values; and is also used to formulate a storage plan for different data to be stored based on the number of currently used node modules and the number of all data groups. The number of currently used node modules corresponds to one of the preset values. All data in several data groups are stored in the corresponding node modules in a hidden storage manner according to the storage plan; and it is used to determine whether it is necessary to adjust the number of currently used node modules; and is also used to calculate a first ratio of the characteristic value to the number of node modules after adjustment and a second ratio of the characteristic value to the number of node modules before adjustment in the case of reducing the number of currently used node modules, and transfer all the hidden data in the data groups with the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules respectively. In the case of increasing the number of currently used node modules, calculate a third ratio of the characteristic value to the number of node modules before adjustment and a fourth ratio of the characteristic value to the number of node modules after adjustment, and transfer all the hidden data in the data groups with the difference between the third ratio and the fourth ratio from the currently used node modules to the corresponding node modules to be added respectively;

[0084] Node modules, which are used to store the hidden data corresponding to the data to be stored;

[0085] An extraction module, which is used to extract the stored hidden data from the corresponding node modules and perform recovery processing on the hidden data;

[0086] An analysis module, which is used to perform analysis processing on the data after the recovery processing by the extraction module.

[0087] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0088] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0089] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A method for automatically managing the capacity of a distributed data analysis engine, characterized in that: The method comprises the following steps: S1. The management module obtains a number of preset values, determines a characteristic value based on the obtained preset values, and divides different data to be stored into data groups with the same number of characteristic values, where the characteristic value is a common multiple of the preset values; S2, the management module formulates a storage plan table for different data to be stored based on the number of node modules currently in use and the number of all data groups, the number of node modules currently in use corresponds to one of several preset values, and stores all data in several data groups in the corresponding node modules in a secret storage manner according to the storage plan table; S3, the management module determines whether the number of currently used node modules needs to be adjusted, and if not, repeats this step, and if yes, proceeds to the next step; S4. In the case of reducing the number of node modules currently in use, the management module calculates a first ratio of the characteristic value to the number of node modules after adjustment, calculates a second ratio of the characteristic value to the number of node modules before adjustment, and transfers all secret data in the data group of the difference between the first ratio and the second ratio from the node modules to be reduced to the corresponding other node modules; in the case of increasing the number of node modules currently in use, the management module calculates a third ratio of the characteristic value to the number of node modules before adjustment, calculates a fourth ratio of the characteristic value to the number of node modules after adjustment, and transfers all secret data in the data group of the difference between the third ratio and the fourth ratio from the node modules currently in use to the corresponding node modules to be added; S5. The extraction module extracts the stored secret data from the corresponding node module, performs recovery processing on the secret data, and the analysis module uses the recovered data for analysis processing.

2. The method according to claim 1, characterized in that In the above S4, when the number of currently used node modules is reduced, the number of node modules after reduction corresponds to one of several preset values.

3. The method according to claim 2, characterized in that In the above S4, when the number of currently used node modules is increased, the increased number of node modules corresponds to one of several preset values.

4. The method according to claim 3, characterized in that The management module stores all data in several data groups in corresponding node modules in a secret storage manner, including the following steps: S21, for each data in the plurality of data groups, the management module performs secret processing on the data to obtain secret data; S22, with respect to the secret data to be stored corresponding to each data in the plurality of data groups, the management module compares the secret data to be stored with the different secret data stored in the corresponding node module; S23. With respect to the secret data to be stored corresponding to each data in several data groups, the management module determines whether there is stored secret data that has the same part as the secret data to be stored. If so, storage processing is performed on different parts of the secret data to be stored; if not, storage processing is performed on the secret data to be stored.

5. The method according to claim 4, characterized in that In S21, the management module performs confidentiality processing on the data to obtain confidential data, including the following steps: S211, the management module divides the data into a plurality of partial data, and the management module sequentially generates a sequence number for each partial data in the order in which the plurality of partial data are obtained; S212, the management module uses all the partial data except the partial data corresponding to the largest sequence number among the plurality of partial data to form intermediate data, and performs a first operation processing on the intermediate data to obtain first operation result data; S213, the management module performs a second operation on the first operation result data and the partial data corresponding to the largest sequence number among the plurality of partial data to obtain second operation result data, and generates obfuscated data according to the second operation result data; S214, the management module performs a second operation on the obfuscated data and the intermediate data to obtain the first secret data, and performs encryption on the partial data corresponding to the largest sequence number among the partial data to obtain the second secret data, and connects the first secret data and the second secret data to obtain the secret data.

6. The method according to claim 5, characterized in that In S22, the management module compares the secret data to be stored with the different secret data stored in the corresponding node modules, including the following steps: S221, with respect to each secret data stored in the corresponding node module, the management module performs segmentation processing on the secret data to obtain a first portion of data and a second portion of data, and performs the same segmentation processing on the secret data to be stored to obtain a third portion of data and a fourth portion of data; S222, the management module determines whether the first part of data is the same as the third part of data. If yes, the third part of data is regarded as the same part of the secret data to be stored, and the fourth part of data is regarded as the different part of the secret data to be stored.

7. The method according to claim 6, characterized in that The extraction module performs recovery processing on the secret data, including the following steps: S51, the extraction module performs segmentation processing on the secret data to obtain partial secret data 1 and partial secret data 2, and continues to perform decryption processing on the partial secret data 2 to obtain decryption result data; S52, the extraction module generates mixed data according to the decryption result data, and performs a second operation on the mixed data and part of the secret data to obtain first restored data; S53, the extraction module performs a first operation on the first recovery data to obtain intermediate recovery data, and continues to perform a second operation on the intermediate recovery data and the decryption result data to obtain second recovery data, and connects the first recovery data and the second recovery data to obtain the recovered data.

8. The method according to claim 1, characterized in that The extraction module extracts metadata of the secret data while extracting the stored secret data from the corresponding node module. Extracting metadata of the secret data includes: judging whether the amount of metadata in the memory exceeds a preset threshold value, if not, extracting metadata of the secret data; if exceeded, unloading several metadata that have not been accessed in the last 7 days from the memory, continuing to judge whether the amount of metadata in the memory exceeds a preset threshold value, if not exceeded, extracting metadata of the secret data; if exceeded, determining several metadata using an LRU algorithm, unloading the determined several metadata from the memory, and extracting metadata of the secret data.

9. A capacity automatic management system for a distributed data analysis engine, used to implement the method according to any one of claims 1 to 8, characterized in that: Includes the following modules: The management module is used to obtain a number of preset values, determine characteristic values ​​based on the obtained preset values, and divide different data to be stored into data groups with the number of characteristic values, where the characteristic value is a common multiple of the preset values; and is also used to formulate a storage plan for different data to be stored based on the number of node modules currently in use and the number of all data groups, where the number of node modules currently in use corresponds to one of the preset values, and all data in the data groups are stored in the corresponding node modules in a secret storage manner according to the storage plan; and is used to determine whether it is necessary to adjust the number of node modules currently in use; and is also used to reduce the number of node modules currently in use. In the case of reducing the number of node modules, a first ratio of the characteristic value to the number of node modules after adjustment is calculated, a second ratio of the characteristic value to the number of node modules before adjustment is calculated, and all secret data in the data group of the difference between the first ratio and the second ratio are transferred to the corresponding other node modules from the node modules to be reduced. In the case of increasing the number of node modules currently in use, a third ratio of the characteristic value to the number of node modules before adjustment is calculated, and a fourth ratio of the characteristic value to the number of node modules after adjustment is calculated, and all secret data in the data group of the difference between the third ratio and the fourth ratio are transferred to the corresponding node modules to be added from the node modules currently in use. A node module, used to store secret data corresponding to the data to be stored; An extraction module, used to extract the stored secret data from the corresponding node module and perform recovery processing on the secret data; The analysis module is used to perform analysis on the data recovered by the extraction module.

Citation Information

Patent Citations

  • Distributed scheduling method of multi-source data analysis engine, computing node and distributed scheduling system of the multi-source data analysis engine

    CN111625696A

  • Real-time data analysis engine of distributed tracking system

    CN112099977A

  • Distributed data management system and data storage method

    CN112988904A

  • Distributed storage system based on high-speed encryption algorithm

    CN116915510A