Big Data Storage and Backup Management System Based on Cloud Services
The backup feature information is obtained through the data acquisition module, a weighted analysis model is established, and the backup interval is dynamically adjusted, which solves the problem of parameter library that cannot be called dynamically in the existing technology, realizes the optimization of system stability and resource utilization, and reduces the risk of data loss.
Patent Information
- Application Number
- CN202510147830.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In the face of sudden data growth or abnormal network fluctuations, the parameters related to the backup time interval cannot be accurately called, resulting in the parameters collected in the parameter library being dynamically called to balance system stability, reduce the adaptability of the cloud storage environment for long-term operation of the system, and increase the waste of computing resources.
The backup feature information is obtained through the data acquisition module, a weighted analysis model is established, a deviation coefficient is generated, and a data hot and cold mechanism is set. Fuzzy reasoning is performed based on the frequency characteristics of the cold data parameters in the backup interval, and the backup interval is dynamically adjusted.
Improve the dynamic adaptability of backup strategies, optimize storage resource utilization, reduce the risk of data loss, and balance system stability.
Smart Images

Figure CN119669234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data backup management, and more specifically, to a big data storage and backup management system based on cloud services. Background Art
[0002] In a big data storage and backup management system based on cloud services, backup is an important means to ensure data security. A reasonable backup interval can ensure that when unexpected data failures or anomalies occur, the system has recent backup data, thus minimizing the risk of data loss. Reasonable adjustment of the backup interval is crucial for the stability of the system, the optimal utilization of storage resources, and the efficiency of data recovery.
[0003] The prior art has the following deficiencies:
[0004] Currently, by dynamically adjusting the backup frequency to balance system stability and backup performance, however, in the face of sudden data growth or abnormal network fluctuations, it is impossible to accurately call the parameters related to the backup duration interval for substitution and calculation. This results in the inability to dynamically adjust the parameters collected in the parameter library to balance system stability, reduces the adaptability of the cloud storage environment for the long-term operation of the system, and increases the waste of computing resources caused by unreasonable adjustment. Therefore, a big data storage and backup management system based on cloud services is proposed.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a big data storage and backup management system based on cloud services, and solves the problems raised in the above background art by using different product inspection methods.
[0007] To achieve the above object, the present invention provides the following technical solution: A big data storage and backup management system based on cloud services, including a data acquisition module, an adjustment determination module, a parameter management module, and a data call module; the modules are signal-connected to each other;
[0008] The data acquisition module is used to collect backup feature information and backup round information, and send them to the adjustment determination module;
[0009] The adjustment determination module is used to obtain backup feature information and backup round information, establish a weighted analysis model, obtain the total number of backup interval adjustments, and collect the parameters affecting the backup interval adjustment range, and integrate them into a parameter library and send it to the parameter management module;
[0010] The parameter management module is used to obtain the total number of backup interval adjustments and the parameter library, obtain all the parameters in the parameter library, and set the data cold and hot mechanism according to the data corresponding to all the parameters in the parameter library, classify the parameters in the parameter library, and send them to the data call module;
[0011] The data call module is used to obtain the parameter classification result, and substitute the frequency characteristics of the cold data parameters within the backup interval into the fuzzy logic to determine the cold data parameter call result.
[0012] In a preferred embodiment, the backup feature information includes the deviation between the estimated backup time interval and the reference value, and the deviation between the estimated backup data volume and the reference data volume;
[0013] Based on the backup time interval data recorded in the previous backup operation and combined with the current system storage utilization rate, calculate the estimated backup time interval, and subtract it from the pre-set backup time interval reference value to obtain the deviation between the estimated backup time interval and the reference value;
[0014] Based on the previous backup data volume, the current data growth rate, and the data change frequency, obtain the estimated backup data volume, and subtract it from the pre-set reference backup data volume to obtain the deviation between the estimated backup data volume and the reference data volume;
[0015] Calculate the overall round duration according to the real-time monitoring data.
[0016] In a preferred embodiment, establish a weighted analysis model for the deviation between the estimated backup time interval and the reference value and the deviation between the estimated backup data volume and the reference data volume to generate a deviation coefficient;
[0017] Substitute the deviation coefficient and the overall round duration into the adaptive backup interval adjustment formula to obtain the total number of backup interval adjustments.
[0018] In a preferred embodiment, the parameter library includes the data growth ratio, the time deviation ratio, and the system load deviation ratio;
[0019] The backup interval is a preset fixed backup interval during the first run, and the length of the subsequent backup interval is determined according to the corresponding parameters in the parameter library;
[0020] The data growth ratio, the time deviation ratio, and the system load deviation ratio are all the cumulative data growth ratio, time deviation ratio, and system load deviation ratio of the second time and the first time after the first initial backup interval of the current round has been run;
[0021] The network delay effect corresponding to all parameters in the parameter library and the data change frequency jointly determine the adjustment of the data growth ratio, time deviation ratio, and system load deviation ratio.
[0022] In a preferred embodiment, the round-trip time and the one-way transmission delay are used to perform a weighted calculation on the ratio calculation result of the round-trip time to the currently backed-up data volume and the current network bandwidth to obtain the network delay effect;
[0023] Obtain the data change amount within a unit time and perform a ratio calculation with the time intervals accumulated in a preset number of unit times to obtain the data change frequency.
[0024] In a preferred embodiment, the network delay effect and the data change frequency corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio are multiplied and substituted into the data growth ratio, time deviation ratio, and system load deviation ratio to obtain the data cold and hot evaluation criteria;
[0025] Sort the data cold and hot evaluation criteria corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio from largest to smallest. The two parameters with higher rankings are used as hot data parameters, and the parameter with the lowest ranking is used as a cold data parameter.
[0026] In a preferred embodiment, the frequency characteristics of the cold data parameter within the backup interval include the cold data parameter growth frequency and the cold data parameter backup interval times;
[0027] Record the value of the cold data parameter, calculate the growth amplitude using the cold data parameter values of two consecutive backups, and then perform a ratio calculation on two adjacent growth amplitudes to obtain the cold data parameter growth frequency;
[0028] Count the number of times the hot data parameter is generated to obtain the cold data parameter backup interval times.
[0029] In a preferred embodiment, the cold data parameter growth frequency and the cold data parameter backup interval times are defined as input variables and divided into different fuzzy sets;
[0030] Define the cold data parameter call result as an output variable and divide it into a fuzzy set;
[0031] Formulate fuzzy rules to describe the influence of the cold data parameter growth frequency and the cold data parameter backup interval times on the cold data parameter call result;
[0032] Perform fuzzy reasoning according to the fuzzy rules to determine the cold data parameter call result.
[0033] The technical effects and advantages of the present invention:
[0034] 1. The present invention collects the deviation amount between the estimated backup time interval and the reference value, the deviation amount between the estimated backup data volume and the reference data volume, and the overall round duration, establishes a weighted analysis model to obtain the total number of backup interval adjustments, collects the parameters affecting the backup interval adjustment range, and integrates them into a parameter library. According to the network delay effect and data change frequency corresponding to all the parameters in the parameter library, a data hot and cold mechanism is set, and the hot data parameters and cold data parameters are classified, so as to improve the dynamic adaptability of the backup strategy, optimize the utilization rate of storage resources, dynamically adjust the backup interval, and balance the system stability.
[0035] 2. The present invention obtains the hot data parameters and cold data parameters, and based on the frequency characteristics of the cold data parameters within the backup interval, obtains the growth frequency of the cold data parameters and the number of backup intervals of the cold data parameters, formulates a set of fuzzy rules for fuzzy reasoning, and determines the call result of the cold data parameters, so as to adapt to a large-scale data backup environment, avoid long-term non-backup of cold data, and reduce the risk of data loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the modules of the big data storage and backup management system based on cloud service of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] Embodiment 1
[0039] The present invention discloses a big data storage and backup management system based on cloud service, as Figure 1 shown, including a data collection module, an adjustment and determination module, a parameter management module, and a data call module; the modules are connected by signals;
[0040] The data collection module is used to collect backup feature information and backup round information, obtain the deviation amount between the estimated backup time interval and the reference value, the deviation amount between the estimated backup data volume and the reference data volume, and the overall round duration through data processing, and send them to the adjustment and determination module;
[0041] Among them, data processing includes data type conversion, which refers to converting the original data collected by the system into a standardized data type, and missing value processing, which refers to filling in the possible data missing during the system collection process through interpolation method and mean filling to ensure the continuity and integrity of the data, etc. The purpose of data processing operations is to make the data clearer and easier to understand, and make the system calculations faster;
[0042] Specifically, the above data operation methods are all existing technologies and will not be elaborated here;
[0043] Among them, the backup feature information includes the deviation between the estimated backup time interval and the reference value, and the deviation between the estimated backup data volume and the reference data volume. The backup round information includes the overall round duration;
[0044] Specifically, the system backup operation is divided into multiple rounds, and each round contains multiple backup times. The initial backup interval is set preferentially. The specific number of rounds and the initial backup interval set are both determined by the experimenter based on historical data and the big data coverage, and will not be elaborated here;
[0045] The acquisition logic of the deviation between the estimated backup time interval and the reference value is to calculate the estimated backup time interval through the backup time interval data recorded in the previous round of backup operation, and combine the current system storage utilization rate. According to the trend prediction algorithm, the estimated backup time interval is calculated, and then according to the system design requirements and historical operation statistics, the backup time interval reference value is preset, and the difference between the estimated backup time interval and the backup time interval reference value is obtained to get the deviation between the estimated backup time interval and the reference value. ;
[0046] Among them, the trend prediction algorithm calculates the estimated next backup time interval through the exponential smoothing method. Let the actual time interval of the i-th backup in the previous round be , and assume that there are n backup records. These data form a time series:
[0047] ;
[0048] Specifically, the exponential smoothing method is a simple and effective trend prediction algorithm. Its core idea is to assign higher weights to the latest observed values and lower weights to earlier data. Let the smoothing coefficient be (where ), and the prediction model is described as:
[0049] ;
[0050] In the formula, is the predicted value of the (i + 1)-th backup time interval, is the actual backup time interval of the i-th time, is the predicted value of the i-th backup time interval;
[0051] It should be noted that the initial value can be set as the first actual observation value or a reasonable value selected by the experimenter according to prior knowledge, etc., which will not be elaborated here;
[0052] Specifically, regarding the subscripts and superscripts in the above calculation formula, b refers to the parameter related to "backup", that is, the time interval in the backup operation, and p represents "estimation", that is, the estimated backup time interval obtained by the parameter through the prediction model;
[0053] Using the above exponential smoothing formula, iterative calculations are performed on the data of the previous round to sequentially obtain the predicted time intervals for the second to the n + 1-th backups, and the result of the last iterative calculation is obtained After that, it is used as the estimated backup time interval;
[0054] It should be noted that the estimated backup time interval is actually the first backup time interval of the current round, which can be considered as the initial backup interval of the current round;
[0055] Among them, the smoothing coefficient is to control the weight of the latest observation value in the prediction. A larger smoothing coefficient makes the prediction more sensitive to the latest data, while a smaller smoothing coefficient makes the predicted value smoother. The optimal smoothing coefficient is determined through historical data fitting and error analysis. Therefore, the smoothing coefficient is continuously iteratively updated;
[0056] Subtract the estimated backup time interval from the reference value of the backup time interval. The specific formula is:
[0057] ;
[0058] In the formula, is the reference value of the backup time interval, and its subscript re represents the "reference" value, that is, the preset reference value of the backup time interval;
[0059] The logic for obtaining the deviation amount between the estimated backup data volume and the reference data volume is based on the backup data volume of the previous round, the current data growth rate, and the data change frequency. The estimated backup data volume is calculated by substituting into the prediction model, and then based on the system storage plan and historical statistical data, the reference backup data volume is preset. Subtract the estimated backup data volume from the reference backup data volume to obtain the deviation amount between the estimated backup data volume and the reference data volume ;
[0060] Specifically, the backup data volume of the previous round is the historical backup data volume obtained by weighted calculation of the backup file size and the number of backup records;
[0061] Subtract the estimated backup data volume from the reference backup data volume, and the specific formula is expressed as:
[0062] ;
[0063] In the formula, is the reference backup data volume;
[0064] It should be noted that substituting into the prediction model to calculate the estimated next backup data volume can be based on the above prediction model and so on, which will not be elaborated here;
[0065] The overall round duration refers to the time span from the start to the end of a round in the big data storage and backup management system based on cloud services, which includes a series of backup operations and possible backup interval adjustments during this period. Its acquisition logic calculates the overall round duration based on real-time monitoring data, and substitutes the backup data volume of the previous round and the target data transfer rate of the previous round into the optimization model to obtain the overall round duration ;
[0066] Specifically, the calculation formula of the optimization model is:
[0067] ;
[0068] In the formula, is the backup data volume of the previous round, is the target data transfer rate of the previous round, t is the duration of the previous round, is to find the t value that minimizes the objective function;
[0069] It should be noted that a round represents all relevant scheduling, adjustment, and monitoring operations throughout the backup cycle. The entire optimization process aims to select a suitable time window for the new round based on the actual operation data of the previous round, so as to achieve the adaptive optimization and efficient operation of the system;
[0070] Specifically, the reference values corresponding to all the above parameters are obtained by the experimenters based on historical data, which will not be elaborated here;
[0071] The adjustment determination module is used to obtain the deviation amount between the estimated backup time interval and the reference value, the deviation amount between the estimated backup data volume and the reference data volume, establish a weighted analysis model, obtain the deviation coefficient, substitute the deviation coefficient and the overall round duration into the calculation to obtain the total number of backup interval adjustments, and collect the parameters affecting the backup interval adjustment range, and integrate them into a parameter library and the total number of backup interval adjustments and send them to the parameter management module;
[0072] Specifically, the weighted analysis model refers to generating the deviation coefficient through weighted calculation;
[0073] Generally, during the weighted calculation process, the above parameters are preferentially normalized to normalize parameters with different dimensions and different units of measurement so that they are within the same relative scale range for reasonable weighted calculation;
[0074] Obtain the deviation between the estimated backup time interval and the reference value, and the deviation between the estimated backup data volume and the reference data volume, establish a weighted analysis model, and generate a deviation coefficient , and the formula is:
[0075] ;
[0076] In the formula, is the deviation coefficient, and are respectively the preset proportionality coefficients of the ratios of the deviation between the estimated backup time interval and the reference value and the deviation between the estimated backup data volume and the reference data volume to the corresponding reference values, and and are both greater than 0;
[0077] Substitute the deviation coefficient and the overall round duration into the calculation to obtain, and the specific formula expression is:
[0078] ;
[0079] In the formula, is the total number of adjustments of the backup interval; is the number of reference backup cycles theoretically included in this round;
[0080] Furthermore, to prevent insufficient system adjustment due to a negative deviation coefficient or too low adjustment times, usually at least one adjustment is set, that is ;
[0081] It should be noted that when the system is in the initial state of round operation, different initial fixed backup intervals are set. The specific fixed backup time can be determined according to the transmission status of big data or the security level. The specific security level can be obtained according to the preset hash algorithm, and the transmission status of big data can be obtained according to the transmission network status and data size, etc., which will not be elaborated here;
[0082] Specifically, the parameter library contains a data growth ratio, a time deviation ratio, and a system load deviation ratio;
[0083] Then, as can be seen from the above, the backup interval is the preset fixed backup interval during the first run, and the length of the subsequent backup interval is determined according to the corresponding parameters in the parameter library;
[0084] Meanwhile, based on the length of the subsequent backup interval and the above calculations, the total number of backup interval adjustments is obtained, so as to ensure that the parameters selected and called from the parameter library are different in each backup interval. The total number of backup interval adjustments represents that in the current round, the parameters obtained through the prediction model are substituted into the data model for calculation to determine the deviation coefficient, and then calculated with the overall round duration to determine in advance the total number of backup interval adjustments in the current round;
[0085] It should be noted that the total number of backup interval adjustments and the total number of backup intervals are not the same concept. The total number of backup intervals is related to parameters such as the duration of the current round and usually includes the number of backup intervals that do not require interval adjustment;
[0086] The parameter management module is used to obtain the total number of backup interval adjustments, connect to the parameter library to obtain all parameters in the parameter library, and set the data hot and cold mechanism according to the network delay effect and data change frequency corresponding to all parameters in the parameter library, classify the hot data parameters and cold data parameters, and send them to the data call module;
[0087] Among them, the parameter library includes the data growth ratio, time deviation ratio, and system load deviation ratio;
[0088] Specifically, the data growth ratio, time deviation ratio, and system load deviation ratio are all the cumulative data growth ratio, time deviation ratio, and system load deviation ratio of the second time and the first time after the first initial backup interval of the current round has been run;
[0089] For example, if the data volume is m, the time is t1, and the load volume is v1 for the first time, and the data volume is u, the time is t2, and the load volume is v2 for the second time, then the data growth ratio is (u - m) / m, the time deviation ratio is (t2 - t1) / t1, and the system load deviation ratio is (v2 - v1) / v1, and so on;
[0090] Among them, the network delay effect and data change frequency corresponding to all parameters in the parameter library can jointly determine the increase and decrease of the data growth ratio, time deviation ratio, and system load deviation ratio;
[0091] The acquisition logic of the network delay effect is through the round-trip time and the one-way transmission delay. The round-trip time from the backup system to the storage server is measured by the ping command, and the ratio of the round-trip time to the current data volume and the current network bandwidth is weighted to obtain the network delay effect. ; where i is the i-th parameter in the parameter library;
[0092] Specifically, when the network delay effect is greater, the data growth ratio, time deviation ratio, and system load deviation ratio, that is, the delay increases, resulting in a longer backup data transmission time. Then the time deviation ratio increases, causing the data not to be backed up on time. When backing up next time, the data volume becomes larger, so the data growth ratio increases, resulting in a longer system waiting time and the CPU being occupied for a long time, then the system load deviation ratio increases;
[0093] The acquisition logic of the data change frequency is to set a unit time in the backup interval, obtain the data change amount within the unit time, and calculate the ratio with the time interval accumulated by a preset number of unit times to obtain the data change frequency. ;
[0094] Specifically, when the data change frequency increases, the amount of data within the unit time increases. Therefore, compared with the previous backup, the current data volume will become larger, resulting in an increase in the data growth ratio, indirectly causing an increase in the time required for the backup process, thereby increasing the time deviation ratio. Secondly, it will cause the storage system to need to process a larger amount of data, increasing the CPU calculation, disk, and network bandwidth occupancy, thereby increasing the system load;
[0095] It should be noted that the preset unit time is set by the experimenter according to the total number of specific backup intervals, the current backup interval, and the total number of backup intervals in the previous round. Specifically, for the preset unit time in the data change frequency, it needs to be less than the current backup interval. Since the backup interval is in an adjusted state, the preset unit time will also be adjusted with the adjustment of the current backup interval. The specific adjustment mechanism is not limited and will not be elaborated here;
[0096] Perform a product calculation on the network delay effect and data change frequency corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio, and substitute them into the data growth ratio, time deviation ratio, and system load deviation ratio to obtain the data cold and hot evaluation criteria;
[0097] The specific formula for the data cold and hot evaluation criteria is:
[0098] ;
[0099] Since this example lists three parameters in the parameter library, in the actual application scenario, multiple parameters can be applied for calculation, and the number of parameters in the specific parameter library is not limited;
[0100] Among them, and the network delay effect and data change frequency corresponding to the parameter of the data growth ratio, is the data cold and hot evaluation criteria corresponding to the data growth ratio, then and The network delay effect and data change frequency corresponding to the time deviation ratio parameter are the data cold and hot evaluation criteria corresponding to the time deviation ratio. Then and the network delay effect and data change frequency corresponding to the system load deviation ratio parameter are the data cold and hot evaluation criteria corresponding to the system load deviation ratio;
[0101] Sort the data cold and hot evaluation criteria corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio from largest to smallest. Take the two parameters with higher rankings as hot data parameters, and take the parameter with the lower ranking as the cold data parameter;
[0102] For example, sort the above , and from largest to smallest to obtain . Then, the time deviation ratio and system load deviation ratio are hot data, and the data growth ratio is cold data. At the next backup interval, execute the adjustment instruction, substitute the time deviation ratio and system load deviation ratio into the calculation to determine the adjustment range of the backup interval, etc., which will not be elaborated here;
[0103] The present invention collects the deviation amount between the estimated backup time interval and the reference value, the deviation amount between the estimated backup data volume and the reference data volume, and the overall round duration, establishes a weighted analysis model to obtain the total number of backup interval adjustments, collects the parameters affecting the backup interval adjustment range, integrates them into a parameter library, and sets a data cold and hot mechanism according to the network delay effect and data change frequency corresponding to all the parameters in the parameter library, classifies the hot data parameters and cold data parameters, improves the dynamic adaptability of the backup strategy, optimizes the storage resource utilization rate, dynamically adjusts the backup interval, and balances the system stability.
[0104] Embodiment 2
[0105] In Embodiment 1 of the present invention, the deviation amount between the collected estimated backup time interval and the reference value, the deviation amount between the estimated backup data volume and the reference data volume, and the overall round duration are mainly exemplified. A weighted analysis model is established to obtain the total number of backup interval adjustments. Parameters affecting the backup interval adjustment range are collected and integrated into a parameter library. According to the network delay effect and data change frequency corresponding to all parameters in the parameter library, a data hot and cold mechanism is set, and the operation strategies for hot data parameters and cold data parameters are classified. However, in Embodiment 1, only hot data parameters are analyzed. Obviously, if cold data parameters are ignored for a long time, cold data will not be backed up for a long time, increasing the risk of data loss and reducing the system's backtracking ability. For the above problems, Embodiment 2 of the present invention is further refined;
[0106] The data call module is used to obtain hot data parameters and cold data parameters, and based on the frequency characteristics of cold data parameters within the backup interval, obtain the growth frequency of cold data parameters and the number of cold data parameter backup intervals, and substitute them into fuzzy logic to determine the cold data parameter call result;
[0107] Specifically, the frequency characteristics of cold data parameters within the backup interval include the growth frequency of cold data parameters and the number of cold data parameter backup intervals;
[0108] The acquisition logic of the growth frequency of cold data parameters is to record the values of cold data parameters within each backup interval, use the values of cold data parameters in two consecutive backups as the comparison basis, calculate the growth amplitude, and then calculate the ratio of two adjacent growth amplitudes to obtain the growth frequency of cold data parameters;
[0109] The acquisition logic of the number of cold data parameter backup intervals is to record the number and corresponding status of the current backup interval within the current round, count the number of times of generating hot data parameters, and record it as the number of cold data parameter backup intervals;
[0110] For example, "High", "Low", "Medium" for the growth frequency of cold data parameters, and "More", "Less", "Average" for the number of cold data parameter backup intervals;
[0111] Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example:
[0112] Mark the growth frequency of cold data parameters as X, the number of cold data parameter backup intervals as U, and the cold data parameter call result as C_Public;
[0113] Then it can be defined as:
[0114] Rule 1: IF (X is High) AND (U is Less) THEN (C_Public is High)
[0115] Rule 2: IF (X is Low) AND (U is More) THEN (C_Public is Low) ...
[0117] Perform fuzzy inference according to the fuzzy rules to determine the call result of the cold data parameters;
[0118] It should be noted that the division of the fuzzy sets can be adjusted according to the actual situation. For example, although three fuzzy sets are taken as an example in this embodiment, in fact, the growth frequency of the cold data parameters and the number of cold data parameter backup intervals can be divided into more than three sets to facilitate more accurate adjustment of classifying data packets according to different life cycle stages.
[0119] Furthermore, for the judgment of high, medium, and low of the growth frequency of the cold data parameters and the number of cold data parameter backup intervals, thresholds can be set for judgment according to the actual situation. For example, when the growth frequency of the cold data parameters exceeds 80%, it is calibrated as "High", and when the number of cold data parameter backup intervals is higher than 75%, it is calibrated as "More", etc., which will not be elaborated here;
[0120] It should be noted that when the call result of the cold data parameter is "High", it means that when executing the adjustment instruction at the next backup interval, not only the hot data parameters need to be called for substitution calculation, but also the cold data needs to be called as the adjustment basis for the backup interval for substitution calculation;
[0121] The present invention obtains the hot data parameters and the cold data parameters, and based on the frequency characteristics of the cold data parameters within the backup interval, obtains the growth frequency of the cold data parameters and the number of cold data parameter backup intervals, formulates a set of fuzzy rules for fuzzy inference, determines the call result of the cold data parameters, adapts to the large-scale data backup environment, avoids long-term non-backup of cold data, and reduces the risk of data loss.
[0122] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0123] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0124] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0125] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0127] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0128] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0130] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0131] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A big data storage and backup management system based on cloud services, characterized in that: It includes a data acquisition module, an adjustment determination module, a parameter management module, and a data call module; the modules are signal-connected to each other; The data acquisition module is used to acquire backup feature information and backup round information, and send them to the adjustment determination module; The adjustment determination module is used to obtain backup feature information and backup round information, establish a weighted analysis model, obtain the total number of backup interval adjustments, and collect parameters affecting the backup interval adjustment range, and integrate them into a parameter library and send it to the parameter management module; The parameter management module is used to obtain the total number of backup interval adjustments and the parameter library, obtain all the parameters in the parameter library, and set a data cold and hot mechanism according to the data corresponding to all the parameters in the parameter library, classify the parameters in the parameter library, and send them to the data call module; The data call module is used to obtain the parameter classification result, and substitute it into fuzzy logic to determine the cold data parameter call result according to the frequency characteristics of the cold data parameter within the backup interval; The frequency characteristics of the cold data parameter within the backup interval include the growth frequency of the cold data parameter and the number of backup intervals of the cold data parameter; Record the value of the cold data parameter, calculate the growth rate based on the cold data parameter values of two consecutive backups, and then calculate the ratio of two adjacent growth rates to obtain the growth frequency of the cold data parameter; Count the number of times the hot data parameter is generated to obtain the number of backup intervals of the cold data parameter; The backup feature information includes the deviation between the estimated backup time interval and the reference value, and the deviation between the estimated backup data volume and the reference data volume; Define the growth frequency of the cold data parameter and the number of backup intervals of the cold data parameter as input variables, and divide them into different fuzzy sets; Define the cold data parameter call result as the output variable and divide it into a fuzzy set; Formulate fuzzy rules to describe the influence of the growth frequency of the cold data parameter and the number of backup intervals of the cold data parameter on the cold data parameter call result; Perform fuzzy reasoning according to the fuzzy rules to determine the cold data parameter call result.
2. The big data storage and backup management system based on cloud service according to claim 1, characterized in that: Calculate the estimated backup time interval through the backup time interval data recorded in the previous round of backup operation, and combine the current system storage utilization rate, and subtract it from the preset backup time interval reference value to obtain the deviation between the estimated backup time interval and the reference value; Based on the backup data volume of the previous round, the current data growth rate, and the data change frequency, obtain the estimated backup data volume, and subtract it from the preset reference backup data volume to obtain the deviation between the estimated backup data volume and the reference data volume; Calculate the overall round duration according to the real-time monitoring data.
3. The big data storage and backup management system based on cloud service according to claim 2, characterized in that: Obtain the deviation between the estimated backup time interval and the reference value, and the deviation between the estimated backup data volume and the reference data volume, establish a weighted analysis model, and generate a deviation coefficient; Substitute the deviation coefficient and the overall round duration into the adaptive backup interval adjustment formula to obtain the total number of backup interval adjustments. The specific formula is expressed as: ; In the formula, is the total number of backup interval adjustments; is the number of reference backup cycles theoretically included in this round, is the reference value of the estimated backup time interval.
4. The big data storage and backup management system based on cloud service according to claim 3, characterized in that: The parameter library includes a data growth ratio, a time deviation ratio, and a system load deviation ratio; The backup interval is a preset fixed backup interval during the first run, and the length of the subsequent backup interval is determined according to the corresponding parameters in the parameter library; The data growth ratio, time deviation ratio, and system load deviation ratio are all the cumulative data growth ratio, time deviation ratio, and system load deviation ratio of the second time and the first time after the first initial backup interval of the current round has been run. The network delay effect and data change frequency corresponding to all parameters in the parameter library jointly determine the adjustment of the data growth ratio, time deviation ratio, and system load deviation ratio.
5. The big data storage and backup management system based on cloud service according to claim 4, characterized in that: Through the round-trip time and the one-way transmission delay, the ratio calculation result of the round-trip time to the amount of data currently backed up and the current network bandwidth is weighted to obtain the network delay effect. Obtain the amount of data change per unit time and calculate the ratio with the time interval accumulated in a preset number of unit times to obtain the data change frequency.
6. The big data storage and backup management system based on cloud service according to claim 5, wherein: Multiply the network delay effect and data change frequency corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio, and substitute them into the data growth ratio, time deviation ratio, and system load deviation ratio to obtain the data hot and cold evaluation criteria. Sort the data hot and cold evaluation criteria corresponding to the data growth ratio, time deviation ratio, and system load deviation ratio from largest to smallest. The two parameters with the highest ranking are used as hot data parameters, and the parameter with the lowest ranking is used as a cold data parameter.
Citation Information
Patent Citations
Data backup method and system
CN114968672A
Textile equipment energy consumption real-time acquisition method and system based on 5G network
CN118382024A