A file information storage space management system and method based on big data

By identifying and migrating the load of abnormal storage nodes in the storage space management system, the problem of load imbalance in distributed storage systems is solved, and system performance improvement and stability guarantee are achieved.

CN119556866BActive Publication Date: 2025-05-23SHENZHEN SHIXINDA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510116971.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In distributed storage systems, unbalanced load of storage nodes leads to degradation of system performance, affecting the access speed of data.

Method used

By building a storage space management system, generating storage logs, preset storage nodes, segmenting and storing archive information separately; based on risk assessment of scheduling records, identifying abnormal storage nodes, extracting abnormal factors, formulating information migration strategies, and achieving load balancing.

Benefits of technology

The balanced distribution of storage node loads is realized, resource utilization is improved, overload frequency is reduced, and the stability and reliability of the storage system are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119556866B_ABST
    Figure CN119556866B_ABST
Patent Text Reader

Abstract

The present invention discloses an archive information storage space management system and method based on big data, and relates to the technical field of storage space management. The management method comprises the following steps: constructing a storage space management system to store each archive information, dividing any archive information and storing them separately; generating a corresponding scheduling record for each scheduling behavior, and performing risk assessment on any scheduling record; identifying abnormal storage nodes with abnormal load phenomena, and extracting abnormal factors of the abnormal storage nodes; formulating an information migration strategy to migrate the archive information stored in the abnormal storage nodes; when a new archive information needs to be stored, presetting a number of target storage nodes, and predicting the load condition of any target storage node after storage is completed; giving abnormal reminders to target storage nodes with abnormal prediction results, and reallocating the target storage nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of storage space management, and in particular to an archive information storage space management system and method based on big data. Background Art

[0002] Archival information storage space refers to the physical or virtual space used to store and manage archival information. With the advancement of digitalization, more and more archival information is converted into electronic format and stored in servers, cloud storage or databases. At the same time, there are more and more ways to store archival information, including distributed storage systems.

[0003] A distributed storage system for archival information storage is a system that stores archival information in a dispersed manner on multiple nodes, aiming to improve data availability, reliability and access speed. However, in a distributed storage system, the loads of different storage nodes may vary significantly. Some nodes may be overloaded due to storing a large amount of archival information, while other nodes may be idle. This imbalance will cause system performance to degrade and affect data access speed. Summary of the invention

[0004] The purpose of the present invention is to provide an archive information storage space management system and method based on big data to solve the problems raised in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solution: a method for managing archive information storage space based on big data, the management method comprising the following steps:

[0006] Step S100: construct a storage space management system to store each file information, generate corresponding storage logs, preset several storage nodes, divide and store any file information separately; generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record;

[0007] Step S200: Based on the risk assessment of any storage node in each scheduling record, an abnormal storage node with abnormal load phenomenon is identified; the risk assessment difference of any abnormal storage node in any two scheduling records is analyzed, and the abnormal factors of the abnormal storage node are extracted;

[0008] Step S300: extracting several abnormal factors of any abnormal storage node, and formulating an information migration strategy to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factor in each scheduling record; re-evaluating the risk of each storage node where information migration occurs, and analyzing the information migration status of each storage node;

[0009] Step S400: When a new archive information needs to be stored, a number of target storage nodes are preset for the new archive information based on the load conditions of each storage node, and the load conditions of any target storage node after the storage is completed are predicted; an abnormal reminder is issued to the target storage node with abnormal prediction results, and the target storage node is reallocated.

[0010] Furthermore, step S100 includes the following steps:

[0011] Step S101: Preset several information segmentation rules and several storage node matching rules. For any archive information, select a certain information segmentation rule to segment the archive information to obtain several information blocks, and store each information block in a corresponding storage node according to a certain storage node matching rule, generate a storage log in the storage space management system, and store each storage node of the archive information; the information segmentation rules include horizontal segmentation, vertical segmentation, time segmentation, random segmentation and other segmentation rules, and the storage node matching rules include random allocation, polling allocation, hash allocation and other matching rules;

[0012] Step S102: When the user schedules the archive information each time, a scheduling record is generated in the storage log; a storage node in a scheduling record is randomly selected from the storage log, and load evaluation rules of several dimensions are preset to obtain the load value presented by the storage node in each dimension, and the load value of the storage node is accumulated to obtain the load value of the storage node; the preset dimensions include CPU usage, network bandwidth, response time, scheduling frequency and other dimensions;

[0013] Step S103: Obtain the information block size stored in the i-th storage node in the scheduling record as size i , according to the formula:

[0014] ;

[0015] Where a is the number of storage nodes for the scheduling record, Load i is the load value of the i-th storage node in the scheduling record; the load risk value r of the i-th storage node is calculated i ; Accumulate the load risk value of each storage node to obtain the risk value of the scheduling record; The load risk value of the storage node will be affected by the size of the stored information block, which will directly affect the load value, thereby generating a risk value.

[0016] Further, step S200 includes the following steps:

[0017] Step S201: Select a scheduling record from any storage log, obtain the risk value of the scheduling record as F, and set the load risk value of the i-th storage node in the scheduling record to r i , the load risk ratio of the i-th storage node in the scheduling record is α i =r i / F;

[0018] Step S202: Obtain the archive information size S stored in the storage log and the information block size size stored in the i-th storage node i , the proportion of information blocks stored in the i-th storage node is calculated to be β i =size i / S; if α i >β i , then the i-th storage node in the scheduling record is marked as abnormal; to determine whether the storage node is abnormal in each scheduling record, it can be compared according to the proportion of the stored information blocks. The larger the proportion, the greater the corresponding load risk proportion. By comparing the two, the abnormality can be determined to a certain extent;

[0019] Step S203: Count the number of scheduling records with abnormal marks for the i-th storage node in the storage log as Num i , get the abnormal frequency f of the i-th storage node in the storage log i ; Set an abnormal proportion threshold f th , if f i >f th , then the i-th storage node is set as an abnormal storage node;

[0020] Step S204: arbitrarily select an abnormal storage node, obtain each scheduling record without an abnormal mark in the abnormal storage node, extract the load value under any dimension respectively, and calculate the average value to obtain the average load value Load of the dimension ave , set the maximum load value under the dimension to Load max , the allowable deviation amplitude η of the dimension is obtained = (Load max -Load ave ) / Load ave ;

[0021] Step S205: respectively obtaining the load values ​​of the abnormal storage node in each dimension from any two scheduling records, and if the load deviation amplitude in one dimension exceeds the allowable deviation amplitude, taking the dimension as an abnormal factor of the abnormal storage node;

[0022] Step S206: Acquire several abnormal factors of each abnormal storage node, and integrate all abnormal factors to obtain a set of abnormal factors that affect the storage node load.

[0023] Further, step S300 includes the following steps:

[0024] Step S301: arbitrarily select an abnormal storage node, and arbitrarily obtain two scheduling records with abnormal marks on the abnormal storage node in any storage log, and obtain the load risk difference between the two scheduling records in the abnormal storage node as Δr; obtain the proportion of information blocks in the abnormal storage node in the storage log as β ’, Set the influence degree of the kth abnormal factor in the abnormal storage node to y k , according to the formula:

[0025] ;

[0026] Among them, η k is the load deviation amplitude of the two scheduling records under the dimension corresponding to the kth abnormal factor, c is the number of abnormal factors between the two scheduling records, and ΔLoad is the load difference of the abnormal storage node in the two scheduling records; the load conditions of any two scheduling records in the storage log are substituted into the formula to determine the influence degree of each influencing factor; because the load risk difference of the storage node is obtained through the load value, if there is no influence of the influencing factor, the two are in a linear relationship; however, due to the existence of the influencing factor, this linear relationship fails, and the ratio generated by the two is calculated by the different influence degrees of the influencing factors;

[0027] Step S302: Obtain the influence degree of each influencing factor in any abnormal storage node, and accumulate them to obtain the comprehensive influence degree of the abnormal storage node; sort the abnormal storage nodes from high to low according to the comprehensive influence degree, and distribute the information blocks in the abnormal storage nodes according to the sorting, and sort the information blocks from large to small and distribute them to the storage nodes without abnormalities;

[0028] Step S303: Whenever an information block is migrated, the storage node to which the information block is migrated is set as the target storage node, and the load risk values ​​of the abnormal storage node and the target storage node are calculated according to the formula:

[0029] ;

[0030] Among them, S ’ is the total amount of information blocks stored in the abnormal storage node, S ’’ is the total amount of information blocks stored in the target storage node, (S’ ) total is the storage capacity of the abnormal storage node, (S ’’ ) total is the storage capacity of the target storage node, S yc The size of the information block to be migrated, Load ’ is the load value of the abnormal storage node in the latest generated scheduling record, Load ’’ is the load value of the target storage node in the latest generated scheduling record; and the expected load value r of the abnormal storage node after the information block migration is calculated. ’ and the expected load value r of the target storage node ’’ ;

[0031] Step S304: Obtain any information block in the abnormal storage node, and obtain the risk value F when the archive information where the information block is located is scheduled. ’ and the information block proportion β of the information block in the abnormal storage node ’ , if r ’ / F ’ >β ’ , then continue to select information blocks from the abnormal storage node in order for migration. If r ’ / F ’ <β ’ , then continue to select abnormal storage nodes in order to migrate information blocks.

[0032] Furthermore, step S400 includes the following steps:

[0033] Step S401: by selecting a certain information segmentation rule, the new archive information is segmented to obtain a number of information blocks, one information block is randomly selected, the information block is stored in any storage node, and the expected risk value of the storage node is calculated to obtain the expected load risk ratio of the storage node;

[0034] Step S402: obtaining the information block ratio of the information block in the storage node, and if the expected load risk ratio is greater than the information block ratio, providing an abnormal reminder to the storage node and allocating a new storage node.

[0035] In order to better implement the above method, an archive information storage space management system is also proposed. The management system includes a storage node analysis module, an abnormal load analysis module, a scheduling strategy analysis module and a real-time storage allocation module;

[0036] The storage node analysis module is used to build a storage space management system to store each file information, generate corresponding storage logs, preset several storage nodes, divide any file information and store them separately; generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record;

[0037] The abnormal load analysis module is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; analyze the risk assessment difference of any abnormal storage node in any two scheduling records, and extract the abnormal factors of the abnormal storage node;

[0038] The scheduling strategy analysis module is used to extract several abnormal factors of any abnormal storage node, formulate information migration strategies to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factors in each scheduling record; re-evaluate the risk of each storage node where information migration occurs, and analyze the information migration status of each storage node;

[0039] The real-time storage allocation module is used to preset several target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and predict the load conditions of any target storage node after the storage is completed; abnormal reminders are given to target storage nodes with abnormal prediction results, and the target storage nodes are reallocated.

[0040] Further, the storage node analysis module includes a storage node partitioning unit and a node load evaluation unit;

[0041] The storage node division unit is used to build a storage space management system to store each piece of archive information, generate corresponding storage logs, preset several storage nodes, divide any archive information and store them separately; the node load assessment unit is used to generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record.

[0042] Further, the abnormal load analysis module includes an abnormal load identification unit and an abnormal factor extraction unit;

[0043] The abnormal load identification unit is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; the abnormal factor extraction unit is used to analyze the risk assessment difference of any abnormal storage node in any two scheduling records and extract the abnormal factors of the abnormal storage node.

[0044] Further, the scheduling strategy analysis module includes a scheduling strategy adjustment unit and a data migration analysis unit;

[0045] The scheduling strategy adjustment unit is used to extract several abnormal factors of any abnormal storage node, and based on the impact of any abnormal factors in each scheduling record, formulate an information migration strategy to migrate the archival information stored in the abnormal storage node; the data migration analysis unit is used to re-evaluate the risk of each storage node where information migration occurs, and analyze the information migration status of each storage node.

[0046] Further, the real-time storage allocation module includes a storage load prediction unit and an abnormal load adjustment unit;

[0047] The storage load prediction unit is used to preset several target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and to predict the load conditions of any target storage node after the storage is completed; the abnormal load adjustment unit is used to issue abnormal reminders to the target storage nodes with abnormal prediction results and reallocate the target storage nodes.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention helps to achieve load balancing of nodes by analyzing the load conditions of storage nodes, ensuring that archive information can be reasonably distributed to different nodes, thereby improving the resource utilization rate of each node and reducing the overload frequency of each node;

[0050] 2. The present invention evaluates the risk of each storage node based on the load of each storage node, helps identify the storage risk of each storage node, facilitates timely adjustment of storage strategy during subsequent storage, and ensures the stability and reliability of the storage system;

[0051] 3. When storing any archival information, the present invention predicts the load conditions of each storage node, helps store the archival information in a suitable storage node, reduces the probability of subsequent abnormalities in each storage node, and ensures the normal operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of the steps of a method for managing archive information storage space based on big data;

[0053] Figure 2 The figure is a structural diagram of an archive information storage space management system based on big data. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] Example: Figure 1 to Figure 2 As shown, the present invention provides an archive information storage space management method based on big data, and the management method comprises the following steps:

[0056] Step S100: construct a storage space management system to store each file information, generate corresponding storage logs, preset several storage nodes, divide and store any file information separately; generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record;

[0057] Wherein, step S100 includes the following steps:

[0058] Step S101: Preset several information segmentation rules and several storage node matching rules. For any archive information, select a certain information segmentation rule to segment the archive information to obtain several information blocks, and store each information block in a corresponding storage node according to a certain storage node matching rule. Generate a storage log in the storage space management system to store each storage node of the archive information.

[0059] Step S102: When the user schedules the archive information each time, a scheduling record is generated in the storage log; a storage node in a scheduling record is randomly selected from the storage log, and load evaluation rules of several dimensions are preset to obtain the load value presented by the storage node in each dimension, and the load value of the storage node is accumulated to obtain the load value of the storage node;

[0060] Step S103: Obtain the information block size stored in the i-th storage node in the scheduling record as size i , according to the formula:

[0061] ;

[0062] Where a is the number of storage nodes for the scheduling record, Load i is the load value of the i-th storage node in the scheduling record; the load risk value r of the i-th storage node is calculated i ; Accumulating the load risk values ​​of each storage node to obtain the risk value of the scheduling record;

[0063] Example 1: Randomly select a storage node, the information block size of the archive information in the storage node is 10, and the total size of the archive information is 50. Based on the load values ​​presented in multiple dimensions, the load value of the storage node is 20, and the load risk value r of the storage node is calculated. i =10 / 50×20=4.

[0064] Step S200: Based on the risk assessment of any storage node in each scheduling record, an abnormal storage node with abnormal load phenomenon is identified; the risk assessment difference of any abnormal storage node in any two scheduling records is analyzed, and the abnormal factors of the abnormal storage node are extracted;

[0065] Wherein, step S200 includes the following steps:

[0066] Step S201: Select a scheduling record from any storage log, obtain the risk value of the scheduling record as F, and set the load risk value of the i-th storage node in the scheduling record to r i , the load risk ratio of the i-th storage node in the scheduling record is α i =r i / F;

[0067] Step S202: Obtain the archive information size S stored in the storage log and the information block size size stored in the i-th storage node i , the proportion of information blocks stored in the i-th storage node is calculated to be β i =size i / S; if α i >β i , then the i-th storage node in the scheduling record is marked as abnormal;

[0068] Step S203: Count the number of scheduling records with abnormal marks for the i-th storage node in the storage log as Num i , get the abnormal frequency f of the i-th storage node in the storage log i ; Set an abnormal proportion threshold f th , if f i >f th , then the i-th storage node is set as an abnormal storage node;

[0069] Step S204: arbitrarily select an abnormal storage node, obtain each scheduling record without an abnormal mark in the abnormal storage node, extract the load value under any dimension respectively, and calculate the average value to obtain the average load value Load of the dimension ave, set the maximum load value under the dimension to Load max , the allowable deviation amplitude η of the dimension is obtained = (Load max -Load ave ) / Load ave ;

[0070] Step S205: respectively obtaining the load values ​​of the abnormal storage node in each dimension from any two scheduling records, and if the load deviation amplitude in one dimension exceeds the allowable deviation amplitude, taking the dimension as an abnormal factor of the abnormal storage node;

[0071] Step S206: Acquire several abnormal factors of each abnormal storage node, and integrate all abnormal factors to obtain a set of abnormal factors that affect the storage node load.

[0072] Step S300: extracting several abnormal factors of any abnormal storage node, and formulating an information migration strategy to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factor in each scheduling record; re-evaluating the risk of each storage node where information migration occurs, and analyzing the information migration status of each storage node;

[0073] Wherein, step S300 includes the following steps:

[0074] Step S301: arbitrarily select an abnormal storage node, and arbitrarily obtain two scheduling records with abnormal marks on the abnormal storage node in any storage log, and obtain the load risk difference between the two scheduling records in the abnormal storage node as Δr; obtain the proportion of information blocks in the abnormal storage node in the storage log as β ’, Set the influence degree of the kth abnormal factor in the abnormal storage node to y k , according to the formula:

[0075] ;

[0076] Among them, η k is the load deviation amplitude of the two scheduling records under the dimension corresponding to the kth abnormal factor, c is the number of abnormal factors between the two scheduling records, and ΔLoad is the load difference of the abnormal storage node in the two scheduling records; the load conditions of any two scheduling records in the storage log are substituted into the formula to determine the influence degree of each influencing factor;

[0077] Example 2: Set the load risk difference of the two selected scheduling records to 2, and the information block ratio to 20%, and the load difference of the abnormal storage nodes in the two scheduling records to 8; set a total of 2 influencing factors between the two scheduling records, and the load deviation amplitudes are 50% and 68% respectively, and the equation is 1.25=(1+0.5×y k1 )×(1+0.68×y k2 ), and then select the comparison between the other two scheduling records to obtain another equation: 1.3=(1+0.5×y 1 )×(1+0.9×y 2 ), and get y 1 =20%,y 2 =20%;

[0078] Step S302: Obtain the influence degree of each influencing factor in any abnormal storage node, and accumulate them to obtain the comprehensive influence degree of the abnormal storage node; sort the abnormal storage nodes from high to low according to the comprehensive influence degree, and distribute the information blocks in the abnormal storage nodes according to the sorting, and sort the information blocks from large to small and distribute them to the storage nodes without abnormalities;

[0079] Step S303: Whenever an information block is migrated, the storage node to which the information block is migrated is set as the target storage node, and the load risk values ​​of the abnormal storage node and the target storage node are calculated according to the formula:

[0080] ;

[0081] Among them, S ’ is the total amount of information blocks stored in the abnormal storage node, S ’’ is the total amount of information blocks stored in the target storage node, (S ’ ) total is the storage capacity of the abnormal storage node, (S ’’ ) total is the storage capacity of the target storage node, S yc The size of the information block to be migrated, Load ’ is the load value of the abnormal storage node in the latest generated scheduling record, Load ’’ is the load value of the target storage node in the latest generated scheduling record; and the expected load value r of the abnormal storage node after the information block migration is calculated. ’ and the expected load value r of the target storage node ’’ ;

[0082] Step S304: Obtain any information block in the abnormal storage node, and obtain the risk value F when the archive information where the information block is located is scheduled. ’ and the information block proportion β of the information block in the abnormal storage node ’ , if r ’ / F ’ >β ’ , then continue to select information blocks from the abnormal storage node in order for migration. If r ’ / F ’ <β ’ , then continue to select abnormal storage nodes in order to migrate information blocks.

[0083] Step S400: When a new archive information needs to be stored, a number of target storage nodes are preset for the new archive information based on the load conditions of each storage node, and the load conditions of any target storage node after the storage is completed are predicted; an abnormal reminder is given to the target storage node with abnormal prediction results, and the target storage node is reallocated;

[0084] Wherein, step S400 includes the following steps:

[0085] Step S401: by selecting a certain information segmentation rule, the new archive information is segmented to obtain a number of information blocks, one information block is randomly selected, the information block is stored in any storage node, and the expected risk value of the storage node is calculated to obtain the expected load risk ratio of the storage node;

[0086] Step S402: obtaining the information block ratio of the information block in the storage node, and if the expected load risk ratio is greater than the information block ratio, providing an abnormal reminder to the storage node and allocating a new storage node.

[0087] An archive information storage space management system, the management system includes a storage node analysis module, an abnormal load analysis module, a scheduling strategy analysis module and a real-time storage allocation module;

[0088] The storage node analysis module is used to build a storage space management system to store each file information, generate corresponding storage logs, preset several storage nodes, divide any file information and store them separately; generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record;

[0089] The abnormal load analysis module is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; analyze the risk assessment difference of any abnormal storage node in any two scheduling records, and extract the abnormal factors of the abnormal storage node;

[0090] The scheduling strategy analysis module is used to extract several abnormal factors of any abnormal storage node, formulate information migration strategies to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factors in each scheduling record; re-evaluate the risk of each storage node where information migration occurs, and analyze the information migration status of each storage node;

[0091] The real-time storage allocation module is used to preset several target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and predict the load conditions of any target storage node after the storage is completed; abnormal reminders are given to target storage nodes with abnormal prediction results, and the target storage nodes are reallocated.

[0092] Wherein, the storage node analysis module includes a storage node division unit and a node load evaluation unit;

[0093] The storage node division unit is used to build a storage space management system to store each piece of archive information, generate corresponding storage logs, preset several storage nodes, divide any archive information and store them separately; the node load evaluation unit is used to generate corresponding scheduling records for each scheduling behavior, and perform risk evaluation on the scheduling records based on the load conditions of each storage node involved in any scheduling record.

[0094] Among them, the abnormal load analysis module includes an abnormal load identification unit and an abnormal factor extraction unit;

[0095] The abnormal load identification unit is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; the abnormal factor extraction unit is used to analyze the risk assessment difference of any abnormal storage node in any two scheduling records and extract the abnormal factors of the abnormal storage node.

[0096] Among them, the scheduling strategy analysis module includes a scheduling strategy adjustment unit and a data migration analysis unit;

[0097] The scheduling strategy adjustment unit is used to extract several abnormal factors of any abnormal storage node, and based on the impact of any abnormal factors in each scheduling record, formulate an information migration strategy to migrate the archival information stored in the abnormal storage node; the data migration analysis unit is used to re-evaluate the risk of each storage node where information migration occurs, and analyze the information migration status of each storage node.

[0098] Wherein, the real-time storage allocation module includes a storage load prediction unit and an abnormal load adjustment unit;

[0099] The storage load prediction unit is used to preset several target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and to predict the load conditions of any target storage node after the storage is completed; the abnormal load adjustment unit is used to issue abnormal reminders to the target storage nodes with abnormal prediction results and reallocate the target storage nodes.

[0100] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A method for managing archival information storage space based on big data, characterized in that: The management method comprises the following steps: Step S100: Each file information is segmented based on a preset information segmentation rule, and stored in a plurality of storage nodes in the storage space management system, respectively, to generate corresponding storage logs; a corresponding scheduling record is generated for each scheduling behavior, and a risk assessment is performed on the scheduling record based on the load conditions of each storage node involved in any scheduling record; Step S200: Based on the risk assessment of any storage node in each scheduling record, an abnormal storage node with abnormal load phenomenon is identified; the risk assessment difference between any abnormal storage node in different scheduling records is analyzed, and the abnormal factors of the abnormal storage node are extracted; Step S300: extracting several abnormal factors of any abnormal storage node, and formulating an information migration strategy to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factor in each scheduling record; re-evaluating the risk of each storage node where information migration occurs, and analyzing the information migration status of each storage node; Step S400: When a new archive information needs to be stored, a number of target storage nodes are preset for the new archive information based on the load conditions of each storage node, and the load conditions of any target storage node after the storage is completed are predicted; an abnormal reminder is given to the target storage node with abnormal prediction results, and the target storage node is reallocated; The step S200 includes the following steps: Step S201: Select a scheduling record from any storage log, obtain the risk value of the scheduling record as F, and set the load risk value of the i-th storage node in the scheduling record to r i , the load risk ratio of the i-th storage node in the scheduling record is α i =r i / F; Step S202: Obtain the archive information size S stored in the storage log and the information block size size stored in the i-th storage node i , the proportion of information blocks stored in the i-th storage node is calculated to be β i =size i / S; if α i >β i , then the i-th storage node in the scheduling record is marked as abnormal; Step S203: Count the number of scheduling records with abnormal marks for the i-th storage node in the storage log as Num i , get the abnormal frequency f of the i-th storage node in the storage log i ; Set an abnormal proportion threshold f th , if f i >f th , then the i-th storage node is set as an abnormal storage node; Step S204: arbitrarily select an abnormal storage node, obtain each scheduling record without an abnormal mark in the abnormal storage node, extract the load value under any dimension respectively, and calculate the average value to obtain the average load value Load of the dimension ave , set the maximum load value under the dimension to Load max , the allowable deviation amplitude η of the dimension is obtained = (Load max -Load ave ) / Load ave ; Step S205: respectively obtaining the load values ​​of the abnormal storage node in each dimension from any two scheduling records, and if the load deviation amplitude in one dimension exceeds the allowable deviation amplitude, taking the dimension as an abnormal factor of the abnormal storage node; Step S206: obtaining several abnormal factors of each abnormal storage node, and integrating all abnormal factors to obtain a set of abnormal factors that affect the load of the storage node; The step S300 includes the following steps: Step S301: arbitrarily select an abnormal storage node, and arbitrarily obtain two scheduling records with abnormal marks on the abnormal storage node in any storage log, and obtain the load risk difference between the two scheduling records in the abnormal storage node as Δr; obtain the proportion of information blocks in the abnormal storage node in the storage log as β ’ ; Obtain each difference abnormal factor of the two scheduling records in the abnormal node, and set the influence degree of the kth difference abnormal factor in the abnormal storage node to y k , according to the formula: ; Among them, η k is the load deviation amplitude of the two scheduling records in the dimension corresponding to the kth difference abnormal factor, c is the number of difference abnormal factors between the two scheduling records, and ΔLoad is the load difference of the abnormal storage node in the two scheduling records; the load conditions of any two scheduling records in the storage log are substituted into the formula to determine the influence degree of each influencing factor; Step S302: Obtain the influence degree of each influencing factor in any abnormal storage node, and accumulate them to obtain the comprehensive influence degree of the abnormal storage node; sort the abnormal storage nodes from high to low according to the comprehensive influence degree, and distribute the information blocks in the abnormal storage nodes according to the sorting, and sort the information blocks from large to small and distribute them to the storage nodes without abnormalities; Step S303: Whenever an information block is migrated, the storage node to which the information block is migrated is set as the target storage node, and the load risk values ​​of the abnormal storage node and the target storage node are calculated according to the formula: ; Among them, S ’ is the total amount of information blocks stored in the abnormal storage node, S ’’ is the total amount of information blocks stored in the target storage node, (S ’ ) total is the storage capacity of the abnormal storage node, (S ’’ ) total is the storage capacity of the target storage node, S yc The size of the information block to be migrated, Load ’ is the load value of the abnormal storage node in the latest generated scheduling record, Load ’’ is the load value of the target storage node in the latest generated scheduling record; and the expected load value r of the abnormal storage node after the information block migration is calculated. ’ and the expected load value r of the target storage node ’’ ; Step S304: Obtain any information block in the abnormal storage node, and obtain the risk value F when the archive information where the information block is located is scheduled. ’ and the information block proportion β of the information block in the abnormal storage node ’ , if r ’ / F ’ >β ’ , then continue to select information blocks from the abnormal storage node in order for migration. If r ’ / F ’ <β ’ , then continue to select abnormal storage nodes in order to migrate information blocks.

2. The method for managing archival information storage space based on big data according to claim 1, characterized in that: The step S100 includes the following steps: Step S101: obtaining the file type of any file information, assigning several information segmentation rules to the file type, and randomly selecting one information segmentation rule to segment the file information to obtain several information blocks; setting a corresponding storage node matching rule for each information segmentation rule in the storage space management system, storing each information block in a corresponding storage node, and obtaining a storage log of the file information; Step S102: When the user schedules the archive information each time, a scheduling record is generated in the storage log; a storage node in a scheduling record is randomly selected from the storage log, and load evaluation rules of several dimensions are set for any storage node to obtain the load value presented by the storage node in each dimension, and the load value of the storage node is accumulated to obtain the load value of the storage node; Step S103: Obtain the information block size stored in the i-th storage node in the scheduling record as size i , according to the formula: ; Where a is the number of storage nodes for the scheduling record, Load i is the load value of the i-th storage node in the scheduling record; the load risk value r of the i-th storage node is calculated i ; Accumulate the load risk value of each storage node to obtain the risk value of the scheduling record.

3. The method for managing archival information storage space based on big data according to claim 2, characterized in that: The step S400 includes the following steps: Step S401: by selecting a certain information segmentation rule, the new archive information is segmented to obtain a number of information blocks, one information block is randomly selected, the information block is stored in any storage node, and the expected risk value of the storage node is calculated to obtain the expected load risk ratio of the storage node; Step S402: obtaining the information block ratio of the information block in the storage node, and if the expected load risk ratio is greater than the information block ratio, providing an abnormal reminder to the storage node and allocating a new storage node.

4. An archive information storage space management system, used to implement an archive information storage space management method based on big data as claimed in any one of claims 1 to 3, characterized in that: The management system includes a storage node analysis module, an abnormal load analysis module, a scheduling strategy analysis module and a real-time storage allocation module; The storage node analysis module is used to segment each archive information based on a preset information segmentation rule, and store them in several storage nodes in the storage space management system respectively, and generate corresponding storage logs; generate corresponding scheduling records for each scheduling behavior, and perform risk assessment on the scheduling records based on the load conditions of each storage node involved in any scheduling record; The abnormal load analysis module is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; Analyze the risk assessment differences between different scheduling records of any abnormal storage node, and extract the abnormal factors of the abnormal storage node; The scheduling strategy analysis module is used to extract several abnormal factors of any abnormal storage node, and formulate an information migration strategy to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factor in each scheduling record; Re-evaluate the risk of each storage node where information migration occurs and analyze the information migration status of each storage node; The real-time storage allocation module is used to preset several target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and predict the load conditions of any target storage node after the storage is completed; to issue abnormal reminders to target storage nodes with abnormal prediction results, and to reallocate the target storage nodes.

5. The archival information storage space management system according to claim 4, characterized in that: The storage node analysis module includes a storage node division unit and a node load evaluation unit; The storage node division unit is used to divide each piece of archive information based on a preset information segmentation rule, and store them respectively in several storage nodes in the storage space management system to generate corresponding storage logs; the node load evaluation unit is used to generate a corresponding scheduling record for each scheduling behavior, and perform risk evaluation on the scheduling record based on the load conditions of each storage node involved in any scheduling record.

6. The archival information storage space management system according to claim 4, characterized in that: The abnormal load analysis module includes an abnormal load identification unit and an abnormal factor extraction unit; The abnormal load identification unit is used to identify abnormal storage nodes with abnormal load phenomena based on the risk assessment of any storage node in each scheduling record; The abnormal factor extraction unit is used to analyze the risk assessment differences between different scheduling records of any abnormal storage node and extract the abnormal factors of the abnormal storage node.

7. The archival information storage space management system according to claim 4, characterized in that: The scheduling strategy analysis module includes a scheduling strategy adjustment unit and a data migration analysis unit; The scheduling strategy adjustment unit is used to extract several abnormal factors of any abnormal storage node, and formulate an information migration strategy to migrate the archive information stored in the abnormal storage node based on the influence of any abnormal factor in each scheduling record; The data migration analysis unit is used to re-evaluate the risk of each storage node where information migration occurs, and analyze the information migration status of each storage node.

8. The archival information storage space management system according to claim 4, characterized in that: The real-time storage allocation module includes a storage load prediction unit and an abnormal load adjustment unit; The storage load prediction unit is used to preset a number of target storage nodes for the new archive information based on the load conditions of each storage node when a new archive information needs to be stored, and to predict the load condition of any target storage node after the storage is completed; The abnormal load adjustment unit is used to issue an abnormal reminder to the target storage node with abnormal prediction results and reallocate the target storage node.

Citation Information

Patent Citations

  • Server data storage management method and system

    CN118036042A

  • KR20220060871A