Method and system for optimizing storage of archival data

By analyzing the access curve and access feature periods in the medical archive storage system, and reasonably allocating backup storage space, the problem of low storage efficiency of medical archive backup in the existing technology is solved, and more efficient storage and transmission is achieved.

CN119781695BActive Publication Date: 2025-06-06BEIJING GO TO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272857.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-06
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

In the prior art, the backup storage efficiency of medical files is low, mainly due to the unreasonable setting of backup storage spaces for each system node.

Method used

By obtaining the access curve of the archive pages in each system node, the access feature period is determined, and the system nodes are grouped. According to the storage performance of each system node in the system node group, the backup reliability is determined, and the backup storage space size of each system node is adjusted according to the reliability and the number of nodes in the group.

Benefits of technology

The storage efficiency of archive backup is improved, and the storage information transmission between system nodes is optimized by reasonably allocating backup storage space, and the backup storage space of system nodes with poor storage performance is appropriately reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781695B_ABST
    Figure CN119781695B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimized storage method and system for archival data, and relates to the technical field of data processing. The method comprises: obtaining access curves of each archive page in each system node of an archive storage system; determining the access characteristic time period of each system node according to each access curve in each system node; grouping the system nodes according to the access characteristic time period of each system node to obtain multiple system node groups; for each system node group, determining the backup reliability of each system node according to the node storage performance of each system node in the system node group; creating a backup storage space for each system node according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs; and backing up and storing each archive page through the backup storage space of each system node. The optimized storage method for archival data provided by the embodiment of the present invention can improve the storage efficiency of archive backup.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to an optimized storage method and system for archival data. Background Art

[0002] Because paper medical records are difficult to preserve and there is a high possibility of information loss, patients' medical records are currently usually stored digitally.

[0003] In the existing method, medical records are stored in different system nodes based on the principle of distributed storage, and medical records that are frequently accessed recently are backed up and stored in the system nodes.

[0004] However, the existing method has the problem that the backup storage space of each system node is not set reasonably, which leads to low efficiency of archive backup storage. Summary of the invention

[0005] The embodiment of the present invention provides a method and system for optimizing storage of archive data, which can improve the storage efficiency of archive backup.

[0006] A first aspect of an embodiment of the present invention provides a method for optimizing storage of archive data, comprising:

[0007] Obtaining access curves of each archive page in each system node of the archive storage system, where the access curves are used to represent a time series curve of the number of accesses to the archive page by users in the system node;

[0008] According to each access curve in each system node, determine the access characteristic time period of each system node, and the access characteristic time period is used to represent the time period that conforms to the access habits of users in the system node to the archives;

[0009] According to the access characteristic time period of each system node, the system nodes are grouped to obtain multiple system node groups;

[0010] For each system node group, determine the backup reliability of each system node according to the node storage performance of each system node in the system node group;

[0011] Creating a backup storage space for each system node according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs;

[0012] Each archive page is backed up and stored through the backup storage space of each system node.

[0013] A second aspect of an embodiment of the present invention provides an optimized storage system for archival data, comprising:

[0014] A curve graph acquisition module is used to acquire the access curve of each archive page in each system node of the archive storage system, and the access curve is used to represent the time series curve of the access volume of the archive page by the users in the system node;

[0015] A time period determination module is used to determine the access characteristic time period of each system node according to each access curve in each system node, and the access characteristic time period is used to represent the time period that conforms to the access habits of users in the system node to the archives;

[0016] A node grouping module is used to group system nodes according to the access characteristic time period of each system node to obtain multiple system node groups;

[0017] A reliability determination module, for determining the backup reliability of each system node according to the node storage performance of each system node in the system node group for each system node group;

[0018] A space creation module, used to create a backup storage space for each system node according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs;

[0019] The archive storage module is used to back up and store each archive page through the backup storage space of each system node.

[0020] In the optimized storage method for archive data provided by the embodiment of the present invention, the access characteristic time period of each system node is first determined according to each access curve in each system node. Then, the system nodes are grouped according to the access characteristic time period of each system node to obtain multiple system node groups. In this way, the system nodes with similar access habits for different archives are divided into the same group, which can facilitate the efficient transmission of storage information between each system node in the system node group and improve the storage efficiency of archive backup. Then, for each system node group, the backup reliability of each system node is determined according to the node storage performance of each system node in the system node group. And according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs, a backup storage space for each system node is created. In this way, the backup storage space size of different system nodes is adjusted according to the backup reliability of each system node, and the backup storage space size of the system node with poor storage performance is reduced. In this way, a larger backup storage space is set for the system node suitable for storing archive backup, and a smaller backup storage space is set for the system node not suitable for storing archive backup, so that the storage efficiency of archive backup can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A schematic flow chart of a first method for optimizing storage of archive data provided by an embodiment of the present invention;

[0023] Figure 2 A schematic diagram of an access curve provided by an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of a system node group provided by an embodiment of the present invention;

[0025] Figure 4 A schematic flow chart of a second method for optimizing storage of archive data provided by an embodiment of the present invention;

[0026] Figure 5 A schematic flow chart of a third method for optimizing storage of archive data provided by an embodiment of the present invention;

[0027] Figure 6 A schematic diagram of the structure of an optimized storage system for archival data provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of an optimized storage method and system for archival data proposed by the present invention, its specific implementation, structure, features and effects, in conjunction with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0029] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0030] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of the present invention are in compliance with the relevant provisions of laws and regulations.

[0031] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned, which should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present invention, but it does not mean that the applicant has or will necessarily use the solution.

[0032] Because paper medical records are difficult to preserve and there is a high possibility of information loss, patients' medical records are currently usually stored digitally.

[0033] In the existing method, medical files are stored in different system nodes based on the principle of distributed storage, and medical files that are frequently accessed recently are backed up and stored in the system nodes. However, the existing method has the problem that the backup storage space of each system node is set irrationally, resulting in low efficiency of file backup storage.

[0034] The object of the present invention is to provide an optimized storage method and system for archival data. In the optimized storage method for archival data provided by the embodiment of the present invention, the access characteristic time period of each system node is first determined according to each access curve in each system node. Then, the system nodes are grouped according to the access characteristic time period of each system node to obtain multiple system node groups. In this way, the system nodes with similar access habits for different archives are divided into the same group, which can facilitate the efficient transmission of storage information between each system node in the system node group and improve the storage efficiency of archive backup. Then, for each system node group, the backup reliability of each system node is determined according to the node storage performance of each system node in the system node group. And according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs, a backup storage space for each system node is created. In this way, the backup storage space size of different system nodes is adjusted according to the backup reliability of each system node, and the backup storage space size of the system node with poor storage performance is reduced. In this way, a larger backup storage space is set for the system node suitable for storing archive backup, and a smaller backup storage space is set for the system node not suitable for storing archive backup, so that the storage efficiency of archive backup can be improved.

[0035] The following describes specific embodiments of the method and system for optimizing storage of archive data provided by the embodiments of the present invention.

[0036] Figure 1 A flow chart of a method for optimizing storage of archival data is provided. The method for optimizing storage of archival data can be applied to a server. The method for optimizing storage of archival data can include the following S101 to S106.

[0037] S101, obtaining access curves of various archive pages in various system nodes of the archive storage system, where the access curves are used to represent a time series curve of the number of accesses to the archive pages by users in the system nodes.

[0038] In this embodiment, each system node of the archive storage system corresponds to a department. For example, in the archive storage system of a hospital, the outpatient department corresponds to a system node, and the inspection department corresponds to a system node.

[0039] The access curve can reflect the changes in the number of visits to the archive page by users in the system node. The horizontal axis of the access curve is time, and the vertical axis of the access curve is the number of visits.

[0040] As an example, the server first collects the access logs of the archive pages of each user in each system node, wherein the access logs contain information such as access time, user ID, and ID of the archive page visited. Then, the collected log data is processed to count the number of visits per minute to each archive page in each system node.

[0041] Finally, using the statistical results, the access curve of each archive page in each system node is generated. Figure 2 As shown in FIG. 1 , a schematic diagram of an access curve is provided, wherein the horizontal axis is time in minutes, and the vertical axis is the number of visits in times.

[0042] S102, determining the access characteristic time period of each system node according to each access curve in each system node, where the access characteristic time period is used to represent the time period that conforms to the access habits of users in the system node to the archives.

[0043] In this embodiment, due to differences in job responsibilities, users in different departments have different access habits to archives. For example, there are differences in the types of access to archives between departments. Users in the laboratory department usually only need to call the basic information page, medical history page, test result page, etc. in the medical archives, and rarely access other archive pages, such as prescription details page, payment details page, etc. At the same time, there will be differences in the access time of archives between departments. Therefore, the difference between the access habits of each system node to archives is characterized by access feature time periods.

[0044] As an example, the server divides multiple time periods according to a preset division rule, such as every five minutes as a time period, and then analyzes the access curve of each system node, identifies the time period with significantly higher access volume than the average level, and determines the time period as the access characteristic period of the system node.

[0045] S103, grouping the system nodes according to the access characteristic time period of each system node to obtain a plurality of system node groups.

[0046] In this embodiment, the access habits of users in different departments to archives are not only different, but also similar. For example, medical staff in the outpatient department need to access the basic information page and prescription page in the patient's medical archive during the patient's visit; since most patients need to get medicine after the visit, the staff of the pharmacy department may also access the patient's prescription page at a similar time. Therefore, system nodes with similar access feature time periods can be divided into the same system node group.

[0047] As an example, the server compares the access feature time periods of each system node and calculates the similarity between them. Specifically, the similarity can be evaluated by comparing the overlap degree of the access feature time periods and the change trend of the access volume in the access feature time periods.

[0048] Then, a clustering algorithm (such as K-means, DBSCAN, etc.) is applied to group the system nodes so that the system nodes in the same system node group have similar access characteristic time periods.

[0049] like Figure 3 As shown, a schematic diagram of a system node group is provided, wherein the system node group 310 includes the system node 301 , the system node 302 , and the system node 303 , and the system node group 320 includes the system node 304 and the system node 305 .

[0050] S104: For each system node group, determine the backup reliability of each system node according to the node storage performance of each system node in the system node group.

[0051] In this embodiment, the backup reliability is used to characterize the reliability and efficiency of the system node in the backup storage task of the archive page.

[0052] As an example, the server evaluates the storage performance of each system node in the system node group, including storage speed, storage capacity, failure rate and other indicators. Then, based on the evaluation results of the storage performance and the access characteristic period of the system node (such as the processing capacity during the access peak period), the backup reliability of each system node is evaluated.

[0053] S105 , creating a backup storage space for each system node according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs.

[0054] In this embodiment, the server allocates corresponding backup storage space resources according to the backup reliability of each system node and the number of system nodes in the system node group to which it belongs.

[0055] Among them, the system node with higher backup reliability can be allocated more backup storage space resources; the more system nodes there are in the system node group to which the system node belongs, the more archive pages the system node group needs to back up and store, that is, more backup storage space resources are allocated.

[0056] S106, backing up and storing each archive page through the backup storage space of each system node.

[0057] In this embodiment, as an example, the server backs up each archive page to the backup storage space of the corresponding system node in the corresponding system node group according to the pre-established backup storage strategy. The integrity and consistency of the data should be ensured during the backup process. After the backup is completed, the backup data is verified and checked to ensure the validity and availability of the backup data.

[0058] As another example, the server may also evaluate the backup necessity of each archive page. When the backup necessity of the archive page is greater than a preset necessity threshold, the archive page is stored; when the backup necessity of the archive page is less than or equal to the preset necessity threshold, the archive page is not stored.

[0059] Specifically, the necessity of backing up the archive page can be determined by the following formula 1:

[0060] Formula 1

[0061] In formula 1, It is used to characterize the necessity of backing up the a-th file page in the i-th file. It is used to represent the maximum number of visits to the a-th archive page in each archive within a preset time period. It is used to represent the number of visits to the a-th archive page in the i-th archive within a preset time period. It is used to represent the time from the last access behavior of the a-th archive page in the i-th archive to the current time. It is used to characterize the normalized function operation. It should be noted that, in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiments of the present invention, when the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to the actual situation, and this application does not impose any special restrictions.

[0062] In the optimized storage method for archival data provided in the present embodiment, the access characteristic time period of each system node is first determined according to each access curve in each system node. Then, the system nodes are grouped according to the access characteristic time period of each system node to obtain multiple system node groups. In this way, the system nodes with similar access habits for different archives are divided into the same group, which can facilitate the efficient transmission of stored information between each system node in the system node group and improve the storage efficiency of archive backup. Then, for each system node group, the backup reliability of each system node is determined according to the node storage performance of each system node in the system node group. And according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs, a backup storage space for each system node is created. In this way, the backup storage space size of different system nodes is adjusted according to the backup reliability of each system node, and the backup storage space size of the system node with poor storage performance is reduced. In this way, a larger backup storage space is set for the system node suitable for storing archive backup, and a smaller backup storage space is set for the system node not suitable for storing archive backup, so that the storage efficiency of archive backup can be improved.

[0063] As an optional embodiment, Figure 4 As shown, S102 may specifically include:

[0064] For each access curve in each system node, the following S401 to S403 are performed respectively:

[0065] S401, dividing the target access curve from the extreme value point to obtain a plurality of first curve segments, the target access curve being any one of the access curves;

[0066] S402, determining the importance of the candidate time periods corresponding to the first curve segments according to the number of visits of the target access curve in the first curve segments;

[0067] S403: Determine each candidate time period whose importance is greater than a preset importance threshold as an access feature time period.

[0068] In this embodiment, the first curve segment is used to represent each sub-curve segment obtained after dividing the target access curve, and the time period corresponding to each sub-curve segment is the candidate time period.

[0069] As an example, the server first traverses the target access curve and finds all local maxima (peaks) and local minima (valleys). This is usually done by comparing the size of each point with its neighboring points. Then, based on the identified extreme points, the curve is divided into multiple first curve segments. Each first curve segment starts from a valley and ends at the next peak (or starts from a peak and ends at the next valley).

[0070] Then, for each first curve segment, the sum, average, standard deviation or other statistical indicators of the number of visits are calculated. According to the business logic, one or more criteria are pre-set to evaluate the importance of the candidate time period corresponding to the first curve segment. For example, a time period with a high average value of the number of visits is considered to be more important. Then, according to the set criteria, an importance is calculated for each first curve segment. Specifically, the importance can be a linear function or a nonlinear function based on the number of visits, or it can be other complex evaluation models.

[0071] Finally, according to business needs or historical data, set a reasonable importance threshold. This threshold can be fixed or dynamically adjusted. Traverse all candidate time periods corresponding to the first curve segment, compare their importance with the preset importance threshold, and mark the candidate time periods above the importance threshold as access feature time periods.

[0072] Through this embodiment, according to the access curve of the system node, the access characteristic period with significant access characteristics in the system node can be effectively identified, which helps to accurately group the system nodes according to the access characteristic period of the system node, thereby improving the storage efficiency of the archive backup.

[0073] As an optional embodiment, S402 may specifically include:

[0074] For each first curve segment, perform the following steps respectively:

[0075] In the target access curve, obtaining the minimum access volume of the first curve segment and the candidate duration of the candidate time period corresponding to the first curve segment;

[0076] According to each reference access curve, determine the reference access volume mean and reference duration mean corresponding to the first curve segment, the reference access curve is each access curve in the system node except the target access curve, the reference access volume mean is used to characterize the access volume mean of the second curve segment corresponding to the first curve segment in each reference access curve, and the reference duration mean is used to characterize the duration mean of each second curve segment;

[0077] The importance of the candidate time period corresponding to the first curve segment is determined by using the minimum access volume, the candidate time period, the reference access volume average and the reference time period average.

[0078] In this embodiment, the reference access curve is other access curves in the system node except the target access curve, and the second curve segment is used to characterize the curve segment in the reference access curve corresponding to the first curve segment. For example, if the first curve segment is the curve segment from the 5th minute to the 10th minute, the second curve segment is the curve segment from the 5th minute to the 10th minute in the reference access curve.

[0079] If there is no curve segment from the 5th minute to the 10th minute in the reference access curve, the middle time of each curve segment in the reference access curve is compared with the middle time of the first curve segment, and the curve segment closest to the middle time of the first curve segment is used as the second curve segment. That is, the middle time of the first curve segment is 7.5 minutes, and the curve segment with the middle time closest to 7.5 minutes among the curve segments of the reference access curve is used as the second curve segment.

[0080] The reference visit volume mean is used to characterize the result obtained by calculating the mean of the visit volume of each second curve segment, and the reference duration mean is used to characterize the result obtained by calculating the mean of the duration of each second curve segment.

[0081] As an example, the importance can be specifically determined by the following formula 2:

[0082] Formula 2

[0083] In formula 2, It is used to characterize the importance of the candidate time segment corresponding to the jth first curve segment. It is used to represent the minimum value of the number of visits in the jth first curve segment. The candidate duration used to represent the candidate time period corresponding to the j-th first curve segment. It is used to characterize the mean value of the reference visits corresponding to the jth first curve segment. It is used to represent the mean value of the reference duration corresponding to the jth first curve segment. Used to characterize normalization function operations.

[0084] in, The larger it is, the faster the change trend of the j-th first curve segment is, and the greater the importance of the candidate time period corresponding to the j-th first curve segment is; The larger the jth first curve segment is, the greater the number of visits to the jth first curve segment is, and the greater the importance of the candidate time period corresponding to the jth first curve segment is.

[0085] Through this embodiment, the importance of the candidate time period corresponding to the first curve segment is comprehensively evaluated by using the minimum access volume, candidate duration, reference access volume average, and reference duration average corresponding to the first curve segment. In this way, by quantitatively calculating the importance, the access feature time period with significant access characteristics in the system node can be effectively identified, thereby improving the identification accuracy of the access feature time period.

[0086] As an optional embodiment, Figure 5 As shown, S103 may specifically include the following S501 to S503:

[0087] S501, determining a first similarity between a first system node and a second system node according to the importance of an access characteristic time period of a first system node and the importance of an access characteristic time period of a second system node, wherein the first system node and the second system node are any two different system nodes;

[0088] S502, determining, based on the first similarity, a possibility that the first system node and the second system node are in the same group;

[0089] S503: When the same-group possibility is greater than a preset possibility threshold, determine that the first system node and the second system node belong to the same system node group.

[0090] In this embodiment, the first system node and the second system node are any two different system nodes among the system nodes of the archive storage system.

[0091] The first similarity is used to characterize the similarity between the first system node and the second system node in their access habits to archives, and the same-group possibility is used to characterize the possibility that the first system node and the second system node belong to the same system node group.

[0092] As an example, the server uses a similarity measurement method (such as cosine similarity, Pearson correlation coefficient, Jaccard similarity, etc.) to calculate the first similarity between the first system node and the second system node in the importance of the access feature time period based on the importance of the access feature time period of the first system node and the importance of the access feature time period of the second system node.

[0093] Then, the first similarity is used as an input feature, and a machine learning model (such as a classifier, a regression model, etc.) is used to predict the possibility that the first system node and the second system node are in the same group.

[0094] Finally, the same-group possibility is compared with the preset possibility threshold. If the same-group possibility is greater than the preset possibility threshold, it is considered that the first system node and the second system node belong to the same system node group; if the same-group possibility is less than or equal to the preset possibility threshold, it is considered that the first system node and the second system node do not belong to the same system node group.

[0095] Through this embodiment, each system node is accurately grouped according to the importance of the access characteristic time period of the first system node and the importance of the access characteristic time period of the second system node, which helps to divide system nodes with similar archive access habits into the same group and improve the storage efficiency of archive backup.

[0096] As an optional embodiment, the number of access characteristic time periods is multiple;

[0097] S501 may specifically include:

[0098] Obtaining a minimum importance value among the importance values ​​of each access feature time period of the first system node;

[0099] The importance of the zth access characteristic period of the first system node is subtracted from the importance of the zth access characteristic period of the second system node, and the absolute value of the difference between the importance of the first system node and the second system node is obtained, where z is a positive integer;

[0100] The first similarity between the first system node and the second system node in the zth access feature time period is determined by using the absolute value of the importance difference and the minimum importance value.

[0101] In this embodiment, the first similarity can be specifically determined by the following formula 3:

[0102] Formula 3

[0103] In formula 3, Used to characterize the first system node k and the second system node In the The first similarity of the access feature time period, The first node k is used to characterize the The importance of the access feature time period, Used to characterize the second system node No. The importance of the access feature time period, The minimum importance value among the importance values ​​of each access feature time period used to characterize the first system node k.

[0104] in, It is used to characterize the absolute value of the importance difference between the first system node and the second system node in the zth access feature time period. The smaller the absolute value of the importance difference is, the closer the importance of the first system node is to that of the second system node, that is, the greater the first similarity between the first system node and the second system node is.

[0105] Through this embodiment, according to the importance of the access characteristic period of the first system node and the importance of the access characteristic period of the second system node, the first similarity between the first system node and the second system node can be accurately determined, thereby helping to classify system nodes with similar file access habits into the same group, thereby improving the storage efficiency of file backup.

[0106] As an optional embodiment, the number of access characteristic time periods is multiple;

[0107] After S501, the optimized storage method for archival data may further include:

[0108] Determining a second similarity between the first system node and the second system node according to a first similarity between the first system node and the second system node in each access feature time period;

[0109] S502 may specifically include:

[0110] According to the second similarity, a possibility that the first system node and the second system node are in the same group is determined.

[0111] In this embodiment, the first similarity is used to characterize the local similarity of the access habits of the first system node and the second system node to the archives evaluated based on a single access feature period, and the second similarity is used to characterize the overall similarity of the access habits of the first system node and the second system node to the archives evaluated based on all access feature periods.

[0112] As an example, after obtaining the local first similarities between the first system node and the second system node in each access feature time period, the server calculates the average of each first similarity to obtain the overall second similarity between the first system node and the second system node.

[0113] As another example, the second similarity between the first system node and the second system node may be determined by the following formula 4:

[0114] Formula 4

[0115] In formula 4, Used to characterize the first system node k and the second system node The second similarity is Used to characterize the first system node k and the second system node The mean of the first similarities of Used to characterize the first system node k and the second system node The maximum value of the first similarity, Used to characterize the first system node k and the second system node The minimum value of the first similarity, Used to characterize the first system node k and the second system node The maximum value of the first similarities of the third system nodes other than .

[0116] in, Used to characterize the first system node k and the second system node The fluctuation range of the first similarity of , the smaller the value, the greater the difference between the first system node k and the second system node The smaller the fluctuation of the similarity of access to different archive pages, the closer the first system node k is to the second system node The greater the second similarity.

[0117] Then, the second similarity is used as an input feature, and a machine learning model (such as a classifier, a regression model, etc.) is used to predict the possibility that the first system node and the second system node are in the same group.

[0118] Through this embodiment, the first similarity in each access feature time period is comprehensively considered to obtain a more accurate second similarity between the first system node and the second system node, thereby improving the calculation accuracy of the possibility of the same group and achieving accurate grouping of each system node.

[0119] As an optional embodiment, S502 may specifically include:

[0120] Acquire the response time between the first system node and the second system node, and acquire the target similarity between the first system node and the third system nodes except the second system node, where the target similarity is used to characterize the maximum value of the second similarities between the first system node and each third system node;

[0121] The possibility that the first system node and the second system node are in the same group is determined by using the second similarity, the response duration, and the target similarity between the first system node and the second system node.

[0122] In this embodiment, the response duration is used to represent the time taken from when a system node sends a request to another system node to when the receiving system node receives the request and returns a response.

[0123] The third system node is used to represent each system node of the archive storage system except the first system node and the second system node.

[0124] As an example, the same group probability can be specifically determined by the following formula 5:

[0125] Formula 5

[0126] In formula 5, Used to characterize the first system node k and the second system node The possibility of the same group, Used to characterize the first system node k and the second system node The second similarity. It is used to characterize the target similarity between the first system node k and the third system nodes. Specifically, the second similarities between the first system node k and each third system node are calculated by the above formula 3, and the maximum value thereof is determined as the target similarity. Used to characterize the first system node k and the second system node Response time.

[0127] in, The larger the value, the greater the difference between the first system node k and the second system node The more similar the access habits between the first system node and the second system node are to those of other system nodes, the greater the possibility that the first system node and the second system node are in the same group.

[0128] Through this embodiment, the second similarity between the first system node and the second system node, the response time between the first system node and the second system node, and the target similarity between the first system node and the third system node other than the second system node are used to accurately evaluate the possibility of the first system node and the second system node being in the same group. This can improve the calculation accuracy of the possibility of being in the same group and achieve accurate grouping of each system node.

[0129] As an optional embodiment, S104 may specifically include:

[0130] For each system node group, perform the following steps:

[0131] The information complexity of the target system node group is determined by using the possibility of the system nodes in the target system node group being in the same group, and the target system node group is any system node group;

[0132] The backup reliability of each system node in the target system node group is determined by utilizing the information complexity of the target system node group and the node storage performance of each system node in the target system node group.

[0133] In this embodiment, the information complexity can be specifically determined by the following formula 6:

[0134] Formula 6

[0135] In formula 6, Used to characterize the information complexity of the a-th system node group, It is used to characterize the maximum value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to characterize the mean value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to represent the number of system nodes contained in the a-th system node group. Used to represent the total number of system nodes in the archive storage system.

[0136] in, The larger it is, the lower the similarity of archive access between system nodes in the a-th system node group is, and the greater the information complexity of the a-th system node group is; The larger it is, the more system nodes are contained in the a-th system node group, and the greater the information complexity of the a-th system node group.

[0137] The backup reliability can be specifically determined by the following formula 7:

[0138] Formula 7

[0139] In formula 7, It is used to characterize the backup reliability of the cth system node in the ath system node group. Used to characterize the information complexity of the a-th system node group. It is used to characterize the average of the maximum throughput of the storage devices of each system node in the a-th system node group. It is used to characterize the maximum throughput of the storage device of the cth system node in the ath system node group. Indicates the total number of storage devices in the a-th system node group.

[0140] in, The smaller it is, the smaller the maximum throughput of the storage device of the cth system node in the ath system node group is, that is, the worse the node storage performance of the storage device of the cth system node in the ath system node group is, and the lower the backup reliability of the cth system node in the ath system node group is; the greater the information complexity of the ath system node group and the more the total number of storage devices in the ath system node group is, the smaller the possibility that the storage device of the cth system node has spare capacity to store archive page backups is, that is, the lower the backup reliability of the cth system node in the ath system node group is.

[0141] Through this embodiment, the backup reliability of each system node in the target system node group can be accurately determined by utilizing the information complexity of the target system node group and the node storage performance of each system node in the target system node group. Thus, the backup storage space of each system node can be accurately determined, and the storage efficiency of archive backup can be improved.

[0142] As an optional embodiment, S105 may specifically include:

[0143] For each system node, perform the following steps:

[0144] Determine a backup evaluation parameter of the system node by using the backup reliability of the system node and the number of system nodes in the system node group to which the system node belongs;

[0145] The backup evaluation parameter is rounded upward to an integer to obtain the number of backup storage devices of the system node;

[0146] Create backup storage space for system nodes based on the number of backup storage devices for system nodes.

[0147] In this embodiment, the backup evaluation parameter can be specifically determined by the following formula 8:

[0148] Formula 8

[0149] In formula 8, The backup evaluation parameter used to characterize the c-th system node, It is used to represent the number of system nodes contained in the a-th system node group where the c-th system node is located. Used to characterize the backup reliability of the cth system node in the ath system node group.

[0150] Then, the server rounds up the backup evaluation parameter of the cth system node to obtain the number of backup storage devices of the cth system node, thereby creating a backup storage space of the cth system node corresponding to the number of backup storage devices in the hash ring.

[0151] Through this embodiment, the backup storage space of each system node is created according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs. In this way, a larger backup storage space is set for the system node suitable for storing archive backups, and a smaller backup storage space is set for the system node not suitable for storing archive backups, thereby improving the storage efficiency of archive backups.

[0152] Based on the optimized storage method for archival data, the present invention also provides a specific embodiment of the optimized storage system for archival data.

[0153] Figure 6 A structural schematic diagram of an optimized storage system for archival data is provided. The optimized storage system 600 for archival data includes a curve graph acquisition module 610, a time period determination module 620, a node grouping module 630, a reliability determination module 640, a space creation module 650 and an archive storage module 660.

[0154] The curve graph acquisition module 610 is used to acquire the access curve of each archive page in each system node of the archive storage system, and the access curve is used to represent the time series curve of the access volume of the archive page by the user in the system node;

[0155] The time period determination module 620 is used to determine the access characteristic time period of each system node according to each access curve in each system node, and the access characteristic time period is used to represent the time period that conforms to the access habits of users in the system node to the archives;

[0156] The node grouping module 630 is used to group the system nodes according to the access characteristic time period of each system node to obtain multiple system node groups;

[0157] A reliability determination module 640 is used to determine the backup reliability of each system node according to the node storage performance of each system node in the system node group for each system node group;

[0158] A space creation module 650, for creating a backup storage space for each system node according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs;

[0159] The archive storage module 660 is used to back up and store each archive page through the backup storage space of each system node.

[0160] In the optimized storage method for archival data provided in the present embodiment, the access characteristic time period of each system node is first determined according to each access curve in each system node. Then, the system nodes are grouped according to the access characteristic time period of each system node to obtain multiple system node groups. In this way, the system nodes with similar access habits for different archives are divided into the same group, which can facilitate the efficient transmission of stored information between each system node in the system node group and improve the storage efficiency of archive backup. Then, for each system node group, the backup reliability of each system node is determined according to the node storage performance of each system node in the system node group. And according to the backup reliability of each system node and the number of system nodes in the system node group to which the system node belongs, a backup storage space for each system node is created. In this way, the backup storage space size of different system nodes is adjusted according to the backup reliability of each system node, and the backup storage space size of the system node with poor storage performance is reduced. In this way, a larger backup storage space is set for the system node suitable for storing archive backup, and a smaller backup storage space is set for the system node not suitable for storing archive backup, so that the storage efficiency of archive backup can be improved.

[0161] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0162] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in a different order from the embodiments, or several steps can be performed simultaneously.

[0163] The above is only a specific implementation of the present invention. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system, module and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be covered within the protection scope of the present invention.

Claims

1. A method for optimizing storage of archival data, characterized in that: The method comprises: Obtaining an access curve of each archive page in each system node of the archive storage system, wherein the access curve is used to represent a time series curve of the number of visits to the archive page by users in the system node; Determine, according to each of the access curves in each of the system nodes, an access characteristic time period of each of the system nodes, wherein the access characteristic time period is used to characterize a time period that conforms to the access habits of users in the system nodes to archives; According to the access characteristic time period of each of the system nodes, the system nodes are grouped to obtain a plurality of system node groups; For each of the system node groups, determining the backup reliability of each of the system nodes according to the node storage performance of each of the system nodes in the system node group; Creating a backup storage space for each of the system nodes according to the backup reliability of each of the system nodes and the number of system nodes in the system node group to which the system node belongs; Backing up and storing each of the archive pages through the backup storage space of each of the system nodes; The step of determining, for each of the system node groups, the backup reliability of each of the system nodes according to the node storage performance of each of the system nodes in the system node group, includes: For each of the system node groups, perform the following steps respectively: Determine the information complexity of the target system node group by using the possibility of the system nodes in the target system node group being in the same group, wherein the target system node group is any one of the system node groups; Determine the backup reliability of each of the system nodes in the target system node group by using the information complexity of the target system node group and the node storage performance of each of the system nodes in the target system node group; The calculation formula for information complexity includes: ;in, Used to characterize the information complexity of the a-th system node group, It is used to characterize the maximum value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to characterize the mean value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to represent the number of system nodes contained in the a-th system node group. Used to represent the total number of system nodes in the archive storage system; The calculation formula for backup reliability includes: ;in, It is used to characterize the backup reliability of the cth system node in the ath system node group. It is used to characterize the average of the maximum throughput of the storage devices of each system node in the a-th system node group. It is used to characterize the maximum throughput of the storage device of the cth system node in the ath system node group. Indicates the total number of storage devices in the a-th system node group.

2. The method for optimizing storage of archival data according to claim 1, characterized in that: The determining, according to each of the access curves in each of the system nodes, an access characteristic time period of each of the system nodes comprises: For each access curve in each system node, the following steps are performed respectively: Dividing the target access curve from the extreme point to obtain a plurality of first curve segments, wherein the target access curve is any one of the access curves; Determining the importance of the candidate time periods corresponding to the first curve segments according to the number of visits of the target access curve in the first curve segments; Each of the candidate time periods whose importance is greater than a preset importance threshold is determined as the access feature time period.

3. The method for optimizing storage of archival data according to claim 2, characterized in that: The determining, according to the access amount of the target access curve in each of the first curve segments, the importance of the candidate time period corresponding to each of the first curve segments, includes: For each of the first curve segments, the following steps are performed respectively: In the target access curve, obtaining the minimum access volume of the first curve segment and the candidate duration of the candidate time period corresponding to the first curve segment; Determine, according to each reference access curve, a reference access volume mean and a reference duration mean corresponding to the first curve segment, wherein the reference access curve is each access curve in the system node except the target access curve, the reference access volume mean is used to characterize the access volume mean of the second curve segment corresponding to the first curve segment in each reference access curve, and the reference duration mean is used to characterize the duration mean of each second curve segment; The importance of the candidate time period corresponding to the first curve segment is determined by using the minimum access volume, the candidate time period, the reference access volume average value and the reference time period average value.

4. The method for optimizing storage of archival data according to claim 1, characterized in that: The system nodes are grouped according to the access characteristic time period of each of the system nodes to obtain a plurality of system node groups, including: Determining a first similarity between the first system node and the second system node according to the importance of the access characteristic time period of the first system node and the importance of the access characteristic time period of the second system node, wherein the first system node and the second system node are any two different system nodes; Determining, based on the first similarity, a possibility that the first system node and the second system node are in the same group; When the same-group possibility is greater than a preset possibility threshold, it is determined that the first system node and the second system node belong to the same system node group.

5. The method for optimizing storage of archival data according to claim 4, characterized in that: The number of the access feature time periods is multiple; The determining, according to the importance of the access characteristic time period of the first system node and the importance of the access characteristic time period of the second system node, the first similarity between the first system node and the second system node includes: Obtaining a minimum importance value among the importance values ​​of each of the access feature time periods of the first system node; The importance of the zth access characteristic period of the first system node is subtracted from the importance of the zth access characteristic period of the second system node to obtain the absolute value of the difference between the importance of the first system node and the second system node, where z is a positive integer; The first similarity between the first system node and the second system node in the zth access feature time period is determined by using the absolute value of the importance difference and the minimum importance value.

6. The method for optimizing storage of archival data according to claim 4, characterized in that: The number of the access feature time periods is multiple; After determining the first similarity between the first system node and the second system node according to the importance of the access characteristic time period of the first system node and the importance of the access characteristic time period of the second system node, the method further includes: determining a second similarity between the first system node and the second system node according to the first similarity between the first system node and the second system node in each of the access feature time periods; The determining, according to the first similarity, a possibility that the first system node and the second system node are in the same group includes: A possibility that the first system node and the second system node are in the same group is determined according to the second similarity.

7. The method for optimizing storage of archival data according to claim 6, characterized in that: The determining, according to the second similarity, a possibility that the first system node and the second system node are in the same group includes: Acquire a response duration between the first system node and the second system node, and acquire a target similarity between the first system node and a third system node other than the second system node, wherein the target similarity is used to characterize a maximum value of the second similarities between the first system node and each of the third system nodes; The possibility that the first system node and the second system node are in the same group is determined by using the second similarity, the response duration, and the target similarity between the first system node and the second system node.

8. The method for optimizing storage of archival data according to claim 1, characterized in that: The creating a backup storage space for each of the system nodes according to the backup reliability of each of the system nodes and the number of system nodes in the system node group to which the system node belongs includes: For each of the system nodes, perform the following steps respectively: Determine a backup evaluation parameter of the system node by using the backup reliability of the system node and the number of system nodes in the system node group to which the system node belongs; Rounding the backup evaluation parameter upward to an integer to obtain the number of backup storage devices of the system node; A backup storage space of the system node is created according to the number of backup storage devices of the system node.

9. An optimized storage system for archival data, characterized in that: The system comprises: A curve graph acquisition module, used to acquire an access curve of each archive page in each system node of the archive storage system, wherein the access curve is used to represent a time series curve of the number of visits to the archive page by users in the system node; A time period determination module, used to determine the access characteristic time period of each of the system nodes according to each of the access curves in each of the system nodes, wherein the access characteristic time period is used to characterize the time period that conforms to the access habits of users in the system nodes to the archives; A node grouping module, used to group the system nodes according to the access characteristic time period of each of the system nodes to obtain a plurality of system node groups; A reliability determination module is used to determine the backup reliability of each system node according to the node storage performance of each system node in the system node group for each system node group; the step of determining the backup reliability of each system node according to the node storage performance of each system node in the system node group for each system node group includes: For each of the system node groups, perform the following steps respectively: Determine the information complexity of the target system node group by using the possibility of the system nodes in the target system node group being in the same group, wherein the target system node group is any one of the system node groups; The calculation formula for information complexity includes: ;in, Used to characterize the information complexity of the a-th system node group, It is used to characterize the maximum value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to characterize the mean value of the possibility of the system nodes in the a-th system node group being in the same group. It is used to represent the number of system nodes contained in the a-th system node group. Used to represent the total number of system nodes in the archive storage system; Determine the backup reliability of each of the system nodes in the target system node group by using the information complexity of the target system node group and the node storage performance of each of the system nodes in the target system node group; The calculation formula for backup reliability includes: ;in, It is used to characterize the backup reliability of the cth system node in the ath system node group. It is used to characterize the average of the maximum throughput of the storage devices of each system node in the a-th system node group. It is used to characterize the maximum throughput of the storage device of the cth system node in the ath system node group. Used to represent the total number of storage devices in the a-th system node group; A space creation module, used to create a backup storage space for each of the system nodes according to the backup reliability of each of the system nodes and the number of system nodes in the system node group to which the system node belongs; The archive storage module is used to back up and store each of the archive pages through the backup storage space of each of the system nodes.

Citation Information

Patent Citations

  • Database performance analysis method and device, electronic equipment and storage medium

    CN114610588A

  • Distributed data storage planning method and system for improving disaster recovery capability of data center

    CN118113526A