An audit data tiering storage method and system based on data heat

CN122614964APending Publication Date: 2026-08-21BEIJING ZHIZHENYUN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610783812.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]但是在审计数据分级存储过程中,存在多重干扰因素的影响,导致分级存储的效果较差,直接影响后续的审计工作的开展;传统的审计数据分级存储方法的分级标准缺乏统一性和客观性,仅通过数据敏感程度主观划分等级,未结合访问频率、业务价值等量化指标,并且审计数据的访问频率会随审计阶段、业务周期发生波动,而静态分级策略无法实时响应这种变化,导致分级结果与实际使用需求差异较大甚至出现分级结果失效;因此传统的审计数据分级存储策略无法及时发现分级偏差并进行调整,长期积累后会导致存储资源浪费或审计效率下降

Benefits of technology

本申请为解决传统的审计数据分级存储策略无法及时发现分级偏差并及时进行调整,导致后续存储资源浪费以及审计效率下降的问题,提出基于数据热度的审计数据分级存储方法,首先基于多源异构审计数据的预处理结果,结合不同审计数据在不同时间段的热度动态变化特征,进行审计数据阶段性热度变化敏感性特征的响应分析,根据分析结果实现各子类审计数据综合访问热度与业务使用活跃度的精准评估,从而可以精准分析数据热度的时序波动以及关联响应特征,提高基于数据热度的审计数据的特征分析结果;进一步根据基于数据热度的审计数据的热度特征分析结果进行分级存储的动态适配与等级调整判断,实现审计数据热度等级与存储资源、安全防护等级的精准匹配,提高后续审计数据检索响应速度以及存储资源利用率,避免出现审计效率下降及存储成本冗余问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614964A_ABST
    Figure CN122614964A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data security, in particular to an audit data hierarchical storage method and system based on data heat, which comprises the following steps: obtaining audit data of each audit data type of an audit institution and preprocessing to obtain each sub-class audit data; counting each heat feature data of each sub-class audit data, obtaining a heat sensitive response value through the change trend of each heat feature data and the correlation between different heat feature data; calculating a heat evaluation value of each group of heat feature data through the significant degree of the heat sensitive response value of each heat feature data in the sub-class audit data; predicting the heat evaluation value corresponding to each sub-class audit data, and dividing the sub-class audit data into heat grades according to the predicted heat evaluation value, and then performing hierarchical storage on the audit data. The application can improve the hierarchical storage efficiency of the audit data and realize efficient and accurate hierarchical storage of the audit data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, specifically to a hierarchical storage method and system for audit data based on data popularity. Background Technology

[0002] In a digital audit system, tiered storage of audit data is crucial for compliance assurance, resource optimization, and efficiency improvement. During the process of classifying and tiering data management, core data, important data, and general data must be stored differently. Core data must meet the requirements of a Level 4 network security rating, while important data must meet Level 3 or higher protection standards. Tiered storage allows frequently accessed data to be deployed on high-performance storage media, while low-frequency data is stored on low-cost archiving media, ensuring audit continuity and reducing IT infrastructure costs. A scientifically effective tiered strategy can reduce storage costs and improve the response speed for accessing core data.

[0003] However, the hierarchical storage of audit data is affected by multiple interfering factors, resulting in poor hierarchical storage performance and directly impacting subsequent audit work. Traditional hierarchical storage methods for audit data lack uniformity and objectivity in their hierarchical standards, subjectively classifying levels based solely on data sensitivity without considering quantitative indicators such as access frequency and business value. Furthermore, the access frequency of audit data fluctuates with audit stages and business cycles, and static hierarchical strategies cannot respond to these changes in real time, leading to significant discrepancies between the hierarchical results and actual usage needs, or even hierarchical results becoming invalid. Therefore, traditional hierarchical storage strategies for audit data cannot promptly detect and adjust hierarchical deviations, which, over time, can lead to wasted storage resources or decreased audit efficiency. Summary of the Invention

[0004] To address the aforementioned technical problems, the purpose of this application is to provide a hierarchical storage method and system for audit data based on data popularity. The specific technical solution adopted is as follows: This application provides a method for hierarchical storage of audit data based on data popularity, including the following steps: Obtain and preprocess audit data of various audit types from the audit firm to obtain audit data of each sub-category; Statistically analyze the popularity characteristics of each sub-category of audit data, and use all the popularity characteristics of each sub-category of audit data within each preset time window as the set of popularity characteristics of each sub-category of audit data; By analyzing the changing trends of various heat characteristic data and the correlation between different heat characteristic data, the heat sensitivity response value of each sub-category of audit data is measured; The popularity of each group of popularity feature data in the sub-category audit data is evaluated by assessing the significance of the popularity sensitivity response values ​​of each popularity feature data in the sub-category audit data, and the popularity evaluation value of each group of popularity feature data is obtained. The popularity assessment value corresponding to each sub-category of audit data is predicted, and the popularity level of the sub-category audit data is divided according to the predicted popularity assessment value, and then the audit data is stored in a hierarchical manner.

[0005] Preferably, the popularity feature data includes the number of accesses to the subclass audit data, the average duration of a single access, and the number of concurrent accesses.

[0006] Preferably, before obtaining the heat sensitivity response value, for any sub-category of audit data, each heat feature data is arranged in chronological order, and trend statistics are performed on the arranged heat feature data to obtain trend statistics.

[0007] Preferably, the average absolute value of the Pearson correlation coefficient between various heat feature data and all other heat features in each sub-category of audit data is calculated, as well as the absolute value of the trend statistic of various heat feature data in each sub-category of audit data. The heat sensitivity response value is positively correlated with the average value and the absolute value, respectively.

[0008] Preferably, before obtaining the heat assessment value, for each type of audit data, the ratio of the heat sensitivity response value of each heat feature data to the sum of the heat sensitivity response values ​​of all heat feature data is calculated and used as the sensitivity response weight of each heat feature data.

[0009] Preferably, the specific formula for obtaining the popularity evaluation value of each group of popularity feature data for any sub-category of audit data is as follows: In the formula, The first subclass of audit data represents the audit data of that subclass. The heat assessment value of the group's heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data The normalization result of the heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data Sensitive response weights for heat signature data; This indicates the number of types of heat feature data.

[0010] Preferably, the heat assessment values ​​of all groups of heat characteristic data of each sub-category of audit data are used to make predictions using a time series prediction algorithm, and the prediction results of the heat assessment values ​​corresponding to each sub-category of audit data are obtained. After normalization processing, the normalized results of the predicted values ​​of each heat assessment value are obtained.

[0011] Preferably, the sub-category audit data is divided into three heat levels using the normalized result of the predicted heat assessment value. Sub-category audit data with a normalized result of the predicted heat assessment value greater than or equal to a first preset value is classified as high-heat data. Sub-category audit data with a normalized result of the predicted heat assessment value greater than or equal to a second preset value and less than the first preset value is classified as medium-heat data. Sub-category audit data with a normalized result of the predicted heat assessment value less than the second preset value is classified as cold data. The first preset value is greater than the second preset value.

[0012] Preferably, all types of audit data are divided into core data, important data, and general data. Core data storage has a higher priority than important data storage, which in turn has a higher priority than general data storage. Among core data, hot data, warm data, and cold data have the same storage priority. Among important data, hot data storage has a higher priority than warm data storage, which in turn has a higher priority than cold data storage. Among general data, hot data storage has a higher priority than warm data storage, which is equal to cold data storage.

[0013] This application also provides a hierarchical storage system for audit data based on data popularity, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any of the above-described hierarchical storage methods for audit data based on data popularity.

[0014] As can be seen from the above, the hierarchical storage method and system for audit data based on data popularity provided in this application has at least the following beneficial effects: This application addresses the problem that traditional hierarchical storage strategies for audit data fail to promptly detect and adjust hierarchical deviations, leading to wasted storage resources and decreased audit efficiency. It proposes a hierarchical audit data storage method based on data popularity. First, based on the preprocessing results of multi-source heterogeneous audit data, and combined with the dynamic changes in the popularity of different audit data over different time periods, a response analysis of the sensitivity characteristics of audit data to periodic popularity changes is conducted. Based on the analysis results, a precise assessment of the comprehensive access popularity and business activity of each sub-category of audit data is achieved. This allows for accurate analysis of the temporal fluctuations and associated response characteristics of data popularity, improving the feature analysis results of audit data based on data popularity. Furthermore, based on the popularity feature analysis results of the audit data, dynamic adaptation and hierarchical storage adjustments are made, achieving precise matching between audit data popularity levels and storage resources and security protection levels. This improves the subsequent audit data retrieval response speed and storage resource utilization, avoiding problems such as decreased audit efficiency and redundant storage costs. Attached Figure Description

[0015] Figure 1 A flowchart illustrating the steps of a hierarchical storage method for audit data based on data popularity, as provided in this application. Detailed Implementation

[0016] The following, in conjunction with the accompanying drawings, details the specific scheme of the hierarchical storage method and system for audit data based on data popularity provided in this application.

[0017] Please see Figure 1 The diagram illustrates a flowchart of a hierarchical storage method for audit data based on data popularity, according to an embodiment of this application, including the following steps: Step 1: Obtain and preprocess the audit data of each audit type from the audit firm to obtain the audit data of each sub-category.

[0018] To ensure that the acquired data meets audit requirements and is usable, in this embodiment, before acquiring data, the auditing firm needs to execute a data requirements specification based on the designed objectives. This clarifies the name of the system to be collected, the data content, format requirements, collection time limit, and the responsibilities of both parties. Furthermore, the auditing firm needs to clarify the scope and nature of core and important data with the auditee through an engagement letter or confirmation letter. The specific data acquisition process is as follows: For financial software conforming to the GB / T19581-2004 interface standard, data is acquired directly through the standard interface; for heterogeneous databases, cross-platform data reading can be achieved through the ODBC interface; for unstructured data or special format files, the auditee can convert them to the specified format under the supervision of the auditors and then acquire them through encrypted file transfer. During the collection process, all operation records must be kept, including the collection time, operators, data source, and collection volume. The auditing firm and the auditee's designated personnel must complete handover procedures to ensure the authenticity and completeness of the data acquired from the auditee.

[0019] Furthermore, to reduce the impact on data quality during data transmission and collection, this embodiment first uses SQL statements to query and obtain missing values ​​and duplicate records in the audit data. For missing key fields, they can be supplemented after confirmation with the audited entity. For duplicate records, duplicates are removed by comparing key fields to ensure data uniqueness. After the above processing, data backup is performed to provide security for subsequent processing.

[0020] Thus, according to the above process in this embodiment, various types of audit data can be obtained by the auditing institution. In this embodiment, the audit data types include core business data, financial data, operational data, and log data, which are denoted as sub-types of audit data.

[0021] Step 2: Statistically analyze the popularity characteristics of each sub-category of audit data, and use all the popularity characteristics of each sub-category of audit data within each preset time window as the popularity characteristic data of each sub-category of audit data.

[0022] Due to the differences in data attributes and storage requirements of audit data, it is necessary to store audit data in a tiered manner. In this embodiment, the audit data types include core business data, financial data, operational data, and log data. In actual application scenarios, the audit data types are not limited to the above-mentioned audit data types in this embodiment, and this embodiment does not impose any restrictions on them. The access popularity, business importance, security level, and compliance retention period of different subcategories of each type of audit data are different. The access frequency, security requirements, and retention periods of different subcategories of data vary greatly. If a uniform storage strategy is adopted, it may affect the response speed of core business. Therefore, it is necessary to achieve matching between audit data and corresponding performance resources through tiered storage to improve the response speed of audit data.

[0023] Furthermore, considering the poor dynamic adaptability of traditional tiered storage of audit data, and that data popularity directly reflects the actual frequency of data use and the urgency of business operations, high-frequency access audit data typically corresponds to key evidence and real-time business data during the audit implementation phase, requiring low-latency storage support. Medium-frequency access audit data usually consists of intermediate results and periodic reports during the audit process, with moderate performance requirements for storage. Low-frequency access audit data is typically historical archives and compliance-retained data. Therefore, to avoid poor tiered storage performance due to relying solely on the sensitivity of audit data, this embodiment combines data popularity to accurately match the storage resources and usage requirements of audit data. By analyzing the dynamic changes in popularity, the tiered strategy is adjusted in real time, enabling high-performance storage media to centrally serve high-value, high-popularity data, thereby improving the overall operational efficiency of the storage architecture.

[0024] In the hierarchical storage of audit data, due to the significant differences in access characteristics among different types of audit data, a single indicator cannot accurately reflect the true popularity of the data. Therefore, historical access characteristic data for each type of audit data is statistically analyzed. This historical access characteristic data includes the number of accesses, average duration of a single access, and number of concurrent accesses for each sub-category of audit data within each type of audit data. The access characteristic data collected above can reflect the actual usage popularity and search activity of the corresponding sub-category of audit data. Furthermore, considering that audit data access has a time-series periodicity, and that short-term access characteristics better reflect real-time popularity value, all popularity characteristic data for each sub-category of audit data within each preset number of days are used as the set of popularity characteristic data for each sub-category of audit data. Preferably, in this embodiment, for each sub-category of audit data within the last 3 years, a statistical time window is set to 10 days. The number of accesses, duration of a single access, and number of concurrent accesses for that sub-category of audit data within 10 days are used as a set of popularity characteristic data for subsequent quantitative modeling, hierarchical classification, and precise hierarchical storage determination of audit data popularity, thereby obtaining all sets of popularity characteristic data for each sub-category of audit data within the last 3 years.

[0025] Step 3: Measure the heat sensitivity response value of each sub-category of audit data by analyzing the changing trends of each heat characteristic data and the correlation between different heat characteristic data.

[0026] Furthermore, since the audit data for each sub-category exhibits dynamic changes in popularity and differences in access characteristics across different time periods, a time-series correlation and popularity sensitivity analysis is performed on the popularity characteristic data of all groups within the past three years for each sub-category of audit data. This determines the popularity fluctuation characteristics of each sub-category of audit data over time. Specifically, for any sub-category of audit data, each type of popularity characteristic data is first sorted according to time sequence. Trend statistics are then performed on the sorted popularity characteristic data. In this implementation, the Mk trend verification algorithm is used to perform trend statistics on the sorted popularity characteristic data, obtaining the corresponding trend statistic. The larger the trend statistic, the greater the trend in the popularity of the data over time. The more significant the growth, the more significant the correlation between the data popularity and access characteristics of the audit data. Secondly, as the data popularity of a sub-category of audit data changes dynamically over time, access to that sub-category of audit data exhibits dynamic correlation characteristics. Therefore, the Pearson correlation coefficient is calculated between each popularity characteristic data of each sub-category of audit data and each other popularity characteristic data. The larger the absolute value of the Pearson correlation coefficient, the more significant the correlation between the multi-dimensional access characteristics of the sub-category of data and the dynamic changes in popularity. Through the above analysis, time-series trend characteristic analysis and dynamic correlation characteristic analysis were performed on each sub-category of audit data under the influence of dynamic changes in popularity, providing a quantitative basis for subsequent dynamic hierarchical storage and adaptive allocation of storage resources.

[0027] Furthermore, based on the above analysis and processing, the heat sensitivity response characteristics of each sub-category of audit data are analyzed, and the heat sensitivity response value of each sub-category of audit data is calculated. The larger the heat sensitivity response value, the more significant the sensitivity of the heat characteristic data in that sub-category of audit data to dynamic changes in heat. Specifically, the average absolute value of the Pearson correlation coefficient between each sub-category of audit data and all other heat characteristics, as well as the absolute value of the trend statistic of each sub-category of audit data, are calculated. The heat sensitivity response value is positively correlated with the average value and the absolute value, respectively. In this embodiment, the positive correlation means that the variables have the same trend of change. The specific calculation relationship can be selected by the implementer at their own discretion, and this embodiment does not impose any restrictions on this.

[0028] Preferably, in this embodiment, the specific calculation formula for the heat-sensitive response value is as follows: ,in Indicates the first The first seed audit data The heat sensitivity response value of heat characteristic data; Indicates the first The first seed audit data The average of the absolute values ​​of the Pearson correlation coefficients between this heat feature data and all other heat features; Indicates the first The first seed audit data The absolute value of the trend statistic for this type of popularity characteristic data should be noted. It should be noted that whether the number of visits increases explosively or drops precipitously, it indicates that the indicator is an active variable in the current audit stage and has extremely high sensitivity. Therefore, the absolute value of the trend statistic is used for popularity sensitivity response analysis.

[0029] Step 4: By assessing the significance of the heat sensitivity response values ​​of each heat feature data in the sub-category audit data, evaluate the heat of each group of heat feature data in the sub-category audit data and obtain the heat evaluation value of each group of heat feature data.

[0030] Based on the above calculations and analysis, a comparative correlation analysis is conducted on the sensitivity response characteristics of each heat feature data in each sub-category of audit data under dynamic changes in heat features to obtain the heat sensitivity response value of each heat feature data. Furthermore, considering that the sensitivity characteristics of each heat feature data in each sub-category of audit data are different under dynamic changes in heat, the heat sensitivity response value of each heat feature data is combined to evaluate the heat of various sub-category audit data, analyze the comprehensive access activity and real-time heat level of each sub-category of audit data, and thus serve as the core quantitative judgment basis for audit data level classification, storage medium matching, and dynamic storage migration scheduling.

[0031] Specifically, for each type of audit data, the ratio of the sensitivity response value of each type of heat feature data to the sum of the sensitivity response values ​​of all heat feature data is calculated and used as the sensitivity response weight of each type of heat feature data. That is, compared with the sensitivity characteristics of other heat feature data under this type of audit data to dynamic changes in heat, the current heat feature data has a more significant sensitivity characteristic to dynamic changes in heat.

[0032] Furthermore, based on the sensitivity response of each heat characteristic data to the dynamic changes in the heat of sub-category audit data determined above, the heat assessment value of each sub-category audit data in different time periods is calculated. Specifically, for each sub-category audit data, the maximum-minimum normalization algorithm is used to normalize various heat characteristic data to obtain normalized heat characteristic data, avoiding the impact of the difference in the dimensions of different heat characteristic data on the accuracy of heat assessment. It should be noted that in the process of performing maximum-minimum normalization, the global maximum and global minimum values ​​of each heat characteristic data of each sub-category audit data are determined by the historical heat characteristic data of the most recent 3 years for normalization processing. Furthermore, for each set of heat characteristic data of any sub-category audit data, combined with the sensitivity response characteristics of each heat characteristic data to the dynamic changes in heat, the heat assessment value of each set of heat characteristic data is calculated, and the calculation relationship is as follows: ,in The first subclass of audit data represents the audit data of that subclass. The heat assessment value of the group's heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data The normalization result of the heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data The sensitivity response weight of the heat feature data, that is, the larger the sensitivity response weight, the better the response in the heat feature data. In the process of evaluating the popularity of a group of popularity feature data, the more significant its contribution and influence on the overall popularity evaluation value; This indicates the number of types of heat characteristic data; the larger the calculated heat assessment value, the higher the overall access heat of the audit data for the corresponding time period of the heat characteristic data, the stronger the business usage activity, and the higher the corresponding storage priority and high-performance storage resource configuration requirements.

[0033] Step 5: Predict the popularity assessment value corresponding to each type of audit data, and classify the audit data of each type of audit data into popularity levels based on the predicted popularity assessment value, and then store the audit data in a hierarchical manner.

[0034] Based on the above calculation and analysis results, for each type of audit data, the popularity characteristic data obtained at different time periods are used to calculate the popularity evaluation value of each group of popularity characteristic data. This reflects the overall access popularity and business activity of the sub-category audit data corresponding to that time period. Furthermore, to determine the dynamic changes in popularity of each type of audit data, the popularity evaluation values ​​of all groups of popularity characteristic data for each sub-category audit data are arranged chronologically and used as input for a time series prediction algorithm. In this embodiment, the ARIMA algorithm is used for predictive analysis to obtain the predicted results of the popularity evaluation values ​​corresponding to each sub-category audit data. This assesses the rise and fall and fluctuation cycle of audit data access popularity in subsequent business cycles. The purpose is to predict data popularity characteristics in advance, realize pre-emptive storage resource allocation and pre-tiered scheduling, and accurately evaluate and analyze the overall access popularity and business activity of each type of audit data based on the obtained predicted popularity evaluation values. In this embodiment, the preset duration is 15 days.

[0035] Furthermore, the audit data is classified into popularity levels based on the predicted popularity assessment value for each sub-category of audit data. Specifically, the popularity assessment value is first normalized to its maximum value. The specific normalization process is a well-known technique and will not be elaborated in this embodiment. Further, the sub-category audit data is divided into three popularity levels using the normalized result of the predicted popularity assessment value: high-popularity data, medium-popularity data, and low-popularity data. Preferably, in this embodiment, the sub-category audit data with a normalized result of the predicted popularity assessment value greater than or equal to 0.7 is considered high-popularity data; the sub-category audit data with a normalized result of the predicted popularity assessment value greater than or equal to 0.4 and less than 0.7 is considered medium-popularity data; and the sub-category audit data with a normalized result of the predicted popularity assessment value less than 0.4 is considered low-popularity data. The audit data is then stored in a tiered manner. In this embodiment, the data representation of high-hot data has a high overall access frequency, strong business dependence, and urgent real-time retrieval needs, requiring priority configuration of high-performance, low-latency storage resources and high-availability online resident storage; the data representation of warm data has a medium access frequency, many periodic business calls, and no extreme low latency requirements, requiring adaptation of mid-range storage resources that balance performance and cost; the data representation of cold data has an extremely low access frequency, is only used for compliance archiving and post-event review, and has weak real-time access needs, requiring large-capacity, low-cost, archive-style long-term storage management.

[0036] Preferably, in this embodiment, to ensure the real-time, dynamic, and accurate classification of audit data popularity levels, thereby enabling the hierarchical storage strategy to adaptively adjust with data access behavior and avoiding the mismatch problem caused by using static classification rules, where hot data is stored in low-end storage and cold data occupies high-performance storage resources, this embodiment updates the popularity feature data hourly through the access logs of the storage system. When the change in popularity assessment value exceeds ±0.1, a temporary popularity level adjustment is triggered, that is, the popularity level is reclassified based on the calculated popularity assessment value according to the above steps. Specifically, if the change in the popularity assessment value predicted by ARIMA exceeds ±0.1, a temporary level adjustment is immediately triggered to compensate for prediction deviations or sudden access.

[0037] To achieve a precise match between the popularity level and data security level of audit data, determine the final hierarchical storage strategy, and realize the optimal configuration of storage performance, security protection level and business data value, this embodiment first divides all types of audit data into three categories: core data, important data and general data, based on the industry standards of the audited entity. For example, core business permission change audit data and financial audit data of fund flow are core data, while routine business process audit data and system key operation status audit data are important data.

[0038] Furthermore, since core data involves the lifeline of business, financial security, and the highest level of confidentiality, it cannot be arbitrarily downgraded in storage based solely on popularity. Therefore, the storage priority of core data is higher than that of important data, which in turn is higher than that of general data. Among core data, high-frequency data, warm data, and cold data have the same storage priority. Among important data, high-frequency data has a higher storage priority than warm data, which in turn is higher than cold data. Among general data, high-frequency data has a higher storage priority than warm data, which is equal to cold data. Preferably, in this embodiment, core data, regardless of its popularity level, is preferentially deployed in Level 1 storage and protected by Level 4 network security, including network physical isolation, minimum access permission authorization, and transmission encryption. For important data, the audit data of each subcategory within the important data is classified by popularity level to determine the corresponding hot data, warm data, and cold data. High-popularity data is deployed in Level 1 storage, warm data in Level 2 storage, and cold data in Level 3 storage, and the storage system for all important data meets the requirements of Level 3 or higher network security protection. For general data, the audit data of each subcategory within the general data is also classified by popularity level to determine the corresponding hot data, warm data, and cold data. High-popularity data is deployed in Level 2 storage, while warm and cold data are deployed in Level 3 storage or archive storage, using basic security protection measures.

[0039] It should be noted that implementing tiered storage deployment in the above-mentioned tiered storage process ensures the effective execution of the storage strategy. Tier 1 storage uses an all-flash array, offering high IOPS and low latency performance, while supporting real-time data read / write and high-frequency access. Tier 2 storage employs a hybrid storage architecture of SSDs and HDDs, balancing performance and cost to meet mid-frequency access requirements. Tier 3 storage uses an object storage system, characterized by high capacity and low cost, supporting massive data storage and on-demand access. Archive storage uses tape libraries or low-cost cloud archiving services to meet long-term compliant retention requirements, with access latency acceptable within 24 hours.

[0040] Based on the same inventive concept as the above method, this application embodiment also provides an audit data hierarchical storage system based on data popularity, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described audit data hierarchical storage methods based on data popularity.

Claims

1. A hierarchical storage method for audit data based on data popularity, characterized in that, Includes the following steps: Obtain and preprocess audit data of various audit types from the audit firm to obtain audit data of each sub-category; Statistically analyze the popularity characteristics of each sub-category of audit data, and use all the popularity characteristics of each sub-category of audit data within each preset time window as the set of popularity characteristics of each sub-category of audit data; By analyzing the changing trends of various heat characteristic data and the correlation between different heat characteristic data, the heat sensitivity response value of each sub-category of audit data is measured; The popularity of each group of popularity feature data in the sub-category audit data is evaluated by assessing the significance of the popularity sensitivity response values ​​of each popularity feature data in the sub-category audit data, and the popularity evaluation value of each group of popularity feature data is obtained. The popularity assessment value corresponding to each sub-category of audit data is predicted, and the popularity level of the sub-category audit data is divided according to the predicted popularity assessment value, and then the audit data is stored in a hierarchical manner.

2. The hierarchical storage method for audit data based on data popularity as described in claim 1, characterized in that, The popularity feature data includes the number of accesses to subclass audit data, the average duration of a single access, and the number of concurrent accesses.

3. The hierarchical storage method for audit data based on data popularity as described in claim 1, characterized in that, Before obtaining the heat sensitivity response value, for any subclass of audit data, each heat feature data is arranged in chronological order, and trend statistics are performed on the arranged heat feature data to obtain the trend statistics.

4. The hierarchical storage method for audit data based on data popularity as described in claim 1, characterized in that, Calculate the average absolute value of the Pearson correlation coefficient between various heat characteristic data and all other heat characteristics in each sub-category of audit data, and the absolute value of the trend statistic of various heat characteristic data in each sub-category of audit data. The heat sensitivity response value is positively correlated with the average value and the absolute value, respectively.

5. A hierarchical storage method for audit data based on data popularity as described in claim 1, characterized in that, Before obtaining the heat assessment value, for each type of audit data, the ratio of the heat sensitivity response value of each heat feature data to the sum of the heat sensitivity response values ​​of all heat feature data is calculated and used as the sensitivity response weight of each heat feature data.

6. The hierarchical storage method for audit data based on data popularity as described in claim 5, characterized in that, The specific formula for obtaining the popularity evaluation value of each set of popularity feature data for any subclass of audit data is as follows: In the formula, The first subclass of audit data represents the audit data of that subclass. The heat assessment value of the group's heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data The normalization result of the heat characteristic data; The first subclass of audit data represents the audit data of that subclass. The first group of heat characteristic data Sensitive response weights for heat signature data; This indicates the number of types of heat feature data.

7. The hierarchical storage method for audit data based on data popularity as described in claim 1, characterized in that, Using the popularity assessment values ​​of all groups of popularity feature data of each sub-category of audit data, a time series prediction algorithm is used to predict the popularity assessment values ​​corresponding to each sub-category of audit data. After normalization processing, the normalized results of the predicted popularity assessment values ​​are obtained.

8. The hierarchical storage method for audit data based on data popularity as described in claim 7, characterized in that, The audit data of the subclass is divided into three heat levels based on the normalized result of the predicted heat assessment value. The audit data of the subclass with a normalized result of the predicted heat assessment value greater than or equal to the first preset value is high heat data. The audit data of the subclass with a normalized result of the predicted heat assessment value greater than or equal to the second preset value and less than the first preset value is medium heat data. The audit data of the subclass with a normalized result of the predicted heat assessment value less than the second preset value is cold heat data. The first preset value is greater than the second preset value.

9. A hierarchical storage method for audit data based on data popularity as described in claim 8, characterized in that, All types of audit data are divided into core data, important data, and general data. Core data storage has a higher priority than important data storage, which in turn has a higher priority than general data storage. Among core data, hot data, warm data, and cold data have the same storage priority. Among important data, hot data storage has a higher priority than warm data storage, which in turn has a higher priority than cold data storage. Among general data, hot data storage has a higher priority than warm data storage, which is equal to cold data storage.

10. A hierarchical storage system for audit data based on data popularity, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the hierarchical storage method for audit data based on data popularity as described in any one of claims 1-9.