Data cold and hot hierarchical storage and data intelligent scheduling method and device
By using a multi-dimensional evaluation model and an intelligent scheduling hub, the problems of misjudging hot and cold data and low efficiency of manual scheduling in financial data storage have been solved, achieving efficient and automated data management and real-time querying.
Patent Information
- Application Number
- CN202511677043.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-10
AI Technical Summary
Existing financial data storage and management technologies fail to accurately distinguish between hot and cold data, resulting in wasted storage resources, difficulties in data traceability, low efficiency of manual scheduling, and inability to meet real-time query needs.
Construct a multi-dimensional hot and cold evaluation model, dynamically allocate weights based on business characteristics, design a storage-deletion/modification feature adaptation architecture, build an intelligent scheduling hub, realize automated data flow and full-process monitoring, and establish an intelligent reporting mechanism.
It enables accurate determination of data hot/cold attributes, optimizes storage resource utilization, improves the automation level and response speed of data management, and meets the requirements of the financial field for data real-time performance and reliability.
Smart Images

Figure CN121501221A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage technology, and in particular to a method and apparatus for cold and hot data tiered storage and intelligent data scheduling. Background Technology
[0002] In the financial and corporate finance sectors, financial data (such as commission accrual and expense posting data) is characterized by its dynamic change in business value over time, the need for frequent adjustments to some data, and extremely high requirements for real-time performance and traceability. Its full lifecycle management directly impacts the efficiency of core business processes such as financial settlement and auditing. Current financial data storage and management technologies often use a single dimension (such as data generation time) to classify data as hot or cold, ignoring the unique business attributes of financial data, such as "posting status" and "rate validity." This easily leads to frequently accessed valid data being misclassified as cold data, and infrequently accessed invalid data being misclassified as hot data, resulting in a serious waste of storage resources. Furthermore, existing technologies do not differentiate between the deletion and modification characteristics of storage systems, mixing dynamically adjustable temporary data with non-deletable archived data in the same database. This not only increases the difficulty of data traceability but also increases the risk of data errors due to manual maintenance of intermediate data.
[0003] Furthermore, current financial data scheduling relies heavily on manual migration tasks performed by operations and maintenance personnel. This not only results in migration delays often exceeding 24 hours but also lacks a full-process monitoring mechanism. Data loss requires manual investigation, leading to high operations and maintenance costs. In the report query stage, data needs to be manually extracted from multiple databases, such as business databases and archive databases, and integrated to generate reports. The response time often reaches several hours, which cannot meet the real-time data query needs of scenarios such as auditing and financial settlement. These technical deficiencies severely restrict the efficiency and reliability of financial data management, and there is an urgent need for a fully automated data management solution that is adapted to the characteristics of financial business. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a method and apparatus for data cold and hot tiered storage and intelligent data scheduling, which can realize automated management of the entire data lifecycle, improve management accuracy and efficiency, and is applicable to fields with high data requirements such as finance.
[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a method for tiered cold and hot data storage and intelligent data scheduling, the method comprising: Construct a multi-dimensional hot and cold assessment model, collect at least two types of assessment dimension data related to business characteristics from the data source layer, quantify the data of each assessment dimension, dynamically allocate the weight of each assessment dimension according to the target business scenario, calculate the comprehensive score of the data based on the quantification results and weights, and determine the hot and cold attributes of the data according to the preset threshold. Based on the determined hot / cold attributes and the data deletion / modification requirements, the data is allocated to storage systems with corresponding characteristics. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and an intermediate data storage system that supports deletion / modification. An intelligent scheduling hub is established, which includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automated flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. Create intelligent reports, integrate data from various storage systems to construct thematic datasets, and automatically complete report queries, calculations, generation, push, and updates based on preset report rules.
[0006] Secondly, embodiments of this application also provide a data cold and hot tiered storage and intelligent data scheduling device, the device comprising: The judgment module is used to build a multi-dimensional hot and cold evaluation model. It collects at least two types of evaluation dimension data related to business characteristics from the data source layer, quantifies the data of each evaluation dimension, dynamically allocates the weight of each evaluation dimension according to the target business scenario, calculates the comprehensive score of the data based on the quantification results and weights, and determines the hot and cold attributes of the data according to the preset threshold. The allocation module is used to allocate data to storage systems with corresponding characteristics based on the determined hot / cold attributes and the data deletion / modification requirements. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and a deletion / modification-supporting intermediate data storage system. The scheduling module is used to build an intelligent scheduling hub. The intelligent scheduling hub includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automated flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. The reporting module is used to create intelligent reports, integrate data from various storage systems to build thematic datasets, and automatically complete the querying, calculation, generation, push and updating of reports based on preset reporting rules.
[0007] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the data cold and hot tiered storage and intelligent data scheduling method described in any of the first aspects.
[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the data cold and hot tiered storage and intelligent data scheduling method described in any one of the first aspects.
[0009] The embodiments of this application have the following beneficial effects: By constructing a multi-dimensional hot / cold data assessment model (breaking through the limitations of traditional single-dimensional models, dynamically adapting weights and judgment thresholds based on business characteristics to achieve accurate determination of data hot / cold attributes), designing a storage-deletion / modification feature adaptation architecture (matching corresponding storage systems according to hot / cold attributes and deletion / modification requirements, covering the storage needs of non-deletable cold / hot data and deletable intermediate data), building an intelligent scheduling hub that includes data flow scheduling, full-process monitoring, dynamic resource control, and data traceability (realizing automated data flow, real-time status monitoring, on-demand resource optimization, and full lifecycle trajectory traceability), and establishing an intelligent reporting mechanism (integrating multi-source data to construct thematic datasets, and automatically reporting based on preset rules), the system achieves accurate determination of data hot / cold attributes. It automatically completes report querying, calculation, generation, and updating, forming an end-to-end closed-loop management of data from cold / hot data determination, storage allocation, flow management to report output. It effectively solves problems in existing technologies such as inaccurate cold / hot data assessment leading to storage resource waste, data traceability difficulties caused by mismatch between storage and deletion / modification needs, low efficiency and high operation and maintenance costs of manual scheduling, and long report query time that cannot meet real-time needs. It significantly improves the accuracy, automation level, and resource utilization of data management, while ensuring data traceability and real-time business response. It provides efficient and stable data management solutions for fields such as finance with high requirements for data real-time performance, reliability, and compliance. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating steps S101-S104 provided in the embodiments of this application; Figure 2This is a flowchart of the multi-dimensional weighted evaluation provided in the embodiments of this application; Figure 3 This is a storage-deletion / modification flowchart provided in an embodiment of this application; Figure 4 This is a block diagram of the control layer principle provided in the embodiments of this application; Figure 5 This is a flowchart of the report generation process provided in this application embodiment; Figure 6 This is a schematic diagram of the structure of the data cold and hot tiered storage and intelligent data scheduling device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0015] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0016] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application and is not intended to limit this application.
[0018] See Figure 1 , Figure 1 This is a flowchart illustrating steps S101-S104 of the data cold and hot tiered storage and intelligent data scheduling method provided in this application embodiment, which will be combined with... Figure 1 Steps S101-S104 are explained below.
[0019] In step S101, a multi-dimensional hot and cold evaluation model is constructed. At least two types of evaluation dimension data related to business characteristics are collected from the data source layer. The data of each evaluation dimension are quantified. The weight of each evaluation dimension is dynamically allocated according to the target business scenario. The comprehensive score of the data is calculated based on the quantification results and weights. The hot and cold attributes of the data are determined according to the preset threshold. Please see Figure 2 , Figure 2 This is a flowchart of the multi-dimensional weighted evaluation provided in the embodiments of this application, such as... Figure 2 As shown, considering the characteristics of financial data that "business value decays with posting time and access demand changes with rate value status", the data collected by the data source layer specifically includes three categories: "financial posting time", "rate value effective date" and "access frequency in the past 30 days" (which can be extended to other dimensions related to financial business). The collection scope covers all business data in the financial system, such as commission accrual and expense posting, to ensure that the evaluation dimensions are strongly correlated with the financial business scenario.
[0020] To avoid subjective judgment bias, a quantitative method of "segmented assignment" is adopted, which transforms the qualitative descriptions of each dimension into quantitative scores. For example, in the "financial posting time" dimension, posting time ≤ 3 months is assigned 100 points, and posting time > 12 months is assigned 10 points. Through quantification, the evaluation standards are standardized and traceable.
[0021] Financial operations involve different scenarios such as "settlement" and "audit," and each scenario focuses on different dimensions (e.g., settlement scenarios focus more on posting time, while audit scenarios focus more on historical data access needs). Therefore, the target scenario is automatically identified by "application layer call tags" (without manual intervention), and weights are allocated accordingly to ensure that the evaluation results are adapted to the needs of the scenario.
[0022] The overall score is calculated using a weighted summation method (score = posting time score × weight + rate value effectiveness score × weight + access popularity score × weight). The preset thresholds are verified through business testing (hot data ≥ 60 points, cold data < 50 points, transitional data 50 ≤ score < 60 points). The judgment results directly provide a basis for subsequent storage allocation, avoiding problems such as "old rate value data being misjudged as cold data" and "invalid data being misjudged as hot data".
[0023] In step S102, based on the determined hot / cold attributes and the data deletion / modification requirements, the data is allocated to a storage system with corresponding characteristics. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and a deletion / modification-supporting intermediate data storage system. Existing technologies mix temporary data that needs to be deleted with archived data that cannot be deleted, making data traceability difficult. This step achieves accurate allocation through "dual-dimensional judgment"—first, it determines whether the data has "deletion and modification needs" (such as rate value data to be corrected or intermediate data to be reviewed that needs to be deleted and modified), and then combines the hot and cold attributes to ensure that the "deletion and modification characteristics" of the storage system are fully matched with the data requirements.
[0024] Please see Figure 3 , Figure 3 This is a storage-deletion / modification flowchart provided in the embodiments of this application, such as... Figure 3 As shown, based on the characteristics of the ecosystem storage products, MRS is specifically selected as the non-deletable cold data storage system (low cost, high reliability, and support for massive archiving), DWS as the non-deletable hot data storage system (near real-time computing, ad-hoc querying, and response latency ≤1 second), and GaussDB as the intermediate data storage system that supports deletion and modification (row-level deletion and modification, transaction consistency, and ACID compliance). The three types of systems form a three-level architecture of "cold-hot-intermediate" to cover the storage needs of financial data throughout its entire lifecycle.
[0025] In step S103, an intelligent scheduling center is established. The intelligent scheduling center includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automated flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts the storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. Please see Figure 4 , Figure 4 This is a block diagram of the control layer principle provided in the embodiments of this application, such as... Figure 4 As shown, the intelligent scheduling center includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. These four modules work together to form a closed loop of "scheduling-monitoring-control-traceability," avoiding delays and errors associated with manual scheduling. Data flow scheduling module: Based on the evaluation and allocation results of the preceding steps, it automatically generates flow tasks (timed / trigger-based) to ensure that data flows in an orderly manner between various storage systems; The end-to-end monitoring module collects data flow status (such as task success / failure, synchronization delay) and storage resource status (such as CPU utilization, storage utilization) in real time through log collectors and resource monitoring agents, and triggers multi-level alarms when anomalies occur; Resource dynamic control module: Automatically adjusts resources based on monitoring data (such as expansion, cleanup, and connection release) to avoid resource waste or shortage; Data traceability module: Generates a unique flow ID for each piece of data, records the entire lifecycle trajectory, and meets the traceability requirements of financial data.
[0026] The transfer methods include "timed transfer" (such as daily DWS→MRS cold data migration at dawn) and "triggering transfer" (such as GaussDB data synchronization immediately after deletion and modification confirmation), and support "breakpoint resume" and "retry mechanism" (such as automatically retrying 3 times after task failure, and triggering an alarm if the retry fails), to ensure that the transfer process is reliable and there is no data loss.
[0027] Monitoring data is pushed to the visualization dashboard in real time, allowing financial operations personnel to intuitively view task progress and resource load, such as CPU utilization of DWS and storage utilization of MRS. Abnormal situations (such as synchronization delay exceeding 10 minutes or storage utilization exceeding 90%) trigger "SMS + email" alarms to ensure that problems are detected in a timely manner.
[0028] Automatic control operations are performed based on preset thresholds. For example, when the DWS CPU utilization exceeds 80% for 5 consecutive minutes, the cloud service API is automatically called to expand the computing nodes; when the MRS storage utilization exceeds 90%, cold data that has been archived for more than 7 years and has no access records is automatically deleted without manual intervention, reducing operation and maintenance costs.
[0029] Each data transfer ID is associated with "transfer path (e.g., GaussDB→DWS→MRS), time node (e.g., synchronization time, migration time), and operator (automatic operations by the system are marked as 'scheduling center')". The complete trajectory of the data can be queried on the visualization platform through the transfer ID, which meets the compliance requirements for data traceability in the financial and accounting fields.
[0030] In step S104, an intelligent report is created, data from various storage systems is integrated to construct a thematic dataset, and the report is automatically queried, calculated, generated, pushed, and updated based on preset report rules.
[0031] Please see Figure 5 , Figure 5 This is a flowchart of the report generation process provided in an embodiment of this application, such as... Figure 5 As shown, existing technologies require manual data extraction across databases, which is inefficient. This step utilizes ClickHouse columnar storage databases to integrate DWS hot data, MRS cold data, and GaussDB intermediate data according to "business and financial themes" (such as monthly commission settlement and annual expense audit). First, multi-source data is extracted, then cleaned (unified format, removal of duplicate data), summarized (dimensional modeling to generate fact tables), and wide table designed (primary and secondary classification, separation of hot and cold data) to ensure that the dataset directly matches the financial statement requirements and avoid delays in cross-database queries.
[0032] By adopting IPA (Intelligent Process Automation) technology, finance personnel only need to configure report rules on the IPA platform (such as fields, generation time, and recipients for the "Monthly Commission Report"), and the system can automatically perform subsequent operations—automatically connecting to ClickHouse to execute query statements, aggregating and calculating the results (such as SUM commission amount), generating Excel / PDF format reports, and pushing them to preset email addresses. When the ClickHouse dataset is updated (such as when new posting data is added), the report is automatically regenerated and the update time is marked, eliminating the need for manual repetition. The report response time is reduced from "hours" to "minutes," meeting the real-time query needs of auditing and financial settlement.
[0033] In some embodiments, the evaluation dimensions include financial posting time, rate value effective date, and data access frequency; the quantification process specifically includes: Financial posting time: 100 points for posting within 3 months, 60 points for posting within 6 months, 30 points for posting within 12 months, and 10 points for posting within 12 months. Effective date of rate value: 100 points are assigned when the rate is active, 70 points are assigned when the rate is active and the remaining talent value is ≤30, 40 points are assigned when the rate has expired for ≤6 months, and 10 points are assigned when the rate has expired for >6 months. Data access popularity: ≥10 visits in the past 30 days are assigned 100 points, 5 ≤ visits <10 visits are assigned 70 points, 1 ≤ visits <5 visits are assigned 30 points, and <1 visit is assigned 10 points.
[0034] The evaluation dimensions in this application embodiment include "financial posting time", "rate value effective date", and "data access popularity". This selection is based on the core business characteristics of financial data. "Financial posting time": After financial data is posted, its business value decreases over time (e.g., data posted in the current month needs to be used frequently for settlement, while data posted a year ago only needs to be archived). Therefore, this dimension directly reflects the "timeliness value" of the data. "Rate value effective date": Financial data (such as commission data) are associated with fee rates / commission rates. Rate value data that is in effect needs to be queried frequently (such as real-time commission calculation), while expired rate value data only needs to be archived. This dimension reflects the "business effectiveness value" of the data. "Data Access Popularity": The query frequency in the past 30 days directly reflects the "actual access demand" for the data, avoiding misjudgments caused by relying solely on the time dimension (such as historical rate data that has expired but is still frequently queried).
[0035] The three dimensions provide a comprehensive assessment from the perspectives of "timeliness," "effectiveness," and "access requirements," covering the core characteristics of financial data and avoiding the limitations of existing technologies that rely on a single time dimension.
[0036] The "segmented assignment" rules for each dimension are as follows: The "Financial Posting Time" is quantified as follows: 100 points (highest score) are assigned to postings ≤ 3 months, as this period of data is needed for high-frequency scenarios such as monthly settlement and reconciliation; 60 points are assigned to postings ≤ 6 months, as this period of data is still needed for occasional queries (such as quarterly audits); 30 points are assigned to postings ≤ 12 months, which only require low-frequency queries; and 10 points (lowest score) are assigned to postings > 12 months, which only require archiving. The score gradient is perfectly matched with the business value decay law.
[0037] The “Rate Value Effective Date” is quantified as follows: 100 points are assigned to those in effect, as the rate value needs to be used in real time for commission calculation; 70 points are assigned to those with ≤30 remaining active talent values, and rate value switching needs to be prepared in advance; 40 points are assigned to those that have expired for ≤6 months, which may be used for historical data backtracking; and 10 points are assigned to those that have expired for >6 months, which only need to be archived. The score gradient is completely matched with the business effectiveness of the rate value.
[0038] The "Data Access Popularity" quantification is as follows: ≥10 accesses in the past 30 days are assigned 100 points, which is high-frequency access data (must be stored in DWS); 5 ≤ accesses <10 accesses are assigned 70 points, which is medium-frequency access data; 1 ≤ accesses <5 accesses are assigned 30 points, which is low-frequency access data; and <1 accesses are assigned 10 points, which is very infrequent access data (must be stored in MRS). The score gradient perfectly matches the actual access needs.
[0039] This quantification rule can transform the vague "hot / cold" judgment into a precise "score judgment," ensuring that the judgment results of different maintenance personnel and at different times are consistent, thereby improving the objectivity and repeatability of the assessment.
[0040] In some embodiments, the dynamic allocation of weights specifically refers to: In settlement scenarios, the weighting is as follows: financial posting time 40%, rate value effective date 35%, and data access popularity 25%. In an auditing scenario, the weightings are: financial posting time (25%), rate effective date (25%), and data access frequency (50%). The preset threshold determination rule is as follows: a comprehensive score ≥ 60 is considered hot data, a score < 50 is considered cold data, and a score ≤ 50 < 60 is considered transitional data.
[0041] Here, the weighting rules for "settlement scenarios" and "audit scenarios" are based on the core business needs of the two scenarios: "Settlement Scenarios (such as monthly commission settlement)": The core requirement is "timeliness". Priority should be given to ensuring high-frequency access to monthly posting data and efficiency value data. Therefore, "financial posting time" has the highest weight (40%), followed by "efficiency value effective date" (35%), and "access popularity" has the lowest weight (25%). For example, a piece of data that has been posted for 1 month (100 points), lost efficiency value (10 points), and accessed 5 times (70 points) has a comprehensive score of 100×40%+10×35%+70×25%=40+3.5+17.5=61 points, which is judged as hot data (stored in DWS) and meets the timeliness requirements of the settlement scenario.
[0042] "Audit Scenarios (such as annual financial audits)": The core requirement is "historical data verifiability". It is necessary to frequently query historical posting data and archiving rate data. Therefore, "data access popularity" has the highest weight (50%), while "financial posting time" and "rate value effective date" are both 25%. For example, a piece of data that has been posted for 10 months (30 points), has lost its efficiency value (10 points), and has been accessed 12 times (100 points) has a comprehensive score of 30×25%+10×25%+100×50%=7.5+2.5+50=60 points, which is judged as hot data (stored in DWS) to meet the high-frequency access requirements of historical data in audit scenarios.
[0043] The dynamic weight allocation is automatically achieved through "application layer call tags". When the financial system calls the scheduling center, the tag is marked as "settlement scenario"; when the audit system calls it, the tag is marked as "audit scenario". No manual switching is required, ensuring the automation of scenario adaptation.
[0044] The hot and cold thresholds (hot data ≥ 60 points, cold data < 50 points, transition data 50 ≤ score < 60 points) of this application embodiment are verified through the following tests: 1000 financial data entries were selected, and their storage locations were determined according to thresholds. The proportion of "high-frequency access data stored in DWS" (target ≥ 98%) and the proportion of "low-frequency access data stored in MRS" (target ≥ 98%) were statistically analyzed. The test results showed that both proportions reached 99.2%, meeting business requirements. After allocating data according to thresholds, the CPU utilization of DWS stabilized at 60%-70% (no overload), and the storage utilization of MRS stabilized at 70%-80% (no waste), achieving optimal resource utilization. Transitional data (50 ≤ score < 60) was temporarily stored in DWS and re-evaluated the following month. Statistics showed that 30% of the transitional data scored ≥ 60 points the following month (still hot data), and 70% of the transitional data scored < 50 points the following month (migrated to MRS), avoiding resource waste caused by mis-storing transitional data.
[0045] The threshold determination rules ensure that the data storage location is fully matched with business needs and resource utilization, avoiding the problem of wasting storage resources in existing technologies.
[0046] In some embodiments, the storage system characteristics and data allocation rules are as follows: The immutable cold data storage system is called MRS, which is used to store cold data and final archived data, and has the capabilities of low-cost massive storage and offline analysis. The non-deletable hot data storage system is DWS, which is used to store hot data and transition data. It has near real-time computing and ad-hoc query capabilities, with a response latency of ≤1 second. The intermediate data storage system that supports deletion and modification is GaussDB, which is used to store business data that needs to be dynamically adjusted and has row-level deletion and modification capabilities as well as transaction consistency capabilities. The data allocation is achieved through a two-step determination: the first step is to identify whether there is a need to delete or modify the data, and if so, it is preferentially allocated to GaussDB; the second step is to determine whether the data belongs to MRS or DWS based on its hot or cold attribute.
[0047] Here, the selection of the three types of storage systems (MRS, DWS, GaussDB) in this application embodiment is based on the "deletion and modification requirements" and "access frequency" of financial data, specifically according to the following: “MRS (Immutable Cold Data Storage System)”: MapReduce Service is selected, which is characterized by “immutability, low cost, massive storage, and offline analysis”. Financial cold data (such as posting data archived for more than 1 year) does not need to be deleted or modified, only low-cost long-term storage is required, and offline analysis (such as annual data statistics) needs to be supported. MRS perfectly matches this requirement. “DWS (Immutable Hot Data Storage System)”: Select data warehouse service, which has the characteristics of “immutable, near real-time computing, ad-hoc query, and response latency ≤1 second”. Financial hot data (such as monthly posting data) does not need to be deleted or modified and needs to be used frequently in scenarios such as settlement and reconciliation. Near real-time computing and low-latency query can ensure efficient business operation. DWS is a perfect match for this requirement. "GaussDB (an intermediate data storage system that supports deletion and modification)": The GaussDB relational database is selected because it supports row-level deletion and modification and transaction consistency (ACID). Financial intermediate data (such as rate value data to be corrected and commission data to be reviewed) needs to be frequently deleted and modified (such as correcting rate value errors and rejecting review data), and data consistency needs to be ensured (such as deletion and modification not affecting other related data). GaussDB perfectly meets this requirement.
[0048] The characteristics of the three types of storage systems correspond one-to-one with the "deletion, modification, and access" requirements of financial data, avoiding the problem of mixing existing storage media.
[0049] Data allocation is achieved through a "two-step decision," a logic designed to prioritize "deletion and modification needs" (one of the core business needs for financial data). The specific logic is as follows: The first step, "Deletion and Modification Requirement Determination," prioritizes determining whether data has "deletion and modification requirements." For example, "rate value data to be corrected" requires frequent modification, and "intermediate data to be reviewed" may be rejected and deleted. Such data is directly assigned to GaussDB (which supports deletion and modification). Data without deletion and modification requirements (such as confirmed posting data and archived data) proceeds to the second step of determination to avoid storing data that needs to be deleted or modified in the non-deletable MRS / DWS, which would prevent the data from being adjusted.
[0050] The second step, "Cold and Hot Attribute Determination," is based on the cold and hot data assessment results. For data that does not require deletion or modification, cold data (score < 50 points) is allocated to MRS (low-cost archiving), while hot data (score ≥ 60 points) and transitional data (50 ≤ score < 60 points) are allocated to DWS (near real-time query), ensuring that the data access frequency matches the storage system performance.
[0051] For example: "Confirmed monthly posting data" (no deletion or modification required, score 80 points) → assigned to DWS; "Rate value data to be corrected" (with deletion or modification required) → assigned to GaussDB; "Posting data archived for more than 2 years" (no deletion or modification required, score 15 points) → assigned to MRS. This dual-dimensional judgment achieves accurate matching between data and storage systems.
[0052] In some embodiments, a data synchronization step is also included: Once data deletion or modification in GaussDB is confirmed, it is automatically synchronized to the corresponding hot / cold attribute DWS or MRS. After synchronization, GaussDB retains a copy of the data and marks it as archived. When hot data in DWS is reassessed as cold data the following month, it is automatically migrated to MRS. After migration, DWS retains the index of the MRS storage path.
[0053] GaussDB stores intermediate data that needs to be deleted or modified. Once the data has been "confirmed for deletion or modification" (e.g., fixed rate values, approved), its status changes from "dynamic adjustment" to "stable archiving," and it needs to be synchronized to the non-deletable MRS / DWS. This synchronization mechanism uses "transaction-level synchronization." Before synchronization, the data in GaussDB is locked (to prevent modification during synchronization). After synchronization, a "synchronization successful" log is generated. If synchronization fails, the GaussDB data is rolled back to ensure that the data in MRS / DWS and GaussDB are consistent and to avoid data deviation. After synchronization, GaussDB retains a copy of the data and marks it as "archived," without directly deleting it. Because financial data needs to be traceable (e.g., subsequent audits require querying the original intermediate data), retaining a copy meets compliance requirements. At the same time, the "archived" mark makes it easier to distinguish between active data and historical data, improving the query efficiency of GaussDB. The synchronization task is automatically triggered by the "trigger-based flow" module of the intelligent scheduling center. When the data status in GaussDB changes to "confirmed," a synchronization task is generated immediately without manual execution. The synchronization delay is ≤5 minutes, meeting the timeliness requirements of financial data.
[0054] For example, after the "rate value data to be corrected" is corrected and confirmed in GaussDB, the scheduling center immediately triggers a synchronization task. If the data scores 75 points (hot data), it is synchronized to DWS; if the score is 45 points (cold data), it is synchronized to MRS. After synchronization, GaussDB retains a copy of the data and marks it as "archived".
[0055] DWS stores both hot and transitional data. After monthly reassessment, some hot data may become cold data (e.g., if the posting time exceeds 3 months, the score drops from 70 to 45 points), requiring migration to MRS. After migrating cold data to MRS, DWS retains only frequently accessed hot data, reducing the storage and computational load on DWS and ensuring that DWS's near real-time computing capabilities (response latency ≤ 1 second) are not affected. After migration, DWS does not delete data, only retaining the "MRS storage path index" (e.g., "MRS: / / archive / 202401 / fee_data.csv"). When the cold data needs to be queried later, the system directly jumps to the MRS storage path through the index, without needing to traverse across databases, reducing query latency from "minutes" to "seconds". The migration task is automatically executed by the intelligent scheduling center's "timed transfer" module. After the monthly reassessment is performed every morning, migration tasks are generated for DWS data determined to be cold data, supporting "breakpoint resume" (if the migration is interrupted, it will resume from the breakpoint), ensuring the reliability of the migration process.
[0056] For example, in DWS, “posting data in January 2024” (original score 70 points, reassessed score in May 2024 40 points) is automatically migrated to MRS by the scheduling center. DWS retains the MRS storage path index of this data, and subsequent queries can quickly locate it through the index.
[0057] In some embodiments, the data flow scheduling module includes two flow methods: timed flow and triggered flow. The timed flow performs cold data migration from DWS to MRS every day at midnight. The triggered flow immediately triggers the task of synchronizing to DWS or MRS after the GaussDB data deletion and modification are confirmed. The flow task supports breakpoint resumption and retry mechanisms. The full-process monitoring module collects the flow task status, data flow volume, synchronization delay time, and CPU utilization, storage utilization, and connection number of each storage system in real time through the data flow log collector and storage resource monitoring agent. It triggers SMS and email alarms when there are abnormalities. The resource dynamic control module executes the following based on monitoring data: automatically expands computing nodes when DWS CPU utilization exceeds 80% for 5 consecutive minutes; automatically deletes archived data older than 7 years with no access records when MRS storage utilization exceeds 90%; and automatically closes idle connections when GaussDB connection count exceeds 80%.
[0058] Here, you can choose to execute the migration at midnight (0:00-2:00) every day, as this period is the off-peak time for financial business (no settlement or reconciliation operations), and the migration process will not affect business operations; the migration scope is all data in DWS that are re-evaluated as cold data in the monthly period, and the migration order is in "reverse order of data generation time" (migrate the newer cold data first to ensure query priority).
[0059] Alternatively, after the deletion and modification of GaussDB data are confirmed, a task to synchronize to DWS or MRS can be triggered immediately. The trigger condition is that "GaussDB data status changes to 'confirmed'" (implemented through a database trigger). After triggering, the scheduling center immediately generates a synchronization task. The synchronization target is determined based on the data hot / cold score (hot data → DWS, cold data → MRS). The synchronization delay is ≤5 minutes, meeting the real-time requirements of financial data.
[0060] The workflow task supports breakpoint resumption and retries. Breakpoint resumption breaks the workflow task into "data block level" (e.g., every 1000 data items constitute a data block). After each data block is synchronized, a "completion mark" is recorded. If the task is interrupted (e.g., due to network failure), it will start from the "incomplete data blocks" the next time it restarts, avoiding repeated synchronization. The retry mechanism automatically retryes after a task failure (default 3 times, retry interval of 5 minutes). If all 3 retries fail, an "SMS + email" alarm is triggered and a failure log is recorded (including the reason for failure, such as incorrect data format), which is convenient for maintenance personnel to troubleshoot and ensure the reliability of the workflow task.
[0061] In addition, the monitoring content of this embodiment includes "data flow status" and "storage resource status", and triggers "SMS + email" alarms. At the data level, the flow task status (pending execution / in execution / success / failure), data flow volume, and synchronization delay time are collected in real time by "data flow log collectors" deployed in each storage system (collection frequency 1 second / time). For example, when the task status is "failed", the log records the failure time, failed data block ID, and failure reason; the data flow volume is calculated as the total migration volume daily / monthly for resource planning; the synchronization delay time is calculated as "task trigger time - task completion time". If the interval exceeds 10 minutes, it is considered abnormal. At the resource level, the following metrics are monitored: DWS CPU utilization, MRS storage utilization, and GaussDB connection count: These are collected in real-time by the "Storage Resource Monitoring Agent" deployed on each storage system (collection frequency: 5 seconds / time). For example, DWS CPU utilization is calculated as "Average CPU Utilization of Compute Nodes," and if it exceeds 80%, it is considered overloaded; MRS storage utilization is calculated as "Actual Storage Amount / Total Storage Amount," and if it exceeds 90%, it is considered insufficient storage; GaussDB connection count is calculated as "Current Active Connections / Maximum Connections," and if it exceeds 80%, it is considered connection overloaded. Monitoring data is pushed to the visualization dashboard in real time, and multi-level alarms (SMS + email) are triggered when anomalies occur. The visualization dashboard supports real-time refresh (refresh frequency 10 seconds / time), and maintenance personnel can view task progress and resource load curves. Anomaly alarms are triggered in a tiered manner—minor anomalies (such as synchronization delay of 5-10 minutes) only trigger email alarms, while severe anomalies (such as task failure, CPU utilization exceeding 90%) trigger both SMS and email alarms, ensuring timely response to anomalies.
[0062] The control rules for the three types of storage systems in this application embodiment are automated through "cloud service API + preset threshold". Specifically: "Automatically expand compute nodes when DWS CPU utilization exceeds 80% for 5 consecutive minutes": This is achieved by calling the "Elastic Expansion API" of Cloud DWS. The expansion quantity is calculated as "current CPU utilization - target utilization" (e.g., if the current utilization is 85% and the target utilization is 70%, then expand by 1 compute node). After expansion, CPU utilization is monitored in real time, and expansion stops when it drops below 70% to avoid excessive expansion and waste of resources.
[0063] "Automatically delete cold data in MRS that is archived for more than 7 years and has no access records when MRS storage utilization exceeds 90%": Before deletion, a "pre-deletion alert" is triggered (an email is sent to the operation and maintenance personnel 24 hours in advance), and the deletion is performed after confirmation that there are no objections; the deletion scope is cold data with "archived time > 7 years" and "access count in the last 365 days = 0", and a deletion log (including data ID and deletion time) is recorded after deletion to meet the compliance requirement of "archived for 7 years" for financial data (the normal archiving cycle in the financial industry).
[0064] "Automatically close idle connections when GaussDB connection count exceeds 80%": This is implemented through GaussDB's "Connection Management API". Idle connections are defined as connections that have not been used for SQL operations for more than 30 minutes. Before closing a connection, the system checks whether the connection is associated with any incomplete transactions (if so, it is skipped and closed after the transaction is completed) to avoid affecting business operations. After closing a connection, the system monitors the number of connections in real time and stops closing connections when the count drops below 70%, ensuring that GaussDB connection resources are allocated reasonably.
[0065] In some embodiments, the construction of the subject dataset specifically involves: extracting hot data from DWS, cold data from MRS, and data to be deleted or modified, as well as intermediate data to be confirmed, from GaussDB every morning; generating a fact table through data cleaning, summarization, and dimensional modeling; and then designing a financial sub-sub-theme wide table based on the fact table. The intelligent reporting mechanism is based on IPA and includes: configuring report rules, automatically connecting to the topic dataset to execute query statements, performing aggregation calculations on the query results, generating Excel or PDF reports and pushing them to preset recipients, and automatically regenerating reports and marking the update time when the topic dataset is updated.
[0066] Here, hot data is extracted from DWS, cold data from MRS, and data to be deleted / modified and intermediate data to be confirmed from GaussDB every day at midnight: the extraction time is selected at midnight every day (2:00-4:00) to avoid affecting business operations; the extraction method adopts "incremental extraction + full verification" - hot data (DWS) and intermediate data (GaussDB) are extracted incrementally (only the newly added / changed data of the day is extracted), and cold data (MRS) is verified in full (to ensure that no historical data is lost); the extracted data is temporarily stored in the "temporary table" of ClickHouse, and written to the formal table after cleaning is completed.
[0067] The cleaning rules are designed for the characteristics of financial data, for example: Unify data format: unify "posting time" to "YYYY-MM-DD HH:MM:SS" format, and unify "commission amount" to "numeric type with 2 decimal places"; Remove duplicate data: Based on the "unique data ID" (such as bill number + data type), duplicate data is marked as "invalid data" and logged; Complete missing data: For data with missing key fields such as "rate value" and "business department", the system automatically links to the intermediate data table of GaussDB to complete the missing data. If the completion fails, the data is marked as "pending manual processing" and an alarm is triggered to ensure data quality.
[0068] Based on business themes, declaration granularity, confirmation dimensions, and confirmation facts, dimensional modeling is performed to generate fact tables. Business themes include "monthly commission settlement," "annual expense audit," and "departmental financial statements," etc. The summary design for each theme is as follows: Partitioning: Partition by "time" (e.g., monthly partitions, 202401, 202402) for easy querying by time range; Granularity: Ensure the granularity of aggregated data by “smallest business unit” (such as “each posting data” or “monthly data for each business department”); Dimensions include "Time Dimension" (posting month, effective month), "Business Dimension" (business department, product type), and "Data Dimension" (rate value type, data status). Facts: Includes "measures" (commission amount, fee amount) and "counts" (number of postings, amount of valid data); Dimensional modeling: A "star schema" is used (fact table is associated with multiple dimension tables). For example, the "monthly commission settlement fact table" is associated with the "time dimension table", "business department dimension table" and "rate value dimension table", which facilitates multi-dimensional analysis of reports.
[0069] Sub-topic wide tables are segmented by topic (e.g., "Monthly Commission Settlement Wide Table" and "Monthly Expense Posting Wide Table"). High-frequency query fields (e.g., business department, commission amount, posting time) are placed as primary fields at the front of the table, while low-frequency query fields (e.g., remarks, operator) are placed as secondary fields at the back. Hot data fields (current month data) and cold data fields (historical data) are stored separately. For example, a "data hot / cold flag" field is added to the wide table, allowing filtering by flag during queries and reducing the data scanning scope. Utilizing ClickHouse's columnar storage feature, the wide tables are partitioned by business department and sorted by posting time. Queries only scan the target partition and target column, improving query efficiency by 5-10 times.
[0070] The intelligent reporting mechanism in this application embodiment is based on IPA technology to achieve a closed loop of "configuration → query → calculation → push → update", and the specific process is as follows: Step 1: Report Configuration. Finance personnel configure report rules on the IPA platform. The IPA platform provides a visual configuration interface, requiring no code development. Configuration content includes: Basic report information: Report name (e.g., "Commission Settlement Report for May 2024"), generation frequency (e.g., the 1st of each month), format (Excel / PDF); Data source: Related target wide tables in ClickHouse (such as "Monthly Commission Settlement Wide Table"); Field selection: Select the fields to be displayed (such as business department, commission amount, rate value, posting time); Filtering criteria: Set data filtering rules (e.g., "posting month = current month - 1"); Recipient Configuration: Enter the recipient's email address (multiple email addresses are supported, grouped by department).
[0071] Step 2: Automatic Query. IPA automatically connects to ClickHouse according to the configuration rules and executes the query statement. The IPA platform has a built-in "SQL Auto-Generation Engine" that automatically generates ClickHouse query statements based on the configuration rules. For example, after configuring "Monthly Commission Report", the engine automatically generates: "SELECT Business Department, SUM (Commission Amount), MAX (Rate Value), MIN (Posting Time) FROM Monthly Commission Settlement Wide Table WHERE Posting Month = DATE_FORMAT (CURRENT_DATE-INTERVAL1 MONTH, '% Y% m') GROUP BY Business Department". The query uses "Pre-query Validation" - first execute the "count(*)" of the query statement to confirm that the data volume is normal before executing the complete query to avoid empty reports.
[0072] Step 3: Automatic Calculation: Perform aggregation calculations on the query results. The aggregation calculation rules are configured according to the report requirements, for example: Numeric fields: Perform calculations such as SUM (total commission amount), AVG (average rate value), MAX (maximum single commission), and MIN (minimum single commission); Category fields: Perform calculations such as COUNT (number of postings for each business department) and DISTINCTCOUNT (quantity for each product type); Automatic verification of calculation results: The calculation results are compared with the original data in GaussDB (e.g., the deviation between the SUM commission amount and the total commission in GaussDB for the current month is ≤0.1%). If the verification passes, proceed to the next step; if the verification fails, an alarm is triggered.
[0073] Step 4: Automatic push. Generate Excel / PDF reports and automatically push them to the configured recipients' email addresses. When generating reports, a "template-based design" is used—finance personnel can upload custom Excel templates (including headers, formats, and formulas), and the IPA platform fills in the data according to the template. A "push log" (including push time, recipients, and report size) is recorded during the push. If the push fails (e.g., due to an incorrect email address), it will automatically retry 3 times. If the retry fails, an SMS alarm will be triggered.
[0074] Step 5: Automatic Updates. When the ClickHouse dataset is updated, IPA automatically regenerates the report and marks the update time. The ClickHouse dataset update trigger condition is "data extraction and cleaning completed every morning at midnight". After the update, the IPA platform automatically identifies the related reports, re-executes the "query → calculation → generation → push" process, and marks the update time after the report title (e.g., "Commission Settlement Report for May 2024 (updated on 2024-06-02 04:30)"). At the same time, a "Report Update Notification" email is sent to the recipient to avoid the recipient using the old report and to ensure the real-time and accuracy of the report data.
[0075] This automated process reduces the time for generating financial statements from "several hours" to "10-30 minutes" without human intervention, fully meeting the real-time query needs of scenarios such as auditing and financial settlement.
[0076] In summary, the embodiments of this application have the following beneficial effects: By constructing a multi-dimensional cold / hot data assessment model specifically for finance (integrating posting time, rate value effective date, and access popularity, and dynamically adapting to settlement / audit scenarios), an MRS-DWS-GaussDB three-level storage architecture with delete / modify feature adaptation (achieving precise allocation and bidirectional synchronization of undeletable cold / hot data and delegateable intermediate data), an intelligent scheduling hub including data flow scheduling / full-process monitoring / dynamic resource control / data traceability (supporting timed and triggered flow, real-time alarms, and automated resource optimization), and a ClickHouse topic dataset + IPA intelligent reporting mechanism (integrating... By integrating multiple data sources, the system enables automatic querying, calculation, generation, and push of reports, forming a closed loop for the full lifecycle management of financial data. This not only improves the efficiency of determining hot and cold data by over 90% and reduces report generation time from several hours to minutes, effectively meeting the real-time needs of auditing, settlement, and other scenarios, but also reduces storage resource waste by over 30% through precise storage allocation. At the same time, it ensures the traceability of financial data by relying on unique transfer IDs and data trajectory records, comprehensively improving the efficiency, resource utilization, data quality, and compliance security of financial data management, and providing key technical support for efficient data management in the financial and corporate finance fields.
[0077] Based on the same inventive concept, this application also provides a data cold and hot tiered storage and data intelligent scheduling device corresponding to the data cold and hot tiered storage and data intelligent scheduling method in the first embodiment. Since the principle of the device in this application is similar to the above-mentioned data cold and hot tiered storage and data intelligent scheduling method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0078] like Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of the data cold and hot tiered storage and intelligent data scheduling device 600 provided in this application embodiment. The data cold and hot tiered storage and intelligent data scheduling device 600 includes: The judgment module 601 is used to construct a multi-dimensional hot and cold evaluation model. It collects at least two types of evaluation dimension data related to business characteristics from the data source layer, performs quantitative processing on the data of each evaluation dimension, dynamically allocates the weight of each evaluation dimension according to the target business scenario, calculates the comprehensive score of the data based on the quantification results and weights, and determines the hot and cold attributes of the data according to the preset threshold. The allocation module 602 is used to allocate data to a storage system with corresponding characteristics based on the determined hot / cold attributes and the data deletion / modification requirements. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and a deletion / modification-supporting intermediate data storage system. The scheduling module 603 is used to build an intelligent scheduling center. The intelligent scheduling center includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automatic flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts the storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. The reporting module 604 is used to create intelligent reports, integrate data from various storage systems to construct thematic datasets, and automatically complete the querying, calculation, generation, push and updating of reports based on preset reporting rules.
[0079] Those skilled in the art should understand that Figure 6 The functions of each unit in the data cold and hot tiered storage and intelligent data scheduling device 600 shown can be understood by referring to the relevant description of the aforementioned data cold and hot tiered storage and intelligent data scheduling method. Figure 6 The functions of each unit in the data cold and hot tiered storage and intelligent data scheduling device 600 shown can be implemented by a program running on a processor or by specific logic circuits.
[0080] In one possible implementation, the evaluation dimensions include financial posting time, rate value effective date, and data access frequency; the quantification process specifically includes: Financial posting time: 100 points for posting within 3 months, 60 points for posting within 6 months, 30 points for posting within 12 months, and 10 points for posting within 12 months. Effective date of rate value: 100 points are assigned when the rate is active, 70 points are assigned when the rate is active and the remaining talent value is ≤30, 40 points are assigned when the rate has expired for ≤6 months, and 10 points are assigned when the rate has expired for >6 months. Data access popularity: ≥10 visits in the past 30 days are assigned 100 points, 5 ≤ visits <10 visits are assigned 70 points, 1 ≤ visits <5 visits are assigned 30 points, and <1 visit is assigned 10 points.
[0081] In one possible implementation, the dynamic allocation of weights specifically involves: In settlement scenarios, the weighting is as follows: financial posting time 40%, rate value effective date 35%, and data access popularity 25%. In an auditing scenario, the weightings are: financial posting time (25%), rate effective date (25%), and data access frequency (50%). The preset threshold determination rule is as follows: a comprehensive score ≥ 60 is considered hot data, a score < 50 is considered cold data, and a score ≤ 50 < 60 is considered transitional data.
[0082] In one possible implementation, the storage system characteristics and data allocation rules are as follows: The immutable cold data storage system is called MRS, which is used to store cold data and final archived data, and has the capabilities of low-cost massive storage and offline analysis. The non-deletable hot data storage system is DWS, which is used to store hot data and transition data. It has near real-time computing and ad-hoc query capabilities, with a response latency of ≤1 second. The intermediate data storage system that supports deletion and modification is GaussDB, which is used to store business data that needs to be dynamically adjusted and has row-level deletion and modification capabilities as well as transaction consistency capabilities. The data allocation is achieved through a two-step determination: the first step is to identify whether there is a need to delete or modify the data, and if so, it is preferentially allocated to GaussDB; the second step is to determine whether the data belongs to MRS or DWS based on its hot or cold attribute.
[0083] In one possible implementation, a data synchronization step is also included: Once data deletion or modification in GaussDB is confirmed, it is automatically synchronized to the corresponding hot / cold attribute DWS or MRS. After synchronization, GaussDB retains a copy of the data and marks it as archived. When hot data in DWS is reassessed as cold data the following month, it is automatically migrated to MRS. After migration, DWS retains the index of the MRS storage path.
[0084] In one possible implementation, the data flow scheduling module includes two flow methods: timed flow and triggered flow. The timed flow performs cold data migration from DWS to MRS every day at midnight. The triggered flow immediately triggers the task of synchronizing to DWS or MRS after the GaussDB data deletion and modification are confirmed. The flow task supports breakpoint resumption and retry mechanisms. The full-process monitoring module collects the flow task status, data flow volume, synchronization delay time, and CPU utilization, storage utilization, and connection count of each storage system in real time through a data flow log collector and storage resource monitoring agent. It triggers SMS and email alarms when an anomaly occurs. The resource dynamic control module executes the following based on monitoring data: automatically expands computing nodes when DWS CPU utilization exceeds 80% for 5 consecutive minutes; automatically deletes archived data older than 7 years with no access records when MRS storage utilization exceeds 90%; and automatically closes idle connections when GaussDB connection count exceeds 80%.
[0085] In one possible implementation, the construction of the topic dataset specifically involves: extracting hot data from DWS, cold data from MRS, and data to be deleted or modified, as well as intermediate data to be confirmed, from GaussDB every morning; generating a fact table through data cleaning, summarization, and dimensional modeling; and then designing a financial sub-topic wide table based on the fact table. The intelligent reporting mechanism is based on IPA and includes: configuring report rules, automatically connecting to the topic dataset to execute query statements, performing aggregation calculations on the query results, generating Excel or PDF reports and pushing them to preset recipients, and automatically regenerating reports and marking the update time when the topic dataset is updated.
[0086] The aforementioned data cold and hot tiered storage and intelligent data scheduling device constructs a multi-dimensional cold and hot assessment model specifically for finance (integrating posting time, rate value effective date, and access popularity, and dynamically adapting to settlement / audit scenarios), an MRS-DWS-GaussDB three-level storage architecture with deletion and modification feature adaptation (achieving accurate allocation and bidirectional synchronization of undeletable cold / hot data and delegateable intermediate data), an intelligent scheduling hub including data flow scheduling / full-process monitoring / dynamic resource control / data traceability (supporting timed and triggered flow, real-time alarms, and automated resource optimization), and ClickHouse topic datasets + IP. An intelligent reporting mechanism (integrating multi-source data to automatically query, calculate, generate, and push reports) forms a closed loop for the full lifecycle management of financial data. It not only improves the efficiency of hot / cold data determination by more than 90% and reduces report generation time from several hours to minutes, effectively meeting the real-time needs of auditing, settlement, and other scenarios, but also reduces storage resource waste by more than 30% through precise storage allocation. At the same time, it ensures the traceability of financial data by relying on unique flow IDs and data trajectory records, comprehensively improving the efficiency, resource utilization, data quality, and compliance security of financial data management, and providing key technical support for efficient data management in the financial and corporate finance fields.
[0087] like Figure 7 As shown, Figure 7 This is a schematic diagram of the composition structure of the electronic device 700 provided in the embodiments of this application. The electronic device 700 includes: The device 700 includes a processor 701, a storage medium 702, and a bus 703. The storage medium 702 stores machine-readable instructions that can be executed by the processor 701. When the electronic device 700 is running, the processor 701 communicates with the storage medium 702 via the bus 703. The processor 701 executes the machine-readable instructions to perform the steps of the data cold and hot tiered storage and intelligent data scheduling method described in the embodiments of this application.
[0088] In practical applications, the various components in the electronic device 700 are coupled together via a bus 703. It is understood that the bus 703 is used to achieve communication between these components. In addition to a data bus, the bus 703 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 7 The general designated all buses as Bus 703.
[0089] The aforementioned electronic devices utilize a multi-dimensional cold / hot data assessment model specifically designed for finance (integrating posting time, rate value effective date, and access popularity, and dynamically adapting to settlement / audit scenarios), a three-tier storage architecture of MRS-DWS-GaussDB with duplicate / modifiable features (achieving precise allocation and bidirectional synchronization of undeletable cold / hot data and duplicate intermediate data), an intelligent scheduling hub including data flow scheduling, full-process monitoring, dynamic resource control, and data traceability (supporting scheduled and triggered flow, real-time alarms, and automated resource optimization), and a ClickHouse theme dataset + IPA intelligent reporting machine. This system (integrating multi-source data to achieve automatic report querying, calculation, generation, and push) forms a closed loop for the full lifecycle management of financial data. It not only improves the efficiency of hot and cold data determination by more than 90% and reduces report generation time from several hours to minutes, effectively meeting the real-time needs of auditing, settlement, and other scenarios, but also reduces storage resource waste by more than 30% through precise storage allocation. At the same time, it ensures the traceability of financial data by relying on unique flow IDs and data trajectory records, comprehensively improving the efficiency, resource utilization, data quality, and compliance security of financial data management, and providing key technical support for efficient data management in the financial and corporate finance fields.
[0090] This application also provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by at least one processor 701, the data cold and hot tiered storage and intelligent data scheduling method described in this application is implemented.
[0091] In some embodiments, the storage medium may be a magnetic random access memory (FRAM), a read-only memory (ROM), or a programmable read-only memory (PROM). Erasable Programmable Read-Only Memory (EPROM) Electrically Erasable Programmable Read-Only Memory (EEPROM) Read-only memory, flash memory, magnetic surface storage, optical disc, or CD-ROM ROM, Compact Disc Read It can be a memory such as a memory only; or it can be a device that includes one or any combination of the above-mentioned memories.
[0092] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0093] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0094] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0095] The aforementioned computer-readable storage media utilizes a multi-dimensional cold / hot data assessment model specific to finance (integrating posting time, rate value effective date, and access popularity, and dynamically adapting to settlement / audit scenarios), an MRS-DWS-GaussDB three-level storage architecture with deletion / modification feature adaptation (achieving precise allocation and bidirectional synchronization of undeletable cold / hot data and delegateable intermediate data), an intelligent scheduling hub including data flow scheduling / full-process monitoring / dynamic resource control / data traceability (supporting timed and triggered flow, real-time alarms, and automated resource optimization), and ClickHouse topic datasets + IPA intelligent reporting. The report mechanism (integrating multi-source data to achieve automatic report querying, calculation, generation, and push) forms a closed loop for the full lifecycle management of financial data. It not only improves the efficiency of hot and cold data determination by more than 90% and reduces the report generation time from several hours to minutes, effectively meeting the real-time needs of auditing, settlement, and other scenarios, but also reduces storage resource waste by more than 30% through precise storage allocation. At the same time, it ensures the traceability of financial data by relying on unique flow IDs and data trajectory records, comprehensively improving the efficiency, resource utilization, data quality, and compliance security of financial data management, and providing key technical support for efficient data management in the financial and corporate finance fields.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0097] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0099] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0100] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for tiered cold and hot data storage and intelligent data scheduling, characterized in that, The method includes: Construct a multi-dimensional hot and cold assessment model, collect at least two types of assessment dimension data related to business characteristics from the data source layer, quantify the data of each assessment dimension, dynamically allocate the weight of each assessment dimension according to the target business scenario, calculate the comprehensive score of the data based on the quantification results and weights, and determine the hot and cold attributes of the data according to the preset threshold. Based on the determined hot / cold attributes and the data deletion / modification requirements, the data is allocated to storage systems with corresponding characteristics. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and an intermediate data storage system that supports deletion / modification. An intelligent scheduling hub is established, which includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automated flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. Create intelligent reports, integrate data from various storage systems to construct thematic datasets, and automatically complete report queries, calculations, generation, push, and updates based on preset report rules.
2. The method according to claim 1, characterized in that, The evaluation dimensions include financial posting time, rate value effective date, and data access frequency; the quantitative processing specifically involves: Financial posting time: 100 points for posting within 3 months, 60 points for posting within 6 months, 30 points for posting within 12 months, and 10 points for posting within 12 months. Effective date of rate value: 100 points are assigned when the rate is active, 70 points are assigned when the rate is active and the remaining talent value is ≤30, 40 points are assigned when the rate has expired for ≤6 months, and 10 points are assigned when the rate has expired for >6 months. Data access popularity: ≥10 visits in the past 30 days are assigned 100 points, 5 ≤ visits <10 visits are assigned 70 points, 1 ≤ visits <5 visits are assigned 30 points, and <1 visit is assigned 10 points.
3. The method according to claim 2, characterized in that, The dynamic weight allocation specifically refers to: In settlement scenarios, the weighting is as follows: financial posting time 40%, rate value effective date 35%, and data access popularity 25%. In an auditing scenario, the weightings are: financial posting time (25%), rate effective date (25%), and data access frequency (50%). The preset threshold determination rule is as follows: a comprehensive score ≥ 60 is considered hot data, a score < 50 is considered cold data, and a score ≤ 50 < 60 is considered transitional data.
4. The method according to claim 1, characterized in that, The storage system characteristics and data allocation rules are as follows: The immutable cold data storage system is called MRS, which is used to store cold data and final archived data, and has the capabilities of low-cost massive storage and offline analysis. The non-deletable hot data storage system is DWS, which is used to store hot data and transition data. It has near real-time computing and ad-hoc query capabilities, with a response latency of ≤1 second. The intermediate data storage system that supports deletion and modification is GaussDB, which is used to store business data that needs to be dynamically adjusted and has row-level deletion and modification capabilities as well as transaction consistency capabilities. The data allocation is achieved through a two-step determination: the first step is to identify whether there is a need to delete or modify the data, and if so, it is preferentially allocated to GaussDB; the second step is to determine whether the data belongs to MRS or DWS based on its hot or cold attribute.
5. The method according to claim 4, characterized in that, It also includes the data synchronization step: Once data deletion or modification in GaussDB is confirmed, it is automatically synchronized to the corresponding hot / cold attribute DWS or MRS. After synchronization, GaussDB retains a copy of the data and marks it as archived. When hot data in DWS is reassessed as cold data the following month, it is automatically migrated to MRS. After migration, DWS retains the index of the MRS storage path.
6. The method according to claim 1, characterized in that, The data flow scheduling module includes two flow methods: timed flow and triggered flow. Timed flow performs cold data migration from DWS to MRS every day at midnight. Triggered flow immediately triggers the task of synchronizing to DWS or MRS after the data in GaussDB is deleted and modified. The flow task supports breakpoint resumption and retry mechanisms. The full-process monitoring module collects the flow task status, data flow volume, synchronization delay time, and CPU utilization, storage utilization, and connection count of each storage system in real time through a data flow log collector and storage resource monitoring agent. It triggers SMS and email alarms when an anomaly occurs. The resource dynamic control module executes the following based on monitoring data: automatically expands computing nodes when DWS CPU utilization exceeds 80% for 5 consecutive minutes; automatically deletes archived data older than 7 years with no access records when MRS storage utilization exceeds 90%; and automatically closes idle connections when GaussDB connection count exceeds 80%.
7. The method according to claim 1, characterized in that, The construction of the subject dataset specifically involves: extracting hot data from DWS, cold data from MRS, and data to be deleted or modified, as well as intermediate data to be confirmed, from GaussDB every morning; generating fact tables through data cleaning, summarization, and dimensional modeling; and then designing financial sub-sub-theme wide tables based on the fact tables. The intelligent reporting mechanism is based on IPA and includes: configuring report rules, automatically connecting to the topic dataset to execute query statements, performing aggregation calculations on the query results, generating Excel or PDF reports and pushing them to preset recipients, and automatically regenerating reports and marking the update time when the topic dataset is updated.
8. A data cold and hot tiered storage and intelligent data scheduling device, characterized in that, The device includes: The judgment module is used to build a multi-dimensional hot and cold evaluation model. It collects at least two types of evaluation dimension data related to business characteristics from the data source layer, quantifies the data of each evaluation dimension, dynamically allocates the weight of each evaluation dimension according to the target business scenario, calculates the comprehensive score of the data based on the quantification results and weights, and determines the hot and cold attributes of the data according to the preset threshold. The allocation module is used to allocate data to storage systems with corresponding characteristics based on the determined hot / cold attributes and the data deletion / modification requirements. The storage system includes at least a non-deletable cold data storage system, a non-deletable hot data storage system, and a deletion / modification-supporting intermediate data storage system. The scheduling module is used to build an intelligent scheduling hub. The intelligent scheduling hub includes a data flow scheduling module, a full-process monitoring module, a resource dynamic control module, and a data traceability module. The data flow scheduling module realizes the automated flow of data between various storage systems. The full-process monitoring module collects the data flow status and storage resource status in real time. The resource dynamic control module adjusts storage resources according to the monitoring results. The data traceability module records the entire life cycle trajectory of the data. The reporting module is used to create intelligent reports, integrate data from various storage systems to build thematic datasets, and automatically complete the querying, calculation, generation, push and updating of reports based on preset reporting rules.
9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the data cold and hot tiered storage and intelligent data scheduling method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the data cold and hot tiered storage and intelligent data scheduling method as described in any one of claims 1 to 7.
Citation Information
Cited By
Cold and hot data layering method and system based on heterogeneous storage medium
CN121957508A