Data storage method and device and electronic equipment

By dynamically adjusting the storage identifier of the data table, the problem of uneven resource utilization and poor system stability caused by the binding of storage and computing resources in the financial industry is solved, achieving more efficient resource utilization and system stability.

CN120743864APending Publication Date: 2025-10-03INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510854269.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing data storage in the financial industry, storage and computing resources are bound to the same node, resulting in uneven resource utilization, poor system stability, and difficulty in adapting to dynamic business needs.

Method used

By responding to the job logs obtained from the scan, the access frequency and access type of the data table are determined, the storage identifier of the data table is dynamically adjusted, and it is stored in a matching storage cluster, including hot storage, warm storage, and cold storage, to achieve decoupling of computing nodes and storage nodes.

Benefits of technology

It improves resource utilization, enhances system flexibility and stability, and reduces storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743864A_ABST
    Figure CN120743864A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data storage method and device and electronic equipment, and relates to the field of big data. The method comprises the steps of determining access data of any data table in a job log in response to the job log obtained through scanning, determining the access frequency and the access type of the data table according to the access data of the data table, and determining the access data of any data table according to the access frequency and the access type of the data table. The storage identifier of the data table is determined, the data table is stored in the storage cluster matched with the storage identifier, system resources are fully utilized, the resource utilization rate is increased, and the system stability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data, and in particular to a data storage method, device and electronic device. Background Art

[0002] With the rapid development of the financial industry, data volumes are rapidly increasing. This includes not only traditional transaction data but also data from various new channels, such as mobile banking, online payment platforms, and other forms of financial services enabled by fintech innovation. Faced with this growing demand for data storage, financial institutions must consider how to store data efficiently.

[0003] Currently, data storage in the financial industry mostly utilizes a storage-computing architecture. In this model, storage and computing resources are bundled together on the same node. Resource allocation is typically static, meaning that storage and computing resources are allocated at the initial stage of system design to ensure that business processing is executed according to pre-set procedures.

[0004] However, because storage and computing resources are tied to the same node, a single node carries both storage and computing responsibilities. A failure in that node not only impacts the execution of current tasks but can also render data inaccessible, impacting the stability of the entire system. Static resource allocation struggles to adapt to dynamically changing business needs and can easily lead to uneven resource utilization. Summary of the Invention

[0005] The present application provides a data storage method, device and electronic device to solve the problems caused by the static allocation method in the existing data storage process, especially the technical difficulties of low resource utilization and poor system stability when the storage and computing nodes are not decoupled.

[0006] In a first aspect, the present application provides a data storage method, comprising:

[0007] In response to the scanned job log, determining access data of any data table in the job log, wherein the job log is used to indicate operation records of multiple data tables;

[0008] Determining the access frequency and access type of the data table according to the data table access data;

[0009] Determining a storage identifier of the data table according to the access frequency and access type of the data table;

[0010] The data table is stored in a storage cluster that matches the storage identifier.

[0011] Optionally, determining the access frequency of the data table according to accessing data in the data table includes:

[0012] Determining whether the access frequency is greater than a first access threshold;

[0013] If so, determining that the access frequency of the data table is high frequency;

[0014] If not, determining whether the access frequency is less than a second access threshold;

[0015] When the access frequency is less than a second access threshold, determining that the access frequency of the data table is low frequency;

[0016] When the access frequency is not less than a second access threshold, it is determined that the access frequency of the data table is a medium frequency.

[0017] Optionally, accessing data in the data table and determining the access type of the data table includes:

[0018] Determining a read access count and a write access count of the data table according to the data table access data;

[0019] Determining whether the read access ratio is greater than a first ratio threshold, where the read access ratio is a ratio of the number of read accesses to the total number of accesses;

[0020] If so, determining that the access type of the data table is read-primarily;

[0021] If not, determining whether the write access ratio is greater than a second ratio threshold, where the write access ratio is a ratio of the number of write accesses to the total number of accesses;

[0022] When the write access ratio is greater than a second ratio threshold, determining that the access type of the data table is mainly write;

[0023] When the write access ratio is not greater than a second ratio threshold, it is determined that the access type of the data table is mixed access.

[0024] Optionally, determining the storage identifier of the data table according to the access frequency and access type of the data table includes:

[0025] If the access frequency of the data table is high, determining that the storage identifier of the data table is a hot storage cluster;

[0026] If the access frequency of the data table is medium frequency and the access type is not mixed access, determining that the storage identifier of the data table is a hot storage cluster;

[0027] If the access frequency of the data table is medium and the access type is mixed access, determining that the storage identifier of the data table is a warm storage cluster;

[0028] If the access frequency of the data table is low, it is determined that the storage identifier of the data table is a cold storage cluster.

[0029] Optionally, the method further includes:

[0030] Regularly obtain access frequency change data and data volume change data of any data table in the database, wherein the access frequency change data is the change in access frequency of the data table within a certain period, and the data volume change data is the change in data volume of the data table within a certain period;

[0031] Determine, based on the access frequency change data and the data volume change data, a data table in the cold storage cluster whose access frequency is higher than a first threshold;

[0032] Determine, based on the access frequency change data and the data volume change data, a data table in the hot storage cluster whose access frequency is lower than a second threshold;

[0033] An operation and maintenance list is generated based on the data tables with access frequencies higher than a first threshold and the data tables with access frequencies lower than a second threshold. The operation and maintenance list is used to migrate corresponding data tables in the cold storage cluster and the hot storage cluster.

[0034] Optionally, after generating the operation and maintenance list based on the data table with an access frequency higher than a first threshold and the data table with an access frequency lower than a second threshold, the method further includes:

[0035] When the data volume of the scanned job log is less than the first data volume, based on the operation and maintenance list, data tables in the cold storage cluster with an access frequency higher than a first threshold are migrated to the hot storage cluster, and data tables in the hot storage cluster with an access frequency lower than a second threshold are migrated to the cold storage cluster.

[0036] In a second aspect, the present application provides a data storage device, comprising:

[0037] a determination module, configured to determine access data of any data table in a job log obtained by scanning, wherein the job log is used to indicate operation records of multiple data tables;

[0038] The determining module is further configured to determine the access frequency and access type of the data table according to access data in the data table;

[0039] The determining module is further configured to determine a storage identifier of the data table according to the access frequency and access type of the data table;

[0040] The processing module is further configured to store the data table in a storage cluster that matches the storage identifier.

[0041] Optionally, the device further includes: a judgment module;

[0042] The judging module is configured to judge whether the access frequency is greater than a first access threshold;

[0043] The determining module is further configured to determine that the access frequency of the data table is high frequency when the access frequency is greater than a first access threshold;

[0044] The judgment module is further configured to judge whether the access frequency is less than a second access threshold when the access frequency is not greater than the first access threshold;

[0045] The determining module is further configured to determine that the access frequency of the data table is low frequency when the access frequency is less than a second access threshold;

[0046] The determining module is further configured to determine that the access frequency of the data table is a medium frequency when the access frequency is not less than a second access threshold.

[0047] Optionally, the determining module is further configured to determine the number of read accesses and the number of write accesses to the data table based on accessing data in the data table;

[0048] The judgment module is further configured to judge whether the read access ratio is greater than a first ratio threshold, wherein the read access ratio is a ratio of the number of read accesses to the total number of accesses;

[0049] The determining module is further configured to determine that the access type of the data table is mainly read when the read access ratio is greater than a first ratio threshold;

[0050] The judgment module is further configured to judge whether the write access ratio is greater than a second ratio threshold when the read access ratio is not greater than the first ratio threshold, wherein the write access ratio is a ratio of the number of write accesses to the total number of accesses;

[0051] The confirmation module is further configured to determine that the access type of the data table is mainly write-based when the write access ratio is greater than a second ratio threshold;

[0052] The confirmation module is further configured to determine that the access type of the data table is mixed access when the write access ratio is not greater than a second ratio threshold.

[0053] Optionally, the determining module is further configured to determine that the storage identifier of the data table is a hot storage cluster when the access frequency of the data table is high frequency;

[0054] The determining module is further configured to determine, when the access frequency of the data table is medium frequency and the access type is non-mixed access, that the storage identifier of the data table is a hot storage cluster;

[0055] The determining module is further configured to determine that the storage identifier of the data table is a warm storage cluster when the access frequency of the data table is medium frequency and the access type is mixed access;

[0056] The determining module is further configured to determine that the storage identifier of the data table is a cold storage cluster when the access frequency of the data table is low.

[0057] Optionally, the device further includes: an acquisition module and a generation module;

[0058] The acquisition module is used to regularly acquire access frequency change data and data volume change data of any data table in the database, wherein the access frequency change data is the change in access frequency of the data table within a certain period, and the data volume change data is the change in data volume of the data table within a certain period;

[0059] The determining module is further configured to determine, based on the access frequency change data and the data volume change data, a data table in the cold storage cluster whose access frequency is higher than a first threshold;

[0060] The determining module is further configured to determine, based on the access frequency change data and the data volume change data, a data table in the hot storage cluster whose access frequency is lower than a second threshold;

[0061] The generation module is used to generate an operation and maintenance list based on the data table with an access frequency higher than a first threshold and the data table with an access frequency lower than a second threshold, and the operation and maintenance list is used to migrate the corresponding data tables in the cold storage cluster and the hot storage cluster.

[0062] Optionally, the processing module is also used to migrate data tables in the cold storage cluster with an access frequency higher than a first threshold to the hot storage cluster, and migrate data tables in the hot storage cluster with an access frequency lower than a second threshold to the cold storage cluster based on the operation and maintenance list when the data volume of the scanned job log is less than the first data volume.

[0063] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0064] The memory stores computer-executable instructions;

[0065] The processor executes the computer-executable instructions stored in the memory to implement the data storage method as described in the first aspect and various possible implementations of the first aspect.

[0066] In a fourth aspect, the present application provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, are used to implement the data storage method as described in the first aspect and various possible implementations of the first aspect.

[0067] In a fifth aspect, the present application provides a program product, comprising a computer program, which implements the data storage method described above when executed by a processor.

[0068] The data storage method, device and electronic device provided in the present application determine the access data of any data table in the job log in response to the scanned job log, the job log is used to indicate the operation records of multiple data tables, and the access frequency and access type of the data table are determined based on the data table access data. The storage identifier of the data table is determined based on the access frequency and access type of the data table, and the data table is stored in a storage cluster that matches the storage identifier, so as to make full use of system resources, improve resource utilization and improve system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0070] Figure 1 A schematic diagram of a data storage method provided in this application Figure 1 ;

[0071] Figure 2 A schematic diagram of a data storage method provided in this application Figure 2 ;

[0072] Figure 3 A schematic diagram of a data storage method provided in this application Figure 3 ;

[0073] Figure 4 A schematic diagram of a data storage method provided in this application Figure 4 ;

[0074] Figure 5 A schematic structural diagram of a data storage device provided in this application;

[0075] Figure 6 A schematic structural diagram of a data storage device provided in this application.

[0076] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0077] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0078] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0079] In addition, this application involves conducting big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and using artificial intelligence technology to make automated decisions, and making technical solutions that have a significant impact on personal rights and interests based on the results of automated decisions. The application provides users with corresponding operation entrances for them to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered.

[0080] It should be noted that the data storage method, device and electronic device provided in this application can be used in the field of big data, and can also be used in any field other than big data. The application field of the data storage method, device and electronic device in this application is not limited.

[0081] First, let’s explain the terms that appear in this application:

[0082] Financial system: The financial system in this application refers to an integrated system for financial industries such as banking and securities that uses computer hardware and software and network equipment to collect and process financial information and continuously provide financial information services.

[0083] As the digital transformation of the financial industry accelerates, data sources and volumes are exploding. Financial institutions not only face storage demands for traditional transaction data, but also face the challenge of processing massive amounts of data from emerging channels and innovative services. An efficient data storage strategy is crucial for financial institutions, as it not only impacts data security and availability but also directly impacts business efficiency and competitiveness.

[0084] Currently, data storage in the financial industry mostly utilizes a storage-computing integrated architecture. In this model, storage and computing resources are bundled together on the same node. Resource allocation is typically static, meaning that storage and computing resources are allocated statically at the initial stage of system design. This simplifies the system architecture and ensures that business processing is executed according to pre-set procedures.

[0085] While integrating storage and computing functions on the same node reduces data transmission latency, the fact that storage and computing resources are tied to the same node makes it difficult to dynamically adjust resource allocation based on actual business needs. Changing business needs often require scaling or upgrading the entire node, resulting in underutilized resources. Furthermore, a single node failure can cause both storage and computing functions to fail, impacting the stability of the entire system.

[0086] In response to the above problems, the present application proposes a data storage method, which determines the access data of any data table in the job log in response to the scanned job log. The job log is used to indicate the operation records of multiple data tables. According to the data table access data, the access frequency and access type of the data table are determined. According to the access frequency and access type of the data table, the storage identifier of the data table is determined, and the data table is stored in a storage cluster that matches the storage identifier, so as to make full use of system resources, improve resource utilization and improve system stability.

[0087] This application is specifically applied to the efficient data storage process in financial systems, especially banking systems. By adopting strategies such as separated storage and computing architecture, dynamic resource management, hybrid cloud architecture and data lifecycle management, it optimizes data storage and management, improves resource utilization, enhances system flexibility and stability, and reduces storage costs.

[0088] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0089] Figure 1 A schematic diagram of a data storage method provided in an embodiment of the present application Figure 1 .like Figure 1As shown, the data storage method provided in this embodiment includes:

[0090] S101: In response to a job log obtained by scanning, determining access data of any data table in the job log.

[0091] Among them, the job log is used to indicate the operation records of multiple data tables, including but not limited to query, update, and deletion. The job log can be obtained by system scheduled scanning or by users performing random scanning. The job log contains various types of relevant data of multiple data tables. By parsing the job log, records related to data table operations can be extracted to obtain access data of any data table.

[0092] One possible implementation method is to analyze and extract the job logs obtained by scanning, obtain the access timestamp, operation type, operation content and other fields in the job log of any data table, and analyze and process the access timestamp, operation type, specific operation and other fields to obtain the access frequency, access type and access time data of each data table.

[0093] By parsing job logs, extracting access records for any data table, organizing key information, and generating reports, we can effectively determine the access data for any data table, making it easier to subsequently match storage clusters to each data table based on its access data.

[0094] S102: Determine the access frequency and access type of the data table according to the data table access data.

[0095] The data table access data records data such as the timestamp of the data table access, the operation type, and the operation content. Based on the data table access data, the timestamp information is extracted from the data table access data, and the number of times any data table is accessed within a certain period is calculated. Based on the number of accesses and a threshold, the access frequency of the data table is determined, e.g., high frequency, medium frequency, or low frequency. Based on the data table access data, the operation type and operation content are extracted from the data table access data, and the access type corresponding to each access to any data table is determined, e.g., data table write, data table read, or data table delete.

[0096] By analyzing table access frequency and access types, you can gain a comprehensive understanding of any table's usage. This facilitates subsequent adjustments to storage resource allocation based on table usage. For example, frequently accessed data or data highly relevant to business operations can be stored in a hot storage cluster, while infrequently accessed data or data with low business relevance can be stored in a cold storage cluster.

[0097] S103: Determine a storage identifier of the data table according to the access frequency and access type of the data table.

[0098] The storage identifier for each data table is determined by analyzing the access count and access type of each data table. Different storage identifiers correspond to different storage clusters, which have different performance and corresponding costs. The storage identifier indicates the storage cluster or storage medium in which the data table should be stored. Therefore, the most suitable storage cluster can be selected for the data table based on the access frequency and access type.

[0099] In one possible implementation, data tables with high access frequency should be stored in a high-performance storage cluster, corresponding to the hot storage cluster identifier, while data tables with low access frequency should be stored in a low-cost storage cluster, corresponding to the cold storage cluster identifier.

[0100] By analyzing the access frequency and access type of each data table, you can determine the most suitable storage cluster for the data table, thereby optimizing storage performance and reducing costs.

[0101] S104: Storing the data table in a storage cluster with a matching storage identifier.

[0102] A storage cluster refers to a group of storage devices with the same or similar storage characteristics, such as a high-performance SSD cluster, a medium-performance HDD cluster, or a low-cost cloud storage cluster.

[0103] The storage ID of a data table represents the storage requirement of the data table, that is, whether the data table needs to be stored in a hot storage cluster or a cold storage cluster. Based on the storage ID of the data table, the data table is stored in the storage cluster that is most suitable for it.

[0104] By storing data tables in a storage cluster that matches their storage requirements, you can optimize performance, reduce costs, and improve system scalability.

[0105] This embodiment provides a data storage method. The method determines the access data of any data table in the job log in response to a scanned job log. The job log is used to indicate the operation records of multiple data tables. The access frequency and access type of the data table are determined based on the data table access data. The storage identifier of the data table is determined based on the access frequency and access type of the data table. The data table is stored in a storage cluster that matches the storage identifier, thereby fully utilizing system resources, improving resource utilization, and improving system stability.

[0106] Figure 2 A schematic diagram of a data storage method provided in an embodiment of the present application Figure 2 .like Figure 2 As shown, in Figure 1Based on the embodiment, a possible implementation method for determining the access frequency and access type of a data table based on accessing data in the data table is described in detail, including:

[0107] S201: In response to the scanned job log, determine access data of any data table in the job log.

[0108] Among them, step S201 is similar to step S101 and will not be repeated here.

[0109] S202: Determine whether the access frequency is greater than a first access threshold. If so, execute step S203; if not, execute step S204.

[0110] S203: Determine that the access frequency of the data table is high frequency.

[0111] Access frequency refers to the number of times a data table is accessed within a certain period of time, such as a day or a week. The first access threshold is a preset value, set based on experience, and can be adjusted dynamically based on data table usage. The first access threshold is used to distinguish between "frequently accessed" and "infrequently accessed" data tables.

[0112] By setting an access frequency threshold, namely the first access threshold, it is determined whether the access frequency of the data table reaches the high frequency standard. If the access frequency of the data table exceeds the first access threshold, the data table described in the task is accessed at a high frequency, that is, the access frequency of the data table is determined to be high frequency.

[0113] One possible implementation method is, for example, that the access frequency of a data table is 10 times per second, and the first access threshold is: 5 times per second. Then, a data table with an access frequency exceeding 5 times per second is considered to be accessed at a high frequency, that is, the access frequency of the data table is determined to be high frequency.

[0114] S204: Determine whether the access frequency is less than a second access threshold. If so, execute step S205; if not, execute step S206.

[0115] S205: Determine that the access frequency of the data table is low.

[0116] S206: Determine that the access frequency of the data table is medium frequency.

[0117] The second access threshold is a preset value, set based on experience, and can be dynamically adjusted based on the usage of the data table. The second access threshold is used to distinguish whether the data table is "low-frequency access" or "non-low-frequency access".

[0118] When the access frequency is not greater than the first access threshold, that is, the access frequency of the data table is not high frequency and the data table is "non-high frequency access", it is further determined whether the access frequency is less than the second access threshold, that is, by comparing the access frequency of the data table with the second access threshold, it is determined whether the data table is medium frequency access or low frequency access.

[0119] If the access frequency is not greater than the first access threshold, if the access frequency is less than the second access threshold, it is determined that the data table is accessed at a low frequency, that is, the access frequency of the data table is determined to be low. If the access frequency is greater than the second access threshold, it is determined that the access frequency of the data table is at an intermediate value, that is, the access frequency of the data table is determined to be medium.

[0120] In one possible implementation, the first access threshold is 5 times per second, and the second access threshold is 2 times per second. If the access frequency of a data table is no more than 5 times per second, and the access frequency of the data table is less than 2 times per second, the access frequency of the data table is determined to be low. If the access frequency of the data table is greater than 2 times per second, the access frequency of the data table is determined to be medium.

[0121] By determining the access frequency of a data table as high, medium, or low, you can better manage the storage and access strategies of the data table and ensure efficient operation of the system.

[0122] S207: Determine the number of read accesses and the number of write accesses to the data table according to the data table access data.

[0123] The number of read accesses refers to the total number of query operations on the data table within a certain period of time. The number of inbound accesses refers to the total number of insert, update, or delete operations on the data table within a certain period of time.

[0124] By analyzing the access records of the data tables, the number of read operations on each data table, such as query operations and write operations, such as the number of insert, update, and delete operations, are counted.

[0125] Data tables with high read frequencies may require more cache resources, while data tables with high write frequencies may require higher storage performance. Understanding the read and write access times of a data table can help you match the data table with the optimal storage cluster.

[0126] S208: Determine whether the read access ratio is greater than a first ratio threshold. If so, execute step S209; if not, execute step S210.

[0127] S209: Determine that the access type of the data table is mainly read.

[0128] The read access ratio is the ratio of the number of read accesses to the total number of accesses. The first ratio threshold is a preset ratio value used to determine whether the read operation is dominant.

[0129] The ratio of read accesses to the total accesses to the data table is calculated and compared with a preset threshold, namely a first ratio threshold. If the read access ratio is greater than the first ratio threshold, it indicates that read operations on the data table account for a larger proportion of the total operations, and the access type of the data table can be considered to be read-dominated.

[0130] Understanding the access types of data tables can help design a more reasonable database architecture, such as choosing an appropriate storage cluster or partitioning strategy.

[0131] S210: Determine whether the write access ratio is greater than a second ratio threshold. If so, execute step S211; if not, execute step S212.

[0132] S211: Determine that the access type of the data table is primarily write-based.

[0133] S212: Determine that the access type of the data table is mixed access.

[0134] The write access ratio is a ratio of the number of write accesses to the total number of accesses. The second ratio threshold is a preset ratio value used to determine whether the write operation is dominant.

[0135] The ratio of write accesses to the total number of accesses to the data table is calculated and compared with a preset threshold, namely a second ratio threshold. If the write access ratio is greater than the second ratio threshold, it means that write operations account for a larger proportion of the total operations on the data table, and the access type of the data table can be considered as write-dominant. Otherwise, the access type of the data table is considered to be mixed access, that is, read and write operations are relatively balanced.

[0136] This embodiment provides a data storage method that determines the access frequency and number of accesses to any data table in a scanned job log. Based on first and second access thresholds and the access frequency, the data table's access frequency is determined to be high, medium, or low frequency. Based on first and second percentage thresholds and the access frequency, the data table's access type is determined to be read-primarily, write-primarily, or mixed access. This facilitates subsequent matching of storage identifiers for data tables based on the access frequency and access type. This method determines the data table's access frequency and access type based on data accessed from the data table, facilitating subsequent determination of corresponding storage identifiers. This method allows the data table to be stored in an appropriate storage, decoupling computing nodes and storage nodes, fully utilizing system resources, and improving resource utilization while also enhancing system stability.

[0137] Figure 3 A schematic diagram of a data storage method provided in an embodiment of the present application Figure 3 .like Figure 3 As shown, in Figure 1 Based on the embodiment, a possible implementation method for determining the storage identifier of the data table according to the access frequency and access type of the data table is described in detail, including:

[0138] S301: If the access frequency of the data table is high, determine that the storage identifier of the data table is a hot storage cluster.

[0139] S302: If the access frequency of the data table is medium frequency and the access type is not mixed access, determine that the storage identifier of the data table is a hot storage cluster.

[0140] S303: If the access frequency of the data table is medium frequency and the access type is mixed access, determine that the storage identifier of the data table is a warm storage cluster.

[0141] S304: If the access frequency of the data table is low, determine that the storage identifier of the data table is a cold storage cluster.

[0142] S305: Store the data table in a storage cluster that matches the storage identifier.

[0143] Among them, data tables with high frequency access and data tables with medium frequency access and non-mixed access type are stored in the hot storage cluster. The hot storage cluster is high-performance storage. Storing data tables with high frequency access and data tables with medium frequency access and non-mixed access type in the hot storage cluster can ensure fast response.

[0144] Data tables with medium-frequency and mixed access are stored in warm storage clusters. Warm storage clusters provide medium-performance storage. Storing data tables with medium-frequency and mixed access in warm storage clusters can balance performance and cost.

[0145] Infrequently accessed data tables are stored in cold storage clusters. Cold storage clusters are low-cost storage. Storing infrequently accessed data tables in cold storage clusters can save resources.

[0146] This embodiment provides a data storage method that matches storage identifiers to data tables with different access frequencies and access types, specifically high, medium, or low, and read-primarily, write-primarily, or read-write balanced, according to matching rules. By matching frequently accessed data tables to hot storage clusters and infrequently accessed data tables to cold storage clusters, system resources are fully utilized, improving resource utilization and system stability.

[0147] Figure 4 A schematic diagram of a data storage method provided in an embodiment of the present application Figure 4 .like Figure 4 As shown, in Figure 1 Based on the embodiment, the data storage method is described in detail, including:

[0148] S401: Regularly obtain access frequency change data and data volume change data of any data table in the database.

[0149] The access frequency change data is the change in the access frequency of the data table within a certain period, and the data volume change data is the change in the data volume within the data table within a certain period. Access frequency change data: The change in the access frequency of a data table within a specified period. For example, an increase from 5 accesses per day to 30 accesses per day. Data volume change data: The change in the data volume within a data table within a specified period. For example, an increase from 200 records to 500 records.

[0150] Regularly collect and analyze data on changes in table access frequency and data volume at regular intervals, such as daily, weekly, or monthly. If a table's access frequency increases significantly, you may need to adjust the storage cluster, migrating the table from a cold storage cluster to a hot storage cluster. If a table's data volume increases significantly, you may need to optimize storage performance or increase storage resources.

[0151] By monitoring the data on changes in access frequency and data volume, storage resources can be planned rationally to avoid unnecessary cost expenditures.

[0152] S402: Determine, based on the access frequency change data and the data volume change data, a data table in the cold storage cluster whose access frequency is higher than a first threshold.

[0153] Cold storage clusters are typically used to store data tables with low access frequency or small data volumes, such as backup data. The first threshold is a preset access frequency value used to determine whether a data table is a frequently accessed "hot data table" in the cold storage cluster. If the access frequency of a data table exceeds the first threshold, it indicates that the data table has high access frequency and large data volume fluctuations, making it no longer suitable for storage in the cold storage cluster and requiring migration to a higher-performance storage cluster. Data tables with high access frequency and large data volume fluctuations should be migrated to a higher-performance storage cluster to ensure rapid response.

[0154] S403: Determine, based on the access frequency change data and the data volume change data, a data table in the hot storage cluster whose access frequency is lower than a second threshold.

[0155] Hot storage clusters are typically used to store data tables with high access frequencies or large data volumes, such as real-time bank transaction data. The second threshold is a preset access frequency value used to determine whether a data table is a "cold data table" with relatively low access frequency in the hot storage cluster. If the access frequency of a data table does not exceed the second threshold, it indicates that the data table has low access frequency and small data volume fluctuations, making it no longer suitable for storage in the hot storage cluster and requiring migration to a lower-performance storage cluster. Migrating data tables with low access frequencies and small data volume fluctuations to a lower-performance storage cluster conserves storage resources.

[0156] S404: Generate an operation and maintenance list according to the data table whose access frequency is higher than the first threshold and the data table whose access frequency is lower than the second threshold.

[0157] S405: When the data volume of the scanned job log is less than the first data volume, based on the operation and maintenance list, migrate the data tables in the cold storage cluster whose access frequency is higher than the first threshold to the hot storage cluster, and migrate the data tables in the hot storage cluster whose access frequency is lower than the second threshold to the cold storage cluster.

[0158] By analyzing changes in data table access frequency, the system identifies data tables with access frequencies above a first threshold, indicating that they require higher-performance storage, and data tables with access frequencies below a second threshold, indicating that they can be migrated to lower-performance storage. Information about these tables is compiled into an operations and maintenance checklist, which is used to migrate the corresponding data tables in the cold and hot storage clusters, thereby rationally allocating storage resources and avoiding waste in high-performance storage clusters.

[0159] When the data volume of the scanned job log is less than the first data volume, that is, during the idle job period, based on the operation and maintenance list, the data tables with access frequency higher than the first threshold in the cold storage cluster are migrated to the hot storage cluster; the data tables with access frequency lower than the second threshold in the hot storage cluster are migrated to the cold storage cluster.

[0160] This embodiment provides a data storage method. The method periodically obtains the access frequency change and data volume change of any table in the database within a certain period, and based on the access frequency change and data volume change, determines data tables with high access frequency and / or large data volume change in a cold storage cluster and data tables with low access frequency and / or small data volume change in a hot storage cluster. The method then migrates the data tables with high access frequency and / or large data volume change in the cold storage cluster to the hot storage cluster, and migrates the data tables with low access frequency and / or small data volume change in the hot storage cluster to the cold storage cluster, thereby fully utilizing system resources, improving resource utilization, and enhancing system stability.

[0161] Figure 5This is a schematic diagram of the structure of a data storage device provided by this application. Figure 5 As shown, the present application provides a data storage device, the data storage device 500 comprising:

[0162] A determination module 501 is configured to determine access data of any data table in a job log obtained by scanning, wherein the job log is used to indicate operation records of multiple data tables;

[0163] The determining module 501 is further configured to determine the access frequency and access type of the data table according to the access data in the data table;

[0164] The determining module 501 is further configured to determine a storage identifier of the data table according to the access frequency and access type of the data table;

[0165] The processing module 502 is further configured to store the data table in a storage cluster that matches the storage identifier.

[0166] Optionally, the device further includes: a judgment module 503;

[0167] The judging module 503 is configured to judge whether the access frequency is greater than a first access threshold;

[0168] The determining module 501 is further configured to determine that the access frequency of the data table is high frequency when the access frequency is greater than a first access threshold;

[0169] The judging module 503 is further configured to, when the access frequency is not greater than the first access threshold, judge whether the access frequency is less than a second access threshold;

[0170] The determining module 501 is further configured to determine that the access frequency of the data table is low frequency when the access frequency is less than a second access threshold;

[0171] The determining module 501 is further configured to determine that the access frequency of the data table is a medium frequency when the access frequency is not less than a second access threshold.

[0172] Optionally, the determining module 501 is further configured to determine the number of read accesses and the number of write accesses to the data table according to access data in the data table;

[0173] The judgment module 503 is further configured to judge whether the read access ratio is greater than a first ratio threshold, where the read access ratio is a ratio of the number of read accesses to the total number of accesses;

[0174] The determining module 501 is further configured to determine that the access type of the data table is mainly read when the read access ratio is greater than a first ratio threshold;

[0175] The judgment module 503 is further configured to judge whether the write access ratio is greater than a second ratio threshold when the read access ratio is not greater than the first ratio threshold, wherein the write access ratio is a ratio of the number of write accesses to the total number of accesses;

[0176] The determining module 501 is further configured to determine that the access type of the data table is primarily write-based when the write access ratio is greater than a second ratio threshold;

[0177] The determining module 501 is further configured to determine that the access type of the data table is mixed access when the write access ratio is not greater than a second ratio threshold.

[0178] Optionally, the determining module 501 is further configured to determine that the storage identifier of the data table is a hot storage cluster when the access frequency of the data table is high frequency;

[0179] The determining module 501 is further configured to determine that the storage identifier of the data table is a hot storage cluster when the access frequency of the data table is medium frequency and the access type is non-mixed access;

[0180] The determining module 501 is further configured to determine that the storage identifier of the data table is a warm storage cluster when the access frequency of the data table is medium frequency and the access type is mixed access;

[0181] The determining module 501 is further configured to determine that the storage identifier of the data table is a cold storage cluster when the access frequency of the data table is low.

[0182] Optionally, the device further includes: an acquisition module 504, a generation module 505;

[0183] The acquisition module 504 is used to periodically acquire access frequency change data and data volume change data of any data table in the database, wherein the access frequency change data is the change in access frequency of the data table within a certain period, and the data volume change data is the change in data volume of the data table within a certain period;

[0184] The determining module 501 is further configured to determine, based on the access frequency change data and the data volume change data, a data table in the cold storage cluster whose access frequency is higher than a first threshold;

[0185] The determining module 501 is further configured to determine a data table in the hot storage cluster whose access frequency is lower than a second threshold value based on the access frequency change data and the data volume change data;

[0186] The generation module 505 is used to generate an operation and maintenance list based on the data table with access frequency higher than the first threshold and the data table with access frequency lower than the second threshold, and the operation and maintenance list is used to migrate the corresponding data tables in the cold storage cluster and the hot storage cluster.

[0187] Optionally, the processing module 502 is also used to migrate data tables in the cold storage cluster with an access frequency higher than a first threshold to the hot storage cluster, and to migrate data tables in the hot storage cluster with an access frequency lower than a second threshold to the cold storage cluster based on the operation and maintenance list when the data volume of the scanned job log is less than the first data volume.

[0188] The implementation principle and technical effects of the data storage device provided in the embodiment of the present application are similar to the implementation methods of each part of the aforementioned data storage method, and will not be repeated here.

[0189] Figure 6 This is a schematic diagram of the structure of a data storage device provided by this application. Figure 6 As shown, the present application provides a data storage device, which includes a data storage device 600 including a receiver 601 , a transmitter 602 , a processor 603 and a memory 604 .

[0190] Receiver 601, for receiving instructions and data;

[0191] Transmitter 602, used to send instructions and data;

[0192] Memory 604, for storing computer-executable instructions;

[0193] The processor 603 is configured to execute the computer-executable instructions stored in the memory 604 to implement the various steps of the data storage method in the above embodiment. For details, please refer to the relevant description of the above embodiment of the data storage method.

[0194] Optionally, the memory 604 may be independent or integrated with the processor 603 .

[0195] When the memory 604 is independently provided, the electronic device further includes a bus for connecting the memory 604 and the processor 603 .

[0196] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.

[0197] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.

[0198] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.

[0199] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0200] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0201] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0202] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0203] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0204] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0205] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data storage method, characterized in that: The method comprises: In response to the scanned job log, determining access data of any data table in the job log, wherein the job log is used to indicate operation records of multiple data tables; Determining the access frequency and access type of the data table according to the data table access data; Determining a storage identifier of the data table according to the access frequency and access type of the data table; The data table is stored in a storage cluster that matches the storage identifier.

2. The method according to claim 1, characterized in that The accessing data in the data table to determine the access frequency of the data table includes: Determining whether the access frequency is greater than a first access threshold; If so, determining that the access frequency of the data table is high frequency; If not, determining whether the access frequency is less than a second access threshold; When the access frequency is less than a second access threshold, determining that the access frequency of the data table is low frequency; When the access frequency is not less than a second access threshold, it is determined that the access frequency of the data table is a medium frequency.

3. The method according to claim 1, characterized in that The accessing data from the data table to determine the access type of the data table includes: Determining a read access count and a write access count of the data table according to the data table access data; Determining whether the read access ratio is greater than a first ratio threshold, where the read access ratio is a ratio of the number of read accesses to the total number of accesses; If so, determining that the access type of the data table is read-primarily; If not, determining whether the write access ratio is greater than a second ratio threshold, where the write access ratio is a ratio of the number of write accesses to the total number of accesses; When the write access ratio is greater than a second ratio threshold, determining that the access type of the data table is mainly write; When the write access ratio is not greater than a second ratio threshold, it is determined that the access type of the data table is mixed access.

4. The method according to claims 2 and 3, characterized in that Determining a storage identifier of the data table according to the access frequency and access type of the data table includes: If the access frequency of the data table is high, determining that the storage identifier of the data table is a hot storage cluster; If the access frequency of the data table is medium frequency and the access type is not mixed access, determining that the storage identifier of the data table is a hot storage cluster; If the access frequency of the data table is medium and the access type is mixed access, determining that the storage identifier of the data table is a warm storage cluster; If the access frequency of the data table is low, it is determined that the storage identifier of the data table is a cold storage cluster.

5. The method according to claim 1, wherein The method further comprises: Regularly obtain access frequency change data and data volume change data of any data table in the database, wherein the access frequency change data is the change in access frequency of the data table within a certain period, and the data volume change data is the change in data volume of the data table within a certain period; Determine, based on the access frequency change data and the data volume change data, a data table in the cold storage cluster whose access frequency is higher than a first threshold; Determine, based on the access frequency change data and the data volume change data, a data table in the hot storage cluster whose access frequency is lower than a second threshold; An operation and maintenance list is generated based on the data tables with access frequencies higher than a first threshold and the data tables with access frequencies lower than a second threshold. The operation and maintenance list is used to migrate corresponding data tables in the cold storage cluster and the hot storage cluster.

6. The method according to claim 5, characterized in that After generating the operation and maintenance list based on the data table with an access frequency higher than the first threshold and the data table with an access frequency lower than the second threshold, the method further includes: When the data volume of the scanned job log is less than the first data volume, based on the operation and maintenance list, data tables in the cold storage cluster with an access frequency higher than a first threshold are migrated to the hot storage cluster, and data tables in the hot storage cluster with an access frequency lower than a second threshold are migrated to the cold storage cluster.

7. A data storage device comprising: a determination module, configured to determine access data of any data table in a job log obtained by scanning, wherein the job log is used to indicate operation records of multiple data tables; The determining module is further configured to determine the access frequency and access type of the data table according to access data in the data table; The determining module is further configured to determine a storage identifier of the data table according to the access frequency and access type of the data table; The processing module is further configured to store the data table in a storage cluster that matches the storage identifier.

8. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.