Storage management method and device of database, equipment and medium
By performing multi-dimensional analysis and automatic optimization of database storage performance, a list of hot and cold data distributions and optimization instructions are generated, which solves the problems of low efficiency of manual operation and maintenance and insufficient accuracy of automatic optimization in existing technologies, thereby improving database performance and reducing labor costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINZHUAN INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
In scenarios with massive amounts of data, existing technologies suffer from low efficiency in manual operation and maintenance optimization, while automatic optimization tools cannot comprehensively solve problems such as slow cross-node access and mismatch between software and hardware resources, resulting in limited optimization effects.
By performing multi-dimensional analysis of database storage performance, a list of hot and cold data distributions, a cross-resource access topology diagram, and storage performance analysis results are generated. Based on the results, optimization instructions are generated to perform automatic optimization operations, and an optimization evaluation report is generated.
It improves database performance, identifies issues such as mixed hot and cold data, implicit redundancy, and distributed storage, reduces labor costs, and is suitable for large-scale cluster operation and maintenance.
Smart Images

Figure CN121880327A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database management technology, and in particular to a database storage management method, apparatus, device, and medium. Background Technology
[0002] As the data volume across multiple business areas grows exponentially, in order to avoid the imbalance in storage resource allocation caused by the mixed storage of massive amounts of data, as well as the performance bottlenecks caused by complex access links across nodes and hardware, it is necessary to optimize database management to achieve dynamic adaptation between data storage formats and access requirements.
[0003] In existing technologies, optimization can be achieved through manual operation and maintenance or through automatic optimization tools. Manual optimization involves operation and maintenance personnel manually analyzing storage problems based on database logs and performance monitoring tools to generate optimization solutions. Automatic optimization tools typically monitor the performance of a specific storage component in the database and generate alerts based on thresholds.
[0004] However, manual operation and maintenance optimization is inefficient, especially in scenarios with massive amounts of data. Manual analysis is costly, slow to respond, and cannot adapt to the dynamic changes in data access behavior in a timely manner. Automated optimization through optimization tools can generally only monitor a single dimension, resulting in insufficient optimization accuracy and an inability to comprehensively solve problems such as slow cross-node access and mismatch between software and hardware resources, thus limiting the optimization effect. Summary of the Invention
[0005] This invention provides a database storage management method, apparatus, device, and medium that can perform multi-dimensional analysis and optimization of database storage performance, thereby improving overall data storage efficiency and access speed.
[0006] According to one aspect of the present invention, a database storage management method is provided, comprising:
[0007] When an optimization task request is received, the optimization task request is parsed to obtain task configuration parameters; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode;
[0008] Based on the pre-established correlation between each granularity of data and storage resources, and the monitoring cycle configuration, access behavior data of each granularity of data is collected periodically in the target database cluster.
[0009] Based on the access behavior data of each granularity, a cold and hot data distribution list, a cross-resource access topology map, and storage performance analysis results are generated.
[0010] If the optimization mode is automatic optimization, then optimization instructions are generated based on the cold and hot data distribution list, cross-resource access topology diagram and storage performance analysis results, and optimization operations are performed according to the optimization instructions. After the optimization is completed, an optimization evaluation report is generated.
[0011] According to another aspect of the present invention, a database storage management apparatus is provided, comprising:
[0012] The task configuration parameter acquisition module is used to parse the optimization task request and acquire the task configuration parameters when an optimization task request is received; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode;
[0013] The access behavior data acquisition module is used to periodically collect access behavior data of each granularity of data in the target database cluster according to the pre-established association between each granularity of data and storage resources and the monitoring period configuration.
[0014] The data analysis module is used to generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results based on access behavior data at various granularities.
[0015] An automatic optimization module is used to generate optimization instructions based on the cold and hot data distribution list, cross-resource access topology map, and storage performance analysis results if the optimization mode is automatic optimization. The module then performs optimization operations according to the optimization instructions and generates an optimization evaluation report after determining that the optimization is complete.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the database storage management method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the database storage management method described in any embodiment of the present invention.
[0021] The technical solution of this invention, upon receiving an optimization task request, parses the request to obtain task configuration parameters. Based on pre-established associations between data at various granularities and storage resources, and the monitoring cycle configuration, it periodically collects access behavior data for each granularity of data in the target database cluster. Based on the access behavior data, it generates a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results. If the optimization mode is automatic optimization, it generates optimization instructions based on the list of hot and cold data distributions, the cross-resource access topology map, and the storage performance analysis results. The optimization operation is then executed according to the optimization instructions, and an optimization evaluation report is generated after the optimization is completed. This approach can automatically identify problems such as mixed hot and cold data, implicit redundancy, and distributed storage that are difficult for manual maintenance to detect, improving identification accuracy. Targeted optimization is performed based on the analysis results, effectively improving database performance. This method is suitable for large-scale cluster maintenance and reduces labor costs.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a database storage management method according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of another database storage management method provided according to Embodiment 2 of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of a database storage management device according to Embodiment 3 of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the database storage management method of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1 This is a flowchart of a database storage management method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations involving monitoring and optimizing database storage. The method can be executed by a database storage management device, which can be implemented in hardware and / or software, and is generally configured in a computer or processor with data processing capabilities. Figure 1 As shown, the method includes:
[0032] S110. When an optimization task request is received, the optimization task request is parsed to obtain the task configuration parameters.
[0033] The task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode.
[0034] Optionally, optimization task requests can refer to instructions initiated by the user or automatically triggered by the system based on performance thresholds, used to optimize database storage structure and access efficiency.
[0035] Optionally, the task configuration parameters can describe the target database cluster by the database cluster identifier and node information to be optimized. The data granularity range can include single or combined granularities such as tables, pages, rows, and columns. The monitoring period configuration can include the length of the entire monitoring period and the collection frequency of data at each granularity. The optimization mode can include automatic optimization or monitoring only.
[0036] Optionally, after receiving an optimization task request, you can first verify the interface protocol and permissions, reject illegal or insufficient requests and return a clear prompt, and then parse the request after it passes the verification.
[0037] S120. Based on the pre-established association between each granularity of data and storage resources, and the monitoring cycle configuration, periodically collect access behavior data of each granularity of data in the target database cluster.
[0038] Optionally, the association between data at each granularity and storage resources can refer to the dynamic mapping relationship between pre-established data logical identifiers and hardware physical resources. This is used to record the correspondence between data units such as tables, pages, rows, and columns and storage nodes, disk partitions, and physical addresses, and is updated in real time as data storage location or hardware resources change.
[0039] Optionally, data logical identifiers can be generated in advance for each granularity of data. That is, a data logical identifier is generated for each granularity of data in the database, such as each table, each page, and each row. Then, the association between the data logical identifier and the storage resources is established, and the corresponding access behavior data is collected according to the data logical identifier during data collection.
[0040] Optionally, access behavior data can refer to multi-dimensional data generated when data at various granularities are accessed during database operation, which may include access frequency, access time period, access link, joint query records, data replica distribution, etc.
[0041] Optionally, hardware changes can be detected in real time through the storage resource monitoring interface, and the correlation between data at each granularity and storage resources can be dynamically updated to ensure effectiveness.
[0042] Optionally, database points can be pre-installed when the database is idle, linking granular data such as tables, pages, rows, and columns on the database side to hardware resources.
[0043] S130. Based on the access behavior data of each granularity, generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results.
[0044] Optionally, the hot and cold data distribution list can refer to a structured list formed by classifying data of various granularities into hot data, ordinary data, and cold data based on data popularity values and preset thresholds, clarifying the popularity level and storage location of each data, and thus forming a structured list.
[0045] Optionally, the cross-resource access topology map is a graphical representation of the data access path across storage resources. It can include information such as access links, link access frequency, and latency levels. Through the cross-resource access topology map, high-latency bottleneck links can be located intuitively.
[0046] Optionally, storage performance analysis results may refer to comprehensive evaluation results generated based on access behavior data, which may include information such as redundant data identification results, scattered data determination results, and storage resource utilization.
[0047] Optionally, after acquiring access behavior data at various granularities, the collected access behavior data can be cleaned and standardized, and a cold and hot data distribution list, a cross-resource access topology map, and storage performance analysis results can be generated based on the cleaned data.
[0048] Optionally, the analysis of access behavior data can adopt an asynchronous processing mechanism, allowing for some missing data and sporadic access uncertainties, and using algorithms to filter core data to ensure the accuracy of the analysis results.
[0049] S140. If the optimization mode is automatic optimization, an optimization instruction is generated based on the cold and hot data distribution list, the cross-resource access topology diagram, and the storage performance analysis results. The optimization operation is executed according to the optimization instruction, and an optimization evaluation report is generated after the optimization is completed.
[0050] Optionally, optimization instructions can refer to structured execution commands generated in automatic optimization mode, which may include specific operational logic such as data migration paths and storage form adjustments.
[0051] Optionally, the optimization evaluation report can refer to the quantitative report of the effect generated after the optimization is completed. It can include details of the optimization operation, comparison of performance indicators before and after optimization, explanation of unoptimized items, etc., for subsequent auditing.
[0052] After generating a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results based on access behavior data at each granularity, the process may also include:
[0053] If the optimization mode is monitoring only, an optimization report will be generated and sent to the user terminal based on the cold and hot data distribution list, cross-resource access topology map, and storage performance analysis results.
[0054] When an optimization instruction is received from the user terminal, the optimization operation is performed according to the optimization instruction.
[0055] Optionally, the report to be optimized may refer to the structured report generated in the monitoring mode only, which integrates information such as hot and cold data distribution, access bottlenecks, redundant or scattered data, and provides corresponding suggested optimization strategies for users to decide whether to perform optimization.
[0056] Optionally, after uploading the report to be optimized, you can listen for user commands in real time. If you receive an optimization command from the user, you can perform the optimization operation according to the command after verifying the legality of the command.
[0057] The database storage management method may further include:
[0058] During the optimization process, the status data of the current optimization database is monitored in real time;
[0059] If an abnormal state of the currently optimized database is detected, the optimization operation will be paused and a rollback mechanism will be triggered to restore the currently optimized database to its initial state before optimization.
[0060] Optionally, it is possible to determine whether there are any abnormalities by monitoring the status data of the current optimized database. The status data can refer to the real-time running data of the database when the optimization operation is executed, which may include information such as CPU utilization, data transmission interface usage, data consistency verification results, and access latency.
[0061] Optionally, the rollback mechanism can refer to the recovery mechanism triggered when an optimization fails. Specifically, it can include restoring the database to its pre-optimization state based on the pre-backed-up original data, thereby ensuring data security and business continuity.
[0062] Optionally, when an anomaly is detected in the currently optimized database state, the following steps are also included: recording an anomaly log, including the time of the anomaly, anomaly indicators, and the optimization steps that have been executed, and providing feedback to the user and prompting them to investigate the cause.
[0063] The technical solution of this invention, upon receiving an optimization task request, parses the request to obtain task configuration parameters. Based on pre-established associations between data at various granularities and storage resources, and the monitoring cycle configuration, it periodically collects access behavior data for each granularity of data in the target database cluster. Based on the access behavior data, it generates a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results. If the optimization mode is automatic optimization, it generates optimization instructions based on the list of hot and cold data distributions, the cross-resource access topology map, and the storage performance analysis results. The optimization operation is then executed according to the optimization instructions, and an optimization evaluation report is generated after the optimization is completed. This approach can automatically identify problems such as mixed hot and cold data, implicit redundancy, and distributed storage that are difficult for manual maintenance to detect, improving identification accuracy. Targeted optimization is performed based on the analysis results, effectively improving database performance. This method is suitable for large-scale cluster maintenance and reduces labor costs.
[0064] Example 2
[0065] Figure 2 This is a flowchart illustrating a database storage management method according to Embodiment 2 of the present invention. Based on the above embodiments, this embodiment specifically describes the database storage management method. Figure 2 As shown, the method includes:
[0066] S210. When an optimization task request is received, the optimization task request is parsed to obtain the task configuration parameters.
[0067] S220. Based on the pre-established association between each granularity of data and storage resources, and the monitoring cycle configuration, periodically collect access behavior data of each granularity of data in the target database cluster.
[0068] S230. Based on the access behavior data of each granularity, generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results.
[0069] Based on access behavior data at each granularity, a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results are generated, which may include:
[0070] Based on the access frequency and access time distribution data of the first data, the popularity value of the first data is calculated, and the first data is classified and labeled according to the pre-set popularity classification threshold.
[0071] A hot and cold data distribution list is generated based on the popularity classification label of the first data and the association between the first data and storage resources.
[0072] Optionally, the first data, second data, third data, fourth data, and fifth data can refer to data of any granularity stored in the database.
[0073] Optionally, by extracting the access behavior data of the first data, information such as the number of accesses within the period and the distribution of access time periods of the first data can be obtained. Then, a weighted algorithm can be used to calculate the popularity value: Popularity value = (number of accesses within the period / maximum number of accesses in each granularity of data within the period) × first coefficient + (percentage of accesses during peak periods) × second coefficient. The sum of the first coefficient and the second coefficient can be 1.
[0074] Optionally, based on the heat value and the heat value range pre-divided for hot data, ordinary data and cold data, the classification of the first data can be determined and the first data can be classified and labeled.
[0075] Optionally, the correspondence between the first data and storage resources can be linked to generate a structured cold and hot data distribution list, which includes information such as data identifier, heat value, storage location, and associated hardware resources.
[0076] Based on the access behavior data at each granularity, a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results are generated, which may include:
[0077] Based on the access behavior data of the second data, determine the access links of the second data across storage resources, and obtain the access frequency and latency level of each access link;
[0078] A cross-resource access topology map is generated based on the access frequency and latency level of each access link.
[0079] Optionally, by parsing the access behavior data of the second data, access records across storage resources such as nodes and disks can be extracted to determine the source storage address, target storage address, number of accesses, and average latency of each access. Then, according to the predefined latency level standards of low latency, medium latency, and high latency, the latency level of each access link can be marked.
[0080] Optionally, nodes in the cross-resource access topology map can represent hardware resources such as storage nodes or disks, links represent access paths, link thickness is positively correlated with access frequency, and link color corresponds to latency level, thereby generating an intuitive cross-resource access topology map.
[0081] The process of generating a hot and cold data distribution list, a cross-resource access topology map, and storage performance analysis results based on access behavior data at each granularity may include at least one of the following:
[0082] The storage distribution and associated access ratio of the same logical identifier of data are statistically analyzed. When it is determined that the third data has a copy in multiple storage resources and the associated access ratio is lower than the preset access ratio value, the third data is determined to be redundant data.
[0083] The frequency ratio of joint queries between the fourth and fifth data is statistically analyzed. If the frequency ratio of joint queries is higher than the preset joint query ratio, and the fourth and fifth data are stored in different tables or storage resources, then the fourth and fifth data are determined to be distributed data.
[0084] Optionally, based on the distribution of the same data logical identifier in storage resources, the number of data replicas and the hardware resources where they are located can be determined, and then the proportion of related accesses to the data can be calculated, that is, the proportion of the frequency of related queries to different replicas to the total access frequency.
[0085] Optionally, the relationship between the associated access ratio and the preset access ratio can be determined. If the associated access ratio is lower than the threshold and the number of replicas is ≥2, it is determined to be redundant data.
[0086] Optionally, the joint query records of the fourth and fifth data can be extracted to determine the proportion of the joint query frequency to the total access frequency of the two data. If the joint query proportion is higher than the preset joint query proportion threshold, and the two data are stored in different tables or different storage resources, they are determined to be scattered data.
[0087] S240. If the optimization mode is automatic optimization, an optimization instruction is generated based on the cold and hot data distribution list, the cross-resource access topology diagram, and the storage performance analysis results. The optimization operation is executed according to the optimization instruction, and an optimization evaluation report is generated after the optimization is completed.
[0088] The optimization instructions generated based on the hot and cold data distribution list, cross-resource access topology map, and storage performance analysis results may include:
[0089] Based on the cold and hot data distribution list, cold data is migrated to low-speed storage, and hot data is migrated to high-speed storage;
[0090] Based on the cross-resource access topology, data with excessive cross-node access latency will be migrated to nodes in the access source set.
[0091] Based on the storage performance analysis results, redundant data was deleted, and scattered data was stored in a combined table.
[0092] Optionally, the optimization instructions may include hot and cold data separation instructions, node migration instructions, and redundant / distributed data optimization instructions.
[0093] Optionally, the hot and cold data separation instruction is a data migration instruction generated based on the hot and cold data distribution list. Cold data is migrated to low-speed storage, such as mechanical hard drives, and hot data is migrated to high-speed storage, such as SSDs. The hot and cold data separation instruction may include information such as migration source address, target address, and execution window period.
[0094] Optionally, the node migration instruction is based on the cross-resource access topology map, identifies high-latency cross-node access links, generates node migration instructions, and is used to migrate the target data corresponding to the link to the node in the access source set, thereby reducing cross-network transmission.
[0095] Optionally, redundant / distributed data optimization instructions are generated based on storage performance analysis results, including delete instructions and / or merge instructions. Delete instructions can be used to delete redundant data copies and retain the copy with the highest proportion of related accesses. Merge instructions can be used to merge distributed data into the same table or the same storage resource to optimize the efficiency of join queries.
[0096] S250. During the optimization operation, the status data of the current optimization database is monitored in real time.
[0097] S260. If an abnormal state of the current optimized database is detected, the optimization operation is paused and a rollback mechanism is triggered to restore the current optimized database to its initial state before optimization.
[0098] The technical solution of this invention, upon receiving an optimization task request, parses the request to obtain task configuration parameters. Based on pre-established associations between data at various granularities and storage resources, and the monitoring cycle configuration, it periodically collects access behavior data for each granularity of data in the target database cluster. Based on the access behavior data, it generates a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results. If the optimization mode is automatic optimization, it generates optimization instructions based on the list of hot and cold data distributions, the cross-resource access topology map, and the storage performance analysis results. The optimization operation is then executed according to the optimization instructions, and an optimization evaluation report is generated after the optimization is completed. This approach can automatically identify problems such as mixed hot and cold data, implicit redundancy, and distributed storage that are difficult for manual maintenance to detect, improving identification accuracy. Targeted optimization is performed based on the analysis results, effectively improving database performance. This method is suitable for large-scale cluster maintenance and reduces labor costs.
[0099] Example 3
[0100] Figure 3 This is a schematic diagram of the structure of a database storage management device provided in Embodiment 3 of the present invention. Figure 3As shown, the device includes: a task configuration parameter acquisition module 310, an access behavior data acquisition module 320, a data analysis module 330, and an automatic optimization module 340.
[0101] The task configuration parameter acquisition module 310 is used to parse the optimization task request and acquire task configuration parameters when an optimization task request is received; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode.
[0102] The access behavior data acquisition module 320 is used to periodically collect access behavior data of each granularity of data in the target database cluster according to the pre-established association between each granularity of data and storage resources and the monitoring period configuration.
[0103] The data analysis module 330 is used to generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results based on the access behavior data of data at various granularities.
[0104] The automatic optimization module 340 is used to generate optimization instructions based on the cold and hot data distribution list, cross-resource access topology map and storage performance analysis results if the optimization mode is automatic optimization, to perform optimization operations according to the optimization instructions, and to generate an optimization evaluation report after determining that the optimization is completed.
[0105] The technical solution of this invention, upon receiving an optimization task request, parses the request to obtain task configuration parameters. Based on pre-established associations between data at various granularities and storage resources, and the monitoring cycle configuration, it periodically collects access behavior data for each granularity of data in the target database cluster. Based on the access behavior data, it generates a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results. If the optimization mode is automatic optimization, it generates optimization instructions based on the list of hot and cold data distributions, the cross-resource access topology map, and the storage performance analysis results. The optimization operation is then executed according to the optimization instructions, and an optimization evaluation report is generated after the optimization is completed. This approach can automatically identify problems such as mixed hot and cold data, implicit redundancy, and distributed storage that are difficult for manual maintenance to detect, improving identification accuracy. Targeted optimization is performed based on the analysis results, effectively improving database performance. This method is suitable for large-scale cluster maintenance and reduces labor costs.
[0106] Based on the above embodiments, the data analysis module 330 can be used for:
[0107] Based on the access frequency and access time distribution data of the first data, the popularity value of the first data is calculated, and the first data is classified and labeled according to the pre-set popularity classification threshold.
[0108] A hot and cold data distribution list is generated based on the popularity classification label of the first data and the association between the first data and storage resources.
[0109] Based on the above embodiments, the data analysis module 330 can be used for:
[0110] Based on the access behavior data of the second data, determine the access links of the second data across storage resources, and obtain the access frequency and latency level of each access link;
[0111] A cross-resource access topology map is generated based on the access frequency and latency level of each access link.
[0112] Based on the above embodiments, the data analysis module 330 can be used to perform any of the following:
[0113] The storage distribution and associated access ratio of the same logical identifier of data are statistically analyzed. When it is determined that the third data has a copy in multiple storage resources and the associated access ratio is lower than the preset access ratio value, the third data is determined to be redundant data.
[0114] The frequency ratio of joint queries between the fourth and fifth data is statistically analyzed. If the frequency ratio of joint queries is higher than the preset joint query ratio, and the fourth and fifth data are stored in different tables or storage resources, then the fourth and fifth data are determined to be distributed data.
[0115] Based on the above embodiments, the automatic optimization module 340 may include:
[0116] Based on the cold and hot data distribution list, cold data is migrated to low-speed storage, and hot data is migrated to high-speed storage;
[0117] Based on the cross-resource access topology, data with excessive cross-node access latency will be migrated to nodes in the access source set.
[0118] Based on the storage performance analysis results, redundant data was deleted, and scattered data was stored in a combined table.
[0119] Based on the above embodiments, an optimized reporting module may also be included, used for:
[0120] If the optimization mode is monitoring only, an optimization report will be generated and sent to the user terminal based on the cold and hot data distribution list, cross-resource access topology map, and storage performance analysis results.
[0121] When an optimization instruction is received from the user terminal, the optimization operation is performed according to the optimization instruction.
[0122] Based on the above embodiments, an optimized monitoring module may also be included, for:
[0123] During the optimization process, the status data of the current optimization database is monitored in real time;
[0124] If an abnormal state of the currently optimized database is detected, the optimization operation will be paused and a rollback mechanism will be triggered to restore the currently optimized database to its initial state before optimization.
[0125] The database storage management device provided in the embodiments of the present invention can execute the database storage management method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0126] Example 4
[0127] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0128] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0129] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0130] Processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the database storage management method described in the embodiments of the present invention. That is:
[0131] When an optimization task request is received, the optimization task request is parsed to obtain task configuration parameters; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode;
[0132] Based on the pre-established correlation between each granularity of data and storage resources, and the monitoring cycle configuration, access behavior data of each granularity of data is collected periodically in the target database cluster.
[0133] Based on the access behavior data of each granularity, a cold and hot data distribution list, a cross-resource access topology map, and storage performance analysis results are generated.
[0134] If the optimization mode is automatic optimization, then optimization instructions are generated based on the cold and hot data distribution list, cross-resource access topology diagram and storage performance analysis results, and optimization operations are performed according to the optimization instructions. After the optimization is completed, an optimization evaluation report is generated.
[0135] In some embodiments, the database storage management method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the database storage management method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the database storage management method by any other suitable means (e.g., by means of firmware).
[0136] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0140] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0141] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0142] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0143] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A storage management method of a database, characterized by, include: When an optimization task request is received, the optimization task request is parsed to obtain task configuration parameters; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode; Based on the pre-established correlation between each granularity of data and storage resources, and the monitoring cycle configuration, access behavior data of each granularity of data is collected periodically in the target database cluster. Based on the access behavior data of each granularity, a cold and hot data distribution list, a cross-resource access topology map, and storage performance analysis results are generated. If the optimization mode is automatic optimization, then optimization instructions are generated based on the cold and hot data distribution list, cross-resource access topology diagram and storage performance analysis results, optimization operations are executed according to the optimization instructions, and an optimization evaluation report is generated after the optimization is completed.
2. The method of claim 1, wherein, Based on access behavior data at each granularity, a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results are generated, including: Based on the access frequency and access time distribution data of the first data, the popularity value of the first data is calculated, and the first data is classified and labeled according to the pre-set popularity classification threshold. A hot and cold data distribution list is generated based on the popularity classification label of the first data and the association between the first data and storage resources.
3. The method of claim 1, wherein, Based on access behavior data at each granularity, a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results are generated, including: Based on the access behavior data of the second data, determine the access links of the second data across storage resources, and obtain the access frequency and latency level of each access link; A cross-resource access topology map is generated based on the access frequency and latency level of each access link.
4. The method of claim 1, wherein, Based on access behavior data at each granularity, generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results, including at least one of the following: The storage distribution and associated access ratio of the same logical identifier of data are statistically analyzed. When it is determined that the third data has a copy in multiple storage resources and the associated access ratio is lower than the preset access ratio value, the third data is determined to be redundant data. The frequency ratio of joint queries between the fourth and fifth data is statistically analyzed. If the frequency ratio of joint queries is higher than the preset joint query ratio, and the fourth and fifth data are stored in different tables or storage resources, then the fourth and fifth data are determined to be distributed data.
5. The method of claim 1, wherein, Based on the hot and cold data distribution list, cross-resource access topology map, and storage performance analysis results, optimization instructions are generated, including: Based on the cold and hot data distribution list, cold data is migrated to low-speed storage, and hot data is migrated to high-speed storage; Based on the cross-resource access topology, data with excessive cross-node access latency will be migrated to nodes in the access source set. Based on the storage performance analysis results, redundant data was deleted, and scattered data was stored in a combined table.
6. The method of claim 1, wherein, After generating a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results based on access behavior data at various granularities, the process also includes: If the optimization mode is monitoring only, an optimization report will be generated and sent to the user terminal based on the cold and hot data distribution list, cross-resource access topology map, and storage performance analysis results. When an optimization instruction is received from the user terminal, the optimization operation is performed according to the optimization instruction.
7. The method according to any one of claims 1 to 6, characterized in that, Also includes: During the optimization process, the status data of the current optimization database is monitored in real time; If an abnormal state of the currently optimized database is detected, the optimization operation will be paused and a rollback mechanism will be triggered to restore the currently optimized database to its initial state before optimization.
8. A storage management apparatus of a database, characterized by comprising: include: The task configuration parameter acquisition module is used to parse the optimization task request and acquire the task configuration parameters when an optimization task request is received; wherein, the task configuration parameters include the target database cluster, data granularity range, monitoring cycle configuration, and optimization mode; The access behavior data acquisition module is used to periodically collect access behavior data of each granularity of data in the target database cluster according to the pre-established association between each granularity of data and storage resources and the monitoring period configuration. The data analysis module is used to generate a list of hot and cold data distributions, a cross-resource access topology map, and storage performance analysis results based on access behavior data at various granularities. An automatic optimization module is used to generate optimization instructions based on the cold and hot data distribution list, cross-resource access topology map, and storage performance analysis results if the optimization mode is automatic optimization. The module then performs optimization operations according to the optimization instructions and generates an optimization evaluation report after determining that the optimization is complete.
9. An electronic device, comprising: The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the storage management method of the database according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the storage management method of the database according to any one of claims 1-7.