Storage architecture based on multi-dimensional indexes and logic data partitions
Through the logical data partition storage architecture based on multi-dimensional indicators, the problem of low data storage and access efficiency in the existing storage architecture is solved, intelligent allocation and storage optimization of data are realized, the scalability and availability of the system are improved, and the storage cost is reduced.
Patent Information
- Application Number
- CN202510353967.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-29
AI Technical Summary
When processing massive data, existing storage architectures face problems such as low data storage and access efficiency, unreasonable allocation of storage resources, poor system scalability, and lack of unified management and access mechanisms.
The storage architecture based on multi-dimensional indicators and logical data partitions is adopted, including physical storage component management module, data partition management module, data index management module, data migration management module and unified data access service module. The logical data partition is bound through multi-dimensional data evaluation indicators to realize intelligent allocation and storage optimization of data, and provide a unified data reading and writing interface.
It improves data access efficiency and storage resource utilization, reduces storage costs, enhances the scalability and high availability of the system, and meets the diversified needs in different business scenarios.
Smart Images

Figure CN120386486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage and management, and more specifically, the present invention relates to a storage architecture based on multi-dimensional metrics and logical data partitioning. Background Art
[0002] With the rapid development of big data technology, the scale of data has shown an explosive growth, and traditional storage architectures face many challenges when dealing with massive data. Existing storage systems usually adopt a single storage medium or a simple hierarchical storage strategy, which is difficult to meet the diverse requirements for data access performance, storage cost, and scalability in different business scenarios. Although some technologies attempt to optimize storage efficiency through data partitioning or migration, these methods often lack a multi-dimensional evaluation of data characteristics, resulting in unreasonable data allocation and low utilization of storage resources. In addition, the management of physical storage components in existing technologies is relatively scattered, lacking a unified access interface, which increases the complexity of system maintenance and expansion.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the existing technology: low data storage and access efficiency, unreasonable storage resource allocation, poor system scalability, and lack of a unified management and access mechanism. Summary of the Invention
[0004] The present invention provides a storage architecture based on multi-dimensional metrics and logical data partitioning, including:
[0005] A physical storage component management module, a data partitioning management module, a data metric management module, a data migration management module, and a unified data access service module;
[0006] The physical storage component management module is used to select and associate physical storage components according to business scenarios and data access characteristics;
[0007] The data partitioning management module is used to define logical data partitions and their priority scores, and associate the logical data partitions with physical storage components;
[0008] The data metric management module is used to define multi-dimensional data evaluation metrics, and bind logical data partitions according to the calculation results of the multi-dimensional data evaluation metrics;
[0009] The data migration management module is used to periodically scan data and perform data migration based on the calculation results of the multi-dimensional data evaluation metrics;
[0010] The unified data access service module is used to provide a unified data read / write interface externally, shielding the differences of physical storage components.
[0011] Further, the specific implementation steps of the physical storage component management module include:
[0012] Select a physical storage component type based on the business scenario and data access characteristics. The physical storage component types include relational databases, non-relational databases, or custom storage components;
[0013] Write data to or read data from the physical storage component by inheriting and implementing the standard data interface of the physical storage component.
[0014] Further, the specific implementation steps of the data partitioning management module include:
[0015] Define multiple logical data partitions and set a priority score for each logical data partition. The priority score ranges from 1 to 10, and the higher the score, the higher the priority;
[0016] Associate each logical data partition with the corresponding physical storage component;
[0017] Create a data temporary partition, which is used to cache the written data and is implemented by using a high-IO sequential read / write storage component.
[0018] Further, the specific implementation steps of the data metric management module include:
[0019] Define at least two multi-dimensional data evaluation metrics;
[0020] Bind the data to the corresponding logical data partition based on the calculation results of the multi-dimensional data evaluation metrics;
[0021] When a data write request is initiated, calculate the results of the multi-dimensional data evaluation metrics in real time and determine the final logical data partition to be written according to the priority score of the logical data partition.
[0022] Further, the specific implementation steps of the data migration management module include:
[0023] Periodically calculate the multi-dimensional metric results of the data through a scheduled scan task;
[0024] When the multi-dimensional metric results trigger a logical data partition change, generate a migration task;
[0025] Migrate the data from the physical storage component corresponding to the original logical data partition to the physical storage component corresponding to the target logical data partition through a two-phase transaction mechanism to ensure data consistency and integrity.
[0026] Further, the generation condition of the migration task is:
[0027] Data migration is triggered when the priority score of the target logical data partition is higher than that of the current logical data partition, and the physical storage components associated with the target logical data partition are different from those associated with the current logical data partition.
[0028] Further, the specific implementation steps of the unified data access service module include:
[0029] Receive read and write requests from external clients and forward the read and write requests to the data scheduling center;
[0030] The data scheduling center determines the logical data partition to which the data belongs based on the calculation results of the multi-dimensional metrics and returns the corresponding physical storage component information;
[0031] The client performs data read and write operations according to the physical storage component information.
[0032] Further, the usage steps of the data temporary partition include:
[0033] When a data write request triggers multi-dimensional metric calculation, temporarily store the data in the data temporary partition;
[0034] After completing the metric calculation, migrate the data from the data temporary partition to the physical storage component corresponding to the target logical data partition according to the priority score of the logical data partition.
[0035] Further, the multi-dimensional data evaluation metrics include at least two of the following:
[0036] Real-time metric, access frequency metric, data type metric, where the determination result of each metric is directly mapped to the corresponding logical data partition.
[0037] Further, the dynamic adjustment rule of the priority score of the logical data partition is:
[0038] If multiple logical data partitions all meet the metric binding conditions, select the logical data partition with the highest priority score as the final storage location;
[0039] If the priority scores are the same, select the logical data partition according to the preset default rule.
[0040] The above embodiments of the present invention have at least the following beneficial effects: The present invention can comprehensively analyze data based on a multi-dimensional index evaluation system, and combine logical data partitioning to achieve intelligent allocation and storage optimization of data, thereby improving data access efficiency and storage resource utilization. Through the collaborative work of the physical storage component management module and the data partitioning management module, the system can flexibly adapt to different storage media according to the business scenario, ensuring the rationality and efficiency of data storage. In addition, the data migration management module can scan and dynamically migrate data regularly to further optimize storage performance and reduce storage costs.
[0041] The present invention can also provide a standardized data read / write interface to the outside world through the unified data access service module, shielding the differences of underlying physical storage components and simplifying the system maintenance and expansion process. This design can enhance the scalability and high availability of the system, meeting the diverse needs under different business scenarios. At the same time, based on the binding mechanism of logical data partitioning and multi-dimensional indexes, the system can achieve refined management of data, providing more flexible and efficient support for subsequent data analysis and applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, wherein:
[0043] Figure 1 FIG. is a schematic structural diagram of a storage architecture based on multi-dimensional indexes and logical data partitioning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to convey the scope of the present invention to those skilled in the art completely.
[0045] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0046] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0047] The following refers to Figure 1 , Figure 1 which is a schematic structural diagram of a storage architecture based on multi-dimensional metrics and logical data partitioning provided by an embodiment of the present invention. As Figure 1 shown, a storage architecture based on multi-dimensional metrics and logical data partitioning includes:
[0048] a physical storage component management module, a data partitioning management module, a data metric management module, a data migration management module, and a unified data access service module;
[0049] The physical storage component management module is used to select and associate physical storage components according to business scenarios and data access characteristics;
[0050] The data partitioning management module is used to define logical data partitions and their priority scores, and associate the logical data partitions with physical storage components;
[0051] The data metric management module is used to define multi-dimensional data evaluation metrics, and bind logical data partitions according to the calculation results of the multi-dimensional data evaluation metrics;
[0052] The data migration management module is used to periodically scan data and perform data migration based on the calculation results of the multi-dimensional data evaluation metrics;
[0053] The unified data access service module is used to provide a unified data read / write interface externally, shielding the differences of physical storage components.
[0054] It should be noted that the physical storage component management module can select and associate physical storage components according to business scenarios and data access characteristics. The physical storage component management module is one of the core modules of the system, responsible for selecting appropriate storage media according to business requirements and data access patterns, such as high-speed solid-state drives, mechanical hard drives, or cloud storage, etc., and associating these components with the system to ensure the efficiency and flexibility of data storage. Business scenarios include but are not limited to real-time data processing, offline analysis, or hybrid modes, while data access characteristics involve factors such as read / write frequency, latency requirements, and data volume size.
[0055] Specifically, when selecting a storage medium, the physical storage component management module will comprehensively consider parameters such as performance, cost, and scalability. For example, for hot data with high access frequency, a high-speed solid-state drive can be selected to reduce latency; for cold data with low access frequency, a mechanical hard drive or cloud storage with lower cost can be selected. In addition, the module can also dynamically adjust the storage strategy according to business requirements, such as temporarily increasing high-performance storage resources during peak business periods, or reducing resources during off-peak business periods to reduce costs. The association method of physical storage components can be achieved through configuration files, API interfaces, or automated scripts, ensuring that the system can flexibly adapt to different storage environments.
[0056] More specifically, the physical storage component management module can optimize storage selection through a priority scoring mechanism. For example, assign a priority score to each business scenario and data access characteristic, and then select the most suitable storage component according to the score. In an alternative solution, the module can also introduce machine learning algorithms to predict future storage requirements through historical data analysis and automatically adjust the storage strategy. In addition, the module can support the mixed use of multiple storage media. For example, store hot data on local high-speed solid-state drives while migrating cold data to the cloud to achieve optimal configuration of storage resources and cost control.
[0057] In some embodiments, the specific implementation steps of the physical storage component management module include:
[0058] Select the type of physical storage component based on the business scenario and data access characteristics, where the type of physical storage component includes relational databases, non-relational databases, or custom storage components;
[0059] Write data to or read data from the physical storage component by inheriting and implementing the standard data interface of the physical storage component.
[0060] It should be noted that the data partition management module can define logical data partitions and their priority scores and associate them with physical storage components. A logical data partition refers to dividing data into different logical units according to business requirements or data characteristics, such as dividing by time, region, or business type. The priority score is used to evaluate the importance or access frequency of each logical data partition in order to allocate the most suitable physical storage resources for it. The data partition management module ensures the efficiency and flexibility of data storage by associating logical data partitions with physical storage components.
[0061] Specifically, the definition of logical data partitions can be based on multiple dimensions, such as the timestamp, geographical location, business type, or access pattern of the data. The setting of the priority score can be achieved through a scoring model. For example, score according to access frequency, data importance, or business requirements. The higher the score, the higher the priority. When associating with physical storage components, the module can select the most suitable storage medium according to the priority score. For example, associate high-priority partitions with high-speed solid-state drives, while associating low-priority partitions with lower-cost mechanical hard drives or cloud storage. In addition, the module also supports dynamically adjusting the association relationship between partitions and storage components to adapt to changes in business requirements.
[0062] More specifically, the data partition management module can implement the definition and priority scoring of logical data partitions through automated tools. For example, the module can integrate data analysis functions to automatically divide logical partitions and calculate priority scores based on historical access patterns. In an alternative solution, the module can also support manual configuration, allowing administrators to customize partition rules and priority scoring criteria according to business requirements. In addition, the module can introduce intelligent algorithms, such as machine learning-based partition optimization strategies, to dynamically adjust the association relationship between partitions and storage components by analyzing data access trends, so as to further improve storage efficiency and system performance.
[0063] In some embodiments, the specific implementation steps of the data partition management module include:
[0064] Define multiple logical data partitions and set a priority score for each logical data partition. The priority score ranges from 1 to 10, and the higher the score, the higher the priority.
[0065] Associate each logical data partition with a corresponding physical storage component.
[0066] Create a data temporary partition, which is used to cache the written data and is implemented by high-IO sequential read and write of the storage component.
[0067] It should be noted that the data metric management module can define multi-dimensional data evaluation metrics and bind logical data partitions according to the calculation results. Multi-dimensional data evaluation metrics refer to metrics that evaluate data characteristics from multiple perspectives, such as access frequency, data size, storage cost, or business importance, etc. By calculating the comprehensive results of these metrics, the module can bind logical data partitions to the most suitable storage strategy, thereby optimizing data storage and access efficiency.
[0068] Specifically, the multi-dimensional data evaluation metrics can include but are not limited to access frequency, data life cycle, storage cost, data security level, or business priority, etc. For example, the access frequency can be evaluated by counting historical access times, the data life cycle can be calculated based on the data creation time and the expected expiration time, and the storage cost can be estimated based on the cost of the storage medium and the occupied space. The module can synthesize these metrics into an evaluation result through weighted calculation or a scoring model, and then bind the logical data partition to the corresponding storage strategy according to the result, such as binding frequently accessed data to high-performance storage and binding infrequently accessed data to low-cost storage.
[0069] More specifically, the data metric management module can achieve the evaluation and binding of multi-dimensional metrics through automated tools. For example, the module can integrate real-time monitoring functions, dynamically collect data access and storage information, and automatically calculate evaluation metrics. In an alternative solution, the module can also support manual configuration, allowing administrators to customize metric weights and binding rules according to business requirements. In addition, the module can introduce intelligent algorithms, such as machine learning-based metric optimization strategies, to dynamically adjust the binding relationship by analyzing data characteristics and access trends, so as to further improve storage efficiency and system performance. For example, for periodically accessed data, the module can predict its access peak and adjust the storage strategy in advance to ensure the efficient operation of the system during the peak period.
[0070] In some embodiments, the specific implementation steps of the data metric management module include:
[0071] Define at least two multi-dimensional data evaluation metrics;
[0072] Based on the calculation results of the multi-dimensional data evaluation metrics, bind the data to the corresponding logical data partitions;
[0073] When a data write request is initiated, calculate the results of the multi-dimensional data evaluation metrics in real time, and determine the final logical data partition to be written according to the priority scores of the logical data partitions.
[0074] It should be noted that the data migration management module can scan the data regularly and perform data migration based on the calculation results. The data migration management module is one of the key components of the system, responsible for migrating data from the current storage location to a more suitable storage location according to the calculation results of multi-dimensional metrics to optimize storage efficiency and access performance. The regular scanning mechanism can ensure that the system can dynamically respond to changes in data characteristics, such as fluctuations in access frequency or changes in the data life cycle, so as to adjust the data storage strategy in a timely manner.
[0075] Specifically, the regular scanning period of the data migration management module can be set according to business requirements, such as performing a scan once an hour, once a day, or once a week. The scanning scope can cover all logical data partitions, or selectively scan some partitions according to the priority scores. When performing data migration based on the calculation results, the module will migrate the data from high-cost or low-performance storage media to low-cost or high-performance storage media according to the evaluation results of multi-dimensional metrics. For example, migrate cold data with decreasing access frequency from high-speed solid-state drives to mechanical hard drives, or migrate hot data with increasing access frequency from mechanical hard drives to high-speed solid-state drives. During the migration process, the module will ensure the integrity and consistency of the data, avoiding data loss or damage caused by migration operations.
[0076] More specifically, the data migration management module can implement scanning and migration operations through automation tools. For example, the module can integrate a task scheduling function to automatically execute scanning and migration tasks according to a preset schedule. In an alternative approach, the module can also support manual triggering, allowing administrators to initiate scanning and migration operations at any time according to business requirements. Additionally, the module can introduce intelligent algorithms, such as machine learning-based migration optimization strategies, which dynamically adjust the migration strategy by analyzing data access trends and storage resource utilization to further improve system performance. For example, for data with a periodic access pattern, the module can predict its access peak and migrate it to a high-performance storage medium in advance to ensure the efficient operation of the system during peak periods.
[0077] In some embodiments, the specific implementation steps of the data migration management module include:
[0078] Periodically calculate the multi-dimensional metric results of the data through a timed scanning task;
[0079] When the multi-dimensional metric results trigger a logical data partition change, generate a migration task;
[0080] Migrate the data from the physical storage component corresponding to the original logical data partition to the physical storage component corresponding to the target logical data partition through a two-phase transaction mechanism to ensure data consistency and integrity.
[0081] It should be noted that the unified data access service module can provide a unified data read / write interface externally, shielding the differences of physical storage components. The unified data access service module is one of the core modules of the system, responsible for providing a standardized data access interface for external applications or users, so that they do not need to pay attention to the specific implementation details of the underlying physical storage components. By shielding the differences of storage components, the module can simplify the data access process and improve the usability and maintainability of the system.
[0082] Specifically, the interface design of the unified data access service module can include basic operations such as data reading, writing, updating, and deleting, and at the same time support functions such as batch processing and transaction management. The implementation of the interface can be completed through a standardized protocol (such as RESTful API or GraphQL) or a custom protocol to ensure compatibility with different storage components. The mechanism for shielding the differences of physical storage components can be implemented through an abstraction layer. For example, define an adapter for each storage component to convert the underlying storage operations into a unified interface call. Additionally, the module can also support the hybrid access of multiple storage media, such as accessing local storage and cloud storage simultaneously, and automatically selecting the optimal storage component according to business requirements.
[0083] More specifically, the unified data access service module can achieve flexible extension and optimization of interfaces through a configuration management tool. For example, the module can support dynamic loading of storage component adapters to add or replace storage components during system operation. In an alternative solution, the module can also introduce a caching mechanism to cache frequently accessed data in memory to further improve data access performance. In addition, the module can integrate an intelligent routing function to automatically select the optimal storage component according to data characteristics and access patterns. For example, frequently accessed data can be routed to high-performance storage media, while less frequently accessed data can be routed to low-cost storage media. This design can further enhance the flexibility and efficiency of the system.
[0084] In some embodiments, the generation condition of the migration task is as follows:
[0085] When the priority score of the target logical data partition is higher than that of the current logical data partition, and the physical storage component associated with the target logical data partition is different from the physical storage component associated with the current logical data partition, data migration is triggered.
[0086] It should be noted that the physical storage component management module can select and associate physical storage components according to the business scenario and data access characteristics. The physical storage component management module is one of the core components of the system and is responsible for selecting the most suitable storage media according to business requirements and data access patterns, such as high-speed solid-state drives, mechanical hard drives, or cloud storage, and associating these components with the system. Business scenarios include, but are not limited to, real-time data processing, offline analysis, or hybrid modes, while data access characteristics involve factors such as read / write frequency, latency requirements, and data volume size.
[0087] Specifically, when selecting a storage medium, the physical storage component management module will comprehensively consider parameters such as performance, cost, and scalability. For example, for hot data with high access frequency, a high-speed solid-state drive can be selected to reduce latency; for cold data with lower access frequency, a lower-cost mechanical hard drive or cloud storage can be selected. When associating physical storage components, the module can be implemented through configuration files, API interfaces, or automation scripts to ensure that the system can flexibly adapt to different storage environments. In addition, the module can also support dynamic adjustment of storage policies, such as temporarily increasing high-performance storage resources during business peak periods or reducing resources during business off-peak periods to reduce costs.
[0088] More specifically, the physical storage component management module can optimize storage selection through a priority scoring mechanism. For example, assign a priority score to each business scenario and data access characteristic, and then select the most suitable storage component based on the score. In an alternative, the module can also introduce machine learning algorithms to predict future storage requirements through historical data analysis and automatically adjust the storage strategy. In addition, the module can support the mixed use of multiple storage media. For example, store hot data on local high-speed solid-state drives while migrating cold data to the cloud to achieve optimal configuration of storage resources and cost control. This design can further enhance the flexibility and efficiency of the system.
[0089] In some embodiments, the specific implementation steps of the unified data access service module include:
[0090] Receive read and write requests from external clients and forward the read and write requests to the data scheduling center;
[0091] The data scheduling center determines the logical data partition to which the data belongs based on the calculation results of the multi-dimensional metrics and returns the corresponding physical storage component information;
[0092] The client performs data read and write operations according to the physical storage component information.
[0093] It should be noted that the data partition management module can define logical data partitions and their priority scores and associate them with physical storage components. A logical data partition refers to dividing data into different logical units according to business requirements or data characteristics, such as dividing by time, region, or business type. The priority score is used to evaluate the importance or access frequency of each logical data partition in order to allocate the most suitable physical storage resources for it. The data partition management module ensures the efficiency and flexibility of data storage by associating logical data partitions with physical storage components.
[0094] Specifically, the definition of logical data partitions can be based on multiple dimensions, such as the timestamp, geographical location, business type, or access pattern of the data. The setting of the priority score can be achieved through a scoring model. For example, score according to the access frequency, data importance, or business requirements. The higher the score, the higher the priority. When associating with physical storage components, the module can select the most suitable storage medium according to the priority score. For example, associate high-priority partitions with high-speed solid-state drives, while associate low-priority partitions with lower-cost mechanical hard drives or cloud storage. In addition, the module also supports dynamically adjusting the association relationship between partitions and storage components to adapt to changes in business requirements.
[0095] More specifically, the data partition management module can implement the definition and priority scoring of logical data partitions through automated tools. For example, the module can integrate data analysis functions to automatically divide logical partitions and calculate priority scores based on historical access patterns. In an alternative solution, the module can also support manual configuration, allowing administrators to customize partition rules and priority scoring criteria according to business requirements. In addition, the module can introduce intelligent algorithms, such as machine learning-based partition optimization strategies, to dynamically adjust the association relationship between partitions and storage components by analyzing data access trends, so as to further improve storage efficiency and system performance. For example, for periodically accessed data, the module can predict its access peak and adjust the storage strategy in advance to ensure the efficient operation of the system during the peak period.
[0096] In some embodiments, the steps of using the data temporary partition include:
[0097] When a data write request triggers multi-dimensional metric calculation, temporarily store the data in the data temporary partition;
[0098] After completing the metric calculation, migrate the data from the data temporary partition to the physical storage component corresponding to the target logical data partition according to the priority score of the logical data partition.
[0099] It should be noted that the data metric management module can define multi-dimensional data evaluation metrics and bind logical data partitions according to the calculation results. Multi-dimensional data evaluation metrics refer to metrics that evaluate data characteristics from multiple perspectives, such as access frequency, data size, storage cost, or business importance, etc. By calculating the comprehensive results of these metrics, the module can bind logical data partitions to the most appropriate storage strategy, thereby optimizing data storage and access efficiency.
[0100] Specifically, the multi-dimensional data evaluation metrics can include, but are not limited to, access frequency, data life cycle, storage cost, data security level, or business priority, etc. For example, the access frequency can be evaluated by counting historical access times, the data life cycle can be calculated based on the data creation time and the expected expiration time, and the storage cost can be estimated according to the cost of the storage medium and the occupied space. The module can synthesize these metrics into an evaluation result through weighted calculation or a scoring model, and then bind the logical data partition to the corresponding storage strategy according to the result, such as binding frequently accessed data to high-performance storage and binding infrequently accessed data to low-cost storage.
[0101] More specifically, the data metric management module can implement the evaluation and binding of multi-dimensional metrics through automated tools. For example, the module can integrate real-time monitoring functions, dynamically collect data access and storage information, and automatically calculate evaluation metrics. In an alternative solution, the module can also support manual configuration, allowing administrators to customize metric weights and binding rules according to business requirements. In addition, the module can introduce intelligent algorithms, such as machine learning-based metric optimization strategies, to dynamically adjust the binding relationship by analyzing data characteristics and access trends, so as to further improve storage efficiency and system performance. For example, for periodically accessed data, the module can predict its access peak and adjust the storage strategy in advance to ensure the efficient operation of the system during the peak period.
[0102] In some embodiments, the multi-dimensional data evaluation metrics include at least two of the following:
[0103] Real-time metrics, access frequency metrics, data type metrics, where the determination result of each metric is directly mapped to the corresponding logical data partition.
[0104] It should be noted that the data migration management module can periodically scan the data and perform data migration based on the calculation results. The data migration management module is one of the key components of the system, responsible for migrating data from the current storage location to a more suitable storage location according to the calculation results of multi-dimensional metrics, so as to optimize storage efficiency and access performance. The periodic scanning mechanism can ensure that the system can dynamically respond to changes in data characteristics, such as fluctuations in access frequency or changes in the data life cycle, so as to adjust the data storage strategy in a timely manner.
[0105] Specifically, the periodic scanning period of the data migration management module can be set according to business requirements, such as performing a scan once per hour, per day, or per week. The scanning scope can cover all logical data partitions, or selectively scan some partitions according to the priority score. When performing data migration based on the calculation results, the module will migrate data from high-cost or low-performance storage media to low-cost or high-performance storage media according to the evaluation results of multi-dimensional metrics. For example, migrating cold data with a decreasing access frequency from a high-speed solid-state drive to a mechanical hard drive, or migrating hot data with an increasing access frequency from a mechanical hard drive to a high-speed solid-state drive. During the migration process, the module will ensure the integrity and consistency of the data, avoiding data loss or damage caused by the migration operation.
[0106] More specifically, the data migration management module can implement scanning and migration operations through automated tools. For example, the module can integrate task scheduling functions to automatically execute scanning and migration tasks according to a preset schedule. In an alternative solution, the module can also support manual triggering, allowing administrators to initiate scanning and migration operations at any time according to business requirements. In addition, the module can introduce intelligent algorithms, such as machine learning-based migration optimization strategies, to dynamically adjust the migration strategy by analyzing data access trends and storage resource utilization, so as to further improve system performance. For example, for data with a periodic access pattern, the module can predict its access peak and migrate it to a high-performance storage medium in advance to ensure the efficient operation of the system during the peak period.
[0107] In some embodiments, the dynamic adjustment rule for the priority score of the logical data partition is as follows:
[0108] If multiple logical data partitions all meet the index binding conditions, select the logical data partition with the highest priority score as the final storage location;
[0109] If the priority scores are the same, select the logical data partition according to the preset default rule.
[0110] It should be noted that the unified data access service module can provide a unified data read-write interface externally, shielding the differences of physical storage components. The unified data access service module is one of the core modules of the system, responsible for providing a standardized data access interface for external applications or users, so that they do not need to pay attention to the specific implementation details of the underlying physical storage components. By shielding the differences of storage components, the module can simplify the data access process and improve the usability and maintainability of the system.
[0111] Specifically, the interface design of the unified data access service module can include basic operations such as data reading, writing, updating, and deleting, and at the same time support functions such as batch processing and transaction management. The implementation of the interface can be completed through a standardized protocol (such as RESTful API or GraphQL) or a custom protocol to ensure compatibility with different storage components. The mechanism for shielding the differences of physical storage components can be implemented through an abstraction layer. For example, an adapter is defined for each storage component to convert the underlying storage operations into unified interface calls. In addition, the module can also support the hybrid access of multiple storage media, such as accessing local storage and cloud storage simultaneously, and automatically selecting the optimal storage component according to business requirements.
[0112] More specifically, the unified data access service module can achieve flexible extension and optimization of interfaces through a configuration management tool. For example, the module can support dynamic loading of storage component adapters to add or replace storage components during system operation. In an alternative solution, the module can also introduce a caching mechanism to cache frequently accessed data in memory to further improve data access performance. In addition, the module can integrate an intelligent routing function to automatically select the optimal storage component according to data characteristics and access patterns. For example, frequently accessed data can be routed to high-performance storage media, while infrequently accessed data can be routed to low-cost storage media. This design can further enhance the flexibility and efficiency of the system.
[0113] The above-mentioned various embodiments of the present invention have the following beneficial effects: The present invention can comprehensively analyze data based on a multi-dimensional index evaluation system, and combine logical data partitioning to achieve intelligent allocation and storage optimization of data, thereby improving data access efficiency and storage resource utilization rate. Through the collaborative work of the physical storage component management module and the data partitioning management module, the system can flexibly adapt to different storage media according to business scenarios to ensure the rationality and efficiency of data storage. In addition, the data migration management module can periodically scan and dynamically migrate data to further optimize storage performance and reduce storage costs.
[0114] The present invention can also provide a standardized data read / write interface through the unified data access service module, shielding the differences of underlying physical storage components and simplifying the system maintenance and extension process. This design can enhance the scalability and high availability of the system to meet the diverse needs in different business scenarios. At the same time, based on the binding mechanism of logical data partitioning and multi-dimensional indexes, the system can achieve refined management of data, providing more flexible and efficient support for subsequent data analysis and applications. In addition, through the multi-dimensional evaluation indexes of the data index management module, the system can dynamically adjust the data storage strategy to ensure efficient access to data and optimal configuration of storage resources.
[0115] Furthermore, the storage medium of the embodiment of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.
[0116] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present invention.
Claims
1. A storage architecture based on multi-dimensional metrics and logical data partitioning, characterized in that, It includes the following modules: Physical storage component management module, data partition management module, data metric management module, data migration management module, unified data access service module; The physical storage component management module is used to select and associate physical storage components according to the business scenario and data access characteristics; The data partition management module is used to define logical data partitions and their priority scores, and associate the logical data partitions with physical storage components; The data metric management module is used to define multi-dimensional data evaluation metrics, and bind logical data partitions according to the calculation results of the multi-dimensional data evaluation metrics; The data migration management module is used to periodically scan data and perform data migration based on the calculation results of the multi-dimensional data evaluation metrics; The unified data access service module is used to provide a unified data read-write interface externally, shielding the differences of physical storage components.
2. The storage architecture according to claim 1, wherein The specific implementation steps of the physical storage component management module include: Select the physical storage component type based on the business scenario and data access characteristics, and the physical storage component type includes relational databases, non-relational databases, or custom storage components; Write data to the physical storage component or read data from the physical storage component by inheriting and implementing the standard data interface of the physical storage component.
3. The storage architecture according to claim 1, wherein The specific implementation steps of the data partition management module include: Define multiple logical data partitions, and set a priority score for each logical data partition. The priority score ranges from 1 to 10, and the larger the score, the higher the priority; Associate each logical data partition with the corresponding physical storage component; Create a data temporary partition, which is used to cache the written data and is implemented by using a high-IO sequential read-write storage component.
4. The storage architecture according to claim 1, wherein The specific implementation steps of the data metric management module include: Define at least two multi-dimensional data evaluation metrics; Based on the calculation results of the multi-dimensional data evaluation metrics, bind the data to the corresponding logical data partition; When a data write request is initiated, calculate the results of the multi-dimensional data evaluation metrics in real time, and determine the final logical data partition to be written according to the priority score of the logical data partition.
5. The storage architecture according to claim 1, wherein The specific implementation steps of the data migration management module include: Periodically calculate the multi-dimensional metric results of the data through a scheduled scan task; Generate a migration task when the multi-dimensional metric results trigger a change in the logical data partition; Migrate the data from the physical storage component corresponding to the original logical data partition to the physical storage component corresponding to the target logical data partition through a two-phase transaction mechanism to ensure data consistency and integrity.
6. The storage architecture according to claim 5, characterized in that, The generation condition of the migration task is: When the priority score of the target logical data partition is higher than that of the current logical data partition, and the physical storage component associated with the target logical data partition is different from the physical storage component associated with the current logical data partition, data migration is triggered.
7. The storage architecture according to claim 1, characterized in that, The specific implementation steps of the unified data access service module include: Receive read-write requests from external clients and forward the read-write requests to the data scheduling center; The data scheduling center determines the logical data partition to which the data belongs based on the calculation results of the multi-dimensional metrics and returns the corresponding physical storage component information; The client performs data read and write operations according to the physical storage component information.
8. The storage architecture according to claim 3, wherein The usage steps of the data temporary partition include: When a data write request triggers the calculation of multi-dimensional metrics, the data is temporarily stored in the data temporary partition; After the metrics calculation is completed, the data is migrated from the data temporary partition to the physical storage component corresponding to the target logical data partition according to the priority score of the logical data partition.
9. The storage architecture according to claim 4, wherein The multi-dimensional data evaluation metrics include at least two of the following: Real-time metric, access frequency metric, data type metric, where the determination result of each metric is directly mapped to the corresponding logical data partition.
10. The storage architecture according to claim 1, wherein The dynamic adjustment rule of the priority score of the logical data partition is: If multiple logical data partitions all meet the metric binding conditions, the logical data partition with the highest priority score is selected as the final storage location; If the priority scores are the same, the logical data partition is selected according to the preset default rule.