A hot data identification and migration method for a hybrid storage architecture
Patent Information
- Application Number
- CN202611303635.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,现有混合存储热点数据识别与迁移方案仍存在诸多固有技术缺陷,严重制约混合存储系统的运行效率与稳定性,具体体现在以下多个方面:
通过构建融合四维特征的自适应热点评价体系并动态修正权重系数,有效提升了热点识别的准确率,并通过数据粒度自适应适配处理,消除了超大尺寸文件和碎片化小文件的识别盲区;
Smart Images

Figure CN122816554A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer storage technology, and in particular relates to a method for identifying and migrating hot data for hybrid storage architectures. Background Technology
[0002] With the rapid development of big data services, cloud storage services, and high-concurrency online services, single storage media can hardly simultaneously meet the demands for high IO read / write performance, large-capacity storage, and hardware cost control. Therefore, a multi-level heterogeneous hybrid storage architecture consisting of memory, NVMe SSDs, SATA SSDs, and HDDs has become the mainstream storage deployment solution. Hybrid storage effectively balances storage performance and cost by deploying frequently accessed hot data on high-speed media and storing infrequently accessed cold data on low-speed, high-capacity media. Accurate identification and intelligent migration scheduling of hot data are the core factors determining the overall IO performance, resource utilization, and hardware lifespan of a hybrid storage architecture.
[0003] However, existing hybrid storage hotspot data identification and migration solutions still have many inherent technical defects, which seriously restrict the operating efficiency and stability of hybrid storage systems, specifically in the following aspects: First, existing methods for collecting hotspot data are costly, highly intrusive, and poorly adaptable. Currently, the industry commonly uses a global, timed polling scan to collect statistical data access characteristics. This requires traversing the entire stored data and metadata, consuming significant CPU and I / O resources, easily interfering with normal business read / write performance, and failing to dynamically adapt sampling strategies based on cluster load. A fixed sampling frequency during peak load periods exacerbates system resource consumption, while sampling accuracy is insufficient during off-peak periods, making it difficult to achieve low-overhead, high-precision access characteristic collection. This results in lag and distortion of the corresponding basic identification data.
[0004] Second, hotspot identification suffers from limited dimensions, fixed weights, and a lack of ability to identify false hotspots. Traditional hotspot identification often relies solely on data access frequency as the sole criterion, failing to comprehensively consider key factors such as media characteristics, access time-series attenuation, and business priorities, resulting in a one-sided identification approach. Furthermore, the existing identification process cannot dynamically adjust parameters based on real-time storage media load, making it difficult to adapt to dynamically changing business scenarios.
[0005] Thirdly, the fixed granularity of migration identification leads to blind spots in hotspot identification. Existing technologies all use a uniform and fixed data block granularity for hot / cold file determination and migration scheduling. For ultra-large files, they can only determine the hot / cold status of the entire file and cannot identify local hotspot areas within the file. For massive amounts of fragmented small files, determining each block individually will generate a large amount of ineffective scheduling overhead, making it impossible to achieve aggregated identification of hotspots in small file clusters. This results in significant blind spots in identification and scheduling, making it difficult to adapt to diverse file storage formats.
[0006] Fourth, hotspot identification is delayed and lacks predictive and pre-scheduling capabilities. Traditional migration mechanisms are all passive response modes of access first, identification later, and migration then. They can only schedule existing hotspot data that has already generated high-frequency access, and cannot discover potential hotspot data with an upward trend in access. When new sudden business or periodic business is launched, potential hotspot data has not been pre-loaded into the high-speed medium, which will generate a lot of read and write latency spikes and cannot meet the operational requirements of high-concurrency, low-latency business.
[0007] Fifth, the migration strategies are homogenized, lacking media-specific adaptation, and resource scheduling is coarse. Existing hybrid storage migration solutions generally adopt uniform migration granularity, switching thresholds, and scheduling logic, failing to differentiate between the performance differences and wear characteristics of high-speed flash memory and low-speed mechanical storage. Furthermore, they lack media wear constraint mechanisms, frequently performing hot and cold switching on highly worn media such as SSDs, which easily leads to accelerated media wear and shortened lifespan, failing to balance system performance and hardware reliability.
[0008] Sixth: Frequent bidirectional switching between hot and cold data leads to serious migration oscillation problems. Existing technologies do not have a constraint mechanism for cold and hot migration residence. After data switches between hot and cold media, it is very easy for reverse repeated migration to occur due to short-term access fluctuations, causing frequent data jitter between media, generating a large amount of invalid migration IO overhead, and further aggravating system load and media wear.
[0009] In summary, existing hybrid storage hotspot identification and migration technologies suffer from numerous shortcomings, including low identification accuracy, high system overhead, rigid scheduling strategies, poor operational stability, high media loss, and weak adaptability. These limitations make it difficult to meet the intelligent, refined, and highly reliable operational requirements of current high-concurrency, multi-service, heterogeneous, multi-level hybrid storage clusters. Therefore, this invention aims to provide a hotspot data identification and migration method for hybrid storage architectures to address the aforementioned technical problems. Summary of the Invention
[0010] The purpose of this invention is to provide a method for identifying and migrating hot data in a hybrid storage architecture, so as to solve the technical problems existing in the prior art.
[0011] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A method for identifying and migrating hotspot data in hybrid storage architectures includes the following steps: S1: Employing a non-intrusive bypass monitoring approach, it streams and captures IO access requests from heterogeneous storage media at various levels in a hybrid storage architecture. It only collects and caches the access metadata corresponding to each data block, without reading complete business data. It dynamically adjusts the sampling window and sampling frequency based on the real-time load status of the hybrid storage cluster. During peak cluster load periods, it reduces the sampling frequency to reduce noise and load, and during off-peak cluster load periods, it improves sampling accuracy to obtain the real-time access feature dataset for each data block.
[0012] The output real-time access feature dataset provides complete, low-noise, and real-time basic data input for subsequent four-dimensional hotspot evaluation system scoring, feature weight correction, and pseudo-hotspot filtering. It serves as the data foundation for subsequent full-process hotspot identification and migration scheduling.
[0013] S2: Based on the real-time access feature dataset, construct a four-dimensional hotspot evaluation system that integrates static media features, dynamic access features, time-series decay features, and business priority features, and perform normalized weighted scoring on each data block; dynamically and adaptively adjust the weight coefficients corresponding to each feature in combination with the real-time load status of each level of storage media, and abandon fixed weight configuration; at the same time, filter invalid access data such as instantaneous batch reads, short-term pulse accesses, and one-time crawler accesses, identify and eliminate pseudo-hotspot data, and output the real-time hot / cold level corresponding to each data block.
[0014] The real-time hot and cold rating provides a precise quantitative basis for determining hot and cold conditions for subsequent granular adaptive calibration, potential hot spot prediction, media-differentiated migration, and migration priority scheduling, which can avoid ineffective scheduling and erroneous migration.
[0015] S3: Based on the file size and storage format of each data block, perform dynamic and uniform block processing on ultra-large files, and aggregate and merge fragmented small files. Based on the smallest hot spot unit after adaptive adaptation, match the real-time hot and cold levels to complete the identification of local hot spots and cluster hot spots.
[0016] Precise identification of local hotspots and cluster hotspots eliminates the blind spots of traditional fixed-granularity hot and cold identification, providing the smallest precise scheduling unit for subsequent refined time-series trend prediction, differential granularity migration of heterogeneous media, and oscillation suppression.
[0017] S4: Construct a sliding fitting window based on the historical access time series data of each data block, extract access cycle features and access trend features, fit the data access change pattern, predict potential hotspot data that is about to heat up, and add pre-migration tags to the potential hotspot data.
[0018] Potential hotspot data with pre-migration tags enables a shift from passive migration to proactive pre-scheduling, providing a priority migration basis for subsequent bidirectional migration scheduling and eliminating business latency spikes caused by hotspot switching.
[0019] S5: Pre-configure the capacity high-water level threshold, capacity low-water level threshold, and media wear threshold for each level of heterogeneous media in the hybrid storage architecture, including memory, NVMeSSD, SATASSD, and mechanical hard drive. Combine the media performance characteristics to match differentiated identification thresholds, migration granularity, and scheduling strategies. Perform bidirectional closed-loop migration based on the real-time hot and cold levels, pre-migration markers, and media water level status of each data block.
[0020] Bidirectional hot and cold data migration enables dynamic and balanced resource allocation across multi-level heterogeneous storage media, providing real-world migration scenario sample data for subsequent bandwidth limiting, oscillation control, and full-process parameter iteration.
[0021] S6: During the data migration process, the remaining available bandwidth of the hybrid storage cluster is detected in real time, and the migration IO rate is dynamically and adaptively adjusted to avoid migration traffic from competing for the read and write bandwidth of upper-layer services; at the same time, the migration tasks are sorted according to the service priority tags, and the data migration tasks corresponding to high-priority services are scheduled first.
[0022] Adaptive bandwidth limiting and priority scheduling results ensure stable service I / O during the migration process, providing a low-interference and highly reliable operating condition foundation for subsequent migration oscillation suppression, media wear control, and strategy iteration optimization.
[0023] S7: For data blocks that have completed the cold / hot media migration switch, configure a dynamic residence cooling time window. During the cooling window period, repeated reverse cold / hot migration operations for the same data block are prohibited, suppressing frequent jitter and oscillation of data between multiple storage media. This stabilizes the cold / hot residence state, eliminates the invalid IO overhead of frequent data migration, and provides stable, low-noise, and effective training samples for subsequent full-process self-learning iterative optimization, ensuring reliable parameter iteration convergence.
[0024] S8: Collect media load data, service read / write latency data, hot / cold switching frequency data, and media wear data after each data migration is completed. Based on the collected operation and maintenance data, iteratively correct the feature weights, media water level thresholds, residence cooling window parameters, and migration current limiting parameters of the hot spot evaluation system to achieve adaptive iterative upgrades of system strategies.
[0025] Preferably, the specific process of streaming I / O access requests in step S1 is as follows: Deploy a bypass listening agent module on the IO data path of the hybrid storage cluster to capture IO request packets passing through each level of storage media in real time through mirrored ports or drop-capture (bypass mirroring interception) mechanism.
[0026] The drop-capture mechanism is a non-intrusive message copying mechanism for the IO path: it generates a mirror copy of the business message flowing through the IO path, sends the copy to the listening agent module to parse and extract metadata; the original business message continues to be transmitted along the original path, will not be dropped or modified, and will not interfere with the normal IO business execution of the upper layer.
[0027] The captured IO request messages are parsed to extract five core access metadata parameters: data block identifier, access operation type, access offset, access data length, and access timestamp. The extracted access metadata parameters are written to an in-memory cache queue, and then cleaned, standardized, and deduplicated through a streaming pipeline to generate structured access metadata records.
[0028] The specific process of dynamically adjusting the sampling window and sampling frequency in step S1 is as follows: Based on the real-time load monitoring metrics of the hybrid storage cluster, including CPU utilization, IO queue depth, and network bandwidth utilization, the sampling adjustment coefficient is dynamically calculated. When the CPU utilization exceeds the upper limit threshold of 85% or the IO queue depth exceeds the threshold, the sampling adjustment coefficient is lowered to the first interval [0.3, 0.5] to achieve sampling frequency reduction; When the CPU utilization is detected to be below the lower limit threshold of 30% and the IO queue depth is below the threshold, the sampling adjustment coefficient is increased to the second interval [0.8, 1.0] to improve the sampling accuracy.
[0029] Preferably, the specific process of constructing the four-dimensional hotspot evaluation system in step S2 is as follows: S21: Feature Dimension Definition: Define the feature dimensions of static media. Characterizes the type and performance level of the storage medium where the data block is currently located, and defines the dynamic access feature dimension. Characterize the access frequency, access throughput, and IO concurrency of data blocks within a statistical period, and define the time-series decay feature dimension. Characterize the time-decrease properties of data block access behavior and define business priority feature dimensions. The priority label weight represents the business to which the data block belongs; S22: Normalization: Normalize the original values of each feature dimension and map them to... For the interval, a modified Min-Max standardization method is used to eliminate the impact of dimensional differences on the comprehensive score; S23: Adaptive Weight Calculation: The adaptive weight coefficients for each feature dimension are dynamically calculated based on the real-time load status of each level of storage media. The weight calculation formula is as follows: ; in, For the first Real-time adaptive weight coefficients for each feature dimension; , feature dimensions , The basic weights, , feature dimensions , The extent to which the medium load deviates from the equilibrium state. , For feature dimensions , The load sensitivity coefficient is a preset constant used to characterize the sensitivity of the feature weights to changes in the medium load. The larger the value, the more obvious the weight fluctuation with the load. These correspond to four feature dimensions respectively; S24: Pseudo-hotspot filtering rule configuration: Configure instantaneous batch read identification rules, the judgment criterion is that the data volume of a single access request exceeds... Furthermore, the standard deviation of the access interval within the time window is less than the threshold; configure short-time pulse access identification rules, with the judgment criteria being that the access frequency exhibits a pulse distribution within the time window and the attenuation rate exceeds the attenuation rate threshold. Configure a one-time crawler access identification rule, the judgment criterion is that the same client accesses the same data block only once and there is no subsequent access behavior; S25: Comprehensive Score Calculation: Calculate the comprehensive hotspot score for each data block based on normalized eigenvalues and adaptive weighting coefficients. : ; in, For the first The normalized eigenvalues of each feature dimension are obtained. All eigenvalues are standardized using Min-Max and take values in the range [0,1].
[0030] The hot / cold category is determined based on a comprehensive score; a score greater than or equal to... This data was identified as trending, and the score was... to The interval is determined to be warm data, and the score is less than This data is classified as cold data.
[0031] Preferably, the specific process of data granularity adaptive adaptation in step S3 is as follows: S31: File Size Detection and Classification: Extract file metadata corresponding to data blocks and obtain total file size information; classify data blocks into three categories according to file size: oversized files, regular-sized files, and fragmented small files. The criteria for determining oversized files is that the file size is greater than or equal to 2GB, and the criteria for determining fragmented small files is that the file size is less than 64KB. S32: Dynamic block processing for ultra-large files: For ultra-large files, the file is divided into multiple data sub-blocks with a uniform block granularity of 256MB to 512MB. Each data sub-block is independently allocated a block-level access statistics counter, and the access frequency of each sub-block is collected separately to realize the identification of local hot spots inside the file. S33: Fragmented small file aggregation and merging processing: For fragmented small files, multiple small files are aggregated into logical aggregation units according to storage path or access time window. Access heat statistics are established for the entire aggregation unit. Hot spot determination and migration scheduling are performed at the aggregation unit granularity to avoid the huge IO overhead caused by judging a large number of small files block by block. S34: Hotspot calibration and mapping: Match and calibrate the smallest hotspot unit after granularity adaptation with the real-time hot and cold levels output in step S25 to generate local hotspot calibration results and cluster hotspot calibration results.
[0032] Preferably, the specific process of predicting potential hotspots in step S4 is as follows: S41: Sliding window construction: Starting from the current moment, backtracking along the historical timeline. The duration is used to construct a sliding fitting window, where The value is adaptively determined based on the access cycle characteristics of the data block, with a typical value being... or .
[0033] S42: Access cycle feature extraction: Perform spectral analysis on the access time series data within the sliding window, use Fast Fourier Transform to extract the main cycle frequency of the access behavior, and generate an access cycle fingerprint.
[0034] S43: Access Trend Feature Fitting: The access frequency time series data is fitted with an exponentially weighted moving average algorithm to calculate the access slope. : ; in, The access slope is the slope of the time sequence change in access frequency. A value greater than 0 indicates that the access popularity is on the rise, while a value less than or equal to 0 indicates that the popularity is stable or declining. For time decay weight, For a moment Access frequency, This represents the number of data points within the window. This represents the average access frequency across all times within the sliding window. This represents the average of the time node numbers within the sliding window.
[0035] S44: Potential Hotspot Prediction Rules: When the access slope Greater than the preset trend threshold Furthermore, when the current hot spot score is lower than the hot spot determination threshold, the data block is determined to be potential hot spot data and a pre-migration mark is automatically added; potential hot spot data enters the pre-migration priority queue, waiting to be migrated to high-speed storage media.
[0036] Preferably, the specific process of heterogeneous media differentiated bidirectional migration scheduling in step S5 is as follows: S51: Media Parameter Pre-configuration: Configure a high-water threshold for memory media capacity. for Low water level threshold for Media wear threshold for Configure a high-water mark for NVMe SSD capacity. for Low water level threshold for Media wear threshold for Configure a high-water mark for the capacity of SATA SSDs. for Low water level threshold for Media wear threshold for Configure high watermark capacity for mechanical hard drives. for Low water level threshold for Media wear threshold for ; S52: Media Differentiation Strategy Matching: Identification thresholds and migration granularity are configured differently based on media type. The hotspot identification threshold for memory media is... Migration granularity is to The hotspot identification threshold for NVMe SSDs is Migration granularity is to The hotspot identification threshold for SATA SSDs is Migration granularity is to The hotspot detection threshold for mechanical hard drives is Migration granularity is to ; S53: Bidirectional migration trigger determination: Real-time monitoring of the capacity utilization rate of storage media at all levels. When the capacity of high-speed storage media reaches the corresponding high water level threshold, the migration scheduling process is triggered; when the capacity of high-speed storage media falls back to the corresponding low water level threshold, warm and cold data migration is stopped to maintain the steady state of storage resources. S54: Migration Execution Logic: Prioritize migrations with higher overall scores. The thermal data is transferred to the high-speed medium, while the overall score is lower than that of the medium. Cold data is migrated out of high-speed media; for potential hot data carrying pre-migration tags, migration resources are prioritized to ensure that potential hot data is installed in high-speed media in advance.
[0037] Preferably, the specific process of bandwidth awareness and service priority coordinated rate limiting migration in step S6 is as follows: S61: Real-time Remaining Bandwidth Detection: The system network interface monitoring module collects the network interface throughput, IO controller bandwidth, and storage medium IO bandwidth of the hybrid storage cluster in real time, and calculates the current remaining available bandwidth. : ; in Total bandwidth As of the current bandwidth utilization, Reserve bandwidth for business read and write operations; S62: Dynamic migration rate adjustment: Dynamically adjusts the migration I / O rate based on the remaining available bandwidth. The calculation formula is: ; in The upper limit of the migration rate set for the system. This is the bandwidth allocation coefficient, and its value range is... to ; S63: Business Priority Tag Extraction: Extract business priority tags from the metadata of data blocks. Priority tags are divided into four levels: urgent, high, normal, and low. S64: Migration Task Hierarchical Sorting: Sort the migration task queue in descending order according to business priority tags. Urgent migration tasks are scheduled and executed first. Tasks within the same priority are sorted by hot / cold level and pre-migration tag.
[0038] Preferably, the specific process of cold and hot migration oscillation suppression and stabilization treatment in step S7 is as follows: S71: Dynamic Calculation of Cooldown Time Window: Dynamically calculates the duration of the cooldown time window based on the historical hot / cold switching frequency of the data block. : ; in The base cooldown time is set to a value of , The frequency sensitivity coefficient has a value of [value missing]. , For history Number of times the temperature changes within the room; S72: Cooling window locking mechanism: The data block that completes the switching of hot and cold media locks its current hot and cold state within the cooling time window, and prohibits repeated reverse migration operations; S73: Cooling window end determination: When the cooling time window expires, the state lock is released, allowing the data block to re-determine its hot / cold status based on real-time access characteristics and participate in subsequent migration scheduling.
[0039] Preferably, the specific process of the full-process closed-loop self-learning iterative optimization in step S8 is as follows: S81: Operation and maintenance data collection: After each data migration is completed, automatically collect and record media load change data, service read and write latency change data, cold and hot switching frequency data, and media wear increment data to form a migration execution log; S82: Strategy Deviation Assessment: Compare and analyze the actual effect data after migration execution with the expected strategy objectives, and calculate the actual deviation values of each strategy parameter; S83: Parameter Iterative Correction: Based on the strategy deviation value, the feedback adjustment algorithm is used to iteratively correct the characteristic weight coefficients, medium water level threshold parameters, residence cooling window parameters, and migration flow restriction parameters of the hot spot evaluation system; S84: Policy Version Release: Generates the revised policy parameters into a new version policy configuration and dynamically distributes it to each node of the hybrid storage cluster, realizing online hot updates of the policy and completing the entire closed-loop self-learning iterative optimization process.
[0040] The beneficial effects of this invention include: By using a non-intrusive bypass monitoring method, we can achieve streaming dynamic collection of IO access metadata and obtain accurate access characteristic data while ensuring normal business read and write performance. By constructing an adaptive hotspot evaluation system that integrates four-dimensional features and dynamically adjusting the weight coefficients, the accuracy of hotspot identification is effectively improved. Furthermore, through adaptive data granularity processing, the blind spots in the identification of ultra-large files and fragmented small files are eliminated. By predicting potential hotspots through time-series trends, the pre-migration of hotspot data was achieved. Through differentiated bidirectional migration scheduling of heterogeneous media, both storage performance and media lifespan were taken into account. By using bandwidth awareness and service priority-based rate limiting, we prevented migration traffic from monopolizing service bandwidth; and by using a cold / hot migration oscillation suppression mechanism, we reduced ineffective migration overhead. Through a closed-loop self-learning iteration throughout the entire process, adaptive optimization of the strategy is achieved. The overall solution significantly improves the operating efficiency, resource utilization, and hardware lifespan of the hybrid storage system. Attached Figure Description
[0041] Figure 1 This is a flowchart illustrating the hotspot data identification and migration method for hybrid storage architectures according to the present invention. Figure 2 This is a schematic diagram of the logical structure of the four-dimensional hotspot evaluation system of the present invention. Detailed Implementation
[0042] The following is in conjunction with the appendix Figures 1-2 The present invention will be further described in detail below: The system architecture of this invention mainly includes a hybrid storage cluster, a bypass monitoring proxy module, an access metadata cache queue, a four-dimensional hotspot evaluation engine, a granular adaptive processing module, a time-series trend prediction module, a heterogeneous media scheduling controller, a bandwidth-aware rate limiting module, an oscillation suppression controller, and a self-learning iterative optimization engine. The hybrid storage cluster consists of four levels of heterogeneous storage media: memory, NVMe SSDs, SATA SSDs, and hard disk drives, interconnected via an IO data path. The bypass monitoring proxy module is deployed on the IO data path, capturing IO request packets passing through each level of storage media in real time via mirroring ports or a drop-capture mechanism.
[0043] The access metadata cache queue receives and caches access metadata parameters. The four-dimensional hotspot evaluation engine calculates the popularity score of each data block based on the cached access metadata. The granular adaptive processing module performs block or aggregation processing on files of different sizes. The time-series trend prediction module mines potential hotspot data. The heterogeneous media scheduling controller performs bidirectional migration scheduling. The bandwidth-aware rate limiting module dynamically adjusts the migration rate. The oscillation suppression controller locks data blocks within the cooling window. The self-learning iterative optimization engine iteratively corrects strategy parameters based on operation and maintenance data. See appendix. Figure 1 As shown, the hotspot data identification and migration method for hybrid storage architectures includes the following steps: S1: Employing a non-intrusive bypass monitoring approach, it streams and captures IO access requests from heterogeneous storage media at various levels in a hybrid storage architecture. It only collects and caches the access metadata corresponding to each data block, without reading complete business data. It dynamically adjusts the sampling window and sampling frequency based on the real-time load status of the hybrid storage cluster. During peak cluster load periods, it reduces the sampling frequency to reduce noise and load, and during off-peak cluster load periods, it improves sampling accuracy to obtain the real-time access feature dataset for each data block.
[0044] S2: Based on the real-time access feature dataset, construct a four-dimensional hotspot evaluation system that integrates static media features, dynamic access features, time-series decay features, and business priority features, and perform normalized weighted scoring on each data block; dynamically and adaptively adjust the weight coefficients corresponding to each feature in combination with the real-time load status of each level of storage media, and abandon fixed weight configuration; at the same time, filter invalid access data such as instantaneous batch reads, short-term pulse accesses, and one-time crawler accesses, identify and eliminate pseudo-hotspot data, and output the real-time hot / cold level corresponding to each data block.
[0045] S3: Based on the file size and storage format of each data block, perform dynamic and uniform block processing on ultra-large files, and aggregate and merge fragmented small files. Based on the smallest hot spot unit after adaptive adaptation, match the real-time hot and cold levels to complete the identification of local hot spots and cluster hot spots.
[0046] S4: Construct a sliding fitting window based on the historical access time series data of each data block, extract access cycle features and access trend features, fit the data access change pattern, predict potential hotspot data that is about to heat up, and add pre-migration tags to the potential hotspot data.
[0047] S5: Pre-configure the capacity high-water level threshold, capacity low-water level threshold, and media wear threshold for each level of heterogeneous media in the hybrid storage architecture, including memory, NVMeSSD, SATASSD, and mechanical hard drive. Combine the media performance characteristics to match differentiated identification thresholds, migration granularity, and scheduling strategies. Perform bidirectional closed-loop migration based on the real-time hot and cold levels, pre-migration markers, and media water level status of each data block.
[0048] S6: During the data migration process, the remaining available bandwidth of the hybrid storage cluster is detected in real time, and the migration IO rate is dynamically and adaptively adjusted to avoid migration traffic from competing for the read and write bandwidth of upper-layer services; at the same time, the migration tasks are sorted according to the service priority tags, and the data migration tasks corresponding to high-priority services are scheduled first.
[0049] S7: Configure a dynamic residence cooling time window for data blocks that have completed the cold and hot media migration and switching. During the cooling window period, repeated reverse cold and hot migration operations of the same data block are prohibited, suppressing frequent jitter and oscillation of data between multiple storage media.
[0050] S8: Collect media load data, service read / write latency data, hot / cold switching frequency data, and media wear data after each data migration is completed. Based on the collected operation and maintenance data, iteratively correct the feature weights, media water level thresholds, residence cooling window parameters, and migration current limiting parameters of the hot spot evaluation system to achieve adaptive iterative upgrades of system strategies.
[0051] The specific process of streaming I / O access requests in step S1 is as follows: A bypass listening agent module is deployed on the IO data path of the hybrid storage cluster to capture IO request packets passing through each level of storage media in real time via mirrored ports or drop-capture mechanism. The captured IO request packets are parsed to extract five core access metadata parameters: data block identifier, access operation type, access offset, access data length, and access timestamp. These extracted access metadata parameters are written to an in-memory cache queue and cleaned, standardized, and deduplicated through a streaming pipeline to generate structured access metadata records. Based on real-time load monitoring metrics of the hybrid storage cluster, including CPU utilization, IO queue depth, and network bandwidth utilization, a sampling adjustment coefficient is dynamically calculated. When CPU utilization exceeds 85% or IO queue depth exceeds a threshold, the sampling adjustment coefficient is lowered to the 0.3-0.5 range to reduce sampling frequency. When CPU utilization is below 30% and IO queue depth is below a threshold, the sampling adjustment coefficient is increased to the 0.8-1.0 range to improve sampling accuracy.
[0052] The specific process of constructing the four-dimensional hotspot evaluation system in step S2 is as follows: S21: Feature Dimension Definition: Define the feature dimensions of static media. Characterizes the type and performance level of the storage medium where the data block is currently located, and defines the dynamic access feature dimension. Characterize the access frequency, access throughput, and IO concurrency of data blocks within a statistical period, and define the time-series decay feature dimension. Characterize the time-decrease properties of data block access behavior and define business priority feature dimensions. The priority label weight represents the business to which the data block belongs; S22: Normalization: Normalize the original values of each feature dimension and map them to... For the interval, a modified Min-Max standardization method is used to eliminate the impact of dimensional differences on the comprehensive score; S23: Adaptive Weight Calculation: The adaptive weight coefficients for each feature dimension are dynamically calculated based on the real-time load status of each level of storage media. The weight calculation formula is as follows: in, For the first Real-time adaptive weight coefficients for each feature dimension; , feature dimensions , The basic weights, , feature dimensions , The extent to which the medium load deviates from the equilibrium state. , For feature dimensions , The load sensitivity coefficient can be pre-configured according to the media type. For example, it is 0.2 for memory media, 0.15 for NVMe SSD, 0.1 for SATA SSD, and 0.05 for mechanical hard drive. It is used to characterize the sensitivity of feature weight to changes in media load. The larger the value, the more obvious the weight fluctuation with load. These correspond to four feature dimensions.
[0053] ; in, For the current load of the medium, To balance the baseline load of the system, This is the upper limit of the media load.
[0054] S24: Pseudo-hotspot filtering rule configuration: Configure instantaneous batch read identification rules, the judgment criterion is that the data volume of a single access request exceeds... P th1 And the standard deviation of the access interval within the time window is less than the threshold. σ th Configure short-time pulse access identification rules, with the criteria being that the access frequency exhibits a pulse distribution within the time window and the attenuation rate exceeds [a certain threshold]. δ th Configure a one-time crawler access identification rule, the judgment criterion is that the same client accesses the same data block only once and there is no subsequent access behavior.
[0055] P th1 Dynamically calibrated based on 120% of the maximum single IOPS throughput of the current storage medium; σ th =0.3; attenuation rate The calculation formula is ,in, The maximum number of I / O accesses to the data within a single period represents the highest popularity level of the data. The real-time access frequency of the data block at the end of the statistical period represents the actual popularity status at the end of the data period. δ th The value is fixed at 0.7. Pulse distribution identification uses a sliding window local variance detection method; a pulse is considered valid when the variance suddenly increases by more than twice the baseline value and lasts for less than 50ms. All thresholds are initialized using offline historical log training sets via quantile regression and adaptively fluctuate by ±10% during online operation.
[0056] S25: Comprehensive Score Calculation: Calculate the comprehensive hotspot score for each data block based on normalized eigenvalues and adaptive weighting coefficients. : in, For the first The normalized eigenvalues of each feature dimension are obtained. All eigenvalues are standardized using Min-Max and take values in the range [0,1].
[0057] The hot / cold category is determined based on the overall score, with scores greater than or equal to the upper limit threshold. Data identified as trending, with ratings below the lower limit threshold. Up to the upper limit of the rating threshold The interval is determined to be warm data, and the score is less than the lower limit threshold. This data is classified as cold data.
[0058] The specific process of data granularity adaptive adaptation in step S3 is as follows: S31: File Size Detection and Classification: Extract file metadata corresponding to data blocks to obtain total file size information; classify data blocks into three categories according to file size: oversized files, regular-sized files, and fragmented small files. The criterion for determining oversized files is that the file size is greater than or equal to... The criterion for determining fragmented small files is that the file size is less than [a certain value]. ; S32: Dynamic block processing for ultra-large files: For ultra-large files, according to... to The uniform block granularity divides the file into multiple data sub-blocks, allocates a block-level access statistics counter to each data sub-block independently, and collects the access popularity of each sub-block separately to achieve local hot spot identification within the file. S33: Fragmented small file aggregation and merging processing: For fragmented small files, multiple small files are aggregated into logical aggregation units according to storage path or access time window. Access heat statistics are established for the entire aggregation unit. Hot spot determination and migration scheduling are performed at the aggregation unit granularity to avoid the huge IO overhead caused by judging a large number of small files block by block. S34: Hotspot calibration and mapping: Match and calibrate the smallest hotspot unit after granularity adaptation with the real-time hot and cold levels output in step S25 to generate local hotspot calibration results and cluster hotspot calibration results.
[0059] The specific process of predicting potential hotspots based on time-series trends in step S4 is as follows: S41: Sliding window construction: Starting from the current moment, backtracking along the historical timeline. The duration is used to construct a sliding fitting window, where The value is adaptively determined based on the access cycle characteristics of the data block, with a typical value being... or ; S42: Access cycle feature extraction: Perform spectral analysis on the access time series data within the sliding window, use Fast Fourier Transform to extract the main cycle frequency of the access behavior, and generate an access cycle fingerprint; S43: Access Trend Feature Fitting: The access frequency time series data is fitted with an exponentially weighted moving average algorithm to calculate the access slope. : in, The access slope is the slope of the time sequence change in access frequency. A value greater than 0 indicates that the access popularity is on the rise, while a value less than or equal to 0 indicates that the popularity is stable or declining. For time decay weight, For a moment Access frequency, This represents the number of data points within the window. This represents the average access frequency across all times within the sliding window. This represents the average of the time node numbers within the sliding window.
[0060] in For the preset attenuation coefficient, For the current moment, Historical visit times; the more recent the time, the higher the weight, highlighting the dominance of recent visit trends. S44: Potential Hotspot Prediction Rules: When the access slope Greater than the preset trend threshold Furthermore, when the current hot spot score is lower than the hot spot determination threshold, the data block is determined to be potential hot spot data and a pre-migration mark is automatically added; potential hot spot data enters the pre-migration priority queue, waiting to be migrated to high-speed storage media.
[0061] The specific process of heterogeneous media differentiated bidirectional migration scheduling in step S5 is as follows: S51: Media Parameter Pre-configuration: Configure a high-water threshold for memory media capacity. for Low water level threshold for Media wear threshold for Configure a high-water mark for NVMe SSD capacity. for Low water level threshold for Media wear threshold for Configure a high-water mark for the capacity of SATA SSDs. for Low water level threshold for Media wear threshold for Configure high watermark capacity for mechanical hard drives. for Low water level threshold for Media wear threshold for ; S52: Media Differentiation Strategy Matching: Identification thresholds and migration granularity are configured differently based on media type. The hotspot identification threshold for memory media is... Migration granularity is to The hotspot identification threshold for NVMe SSDs is Migration granularity is to The hotspot identification threshold for SATA SSDs is Migration granularity is to The hotspot detection threshold for mechanical hard drives is Migration granularity is to ; S53: Bidirectional migration trigger determination: Real-time monitoring of the capacity utilization rate of storage media at all levels. When the capacity of high-speed storage media reaches the corresponding high water level threshold, the migration scheduling process is triggered; when the capacity of high-speed storage media falls back to the corresponding low water level threshold, warm and cold data migration is stopped to maintain the steady state of storage resources. S54: Migration Execution Logic: Prioritize migrations with higher overall scores. The thermal data is transferred to the high-speed medium, while the overall score is lower than that of the medium. Cold data is migrated out of high-speed media; for potential hot data carrying pre-migration tags, migration resources are prioritized to ensure that potential hot data is installed in high-speed media in advance.
[0062] The specific process of bandwidth awareness and service priority-based coordinated rate limiting migration in step S6 is as follows: S61: Real-time Remaining Bandwidth Detection: The system network interface monitoring module collects the network interface throughput, IO controller bandwidth, and storage medium IO bandwidth of the hybrid storage cluster in real time, and calculates the current remaining available bandwidth. : in Total bandwidth As of the current bandwidth utilization, Reserve bandwidth for business read and write operations; S62: Dynamic migration rate adjustment: Dynamically adjusts the migration I / O rate based on the remaining available bandwidth. The calculation formula is: in The upper limit of the migration rate set for the system. This is the bandwidth allocation coefficient, and its value range is... to ; S63: Business Priority Tag Extraction: Extract business priority tags from the metadata of data blocks. Priority tags are divided into four levels: urgent, high, normal, and low. S64: Migration Task Hierarchical Sorting: Sort the migration task queue in descending order according to business priority tags. Urgent migration tasks are scheduled and executed first. Tasks within the same priority are sorted by hot / cold level and pre-migration tag.
[0063] The specific process of cold and hot migration oscillation suppression and stabilization in step S7 is as follows: S71: Dynamic Calculation of Cooldown Time Window: Dynamically calculates the duration of the cooldown time window based on the historical hot / cold switching frequency of the data block. : in The base cooldown time is set to a value of , The frequency sensitivity coefficient has a value of [value missing]. , For history Number of times the temperature changes within the room; S72: Cooling window locking mechanism: The data block that completes the switching of hot and cold media locks its current hot and cold state within the cooling time window, and prohibits repeated reverse migration operations; S73: Cooling window end determination: When the cooling time window expires, the state lock is released, allowing the data block to re-determine its hot / cold status based on real-time access characteristics and participate in subsequent migration scheduling.
[0064] The specific process of the full-process closed-loop self-learning iterative optimization in step S8 is as follows: S81: Operation and maintenance data collection: After each data migration is completed, automatically collect and record media load change data, service read and write latency change data, cold and hot switching frequency data, and media wear increment data to form a migration execution log; S82: Strategy Deviation Assessment: Compare and analyze the actual effect data after migration execution with the expected strategy objectives, and calculate the actual deviation values of each strategy parameter; S83: Parameter Iterative Correction: Based on the strategy deviation value, the feedback adjustment algorithm is used to iteratively correct the characteristic weight coefficients, medium water level threshold parameters, residence cooling window parameters, and migration flow restriction parameters of the hot spot evaluation system; S84: Policy Version Release: Generates the revised policy parameters into a new version policy configuration and dynamically distributes it to each node of the hybrid storage cluster, realizing online hot updates of the policy and completing the entire closed-loop self-learning iterative optimization process.
[0065] In another embodiment, the streaming capture of IO access requests in step S1 further includes a load adaptive sampling adjustment mechanism. This mechanism monitors the operating status of the hybrid storage cluster in real time and calculates the overall load index of the current cluster using a preset load assessment model. The overall load index is calculated by combining the weighted values of CPU utilization, IO queue depth, and network bandwidth utilization. When the overall load index exceeds a first preset threshold, the cluster is determined to be in a peak load state. At this time, a downsampling mode is activated, setting the sampling adjustment coefficient to a lower value to reduce the number of samples per unit time and reduce system overhead. When the overall load index is below a second preset threshold, the cluster is determined to be in a low load state. At this time, a fine sampling mode is activated, setting the sampling adjustment coefficient to a higher value to increase the sampling frequency and data accuracy.
[0066] In another embodiment, the four-dimensional hotspot evaluation system construction process in step S2 also includes a dynamic feature dimension expansion mechanism. This mechanism dynamically expands the feature dimensions based on the media type composition and business characteristics of the hybrid storage cluster. For example, when persistent memory media is detected in the hybrid storage cluster, an additional persistent memory feature dimension is introduced. This characterizes the degree to which data blocks adapt to persistent memory characteristics; when a clear spatiotemporal distribution feature of access frequency is detected in a business scenario, an additional spatiotemporal distribution feature dimension is introduced. This characterizes the distribution pattern of data block access behavior in the temporal and spatial dimensions. The dynamically expanded feature dimension set is reduced in dimensionality using principal component analysis to eliminate collinearity among features and improve the robustness of the comprehensive score.
[0067] In another embodiment, the data granularity adaptive adaptation process in step S3 also includes a hybrid granularity collaborative identification mechanism. This mechanism targets large mixed files containing both hot and cold regions, comprehensively utilizing both file-level heat assessment and block-level heat identification strategies. At the file level, the overall file access heat distribution is statistically analyzed; at the block level, the file is subdivided according to a preset block granularity, and the access heat of each sub-block is collected individually. The system integrates the file-level and block-level assessment results. For cases where the file-level data is determined to be warm data but the block level shows obvious hot regions, a local migration process is automatically triggered, migrating only the hot sub-blocks to the high-speed medium while retaining the cold sub-blocks in the original medium, thus achieving refined heat layering management.
[0068] In another embodiment, the time-series trend prediction and potential hotspot mining process in step S4 also includes a multi-period collaborative prediction mechanism. This mechanism simultaneously constructs daily, weekly, and monthly sliding windows to extract access cycle features and access trend features at different time scales. The multi-period collaborative prediction algorithm integrates the trend analysis results of each period window to calculate the comprehensive trend score of the data block. When the comprehensive trend score exceeds a preset threshold, even if the current hotspot score does not reach the hotspot determination threshold, the data block is marked as potential hotspot data and included in the pre-migration priority queue. The formula for calculating the comprehensive trend score is: in , , The access slopes are calculated for daily, weekly, and monthly timeframes, respectively. , , For the corresponding periodic weight coefficients, satisfying .
[0069] In another embodiment, the heterogeneous media differentiated bidirectional migration scheduling process in step S5 also includes a media health awareness mechanism. This mechanism monitors the health status parameters of storage media at all levels in real time, including the remaining write cycles of NVMeSSDs, the bad block growth rate of SATASSDs, and the number of reallocated sectors of mechanical hard drives. When the health index of a certain medium is detected to be lower than a preset threshold, the migration priority of that medium is automatically reduced, while the migration priority is increased, accelerating the migration of hot data from media with lower health. The media health score calculation formula is as follows: in This represents the current number of erase / write cycles. This is the rated maximum number of erase / write cycles. and These are the bad block factor and the reallocation sector factor, respectively. Factor value rules: : 1 for no bad blocks, the value decreases linearly to 0 as the number of bad blocks increases; : No remapping sectors are 1, and the value decreases linearly to 0 as the number of remapping sectors increases; health value is [0,1].
[0070] In another embodiment, the bandwidth awareness and service priority-based rate limiting migration process in step S6 also includes a migration task preemption mechanism. This mechanism automatically triggers a rate reduction or pause of the migration task when it detects a sharp increase in service bandwidth demand. When the remaining available bandwidth is lower than the service's reserved bandwidth... When a sudden surge in I / O requests is detected, the migration rate will be forcibly reduced to the minimum available rate. All non-urgent migration tasks will be paused until the business load subsides, at which point migration will resume. A gradual acceleration strategy will be employed when resuming migration tasks to avoid sudden bandwidth fluctuations causing instantaneous impacts on the business.
[0071] In another embodiment, the cold / hot migration oscillation suppression and stabilization process in step S7 further includes a dynamic adjustment mechanism for the cooling window. This mechanism dynamically adjusts the duration of the cooling time window based on the historical migration behavior pattern of the data block. For data blocks that frequently switch between cold and hot modes, the cooling time window is automatically extended, increasing the data block's residence time in the current medium and reducing the probability of oscillations. For stable data blocks, i.e., data blocks with low historical switching frequency, a shorter cooling time window is used to ensure the real-time performance of hotspot identification. The dynamic adjustment coefficient of the cooling time window is determined based on the ratio of the historical switching frequency of the data block to the global average switching frequency.
[0072] In another implementation, the full-process closed-loop self-learning iterative optimization process in step S8 also includes a strategy effectiveness verification mechanism. This mechanism does not immediately apply the new strategy to the production environment after iterative correction of the strategy parameters. Instead, it first verifies the strategy's effectiveness in an isolated simulation environment. The verification process is based on historical operational datasets, simulating the migration scheduling of the new strategy and evaluating key indicators such as the accuracy of hotspot identification, migration efficiency, and media load balancing. Only when the verification indicators meet preset standards will the new strategy be approved for release to the production environment. This strategy effectiveness verification mechanism effectively avoids systemic risks caused by deviations in strategy parameters.
[0073] The hotspot data identification and migration method of this invention further includes step S9: distributed cluster global load balancing scheduling. This step, based on steps S5 to S8, introduces a global load balancing scheduling process to collaboratively manage the heat distribution of multiple storage nodes in the hybrid storage cluster. Media load data and hotspot distribution data of each node are collected periodically to calculate a global hotspot tilt index. When the hotspot density of a node exceeds a set multiple of the global average, a cross-node data redistribution process is triggered, migrating some hotspot data evenly to adjacent nodes with lower loads, thus preventing a single node from becoming a performance bottleneck. Cross-node migration adopts the same strategy framework as intra-node migration, only adding node-level filtering conditions when selecting migration targets.
[0074] The hotspot data identification and migration method of the present invention further includes step S10: migration fault tolerance and data consistency assurance. This step introduces a transactional migration mechanism and a data verification mechanism during the bidirectional migration execution process in step S5. The transactional migration mechanism ensures data consistency during the migration process: before a data block is migrated from the source medium, it must first be written to the target medium and pass verification before the data in the source medium can be deleted; if writing or verification fails, it automatically rolls back to the state before migration, retaining a copy of the data in the source medium. The data verification mechanism automatically triggers data integrity verification after migration is completed, ensuring that no data corruption or loss occurs during the migration process by comparing the checksum of the data block with the checksum recorded before migration.
[0075] The hot data identification and migration method of this invention further includes step S11: refined data lifecycle management. This step works in conjunction with steps S3 to S8 to establish a multi-level data lifecycle management strategy. Based on the access frequency trends of data blocks and preset lifecycle rules, the aging stage of the data is automatically identified. When a data block remains in a cold data state for several consecutive statistical periods and reaches a preset aging time threshold, the system automatically marks it as archived data, executes the archive migration process, and migrates the data from high-speed media to low-cost archive storage media. The archiving process supports configuring different archiving levels, realizing hierarchical and tiered data storage management.
Claims
1. A method for identifying and migrating hotspot data in a hybrid storage architecture, characterized in that, Includes the following steps: S1: Capture IO access requests from heterogeneous storage media at all levels in a hybrid storage architecture, collect and cache access metadata corresponding to each data block, dynamically adjust the sampling window and sampling frequency, and obtain a real-time access feature dataset for each data block. S2: Based on the real-time access feature dataset, construct a four-dimensional hotspot evaluation system that integrates static media features, dynamic access features, time-series decay features, and business priority features. Normalize and weight the scores of each data block. Dynamically and adaptively adjust the weight coefficients of each feature in combination with the real-time load status of each level of storage media. Filter invalid access data, identify and eliminate pseudo-hotspot data, and output the real-time hot / cold level of each data block. S3: Based on the file size and storage format of each data block, perform block and aggregation processing on ultra-large files and fragmented small files respectively. Based on the smallest hot spot unit after adaptive adaptation, match the real-time hot and cold levels to complete the identification of local hot spots and cluster hot spots. S4: Construct a sliding fitting window based on the historical access time series data of each data block, extract access cycle features and access trend features, fit the data access change pattern, predict potential hotspot data that will heat up, and add pre-migration tags to them. S5: Pre-configure capacity high-water level thresholds, capacity low-water level thresholds, and media wear thresholds for each level of heterogeneous media in the hybrid storage architecture; match differentiated identification thresholds, migration granularity, and scheduling strategies based on media performance characteristics; and perform bidirectional closed-loop migration based on the real-time hot / cold level, pre-migration markers, and media water level status of each data block. S6: During the data migration process, the remaining available bandwidth of the hybrid storage cluster is detected in real time, and the migration IO rate is dynamically and adaptively adjusted. The migration tasks are then sorted and migrated according to their business priority tags. S7: Configure a dynamic residence cooling time window for data blocks that have completed the cold and hot media migration and switching. During the cooling window period, repeated reverse cold and hot migration operations of the same data block are prohibited to suppress the frequent jitter and oscillation of data between multiple storage media. S8: Collect media load data, service read / write latency data, hot / cold switching frequency data, and media wear data after each data migration is completed. Based on the collected operation and maintenance data, iteratively correct the feature weights, media water level thresholds, residence cooling window parameters, and migration current limiting parameters of the hot spot evaluation system to achieve adaptive iterative upgrades of system strategies.
2. The method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of streaming I / O access requests in step S1 is as follows: Deploy a bypass listening agent module on the IO data path of the hybrid storage cluster to capture IO request packets passing through each level of storage media in real time through mirror port or bypass mirror interception mechanism; The captured IO request messages are parsed to extract five core access metadata parameters: data block identifier, access operation type, access offset, access data length, and access timestamp. The extracted access metadata parameters are written to an in-memory cache queue, and then cleaned, standardized, and deduplicated through a streaming pipeline to generate structured access metadata records.
3. The method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 2, characterized in that, The specific process of dynamically adjusting the sampling window and sampling frequency in step S1 is as follows: Based on the real-time load monitoring metrics of the hybrid storage cluster, including CPU utilization, IO queue depth, and network bandwidth utilization, the sampling adjustment coefficient is dynamically calculated. When the CPU utilization exceeds the upper limit threshold or the IO queue depth exceeds the threshold, the sampling adjustment coefficient is lowered to the first interval to achieve sampling frequency reduction. When the CPU utilization is detected to be below the lower threshold and the IO queue depth is below the threshold, the sampling adjustment coefficient is increased to the second interval to improve sampling accuracy.
4. The method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of constructing the four-dimensional hotspot evaluation system in step S2 is as follows: S21: Define the characteristic dimensions of static media Characterizes the type and performance level of the storage medium where the data block is currently located, and defines the dynamic access feature dimension. Characterize the access frequency, access throughput, and IO concurrency of data blocks within a statistical period, and define the time-series decay feature dimension. Characterize the time-decrease properties of data block access behavior and define business priority feature dimensions. The priority label weight represents the business to which the data block belongs; S22: Normalize and map the original values of each feature dimension to... For the interval, a modified Min-Max standardization method is used to eliminate the impact of dimensional differences on the comprehensive score; S23: Dynamically calculate the adaptive weighting coefficients for each feature dimension based on the real-time load status of each level of storage media. , These correspond to four feature dimensions respectively; S24: Configure instantaneous batch read identification rules, the judgment criteria are that the amount of data in a single access request exceeds the preset value and the standard deviation of the access interval within the time window is less than the threshold; configure short-term pulse access identification rules, the judgment criteria are that the access frequency is pulsed within the time window and the attenuation rate exceeds the attenuation rate threshold; configure one-time crawler access identification rules, the judgment criteria are that the same client accesses the same data block only once and there is no subsequent access behavior. S25: Calculate the comprehensive hotspot score for each data block based on normalized eigenvalues and adaptive weighting coefficients. Data is classified into hot and cold categories based on its comprehensive score. Data with a score greater than or equal to the upper limit of the score threshold is considered hot data. Data with a score between the lower limit and the upper limit of the score threshold is considered warm data. Data with a score less than the lower limit of the score threshold is considered cold data.
5. A method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 4, characterized in that, The specific process of data granularity adaptive adaptation in step S3 is as follows: S31: Divide data blocks into three categories according to file size: oversized files, regular-sized files, and fragmented small files; S32: For ultra-large files, the file is divided into multiple data sub-blocks according to a uniform block granularity of a specified size. Each data sub-block is independently allocated a block-level access statistics counter, and the access frequency of each sub-block is collected separately to realize the identification of local hot spots within the file. S33: For fragmented small files, multiple small files are aggregated into logical aggregation units according to storage path or access time window. Access heat statistics are established for the aggregation unit as a whole, and hot spot determination and migration scheduling are performed at the aggregation unit as the granularity. S34: Match and calibrate the smallest hot spot unit after granularity adaptation with the real-time hot and cold levels output in step S25 to generate local hot spot calibration results and cluster hot spot calibration results.
6. The method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of predicting potential hotspots based on time-series trends in step S4 is as follows: S41: Starting from the current moment, rewind along the historical timeline. The duration is used to construct a sliding fitting window, where It is adaptively determined based on the access cycle characteristics of the data block; S42: Perform spectral analysis on the access time series data within the sliding fitting window, and use Fast Fourier Transform to extract the main period frequency of the access behavior to generate an access period fingerprint; S43: Use the exponentially weighted moving average algorithm to fit the trend of the access frequency time series data and calculate the access slope. ; S44: When access slope Greater than the preset trend threshold Furthermore, when the current hot spot score is lower than the hot spot determination threshold, the data block is determined to be potential hot spot data and a pre-migration mark is automatically added; potential hot spot data enters the pre-migration priority queue, waiting to be migrated to high-speed storage media.
7. The method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of heterogeneous media differentiated bidirectional migration scheduling in step S5 is as follows: S51: Configure a high-water threshold for memory media capacity. Low water level threshold Media wear threshold Configure a high-water mark for NVMe SSD capacity. Low water level threshold Media wear threshold Configure a high-water mark for the capacity of SATA SSDs. Low water level threshold Media wear threshold ; Configure high watermark capacity for mechanical hard drives Low water level threshold Media wear threshold ; S52: Configure the identification threshold and migration granularity differently according to the media type, including the hotspot identification threshold and corresponding preset migration granularity for NVMeSSD; the hotspot identification threshold and corresponding preset migration granularity for SATASSD; and the hotspot identification threshold and corresponding preset migration granularity for mechanical hard drives. S53: Real-time monitoring of the capacity utilization of storage media at all levels. When the capacity of high-speed storage media reaches the corresponding high water level threshold, the migration scheduling process is triggered. When the capacity of the high-speed storage medium drops back to the corresponding low water level threshold, the warm and cold data migration is stopped to maintain the steady state of storage resources. S54: Prioritize migrating hot data into high-speed media, while migrating cold data out of high-speed media; for potential hot data carrying pre-migration tags, prioritize migration resources to ensure that potential hot data is installed in high-speed media in advance.
8. A method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 7, characterized in that, The specific process of bandwidth awareness and service priority-based coordinated rate limiting migration in step S6 is as follows: S61: Real-time data collection of network interface throughput, IO controller bandwidth, and storage medium IO bandwidth of the hybrid storage cluster via the system network interface monitoring module; calculation of the current remaining available bandwidth. ; S62: Dynamically adjust the migration I / O rate based on the remaining available bandwidth; S63: Extract business priority tags from the metadata of data blocks. Priority tags are divided into four levels: urgent, high, normal, and low. S64: Sort the migration task queue in descending order according to the business priority label. Urgent migration tasks are scheduled and executed first. Tasks within the same priority are sorted according to the hot / cold level and the pre-migration mark.
9. A method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of cold and hot migration oscillation suppression and stabilization in step S7 is as follows: S71: Dynamically calculate the cooling time window duration based on the historical hot / cold switching frequency of the data block. ; S72: To lock the current hot / cold state of a data block that has completed the switching of hot and cold media within the cooling time window, repeated reverse migration operations are prohibited. S73: When the cooldown window expires, release the state lock, allowing the data block to re-determine its hot / cold status based on real-time access characteristics and participate in subsequent migration scheduling.
10. A method for identifying and migrating hotspot data in a hybrid storage architecture according to claim 1, characterized in that, The specific process of the full-process closed-loop self-learning iterative optimization in step S8 is as follows: S81: After each data migration is completed, automatically collect and record media load change data, service read / write latency change data, cold / hot switching frequency data, and media wear increment data to form a migration execution log; S82: Compare and analyze the actual results data after the migration is executed with the expected strategy objectives, and calculate the actual deviation values of each strategy parameter; S83: Iterative Correction: Based on the strategy deviation value, the feedback adjustment algorithm is used to iteratively correct the characteristic weight coefficients, medium water level threshold parameters, residence cooling window parameters, and migration flow restriction parameters of the hot spot evaluation system; S84: Generates the revised policy parameters into a new version policy configuration, dynamically distributes it to each node of the hybrid storage cluster, realizes online hot updates of the policy, and completes the closed-loop self-learning iterative optimization of the entire process.