Data synchronization method and electronic equipment
By calculating the heat statistics index movement value in remote replication scenarios to perform data synchronization and tiered migration, the problem of cold data occupying SSD resources at remote sites is solved, data synchronization efficiency and storage costs are improved, and critical data can maintain high performance after failover, optimizing tiering strategies and WAN bandwidth utilization.
Patent Information
- Application Number
- CN202511175056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In remote replication scenarios with tiered storage devices, the remote site's SSD resources are often occupied by a large amount of cold data, resulting in low data synchronization efficiency and high WAN bandwidth consumption. Furthermore, the tiering policy at the remote site is inconsistent with that at the local site, impacting performance and synchronization efficiency after failover.
By calculating the heat statistics index movement value at the local site, determining the data type and sending it to the remote site, the remote site calculates and performs tiered migration processing based on the received data, prioritizes synchronizing hot data, utilizes WAN bandwidth, compresses or limits cold data, ensures that critical data is located in the high-speed storage layer, and combines the tiered preheating mechanism to optimize the tiered status of the remote site.
It improves data synchronization efficiency at remote sites, saves WAN bandwidth costs, ensures that critical business data can maintain high performance after failover, optimizes storage costs and tiering strategies at remote sites, reduces performance fluctuations after synchronization, and improves recovery efficiency after failover.
Smart Images

Figure CN120676005A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technology, and in particular to a data synchronization method and electronic device. Background Art
[0002] Currently, storage device tiering is a technology that allocates data to different storage media based on access frequency, performance requirements, and cost-effectiveness, aiming to optimize the overall performance, capacity, and cost of storage systems. Remote replication, on the other hand, is a key technology that replicates data from a local storage location to a remote device in real time or periodically to achieve data protection, disaster recovery, and business continuity.
[0003] In remote replication scenarios between two or more sets of tiered storage devices, related remote replication solutions typically ignore the local tier status and simply copy all local source data blocks at all tiers to a single target tier at the remote site (usually a storage tier that compromises performance or capacity). The remote site uses a fixed high-performance tier of solid-state drives (SSDs) for data synchronization. As a result, the remote site's SSD resources are occupied by a large amount of rarely accessed cold data, resulting in low data synchronization storage efficiency. Summary of the Invention
[0004] The present application provides a data synchronization method and electronic device, the method comprising: a local site receiving and storing read and write request data; the local site calculating a first heat statistical index movement value based on the read and write request data, and sending the first heat statistical index movement value to a remote site; a local tiering engine performing local tiering migration processing based on a heat analysis result of the first heat statistical index movement value; the local site determining whether the type of the requested data is write data, and in response to the request data type being write data, the local site sending the write data to the remote site, and the remote site calculating a second heat statistical index movement value based on the received write data; a remote tiering engine performing remote tiering migration processing based on a full system heat analysis result of the second heat statistical index movement value and the first heat statistical index movement value; and performing data synchronization between the local site and the remote site based on a priority identifier of the requested data. The present application can solve the problem of low data synchronization efficiency caused by the limited synchronization capability of the local site and the expensive wide area network bandwidth consumed by the synchronization of a large amount of cold spot data.
[0005] The present application provides a data synchronization method, which is applied to a data synchronization system. The system includes a local site and a remote site. The local site includes a local tiering engine, and the remote site includes a remote tiering engine. The method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0006] The present application also provides a data synchronization system, the system including a local site and a remote site, the local site including a local layering engine, the remote site including a remote layering engine, and the system including a data synchronization module; The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. The data synchronization module is used to synchronize data between the local site and the remote site according to the priority identifier of the requested data.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of a data synchronization method when executing the computer program. The method comprises: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0008] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data synchronization method are implemented, and the method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0009] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the data synchronization method are implemented. The method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0010] Through the present application, the method includes the following steps: a local site receives and stores read and write request data; the local site calculates a first heat statistical index moving value based on the read and write request data, and sends the first heat statistical index moving value to the remote site; a local tiering engine performs local tiering migration processing based on a heat analysis result of the first heat statistical index moving value; the local site determines whether the type of the request data is write data, and in response to the request data type being write data, the local site sends the write data to the remote site, and the remote site calculates a second heat statistical index moving value based on the received write data; the remote tiering engine performs remote tiering migration processing based on a full system heat analysis result of the second heat statistical index moving value and the first heat statistical index moving value; and data is synchronized between the local site and the remote site based on the priority identifier of the request data. The present application can solve the problem of low data synchronization efficiency caused by the limited synchronization capability of the local site and the expensive wide area network bandwidth consumed by the synchronization of a large amount of cold spot data.
[0011] The technical solution of this application can significantly improve the performance of remote sites: ensure that key business data is located in the high-speed storage layer after data synchronization to maintain application performance; optimize remote site storage costs: avoid cold data occupying expensive high-speed storage resources; efficiently utilize wide area network bandwidth: prioritize the replication of hot data, compress, limit or selectively replicate cold data, and save channel bandwidth costs; accelerate remote tiering optimization: utilize local data heat information to enable the remote site tiering engine to reach the optimal storage state faster; improve fault switching / data synchronization quality: reduce performance fluctuations after data synchronization through tiered preheating; improve the efficiency of recovery after data synchronization interruption: tiered information synchronization accelerates data layout optimization after incremental data synchronization. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 A first flowchart of data synchronization provided in an embodiment of the present application; Figure 2 A second flowchart of data synchronization provided in an embodiment of the present application; Figure 3 A QoS flow chart of remote replication of tiered storage provided by an embodiment of the present application; Figure 4 A flowchart of the hierarchical storage remote replication write process provided in an embodiment of the present application; Figure 5 A structural diagram of a data synchronization system provided in an embodiment of the present application; Figure 6 The exemplary system provided for the embodiments of the present application can be used to implement various embodiments in the present application. DETAILED DESCRIPTION
[0014] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0015] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0016] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0017] Related remote replication methods typically do not consider the tiered state of the source end, but simply replicate all local source data blocks at all tiers to a single target tier at the remote site. This technical solution also leads to the following problems: Performance loss: When the remote site uses a low-performance tier of fixed mechanical hard disks (HDDs) for data synchronization, hot data on the local solid-state drive (SSD) is copied to the remote site and then placed on the slower mechanical HDD. When failover to the remote site is required, the performance of critical applications drops sharply, defeating the original purpose of disaster recovery.
[0018] Low replication efficiency: Because the data in the remote storage target area is stored in a low-performance tier, the primary-side data IO capability is limited in synchronous replication. The replication of large amounts of cold data consumes expensive WAN bandwidth and remote storage, resulting in low replication efficiency.
[0019] Split tiering strategies: The tiering strategies of local and remote sites usually run independently and are unaware of each other; "hot" data identified locally may be considered "cold" and downgraded at the remote site, and vice versa.
[0020] Inconsistent performance after failover: When the local site fails and services are switched to the remote site, application performance cannot reach the same level as the local site because the data tiering status of the remote site is different from that of the local site.
[0021] Low synchronization efficiency: When recovering from a link interruption or synchronizing incremental data, the synchronization efficiency of full or coarse-grained data without hierarchical layers is relatively low.
[0022] The embodiment of the present application provides a data synchronization method, such as Figure 1 As shown, the method is applied to a data synchronization system, the system includes a local site and a remote site, the local site includes a local tiering engine, the remote site includes a remote tiering engine, and the method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0023] It can be understood that the storage system of the present application is composed of several underlying block resource extents within the storage. Each data block resource can be provided by a raid composed of disks with different attributes such as HDD or SSD. When writing data blocks to local volumes, the storage system performs information statistics on each resource unit extent, recording not only its physical location (LBA), but also its storage tier identifier (Tier ID - such as SSD_Tier_1, HDD_Tier_2, Archive_Tier) and heat metadata (such as access frequency score, last access timestamp, business importance label).
[0024] Here, the popularity metadata is equivalent to the popularity statistical index moving value.
[0025] When performing remote replication, the system not only replicates the data block itself, but also associates the replication with the corresponding Extents storage tier identifier and heat metadata, which is sent to the remote site as an independent heat metadata update of the replication data stream.
[0026] Optional remote replication tiering policies: During the remote replication service creation phase, the source (local site) specifies the target tier where data blocks should be placed at the remote site based on the source data policy and replication importance, limiting inter-tier data migration to the backend data of the target tier on the slave side.
[0027] During the creation phase of a remote replication service, the target tier of the slave data is not mandatory. During data synchronization, the source side provides Extent tier identifiers and popularity metadata as suggestions, and the remote site makes the final decision based on its own tiering strategy and resource conditions.
[0028] Remote Site Tiering Engine Integration: The tiering engine at the remote site receives the data blocks copied from the source and the storage tier identifiers and heat information associated with their corresponding extents. The remote end receives the volume identifier LUN data and synchronously places the newly copied data blocks directly into the corresponding block resources of the remote replication volume, without considering the target tier recommended or forcibly specified by the source.
[0029] The remote tiering engine uses the received heat metadata as key input and combines it with the locally monitored access patterns to preheat the data tiers, quickly and accurately migrating data blocks to the appropriate local tier. If the target volume LUN is configured for forced tier mapping, data tier migration will not be performed based on heat metadata.
[0030] Finally, the local storage synchronizes data with the slave end, and the Quality of Service (QOS) policy can be configured based on the layer, limiting the upper and lower bandwidth limits of each layer to achieve efficient utilization of remote replication bandwidth resources. Data in high-priority layers (such as solid-state drives (SSDs)) are prioritized for replication to ensure low latency and high bandwidth allocation; synchronous or low-latency asynchronous replication can be used; cold data in low-priority layers (such as mechanical hard drives (HDDs)) is allowed higher replication latency and can be bandwidth-restricted.
[0031] It can be understood that the local storage site monitors the access frequency of data blocks in the local storage volume; based on the access pattern, the data blocks are dynamically allocated to local storage media with different performance levels; after the level is allocated or migrated for the data block, the level identification information of the data block is recorded; when the data block replication to the remote storage site is started or performed, the data block is transferred to the remote storage site, and its associated storage level heat identification information is also passed to the remote storage site.
[0032] The embodiment of the present application provides a data synchronization method, such as Figure 2 As shown, the method is applied to a data synchronization system, the system includes a local site and a remote site, the local site includes a local tiering engine, the remote site includes a remote tiering engine, and the method includes: Step S01: The local site receives storage read / write request data.
[0033] Specifically, read and write request data is sent: the IO client sends read and write IO requests to the local storage; Local end reception and processing: local storage receives and processes IO request data; Determine whether the requested data IO is a write operation; if it is a write IO, continue to synchronize data to the remote site; if it is not a write IO (i.e., a read IO), directly read the local storage resource and end the process after the local IO is processed.
[0034] In step S02, the local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tiering migration processing based on the heat analysis result of the first heat statistics index movement value.
[0035] Step S021: The local tiering engine determines whether the local site should perform local tiering migration based on the heat analysis result and the first period threshold; If yes, the read / write request data is migrated to the corresponding storage tier based on the heat analysis results; if no, the current tier status of the storage is maintained; Migrate read and write request data to the corresponding storage tier based on heat analysis results, including: When the first heat statistical index movement value of the read / write request data is greater than a first preset value, the read / write request data is determined to be hot business data, and the hot business data is migrated to the first storage tier (SSD tier) of the local site; When the first heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold spot business data, and the cold spot business data is migrated to the second storage tier (HDD tier) of the local site.
[0036] Specifically, count IO heat, generate storage LUN-extent heat and calculate EMA: perform heat analysis on read and write IO, and generate the first heat statistical exponential moving average (EMA) corresponding to the read and write data; Here, the heat statistics exponential moving value is preferably the heat statistics exponential moving average.
[0037] Send the generated popularity metadata (the first popularity statistical index moving value) to the remote site for subsequent read and write data synchronization; Performing further heat analysis based on the generated first heat statistical index movement value; The local tiering engine determines whether to perform tiered migration at the local site based on the first cycle threshold (every 5 minutes or every hour): If tiered migration is required, data is migrated to the appropriate storage tier based on the heat analysis results. If not required, the current tiering state of the local site storage pool is maintained; After completing a stratification adjustment cycle, wait for the next cycle to arrive.
[0038] The first preset value and the second preset value are different values calculated according to different scenarios. For example, the first preset value is 80, and the second preset value is 20.
[0039] In step S03, the local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site, and the remote site calculates the second heat statistical index movement value based on the received write data; the remote tiering engine performs remote tiering migration processing based on the full system heat analysis results of the second heat statistical index movement value and the first heat statistical index movement value.
[0040] Specifically, the remote site receives data synchronization replication IO from the local site; writes the received data into a storage device at the remote site; and completes data synchronization IO processing at the remote site.
[0041] It can be understood that the remote storage site receives data blocks and their associated hierarchical identification information from the source storage site; based on the storage hierarchical identification information of the received data blocks, the data blocks are stored on the target hierarchical storage medium corresponding to the storage hierarchical identification information on the remote storage site.
[0042] The storage tier identification information includes popularity metadata. The remote site uses the popularity metadata as input and combines it with its own tiering strategy to decide and migrate data blocks to the appropriate storage tier.
[0043] Step S031: The remote tiering engine determines whether the remote site should perform remote tiering migration based on the first period threshold according to the system-wide heat analysis result; If yes, the read and write request data will be migrated to the corresponding storage tier based on the system-wide heat analysis results; if no, the current tier status of the storage will be maintained; Migrate read and write request data to the corresponding storage tier based on the system-wide heat analysis results, including: When the heat statistical index movement value of the read / write request data is greater than a first preset value, the read / write request data is determined to be hot business data; and the hot business data is migrated to the first storage tier of the remote site; When the heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold-spot business data; and the cold-spot business data is migrated to the second storage tier of the remote site.
[0044] Specifically, the remote site receives write data from the local site; based on the received write data, a corresponding second heat statistical index movement value is generated at the remote site; combined with the first heat statistical index movement value, a heat analysis of the entire system is performed; the heat analysis of the entire system breaks the limitations of the "local perspective" and constructs a unified global data heat view.
[0045] The local site continuously uses EMA (Exponential Moving Average) to count the access frequency of each data block (such as a LUN extent). In remote replication scenarios, hotness metadata is synchronized to the remote site. After receiving this metadata, the remote site combines it with the locally recorded write / read behavior to form a more complete hotness profile, avoiding remote "cold start" issues caused by failover. This achieves unified data hotness awareness across sites, providing a basis for intelligent tiering. The local tiering engine periodically determines whether to perform tier migration: if tier migration is required, it is performed based on the data type; if not, the current storage tier status is maintained. Based on the hotness analysis results, data is migrated to the appropriate storage tier. Hot data: For data objects with an EMA value higher than a first preset value, migrate them to a higher-performance storage tier, such as a solid-state drive (SSD), to speed up access. Cold data: For data objects with an EMA value lower than the second preset value, consider migrating them to a lower-performance but more cost-effective storage tier, such as a mechanical hard disk (HDD), to save expensive high-performance storage space. Moderately hot data: Data between two preset values can be migrated or remain unchanged based on specific circumstances. The current tier adjustment cycle ends: After completing a tier adjustment cycle, wait for the arrival of the next cycle.
[0046] The entire process ensures data consistency and efficient access through data synchronization and hierarchical management between local and remote sites. Each site independently performs heat analysis and hierarchical adjustment, while maintaining data consistency between the two locations through periodic data synchronization.
[0047] Step S04, obtaining the exponential moving value EMA(old) of the last read / write request data, the new read / write request data point NewValue at the current moment, and a smoothing coefficient alpha; wherein the smoothing coefficient is a value between 0 and 1; Calculate the first heat statistical exponential moving value EMA(new) of the read / write request data at the current moment using the formula: EMA(new)=EMA(old)×(1-alpha)+NewValue×alpha; Calculate the second heat statistics index moving value based on the write data, including: Get the exponential moving value EMA(old)^ of the heat statistics of the data written at the previous moment and the new data point NewValue^ at the current moment; The second heat statistical exponential moving value EMA(new)^ of the data written at the current moment is calculated using the formula: EMA(new)^=EMA(old)^×(1-alpha)+NewValue^×alpha.
[0048] Specifically, the local storage captures the IO statistics of read and write request data. At predetermined time intervals, the number of IO accesses to each extent unit is periodically counted and stored. This is used as the heat information input. The exponential moving average of the heat statistics is periodically calculated to obtain the EMA (Exponential Moving Average): EMA(new)=EMA(old) ×(1-alpha)+NewValue×alpha; By smoothing heat statistics, abnormal I / O can be avoided to avoid layer jumps.
[0049] alpha (α): Smoothing coefficient (also called weight factor), a value between 0 and 1 that determines the weight of new data.
[0050] Step S05: performing a layered warm-up operation at the local site and the remote site; Perform tiered warm-up operations at the local site and remote sites, including: Identify hot business data in the corresponding storage pool based on the popularity metadata information of the last data synchronization at the local site or the popularity metadata information accumulated at the remote site, and store the hot business data from the second storage tier of the storage pool to the first storage tier; Among them, the popularity metadata information includes data access frequency score, data access timestamp, and performance label of business data.
[0051] Specifically, the remote storage site performs: based on the storage tier identification information received and stored from the local site, or based on the popularity information accumulated at the remote site to identify key data blocks, and actively migrate the key data blocks to the high-performance tier storage medium of the remote storage site.
[0052] Step S06: Data synchronization is performed between the local site and the remote site according to the priority identifier of the requested data.
[0053] Step S061, setting the hot spot service data as high priority data, and setting the cold spot service data as low priority data; When the requested data is high-priority data, data synchronization is prioritized and the transmission channel bandwidth between the local site and the remote site is widened; When the requested data is low-priority data, data synchronization is delayed and the bandwidth of the transmission channel between the local site and the remote site is limited.
[0054] Step S062: The local site sets a corresponding priority tag for each request data; Get the priority identifier of the current request data; Determine the storage level of the requested data at the local site based on the priority identifier of the requested data; Set corresponding preset thresholds for each storage tier at the local site; Determine whether the resource ratio of the level stored at the local site corresponding to the requested data has reached a preset threshold; If so, data synchronization of the requested data to the remote site is delayed; if not, data synchronization of the requested data is performed between the local site and the remote site.
[0055] Specifically, if Figure 3 The figure shows the tiered storage remote replication QoS (Quality of Service) process, which involves three main parts: local storage IO, local tiered QoS, and remote site IO: 1. Remote replication process Configure remote replication parameters and policies on the local storage; initiate IO synchronization operations based on the configured remote replication relationship and send the requested data to the remote site.
[0056] 2. Local Tiered QoS The local site sets priority tags for different IO request data based on pre-configured QoS policies; obtains the priority identifier associated with the current IO request data; Determine, based on the obtained identifier, which storage layer the requested data belongs to; Insert IO requests into corresponding storage according to their hierarchical attributes; Check whether the IO request in the current layer has reached the preset threshold; Here, the system sets a preset threshold for each, for example: maximum concurrent IO number ≤ 100; CPU / IO weight ratio ≤ 30%; If yes, the execution of the IO request is temporarily delayed and waits for subsequent processing; if no, data synchronization continues.
[0057] 3. Remote Site IO For IO requests that do not reach the threshold value, the request data IO transmission operation is directly performed; after the remote site receives the IO request data, it is processed; and the IO processing process of the remote site is completed.
[0058] The entire process classifies and schedules IO requests through the QoS mechanism on local storage, ensuring that request data of different priorities receive reasonable resource allocation and service quality assurance during the remote replication process; high-priority IO requests will be processed and transmitted first, while low-priority IO requests or IO requests that reach the threshold will be delayed to ensure the performance and stability of the overall system.
[0059] Step S07, compressing and deduplicating the low-priority data; Perform data compression and deduplication on low-priority data, including: Compressing the low-priority data using a compression algorithm, wherein the compression algorithm includes a lossless data compression algorithm; Get the hash value of low-priority data and create a data storage index; Compare the hash value of the low-priority data with the data storage index; When the hash value of low-priority data exists in the data storage index, the low-priority data is deduplicated; In response to data synchronization between the local site and the remote site, synchronizing the difference data blocks and the data storage tier identifier and the heat metadata information of the difference data blocks updated during the interruption of data synchronization; The data storage tier identifier includes: a storage tier identifier allocated by a data storage site, a data heat statistics index movement value, and a data logical unit number identifier.
[0060] Specifically, when incremental data synchronization is resumed after a remote replication link is interrupted, the local storage site not only transmits the differential data blocks, but also transmits the storage layer identification information of these differential data blocks updated during the interruption, ensuring that the remote site can immediately optimize the data layout according to the latest status after the data synchronization is completed.
[0061] Step S08, performing reverse data synchronization between the local site and the remote site; Reverse data synchronization between the local site and the remote site, including: Get the data blocks that changed when the remote site last synchronized data; Synchronize the data blocks that changed during the last data synchronization at the remote site to the local site through incremental synchronization; Verify the integrity and correctness of data blocks after reverse data synchronization; In response to executing reverse data synchronization, the local site obtains the popularity metadata information updated during the data synchronization from the remote site, and optimizes the storage pool tiering state of the local site according to the updated popularity metadata information.
[0062] Specifically, data compression and deduplication are performed on low-priority tier data to further save bandwidth. Before failover to a remote site, the system can perform tier warm-up operations: based on the heat information last synchronized from the source end or the heat information accumulated by the remote site, critical business data is proactively promoted from a slower tier (such as HDD) to a high-performance tier (such as SSD). After reverse replication (failback) or recovery of the primary site, the source site can obtain the updated heat information accumulated during the failover from the remote site to quickly restore or optimize the local tier status.
[0063] Here, as Figure 4 As shown, the master-side storage of this application calculates which extent objects to migrate hierarchically based on the local resource heat by the hierarchical function migration analysis engine; and the hierarchical migration decision generation depends on the extent migration cost and the post-migration benefit.
[0064] The master-side storage heat metadata manager module maintains a LUN-extent-EMA heat information table and generates a change table after each heat statistics cycle to record which extents have heat changes. Heartbeat detection is performed periodically between clusters. When the remote partner status is detected to be normal, the master-side storage synchronizes the heat change table to the remote storage to achieve hierarchical metadata synchronization and clear the heat change table. If the remote site status is abnormal, the change table is continuously updated until the remote site status is restored.
[0065] During data synchronization, the source side provides Extent storage tier identifiers and popularity metadata as recommendations, and the remote site makes the final decision based on its own tiering strategy and resource conditions.
[0066] The remote storage site receives the tiered heat metadata and synchronizes the heat change table, updates the local heat information table that records the master storage heat metadata, and also maintains a local heat information table locally. When the tiered DMP migration analysis cycle arrives, the remote storage will compare and update the local heat information table based on the remote tiering policy configuration, and use the master heat value to replace the remote storage local heat value statistics or maintain the original local heat value. The remote migration analysis engine module calculates and decides which extent objects to migrate tiered based on the updated local heat metadata. The migration logic is consistent with the master storage.
[0067] Data is synchronized from the primary storage to the secondary storage, and QOS policies can be configured based on the layers to limit the upper and lower bandwidth limits of each layer, thereby achieving efficient utilization of remote replication bandwidth resources. Data in high-priority layers (such as SSDs) is prioritized for replication to ensure low latency and high bandwidth allocation, and can use synchronous or low-latency asynchronous replication. Cold spot data in low-priority layers (such as HDDs) is allowed higher replication latency and can be bandwidth-restricted.
[0068] Remote sites support flow control management of data at different layers during data synchronization. SSD and HDD layers are configured separately on the primary storage side, and support limiting the upper and lower bandwidth limits of each layer to achieve efficient utilization of remote replication bandwidth resources.
[0069] Based on the heat information last synchronized from the source or the heat information accumulated by the remote site, critical business data is proactively promoted from a slower tier to a higher-performance tier. After reverse replication (failback) or restoring the primary site, the source site can obtain the updated heat information accumulated during the failover from the remote site to quickly restore or optimize the local tier status. Synchronizing tiered heat information between the primary and remote storage can resolve the issue of inconsistent heat between the primary and remote storage end readings due to the remote storage only having write IO. This allows the remote storage end to synchronously obtain the host's IO heat information for the data, and the remote storage migrates the data to the corresponding heat tier based on the synchronized information. Once the need to switch remote partnerships arises, the data required by the original remote storage end host is already in the corresponding tier, achieving "tiered preheating."
[0070] This involves an advanced data management and optimization strategy, which is mainly used to optimize data tiering by synchronizing popularity information in storage systems, especially in disaster recovery (failover) and reverse replication (failback) scenarios: Refers to information such as data access frequency and patterns recorded when the primary site (source) last synchronized with the remote site. It also includes information about new data access popularity generated by the remote site as it continues to provide services during a failover. This information is used to identify critical business data and proactively promote it from slower storage tiers (such as HDDs) to higher-performance tiers (such as SSDs) based on its importance and access frequency. This ensures that critical data can be quickly accessed when needed, improving system performance. After performing reverse replication or restoring the primary site, the source site can obtain updated heat information accumulated during the failover from the remote site. This enables the source site to quickly adjust its local tiering state after recovery to match the latest usage patterns and needs. This helps quickly restore system performance and reduces service interruption time caused by data redistribution. The tiered heat information synchronization mechanism between the primary storage and remote storage solves the problem of inconsistent read data heat and remote storage heat due to the remote storage only receiving write IO requests. Through this data synchronization, the remote storage can obtain the host's actual IO heat information for the data and migrate the data to the corresponding heat tier accordingly, achieving more accurate data tier management. When a remote site relationship needs to be switched, the data required by the original remote storage host is already in the corresponding tier location, achieving so-called "tier pre-warming." This means that during the switchover process, no additional time is required for data migration or adjustment, thus speeding up the entire failure recovery process. This strategy uses heat information synchronization technology to optimize data tiering in disaster recovery and daily operations. It not only improves the system's response speed and service quality, but also reduces fault recovery time, enhances system availability and flexibility, and demonstrates how modern storage solutions can combine intelligent algorithms to improve user experience and resource utilization.
[0071] In addition, after data synchronization between the local site and the remote site, including: Optimize the remote replication tier system based on the latest data access status information; Obtain the latest data access popularity information from the source site (local site), including all updated popularity data accumulated during the failover period. Combined with the write operations recorded by the remote site itself and the popularity information generated, this information is analyzed to obtain a comprehensive view of data access patterns. Evaluate whether the existing data distribution is reasonable based on the current tiering strategy; see which data is in the high-performance tier and which is in the lower-performance tier; and find data blocks or files that no longer fit in their current location due to recent changes in data activity. Newly identified "hot spot" data (frequently accessed data) is migrated to a higher-performance storage tier; "cold spot" data (rarely accessed data) is migrated to a storage tier with lower costs and lower performance requirements; Prioritize different migration tasks based on business needs and resource availability; Use the storage system's automated tools or scripts to perform data migration according to the optimized plan. Perform large-scale data migrations during off-peak hours and limit the migration traffic rate to minimize the impact on front-end application performance. After the migration is complete, verify that the data is correctly placed on the corresponding storage tier and confirm that no data is lost; Establish a long-term monitoring mechanism to regularly reassess data popularity and adjust tiering strategies accordingly; maintain optimal performance of the storage system; By collecting performance indicators and user feedback from actual operations, we continuously optimize tiering algorithms and strategies, improving the adaptability and efficiency of the entire storage architecture.
[0072] Through the above steps, the remote site can effectively optimize its tiered system layout after data synchronization is completed, ensuring that high levels of service quality and resource utilization are maintained even after a failover; this approach not only improves storage efficiency but also enhances system flexibility and responsiveness.
[0073] The data synchronization method provided in the embodiment of the present application can also be improved and optimized without departing from the technical solution of the present application, and these improvements and optimizations should also be regarded as the scope of protection of the present application.
[0074] The beneficial effects of the technical solution provided by the embodiments of the present application are: This application can solve the problem of low data synchronization efficiency caused by the limited synchronization capability of local sites and the expensive wide area network bandwidth consumed by the synchronization of a large amount of cold spot data.
[0075] The technical solution of this application can significantly improve the performance of remote sites: ensure that key business data is located in the high-speed storage layer after data synchronization to maintain application performance; optimize remote site storage costs: avoid cold data occupying expensive high-speed storage resources; efficiently utilize wide area network bandwidth: prioritize the replication of hot data, compress, limit or selectively replicate cold data, and save channel bandwidth costs; accelerate remote tiering optimization: utilize local data heat information to enable the remote site tiering engine to reach the optimal storage state faster; improve fault switching / data synchronization quality: reduce performance fluctuations after data synchronization through tiered preheating; improve the efficiency of recovery after data synchronization interruption: tiered information synchronization accelerates data layout optimization after incremental data synchronization.
[0076] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0077] The embodiment of the present application also provides a data synchronization system, such as Figure 5 As shown, the system includes a local site and a remote site. The local site includes a local tiering engine, and the remote site includes a remote tiering engine. The system includes a data synchronization module, a calculation module, a tiering preheating module, and a processing module.
[0078] In this embodiment, the local site receives storage read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. The data synchronization module is used to synchronize data between the local site and the remote site according to the priority identifier of the requested data.
[0079] In this embodiment, the local site includes a local storage pool, which includes a first storage tier at the local site and a second storage tier at the local site. The local tiering engine performs local tier migration processing based on a heat analysis result of a first heat statistical index movement value, including: The local tiering engine determines whether to perform local tiering migration at the local site based on the heat analysis results and the first cycle threshold; If yes, the read / write request data is migrated to the corresponding storage tier based on the heat analysis results; if no, the current tier status of the storage is maintained; Migrate read and write request data to the corresponding storage tier based on heat analysis results, including: When the first heat statistical index movement value of the read / write request data is greater than a first preset value, the read / write request data is determined to be hot business data, and the hot business data is migrated to the first storage tier of the local site; When the first heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold spot business data, and the cold spot business data is migrated to the second storage tier of the local site.
[0080] In one embodiment, the remote site includes a remote storage pool, the remote storage pool includes a first storage tier at the remote site and a second storage tier at the remote site, and the remote tiering engine performs remote tier migration processing based on a system-wide heat analysis result of the second heat statistical index movement value and the first heat statistical index movement value, including: The remote tiering engine determines whether to perform remote tiering migration at the remote site based on the first cycle threshold according to the system-wide heat analysis results; If yes, the read and write request data will be migrated to the corresponding storage tier based on the system-wide heat analysis results; if no, the current tier status of the storage will be maintained; Migrate read and write request data to the corresponding storage tier based on the system-wide heat analysis results, including: When the heat statistical index movement value of the read / write request data is greater than a first preset value, the read / write request data is determined to be hot business data; and the hot business data is migrated to the first storage tier of the remote site; When the heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold-spot business data; and the cold-spot business data is migrated to the second storage tier of the remote site.
[0081] In one embodiment, the calculation module is used to obtain the heat statistical exponential moving value EMA(old) of the read / write request data at the previous moment, the new read / write request data point NewValue at the current moment, and the smoothing coefficient alpha; wherein the smoothing coefficient is a value between 0 and 1; Calculate the first heat statistical exponential moving value EMA(new) of the read / write request data at the current moment using the formula: EMA(new)=EMA(old)×(1-alpha)+NewValue×alpha; Calculate the second heat statistics index moving value based on the write data, including: Get the exponential moving value EMA(old)^ of the heat statistics of the data written at the previous moment and the new data point NewValue^ at the current moment; The second heat statistical exponential moving value EMA(new)^ of the data written at the current moment is calculated using the formula: EMA(new)^=EMA(old)^×(1-alpha)+NewValue^×alpha.
[0082] In one embodiment, the data synchronization module is used to set the hot spot business data as high priority data and the cold spot business data as low priority data; When the requested data is high-priority data, data synchronization is prioritized and the transmission channel bandwidth between the local site and the remote site is widened; When the requested data is low-priority data, data synchronization is delayed and the bandwidth of the transmission channel between the local site and the remote site is limited.
[0083] In one embodiment, the data synchronization module is used to set a corresponding priority tag for each request data; Get the priority identifier of the current request data; Determine the storage level of the requested data at the local site based on the priority identifier of the requested data; Set corresponding preset thresholds for each storage tier at the local site; Determine whether the resource ratio of the level stored at the local site corresponding to the requested data has reached a preset threshold; If so, data synchronization of the requested data to the remote site is delayed; if not, data synchronization of the requested data is performed between the local site and the remote site.
[0084] In one embodiment, a tiered preheating module is configured to perform tiered preheating operations at a local site and a remote site; Perform tiered warm-up operations at the local site and remote sites, including: Identify hot business data in the corresponding storage pool based on the popularity metadata information of the last data synchronization at the local site or the popularity metadata information accumulated at the remote site, and store the hot business data from the second storage tier of the storage pool to the first storage tier; Among them, the popularity metadata information includes data access frequency score, data access timestamp, and performance label of business data.
[0085] In one embodiment, the processing module is configured to perform data compression and deduplication processing on low priority data; Perform data compression and deduplication on low-priority data, including: Compressing the low-priority data using a compression algorithm, wherein the compression algorithm includes a lossless data compression algorithm; Get the hash value of low-priority data and create a data storage index; Compare the hash value of the low-priority data with the data storage index; When the hash value of low-priority data exists in the data storage index, the low-priority data is deduplicated; In response to data synchronization between the local site and the remote site, synchronizing the difference data blocks and the data level identifiers and heat metadata information of the difference data blocks updated during the interruption of data synchronization; The data tier identifier includes: a storage tier identifier assigned by a data storage site, a data heat statistics index movement value, and a data logical unit number identifier.
[0086] In one embodiment, the processing module is configured to perform reverse data synchronization between the local site and the remote site; Reverse data synchronization between the local site and the remote site, including: Get the data blocks that changed when the remote site last synchronized data; Synchronize the data blocks that changed during the last data synchronization at the remote site to the local site through incremental synchronization; Verify the integrity and correctness of data blocks after reverse data synchronization; In response to executing reverse data synchronization, the local site obtains the popularity metadata information updated during the data synchronization from the remote site, and optimizes the storage pool tiering state of the local site according to the updated popularity metadata information.
[0087] Specifically, if Figure 5 The figure shows an architecture diagram of a remote replication tiered system, which is designed to improve data availability and performance by performing remote replication and tiered storage of data between local and remote sites: Local site (source end) and remote site: represent the two main nodes in the system, used to achieve data redundancy and backup; Network connection: connects the local site and the remote site to ensure that data can be transmitted between the two; The remote replication relationship configuration, remote replication tiering policy, and replication data tiering state consistency modules are responsible for defining and managing remote replication relationships and policies, as well as ensuring data state consistency between different tiers. Incremental synchronization and hierarchical information synchronization, hierarchical data replication resource scheduling ensure that only changed data is synchronized, and optimize resource scheduling to improve efficiency; IO Monitor, Remote Replication QoS Module, Remote Replication Tiered Metadata Manager: monitor I / O operations, ensure quality of service (QoS), and manage hot metadata related to remote replication; DPA heat analysis, remote replication data synchronization module, and layered heat synchronizer: analyze data access heat, synchronize data, and maintain consistency of data heat between different layers; DMP migration analysis engine and local data migration executor: analyze data migration requirements and perform actual data migration operations; Storage media pool NVME-SSD, storage media pool SSD, and storage media pool HDD: These represent different types of storage media, ranging from high-speed NVMe SSDs to slower HDDs, forming a tiered storage architecture to balance performance and cost. This remote replication tiered system achieves efficient data replication and storage by establishing complex management and control mechanisms between local and remote sites, which not only improves data security and reliability, but also optimizes system performance and cost-effectiveness through tiered storage.
[0088] It can be understood that the local tiering engine is used to monitor data access and allocate data blocks to local storage media of different performance levels; the remote tiering engine is used to store or migrate the received data blocks to the corresponding local tier storage media based on the received tier identification information.
[0089] The beneficial effects of the technical solution provided by the embodiments of the present application are: This application can solve the problem of low data synchronization efficiency caused by the limited synchronization capability of local sites and the expensive wide area network bandwidth consumed by the synchronization of a large amount of cold spot data.
[0090] The technical solution of this application can significantly improve the performance of remote sites: ensure that key business data is located in the high-speed storage layer after data synchronization to maintain application performance; optimize remote site storage costs: avoid cold data occupying expensive high-speed storage resources; efficiently utilize wide area network bandwidth: prioritize the replication of hot data, compress, limit or selectively replicate cold data, and save channel bandwidth costs; accelerate remote tiering optimization: utilize local data heat information to enable the remote site tiering engine to reach the optimal storage state faster; improve fault switching / data synchronization quality: reduce performance fluctuations after data synchronization through tiered preheating; improve the efficiency of recovery after data synchronization interruption: tiered information synchronization accelerates data layout optimization after incremental data synchronization.
[0091] For the description of the features in the embodiment corresponding to the data synchronization system, please refer to the relevant description of the embodiment corresponding to the data synchronization method, and will not be repeated here.
[0092] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps of an embodiment of a data synchronization method, the method comprising: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0093] like Figure 6 As shown, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps of the data synchronization method embodiment when running, the method comprising: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0094] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0095] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the data synchronization method embodiment are implemented. The method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0096] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in the data synchronization method embodiment are implemented. The method includes: The local site receives and stores read and write request data; The local site calculates a first heat statistics index movement value based on the read and write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing based on a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on the second heat statistics index movement value and the full system heat analysis result of the first heat statistics index movement value. Data is synchronized between the local site and the remote site based on the priority of the requested data.
[0097] This application can solve the problem of low data synchronization efficiency caused by the limited synchronization capability of local sites and the expensive wide area network bandwidth consumed by the synchronization of a large amount of cold spot data.
[0098] The technical solution of this application can significantly improve the performance of remote sites: ensure that key business data is located in the high-speed storage layer after data synchronization to maintain application performance; optimize remote site storage costs: avoid cold data occupying expensive high-speed storage resources; efficiently utilize wide area network bandwidth: prioritize the replication of hot data, compress, limit or selectively replicate cold data, and save channel bandwidth costs; accelerate remote tiering optimization: utilize local data heat information to enable the remote site tiering engine to reach the optimal storage state faster; improve fault switching / data synchronization quality: reduce performance fluctuations after data synchronization through tiered preheating; improve the efficiency of recovery after data synchronization interruption: tiered information synchronization accelerates data layout optimization after incremental data synchronization.
[0099] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] The above is a detailed introduction to a data synchronization method, system, device and medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only applicable to help understand the method of the present application and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A data synchronization method, characterized in that: The method is applied to a data synchronization system, the system including a local site and a remote site, the local site including a local tiering engine, the remote site including a remote tiering engine, and the method including: The local site receives storage read and write request data; The local site calculates a first heat statistics index movement value according to the read / write request data, and sends the first heat statistics index movement value to the remote site; the local tiering engine performs local tier migration processing according to a heat analysis result of the first heat statistics index movement value; The local site determines whether the type of the requested data is write data. In response to the type of the requested data being write data, the local site sends the write data to the remote site. The remote site calculates a second heat statistics index movement value based on the received write data. The remote tiering engine performs remote tiering migration processing based on a system-wide heat analysis result of the second heat statistics index movement value and the first heat statistics index movement value. Data synchronization is performed between the local site and the remote site according to the priority identifier of the requested data.
2. The data synchronization method according to claim 1, characterized in that: The local site includes a local storage pool, the local storage pool includes a first storage tier of the local site and a second storage tier of the local site, and the local tiering engine performs local tier migration processing according to a heat analysis result of the first heat statistical index movement value, including: The local tiering engine determines whether the local site performs local tiering migration according to the heat analysis result and the first period threshold; If yes, then the read / write request data is migrated to the corresponding storage tier according to the heat analysis result; if no, then the current tier status of the storage is maintained; Migrating the read / write request data to a corresponding storage tier according to the heat analysis result includes: When the first heat statistical index movement value of the read / write request data is greater than a first preset value, the read / write request data is determined to be hot business data, and the hot business data is migrated to the first storage tier of the local site; When the first heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold spot business data, and the cold spot business data is migrated to the second storage tier of the local site.
3. The data synchronization method according to claim 1, wherein: The remote site includes a remote storage pool, the remote storage pool includes a first storage tier at the remote site and a second storage tier at the remote site, and the remote tiering engine performs remote tiering migration processing according to a system-wide heat analysis result of the second heat statistical index movement value and the first heat statistical index movement value, including: The remote tiering engine determines whether the remote site performs remote tiering migration according to the first period threshold based on the system-wide heat analysis result; If yes, then the read / write request data is migrated to the corresponding storage tier according to the system-wide heat analysis result; if no, then the current tier status of the storage is maintained; Migrating the read and write request data to a corresponding storage tier according to the system-wide heat analysis result includes: When the heat statistical index movement value of the read / write request data is greater than a first preset value, determining that the read / write request data is hot business data; migrating the hot business data to the first storage tier of the remote site; When the heat statistical index movement value of the read / write request data is less than a second preset value, the read / write request data is determined to be cold-spot business data; and the cold-spot business data is migrated to a second storage tier at a remote site.
4. The data synchronization method according to claim 1, wherein: The calculating a first heat statistical index moving value according to the read / write request data includes: Obtain the exponential moving value EMA(old) of the read / write request data at the previous moment, the new read / write request data point NewValue at the current moment, and the smoothing coefficient alpha; wherein the smoothing coefficient is a value between 0 and 1; Calculate the first heat statistical exponential moving value EMA(new) of the read / write request data at the current moment using the formula: EMA(new)=EMA(old)×(1-alpha)+NewValue×alpha; The calculating a second heat statistics index moving value according to the write data includes: Get the exponential moving value EMA(old)^ of the heat statistics of the data written at the previous moment and the new data point NewValue^ at the current moment; The second heat statistical exponential moving value EMA(new)^ of the data written at the current moment is calculated using the formula: EMA(new)^=EMA(old)^×(1-alpha)+NewValue^×alpha.
5. The data synchronization method according to claim 2, characterized in that: The synchronizing data between the local site and the remote site according to the priority identifier of the requested data includes: Setting the hotspot service data as high priority data and setting the coldspot service data as low priority data; When the requested data is high-priority data, data synchronization is performed first, and the bandwidth of the transmission channel between the local site and the remote site is widened; When the requested data is low priority data, data synchronization is delayed and the bandwidth of the transmission channel between the local site and the remote site is limited.
6. The data synchronization method according to claim 5, characterized in that: When the requested data is high priority data, data synchronization is performed first; When the requested data is low priority data, data synchronization is delayed, including: The local site sets a corresponding priority tag for each request data; Get the priority identifier of the current request data; Determining the storage level of the requested data at the local site according to the priority identifier of the requested data; Set corresponding preset thresholds for each storage tier at the local site; Determine whether the resource ratio of the level stored at the local site corresponding to the requested data reaches a preset threshold; If so, the request data is delayed for synchronization to the remote site; if not, the request data is synchronized between the local site and the remote site.
7. The data synchronization method according to claim 1, characterized in that: Before synchronizing data between the local site and the remote site according to the priority mark of the requested data, the method includes: performing a tiered warm-up operation at the local site and the remote site; The performing of the hierarchical preheating operation at the local site and the remote site includes: Identify hot business data in the corresponding storage pool according to the popularity metadata information of the last data synchronization of the local site or the popularity metadata information accumulated by the remote site, and store the hot business data from the second storage tier of the storage pool to the first storage tier; The popularity metadata information includes data access frequency score, data access timestamp, and performance label of business data.
8. The data synchronization method according to claim 5, characterized in that: The method comprises: Performing data compression and deduplication processing on the low-priority data; The compressing and deduplicating the low-priority data includes: Compressing the low-priority data using a compression algorithm, wherein the compression algorithm includes a lossless data compression algorithm; Obtaining a hash value of the low-priority data and creating a data storage index; Comparing the hash value of the low-priority data with the data storage index; When the hash value of the low-priority data exists in the data storage index, deduplication processing is performed on the low-priority data; In response to data synchronization between the local site and the remote site, synchronizing the difference data blocks and the data storage tier identifier and the heat metadata information of the difference data blocks updated during the interruption of data synchronization; The data storage tier identifier includes: a storage tier identifier allocated by a data storage site, a data heat statistical index movement value, and a data logical unit number identifier.
9. The data synchronization method according to claim 1, wherein: The method further comprises: Performing reverse data synchronization between the local site and the remote site; The reverse data synchronization between the local site and the remote site includes: Obtain the data blocks that changed during the last data synchronization of the remote site; Synchronize the data blocks that changed during the last data synchronization of the remote site to the local site through incremental synchronization; Verify the integrity and correctness of data blocks after reverse data synchronization; In response to executing reverse data synchronization, the local site obtains the popularity metadata information updated during the data synchronization from the remote site, and optimizes the storage pool tiering state of the local site according to the updated popularity metadata information.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data synchronization method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Remote data classification storage method, electronic device and storage medium
CN107608627A
Data migration method and device
CN113741810A
Cold and hot data monitoring method and device in storage system, terminal and medium
CN115629708A
Cluster space automatic management method based on data cold and hot separation
CN116450041A
Data-tiering service with multiple cold tier quality of service levels
US10579597B1