Dynamic adaptive memory and hard disk hybrid writing method and device

By dynamically adapting the memory and hard disk hybrid writing method and utilizing spatiotemporal feature vectors and data heat analysis, efficient management of cloud computing platform memory resources is achieved, solving the problems of insufficient memory resources and single point failures, and improving system performance and reliability.

CN120780468APending Publication Date: 2025-10-14JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510854182.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing cloud computing platforms have insufficient memory resource allocation in multi-user, multi-tasking concurrent environments, and traditional static allocation mechanisms cannot adapt to dynamic loads, resulting in system performance degradation and single point failure risks, affecting system availability and reliability.

Method used

A dynamic adaptive memory and hard disk hybrid writing method is proposed. The data heat value is calculated by generating spatiotemporal feature vectors, and the data is divided into four levels of queues: super hot, hot, warm, and cold. The dynamic migration and hybrid writing of data are realized by combining physical memory usage, system idle bandwidth, and queue response delay.

Benefits of technology

It optimizes the hybrid write performance between physical memory and hard disk, improves system operation efficiency and resource utilization, ensures the business continuity of virtual machines in extreme scenarios, and reduces data migration bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780468A_ABST
    Figure CN120780468A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic adaptive memory and hard disk hybrid writing method and device. The method comprises the following steps: generating a spatio-temporal feature vector based on a memory access sequence; obtaining a popularity calculation formula according to the spatial-temporal feature vector, and calculating a data popularity value of the data memory by using the popularity calculation formula; obtaining a data popularity probability according to the data popularity value of the data memory; and dividing the memory data into four levels of superheat, heat, temperature and cold according to the data popularity probability, obtaining a comprehensive utilization rate by integrating three factors of a physical memory utilization rate, a system idle bandwidth and queue response delay, and executing a corresponding migration strategy according to the comprehensive utilization rate. Through real-time data popularity analysis, a difference bit pre-migration strategy and a dynamic priority adjustment algorithm, efficient identification and scheduling of cold and hot data of the virtual memory are realized, and the hybrid writing performance between a physical memory and a hard disk is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing, and in particular to a method and device for dynamically adapting memory and hard disk hybrid writing. Background Art

[0002] With the widespread adoption of cloud computing, more and more enterprise users rely on virtualization platforms for business deployment and resource management, placing higher demands on platform stability and performance. In environments with multiple users and multiple tasks running concurrently, physical memory resources are becoming increasingly scarce, and traditional static allocation mechanisms are no longer able to meet dynamically changing memory demands. To alleviate this problem, cloud management platforms typically employ extended memory mechanisms (such as virtual memory or swap space) to automatically utilize extended memory to maintain normal virtual machine operations when insufficient physical memory is detected. However, current memory scheduling strategies suffer from significant flaws, severely hindering system performance and user experience. First, existing solutions generally rely on hot and cold data classification methods based on fixed thresholds (e.g., based on access counts). This approach is inadequate for dynamic load scenarios, particularly under bursty traffic or short-term high concurrent requests. Second, traditional mechanisms employ a full migration strategy during data migration, failing to fully consider data heterogeneity, resulting in unnecessary data movement and increased system overhead and latency. Furthermore, when a large number of virtual machines simultaneously utilize extended memory, the lack of an effective scheduling optimization mechanism can lead to frequent system stalls, significantly degrading overall performance and severely impacting service quality. Finally, the swap storage used to implement memory overcommit in current cloud platforms typically relies on a shared storage pool. If this storage fails, the entire system's memory overcommit functionality will be ineffective, creating a single point of failure risk and reducing system availability and reliability. Therefore, a new write technology is urgently needed. Summary of the Invention

[0003] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0004] The present invention proposes a method for dynamically adapting memory and hard disk hybrid writing, which writes metadata and high-frequency access data into physical memory; and directly writes large block data and persistent data into the hard disk.

[0005] Another object of the present invention is to provide a device for dynamically adapting memory and hard disk hybrid writing.

[0006] To achieve the above objectives, the present invention provides a method for dynamically adapting memory and hard disk hybrid writing, comprising:

[0007] Generate spatiotemporal feature vectors based on memory access sequences;

[0008] According to the space-time feature vector, a hotness calculation formula is obtained, and the data memory data hotness value is calculated by using the hotness calculation formula;

[0009] According to the data memory data hotness value, a data hotness probability is obtained.

[0010] According to the data hotness probability, the memory data is divided into four levels of super-hot, hot, warm and cold, and is bound to different virtual queues respectively.

[0011] The comprehensive utilization rate is obtained by comprehensively considering the physical memory utilization rate, the system idle bandwidth and the queue response delay, and the corresponding migration strategy is executed according to the comprehensive utilization rate; based on the migration strategy and the hash fingerprint of the memory address, the difference bits of the cold data block and the hard disk target area are compared in real time, and only the difference bit metadata is extracted to generate compression encoding for data migration.

[0012] The dynamic adaptive memory and hard disk hybrid writing method of the embodiment of the application can also have the following additional technical features:

[0013] In an embodiment of the application, the space-time feature vector includes frequency, access distance, re-access frequency and service priority.

[0014] In an embodiment of the application, the hotness calculation formula is:

[0015] A = α * frequency + β * access distance + γ * re-access frequency + δ * service priority

[0016] Wherein, α, β, γ, δ are adjustable coefficients.

[0017] In an embodiment of the application, the data classification binding module is also used for:

[0018] The super-hot 1st queue: the data hotness is greater than 95%, the bound queue type is an exclusive high-speed queue, the migration strategy is to prohibit migration, and the resources of other queues are allowed to be preempted; if the response delay of the super-hot 1st queue exceeds 2us, the low-priority queue resources are immediately preempted, and the cache capacity is dynamically expanded.

[0019] The hot 2nd queue: the data hotness is greater than 85% to 95%, the high-priority queue is shared, and the migration strategy is to move locally within the memory only.

[0020] The warm 3rd queue: the data hotness is greater than 60% to 85%, it is an elastic buffer queue, and the migration strategy is to migrate to the Hot queue as needed in reverse;

[0021] The cold 4th queue: the data hotness is greater than <60%, it is a pre-migration queue, and the migration strategy is to migrate asynchronously after difference bit compression.

[0022] In an embodiment of the application, the comprehensive utilization rate is:

[0023] The comprehensive utilization rate = the host physical memory utilization rate * 40% + the bandwidth utilization rate * 30% + the queue delay rate * 30%.

[0024] In an embodiment of the present application, the migration strategy comprises:

[0025] When the comprehensive utilization rate > 80%, the cold level data is forced to migrate to the hard disk, and the pre-migration is started asynchronously to avoid write peak congestion;

[0026] When the comprehensive utilization rate < 30%, the high-frequency accessed warm 3 level queue data is allowed to migrate reversely to the hot 2 level queue.

[0027] When the super-hot 1 level queue delay is detected to be out of limit, the cache block is automatically recycled from the warm 3 level queue; and the physical address mapping range of the super-hot 1 level queue is expanded.

[0028] In an embodiment of the present application, when the virtual machine a is located on the storage pool D, the storage pool D is mounted by the host ABC, the host ABC is configured with the swap file storage D, when the host A to which the virtual machine belongs is disconnected from the storage pool D, the cold data on the virtual machine a is migrated to the swap storage space belonging to B or C on the storage pool D across hosts;

[0029] Alternatively, the swap storage space on the host A has a utilization rate higher than 90%, and the host B or C has a lower swap storage pool utilization rate, the cold data on the virtual machine a is migrated to the host B or the host C across hosts.

[0030] To achieve the above purpose, another aspect of the present application provides a dynamic adaptive memory and hard disk hybrid writing device, comprising:

[0031] A memory monitoring and collecting module is configured to collect a memory access sequence and generate a space-time feature vector; wherein the space-time feature vector comprises frequency, access distance, re-access frequency and service priority.

[0032] A hotness calculation module is configured to obtain a hotness calculation formula according to the space-time feature vector, and calculate a data memory data hotness value by using the hotness calculation formula.

[0033] A hotness reasoning module is configured to obtain a data hotness probability according to the data memory data hotness value.

[0034] A data classification and binding module is configured to divide the memory data into four levels of super-hot, hot, warm and cold according to the data hotness probability, and bind the memory data to different virtual queues respectively.

[0035] The data migration module is used to obtain a comprehensive utilization rate based on three factors: physical memory utilization, system idle bandwidth, and queue response delay, and to execute a corresponding migration strategy based on the comprehensive utilization rate. Based on the migration strategy and the hash fingerprint of the memory address, the module compares the difference bits between the cold data block and the target area of ​​the hard disk in real time, extracts only the difference bit metadata, and generates a compressed code for data migration.

[0036] The dynamic adaptive memory and hard disk hybrid writing method and device of the embodiment of the present invention realizes efficient identification and scheduling of hot and cold data in virtual memory through real-time data heat analysis, differential bit pre-migration strategy and dynamic priority adjustment algorithm, and optimizes the hybrid writing performance between physical memory and hard disk.

[0037] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0039] Figure 1 is a flowchart of a method for dynamically adapting memory and hard disk hybrid writing according to an embodiment of the present invention;

[0040] Figure 2 2 is a structural diagram of a device for dynamically adapting memory and hard disk hybrid writing according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0042] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0043] The following describes a method and apparatus for dynamically adapting memory and hard disk hybrid writing according to an embodiment of the present invention with reference to the accompanying drawings.

[0044] Figure 1 Flowchart of a method for dynamically adapting memory and hard disk hybrid writing according to an embodiment of the present invention. Figure 1 Shown, including:

[0045] S1, generates spatiotemporal feature vectors based on memory access sequences.

[0046] Specifically, the monitoring module collects memory access sequences and generates a spatiotemporal feature vector, including: frequency (number of accesses in the past 1ms window), access distance (data in adjacent storage locations may be accessed continuously), re-access frequency (number of times the recently accessed data is accessed again in a short period of time), and business priority (whether it is a critical business).

[0047] The monitoring module collects memory access sequences in the system in real time, performs fine-grained analysis of memory access behavior, and extracts representative spatiotemporal feature vectors for subsequent behavior modeling, anomaly detection, or resource scheduling optimization. These features include but are not limited to the following key dimensions:

[0048] Access frequency: This metric reflects the number of accesses to a memory area within a specific time window (e.g., the past 1 millisecond). High access frequency may indicate that the memory area stores hot data or frequently called code segments, helping to identify performance bottlenecks or predict future access trends.

[0049] Access distance: Access distance refers to the offset between the current access address and the previous access address, and is used to measure the spatial locality of memory accesses. If adjacent or close addresses are accessed consecutively, it indicates good spatial locality in the program, which can provide a basis for cache management and prefetching strategies. Conversely, it may indicate a random access pattern, requiring different optimization methods.

[0050] Reaccess frequency: This feature describes the number of times a memory address is accessed again within a short period of time (e.g., milliseconds) after being accessed. This metric helps determine which data exhibits "temporary reuse" characteristics, thus supporting smarter cache replacement mechanisms or memory hierarchical management strategies.

[0051] Service Priority: This dimension identifies whether the service currently accessing memory is mission-critical or a high-priority service. For example, critical services such as database transaction processing and real-time control systems typically require higher response speeds and stability, so their memory access behavior warrants special attention. This information can be obtained through process tags, thread priorities, or Quality of Service (QoS) classes, serving as a crucial basis for resource scheduling and assurance.

[0052] By integrating the information from the above multiple dimensions into a structured spatiotemporal feature vector, the monitoring module can comprehensively model memory access behavior, which can not only be used to identify potential abnormal behaviors (such as memory leaks, malicious attacks, etc.), but also provide data support for dynamic resource allocation, memory optimization, performance tuning, etc., thereby improving the overall operating efficiency and security of the system.

[0053] S2, obtaining a heat calculation formula according to the spatiotemporal feature vector, and calculating the heat value of the data in the data memory using the heat calculation formula.

[0054] Specifically, the heat calculation formula A = α*frequency + β*access distance + γ*re-access frequency + δ*business priority is obtained through the monitoring module in S1. Users can dynamically adjust the desired ratio according to their needs to obtain the data memory data heat value.

[0055] Based on the memory access sequences collected by the monitoring module deployed in S1, the system can further utilize a set of weighted calculation models to comprehensively assess the "data hotness" of each memory region. This hotness value reflects the activity and importance of memory data in the current operating environment and can be used to guide key performance tuning operations such as cache management, resource scheduling, and prefetch optimization.

[0056] This formula fully considers the temporal locality of memory accesses (reflected by frequency and re-access frequency), spatial locality (reflected by access distance), and the importance of business logic (reflected by business priority) to form a comprehensive, multi-dimensional evaluation system. Furthermore, the system supports a dynamic configuration mechanism, allowing administrators or automated policies to adjust various weighting parameters in real time based on changes in the runtime environment. For example, in high-concurrency scenarios, the weights of frequency and re-access frequency can be appropriately increased to prioritize hotspot data. In latency-sensitive systems, the weight of business priority can be increased to ensure high-speed access to mission-critical data. This flexible heat calculation mechanism enables the system to more accurately identify memory resource usage patterns, assisting in efficient memory tiering management, hotspot data migration, and cache replacement decisions, significantly improving overall system performance and resource utilization. This configurable model also provides excellent adaptability and scalability for diverse application scenarios, making it suitable for a wide range of complex environments, from embedded systems to large-scale data centers.

[0057] S3, obtains the data heat probability according to the data heat value in the data memory.

[0058] Specifically, in recommendation systems or data analysis scenarios, the data popularity inference module is typically used to evaluate the popularity of different content or items and convert it into a quantifiable probability distribution, thereby providing a basis for subsequent sorting, recommendation, or scheduling. The core goal of this module is to calculate the "popularity probability" of each data item through a comprehensive analysis of historical user behavior data (such as clicks, views, likes, etc.), time decay factors, contextual features, and other factors.

[0059] It can be understood that in typical application scenarios such as recommendation systems or data analysis, the data heat reasoning module plays a key role. This module is mainly used to evaluate the popularity of different content, items or resources among user groups and abstract it into a quantifiable and comparable "heat probability" indicator. This heat probability not only reflects the current trend of the data item, but also serves as an important basis for subsequent sorting, recommendation, scheduling and resource allocation. The core goal of this module is to build a dynamic heat model by comprehensively analyzing multi-dimensional behavior and context information, so as to realize accurate prediction of the real-time heat state of data items. Its input usually includes the following aspects:

[0060] User behavior data: This is the main basis for heat calculation, including but not limited to user click times, browsing time, likes, collections, comments, shares, purchases and other interaction behaviors. These behaviors can effectively reflect the user's interest intensity and participation in a certain content.

[0061] Time decay factor: In order to reflect the trend of data heat over time, the system introduces a time decay mechanism. That is, the influence of earlier behaviors on current heat will gradually weaken over time.

[0062] Contextual features: In addition to basic user behavior, heat reasoning also considers contextual information such as device type, geographic location, time period (such as weekdays / holidays, daytime / nighttime), network environment, etc. These factors can significantly affect user behavior preferences and content dissemination efficiency.

[0063] Content attribute features: including item category, label, publication time, author influence and other metadata information. These information helps to identify potential high-heat content and improve recommendation accuracy in cold start scenarios. Based on the above information, the heat reasoning module uses statistical modeling or machine learning methods to convert raw behavior data into a normalized heat probability distribution. For example, you can use logistic regression, collaborative filtering, deep interest network (DIN) and other algorithms to model the data and output the probability value of each candidate content being focused or clicked by users in the future period of time.

[0064] Exemplarily, first, collect and clean the relevant behavior data, including the number of interactions and timestamps of each item in the past period of time; second, introduce a time decay function to weight the historical behavior, so that recent behavior has higher influence; third, input the weighted behavior data into the heat reasoning module, and generate the heat score of each item through statistical methods (such as Softmax normalization) or machine learning models (such as logistic regression, deep ranking model, etc.); finally, normalize these scores into a probability distribution, i.e. get the data heat probability of each item.

[0065] For example, in a short video recommendation system, videos with high heat probability will have higher display weight in the homepage recommendation list. This way not only improves user experience, but also enhances the efficiency and intelligence level of platform content distribution.

[0066] S4, according to the data heat probability, the memory data is divided into four levels of super-hot, hot, warm and cold, and is respectively bound to different virtual queues.

[0067] Specifically, the super-hot 1st queue: the data heat is greater than 95%, the bound queue type is exclusive high-speed queue, the migration strategy is prohibited migration, and the other queue resources are allowed to be preempted. If the response delay of the super-hot 1st queue exceeds 2us, the low priority queue resources are immediately preempted, and the cache capacity is dynamically expanded.

[0068] Hot 2nd queue: the data heat is greater than 85% to 95%, shared high priority queue, migration strategy is only local migration within memory;

[0069] Warm 3rd queue: the data heat is greater than 60% to 85%, elastic buffer queue, migration strategy is on-demand reverse migration to Hot queue;

[0070] Cold 4th queue: the data heat is greater than <60%, pre-migration queue, migration strategy is asynchronous migration after differential bit compression;

[0071] The four-level definition and queue binding strategy is shown in Table 1:

[0072] Table 1

[0073]

[0074] S5, the comprehensive physical memory usage rate, system idle bandwidth and queue response delay are combined to obtain a comprehensive usage rate, and a corresponding migration strategy is executed according to the comprehensive usage rate; based on the migration strategy and the hash fingerprint of the memory address, the difference bits of the cold data block and the hard disk target area are compared in real time, only the difference bit metadata is extracted to generate compression encoding for data migration.

[0075] Specifically, the comprehensive physical memory usage rate, system idle bandwidth, and queue response delay are combined to obtain a migration mechanism: comprehensive usage rate = host physical memory usage rate * 40% + bandwidth usage rate * 30% + queue delay rate (probability greater than 1us, number of times greater than 1us in 100 transmissions / 100 times) * 30%,

[0076] When the comprehensive usage rate is greater than 80%, the Cold level data is forcibly migrated to the hard disk, and the pre-migration is started asynchronously to avoid write peak congestion;

[0077] When the comprehensive usage rate is less than 30%, the Warm 3rd queue data with high frequency access is allowed to be reverse migrated to the Hot 2nd queue.

[0078] When the super-hot 1st queue queue delay is detected to be out of limit (e.g., > 1us), the cache block is automatically reclaimed from the warm 3rd queue; and the physical address mapping range of the super-hot 1st queue queue is extended.

[0079] According to the above migration strategy, migration is performed, and based on the hash fingerprint of the memory address, the difference bits of the cold data block and the hard disk target area are compared in real time. Only the difference bit metadata is extracted to generate compressed encoding, thereby reducing the migration data amount (reducing 60% bandwidth occupation compared with full migration). After the migration is completed, the hard disk address mapping table is updated, and the original memory area is released.

[0080] Exemplarily, in the current cloud management platform, when the memory swap partition is a shared storage pool and is mounted by multiple hosts, the virtual machine cannot use the swap partition when the storage fails. The current memory data migration mode is suitable for the swap partition failure scenario. The cold data can be migrated to the swap partition of another host mounted by the current storage, thereby ensuring the business continuity in the extreme scenario. For example, the virtual machine a is located on the storage pool D, the storage pool D is mounted by the hosts ABC, and the hosts ABC are configured with the swap file storage D. When the host A to which the virtual machine belongs is disconnected from the storage pool D, the cold data on the virtual machine a can be migrated to the swap storage space belonging to B or C on the storage pool D, or the usage rate of the swap storage space on the host A is higher than 90%, and the usage rate of the swap storage pool of the host B or C is relatively low. Then, the cold data on the virtual machine a can be migrated to the host B or the host C, thereby ensuring the business continuity in the extreme scenario.

[0081] The dynamic adaptive memory and hard disk hybrid writing method of the embodiment of the application writes metadata and high-frequency access data into the physical memory, and writes large block data and persistent data directly into the hard disk. Real-time data heat analysis, difference bit pre-migration strategy and dynamic priority adjustment algorithm are used to realize efficient identification and scheduling of cold and hot data of the virtual memory, and to optimize the hybrid writing performance between the physical memory and the hard disk. Meanwhile, the data migration mode can ensure the continuity of data transmission in the extreme scenario. Bandwidth is saved, migration efficiency is improved, the performance of the virtual machine is ensured, and the utilization rate of the swap partition resource in the cloud platform is also considered.

[0082] To realize the above embodiment, as shown in Figure 2 The embodiment further provides a dynamic adaptive memory and hard disk hybrid writing device 10, which comprises:

[0083] The memory monitoring and collecting module 100 is configured to collect the memory access sequence and generate a space-time feature vector. The space-time feature vector comprises frequency, access distance, re-access frequency and service priority.

[0084] The heat calculation module 200 is configured to obtain a heat calculation formula according to the space-time feature vector, and calculate a data memory data heat value by using the heat calculation formula.

[0085] The heat reasoning module 300 is configured to obtain a data heat probability according to the data memory data heat value.

[0086] The data classification binding module 400 is configured to divide the memory data into four levels of super-hot, hot, warm and cold according to the data heat probability, and bind the memory data to different virtual queues respectively.

[0087] The data migration module 500 is configured to obtain a comprehensive utilization rate by comprehensively considering three factors of a physical memory utilization rate, a system idle bandwidth and a queue response delay, and execute a corresponding migration strategy according to the comprehensive utilization rate; based on the migration strategy and a hash fingerprint of a memory address, difference bits of a cold data block and a hard disk target area are compared in real time, only difference bit metadata is extracted to generate compression encoding, and data migration is performed.

[0088] Specifically, the memory monitoring and collecting module: the monitoring module collects a memory access sequence, and generates a space-time feature vector, including: frequency (the number of access times in the past 1ms window), access distance (data of adjacent storage positions may be continuously accessed), re-access frequency (the number of times of re-accessing the data recently accessed in a short period), and service priority (whether it is a critical service). The heat calculation module: the heat calculation formula A = a * frequency + β * access distance + γ * re-access frequency + δ * service priority is obtained through the monitoring module in 1, and a user can dynamically adjust the proportion according to the user's own needs to obtain a data memory data heat value. The heat reasoning module: the data heat probability is obtained by using the data heat reasoning module. The data classification binding module: according to the prediction result, the memory data is divided into four levels of super-hot, hot, warm and cold, and is bound to different virtual queues respectively.

[0089] The super-hot 1st level queue: the data heat is greater than 95%, the queue type bound is an exclusive high-speed queue, the migration strategy is prohibition of migration, and the resource of other queues is allowed to be preempted. If the response delay of the super-hot 1st level queue exceeds 2us, the low-priority queue resource is immediately preempted, and the cache capacity is dynamically expanded.

[0090] The hot 2nd level queue: the data heat is greater than 85% to 95%, the high-priority queue is shared, and the migration strategy is only local migration in the memory.

[0091] The warm 3rd level queue: the data heat is greater than 60% to 85%, it is an elastic buffer queue, and the migration strategy is reverse migration to the hot queue on demand.

[0092] The cold 4th level queue: the data heat is greater than 60%, it is a pre-migration queue, and the migration strategy is asynchronous migration after difference bit compression.

[0093] 5. Data migration module: comprehensive physical memory usage, system free bandwidth, queue corresponding delay, etc. Three factors are obtained to get the migration mechanism:

[0094] Comprehensive utilization rate = host physical memory usage * 40% + bandwidth usage * 30% + queue delay rate (probability greater than 1us, number of times greater than 1us in 100 times / 100 times) * 30%;

[0095] When the comprehensive utilization rate is greater than 80%, the cold level data is forced to migrate to the hard disk, and the pre-migration is started asynchronously to avoid write peak congestion;

[0096] When the comprehensive utilization rate is less than 30%, the reverse migration of the high-frequency access warm 3 queue data to the hot 2 queue is allowed.

[0097] When the super-hot 1 queue queue delay is detected to be out of limit (such as >1us), the cache block is automatically recycled from the warm 3 queue; the physical address mapping range of the super-hot 1 queue queue is expanded.

[0098] According to the above migration strategy, the migration is performed, and based on the hash fingerprint of the memory address, the difference bits of the cold data block and the hard disk target area are compared in real time. Only the difference bit metadata is extracted to generate compression encoding, reducing the migration data amount (reducing 60% bandwidth occupation compared with full migration). After the migration is completed, the hard disk address mapping table is updated, and the original memory area is released.

[0099] In the current cloud management platform, when the memory swap partition is a shared storage pool, it is mounted by multiple hosts, and when the storage fails, the virtual machine cannot use the swap partition. The current memory data migration method is suitable for swap partition failure scenarios. Cold data can be migrated across hosts to the swap partition of another host mounted by the current storage, ensuring business continuity in extreme scenarios. For example, virtual machine a is located on storage pool D, storage pool D is mounted by host ABC, and host ABC is configured with swap file storage D. When the host A of the virtual machine disconnects the connection with the storage pool D, the cold data on the virtual machine a can be migrated across the host to the swap storage space belonging to B or C on the storage pool D, or the swap storage space usage rate on host A is higher than 90%, and the swap storage pool usage rate of host B or C is lower. Then the cold data on the virtual machine a can be migrated across the host to the host B or the host C, ensuring business continuity in extreme scenarios.

[0100] According to an embodiment of the present invention, the device for dynamically adapting memory and hard disk hybrid writing writes metadata and frequently accessed data into physical memory; large blocks of data and persistent data are written directly into the hard disk. Real-time data heat analysis, differential bit pre-migration strategy, and dynamic priority adjustment algorithm enable efficient identification and scheduling of hot and cold data in virtual memory, optimizing the hybrid writing performance between physical memory and the hard disk. At the same time, this data migration method can ensure the continuity of data transmission in extreme scenarios. It saves bandwidth, improves migration efficiency, and ensures the performance of virtual machine operations, while also taking into account the utilization of swap partition resources in the cloud platform.

[0101] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0102] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

Claims

1. A method for dynamically adapting memory and hard disk hybrid writing, characterized in that: include: Generate spatiotemporal feature vectors based on memory access sequences; Obtaining a heat calculation formula according to the spatiotemporal feature vector, and calculating the heat value of the data in the data memory using the heat calculation formula; Get the data heat probability based on the data heat value in the data memory; Divide the memory data into four levels: super hot, hot, warm, and cold according to the data heat probability, and bind them to different virtual queues respectively; A comprehensive utilization rate is obtained by integrating three factors: physical memory utilization rate, system idle bandwidth, and queue response delay, and a corresponding migration strategy is executed according to the comprehensive utilization rate; Based on the migration strategy and the hash fingerprint of the memory address, the difference bits between the cold data block and the target area of ​​the hard disk are compared in real time. Only the difference bit metadata is extracted to generate a compressed code for data migration.

2. The method according to claim 1, characterized in that Spatiotemporal feature vectors, including: frequency, access distance, revisit frequency, and service priority.

3. The method according to claim 1, characterized in that The heat calculation formula is: A=α*frequency+β*access distance+γ*revisit frequency+δ*service priority Among them, α, β, γ, and δ are adjustable coefficients.

4. The method according to claim 1, wherein The data classification binding module is also used for: Superhot Level 1 Queue: For data with a popularity greater than 95%, the bound queue type is an exclusive high-speed queue, the migration policy is to prohibit migration, and preemption of other queue resources is allowed. If the response delay of the Superhot Level 1 Queue exceeds 2μs, it immediately preempts the resources of lower-priority queues and dynamically expands the cache capacity. Hot Level 2 queue: Data with a popularity greater than 85% to 95% shares a high-priority queue and uses a migration strategy of local migration within memory only. Warm Level 3 queue: This queue has data popularity greater than 60% to 85%. It is an elastic buffer queue and its migration strategy is to migrate data back to the Hot queue on demand. Cold Level 4 queue: This queue has data heat greater than or less than 60%, is a pre-migration queue, and uses the asynchronous migration strategy after differential bit compression.

5. The method according to claim 4, characterized in that The comprehensive utilization rate is: Comprehensive utilization rate = host physical memory utilization rate * 40% + bandwidth utilization rate * 30% + queue delay rate * 30%.

6. The method according to claim 5, characterized in that Migration strategies, including: When the overall utilization rate exceeds 80%, Cold-level data is forcibly migrated to the hard disk and pre-migration is started asynchronously to avoid write peak congestion. When the combined usage rate is less than 30%, reverse migration of frequently accessed warm level 3 queue data to hot level 2 queue is allowed. When the queue delay of the super-hot level 1 queue exceeds the limit, the cache block is automatically reclaimed from the warm level 3 queue; the physical address mapping range of the super-hot level 1 queue is expanded.

7. The method according to claim 6, characterized in that If VM a is located in storage pool D, which is mounted on hosts ABC and configured with swap files on storage pool D, and host A, to which VM a belongs, loses connection to storage pool D, the cold data on VM a is migrated across hosts to the swap storage space on storage pool D belonging to hosts B or C. Alternatively, the swap storage space usage on host A is higher than 90%, and the swap storage pool usage on host B or C is lower. The cold data on virtual machine a is migrated across hosts to host B or host C.

8. A device for dynamically adapting memory and hard disk hybrid writing, characterized in that: include: A memory monitoring and acquisition module is used to collect memory access sequences and generate a spatiotemporal feature vector; wherein the spatiotemporal feature vector includes frequency, access distance, re-access frequency and service priority; A heat calculation module is used to obtain a heat calculation formula based on the spatiotemporal feature vector and calculate the heat value of the data in the data memory using the heat calculation formula; The heat inference module is used to obtain the data heat probability based on the data heat value in the data memory; A data classification and binding module is used to classify memory data into four levels: super hot, hot, warm, and cold according to the data heat probability, and bind them to different virtual queues respectively; The data migration module is used to obtain a comprehensive utilization rate based on three factors: physical memory utilization, system idle bandwidth, and queue response delay, and to execute a corresponding migration strategy based on the comprehensive utilization rate. Based on the migration strategy and the hash fingerprint of the memory address, the module compares the difference bits between the cold data block and the target area of ​​the hard disk in real time, extracts only the difference bit metadata, and generates a compressed code for data migration.

9. The device according to claim 8, characterized in that Spatiotemporal feature vectors, including: frequency, access distance, revisit frequency, and service priority.

10. The device according to claim 8, characterized in that The heat calculation formula is: A=α*frequency+β*access distance+γ*revisit frequency+δ*service priority Among them, α, β, γ, and δ are adjustable coefficients.

Citation Information

Cited By

  • Memory control method and memory storage device

    CN121412150A