Solid state disk cache management system and method based on HMB technology and host memory cooperation
By establishing an HMB collaborative mapping relationship between the solid-state drive cache and the host memory, resource allocation is dynamically adjusted and abnormal cache units are detected, which solves the problems of resource allocation lag and data transmission bottleneck in traditional cache management, and improves system performance and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINGRUI TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional SSD caching management strategies cannot dynamically adjust resource allocation, causing frequently accessed data to be stored in slower-responding areas, increasing data access latency. Furthermore, they fail to monitor the status of the hard drive and memory in real time, affecting system stability and reliability. The lack of adaptation assessment for HMB channels means that data transmission bottlenecks and potential problems are not detected in a timely manner.
By establishing an HMB collaborative mapping relationship between solid-state drive cache and host memory, collecting operational status data, performing hierarchical partitioning and cache unit detection, filtering abnormal cache units, matching and adapting resources, constructing an HMB channel transmission adaptation evaluation system, and optimizing cache management strategies.
It enables dynamic adjustment of cache tiers and resource allocation, improves data access speed, reduces latency, ensures system stability and reliability, avoids data transmission bottlenecks, and ensures balanced and efficient system operation.
Smart Images

Figure CN122044867A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of storage technology, specifically to a solid-state drive cache management system and method based on HMB technology and host memory collaboration. Background Technology
[0002] Solid State Drive (SSD) is a device used for data storage. SSDs are widely used in personal computers, servers, game consoles and various electronic devices. Due to their excellent performance and reliability, they have become an important part of modern computing devices.
[0003] Currently, traditional systems often employ static caching management strategies, which cannot dynamically adjust resource allocation based on real-time data access patterns and storage requirements. This results in frequently accessed data potentially being stored in slower-responding storage areas, thereby increasing data access latency. Furthermore, many traditional systems fail to implement real-time monitoring of the operating status of solid-state drives and host memory, making it impossible to promptly identify and handle faults when performance degradation or abnormal situations occur, thus affecting the stability and reliability of the system.
[0004] Furthermore, in traditional systems, the identification and replacement of abnormal cache units is often slow, making it difficult to manage normal and abnormal units quickly and effectively, leading to decreased system performance and increased risk of failure. Additionally, traditional cache management methods typically lack adaptation assessment for HMB channels, failing to allocate bandwidth resources reasonably according to different cache levels, easily causing data transmission bottlenecks and affecting overall data flow efficiency. Moreover, traditional systems often have relatively simple health status assessments, failing to comprehensively identify the health status of cache regions at each level, resulting in potential problems not being discovered and resolved in a timely manner, thus affecting the balance and efficiency of system operation. Summary of the Invention
[0005] To achieve the above objectives, the present invention provides the following technical solution: a solid-state drive cache management system based on HMB technology and host memory collaboration, comprising: The data acquisition module is used to establish the HMB collaborative mapping relationship between the solid-state drive cache and the host memory, and synchronously collect the running status of the solid-state drive cache and the available resource status of the host memory to obtain collaborative mapping-status fusion data. The data partitioning module is used to perform hierarchical partitioning of the solid-state drive cache based on the collaborative mapping-state fusion data to obtain multi-level cache regions, extract the basic association features of each level of cache region, and obtain a hierarchical cache-association feature dataset. The unit classification module is used to traverse the cache units in each level of cache region based on the hierarchical cache-association feature dataset, perform cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results. The resource matching module is used to filter core abnormal cache units and extract their belonging features based on the cache unit status classification results, and to match the appropriate idle cache resources in the host memory based on the collaborative mapping-state fusion data to obtain an abnormal replacement-resource matching list. The transmission adaptation module is used to collect HMB channel transmission status data, and based on the hierarchical cache-associative feature dataset and the abnormal replacement-resource matching list, construct an HMB channel transmission adaptation evaluation system to obtain channel transmission adaptation parameters. The storage management module is used to collect the operating temperature status of the solid-state drive core, and input the collaborative mapping-state fusion data, cache unit status classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy.
[0006] Preferably, the solid-state drive cache operating status includes cache occupancy status and cache read / write response status, and the host memory available resource status includes memory free capacity and memory data transfer bandwidth; Establish an HMB (Hardware Memory Management) collaborative mapping relationship between the SSD cache and host memory, synchronously collect the SSD cache running status and the host memory available resource status, and obtain collaborative mapping-status fusion data, including: Construct an HMB co-mapping table to record the mapping relationship between each region of the solid-state drive cache and the host memory address, and obtain basic co-mapping data; The data collection process includes: the relationship between cache occupancy status and the total SSD cache capacity; the real-time feedback latency of cache read / write response status; and the acquisition of cache operation status data. The data collection process involves collecting the ratio of free host memory capacity to total host memory capacity, and assessing the compatibility between host memory data transfer bandwidth and the maximum transmission bandwidth of the HMB channel to obtain memory resource status data. The basic collaborative mapping data, cache running status data, and memory resource status data are correlated and integrated to obtain collaborative mapping-state fusion data.
[0007] Preferably, based on the collaborative mapping-state fusion data, the solid-state drive cache is hierarchically divided to obtain multi-level cache regions. Basic association features of each level of cache region are extracted to obtain a hierarchical cache-association feature dataset, including: Extract historical access frequency, data access dependencies, and data validity duration characteristics of cached data from the collaborative mapping-state fusion data; Based on access frequency characteristics, the solid-state drive cache is divided into multi-level cache regions, which include a high-frequency access cache region, a medium-frequency access cache region, and a low-frequency access cache region. Extract the access frequency threshold, data association strength, and data timeliness range corresponding to the multi-level cache regions as the basic association features of each level; The division results of the multi-level cache region are associated with and stored with the corresponding basic associated features to obtain the hierarchical cache-associated feature dataset.
[0008] Preferably, based on the hierarchical cache-association feature dataset, the cache units within each level of the cache region are traversed to perform cache unit operation status detection and historical anomaly record tracing, distinguishing between normal and abnormal cache units, and obtaining cache unit status classification results, including: Based on the hierarchical cache-association feature dataset, determine the boundary range and cache unit distribution of each level of cache region, and construct the cache unit traversal path; By scanning each cache unit through the traversal path, read and write performance tests are performed on each cache unit to obtain real-time running status data; Retrieve historical exception cache unit records, construct a cache unit status record table, and integrate historical exception information with real-time running status data; Based on preset normal operation standards, the cache units are compared with real-time operation status data to determine whether each cache unit is a normal cache unit or an abnormal cache unit. The determination results are recorded in the status record table to obtain the cache unit status classification results.
[0009] Preferably, based on the cache unit state classification results, core abnormal cache units are screened and their attribution features are extracted. Furthermore, based on the collaborative mapping-state fusion data, suitable idle cache resources in host memory are matched to obtain an abnormal replacement-resource matching list, including: From the cache unit status classification results, select the abnormal cache units in the high-frequency access level cache area as the core abnormal cache units; Locate the hierarchical cache region to which the core exception cache unit belongs and the node positions within that region, and integrate the hierarchical cache region and the node positions within that region into the attribution characteristics of the core exception cache unit; Based on the HMB collaborative mapping relationship in the collaborative mapping-state fusion data, the host memory free cache resources corresponding to the attribution feature are matched, and at the same time, it is verified whether the capacity and transmission performance of the matched resources meet the replacement requirements of the core abnormal cache unit. By associating the resource information that meets the requirements with the corresponding core exception cache unit information, an exception replacement-resource matching list is obtained.
[0010] Preferably, HMB channel transmission status data is collected, and based on the hierarchical cache-associative feature dataset and the anomaly replacement-resource matching list, an HMB channel transmission adaptation evaluation system is constructed to obtain channel transmission adaptation parameters, including: Collect data on transmission delay, channel occupancy, and transmission stability of the HMB channel as HMB channel transmission status data; Extract the access priority features of each level of cache region from the hierarchical cache-association feature dataset, and extract the replacement urgency features of the core abnormal cache unit from the abnormal replacement-resource matching list; Construct an HMB channel transmission adaptation evaluation system, using HMB channel transmission status data, access priority characteristics, and replacement urgency characteristics as input dimensions; The evaluation system assigns weights to each input dimension and performs a comprehensive evaluation, outputting the degree of adaptation of each HMB channel to different levels of buffer areas and abnormal replacement requirements, thus obtaining channel transmission adaptation parameters.
[0011] Preferably, the operating temperature status of the solid-state drive (SSD) core is collected, and the collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and SSD core operating temperature status are input into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy, including: Real-time temperature data of the core area and cache chip area of the solid-state drive controller are collected, compared with the preset safe operating temperature benchmark, temperature data exceeding the benchmark are marked, and temperature difference characteristics are integrated to obtain the core operating temperature status of the solid-state drive. A cache-memory collaborative scheduling model is constructed, and the input dimensions of the cache-memory collaborative scheduling model are determined to be collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status. The data from each input dimension are imported into the cache-memory collaborative scheduling model. Based on the goals of optimal resource configuration and stable system operation, the model outputs a cache layering adjustment scheme, a host memory cache mapping optimization scheme, an HMB channel transmission priority allocation scheme, and an abnormal cache unit replacement execution scheme. Compatibility verification is performed on each output scheme to ensure that there are no conflicts between the schemes. The verified schemes are then integrated to obtain a cache-memory collaborative management strategy.
[0012] Preferably, the execution process of the cache-memory collaborative management strategy further includes: Based on the aforementioned cache tiering adjustment scheme, the tiering of the solid-state drive cache is dynamically adjusted, and the tiered cache-associated feature dataset is updated synchronously. Adjust the HMB collaborative mapping relationship according to the host memory cache mapping optimization scheme, and update the collaborative mapping-state fusion data; Based on the HMB channel transmission priority allocation scheme, the transmission resources of each HMB channel are adjusted to ensure the transmission priority of high-frequency access areas and abnormal replacement data. According to the aforementioned abnormal cache unit replacement execution plan, the core abnormal cache unit is replaced using the adapted idle cache resources in the abnormal replacement-resource matching list, and the cache unit status classification result is updated.
[0013] Preferably, according to the abnormal cache unit replacement execution plan, after replacing the core abnormal cache unit using the suitable idle cache resources in the abnormal replacement-resource matching list and updating the cache unit status classification result, the method further includes: The percentage of normal cache units after replacement in each cache region level is used as a region health assessment indicator. Calculate the average health assessment index of all cache regions at all levels as a benchmark for collaborative balancing; By comparing the health assessment metrics of each cache region with the collaborative balancing benchmark, redundant collaborative cache resources are extracted from regions where the health assessment metrics are higher than the benchmark. Redundant collaborative cache resources are reallocated to areas where health assessment indicators are below the baseline, and the remaining abnormal cache units in those areas are replaced until the health assessment indicators of those areas reach the collaborative balance baseline.
[0014] A solid-state drive (SSD) cache management method based on HMB technology and host memory collaboration is applicable to the aforementioned SSD cache management system based on HMB technology and host memory collaboration, including: Establish an HMB collaborative mapping relationship between solid-state drive cache and host memory, and synchronously collect the running status of solid-state drive cache and the available resource status of host memory to obtain collaborative mapping-status fusion data; Based on the collaborative mapping-state fusion data, the solid-state drive cache is divided into layers to obtain multi-level cache regions. The basic association features of each layer of cache regions are extracted to obtain a layered cache-association feature dataset. Based on the hierarchical cache-association feature dataset, the cache units in each level of cache area are traversed to carry out cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results. Based on the cache unit state classification results, core abnormal cache units are screened and their belonging features are extracted. Based on the collaborative mapping-state fusion data, suitable idle cache resources in the host memory are matched to obtain an abnormal replacement-resource matching list. Collect HMB channel transmission status data, and construct an HMB channel transmission adaptation evaluation system based on the hierarchical cache-associative feature dataset and the abnormal replacement-resource matching list to obtain channel transmission adaptation parameters; The operating temperature status of the solid-state drive (SSD) core is collected. The collaborative mapping-state fusion data, cache unit status classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and SSD core operating temperature status are input into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy.
[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention establishes an HMB collaborative mapping relationship between solid-state drive cache and host memory, enabling the system to manage data flow more effectively, optimize read and write operations, thereby accelerating data access speed and improving overall system performance. Furthermore, by monitoring the operating status and resource availability of solid-state drive and host memory in real time, the system can dynamically adjust cache tiering and resource allocation based on the current status, ensuring that frequently accessed data is always kept in a faster-responding storage area, reducing latency and improving data transmission efficiency. This invention monitors the operating status of cache units in real time and traces historical anomaly records. By classifying and managing normal and abnormal cache units, it can promptly identify and replace faulty units, thus maintaining system stability and reliability. Furthermore, by utilizing collaborative mapping-state fusion data, the system can intelligently match suitable idle cache resources to meet the replacement needs of core abnormal cache units, reducing the risk of performance degradation due to hardware failures. This invention constructs an HMB channel transmission adaptation evaluation system, which enables the system to rationally allocate HMB channel bandwidth resources according to the caching requirements and priorities of different levels, thus avoiding data transmission bottlenecks. Furthermore, by evaluating the health of each level of caching area, the system can identify areas with poor health and perform resource allocation and replacement to ensure the system's balanced and continuously efficient operation. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall system architecture in one embodiment of the present invention; Figure 2 This is a schematic flowchart of the overall method in one embodiment of the present invention.
[0017] In the diagram: 1. Data acquisition module; 2. Data partitioning module; 3. Unit classification module; 4. Resource matching module; 5. Transmission adaptation module; 6. Storage management module. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Example 1, please refer to Figure 1 This invention provides a technical solution: a solid-state drive cache management system based on HMB technology and host memory collaboration, comprising: Data acquisition module 1 is used to establish the HMB collaborative mapping relationship between solid-state drive cache and host memory, and synchronously collect the running status of solid-state drive cache and the available resource status of host memory to obtain collaborative mapping-status fusion data; Data partitioning module 2 is used to partition the solid-state drive cache into layers based on collaborative mapping-state fusion data, obtain multi-level cache regions, extract the basic association features of each level of cache region, and obtain a hierarchical cache-association feature dataset. Unit classification module 3 is used to traverse the cache units in each level of cache region based on the hierarchical cache-association feature dataset, carry out cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results; Resource matching module 4 is used to filter core abnormal cache units and extract their belonging features based on the cache unit status classification results, and to match the appropriate idle cache resources in the host memory based on the collaborative mapping-state fusion data to obtain an abnormal replacement-resource matching list. Transmission adaptation module 5 is used to collect HMB channel transmission status data, and construct an HMB channel transmission adaptation evaluation system based on hierarchical cache-associative feature dataset and abnormal replacement-resource matching list to obtain channel transmission adaptation parameters. Storage management module 6 is used to collect the operating temperature status of the solid-state drive core, and input the collaborative mapping-state fusion data, cache unit status classification results, abnormal replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status into the preset cache-memory collaborative scheduling model to obtain the cache-memory collaborative management strategy.
[0020] It's important to note that the data acquisition module establishes a collaborative mapping relationship between the SSD cache and host memory. It synchronously collects two types of status information: the SSD cache's operational status, including cache occupancy (how much space is used) and cache read / write response status (the speed and success rate of reading or writing data); and the host memory's available resource status, including free memory capacity (how much memory is available) and memory data transfer bandwidth (how fast memory can transfer data). Using this information, the module generates collaborative mapping-state fusion data, providing a foundation for subsequent processing. The data partitioning module then performs hierarchical partitioning of the SSD cache based on this collaborative mapping-state fusion data. The purpose of this hierarchical partitioning is to divide the cache region into multiple levels for better management and utilization. Basic correlation features are extracted from each level of the cache region, forming a hierarchical cache-correlation feature dataset, which will be used for subsequent classification and matching. The unit classification module inspects cache units within each tier of cache region. By analyzing the hierarchical cache-association feature dataset, the module can identify normal and abnormal cache units. This process includes detecting the current running status and tracing historical anomaly records. Finally, the module provides the cache unit status classification results to help identify abnormal cache units that require attention. Based on the cache unit status classification results, the resource matching module filters out the core abnormal cache units and extracts features related to these abnormal units. Simultaneously, the module uses co-mapping-state fusion data to match suitable idle cache resources in the host memory. Ultimately, this process generates an anomaly replacement-resource matching list for resource allocation when necessary. The transmission adaptation module is responsible for collecting transmission status data of the HMB (Host Memory Buffer) channel. It uses a hierarchical cache-association feature dataset and anomaly replacement-resource matching list to construct a transmission adaptation evaluation system for the HMB channel. Through this evaluation, the module can derive a series of channel transmission adaptation parameters to ensure that data can be efficiently and securely transferred between the SSD and host memory. The storage management module integrates all the data from the previous modules, including collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and SSD core operating temperature status, and inputs them into a preset cache-memory collaborative scheduling model. This model outputs cache-memory collaborative management strategies to guide the overall performance optimization and resource scheduling of the system.
[0021] In an optional embodiment, the solid-state drive cache operating status includes cache occupancy status and cache read / write response status, and the host memory available resource status includes memory free capacity and memory data transfer bandwidth; Establish an HMB (Hardware Memory Management) collaborative mapping relationship between the SSD cache and host memory, synchronously collect the SSD cache running status and the host memory available resource status, and obtain collaborative mapping-status fusion data, including: Construct an HMB co-mapping table to record the mapping relationship between each region of the solid-state drive cache and the host memory address, and obtain basic co-mapping data; The data collection process includes: the relationship between cache occupancy status and the total SSD cache capacity; the real-time feedback latency of cache read / write response status; and the acquisition of cache operation status data. The data collection process involves collecting the ratio of free host memory capacity to total host memory capacity, and assessing the compatibility between host memory data transfer bandwidth and the maximum transmission bandwidth of the HMB channel to obtain memory resource status data. By associating and integrating basic collaborative mapping data, cache runtime status data, and memory resource status data, collaborative mapping-state fusion data is obtained.
[0022] It's important to note that the SSD's cache area is divided into several blocks, such as block A, block B, and block C; each cache block is assigned a corresponding host memory address; for example: block A -> host memory address 0x0001; block B -> host memory address 0x0002; block C -> host memory address 0x0003; these mapping relationships are recorded in a table, i.e., the HMB (Hardware Memory Mapping) table; the SSD cache usage is calculated as a percentage of the total capacity. For example, if the total SSD cache capacity is 1GB, and... If the current usage is 300MB, then the percentage is 30%; monitor the response time of cache read and write operations; for example, if reading certain data takes 20 milliseconds and writing takes 15 milliseconds, summarize this data into the status metrics; calculate the ratio of host memory free capacity to total capacity; for example, if the total memory is 8GB and the free capacity is 2GB, then the percentage is 25%; monitor the compatibility between the host memory data transfer bandwidth (e.g., 4GB / s) and the maximum transfer bandwidth of the HMB channel (e.g., 6GB / s); the compatibility can be calculated as 4 / 6, or 66.67%; The system merges the basic collaborative mapping table, cache runtime status data, and memory resource status data to form a comprehensive report containing information such as cache usage, latency response, and memory availability. Through this comprehensive data, the system can analyze SSD performance bottlenecks, assess whether memory resources are sufficient, and determine further optimization measures. For example, if it finds that cache block B has high latency and insufficient memory free space, it can consider increasing the cache or optimizing the data scheduling strategy to improve overall performance.
[0023] In an optional embodiment, based on co-mapping-state fusion data, the solid-state drive cache is hierarchically divided to obtain multi-level cache regions. Basic association features of each level of cache region are extracted to obtain a hierarchical cache-association feature dataset, including: Extract historical access frequency, data access dependencies, and data validity duration characteristics from collaborative mapping-state fusion data of cached data; Based on access frequency characteristics, the solid-state drive cache is divided into multi-level cache regions, which include high-frequency access cache region, medium-frequency access cache region and low-frequency access cache region. Extract the access frequency threshold, data association strength, and data timeliness range corresponding to the multi-level cache regions as the basic association features of each level; The results of the multi-level cache region partitioning are associated with the corresponding basic related features and stored to obtain the hierarchical cache-related feature dataset.
[0024] It's important to note that recording the number of times each cached data item is accessed helps determine which data is most frequently requested. For example, if data A is accessed 100 times while data B is only accessed 20 times, then A can be considered a frequently accessed data item. It also helps identify dependencies between different data items. For instance, data C might be frequently requested simultaneously with data D, indicating a dependency. By analyzing access logs, these dependencies can be established, allowing for optimization of the data prefetching mechanism. Finally, it determines the validity of cached items during their use, i.e., the time period from when data is loaded into the cache until it is considered no longer valid. This is crucial for determining when the cache needs updating. For example, if data E hasn't been accessed in the past hour, its validity period can be considered over, and it should be evicted from the cache. The high-frequency access cache contains data items that have historically been accessed very frequently; for example, if data A and data B are accessed very frequently, they can be stored in the high-frequency access cache. The medium-frequency access cache contains data items that are accessed moderately; for example, if data C has been accessed 50 times and data D has been accessed 40 times, they can be placed in this area. The low-frequency access cache contains data items that are accessed very infrequently; for example, data E has only been accessed 5 times and is suitable for being placed in this area. Determine the frequency criteria required to divide each level of cache; for example, set a threshold so that data accessed more than 60 times enters the high-frequency access cache, data accessed 30 to 60 times enters the medium-frequency access cache, and data accessed less than 30 times enters the low-frequency access cache. Assess the access dependencies between different data by quantifying them by calculating the frequency with which data items are jointly accessed; for example, if data A and data D are frequently requested simultaneously, they can be considered to have a strong correlation. Based on access frequency and validity period, set different validity periods for cache regions; the validity period for high-frequency cached data can be set to one hour, while the validity period for low-frequency cached data may be one day. By associating the aforementioned multi-level cache regions with corresponding features, a complete dataset is formed; specifically, a storage structure can be created where the characteristics of each cache region, as well as its access frequency threshold, correlation strength, and validity period, are clearly visible. For example, suppose the analysis of a set of data yields the following results: the high-frequency access cache contains data A and data B, with an access frequency threshold of 60, an association strength of 0.9, and a data validity period of 1 hour; the medium-frequency access cache contains data C, with an access frequency threshold of 30, an association strength of 0.6, and a data validity period of 3 hours; the low-frequency access cache contains data E, with an access frequency threshold of single access, a low association strength, and a data validity period of 24 hours. By combining this information, the system can achieve more intelligent cache management, ensuring rapid response to high-frequency access data while timely cleaning or updating low-frequency access data.
[0025] In an optional embodiment, based on a hierarchical cache-association feature dataset, cache units within each cache region are traversed to perform cache unit operation status detection and historical anomaly record tracing, distinguishing between normal and abnormal cache units to obtain cache unit status classification results, including: Based on the hierarchical cache-association feature dataset, the boundary range of each level of cache region and the distribution of cache units are determined, and the cache unit traversal path is constructed. By scanning each cache unit through the traversal path, read and write performance tests are performed on each cache unit to obtain real-time running status data; Retrieve historical exception cache unit records, construct a cache unit status record table, and integrate historical exception information with real-time running status data; Based on preset normal operation standards, the cache units are compared with real-time operation status data to determine whether each cache unit is a normal cache unit or an abnormal cache unit. The determination results are recorded in the status record table to obtain the cache unit status classification results.
[0026] It's important to note that in a multi-layered caching architecture (e.g., L1, L2, L3 caches), the size, boundaries, and cache unit division of each layer must first be clearly defined. A cache unit can be a fixed-size data block, such as 64 bytes per unit. Assuming an L1 cache size of 32KB, it can hold 512 cache units (32KB / 64B). Analyzing the distribution of these cache units helps with subsequent performance testing and anomaly detection. For example, in an L2 cache, if the defined unit is 128 bytes, a 256KB L2 cache can be divided into 2048 cache units (256KB / 128B). Once the cache structure is determined, a traversal path needs to be created to ensure that all cache units can be accessed sequentially. This path will be used for subsequent read / write performance testing to comprehensively evaluate the performance of each cache unit. For example, a sequential access path can be designed, starting from the first cache unit in L1, accessing the last one sequentially, then entering the L2 cache, and continuing to traverse each layer in a similar manner. Based on the established traversal path, each cache unit is accessed sequentially, and read and write operations are performed to test its performance. This includes measuring key metrics such as read and write latency and bandwidth. For example, when accessing the first cache unit of L1, the time required to read the unit can be recorded, followed by a write operation, and then the write time can be recorded again. The average value is taken from multiple tests to obtain more accurate performance data. While performing read and write performance tests, the status information of each cache unit, such as the number of accesses, success rate, and latency, is recorded in real time. This data will be used for subsequent analysis. For example, if the read latency of an L1 cache unit exceeds a set threshold (e.g., 10 microseconds) during the test, this status is recorded. By combining historical data, information on previously recorded abnormal cache units is obtained, such as units that have experienced high latency or read / write failures. This helps identify potentially problematic units and perform trend analysis. For example, if historical records show that a certain L2 cache unit frequently experienced timeout errors in past tests, this unit will receive special attention in subsequent analysis. Real-time operational status data is integrated with historical anomaly information to form a complete cache unit status record table for subsequent analysis and decision-making. Pre-defined normal operation standards (such as maximum latency, minimum bandwidth, etc.) are used to determine the status of each cache unit. Real-time data is compared with the standards to mark which cache units are normal and which are abnormal. For example, if a cache unit has an average read latency of 12 microseconds, while the preset normal standard is 10 microseconds, then the unit is marked as abnormal; if another unit has a latency of 8 microseconds, it is marked as normal. Finally, the judgment results (normal or abnormal) are entered into the status record table to form the final cache unit status classification results for subsequent monitoring and management.
[0027] In an optional embodiment, based on the cache unit state classification results, core abnormal cache units are screened and their attribution features are extracted. Furthermore, based on the collaborative mapping-state fusion data, suitable idle cache resources in the host memory are matched to obtain an abnormal replacement-resource matching list, including: Select the abnormal cache units in the high-frequency access level cache area from the cache unit status classification results as the core abnormal cache units; Locate the hierarchical cache region to which the core exception cache unit belongs and the node location within the region, and integrate the hierarchical cache region and the node location within the region into the attribution characteristics of the core exception cache unit; Based on the HMB collaborative mapping relationship in the collaborative mapping-state fusion data, the host memory free cache resources corresponding to the attribution features are matched, and the capacity and transmission performance of the matched resources are verified to meet the replacement requirements of the core abnormal cache unit. By associating the resource information that meets the requirements with the corresponding core exception cache unit information, an exception replacement-resource matching list is obtained.
[0028] It's important to note that, based on the previously established cache unit status records, cache units marked as abnormal and exhibiting high-frequency access characteristics are identified. These units are typically accessed frequently during system operation but exhibit abnormal performance metrics (such as high latency or read failures). For example, suppose several cache units in the L1 cache have shown frequent abnormalities in past tests, such as average latency exceeding a preset standard, and these units are accessed very frequently. These units will be selected as core abnormal cache units. Once the core abnormal cache units are determined, the next step is to clarify which cache region these units belong to and their specific location within that region. This helps with subsequent resource management and replacement decisions. For example, if a core abnormal cache unit is located at byte 256 of the L1 cache, then its affiliation can be defined as "L1 cache region, node position 256B". In this way, the root cause of the problem can be accurately located. The hierarchical cache regions obtained from the location are integrated with the node location information to form the complete ownership characteristics of the core anomaly cache unit. This characteristic will be used in the subsequent resource matching and replacement process. For example, assuming the ownership characteristics of the core anomaly cache unit are "L2 cache region, node location 128KB", this information will become the basis for subsequent operations. Using existing cooperative mapping data, especially the cooperative mapping relationship of HMB, the host memory free cache resources that match the ownership characteristics of the core anomaly cache unit are searched. At this time, it is necessary to verify the capacity and transmission performance of the matched resources to ensure that they can meet the replacement requirements. For example, if there is an HMB resource in the system identified as "256KB, transmission rate 10GB / s", then it must first be confirmed whether this resource can replace the aforementioned core anomaly cache unit and meet the performance requirements. For a found matching resource, its capacity and performance need to be checked to see if they meet the requirements of the core anomaly cache unit. If they do, it is marked as a replaceable resource. For example, if a 128KB core anomaly cache unit requires a replacement resource of at least the same size and a required transfer rate of at least 8GB / s, then if an available resource is found to be "256KB, 10GB / s", it can meet the replacement requirement. Finally, the information of the matching resources is associated with the previously identified core anomaly cache units to form an "anomaly replacement-resource matching list". This list helps decision-makers quickly find suitable resources when a replacement is needed. For example, assuming that a core anomaly cache unit "L1 cache area, node location 256B" has been confirmed to need to be replaced, and there is a suitable resource in the system "L1 cache area, node location 512B, 256KB, 10GB / s", then a replacement record can be generated, indicating that the resource can be used to replace the core anomaly unit, thereby ensuring system stability and performance improvement.
[0029] In an optional embodiment, HMB channel transmission status data is collected, and an HMB channel transmission adaptation evaluation system is constructed based on a hierarchical cache-associated feature dataset and an anomaly replacement-resource matching list to obtain channel transmission adaptation parameters, including: Collect data on transmission delay, channel occupancy, and transmission stability of the HMB channel as HMB channel transmission status data; Extract access priority features of each level of cache region from the hierarchical cache-association feature dataset, and extract replacement urgency features of core exception cache units from the exception replacement-resource matching list; Construct an HMB channel transmission adaptation evaluation system, using HMB channel transmission status data, access priority characteristics, and replacement urgency characteristics as input dimensions; The evaluation system assigns weights to each input dimension and performs a comprehensive evaluation, outputting the degree of adaptation of each HMB channel to different levels of buffer areas and abnormal replacement requirements, thus obtaining channel transmission adaptation parameters.
[0030] It's important to note that, firstly, key performance indicators related to the HMB channel need to be collected. These include: transmission latency: the time required for data to travel from the sender to the receiver; channel occupancy: the usage of the HMB channel within a specific time period, usually expressed as a percentage; and transmission stability: measuring fluctuations or instabilities during transmission, such as packet loss rate or latency fluctuations. For example, assuming that during the monitoring period, the HMB channel exhibits an average transmission latency of 5ms, an occupancy rate of 70%, good transmission stability, and a packet loss rate of less than 1%, these data will be recorded as the transmission status of the HMB channel. Secondly, the access priorities of different cache levels are extracted from the associated feature dataset of the hierarchical cache. This is determined based on system design and application requirements; generally, higher-priority cache areas have greater access frequency and importance. For example, in a multi-level cache architecture, the L1 cache may be assigned the highest access priority (e.g., priority level 1), while the L3 cache may be assigned a lower priority (e.g., priority level 3). These priority characteristics will be used for subsequent evaluation and decision-making. In the anomaly replacement-resource matching list, the replacement urgency feature of core anomaly cache units is extracted. This feature reflects the urgency of quickly replacing the cache unit, usually judged based on the severity of the anomaly and its impact on system performance. For example, if a core anomaly cache unit affects overall system performance due to high latency and frequent errors, its replacement urgency may be marked as "high" (level 1), while minor anomalies may be marked as "low" (level 3). The collected HMB channel transmission status data, access priority features, and replacement urgency features are integrated into the input dimensions of the evaluation system. The goal of this evaluation system is to quantify the adaptability of the HMB channel to different levels of cache regions and anomaly replacement requirements. For example, this evaluation system may include an algorithm that can process multiple input dimensions and finally output an adaptation score. For example, the input dimensions include: HMB channel transmission latency (5ms); L1 cache access priority (level 1); replacement urgency (level 1). Within the evaluation system, weights are assigned to each input dimension. For example, weights can be set for transmission latency, priority, and replacement urgency (e.g., latency 50%, priority 30%, urgency 20%). Then, a comprehensive evaluation algorithm calculates the suitability of each HMB channel for different cache levels and anomaly replacement requirements. For example, based on weighted evaluation, assuming the evaluation result is a transmission suitability score of 85 (out of 100), it indicates that this HMB channel is very suitable for replacing anomaly units in the L1 cache. Finally, based on the comprehensive evaluation results, the transmission suitability parameters of each HMB channel are output. These parameters will help decision-makers understand which HMB channels are most suitable for replacing specific core anomaly cache units. For example, if the evaluation system output shows "Channel A suitability score of 90, Channel B suitability score of 75", then the decision-maker can choose Channel A to replace the anomaly core cache unit, thereby ensuring improved system performance.
[0031] In an optional embodiment, the operating temperature status of the solid-state drive (SSD) core is collected. The collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and the SSD core operating temperature status are input into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy, including: Real-time temperature data of the core area and cache chip area of the solid-state drive controller are collected, compared with the preset safe operating temperature benchmark, temperature data exceeding the benchmark are marked, and temperature difference characteristics are integrated to obtain the core operating temperature status of the solid-state drive. A cache-memory collaborative scheduling model is constructed, and the input dimensions of the cache-memory collaborative scheduling model are determined to be collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status. The data from each input dimension are imported into the cache-memory collaborative scheduling model. Based on the goals of optimal resource configuration and stable system operation, the model outputs a cache tiering adjustment scheme, a host memory cache mapping optimization scheme, an HMB channel transmission priority allocation scheme, and an abnormal cache unit replacement execution scheme. Compatibility verification is performed on each output scheme to ensure that there are no conflicts between the schemes. The verified schemes are then integrated to obtain a cache-memory collaborative management strategy.
[0032] It's important to note that real-time monitoring of the SSD controller core area and cache chip area is necessary. This temperature data is compared to a preset safe operating temperature benchmark to ensure the hardware operates within a safe range. For example, assuming the preset safe operating temperature benchmark is 70 degrees Celsius, during real-time monitoring, the controller core area temperature is 72 degrees Celsius, and the cache chip area temperature is 68 degrees Celsius. Based on the comparison, the controller core area temperature exceeds the safe range, and this data will be marked as abnormal. After collecting the temperature data, it's necessary to integrate the temperature difference characteristics of this data to understand the SSD's operating temperature status. This step is to assess the potential impact of temperature anomalies on hardware performance. For example, the temperature difference characteristics can be represented as: the difference between the controller core temperature and the safe benchmark (72°C - 70°C = 2°C), while the difference in cache chip temperature is (68°C - 70°C = -2°C). Through integration, the core operating temperature status is determined to be "abnormal" because the core temperature exceeds the safe threshold. Next, a cache-memory collaborative scheduling model is constructed. The input dimensions of this model include: collaborative mapping-state fusion data: integrating cache and memory usage status information; cache unit status classification results: classifying the health status of cache units; anomaly replacement-resource matching list: listing anomaly units and alternative resources; channel transmission adaptation parameters: evaluating the adaptation between HMB channels and cache; and solid-state drive core operating temperature status: the current temperature status. For example, we can assume that the collaborative mapping-state fusion data indicates that the current memory utilization rate is 80%, the cache unit status classification results show that there is a cache unit in the "faulty" state, and the anomaly replacement resource matching list lists an alternative cache. The data from each of the above input dimensions are imported into the cache-memory collaborative scheduling model. Based on the goals of optimal resource configuration and stable system operation, the model will generate multiple optimization schemes. For example: cache tier adjustment scheme: hot access data is transferred to a faster cache tier to reduce latency; host memory cache mapping optimization scheme: the mapping relationship of memory cache is dynamically adjusted according to the load to improve efficiency; HMB channel transmission priority allocation scheme: priority is assigned to different channels according to the current load and temperature status to ensure timely transmission of critical data; abnormal cache unit replacement execution scheme: the process of replacing faulty cache units is initiated according to the abnormal replacement resource matching list. Compatibility verification is performed on each output scheme to ensure there are no conflicts between them. For example, cache tiering adjustment schemes and host memory cache mapping optimization schemes need to work in coordination to avoid data inconsistency or performance degradation caused by adjustments. For example, during the verification process, it was found that some cache adjustment schemes would affect the availability of host memory, so the schemes need to be re-evaluated to ensure their mutual compatibility. Finally, all verified schemes are integrated to form a complete cache-memory co-management strategy. This strategy will guide the system's optimization decisions under different loads and operating states. For example, the integrated strategy may include periodically monitoring the temperature and initiating cooling strategies when the temperature exceeds the threshold, while ensuring that cache tiering adjustment and memory mapping optimization are carried out synchronously to maintain system stability and high performance.
[0033] In an optional embodiment, the execution process of the cache-memory collaborative management strategy further includes: The tiered cache division of the solid-state drive is dynamically adjusted based on the cache tiering adjustment scheme, and the tiered cache-related feature dataset is updated synchronously. Adjust the HMB collaborative mapping relationship according to the host memory cache mapping optimization scheme, and update the collaborative mapping-state fusion data; The transmission resources of each HMB channel are adjusted based on the HMB channel transmission priority allocation scheme to ensure the transmission priority of high-frequency access areas and abnormal replacement data. According to the abnormal cache unit replacement execution plan, the core abnormal cache unit is replaced by using the adapted idle cache resources in the abnormal replacement-resource matching list, and the cache unit status classification result is updated.
[0034] It's important to note that the cache levels within the SSD are dynamically adjusted based on real-time data and system requirements. This means reclassifying the cache layers according to factors such as access frequency and data usage to improve storage access efficiency. For example, assuming the SSD has three cache levels: L1 (high-speed cache), L2 (medium-speed cache), and L3 (low-speed cache), during monitoring, if certain data blocks are found to be frequently accessed (such as database indexes), these data can be promoted from L3 to L1, while less frequently used data can be downgraded to L3. After the update, the hierarchical cache-associative feature dataset needs to be updated synchronously. This dataset records the data characteristics and access patterns in each level for better future decision-making. Based on the optimized collaborative mapping scheme between host memory and SSD, the mapping relationship between host memory and SSD is dynamically adjusted. This operation aims to improve data transfer efficiency and response speed. For example, suppose the existing mapping relationship results in low transfer efficiency of some frequently accessed data between memory and SSD. Through the optimization scheme, the mapping can be adjusted so that frequently accessed data blocks are directly mapped to the fast access area of host memory, thereby reducing latency. The collaborative mapping-state fusion data is updated to accurately reflect the new mapping relationship and state, ensuring that the system can monitor the data flow between memory and SSD in real time. The resource allocation of each HMB channel is controlled according to the HMB channel transmission priority allocation scheme. The key is to ensure that frequently accessed areas and data that needs to be replaced can obtain transmission resources first, thereby improving overall performance. For example, suppose multiple HMB channels are transmitting data simultaneously, one channel is handling urgent database read requests, while another channel is used to transmit less frequently accessed data. According to the priority allocation scheme, the channel resources for urgent read requests can be increased to ensure they obtain the required bandwidth, while limiting the bandwidth usage of other channels. This ensures that data transmission in high-frequency access areas is not affected, improving system response speed. Finally, according to the abnormal cache unit replacement execution scheme, the available idle cache resources in the abnormal replacement-resource matching list are used to replace the core abnormal cache unit. This step is to repair system faults and restore normal operation. For example, suppose a cache unit is found to have failed due to overheating during monitoring and is marked as "abnormal". According to the abnormal replacement-resource matching list, the system identifies an idle, high-performance cache unit that can be replaced. The faulty unit is immediately replaced with a new cache unit, and the cache unit status classification result is updated, marking the new unit as "normal". At the same time, the relevant information of the replacement operation is recorded for subsequent maintenance and analysis.
[0035] In an optional embodiment, after replacing the core exception cache unit using the adapted idle cache resources in the exception replacement-resource matching list according to the exception cache unit replacement execution scheme, and updating the cache unit status classification result, the method further includes: The percentage of normal cache units after replacement in each cache region level is used as a region health assessment indicator. Calculate the average health assessment index of all cache regions at all levels as a benchmark for collaborative balancing; By comparing the health assessment metrics of each cache region with the collaborative balancing benchmark, redundant collaborative cache resources are extracted from regions where the health assessment metrics are higher than the benchmark. Redundant collaborative cache resources are reallocated to areas where health assessment indicators are below the baseline, and the remaining abnormal cache units in those areas are replaced until the health assessment indicators of those areas reach the collaborative balance baseline.
[0036] It should be noted that the health status of each cache tier is assessed by statistically analyzing the proportion of normally functioning cache units in each tier. For example, suppose an SSD has three cache tiers: L1, L2, and L3, with 100, 200, and 300 cache units per tier, respectively. In one monitoring session, 90 normally functioning cache units were found in L1, 150 in L2, and 250 in L3. The percentage of normally functioning cache units is calculated as follows: L1 normally functioning percentage = 90 / 100 = 90%; L2 normally functioning percentage = 150 / 200 = 75%; L3 normally functioning percentage = 250 / 300 = 83.33%. These percentages are used as the health assessment indicators for each tier. Next, the average health assessment index of all cache regions needs to be calculated as the collaborative balancing benchmark for subsequent comparisons. For example, according to the calculation results above, the health assessment indices of L1, L2, and L3 are 90%, 75%, and 83.33%, respectively; the average is calculated as (90% + 75% + 83.33%) / 3 = 82.78%; therefore, the collaborative balancing benchmark is 82.78%. Based on the calculated health assessment indices of each level, it is determined which regions have health assessment indices higher than the collaborative balancing benchmark, and redundant collaborative cache resources in these regions are extracted. For example, based on the previous indices: L1: 90% (higher than 82.78%); L2: 75% (lower than 82.78%); L3: 83.33% (higher than 82.78%); therefore, the health assessment indices of L1 and L3 are higher than the benchmark, and redundant cache resources can be extracted. Assuming 20 redundant cache units are identified from L1 and L3, these resources can be used in other regions. Redundant cache resources extracted from regions with health assessment metrics above the baseline are reallocated to regions with health assessment metrics below the baseline (such as L2), and the remaining abnormal cache units in that region are replaced. For example, 20 extracted redundant cache units are allocated to the L2 region to enhance its performance. Assuming L2 originally had 50 abnormal cache units, by reallocating these 20 redundant resources, 20 abnormal units are replaced, leaving 30 abnormal units remaining. At this point, the number of normal cache units in L2 may increase from 150 to 170, and the health assessment metrics will improve accordingly. Finally, necessary replacements and reallocations continue until the health assessment metrics reach the collaborative balance baseline. For example, if after reallocation, the proportion of normal cache units in L2 is calculated to be 85%, which still does not reach the baseline of 82.78%, more cache resources can be extracted from the redundant regions, or other replacement measures can be implemented until the health assessment metrics of L2 meet or exceed 82.78%. In this way, the health assessment metrics of all regions will eventually be improved, achieving overall system optimization and performance enhancement.
[0037] Example 2, please refer to Figure 2 This invention provides a technical solution: a solid-state drive cache management method based on HMB technology and host memory collaboration, applicable to the aforementioned solid-state drive cache management system based on HMB technology and host memory collaboration, comprising: S1. Establish the HMB collaborative mapping relationship between the solid-state drive cache and the host memory, and synchronously collect the running status of the solid-state drive cache and the available resource status of the host memory to obtain collaborative mapping-status fusion data. S2. Based on collaborative mapping-state fusion data, the solid-state drive cache is divided into layers to obtain multi-level cache regions. The basic association features of each level of cache region are extracted to obtain a layered cache-association feature dataset. S3. Based on the hierarchical cache-association feature dataset, traverse the cache units in each level of cache area, carry out cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results. S4. Based on the cache unit status classification results, core abnormal cache units are selected and their belonging characteristics are extracted. Based on the collaborative mapping-state fusion data, the appropriate idle cache resources in the host memory are matched to obtain the abnormal replacement-resource matching list. S5. Collect HMB channel transmission status data, and construct an HMB channel transmission adaptation evaluation system based on the hierarchical cache-associative feature dataset and the abnormal replacement-resource matching list to obtain channel transmission adaptation parameters. S6. Collect the core operating temperature status of the solid-state drive, and input the collaborative mapping-state fusion data, cache unit status classification results, abnormal replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status into the preset cache-memory collaborative scheduling model to obtain the cache-memory collaborative management strategy.
[0038] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A solid-state drive cache management system based on HMB technology and host memory collaboration, characterized in that, include: The data acquisition module is used to establish the HMB collaborative mapping relationship between the solid-state drive cache and the host memory, and synchronously collect the running status of the solid-state drive cache and the available resource status of the host memory to obtain collaborative mapping-status fusion data. The data partitioning module is used to perform hierarchical partitioning of the solid-state drive cache based on the collaborative mapping-state fusion data to obtain multi-level cache regions, extract the basic association features of each level of cache region, and obtain a hierarchical cache-association feature dataset. The unit classification module is used to traverse the cache units in each level of cache region based on the hierarchical cache-association feature dataset, perform cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results. The resource matching module is used to filter core abnormal cache units and extract their belonging features based on the cache unit status classification results, and to match the appropriate idle cache resources in the host memory based on the collaborative mapping-state fusion data to obtain an abnormal replacement-resource matching list. The transmission adaptation module is used to collect HMB channel transmission status data, and based on the hierarchical cache-associative feature dataset and the abnormal replacement-resource matching list, construct an HMB channel transmission adaptation evaluation system to obtain channel transmission adaptation parameters. The storage management module is used to collect the operating temperature status of the solid-state drive core, and input the collaborative mapping-state fusion data, cache unit status classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy.
2. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 1, characterized in that, The solid-state drive cache operating status includes cache occupancy status and cache read / write response status; the host memory available resource status includes memory free capacity and memory data transfer bandwidth. Establish an HMB (Hardware Memory Management) collaborative mapping relationship between the SSD cache and host memory, synchronously collect the SSD cache running status and the host memory available resource status, and obtain collaborative mapping-status fusion data, including: Construct an HMB co-mapping table to record the mapping relationship between each region of the solid-state drive cache and the host memory address, and obtain basic co-mapping data; The data collection process includes: the relationship between cache occupancy status and the total SSD cache capacity; the real-time feedback latency of cache read / write response status; and the acquisition of cache operation status data. The data collection process involves collecting the ratio of free host memory capacity to total host memory capacity, and assessing the compatibility between host memory data transfer bandwidth and the maximum transmission bandwidth of the HMB channel to obtain memory resource status data. The basic collaborative mapping data, cache running status data, and memory resource status data are correlated and integrated to obtain collaborative mapping-state fusion data.
3. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 2, characterized in that, Based on the aforementioned collaborative mapping-state fusion data, the solid-state drive cache is hierarchically divided to obtain multi-level cache regions. Basic association features of each level of cache region are extracted to obtain a hierarchical cache-association feature dataset, including: Extract historical access frequency, data access dependencies, and data validity duration characteristics of cached data from the collaborative mapping-state fusion data; Based on access frequency characteristics, the solid-state drive cache is divided into multi-level cache regions, which include a high-frequency access cache region, a medium-frequency access cache region, and a low-frequency access cache region. Extract the access frequency threshold, data association strength, and data timeliness range corresponding to the multi-level cache regions as the basic association features of each level; The division results of the multi-level cache region are associated with and stored with the corresponding basic associated features to obtain the hierarchical cache-associated feature dataset.
4. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 3, characterized in that, Based on the hierarchical cache-association feature dataset, the cache units within each level of the cache region are traversed to perform cache unit operation status detection and historical anomaly record tracing, distinguishing between normal and abnormal cache units, and obtaining cache unit status classification results, including: Based on the hierarchical cache-association feature dataset, determine the boundary range and cache unit distribution of each level of cache region, and construct the cache unit traversal path; By scanning each cache unit through the traversal path, read and write performance tests are performed on each cache unit to obtain real-time running status data; Retrieve historical exception cache unit records, construct a cache unit status record table, and integrate historical exception information with real-time running status data; Based on preset normal operation standards, the cache units are compared with real-time operation status data to determine whether each cache unit is a normal cache unit or an abnormal cache unit. The determination results are entered into the status record table to obtain the cache unit status classification results.
5. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 4, characterized in that, Based on the cache unit status classification results, core abnormal cache units are selected and their attribution features are extracted. And based on the collaborative mapping-state fusion data, matching the adapted idle cache resources in the host memory yields an abnormal replacement-resource matching list, including: From the cache unit status classification results, select the abnormal cache units in the high-frequency access level cache area as the core abnormal cache units; Locate the hierarchical cache region to which the core exception cache unit belongs and the node positions within that region, and integrate the hierarchical cache region and the node positions within that region into the attribution characteristics of the core exception cache unit; Based on the HMB collaborative mapping relationship in the collaborative mapping-state fusion data, the host memory free cache resources corresponding to the attribution feature are matched, and at the same time, it is verified whether the capacity and transmission performance of the matched resources meet the replacement requirements of the core abnormal cache unit. By associating the resource information that meets the requirements with the corresponding core exception cache unit information, an exception replacement-resource matching list is obtained.
6. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 5, characterized in that, Collect HMB channel transmission status data, and based on the hierarchical cache-associative feature dataset and the anomaly replacement-resource matching list, construct an HMB channel transmission adaptation evaluation system to obtain channel transmission adaptation parameters, including: Collect data on transmission delay, channel occupancy, and transmission stability of the HMB channel as HMB channel transmission status data; The access priority features of each level of cache region are extracted from the hierarchical cache-association feature dataset, and the replacement urgency features of the core abnormal cache unit are extracted from the abnormal replacement-resource matching list. Construct an HMB channel transmission adaptation evaluation system, using HMB channel transmission status data, access priority characteristics, and replacement urgency characteristics as input dimensions; The evaluation system assigns weights to each input dimension and performs a comprehensive evaluation, outputting the degree of adaptation of each HMB channel to different levels of buffer areas and abnormal replacement requirements, thus obtaining channel transmission adaptation parameters.
7. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 6, characterized in that, The operating temperature status of the solid-state drive (SSD) core is collected. The collaborative mapping-state fusion data, cache unit status classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and the SSD core operating temperature status are input into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy, including: Real-time temperature data of the core area and cache chip area of the solid-state drive controller are collected, compared with the preset safe operating temperature benchmark, temperature data exceeding the benchmark are marked, and temperature difference characteristics are integrated to obtain the core operating temperature status of the solid-state drive. A cache-memory collaborative scheduling model is constructed, and the input dimensions of the cache-memory collaborative scheduling model are determined to be collaborative mapping-state fusion data, cache unit state classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and solid-state drive core operating temperature status. The data from each input dimension are imported into the cache-memory collaborative scheduling model. Based on the goals of optimal resource configuration and stable system operation, the model outputs a cache layering adjustment scheme, a host memory cache mapping optimization scheme, an HMB channel transmission priority allocation scheme, and an abnormal cache unit replacement execution scheme. Compatibility verification is performed on each output scheme to ensure that there are no conflicts between the schemes. The verified schemes are then integrated to obtain the cache-memory collaborative management strategy.
8. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 7, characterized in that, The execution process of the cache-memory collaborative management strategy also includes: Based on the aforementioned cache tiering adjustment scheme, the tiering of the solid-state drive cache is dynamically adjusted, and the tiered cache-associated feature dataset is updated synchronously. Adjust the HMB collaborative mapping relationship according to the host memory cache mapping optimization scheme, and update the collaborative mapping-state fusion data; Based on the HMB channel transmission priority allocation scheme, the transmission resources of each HMB channel are adjusted to ensure the transmission priority of high-frequency access areas and abnormal replacement data. According to the abnormal cache unit replacement execution plan, the core abnormal cache unit is replaced by the appropriate idle cache resources in the abnormal replacement-resource matching list, and the cache unit status classification result is updated.
9. The solid-state drive cache management system based on HMB technology and host memory collaboration as described in claim 8, characterized in that, According to the aforementioned abnormal cache unit replacement execution plan, after replacing the core abnormal cache unit using the adapted idle cache resources in the abnormal replacement-resource matching list and updating the cache unit status classification result, the process further includes: The percentage of normal cache units after replacement in each cache region level is used as a region health assessment indicator. Calculate the average health assessment index of all cache regions at all levels as a benchmark for collaborative balancing; By comparing the health assessment metrics of each cache region with the collaborative balancing benchmark, redundant collaborative cache resources are extracted from regions where the health assessment metrics are higher than the benchmark. Redundant collaborative cache resources are reallocated to areas where health assessment indicators are below the baseline, and the remaining abnormal cache units in those areas are replaced until the health assessment indicators of those areas reach the collaborative balance baseline.
10. A solid-state drive cache management method based on HMB technology and host memory collaboration, applicable to the solid-state drive cache management system based on HMB technology and host memory collaboration as described in any one of claims 1-9, characterized in that, include: Establish an HMB collaborative mapping relationship between solid-state drive cache and host memory, and synchronously collect the running status of solid-state drive cache and the available resource status of host memory to obtain collaborative mapping-status fusion data; Based on the collaborative mapping-state fusion data, the solid-state drive cache is divided into layers to obtain multi-level cache regions. The basic association features of each layer of cache regions are extracted to obtain a layered cache-association feature dataset. Based on the hierarchical cache-association feature dataset, the cache units in each level of cache area are traversed to carry out cache unit operation status detection and historical anomaly record tracing, distinguish normal cache units from abnormal cache units, and obtain cache unit status classification results. Based on the cache unit state classification results, core abnormal cache units are screened and their belonging features are extracted. Based on the collaborative mapping-state fusion data, suitable idle cache resources in the host memory are matched to obtain an abnormal replacement-resource matching list. Collect HMB channel transmission status data, and construct an HMB channel transmission adaptation evaluation system based on the hierarchical cache-associative feature dataset and the abnormal replacement-resource matching list to obtain channel transmission adaptation parameters; The operating temperature status of the solid-state drive (SSD) core is collected. The collaborative mapping-state fusion data, cache unit status classification results, anomaly replacement-resource matching list, channel transmission adaptation parameters, and SSD core operating temperature status are input into a preset cache-memory collaborative scheduling model to obtain a cache-memory collaborative management strategy.