Garbage collection method and device based on time series causal model and electronic equipment

CN122653546BActive Publication Date: 2026-09-29JINAN MAIWEI INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611130988.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-29
Estimated Expiration
2046-07-29

AI Technical Summary

Technical Problem

[0003]相关技术中,GC方法包括直接采用贪心算法(Greedy Algorithm),仅选择无效数据页最多的块回收,并未考虑数据页的冷热属性,会频繁迁移短期内会被更新的热数据,导致写放大(Write Amplification,WA)急剧升高,严重缩短NAND闪存使用寿命;或采用成本收益算法(Cost-Benefit,CB),其虽然综合了无效数据页比例与数据更新频率,但核心阈值固定,无法动态适配消费级办公、工业级宽温、数据中心高并发等多变的输入输出(Input/Output,IO)负载场景,GC触发时机与块选择的合理性有待提升;此外基于成本收益算法优化演进的成本时长加权的GC算法(Cost Age Times,CAT),其核心目标是平衡三大诉求:最小化回收开销、控制写放大以及实现磨损均衡,同时天然适配冷热数据分离的需求,但是固定权重无法动态适配负载,极端场景存在短板

Benefits of technology

[0027]本公开的基于时序因果模型的垃圾回收方法,通过时序因果模型同步完成GC触发时机预判、存储块回收优先级排序、数据页冷热分类多任务决策,能够精准捕捉时序维度IO负载与数据冷热演变规律,有效减少热数据迁移,显著降低写放大、延长NAND闪存使用寿命;模型采用轻量化时序因果卷积结构,计算开销低,可直接部署于SSD主控,适配消费、工业、数据中心各类负载场景,无需额外算力资源;同时支持算力自适应特征裁剪,搭配完整推理容错机制,在保障GC决策合理性的同时,稳定维持SSD持续读写性能与磨损均衡效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653546B_ABST
    Figure CN122653546B_ABST
Patent Text Reader

Abstract

The disclosure provides a garbage collection method and device based on a time series causal model and an electronic device, applied to the technical field of data storage, wherein the time series causal model comprises a time series feature enhancement layer, a time series feature extraction layer and a multi-task output layer; the method comprises: determining a first feature matrix based on storage block state data, data page state data, input / output load data and garbage collection history data of a to-be-predicted storage block; inputting the first feature matrix into the time series causal model to obtain a corresponding garbage collection decision result of the to-be-predicted storage block; and determining whether to take a garbage collection operation on the to-be-predicted storage block based on the garbage collection decision result, so that the garbage collection decision can be reasonable and the continuous read / write performance and wear leveling effect of the storage block can be stably maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data storage technology, and in particular to a garbage collection method, apparatus and electronic device based on a time-series causal model. Background Technology

[0002] Garbage collection (GC) technology in solid-state drives (SSDs) is a key background management mechanism used to reclaim storage space occupied by invalid data, maintaining the SSD's write performance and lifespan. SSDs are based on a NAND flash memory architecture, where writes are performed on data pages and erases on storage blocks. Offsite update mechanisms continuously generate invalid data pages, requiring GC to reclaim invalid storage blocks, migrate valid data pages, and finally complete the storage block erasure.

[0003] Among related technologies, GC methods include directly using a greedy algorithm, which only selects the block with the most invalid data pages for recycling, without considering the hot and cold attributes of data pages. This leads to frequent migration of hot data that will be updated in the short term, resulting in a sharp increase in write amplification (WA) and severely shortening the lifespan of NAND flash memory. Alternatively, a cost-benefit (CB) algorithm can be used, which, although combining the proportion of invalid data pages and the data update frequency, has a fixed core threshold and cannot dynamically adapt to the variable input / output (IO) load scenarios such as consumer office, industrial wide-temperature, and high-concurrency data center scenarios. The rationality of GC triggering timing and block selection needs to be improved. In addition, the cost-age-time weighted GC algorithm (CAT), which is an optimization and evolution of the cost-benefit algorithm, aims to balance three major requirements: minimizing recycling overhead, controlling write amplification, and achieving wear leveling. It is also naturally adapted to the requirement of separating hot and cold data. However, the fixed weight cannot dynamically adapt to the load, and there are shortcomings in extreme scenarios.

[0004] In summary, the GC methods in related technologies can only make judgments based on historical statistical data at a single moment, and cannot capture the load changes and data hot and cold evolution patterns in the time-series dimension. They can only respond passively and cannot actively predict. In addition, many weight parameters are static, resulting in poor decision rationality. Summary of the Invention

[0005] This disclosure provides a waste recycling method, apparatus, and electronic device based on a time-series causal model, to at least solve the above-mentioned technical problems existing in the prior art.

[0006] According to a first aspect of this disclosure, a garbage collection method based on a temporal causal model is provided, wherein the temporal causal model includes a temporal feature enhancement layer, a temporal feature extraction layer, and a multi-task output layer; the method includes: Based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted, the first feature matrix is ​​determined; The first feature matrix is ​​input into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted. The first feature map is input into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted. The second feature matrix is ​​input into the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted. Based on the garbage collection decision results, it is determined whether to perform garbage collection operations on the storage block to be predicted.

[0007] In the above scheme, the dimension of the first feature matrix is ​​T×D×C, where C represents the number of time windows or the dimension of input channels; D represents the total dimension of the storage block status data, data page status data, input / output load data and garbage collection history data of the storage block to be predicted; and T represents the total sampling time of any time window. The storage block status data includes at least one of the following: the number of valid data pages in the storage block, the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, the storage time of data in the storage block, and bad block risk markers. The data page status data includes at least one of the following: the storage time of data in the data page, the number of historical writes to the data page, the number of times the data page was accessed within a period, historical hot / cold markers, and valid data page information markers. The input / output load data includes at least one of the following: read input / output operations per second (IOPS), write IOPS, and average interval between input / output requests. The garbage collection history data includes at least one of the following: historical average garbage collection time, historical average valid page migration, historical average write amplification factor, average garbage collection trigger interval, and storage block wear leveling.

[0008] In the above scheme, determining the first feature matrix based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted further includes: Determine the computing power of the solid-state drive controller; In response to the computing power of the solid-state drive controller being less than or equal to the computing power threshold, the first feature matrix is ​​pruned; wherein, the temporal dimension of the pruned first feature matrix remains unchanged, but the feature dimension is smaller than the feature dimension of the first feature matrix before pruning. Alternatively, in response to the computing power of the solid-state drive controller being greater than the computing power threshold, a first feature matrix is ​​determined based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted at at least one sampling moment, collected at a preset sampling frequency within the first and second time windows; wherein the second time window is before the first time window, and the duration of the second time window is greater than the duration of the first time window.

[0009] In the above scheme, the temporal feature enhancement layer includes a causal channel attention unit and a causal temporal attention unit. The step of inputting the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted includes: The first feature matrix is ​​input into the causal channel attention unit to obtain the first weight coefficient; the first weight coefficient represents the feature channel weight of the temporal causal model. The first weight coefficient and the first feature matrix are input into the causal temporal attention unit to obtain the first feature map.

[0010] In the above scheme, the causal channel attention unit includes a first global average pooling layer, a first fully connected layer, a second fully connected layer, and a first corrected linear unit; the step of inputting the first feature matrix into the causal channel attention unit to obtain the first weight coefficients includes: The first feature matrix is ​​sequentially input into the first global average pooling layer, the first fully connected layer, the first modified linear unit, and the second fully connected layer to obtain the first weight coefficients, specifically including:

[0011] in, This is the first characteristic matrix; This is the first global average pooling layer; This is the first fully connected layer; This is the second fully connected layer; For the Sigmoid function; This is the first corrected linear unit; The first weighting coefficient has a range of values. .

[0012] In the above scheme, the causal temporal attention unit includes a second global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second modified linear unit; the step of inputting the first weight coefficients and the first feature matrix into the causal temporal attention unit to obtain the first feature map includes: The first feature matrix is ​​weighted based on the first weight coefficient to obtain the channel-weighted features; The transpose of the channel-weighted features is input into the second global average pooling layer, the third fully connected layer, the fourth fully connected layer, and the second modified linear unit to obtain the second weight coefficients; the second weight coefficients characterize the temporal attention weights of the temporal causal model; specifically including:

[0013] The second weighting coefficients after masking are obtained based on the second weighting coefficients and the mask matrix; The channel weighted features are weighted based on the masked second weighting coefficients to obtain the first feature map; in, The transpose of the channel-weighted features; This is the second global average pooling layer; It is the third fully connected layer; It is the fourth fully connected layer; For the Sigmoid function; This is the second corrected linear unit; This is the second weighting coefficient, with a value range of... .

[0014] In the above scheme, the temporal feature extraction layer includes a shallow network and a deep network. Both the shallow and deep networks include multiple bottleneck blocks, and each bottleneck block includes a first pointwise convolutional kernel, a temporal causal convolutional kernel, and a second pointwise convolutional kernel. The step of inputting the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted includes: The first feature map is input into a shallow network to obtain the first temporal feature; The first temporal feature is input into a deep network to obtain the second feature matrix; The number of channels output by the shallow network is less than the number of channels output by the deep network.

[0015] In the above scheme, the multi-task output layer includes a block reclamation priority prediction unit, a page hot / cold classification unit, and a garbage reclamation trigger timing prediction unit. The step of inputting the second feature matrix into the multi-task output layer to obtain the corresponding garbage reclamation decision result for the storage block to be predicted includes: The second feature matrix is ​​input into the block reclamation priority prediction unit to obtain the block reclamation priority; The second feature matrix is ​​input into the page hot / cold classification unit to obtain the page hot / cold classification result for each data page in the storage block to be predicted; The second feature matrix is ​​input into the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted; The garbage collection decision results for the storage block to be predicted include the block collection priority, the page hot / cold classification results of each data page in the storage block to be predicted, and the garbage collection probability of the storage block to be predicted.

[0016] In the above scheme, the step of inputting the second feature matrix into the multi-task output layer to obtain the garbage collection decision result corresponding to the storage block to be predicted includes at least one of the following: The position of the storage block to be predicted in the garbage collection queue is determined based on the block reclamation priority corresponding to the storage block to be predicted. If the position of the storage block to be predicted in the garbage collection queue is greater than the queuing threshold, then the storage block to be predicted is garbage collected. The garbage collection trigger threshold is determined based on the proportion of free storage blocks to all storage blocks; in response to the garbage collection probability of the storage block to be predicted being greater than or equal to the garbage collection trigger threshold, garbage collection is performed on the storage block to be predicted. Based on the page hot / cold classification results of each data page in the storage block to be predicted, data pages whose page hot / cold classification results are hot pages are migrated to free storage blocks.

[0017] The above scheme further includes training a time-series causal model, specifically including: A dataset is determined based on storage block status data, data page status data, input / output load data, and garbage collection history data. Each sample in the dataset includes a storage block with D-dimensional features of a storage block in C time windows, and T consecutive sampling times within each time window. Label each sample in the dataset with block recycling priority label, page hot / cold classification label, and garbage collection trigger label; The training set is determined based on the dataset. The training set is then input into the time-series causal model to obtain the block recycling prediction priority, the page hot / cold prediction classification result, and the garbage recycling prediction probability for each sample. The weights of the time-series causal model are adjusted based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label for each sample, as well as the block recycling prediction priority, page hot / cold prediction classification result, and garbage recycling prediction probability for each data page corresponding to each sample.

[0018] In the above scheme, the block recycling priority label, page hot / cold classification label, and garbage collection trigger label for each sample in the labeled dataset include: The block reclamation priority label for each sample is determined based on the number of invalid data pages, the total number of data pages, the proportion of hot pages, and the average erase error rate. Specifically, it includes:

[0019] in, The number of invalid data pages for sample b. This represents the total number of data pages for sample b. This represents the percentage of hot pages within sample b. The average erase error rate for sample b; , and is the weighting coefficient; where sample b is any sample in the dataset; the block reclamation priority label corresponds one-to-one with the storage block in the dataset; Based on the update information of each sample in the dataset at each time point, determine the page hot / cold classification label of the data page in each sample; the dimension of the page hot / cold classification label is the same as the total number of data pages in each sample; Based on the garbage collection trigger information of each sample in the dataset at each time point, the garbage collection trigger label of each sample is determined; the dimension of the garbage collection trigger label is the same as the sampling time of each sample.

[0020] In the above scheme, adjusting the weights of the time-series causal model based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label for each sample, as well as the block recycling prediction priority corresponding to each sample, the page hot / cold prediction classification result for each data page, and the garbage recycling prediction probability, includes: The first loss is determined based on the block reclamation priority label, block reclamation prediction priority, and total number of samples for each sample. Specifically, it includes:

[0021] in, For the first Block recycling priority label for each sample For the first Block recycling prediction priority for each sample The total number of samples; Based on the page hot / cold classification label of each sample, the page hot / cold prediction classification result of each data page in the sample, and the total number of data pages in each sample, the second loss is determined. Specifically, it includes:

[0022] in, For any sample, the first The page hot / cold category tags for each data page. For any sample, the first The page hot / cold prediction classification results for each data page. This represents the total number of data pages in each sample. The third loss is determined based on the garbage collection trigger label and garbage collection prediction probability for each sample. Specifically, this includes:

[0023] in, The garbage collection trigger tag for the sample, Predict the probability of waste recycling; The training loss of the temporal causal model is determined based on the first loss, the second loss, and the third loss, and the parameters of the temporal causal model are adjusted based on the training loss.

[0024] According to a second aspect of this disclosure, a garbage collection device based on a temporal causal model is provided, wherein the temporal causal model includes a temporal feature enhancement layer, a temporal feature extraction layer, and a multi-task output layer; the device includes: The data matrix determination unit is used to determine the first feature matrix based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted; The feature enhancement unit is used to input the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted. The feature extraction unit is used to input the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted. The inference unit is used to input the second feature matrix into the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted. The processing unit is used to determine, based on the garbage collection decision result, whether to perform garbage collection operation on the storage block to be predicted.

[0025] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods of this disclosure.

[0026] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described in this disclosure.

[0027] This disclosed garbage collection method based on a temporal causal model simultaneously performs multi-task decision-making for GC trigger timing, storage block reclamation priority ranking, and data page hot / cold classification. It can accurately capture the temporal dimension of IO load and data hot / cold evolution patterns, effectively reduce hot data migration, significantly reduce write amplification, and extend the lifespan of NAND flash memory. The model adopts a lightweight temporal causal convolutional structure with low computational overhead, which can be directly deployed on SSD controllers and is suitable for various load scenarios in consumer, industrial, and data center applications without additional computing resources. It also supports adaptive feature pruning based on computing power and is equipped with a complete inference fault tolerance mechanism, which ensures the rationality of GC decisions while maintaining stable SSD continuous read / write performance and wear leveling effect.

[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0029] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which: In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.

[0030] Figure 1 A schematic diagram of an optional process for a garbage collection method based on a time-series causal model provided in an embodiment of this disclosure is shown. Figure 2 This illustration shows another optional process diagram of the garbage collection method based on a time-series causal model provided in an embodiment of this disclosure; Figure 3 A schematic diagram of the process for training a time-series causal model provided in an embodiment of this disclosure is shown; Figure 4 A schematic diagram of the framework of the time-series causal model provided in the embodiments of this disclosure is shown; Figure 5 A schematic diagram of an optional structure of a waste recycling device based on a time-series causal model provided in an embodiment of this disclosure is shown; Figure 6 A schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0031] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0032] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0034] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.

[0035] It should be understood that in the various embodiments of this disclosure, the sequence number of each implementation process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.

[0036] Before providing a further detailed description of the embodiments of this disclosure, the nouns and terms involved in the embodiments of this disclosure will be explained, and the nouns and terms involved in the embodiments of this disclosure shall be interpreted as follows.

[0037] 1) Flash memory.

[0038] Flash memory is a type of non-volatile semiconductor memory that retains data even when power is off. It is mainly divided into two major architectures: NOR flash memory and NAND flash memory. NOR flash memory has a superior random read speed and is often used for device firmware and program code storage; NAND flash memory has a higher degree of integration of stacked storage cells and is the core carrier of high-capacity storage products such as solid-state drives, portable flash drives, and mobile phone built-in storage.

[0039] Based on the number of bits that can be stored in a single storage cell, NAND flash memory has undergone multiple generations of iterations, gradually evolving from Single-Level Cell (SLC) to Quad-Level Cell (QLC).

[0040] 2) Storage blocks Also known as a flash block, it is the smallest physical unit for performing erase operations on NAND flash memory. It consists of several data pages connected in series, and the data page is the smallest unit for reading and writing flash memory.

[0041] A storage block contains a large number of storage cells. As NAND evolves from Single-Level Cell (SLC) to Quad-Level Cell (QLC), flash memory density continues to increase, and the number of data pages that a single block can hold continues to increase.

[0042] 3) Waste recycling Because flash memory cannot be overwritten, when a user wants to write new data, they can only find a free data page to write the new data and point the logical address to the new page. This causes the data in the original data page to expire, forming an invalid page, or "garbage." As more and more expired data accumulates and fills up the storage block space, garbage collection is required if continued writing is needed. Garbage collection involves reading the valid data pages from a storage block, rewriting them to other storage blocks, and then erasing the data in the old storage block, resulting in a new, empty storage block.

[0043] In related technologies, deep learning-based GC optimization methods often employ complex neural network models with large parameters. Due to the limited computing power requirements, these methods are deployed on the host side, increasing additional resource overhead and IO latency. Furthermore, they are prone to preempting the computing and bandwidth resources of front-end services, which in turn exacerbates the performance degradation of SSDs.

[0044] To address the shortcomings of related technologies, this disclosure provides a garbage collection method based on a time-series causal model. It uses lightweight, deep separable convolutional structures adapted to time-series data as its core foundation, combining inverse residuals and linear bottleneck structures to construct a lightweight multi-task neural network model (time-series causal model) for time-series GC scenarios. Through time-series sample construction, a time-specific lightweight causal convolutional architecture, and dynamically adaptive time-series window scaling, a lightweight neural network model targeting SSD GC collection tasks is built. Through multi-task joint learning, it simultaneously achieves three core decision-making processes: GC trigger timing prediction, time-series prioritization of collected blocks, and prediction of page hot / cold trends. Coupled with a full-link time-series fault tolerance and closed-loop optimization mechanism, it ultimately achieves core objectives such as improved GC efficiency, reduced write amplification, minimized impact on overall SSD performance, and extended NAND lifespan.

[0045] This disclosure provides a garbage collection method based on a time-series causal model, which can overcome the static decision-making bottleneck of traditional GC algorithms. It accurately predicts data hot / cold trends, IO load changes, and optimal garbage collection blocks through time-series causal deep learning, reducing write amplification and improving GC execution efficiency. Secondly, it can resolve the core contradiction between the time-series causal model and the SSD controller's computing power constraints, enabling lightweight deployment on embedded devices and basically guaranteeing real-time inference without the need for dedicated hardware acceleration units. Thirdly, this disclosure constructs a closed-loop decision-making system for the entire GC process. The time-series causal model simultaneously covers the time-series prediction of GC trigger timing, time-series sorting of garbage collection block priorities, prediction of effective page hot / cold trends, and wear leveling scheduling, achieving multi-objective joint optimization. Finally, this disclosure designs fault tolerance and backup mechanisms to address risks such as time-series data anomalies and model inference anomalies in embedded environments, ensuring stable operation and data security throughout the SSD's entire lifecycle.

[0046] Figure 1 A schematic diagram of an optional process for a garbage collection method based on a time-series causal model provided in an embodiment of this disclosure is shown, and the steps will be described in detail.

[0047] Step S101: Determine the first feature matrix based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted.

[0048] In some embodiments, the temporal causal model includes at least a temporal feature enhancement layer, a temporal feature extraction layer, and a multi-task output layer.

[0049] In some embodiments, the first feature matrix includes storage block status data, data page status data, input / output load data, and garbage collection history data of storage blocks collected at multiple sampling times. Optionally, the multiple sampling times can be sampling times determined according to a preset sampling frequency within a target time window.

[0050] In some embodiments, the storage block status data includes at least one of the following: the number of valid data pages in the storage block, the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, the storage time of data in the storage block, and a bad block risk marker; the data page status data includes at least one of the following: the storage time of data in the data page, the number of historical writes to the data page, the number of times the data page was accessed within a period, historical hot / cold markers, and data page valid information markers; the input / output load data includes at least one of the following: read IOPS, write IOPS, and average interval between input / output requests; the garbage collection history data includes at least one of the following: historical average garbage collection time, historical average valid page migration, historical average write amplification factor, average garbage collection trigger interval, and storage block wear leveling. The above data are defined as feature data.

[0051] In some embodiments, the dimension of the first feature matrix can be T×D×C, where T is the total number of sampling moments within any time window, D is the total number of each feature dimension in the storage block state data, data page state data, input / output load data, and garbage collection history data, and C is the number of time windows, which is also the input channel dimension.

[0052] For example, using time windows , ... For example, each time window includes T sampling times. The time windows may not overlap completely or may partially overlap. For instance, there may be 0.6T overlap between the time windows.

[0053] Specifically, the carrier can collect feature data of the storage block to be predicted from sampling times t1 to tn. A time window of length T is divided starting from t1 and following the direction of time period growth. , and then with Starting from the time point 0.2T, the time window of length T is divided according to the direction of time period growth. ... until the final time window is tn, resulting in C time windows.

[0054] Each time window includes T×D data points collected from the start time to the end time. That is, the dimension of each time window is T×D, including D-dimensional feature data collected at each of the T sampling times.

[0055] The first feature matrix includes D-dimensional feature data collected at each of the T sampling times within C time windows, i.e., T×D×C.

[0056] In some embodiments, after the carrier (hereinafter referred to as the carrier) that implements the garbage collection based on the time-series causal model collects sampled values ​​of multiple feature dimensions of the storage block to be predicted at a preset frequency within multiple time windows, it also preprocesses the sampled values ​​of multiple feature dimensions, specifically including removing abnormal data and deleting invalid interference features; optionally, after the abnormal data or invalid interference features are deleted, the average value of the time before and after the sampling time corresponding to the abnormal data or invalid interference features can be used to replace the deleted abnormal data or invalid interference features.

[0057] In some alternative embodiments, the carrier can also trim or fuse the first feature matrix based on the computing power of the solid-state drive controller.

[0058] In specific implementation, in response to the computing power of the solid-state drive controller being less than or equal to a computing power threshold, the first feature matrix is ​​pruned. The time-series dimension of the pruned first feature matrix remains unchanged, but the feature dimension is smaller than the feature dimension of the first feature matrix before pruning. For example, reducing the channel dimension or the dimension of feature data (e.g., retaining only 9 feature dimensions) reduces the first feature matrix from T×18×C to T×9×1. Unpruned feature dimensions may include: the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, write IOPS, the proportion of hot data pages in the storage block, the block erase error rate, the number of data page updates, the number of data page accesses within a period, the average garbage collection trigger interval, and the storage block wear leveling. The computing power threshold can be set according to actual needs or experimental results. In this embodiment, the time-series dimension refers to the dimension corresponding to the sampling time; the feature dimension refers to the dimension corresponding to the feature data of the storage block to be predicted, such as storage block status data, data page status data, input / output load data, and garbage collection history data.

[0059] Alternatively, in specific implementation, in response to the computing power of the solid-state drive controller exceeding a computing power threshold, a first feature matrix is ​​determined based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted at at least one sampling moment, collected at a preset sampling frequency within a first and second time window. The second time window precedes the first time window, and the duration of the second time window is greater than the duration of the first time window. The first time window includes 4 sampling periods (4 channels), and the second time window includes 32 sampling periods (32 channels).

[0060] The carrier can be computer programs, electronic circuits, databases, mobile applications, electronic devices, cloud computing platforms, distributed systems, artificial intelligence frameworks, mathematical models, automation tools, and microcontrollers, etc., which are software or hardware capable of implementing algorithms and methods.

[0061] In some optional embodiments, after the first feature matrix is ​​pruned or merged based on the computing power of the solid-state drive controller, the resulting first feature matrix can also be subjected to feature dimensionality upscaling. For example, if the dimension of the first feature matrix is ​​T×18×1, feature dimensionality upscaling can result in a feature matrix with a dimension of T×32×1.

[0062] Step S102: Input the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted.

[0063] In some embodiments, the temporal feature enhancement layer includes a causal channel attention unit and a causal temporal attention unit. The carrier inputs the first feature matrix into the temporal feature enhancement layer to obtain a first feature map corresponding to the storage block to be predicted.

[0064] In specific implementation, the carrier inputs the first feature matrix into the causal channel attention unit to obtain the first weight coefficient; the first weight coefficient represents the feature channel weight of the temporal causal model; the first weight coefficient and the first feature matrix are input into the causal temporal attention unit to obtain the first feature map.

[0065] Furthermore, the causal channel attention unit includes a first global average pooling layer, a first fully connected layer, a second fully connected layer, and a first corrected linear unit; the step of inputting the first feature matrix into the causal channel attention unit to obtain the first weight coefficients includes: (1) in, This is the first characteristic matrix; This is the first global average pooling layer; This is the first fully connected layer; This is the second fully connected layer; For the Sigmoid function; This is the first corrected linear unit; This is the first weighting coefficient, and its value range is... .

[0066] In some embodiments, the causal temporal attention unit includes a second global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second modified linear unit; the step of inputting the first weight coefficients and the first feature matrix into the causal temporal attention unit to obtain the first feature map includes: The first feature matrix is ​​weighted based on the first weight coefficient to obtain the channel-weighted features; specifically including: (2) The channel weighting features The transposed input is fed into the second global average pooling layer, the third fully connected layer, the fourth fully connected layer, and the second modified linear unit to obtain the second weight coefficients; the second weight coefficients characterize the temporal attention weights of the temporal causal model; specifically including:

[0067] in, Transpose of the channel-weighted features; This is the second global average pooling layer; It is the third fully connected layer; It is the fourth fully connected layer; For the Sigmoid function; This is the second corrected linear unit; This is the second weighting coefficient, and its value range is... .

[0068] The second weighting coefficients are obtained based on the second weighting coefficients and the mask matrix; specifically, this includes: (3) Wherein, the mask matrix It can be a lower triangular causality mask matrix. This represents the temporal attention weights after masking, i.e., the second weight coefficients after masking.

[0069] The channel weighted features are weighted based on the second weight coefficient after masking to obtain a first feature map. Specifically, this includes weighting the channel weighted features based on the transpose of the second weight coefficient after masking, specifically including: (4) Among them, the output This is the feature map after dual attention weighting, i.e., the first feature map.

[0070] Step S103: Input the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted.

[0071] In some embodiments, the temporal feature extraction layer includes a shallow network and a deep network, both of which include multiple bottleneck blocks. Each bottleneck block includes a first pointwise convolutional kernel, a temporal causal convolutional kernel, and a second pointwise convolutional kernel. The first and second pointwise convolutional kernels can be 1×1 convolutional kernels; the temporal causal convolutional kernel can be a 3×1 convolutional kernel.

[0072] Each bottleneck block is a 1×1 pointwise convolutional channel expansion, whose output serves as the input for 3×1 temporal causal deep convolutional feature extraction, and whose output serves as the input for 1×1 pointwise convolutional channel compression.

[0073] In some embodiments, the carrier inputs the first feature map into a shallow network to obtain a first temporal feature; and inputs the first temporal feature into a deep network to obtain a second feature matrix; wherein the number of channels output by the shallow network is less than the number of channels output by the deep network.

[0074] In some embodiments, the number of bottleneck blocks in a shallow network is less than the number of bottleneck blocks in a deep network. The shallow network uses fewer output channels to reduce computation; the deep network uses more output channels to ensure temporal feature extraction capabilities. Residual connections are enabled only when the number of input and output channels and feature sizes are exactly the same to avoid feature degradation; the output uses a linear activation function to avoid the ReLU activation function destroying information in low-dimensional features.

[0075] Step S104: Input the second feature matrix into the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted.

[0076] In some embodiments, the multi-task output layer includes a block recycling priority prediction unit, a page hot / cold classification unit, and a garbage recycling trigger timing prediction unit.

[0077] In specific implementation, the carrier inputs the second feature matrix into the block reclamation priority prediction unit to obtain the block reclamation priority. Specifically, the block reclamation priority prediction unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same dimension as the number of storage blocks, and a Sigmoid function. Its output is... , The total number of storage blocks is a fixed value when an SSD is shipped from the factory. Each value corresponds to the block reclamation priority score of the corresponding storage block, with a score range of [missing value]. The closer a value is to 0, the lower its priority for recycling; the closer a value is to 1, the higher its priority for recycling.

[0078] In practice, the carrier inputs the second feature matrix into the page hot / cold classification unit to obtain the page hot / cold classification result for each data page in the storage block to be predicted. Specifically, the page hot / cold classification unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same dimension as the number of data pages in the storage block, and a Sigmoid function. Its output is the page hot / cold classification result. , The number of data pages in the storage block, if If the value exceeds the hot / cold threshold, it is determined to be a hot page. If the value is less than the hot / cold threshold, it is determined to be a cold page. The hot / cold threshold can be set according to actual needs or experimental results, for example, set to 0.5.

[0079] In specific implementation, the carrier inputs the second feature matrix into the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted. Specifically, the garbage collection trigger timing prediction unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same number of sampling times, and a Sigmoid function. Its output is a vector, with the number of elements equal to the number of sampling times, and each element representing the probability of triggering garbage collection at the corresponding sampling time. If the garbage collection probability is greater than or equal to the garbage collection trigger threshold, then garbage collection processing is performed on the storage block to be predicted.

[0080] Step S105: Based on the garbage collection decision result, determine whether to perform garbage collection operation on the storage block to be predicted.

[0081] In some embodiments, the carrier determines the position of the storage block to be predicted in the garbage collection queue based on the block reclamation priority corresponding to the storage block to be predicted. If the position of the storage block to be predicted in the garbage collection queue is greater than the queuing threshold, then the storage block to be predicted is garbage collected.

[0082] In some embodiments, the carrier determines a garbage collection trigger threshold based on the proportion of free storage blocks to all storage blocks; in response to the garbage collection probability of the storage block to be predicted being greater than or equal to the garbage collection trigger threshold, garbage collection is performed on the storage block to be predicted.

[0083] In some embodiments, the carrier migrates data pages that are classified as hot pages to free storage blocks based on the page hot / cold classification results of each data page in the storage block to be predicted, so as to balance the wear of each storage block.

[0084] Thus, the garbage collection method based on the temporal causal model provided in this disclosure synchronously completes GC trigger timing prediction, storage block reclamation priority sorting, and data page hot / cold classification multi-task decision-making through the temporal causal model. It can accurately capture the temporal dimension of IO load and data hot / cold evolution patterns, effectively reduce hot data migration, significantly reduce write amplification, and extend the lifespan of NAND flash memory. The model adopts a lightweight temporal causal convolutional structure with low computational overhead, which can be directly deployed on SSD controllers and is suitable for various load scenarios in consumer, industrial, and data center applications without additional computing resources. At the same time, it supports adaptive feature pruning based on computing power and is equipped with a complete inference fault tolerance mechanism, which ensures the rationality of GC decisions while stably maintaining the continuous read / write performance and wear leveling effect of SSD.

[0085] Figure 2 This illustration shows another optional process diagram of the garbage collection method based on a time-series causal model provided in an embodiment of this disclosure. Figure 3This illustration shows a flowchart of training a time-series causal model according to an embodiment of the present disclosure, which will be combined with Figure 2 and Figure 3 Please provide an explanation.

[0086] Figure 2 The diagram illustrates the method for training a time-series causal model. It should be noted that the structure of the time-series causal model is exactly the same in the training and inference processes. That is, the structure of the time-series causal model used in steps S101 to S105 is the same as the structure of the time-series causal model trained in steps S201 to S205.

[0087] Step S201, Data Acquisition and Processing.

[0088] In some embodiments, such as Figure 3 As shown, the carrier collects raw data from the solid-state drive and filters feature data from multiple dimensions that are strongly correlated with GC; the collected raw data is then noise-removed and standardized, and a dataset is constructed according to the time series growth direction. .

[0089] In some embodiments, the garbage disposal method or training method based on the time-series causal model described herein can be implemented through the main control chip. When collecting feature data of each storage block, sampling is performed based on a fixed sampling period. The sampling period can be set according to actual needs, such as 1ms, that is, feature data of the storage block is collected every 1ms. Optionally, the sampling period can be bound to a timer of the main control chip.

[0090] In some embodiments, the feature data includes at least one of the following: storage block status data (number of valid data pages, number of invalid data pages in the storage block, cumulative number of erases in the storage block, storage time of data in the storage block, bad block risk marker), data page status data (storage time of data in the data page, number of historical writes to the data page, number of times the data page was accessed within a period, historical hot / cold marker, and data page valid information marker), input / output load data (at least one of the following: read IOPS, write IOPS, and average interval of input / output requests), and garbage collection history data (at least one of the following: historical average garbage collection time, historical average valid page migration, historical average write amplification factor, average garbage collection trigger interval, and storage block wear leveling).

[0091] The above feature data is collected in multiples of the period as the statistical unit, such as 100ms (100 sampling periods). In addition to the above feature data, feature data can be added or deleted according to the actual application scenario.

[0092] In some embodiments, the carrier determines samples based on the collected feature data. Each sample must adhere to the core principle of causal alignment, i.e., using historical temporal features at time [tT,t] to predict the label within the future window [t,t+T_predict]. Here, T is the length of the input temporal window (i.e., the time window), and T_predict is the length of the label prediction window, which can be configured from 10 to 1000 depending on the actual training effect and usage scenario (and also depends on the data sampling period). Samples are constructed using a sliding window with a sliding step size of T / 5 to ensure the temporal continuity and feature overlap of adjacent samples, avoiding the loss of temporal information. A single sample consists of 18-dimensional features from T consecutive sampling periods, ultimately reconstructed into a T×18×1 temporal-feature two-dimensional matrix (T is the temporal dimension, 18 is the feature dimension, single-channel input model), serving as the standard input format for the temporal causal model. During training, samples are shuffled only according to scene temporal blocks to ensure temporal continuity within a single sample.

[0093] In some embodiments, the garbage disposal method or training method based on the time-series causal model described herein is applicable to three major categories of SSD application scenarios: consumer-grade, industrial-grade, and data center-grade. The amount of data collected and written in a single scenario is no less than 10TB, and it needs to cover the entire lifecycle of AND flash memory from its initial state to 90% of the nominal PE cycles. Among them, the consumer-grade scenario includes four types: office document reading and writing, video streaming playback, game installation and running, and system disk random reading and writing; the industrial-grade scenario includes four types: 7×24-hour continuous writing, wide-temperature environment reading and writing, burst high-load IO, and low-power intermittent reading and writing; and the data center scenario includes four types: 4K random reading and writing, multi-stream concurrent writing, cold and hot data tiered storage, and snapshot backup continuous writing.

[0094] In some embodiments, after the carrier collects feature data, it will also preprocess the feature data to obtain usable feature data. Specifically, this includes: removing invalid feature data or replacing it with the mean of the preceding and following periods or padding it forward, so as to facilitate the training and prediction of the subsequent neural network model; performing Z-score standardization on multiple time-series feature data, and directly retaining the original 0 / 1 values ​​for binary discrete classification features (such as bad block labels, valid / invalid labels), so that each feature has a uniform scale of computational dimension, which is conducive to the model converging as quickly as possible during the training process.

[0095] Among them, hard constraint rules are set based on SSD hardware and business logic. Samples that do not meet the rules are directly eliminated. For example, data with negative PE counts or exceeding 120% of the NAND nominal maximum PE counts, invalid data pages exceeding the total number of pages in the block, the sum of the number of valid data pages and invalid data pages not equaling the total number of data pages in the storage block, IOPS exceeding the nominal limit of SSD hardware, etc., are considered abnormal characteristic data.

[0096] Step S202, input the time series features.

[0097] In some embodiments, the carrier uses minimizing write amplification after GC operations, minimizing overall IO performance loss of solid-state drives, leveling block wear, and maximizing NAND lifespan as core optimization objectives to complete the labeling of each sample. The labeling benchmark is the optimal GC decision result measured in the entire scenario. In specific implementation, the carrier determines the block reclamation priority label for each sample based on the number of invalid data pages in the sample, the total number of data pages, the proportion of hot pages, and the average erase error rate, specifically including: (5) in, The number of invalid data pages for sample b. This represents the total number of data pages for sample b. This represents the percentage of hot pages within sample b. The average erase error rate for sample b; , and where is the weighting coefficient; where sample b is any sample in the dataset; the block reclamation priority label corresponds one-to-one with the storage block in the dataset, that is, each sample (storage block) corresponds to a block reclamation priority label. , and The values ​​can be configured to 0.5, 0.4, and 0.1, with a sum of 1. This setting of weighting coefficients is intended to highlight the contribution of the invalid data page ratio and the hot / cold attribute to garbage collection priority. The block collection priority label is used to characterize the trend of block revenue changes within a time window.

[0098] In practice, based on the update information of each sample in the dataset at each time step, the hot / cold classification label of the data pages in each sample is determined; the dimension of the hot / cold classification label is the same as the total number of data pages in each sample. Specifically, the hot / cold classification label of a hot page (hot data page) is 1, indicating that it will be updated within the future window [t, t+T_predict]; the hot / cold classification label of a cold page (cold data page) is 0, indicating that it will not be updated within the future window [t, t+T_label].

[0099] In practice, the garbage collection trigger label for each sample is determined based on the garbage collection trigger information of each sample in the dataset at each time point; the dimension of the garbage collection trigger label is the same as the sampling time of each sample. Specifically, since the feature data is the collected historical data, the actual collection time of each data block can be known. Therefore, the garbage collection trigger label at the time when GC is actually triggered is marked as 1, indicating that GC needs to be triggered within the future window [t, t+T_predict], and the garbage collection trigger label at the time when GC is not triggered is marked as 0.

[0100] In some embodiments, the carrier determines the set consisting of all samples as the dataset; the total valid time-series samples are divided into a training set, a validation set, and a test set in strict time-growth order, with a division ratio of 7:2:1, that is, the first 70% of the time-series samples are collected as the training set. The middle 20% is the validation set. The last 10% is the test set. An additional cross-scenario time-series validation set was set up, using full-time-series data from three unfamiliar scenarios that were not used in the training, to verify the generalization ability of the time-series causal model. The training set was used to train the time-series causal model, the validation set was used to verify the model's training effect after each training round, and the test set was used to simulate real-world scenarios to judge the model's training results.

[0101] In some embodiments, after determining the training set, validation set, and test set, the inputs to the time-series causal model are determined based on the training set, validation set, and test set, respectively.

[0102] Specifically, the input channel dimension is set to C (values ​​[1, 8], 1 can be used to reduce the consumption of computing resources), consisting of time series data starting from C adjacent time nodes; the training set is... Divide the data into sequences of length T according to the growth direction of the time period t, with a division step size of T / 5. The input dataset has D=18 feature data dimension attributes (consistent with the feature dimensions in the data collection stage). Then, the input data dimension of the spatiotemporal causal model is T×D×C.

[0103] In some embodiments, such as Figure 3 As shown, after determining the input of the temporal causal model, the input data of the spatiotemporal causal model can be further processed according to the computing power of the main control chip, including pruning or fusion.

[0104] In practical implementation, when the main control chip's computing power is strained, the core nine-dimensional feature data are retained. Feature data with high contribution can be selected for retention, while the nine secondary-dimensional feature data are pruned. At the same time, the timing period T is reduced and the channel dimension is configured to 1, reducing the size of the input data of the spatiotemporal causal model to T×9×1. The core feature data to be retained includes: the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, write IOPS, the proportion of hot data pages in the storage block, the block erase error rate, the number of data page updates, the number of times data pages are accessed within the period, the average garbage collection trigger interval, and the wear leveling of the storage block.

[0105] Alternatively, in specific implementations, the main control chip can be fully utilized during idle periods, and real-time time series features of a short window (4 sampling periods) and historical trend features of a long window (32 sampling periods) can be input simultaneously. These features are then fused through a feature splicing layer and input into the model, taking into account both real-time load response and long-term trend prediction, thereby improving prediction accuracy with a controllable increase in computational load.

[0106] Step S203: Input the sample data into the time-series causal model to obtain the block recycling prediction priority, the page hot / cold prediction classification result, and the garbage recycling prediction probability for each sample.

[0107] In some embodiments, such as Figure 3 As shown, the temporal causal model includes a feature dimensionality enhancement layer, a temporal feature enhancement layer, and a temporal feature extraction layer. The feature dimensionality enhancement layer includes multiple sets of Conv1D convolutional kernels with a size of 1×1.

[0108] In some embodiments, after the carrier further processes the input data of the spatiotemporal causal model based on the computing power of the main control chip to obtain a sample matrix, it can also increase the dimensionality of the sample matrix based on multiple sets of Conv1D pointwise convolutions with a kernel size of 1×1. The increased-dimensional sample matrix is ​​then input into the temporal feature enhancement layer and the temporal feature extraction layer, respectively.

[0109] In some embodiments, the temporal feature enhancement layer includes a causal channel attention unit and a causal temporal attention unit.

[0110] In some embodiments, considering the significant differences in the contribution of core features in GC scenarios, a dual attention enhancement unit combining causal channel attention and causal temporal attention is designed. Feature weighting is achieved through two lightweight fully connected layers, improving feature representation capabilities with minimal impact on overall inference time, while strictly adhering to temporal causality rules. Specifically: The causal channel attention unit is used to perform global average pooling on the input features after dimensionality upscaling. It generates weight coefficients for each feature channel through two fully connected layers, which are then multiplied by the original features to complete the weighting. The formula is as follows: (6) in, As input features, This is the first global average pooling layer during the training process; This is the first fully connected layer during the training process; This is the second fully connected layer during the training process; For the Sigmoid function; This is the first corrected linear unit during the training process; The first weighting coefficient for prediction has a range of values. .

[0111] Based on the predicted first weighting coefficient For input features By performing weighting, the predicted channel weighted features are obtained. Specifically, it includes: (7) The causal temporal attention unit is used to weight features from different sampling times within a temporal window, strictly adhering to causal rules: weights are assigned only to temporal features at the current time and earlier, and the weights of features at future times are forcibly set to 0, thus strengthening the contribution of recent temporal features and weakening the noise interference of distant features. The formula is as follows: (8) in, This is the transpose of the predicted channel-weighted features; This is the second global average pooling layer during the training process; This is the third fully connected layer during the training process; This is the fourth fully connected layer during the training process; For the Sigmoid function; This is the second corrected linear unit during the training process; The second weighting coefficient for prediction has a range of values ​​of [value range missing]. .

[0112] Based on the predicted second weighting coefficient and mask matrix Predicted second weight coefficients after obtaining the mask Specifically, this includes: (9) Wherein, the mask matrix It can be a lower triangular causality mask matrix. This represents the temporal attention weights for the masked predictions, i.e., the second weight coefficients for the masked predictions.

[0113] Predicted second weighting coefficients after masking The predicted channel weighted features Weighting is performed to obtain the first sample image. Specifically, this involves weighting the channel-weighted features based on the transpose of the second weight coefficients after masking. (10) Among them, the output This is the sample image after double attention weighting, i.e., the first sample image.

[0114] The first sample image is then input into the temporal feature extraction layer to obtain the first sample matrix. To address the temporal feature characteristics and embedded computing power constraints of the GC scenario, the temporal inverse residual linear bottleneck block (hereinafter referred to as the bottleneck block) is optimized. In this embodiment, the temporal feature extraction layer includes multiple temporal inverse residual linear bottleneck blocks. Each bottleneck block employs a structure of 1×1 pointwise convolutional channel expansion, 3×1 temporal causal deep convolutional feature extraction, and 1×1 pointwise convolutional channel compression. To address the sparsity of GC features, the shallow network (the first 3 bottleneck blocks) uses an output channel count <= 32 to reduce computation, while the deep network (the last 4 bottleneck blocks) uses an output channel count >= 96 to ensure temporal feature extraction capability. Residual connections are enabled only when the input and output channel counts and feature sizes are completely identical to avoid feature degradation. The output uses a linear activation function to avoid the ReLU activation function destroying information in low-dimensional features.

[0115] Specifically, the bottleneck block uses a 3×1 temporal causal convolution kernel to implement temporal causal deep convolution. It only performs sliding convolution in the temporal dimension (T), with each input channel corresponding to a separate convolution kernel. It only completes the extraction of temporal features for a single channel and does not perform cross-channel feature dimension calculations (to avoid incorporating future temporal features into the current calculation). The padding method is forward causal padding, which only adds zeros to the left side of the temporal dimension (historical direction) to ensure that the temporal length of the convolution output is consistent with the input and does not introduce future data.

[0116] In addition, the bottleneck block uses a 1×1 convolution kernel to perform pointwise convolution and channel fusion in the feature dimension to generate a new feature map without changing the length of the temporal dimension, thereby reducing the amount of computation and model parameters.

[0117] The computational cost of standard convolution is ,in, The spatial size of the convolution kernel. The number of channels in the input feature map. The number of channels in the output feature map. , The height and width of the input feature map are given. The computational complexity of the temporal causal depthwise separable convolution for adaptive modification is divided into: 1. Temporal causal depthwise convolution. The computational cost of pointwise convolution is The total computational cost is .

[0118] The computational compression ratio of temporally causal depthwise separable convolution relative to standard convolution is: When this disclosure sets a 3x1 temporal causal convolution kernel and configures the number of channels in the output feature map to 32, the theoretical compression ratio is 14%, which significantly reduces the computational load. This makes it suitable for scenarios with limited computing power resources of solid-state drive controllers, while also meeting the requirements of temporal causality. Temporal causal depthwise separable convolutions are used throughout the lightweight GC neural network model to replace standard convolutions.

[0119] Step S204: Based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label of each sample, as well as the block recycling prediction priority, page hot / cold prediction classification result, and garbage recycling prediction probability of each data page corresponding to each sample, adjust the weights of the time-series causal model.

[0120] In some embodiments, a three-branch parallel multi-task output head is designed to address the characteristics of solid-state drive (SSD) garbage collection (GC) tasks. This head shares the main timing features, avoiding computational waste caused by redundant calculations across multiple models, while simultaneously covering the timing decision requirements of the entire GC process. The output dimensions of all three branches are aligned with the fixed hardware parameters of the SSD at the factory, ensuring the stability of the network structure. Specifically, the multi-task output layer includes three independent units: a block reclamation priority prediction unit, a page hot / cold classification unit, and a garbage collection trigger timing prediction unit.

[0121] In specific implementation, the carrier inputs the first sample matrix into the block reclamation priority prediction unit to obtain the block reclamation prediction priority. Specifically, the block reclamation priority prediction unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same dimension as the number of storage blocks, and a Sigmoid function. Its output is the block reclamation prediction priority. , The total number of storage blocks is a fixed value when an SSD is shipped from the factory. Each value corresponds to the block reclamation priority score of the corresponding storage block, with a score range of [missing value]. The closer a value is to 0, the lower its priority for recycling; the closer a value is to 1, the higher its priority for recycling.

[0122] The first loss is determined based on the block reclamation priority label, block reclamation prediction priority, and total number of samples for each sample. Specifically, it includes: (11) in, For the first Block recycling priority label for each sample For the first Block recycling prediction priority for each sample The total number of samples.

[0123] In practice, the carrier inputs the first sample matrix into the page hot / cold classification unit to obtain the page hot / cold prediction classification result for each data page in the storage block to be predicted. Specifically, the page hot / cold classification unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same dimension as the number of data pages in the storage block, and a Sigmoid function. Its output is the page hot / cold prediction classification result. , The number of data pages in the storage block, if If the value exceeds the hot / cold threshold, it is determined to be a hot page. If the value is less than the hot / cold threshold, it is determined to be a cold page. The hot / cold threshold can be set according to actual needs or experimental results, for example, set to 0.5.

[0124] Based on the page hot / cold classification label of each sample, the page hot / cold prediction classification result of each data page in the sample, and the total number of data pages in each sample, the second loss is determined. Specifically, this includes: (12) in, For any sample, the first The page hot / cold category tags for each data page. For any sample, the first The page hot / cold prediction classification results for each data page. This represents the total number of data pages in each sample.

[0125] In specific implementation, the carrier inputs the first sample matrix into the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted. Specifically, the garbage collection trigger timing prediction unit includes a 128-dimensional fully connected layer, ReLU, Dropout (0.2), a fully connected layer with the same number of sampling times, and a Sigmoid function. Its output is a vector, with the number of elements equal to the number of sampling times, and each element representing the probability of triggering garbage collection at the corresponding sampling time. If the predicted garbage collection probability is greater than or equal to the garbage collection trigger threshold, then garbage collection processing is performed on the storage block to be predicted.

[0126] The third loss is determined based on the garbage collection trigger label and garbage collection prediction probability for each sample. Specifically, this includes: (13) in, The garbage collection trigger tag for the sample, Predict the probability of waste recycling; The training loss of the temporal causal model is determined based on the first loss, the second loss, and the third loss. The parameters of the temporal causal model are adjusted based on the training loss. The training loss of the temporal causal model includes: (14) in, The loss weights for the three branches are initially set to typical values. , , During training, after reaching the set number of training epochs, it is determined whether the model loss is less than or equal to a loss threshold. If it is less than or equal to the loss threshold, the training of the temporal causal model is complete; otherwise, error backpropagation is performed to continue updating the weights of the temporal causal model using training data. The loss threshold can be set according to actual needs or experimental results.

[0127] In practice, it is determined whether the training termination conditions are met. Training termination conditions may include reaching the required number of training epochs or the prediction error rate being less than or equal to an error rate threshold. The error rate threshold can be set based on actual needs or experimental results.

[0128] If the training termination condition is met, the training is considered complete. Relevant data from the solid-state drive is periodically collected, and the parameters of the time-series causal model are continuously fine-tuned. If the training termination condition is not met, then based on… Backpropagation updates the weights of the time-series causal model.

[0129] If the training termination condition is not met, backpropagation of error is performed based on the training loss of the time-series causal model to update the weights of the time-series causal model.

[0130] In some embodiments, after the training of the temporal causal model is completed and the model inference stage is performed, a garbage collection execution decision is generated for the model inference result. Furthermore, a fault tolerance and backup mechanism is included, specifically: setting a triple anomaly detection rule, namely, input feature validity verification, inference output range verification, and inference timeout monitoring; the inference timeout is set to 500us (which can be customized according to computing power and the size of the temporal causal model), and if it exceeds this timeout, the inference process is immediately terminated; when the temporal causal model continuously exhibits temporal prediction deviations exceeding the threshold, the temporal dimension is first reduced, i.e., the input feature matrix is ​​pruned in the temporal dimension, and a lightweight model is switched to; if the anomaly persists, a seamless switch to a traditional cost-benefit GC algorithm is made as a backup to ensure uninterrupted GC functionality and to avoid impacting host services; simultaneously, anomaly logs are recorded, including the anomaly type, input temporal features, and output results, for subsequent model optimization.

[0131] In addition, it includes a fault tolerance mechanism for model updates and version management: When the firmware of the time-series causal model is updated, the new time-series causal model parameters are first written to the backup storage area. After verifying that the parameter integrity and inference accuracy meet the requirements, the parameters of the main time-series causal model are replaced. If an abnormality occurs during the update process, it is immediately rolled back to the original stable parameters to ensure the availability of the time-series causal model. The time-series causal model update is only performed when the SSD is idle, there are enough free blocks, and there are no host IO requests. During the update process, it is prohibited to respond to host write requests to avoid damage to the time-series causal model caused by update interruption.

[0132] In actual business processes, data related to the solid-state drive (SSD) models used in the business can be selected to continuously fine-tune the time-series causal model. When new SSD data becomes available, this data can be parsed, cleaned, preprocessed, and added to the training set. Hyperparameters can be adjusted adaptively, and historical models can be used as pre-trained models to continuously refine the time-series model.

[0133] Figure 4 A schematic diagram of the framework of the time-series causal model provided in the embodiments of this disclosure is shown.

[0134] like Figure 4 As shown, the temporal causal model includes a data input layer, a feature dimensionality enhancement layer, a temporal feature enhancement layer, a temporal feature extraction layer, a feature fusion layer, and a multi-task output layer.

[0135] The data input layer, corresponding to steps S201 to S202, is mainly used for preprocessing, cropping, fusing, and labeling the collected feature data. Its output dimension is T×18×1.

[0136] The feature upscaling layer includes multiple sets of Conv1D convolutional kernels with a size of 1×1, which are used to upscale the output of the data input layer, with an output dimension of T×32×1.

[0137] The temporal feature enhancement layer corresponds to the temporal feature enhancement layer described in step S203, and its output dimension is T×32×1.

[0138] The temporal feature extraction layer corresponds to the temporal feature extraction layer described in step S203, and includes a multi-layer stack of temporal inverse residual linear bottleneck blocks. That is, it consists of a shallow network and a deep network, each containing multiple bottleneck blocks, and its output dimension is... ×32×96. The feature fusion layer fuses the first sample matrix output by the temporal feature extraction layer, and its output dimension is ×32×96. ×32×192.

[0139] The multi-task output layer is used to independently predict the output of the feature fusion layer in three branches.

[0140] Thus, this disclosed embodiment addresses numerous issues in traditional algorithms and existing neural network solutions for solid-state drive (SSD) garbage collection tasks by designing a temporal causal model. It demonstrates significant improvements across five dimensions: collection efficiency, deployment adaptability, decision rationality, operational reliability, and environmental adaptability. Regarding garbage collection efficiency, compared to traditional algorithms, the solution accurately predicts data hot / cold trends, optimal collection blocks, and GC triggering timing through temporal causal modeling. This effectively reduces invalid migration of hot data, significantly lowers write amplification, improves GC execution efficiency, and reduces NAND flash memory erase / write cycles, further extending flash memory lifespan. In terms of deployment adaptability, the temporal causal model uses temporal causal depthwise separable convolution as its core lightweight foundation. Combined with inverse residuals and linear bottleneck structures to optimize temporal feature extraction capabilities, it completes the entire chain construction for SSD GC collection scenarios. While significantly reducing computational and parameter quantities, it fully adapts to the temporal decision-making requirements of multi-task GC, meets the computing power constraints of SSD controllers, enables lightweight embedded deployment, and essentially guarantees real-time inference without requiring dedicated hardware acceleration units. In terms of decision-making rationality, based on the construction of time-series samples with strict causal alignment and a dedicated convolutional architecture, a single model can simultaneously complete three core decisions: prediction of GC trigger timing, time-series sorting of garbage collection block priority, and prediction of page hot / cold trends. This constructs a closed-loop decision-making system for the entire GC process, avoiding the additional computing power overhead caused by repeated calculations by multiple models, improving decision-making efficiency, achieving multi-objective joint optimization, and significantly enhancing the rationality of the GC collection strategy. In terms of operational reliability, the solution has strong adaptability to all scenarios and a complete inference fault-tolerant backup mechanism, which can cover scenarios such as time-series data anomalies, model inference result anomalies, and inference timeouts. In case of anomalies, it can seamlessly switch to a traditional cost-benefit GC algorithm as a backup, ensuring that the GC function is not interrupted and does not affect the host business. At the same time, it has a model update and version management fault tolerance mechanism. When updating the model firmware, a strategy of replacing after verification in the backup area is adopted. If the update is abnormal, it can be automatically rolled back to a stable version, ensuring the stable operation and data security of the SSD throughout its entire life cycle. In terms of environmental adaptability, the solution supports continuous model maintenance and iterative optimization. Newly generated, cleaned and pre-processed historical data from solid-state drives in actual business operations can be added to the training dataset. The historical model can be used as a pre-trained model to fine-tune and adapt the time-series causal model, which can provide better adaptability and robustness to the actual use environment of solid-state drives.

[0141] Figure 5 A schematic diagram of an optional structure of a garbage collection device based on a time-series causal model provided in an embodiment of this disclosure is shown. The time-series causal model includes a time-series feature enhancement layer, a time-series feature extraction layer, and a multi-task output layer. The device includes a data matrix determination unit, a feature enhancement unit, a feature extraction unit, an inference unit, and a processing unit.

[0142] The data matrix determination unit is used to determine the first feature matrix based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted. The feature enhancement unit is used to input the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted; The feature extraction unit is used to input the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted; The inference unit is used to input the second feature matrix into the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted; The processing unit is used to determine, based on the garbage collection decision result, whether to perform garbage collection operation on the storage block to be predicted.

[0143] In some embodiments, the dimension of the first feature matrix is ​​T×D×C, where C represents the number of time windows or the dimension of input channels; D represents the total dimension of the storage block status data, data page status data, input / output load data and garbage collection history data of the storage block to be predicted; and T represents the total sampling time of any time window. The storage block status data includes at least one of the following: the number of valid data pages in the storage block, the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, the storage time of data in the storage block, and bad block risk markers. The data page status data includes at least one of the following: the storage time of data in the data page, the number of historical writes to the data page, the number of times the data page was accessed within a period, historical hot / cold markers, and data page valid information markers. The input / output load data includes at least one of the following: read IOPS, write IOPS, and average interval between input / output requests. The garbage collection history data includes at least one of the following: historical average garbage collection time, historical average valid page migration, historical average write amplification factor, average garbage collection trigger interval, and storage block wear leveling.

[0144] The data matrix determination unit is also used to determine the computing power of the solid-state drive controller; In response to the computing power of the solid-state drive controller being less than or equal to the computing power threshold, the first feature matrix is ​​pruned; wherein, the temporal dimension of the pruned first feature matrix remains unchanged, but the feature dimension is smaller than the feature dimension of the first feature matrix before pruning. Alternatively, in response to the computing power of the solid-state drive controller being greater than the computing power threshold, a first feature matrix is ​​determined based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted at at least one sampling moment, collected at a preset sampling frequency within the first and second time windows; wherein the second time window is before the first time window, and the duration of the second time window is greater than the duration of the first time window.

[0145] In some embodiments, the temporal feature enhancement layer includes a causal channel attention unit and a causal temporal attention unit. The feature enhancement unit is specifically used to input a first feature matrix into the causal channel attention unit to obtain a first weight coefficient; the first weight coefficient represents the feature channel weight of the temporal causal model. The first weight coefficient and the first feature matrix are input into the causal temporal attention unit to obtain the first feature map.

[0146] In some embodiments, the causal channel attention unit includes a first global average pooling layer, a first fully connected layer, a second fully connected layer, and a first corrected linear unit; the feature enhancement unit is specifically used for The first feature matrix is ​​sequentially input into the first global average pooling layer, the first fully connected layer, the first modified linear unit, and the second fully connected layer to obtain the first weight coefficients, specifically including:

[0147] in, This is the first characteristic matrix; This is the first global average pooling layer; This is the first fully connected layer; This is the second fully connected layer; For the Sigmoid function; This is the first corrected linear unit; The first weighting coefficient has a range of values. .

[0148] In some embodiments, the causal temporal attention unit includes a second global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second modified linear unit; the feature enhancement unit is specifically used to weight the first feature matrix based on a first weight coefficient to obtain channel-weighted features; The transpose of the channel-weighted features is input into the second global average pooling layer, the third fully connected layer, the fourth fully connected layer, and the second modified linear unit to obtain the second weight coefficients; the second weight coefficients characterize the temporal attention weights of the temporal causal model; specifically including:

[0149] The second weighting coefficients after masking are obtained based on the second weighting coefficients and the mask matrix; The channel weighted features are weighted based on the masked second weighting coefficients to obtain the first feature map; in, Transpose of the channel-weighted features; This is the second global average pooling layer; It is the third fully connected layer; It is the fourth fully connected layer; For the Sigmoid function; This is the second corrected linear unit; This is the second weighting coefficient, with a value range of... .

[0150] In some embodiments, the temporal feature extraction layer includes a shallow network and a deep network, both of which include multiple bottleneck blocks. Each bottleneck block includes a first pointwise convolutional kernel, a temporal causal convolutional kernel, and a second pointwise convolutional kernel. The feature extraction unit is specifically used to input the first feature map into the shallow network to obtain the first temporal feature. The first temporal feature is input into a deep network to obtain the second feature matrix; The number of channels output by the shallow network is less than the number of channels output by the deep network.

[0151] In some embodiments, the multi-task output layer includes a block recycling priority prediction unit, a page hot / cold classification unit, and a garbage recycling trigger timing prediction unit. The inference unit is specifically used to input the second feature matrix into the block recycling priority prediction unit to obtain the block recycling priority. The second feature matrix is ​​input into the page hot / cold classification unit to obtain the page hot / cold classification result for each data page in the storage block to be predicted; The second feature matrix is ​​input into the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted; The garbage collection decision results for the storage block to be predicted include the block collection priority, the page hot / cold classification results of each data page in the storage block to be predicted, and the garbage collection probability of the storage block to be predicted.

[0152] The inference unit is specifically used for at least one of the following: The position of the storage block to be predicted in the garbage collection queue is determined based on the block reclamation priority corresponding to the storage block to be predicted. If the position of the storage block to be predicted in the garbage collection queue is greater than the queuing threshold, then the storage block to be predicted is garbage collected. The garbage collection trigger threshold is determined based on the proportion of free storage blocks to all storage blocks; in response to the garbage collection probability of the storage block to be predicted being greater than or equal to the garbage collection trigger threshold, garbage collection is performed on the storage block to be predicted. Based on the page hot / cold classification results of each data page in the storage block to be predicted, data pages whose page hot / cold classification results are hot pages are migrated to free storage blocks.

[0153] In some embodiments, the waste recycling device based on a time-series causal model may further include a training unit.

[0154] The training unit is used to determine the dataset based on storage block status data, data page status data, input / output load data, and garbage collection history data; each sample in the dataset includes a storage block in C time windows, and D-dimensional features of a storage block in each time window for T consecutive sampling times; Label each sample in the dataset with block recycling priority label, page hot / cold classification label, and garbage collection trigger label; The training set is determined based on the dataset. The training set is then input into the time-series causal model to obtain the block recycling prediction priority, the page hot / cold prediction classification result, and the garbage recycling prediction probability for each sample. The weights of the time-series causal model are adjusted based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label for each sample, as well as the block recycling prediction priority, page hot / cold prediction classification result, and garbage recycling prediction probability for each data page corresponding to each sample.

[0155] The training unit is specifically used to determine the block recycling priority label for each sample based on the number of invalid data pages in the sample, the total number of data pages, the proportion of hot pages, and the average erase error rate. Specifically, it includes:

[0156] in, The number of invalid data pages for sample b. This represents the total number of data pages for sample b. This represents the percentage of hot pages within sample b. The average erase error rate for sample b; , and is the weighting coefficient; where sample b is any sample in the dataset; the block reclamation priority label corresponds one-to-one with the storage block in the dataset; Based on the update information of each sample in the dataset at each time point, determine the page hot / cold classification label of the data page in each sample; the dimension of the page hot / cold classification label is the same as the total number of data pages in each sample; Based on the garbage collection trigger information of each sample in the dataset at each time point, the garbage collection trigger label of each sample is determined; the dimension of the garbage collection trigger label is the same as the sampling time of each sample.

[0157] The training unit is specifically used to determine the first loss based on the block reclamation priority label, block reclamation prediction priority, and total number of samples for each sample. Specifically, it includes:

[0158] in, For the first Block recycling priority label for each sample For the first Block recycling prediction priority for each sample The total number of samples; Based on the page hot / cold classification label of each sample, the page hot / cold prediction classification result of each data page in the sample, and the total number of data pages in each sample, the second loss is determined. Specifically, this includes:

[0159] in, For any sample, the first The page hot / cold category tags for each data page. For any sample, the first The page hot / cold prediction classification results for each data page. This represents the total number of data pages in each sample. The third loss is determined based on the garbage collection trigger label and garbage collection prediction probability for each sample. Specifically, this includes:

[0160] in, The garbage collection trigger tag for the sample, Predict the probability of waste recycling; The training loss of the temporal causal model is determined based on the first loss, the second loss, and the third loss, and the parameters of the temporal causal model are adjusted based on the training loss.

[0161] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.

[0162] Figure 6A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0163] like Figure 6 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0164] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0165] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a garbage collection method based on a time-series causal model. For example, in some embodiments, the garbage collection method based on a time-series causal model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the garbage collection method based on a time-series causal model described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform a garbage collection method based on a time-series causal model.

[0166] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0167] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0168] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0170] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0171] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0172] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0173] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.

[0174] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A garbage collection method based on a time-series causal model, characterized in that, The temporal causal model includes a temporal feature enhancement layer, a temporal feature extraction layer, and a multi-task output layer; the multi-task output layer includes a block reclamation priority prediction unit, a page hot / cold classification unit, and a garbage reclamation trigger timing prediction unit; the method includes: Based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted, the first feature matrix is ​​determined; Determine the computing power of the solid-state drive controller; In response to the computing power of the solid-state drive controller being less than or equal to a computing power threshold, the first feature matrix is ​​pruned; wherein the temporal dimension of the pruned first feature matrix remains unchanged, but the feature dimension is smaller than the feature dimension of the first feature matrix before pruning; or, in response to the computing power of the solid-state drive controller being greater than the computing power threshold, the first feature matrix is ​​determined based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted at at least one sampling time, collected at a preset sampling frequency within the first and second time windows; wherein the second time window is before the first time window, and the duration of the second time window is greater than the duration of the first time window; The first feature matrix is ​​input into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted. The first feature map is input into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted. The second feature matrix is ​​input to the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted. Specifically, this includes: inputting the second feature matrix to the block collection priority prediction unit to obtain the block collection priority; inputting the second feature matrix to the page hot and cold classification unit to obtain the page hot and cold classification result of each data page in the storage block to be predicted; and inputting the second feature matrix to the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted. Based on the garbage collection decision results, determine whether to perform garbage collection operations on the storage block to be predicted; Wherein, the first feature matrix has a dimension of T×D×C, where C represents the number of time windows or the dimension of input channels; D represents the total dimension of the storage block status data, data page status data, input / output load data and garbage collection history data of the storage block to be predicted; and T represents the total sampling time of any time window. The corresponding garbage collection decision result of the storage block to be predicted includes the block collection priority, the page hot / cold classification result of each data page in the storage block to be predicted, and the garbage collection probability of the storage block to be predicted.

2. The method according to claim 1, characterized in that, in, The storage block status data includes at least one of the following: the number of valid data pages in the storage block, the number of invalid data pages in the storage block, the cumulative number of erases in the storage block, the storage time of data in the storage block, and bad block risk markers; the data page status data includes at least one of the following: the storage time of data in the data page, the number of historical writes to the data page, the number of times the data page was accessed within a period, historical hot / cold markers, and data page valid information markers; the input / output load data includes at least one of the following: read input / output IOPS, write IOPS, and average interval between input / output requests; the garbage collection history data includes at least one of the following: historical average garbage collection time, historical average valid page migration, historical average write amplification factor, average garbage collection trigger interval, and storage block wear leveling.

3. The method according to claim 1, characterized in that, The temporal feature enhancement layer includes a causal channel attention unit and a causal temporal attention unit. The step of inputting the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted includes: The first feature matrix is ​​input into the causal channel attention unit to obtain the first weight coefficient; the first weight coefficient represents the feature channel weight of the temporal causal model. The first weight coefficient and the first feature matrix are input into the causal temporal attention unit to obtain the first feature map.

4. The method according to claim 3, characterized in that, The causal channel attention unit includes a first global average pooling layer, a first fully connected layer, a second fully connected layer, and a first corrected linear unit; the step of inputting the first feature matrix into the causal channel attention unit to obtain the first weight coefficients includes: The first feature matrix is ​​sequentially input into the first global average pooling layer, the first fully connected layer, the first modified linear unit, and the second fully connected layer to obtain the first weight coefficients, specifically including: in, This is the first characteristic matrix; This is the first global average pooling layer; This is the first fully connected layer; This is the second fully connected layer; For the Sigmoid function; This is the first corrected linear unit; This is the first weighting coefficient, and its value range is... .

5. The method according to claim 3, characterized in that, The causal temporal attention unit includes a second global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second modified linear unit; the step of inputting the first weight coefficients and the first feature matrix into the causal temporal attention unit to obtain the first feature map includes: The first feature matrix is ​​weighted based on the first weight coefficient to obtain the channel-weighted features; The transpose of the channel-weighted features is input into the second global average pooling layer, the third fully connected layer, the fourth fully connected layer, and the second modified linear unit to obtain the second weight coefficients; the second weight coefficients characterize the temporal attention weights of the temporal causal model; specifically including: The second weighting coefficients after masking are obtained based on the second weighting coefficients and the mask matrix; The channel weighted features are weighted based on the masked second weighting coefficients to obtain the first feature map; in, Transpose of the channel-weighted features; This is the second global average pooling layer; It is the third fully connected layer; It is the fourth fully connected layer; For the Sigmoid function; This is the second corrected linear unit; This is the second weighting coefficient, and its value range is... .

6. The method according to claim 1, characterized in that, The temporal feature extraction layer includes a shallow network and a deep network. Both the shallow and deep networks include multiple bottleneck blocks. Each bottleneck block includes a first pointwise convolutional kernel, a temporal causal convolutional kernel, and a second pointwise convolutional kernel. The step of inputting the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted includes: The first feature map is input into a shallow network to obtain the first temporal feature; The first temporal feature is input into a deep network to obtain the second feature matrix; The number of channels output by the shallow network is less than the number of channels output by the deep network.

7. The method according to claim 1, characterized in that, The step of inputting the second feature matrix into the multi-task output layer to obtain the garbage collection decision result corresponding to the storage block to be predicted includes at least one of the following: The position of the storage block to be predicted in the garbage collection queue is determined based on the block reclamation priority corresponding to the storage block to be predicted. If the position of the storage block to be predicted in the garbage collection queue is greater than the queuing threshold, then the storage block to be predicted is garbage collected. The garbage collection trigger threshold is determined based on the proportion of free storage blocks to all storage blocks; if the garbage collection probability of the storage block to be predicted is greater than or equal to the garbage collection trigger threshold, then garbage collection is performed on the storage block to be predicted. Based on the page hot / cold classification results of each data page in the storage block to be predicted, data pages whose page hot / cold classification results are hot pages are migrated to free storage blocks.

8. The method according to claim 1, characterized in that, The method also includes training a time-series causal model, specifically including: A dataset is determined based on storage block status data, data page status data, input / output load data, and garbage collection history data. Each sample in the dataset includes a storage block with D-dimensional features of a storage block in C time windows, and T consecutive sampling times within each time window. Label each sample in the dataset with block recycling priority label, page hot / cold classification label, and garbage collection trigger label; The training set is determined based on the dataset. The training set is then input into the time-series causal model to obtain the block recycling prediction priority, the page hot / cold prediction classification result, and the garbage recycling prediction probability for each sample. The weights of the time-series causal model are adjusted based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label for each sample, as well as the block recycling prediction priority, page hot / cold prediction classification result, and garbage recycling prediction probability for each data page corresponding to each sample.

9. The method according to claim 8, characterized in that, The block recycling priority label, page hot / cold classification label, and garbage collection trigger label for each sample in the labeled dataset include: The block reclamation priority label for each sample is determined based on the number of invalid data pages, the total number of data pages, the proportion of hot pages, and the average erase error rate. Specifically, it includes: in, The number of invalid data pages for sample b. This represents the total number of data pages for sample b. This represents the percentage of hot pages within sample b. The average erase error rate for sample b; , and is the weighting coefficient; where sample b is any sample in the dataset; the block reclamation priority label corresponds one-to-one with the storage block in the dataset; Based on the update information of each sample in the dataset at each time step, determine the page hot / cold classification label of the data page in each sample; the dimension of the page hot / cold classification label is the same as the total number of data pages in each sample. Based on the garbage collection trigger information of each sample in the dataset at each time point, the garbage collection trigger label of each sample is determined; the dimension of the garbage collection trigger label is the same as the sampling time of each sample.

10. The method according to claim 8, characterized in that, The weights of the time-series causal model are adjusted based on the block recycling priority label, page hot / cold classification label, and garbage recycling trigger label for each sample, as well as the block recycling prediction priority, page hot / cold prediction classification result, and garbage recycling prediction probability for each data page. This includes: The first loss is determined based on the block reclamation priority label, block reclamation prediction priority, and total number of samples for each sample. Specifically, it includes: in, For the first Block recycling priority label for each sample For the first Block recycling prediction priority for each sample The total number of samples; Based on the page hot / cold classification label of each sample, the page hot / cold prediction classification result of each data page in the sample, and the total number of data pages in each sample, the second loss is determined. Specifically, it includes: in, For any sample, the first The page hot / cold category tags for each data page. For any sample, the first The page hot / cold prediction classification results for each data page. This represents the total number of data pages in each sample. The third loss is determined based on the garbage collection trigger label and garbage collection prediction probability for each sample. Specifically, it includes: in, The garbage collection trigger tag for the sample, Predict the probability of waste recycling; The training loss of the temporal causal model is determined based on the first loss, the second loss, and the third loss, and the parameters of the temporal causal model are adjusted based on the training loss.

11. A waste recycling device based on a time-series causal model, characterized in that, The temporal causal model includes a temporal feature enhancement layer, a temporal feature extraction layer, and a multi-task output layer; the multi-task output layer includes a block recycling priority prediction unit, a page hot / cold classification unit, and a garbage recycling trigger timing prediction unit; the device includes: A data matrix determination unit is used to determine a first feature matrix based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted; it is also used to determine the computing power of the solid-state drive controller; in response to the computing power of the solid-state drive controller being less than or equal to a computing power threshold, the first feature matrix is ​​truncated; wherein the temporal dimension of the truncated first feature matrix remains unchanged, but the feature dimension is smaller than the feature dimension of the first feature matrix before truncation; or, in response to the computing power of the solid-state drive controller being greater than the computing power threshold, the first feature matrix is ​​determined based on the storage block status data, data page status data, input / output load data, and garbage collection history data of the storage block to be predicted collected at a preset sampling frequency within a first time window and a second time window at at least one sampling time; wherein the second time window is before the first time window, and the duration of the second time window is greater than the duration of the first time window; The feature enhancement unit is used to input the first feature matrix into the temporal feature enhancement layer to obtain the first feature map corresponding to the storage block to be predicted. The feature extraction unit is used to input the first feature map into the temporal feature extraction layer to obtain the second feature matrix corresponding to the storage block to be predicted. The inference unit is used to input the second feature matrix into the multi-task output layer to obtain the corresponding garbage collection decision result of the storage block to be predicted. The inference unit is specifically used to input the second feature matrix into the block reclamation priority prediction unit to obtain the block reclamation priority; input the second feature matrix into the page hot and cold classification unit to obtain the page hot and cold classification result of each data page in the storage block to be predicted; and input the second feature matrix into the garbage collection trigger timing prediction unit to obtain the garbage collection probability of the storage block to be predicted. The processing unit is used to determine, based on the garbage collection decision result, whether to perform garbage collection operation on the storage block to be predicted; Wherein, the first feature matrix has a dimension of T×D×C, where C represents the number of time windows or the dimension of input channels; D represents the total dimension of the storage block status data, data page status data, input / output load data and garbage collection history data of the storage block to be predicted; and T represents the total sampling time of any time window. The corresponding garbage collection decision result of the storage block to be predicted includes the block collection priority, the page hot / cold classification result of each data page in the storage block to be predicted, and the garbage collection probability of the storage block to be predicted.

12. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Data garbage collection method and device, solid state disk, medium and program product

    CN121455840A

  • Shield tunneling parameter prediction method based on multi-source domain transfer learning

    CN121659285A