Cache optimization method and system based on high-bandwidth memory, electronic equipment and medium
By calculating the access frequency of NAND flash memory data blocks and dynamically adjusting the cache capacity allocation, combined with TSV via arrays and crossbar switch matrices, the problem of cache area imbalance in traditional storage systems is solved, thus improving cache efficiency.
Patent Information
- Application Number
- CN202511915426.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2045-12-18
AI Technical Summary
In traditional storage systems, due to the dynamic changes in data access patterns and load characteristics, the fixed cache space allocation leads to an imbalance in resource allocation, causing some cache areas to be overused while other areas are idle, thus reducing the caching efficiency of the storage system.
By acquiring the access frequency, data type, and access time interval of data blocks in NAND flash memory, the access heat value is calculated and priority classification is performed. The capacity allocation ratio between the extended cache layer and the cache area in NAND flash memory is dynamically adjusted. A cache channel is established using the TSV via array and cross switch matrix in the silicon interposer layer, and the transmission bandwidth utilization is monitored to achieve load balancing.
It enables on-demand allocation and load balancing of cache resources, improves the caching efficiency of the storage system, and avoids resource waste under the traditional static cache allocation strategy.
Smart Images

Figure CN121349907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a cache optimization method, system, electronic device, and medium based on high-bandwidth memory. Background Technology
[0002] With the rapid development of big data and artificial intelligence applications, the amount of data that storage systems need to process is growing exponentially, and the requirements for storage performance are also increasing. To achieve a balance between performance and cost, modern storage architectures generally adopt a multi-tiered storage system. Among them, NAND flash memory is widely used in the data storage layer due to its high cost-effectiveness, and data access efficiency is improved by configuring a high-speed cache layer.
[0003] Currently, traditional storage systems employ a static cache allocation strategy, which manages data by pre-allocating fixed cache space and defining preset data block migration rules during system design.
[0004] However, in practical applications, since data access patterns and load characteristics change dynamically over time, a fixed cache space allocation can easily lead to an imbalance in resource allocation, resulting in some cache areas being overused while other areas are idle, thereby reducing the caching efficiency of the storage system. Summary of the Invention
[0005] This application provides a cache optimization method, system, electronic device, and medium based on high-bandwidth memory, which can improve the cache efficiency of storage systems.
[0006] In a first aspect, this application provides a cache optimization method based on high-bandwidth memory, including: Obtain access data for each data block in the NAND flash memory, wherein the access data includes access frequency, data type, and access time interval; Based on the access data, the access popularity value of each data block is calculated, and the data blocks are classified according to the access popularity value to obtain the cache priority corresponding to each data block. Based on the number of data blocks with each cache priority, adjust the capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory, and determine multiple data blocks to be migrated based on the adjusted capacity allocation ratio and the cache priority of each data block. A cache channel is established for each of the data blocks to be migrated by means of a TSV via array and a cross switch matrix integrated within the silicon interposer, and a cache optimization path is determined according to each cache channel. The cache channel is used for data transmission between the preset extended cache layer and the NAND flash memory. When the transmission bandwidth utilization of any cache channel exceeds the bandwidth utilization threshold, the cache optimization path is adjusted according to the deviation rate between the transmission bandwidth utilization and the bandwidth utilization threshold. Based on the cache optimization path, perform cache migration operations on each of the data blocks to be migrated.
[0007] By adopting the above technical solution, based on access data such as access frequency, data type, and access time interval of data blocks in NAND flash memory, the access heat value of data blocks is calculated and priority classification is performed. This allows for dynamic adjustment of the capacity allocation ratio between the extended cache layer and the cache area in NAND flash memory according to the number of data blocks with different priorities. Furthermore, a cache channel is established through a TSV via array and cross-switch matrix integrated within the silicon interposer, achieving on-demand allocation of cache resources. Simultaneously, by monitoring the transmission bandwidth utilization of the cache channel and dynamically adjusting the cache optimization path based on its deviation from the bandwidth utilization threshold, load balancing of the cache channel is achieved. This effectively avoids the problem of some cache areas being overused while others are idle under traditional static cache allocation strategies, thus improving the cache efficiency of the storage system.
[0008] Optionally, for each data block, the fluctuation range of the access frequency of the data block within a preset transmission period is weighted by a time window to obtain a standardized access frequency; the data block type is determined according to the data type, and the type weight factor corresponding to the data block type is determined in a preset data block type mapping table, wherein the data block type includes sequential access data blocks, random access data blocks, and mixed access data blocks; the access interval variance value of the data block is calculated based on each access time interval; the standardized access frequency, access interval variance value, and type weight factor of the data block are weighted and fused to obtain the access popularity value of the data block.
[0009] Optionally, the write rate of the NAND flash memory is collected, and the deviation ratio between the write rate and the baseline write rate is calculated; based on the deviation ratio, the number of data blocks for each cache priority is weighted and corrected to obtain the corrected data block number distribution; the write overhead coefficient corresponding to each cache priority is calculated by combining the average write time and write power consumption of the data blocks corresponding to each cache priority; the dynamic capacity distribution coefficient is calculated according to the data block number distribution and the write overhead coefficient corresponding to each cache priority; the current initial capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory is obtained; the dynamic capacity distribution coefficient is calculated with the initial capacity allocation ratio to obtain the capacity allocation ratio between the extended cache area and the cache area.
[0010] Optionally, the target available capacity of the extended cache is determined according to the capacity allocation ratio; data blocks corresponding to each cache priority are extracted sequentially in descending order of their respective cache priorities, and the data capacity of each data block is accumulated; when the accumulated data capacity reaches the target available capacity of the extended cache, the extracted data blocks are marked as data blocks to be migrated.
[0011] Optionally, the current load status of each TSV via in the TSV via array and the connection status of each switch node in the cross switch matrix are obtained; the channel bandwidth requirement corresponding to each data block to be migrated is determined according to the data capacity and cache priority of each data block to be migrated; based on the channel bandwidth requirement and the current load status, a corresponding TSV via group is allocated to each data block to be migrated in the TSV via array; for each data block, according to the connection status of the TSV via group corresponding to the data block and each switch node, the switch path from the NAND flash memory to the preset extended cache layer is configured through the cross switch matrix to obtain the cache channel corresponding to the data block to be migrated.
[0012] Optionally, based on the cache optimization path, the migration order and corresponding cache channel of each data block to be migrated are determined; according to the migration order, the data blocks to be migrated are read sequentially from the NAND flash memory and transmitted to the preset extended cache layer through the corresponding cache channel; during data transmission, the transmission error rate of each cache channel is monitored; when the transmission error rate exceeds a preset error rate threshold, the transmission of the current data block to be migrated is paused, the switching path of the cache channel is reconfigured through the cross-switch matrix, and the transmission operation of the current data block to be migrated is performed according to the reconfigured cache channel; after each data block to be migrated is migrated, the cache index table of the preset extended cache layer is updated, and the cache status of the migrated data blocks is marked in the NAND flash memory.
[0013] Optionally, the current connection status and load status of each switch node in the cross-connection switch matrix are obtained; based on the transmission error rate, the faulty switch node corresponding to the target cache channel where the transmission error occurred is located; according to the current connection status and load status of each switch node, backup switch nodes with load status below a preset load threshold and no transmission error have occurred are selected from the cross-connection switch matrix; according to the available bandwidth of the backup switch node, a target backup switch node that meets the channel bandwidth requirements of the current data block to be migrated is selected, and a new switch path from the NAND flash memory to the preset extended cache layer is established according to the target backup switch node, thereby completing the reconfiguration of the switch path of the cache channel.
[0014] A second aspect of this application provides a cache optimization system based on high-bandwidth memory, the system comprising: The data acquisition module is used to acquire access data of each data block in the NAND flash memory. The access data includes access frequency, data type and access time interval. The module for determining data blocks to be migrated is used to calculate the access heat value of each data block based on the access data, and classify each data block according to the access heat value to obtain the cache priority corresponding to each data block; adjust the capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory according to the number of data blocks with each cache priority, and determine multiple data blocks to be migrated according to the adjusted capacity allocation ratio and the cache priority of each data block; The path determination module is used to establish a cache channel corresponding to each of the data blocks to be migrated through the TSV via array and cross switch matrix integrated in the silicon interposer, and to determine the cache optimization path according to each cache channel. The cache channel is used for data transmission between the preset extended cache layer and the NAND flash memory. The cache optimization module is used to adjust the cache optimization path according to the deviation rate between the transmission bandwidth utilization rate and the bandwidth utilization threshold when the transmission bandwidth utilization rate of any cache channel exceeds the bandwidth utilization rate threshold; and to perform cache migration operations for each of the data blocks to be migrated based on the cache optimization path.
[0015] A third aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, the program being able to implement a cache optimization method based on high-bandwidth memory when loaded and executed by the processor.
[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a cache optimization method based on high-bandwidth memory.
[0017] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting the above technical solution, based on access data such as access frequency, data type, and access time interval of data blocks in NAND flash memory, the access heat value of data blocks is calculated and priority classification is performed. This allows for dynamic adjustment of the capacity allocation ratio between the extended cache layer and the cache area in NAND flash memory according to the number of data blocks with different priorities. Furthermore, a cache channel is established through a TSV via array and cross-switch matrix integrated within the silicon interposer, achieving on-demand allocation of cache resources. Simultaneously, by monitoring the transmission bandwidth utilization of the cache channel and dynamically adjusting the cache optimization path based on its deviation from the bandwidth utilization threshold, load balancing of the cache channel is achieved. This effectively avoids the problem of some cache areas being overused while others are idle under traditional static cache allocation strategies, thus improving the cache efficiency of the storage system. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a cache optimization method based on high-bandwidth memory provided in an embodiment of this application; Figure 2 This is another flowchart illustrating a cache optimization method based on high-bandwidth memory provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a cache optimization system based on high-bandwidth memory provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0020] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0021] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0022] This application provides a cache optimization method based on high-bandwidth memory. In one embodiment, please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a cache optimization method based on high-bandwidth memory provided in an embodiment of this application. This method can be implemented using a computer program, which can be integrated into an application or run as a standalone utility application. The method can also be implemented using a microcontroller or run on a cache optimization system based on the von Neumann architecture and high-bandwidth memory. Specifically, the method may include the following steps: Step 101: Obtain access data for each data block in the NAND flash memory. The access data includes access frequency, data type, and access time interval.
[0023] NAND flash memory is a non-volatile memory that uses a floating-gate transistor structure to store data, characterized by high density and low power consumption. A data block is the basic unit of data organization in NAND flash memory, typically containing 512 bytes to 8KB of data content, and each data block has a unique logical address identifier. Access data is a collection of statistical information describing the usage status of a data block. Access frequency refers to the number of read and write operations performed on a specific data block per unit of time, usually measured in times per second. Data type is a classification identifier based on the characteristics of the data block content and access patterns, including categories such as system files, user files, and temporary files. Access time interval refers to the time difference between two consecutive access operations on the same data block, reflecting the temporal distribution characteristics of data access. For example, a data block containing operating system kernel files might be accessed 50 times per second, belonging to a system-critical data type, with an average access time interval of 20 milliseconds.
[0024] Specifically, data collection is achieved by deploying an access monitoring module in the NAND flash memory controller. This module runs at the firmware layer of the flash memory controller and collects access data by intercepting and recording all I / O operations targeting the NAND flash memory. Access frequency is obtained by maintaining an access counter for each data block. The counter accumulates access counts within a preset time window (e.g., 1 second), and calculates the access frequency and resets the counter at the end of the time window. Data type identification is based on file system metadata and the logical address range of the data block. The file type of the data block is determined by querying the file allocation table, and the file type is converted into a standardized data type identifier according to predefined mapping rules. The access time interval is calculated by recording the most recent access timestamp of each data block. When a new access operation is detected, the difference between the current time and the last access time is calculated as the access time interval. All collected access data is stored in a structured format in a dedicated metadata buffer, containing fields such as data block ID, timestamp, access type (read / write), access frequency, data type identifier, and time interval, providing a data foundation for subsequent heat calculation and cache optimization.
[0025] Step 102: Based on each access data, calculate the access popularity value of each data block, and classify each data block into cache priorities according to each access popularity value to obtain the cache priority corresponding to each data block.
[0026] Access popularity is a numerical indicator that quantifies the activity level of a data block. It is calculated by considering multiple dimensions such as access frequency, data type, and access time interval. A higher value indicates a more active data block. Cache priority classification is the process of dividing data blocks into different levels based on access popularity, used to guide cache resource allocation strategies. Cache priority indicates the importance level of a data block within the caching system.
[0027] Specifically, the access frequency of each data block is first calculated using a multi-dimensional weighted fusion algorithm. The calculation process consists of four sub-steps: First, the access frequency is standardized using a time window weighting method, with the formula F_norm=Σ(w_i×f_i) / Σ(w_i), where w_i is the weight coefficient of the i-th time window, f_i is the access frequency of the corresponding window, and recent windows have higher weights to reflect timeliness. Second, the type weight factor is determined based on the data type. The corresponding weight is obtained by querying a pre-defined data block type mapping table: sequential access data blocks have a weight of 1.2, random access data blocks have a weight of 0.8, and mixed access data blocks have a weight of 1.0. Third, the access interval variance is calculated to assess access regularity, with the formula V=Σ(t_i-t_avg)² / n, where t_i is the access time interval, t_avg is the average access interval, and n is the number of interval samples. Fourth, the three factors mentioned above are weighted and fused to calculate the popularity value, using the formula H = α × F_norm + β × W_type + γ × (1 / V), where α, β, and γ are preset weight coefficients of 0.5, 0.3, and 0.2, respectively. After the popularity value is calculated, a threshold segmentation method is used to classify cache priorities. The high-priority threshold is set to 0.8, and the medium-priority threshold is set to 0.4. Data blocks with popularity values greater than 0.8 are classified as high priority, those between 0.4 and 0.8 as medium priority, and those less than 0.4 as low priority. Finally, a corresponding cache priority identifier is assigned to each data block and updated in the metadata table.
[0028] In one possible implementation, the access popularity value of each data block is calculated based on each access data, specifically including steps 1021-1023, as follows: Step 1021: For each data block, the fluctuation range of the access frequency of the data block within the preset transmission period is weighted by time window to obtain the standardized access frequency.
[0029] The preset transmission period is a predefined data transmission time unit, typically set to a fixed duration of 1 to 60 seconds. Fluctuation amplitude refers to the degree of change in access frequency over different time periods, reflecting the stability characteristics of data access. Time window weighting is a time-series data analysis method that assigns different weight values to data from different time periods, with more recent data receiving higher weight. Standardized access frequency is the access frequency value after normalization and weight adjustment, eliminating the impact of time fluctuations on frequency calculation.
[0030] Specifically, the access frequency is standardized using a sliding window algorithm. First, the preset transmission period is divided into N equal-length time windows, each with a length of 1 / N of the transmission period; typically, N is between 10 and 20. A frequency record array F[1...N] is created for each data block to record the access frequency within each time window. The weight coefficient for each window is calculated using the exponential decay function w_i=e^(-λ(Ni)), where λ is the decay factor (usually 0.1), and i is the window index; more recent windows have higher weights. Next, the fluctuation amplitude is calculated using the coefficient of variation CV=σ / μ, where σ is the frequency standard deviation and μ is the frequency mean. Then, the weights are normalized to ensure Σw_i=1. Finally, the standardized access frequency is calculated using the formula F_norm=Σ(w_i×F[i])×(1+α×CV), where α is the fluctuation adjustment factor (value 0.2); data blocks with greater fluctuations receive higher standardized frequency values. The algorithm is executed once per transmission cycle, updating the normalized access frequency of all active data blocks and storing the results in an access statistics table.
[0031] Step 1022: Determine the data block type based on the data type, and determine the type weight factor corresponding to the data block type in the preset data block type mapping table. The data block types include sequential access data blocks, random access data blocks, and mixed access data blocks.
[0032] Data block types are functional classifications of data blocks based on their data access patterns. Sequential access data blocks are those read and written sequentially from one address, typically found in streaming media files and large file transfers. Random access data blocks are those accessed at random addresses, typically used in database records and index files. Hybrid access data blocks exhibit both sequential and random access characteristics. A data block type mapping table is a pre-established lookup table structure that stores the correspondence between data types and data block types. Type weighting factors are numerical coefficients reflecting the importance of different data block types in cache optimization.
[0033] Specifically, the data block type is determined by the access pattern analysis algorithm and the weight factor is obtained. First, the access address sequence of the data block is analyzed to calculate the consecutive access ratio and the jump access ratio. The condition for determining consecutive access is that the address difference between two adjacent accesses is less than a preset threshold (usually twice the data block size). The number of consecutive accesses C_seq and the total number of accesses C_total are counted, and the consecutive access rate R_seq = C_seq / C_total is calculated. The type is determined according to the consecutive access rate: when R_seq ≥ 0.8, it is classified as a sequential access data block; when R_seq ≤ 0.3, it is classified as a random access data block; when 0.3 < R_seq < 0.8, it is classified as a mixed access data block. After determining the data block type, the corresponding weight factor is obtained by querying the preset data block type mapping table. The mapping table is stored in a hash table structure and contains three entries: the weight factor of the sequential access data block is 0.8 (relatively low priority because sequential access has less dependence on the cache), the weight factor of the random access data block is 1.5 (the highest priority because random access benefits most from the cache), and the weight factor of the mixed access data block is 1.0 (medium priority). The system records the data block type and the corresponding weight factor in the metadata management table to provide the basic data for subsequent heat calculation.
[0034] Step 1023: Calculate the access interval variance value of the data block based on each access time interval; perform weighted fusion calculation on the normalized access frequency, access interval variance value, and type weight factor of the data block to obtain the access heat value of the data block.
[0035] The access interval variance value is a statistic that measures the degree of dispersion of the access time interval distribution of the data block and reflects the strength of the access regularity. The weighted fusion calculation is a mathematical operation method that linearly combines multiple eigenvalue features in different dimensions according to the preset weights. The access heat value is a comprehensive index for comprehensively evaluating the access activity degree of the data block, and the numerical range is usually between 0 and 1.
[0036] Specifically, the variance of the access interval is calculated as follows: All access interval samples of the data block within the statistical period are collected, forming an interval sequence T = [t1, t2, ..., tn]. The average access interval t_avg = Σti / n is calculated. The interval variance V = Σ(ti-t_avg)² / (n-1) is calculated, using the sample variance formula to ensure statistical accuracy. To avoid division by zero errors, a minimum value of 0.001 is set when the variance is 0. Weighted fusion calculation is performed: A three-dimensional feature vector [F_norm, W_type, 1 / (1+V)] is established, where F_norm is the standardized access frequency, W_type is the type weight factor, and 1 / (1+V) is the interval regularity factor (the smaller the variance, the stronger the regularity, and the greater the contribution of popularity). The fusion weight vector [α, β, γ] = [0.5, 0.3, 0.2] is set, where α corresponds to the access frequency weight, β corresponds to the type weight, and γ corresponds to the regularity weight. Perform a weighted fusion calculation: H = α × F_norm + β × W_type + γ × 1 / (1 + V). Normalize the calculation result using the Sigmoid function H_final = 1 / (1 + e^(-k × (H - θ))), where k = 5 is the kurtosis parameter and θ = 0.5 is the center point parameter, ensuring the popularity value is distributed within the 0-1 range. Update the final access popularity value in the data block's metadata record and trigger a cache priority re-evaluation process.
[0037] Step 103: Based on the number of data blocks for each cache priority, adjust the capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory, and determine multiple data blocks to be migrated based on the adjusted capacity allocation ratio and the cache priority of each data block.
[0038] The default extended cache layer is a high-speed storage tier located above NAND flash memory, implemented using high-bandwidth memory technology. It is used to cache frequently accessed data to improve access performance. The extended cache area is a dedicated storage region within the default extended cache layer for storing data migrated from NAND flash memory. The cache area within NAND flash memory is storage space reserved within the NAND flash memory for temporary data caching. The capacity allocation ratio refers to the storage capacity distribution between the extended cache area and the NAND flash memory cache area, usually expressed as a percentage. The data blocks to be migrated are the set of data blocks selected according to the cache optimization strategy that need to be migrated from NAND flash memory to the extended cache layer.
[0039] Specifically, the optimized configuration of cache resources is achieved through dynamic capacity adjustment algorithms and data block selection algorithms. First, capacity allocation ratio adjustment is performed: the number of data blocks corresponding to each cache priority is counted, and the distribution vector of high, medium, and low priority data blocks [N_high, N_mid, N_low] is calculated. The current write rate W_current of the NAND flash memory is collected and compared with the baseline write rate W_baseline, and the deviation ratio Δ=(W_current-W_baseline) / W_baseline is calculated. Based on the deviation ratio, the number of data blocks is weighted and corrected using the formula N'_i=N_i×(1+α×Δ×P_i), where α is an adjustment coefficient with a value of 0.3, and P_i is the sensitivity coefficient corresponding to each priority [1.2, 1.0, 0.8]. The write overhead coefficient for each priority is calculated, combined with the average write time T_i and write power consumption E_i, using the formula C_i=β×T_i+γ×E_i, where β=0.6 and γ=0.4. The dynamic capacity distribution coefficient S = Σ(N'_i × C_i × P_i) / Σ(N'_i) is calculated based on the corrected data block distribution and write overhead coefficient. The current initial capacity allocation ratio R_initial is obtained, and the adjusted ratio R_new = R_initial × (1 + δ × S), where δ is the adjustment strength coefficient with a value of 0.2. Then, the determination of data blocks to be migrated is performed: the target available capacity V_target of the extended cache is calculated based on the adjusted capacity allocation ratio. All data blocks are sorted in descending order of cache priority, and high-priority data blocks are extracted sequentially, with their data capacity accumulated. When the accumulated capacity is close to but does not exceed V_target, the extracted data blocks are marked as data blocks to be migrated, and their logical address, data capacity, and priority information are recorded in the migration task queue.
[0040] In one possible implementation, the capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory is adjusted according to the number of data blocks for each cache priority. Specifically, this includes steps 1031-1033, as follows: Step 1031: Collect the write rate of NAND flash memory and calculate the deviation ratio between the write rate and the baseline write rate; based on the deviation ratio, perform weighted correction on the number of data blocks for each cache priority to obtain the corrected distribution of the number of data blocks.
[0041] Write rate refers to the amount of data written by NAND flash memory per unit of time, usually measured in MB / s or GB / s. The baseline write rate is a preset write performance reference value under standard load conditions, used to assess the degree of deviation from current performance. The deviation ratio is the relative difference between the current write rate and the baseline write rate, reflecting the trend of system load changes. Weighted correction is a mathematical processing method that adjusts the original data according to the deviation ratio to eliminate the impact of load fluctuations on statistical results. The corrected data block distribution is the statistical result of the number of data blocks corresponding to each cache priority after weighted adjustment. For example, if the current write rate is 150MB / s, the baseline write rate is 100MB / s, and the deviation ratio is 50%, the number of high-priority data blocks is corrected from the original 500 to 580.
[0042] Specifically, the number of data blocks is dynamically adjusted through a performance monitoring module and a statistical correction algorithm. First, write rate data is collected: a performance counter is deployed in the NAND flash controller to collect the total number of bytes and the number of write operations per second, calculating the instantaneous write rate W_instant = number of bytes written / time interval. A moving average algorithm is used to calculate the stable write rate, with the formula W_current = (W_instant + W_prev × (n-1)) / n, where n is the smoothing window size (10), and W_prev is the result of the previous calculation. The deviation ratio Δ = (W_current - W_baseline) / W_baseline is calculated, where W_baseline is the baseline write rate configured by the system. Then, weighted correction processing is performed: a sensitivity weight vector [S_high, S_mid, S_low] = [1.3, 1.0, 0.7] is established for each cache priority, with higher priority data being more sensitive to performance changes. The original number of data blocks N_original for each priority is corrected using the formula N_corrected = N_original × (1 + α × Δ × S_i), where α is the correction strength coefficient (0.25) and S_i is the sensitivity weight for the corresponding priority. When the deviation ratio is positive, the weight of the data block corresponding to that priority is increased; when it is negative, the weight is decreased. The corrected distribution of the number of data blocks [N_high_corrected, N_mid_corrected, N_low_corrected] is stored in the system status table to provide basic data for subsequent capacity allocation calculations.
[0043] Step 1032: Calculate the write overhead coefficient corresponding to each cache priority by combining the average write time and write power consumption of the data blocks corresponding to each cache priority; calculate the dynamic capacity distribution coefficient based on the data block quantity distribution and the write overhead coefficient corresponding to each cache priority.
[0044] Average write time refers to the average time required to complete a single data block write operation, including data transfer time and write acknowledgment time. Write power consumption is the electrical energy consumed during a data block write operation, typically measured in milliwatt-hours (mWh). Write overhead coefficient is a cost assessment metric calculated by considering both write time and power consumption, used to quantify the write cost of data blocks of different priorities. Dynamic capacity distribution coefficient is a capacity allocation adjustment parameter calculated based on data block distribution and write overhead, guiding the dynamic reallocation of cache resources.
[0045] Specifically, various coefficient indicators are calculated through performance analysis and cost modeling algorithms. First, the write overhead coefficient is calculated: write performance data for each cache priority data block is statistically analyzed from historical operation records, and the average write time T_avg_i = Σ(write completion time - write start time) / number of operations is calculated. Write power consumption data is collected through a power monitoring module, and the average write power consumption P_avg_i = Σ(power consumption integral value during write) / number of operations is calculated. A write overhead coefficient calculation model is established, with the formula C_i = w_time × T_avg_i + w_power × P_avg_i, where w_time is the time weight coefficient with a value of 0.6, and w_power is the power consumption weight coefficient with a value of 0.4. The calculation results are normalized to ensure the comparability of overhead coefficients between different priorities. Then, the dynamic capacity distribution coefficient is calculated: a comprehensive evaluation model is established, combining the corrected data block quantity distribution N_corrected_i and the write overhead coefficient C_i. The weighted influence factor I_i = N_corrected_i × C_i for each priority is calculated, reflecting the degree of influence of that priority on the overall capacity requirement. Calculate the overall impact factor I_total = ΣI_i. Finally, calculate the dynamic capacity distribution coefficient D = Σ(I_i × P_i) / I_total, where P_i is the baseline weighting factor for each priority level [0.5, 0.3, 0.2]. This coefficient ranges from 0 to 1; a larger value indicates a need to increase the capacity allocation ratio of the extended cache.
[0046] Step 1033: Obtain the current initial capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory; calculate the dynamic capacity distribution coefficient and the initial capacity allocation ratio to obtain the capacity allocation ratio between the extended cache area and the cache area.
[0047] The initial capacity allocation ratio is the pre-defined capacity allocation relationship between the extended cache and the NAND flash cache during system startup, serving as a baseline reference value for dynamic adjustment. The capacity allocation ratio is the final capacity allocation relationship determined after dynamic adjustment, used to guide the actual cache space partitioning.
[0048] Specifically, dynamic reallocation of cache resources is achieved through a capacity allocation optimization algorithm. First, the initial capacity allocation ratio is obtained: preset capacity allocation parameters are read from the system configuration table, including the initial ratio of the extended cache (R_extend_initial) and the initial ratio of the NAND flash cache (R_nand_initial), whose sum equals 1. The validity of the current allocation state is verified to ensure that the actual usage capacity of each cache does not exceed the allocated capacity. Then, dynamic capacity allocation calculation is performed: a capacity adjustment model is established, using a linear adjustment algorithm R_extend_new = R_extend_initial + β × (D - D_baseline), where β is the adjustment sensitivity coefficient (0.3), D is the currently calculated dynamic capacity distribution coefficient, and D_baseline is the baseline distribution coefficient (0.5). The adjustment range limit is calculated, setting the maximum adjustment range to ±15%, ensuring that R_extend_initial × 0.85 ≤ R_extend_new ≤ R_extend_initial × 1.15. Boundary constraints are applied to the adjustment results to ensure that 0.5 ≤ R_extend_new ≤ 0.9, preventing system instability caused by extreme allocation. The corresponding NAND flash memory cache ratio R_nand_new = 1 - R_extend_new is calculated. The feasibility of the new allocation ratio is verified, and it is checked whether the minimum capacity requirement and hardware constraints are met. The final determined capacity allocation ratio [R_extend_new, R_nand_new] is written to the system configuration table, and a cache reallocation operation is triggered. At the same time, a capacity adjustment log is generated to record the reason for the adjustment and the adjustment range.
[0049] In one possible implementation, multiple data blocks to be migrated are determined based on the adjusted capacity allocation ratio and the cache priority of each data block, specifically including steps 1034-1035, as follows: Step 1034: Determine the target available capacity of the extended cache area according to the capacity allocation ratio; extract the data blocks corresponding to each cache priority in descending order of their priorities, and accumulate the data capacity of each data block.
[0050] The target available capacity is the actual amount of data that the extended cache can store, calculated based on the adjusted capacity allocation ratio, minus the system reserved space and the space occupied by metadata. Descending order refers to arranging data blocks according to their cache priority from highest to lowest, processing data blocks of higher importance first. Data capacity refers to the actual storage space occupied by a single data block, including the data content and necessary metadata information. Accumulation operation refers to the mathematical process of continuously adding the capacity values of multiple data blocks.
[0051] Specifically, the pre-selection process for data to be migrated is implemented through capacity calculation and data block sorting algorithms. First, the target available capacity is determined: the total physical capacity C_total of the extended cache is obtained, typically by querying the hardware configuration table. The allocated capacity C_allocated = C_total × R_extend is calculated based on the capacity allocation ratio R_extend. The system reserved capacity, including metadata storage space (5% of total capacity), error correction space (3% of total capacity), and buffer space (2% of total capacity), is deducted. The calculation formula is C_target = C_allocated × (1 - 0.05 - 0.03 - 0.02) = C_allocated × 0.9. Then, the data block sorting and accumulation process is executed: information for all data blocks is extracted from the data block management table, including data block ID, cache priority, and data capacity. A quicksort algorithm is used to sort the data blocks in descending order of cache priority, with higher priority data blocks at the top. Within the same priority, a secondary sort is performed based on access frequency. An accumulated capacity variable C_accumulated is created and initialized to 0, and a candidate data block list candidate_list is created. The data blocks are traversed sequentially according to the sorting results. For each data block, a capacity accumulation operation C_accumulated += block_capacity is performed. During the accumulation process, the identification information of the current data block is added to the candidate list, recording its logical address, physical address, data capacity, priority, and other key attributes. The accumulation termination condition is set to C_accumulated ≤ C_target to ensure that the target available capacity limit is not exceeded. When approaching the capacity limit, it is checked whether the capacity of the next data block will cause an exceedance. If the limit is exceeded, the accumulation process is stopped, and the current candidate data block list is maintained.
[0052] Step 1035: When the accumulated data capacity reaches the target available capacity of the extended cache, mark each extracted data block as a data block to be migrated.
[0053] The data blocks to be migrated are the set of data blocks that have been filtered for capacity limitations and ultimately determined to need to be migrated from NAND flash memory to the extended cache layer. The marking operation refers to adding a specific identifier to the metadata of the data blocks to be migrated to distinguish the status information of the data blocks to be migrated from that of ordinary data blocks.
[0054] Specifically, the final determination of data blocks to be migrated is achieved through capacity checks and status marking algorithms. First, a capacity compliance check is performed: a capacity check loop is set up, checking whether the current accumulated value C_accumulated meets the termination condition after each accumulation of data block capacity. A predictive check mechanism is used: before accumulating the capacity of the current data block, C_predicted = C_accumulated + current_block_capacity is calculated to determine if it exceeds the target available capacity C_target. If C_predicted > C_target, the accumulation process stops, and the current data block is not included; if C_predicted ≤ C_target, the accumulation continues and the current data block is included. The final accumulated capacity value and the number of included data blocks are recorded, and a capacity usage report is generated. Then, a migration marking operation is performed: batch marking is performed on all data blocks in the candidate list, and the migration status field migration_status = 1 is set in the data block metadata table. A unique migration task ID is assigned to each data block to be migrated for easy subsequent tracking and management. Create a migration task queue and insert the data blocks to be migrated into the queue according to their priority, including fields such as data block ID, source address, destination address, data capacity, and estimated migration time. Update system statistics, recording key indicators such as the total number of data blocks to be migrated, total capacity, and priority distribution. Generate a migration plan document, including migration batch division, estimated completion time, and resource requirement assessment, providing a basis for decision-making in subsequent cache channel establishment and migration execution. Simultaneously, trigger a cache space pre-allocation operation to reserve corresponding storage space for the data blocks to be migrated in the extended cache area.
[0055] Step 104: Establish cache channels corresponding to each data block to be migrated through the TSV via array and cross switch matrix integrated in the silicon interposer, and determine the cache optimization path according to each cache channel. The cache channel is used for data transmission between the preset extended cache layer and NAND flash memory.
[0056] Silicon interposers are an advanced semiconductor packaging technology that uses a silicon substrate as an intermediate carrier to achieve high-density interconnects between different chips. Through-Via (TSV) arrays are conductive via structures that vertically penetrate the silicon wafer, providing vertical electrical connections between chips with low latency and high bandwidth. Cross-connector matrices are programmable network switching structures that configure arbitrary input-to-output connection paths by controlling the on / off states of switch nodes. Cache channels are dedicated data transmission paths connecting NAND flash memory and extended cache layers, encompassing both physical connectivity and logical control. Cache-optimized paths are data transmission strategies optimized based on cache channel configuration and transmission requirements.
[0057] Specifically, cache channels are established and optimized through hardware resource management and path configuration algorithms. First, cache channels are established: the TSV via array status is scanned to obtain the load level, signal integrity, and available bandwidth information of each via, forming a via status matrix. The channel bandwidth requirement is calculated based on the data capacity of the data block to be migrated and the cache priority, using the formula BW_required = block_size × priority_factor / transfer_time_limit, where priority_factor is the priority coefficient [1.5, 1.0, 0.7]. A load balancing algorithm is used to allocate TSV via groups to each data block, prioritizing vias with lower load and sufficient bandwidth. The connection status table of the cross-connector matrix is queried to identify available switch nodes and occupied connection paths. A matrix path planning algorithm is used to configure the switch path from the NAND flash controller to the extended cache layer, establishing a point-to-point data transmission connection. Then, the cache optimization path is determined: the transmission performance parameters of each cache channel are analyzed, including latency, bandwidth utilization, and error rate. A multi-objective optimization algorithm is used to comprehensively evaluate transmission efficiency. The objective function is F = w1 × (1 / latency) + w2 × bandwidth_utilization - w3 × error_rate, with weight coefficients [w1, w2, w3] = [0.4, 0.4, 0.2]. Based on the optimization results, an optimal transmission path is assigned to each data block to be migrated, and a path configuration table is generated, including channel ID, TSV allocation, switch settings, and transmission parameters.
[0058] In one possible implementation, a cache channel corresponding to each data block to be migrated is established through a TSV via array and a crossbar switch matrix integrated within the silicon interposer, specifically including steps 1041-1043, as follows: Step 1041: Obtain the current load status of each TSV via in the TSV via array and the connection status of each switch node in the cross switch matrix; determine the channel bandwidth requirement corresponding to each data block to be migrated based on the data capacity and cache priority of each data block to be migrated.
[0059] Current load status refers to the data transmission load carried by a TSV via at a specific moment, including indicators such as bandwidth utilization, signal strength, and thermal power consumption. Connection status refers to the on / off configuration of each switch node in the crossbar switch matrix, indicating the connection relationship between input and output ports. Channel bandwidth requirement refers to the minimum data transmission bandwidth required for a single data block to complete a migration operation, influenced by data capacity, transmission time requirements, and priority.
[0060] Specifically, basic data for resource allocation is obtained through hardware status monitoring and demand calculation algorithms. First, hardware status information is acquired: real-time status data for each via is collected via the built-in monitoring circuit of the TSV via array, including current transmission bandwidth utilization, signal transmission delay, power consumption level, and temperature status. The monitoring module collects status data every 100 milliseconds and calculates the average load rate over the past second: Load_avg = Σ(instantaneous bandwidth occupancy) / number of samples. The control registers of the crossbar switch matrix are scanned, and the connection status bits of each switch node are read to form a connection status matrix, where 1 indicates connection and 0 indicates disconnection. The fault and maintenance status of the switch nodes are checked, and unusable nodes are marked. Then, the channel bandwidth requirement is calculated: a bandwidth requirement calculation model is established, with the basic formula BW_basic = block_size / target_time, where block_size is the data block capacity and target_time is the target transmission time. Based on the cache priority adjustment factor, high-priority data blocks use 1.5 times bandwidth redundancy, medium-priority blocks use 1.2 times, and low-priority blocks use 1.0 times. The calculation formula is BW_required = BW_basic × priority_multiplier. Considering transmission protocol overhead and error retransmission probability, a 15% bandwidth margin is added, resulting in a final requirement of BW_final = BW_required × 1.15. The bandwidth requirement information for each data block is recorded in the requirement allocation table, including fields such as data block ID, calculated bandwidth requirement value, priority weight, and estimated transmission time.
[0061] Step 1042: Based on the bandwidth requirements of each channel and the current load status, allocate the corresponding TSV via group to each data block to be migrated in the TSV via array.
[0062] A TSV via group is a transmission unit formed by combining multiple TSV vias, providing higher data bandwidth and reliability through parallel transmission. The allocation operation refers to the resource scheduling process of assigning a specific TSV via group to a data block based on demand and resource availability.
[0063] Specifically, optimized allocation of TSV via groups is achieved through resource matching and load balancing algorithms. First, resource assessment and matching are performed: the available bandwidth of each TSV via is analyzed, calculated as Available_BW = Max_BW × (1 - Load_avg), where Max_BW is the maximum bandwidth capacity of the via. A via combination algorithm is established, employing a greedy strategy to combine vias starting with the vias with the highest available bandwidth until the combined bandwidth meets the data block requirements. For each data block to be migrated, the minimum via combination that meets the bandwidth requirements is searched, with the optimization objective being to minimize the number of vias used to improve resource utilization efficiency. The physical proximity of via combinations is checked, prioritizing vias located close to each other on the silicon interposer to reduce signal interference and transmission delay. Then, load balancing allocation is performed: a load balancing algorithm is used to avoid overloading some vias, calculating the global load variance Variance = Σ(Load_i - Load_avg)² / N, where N is the total number of vias. When load imbalance is detected, the via allocation scheme is readjusted, migrating tasks on high-load vias to low-load vias. Establish a conflict detection mechanism to ensure that the same via is not occupied by multiple data blocks simultaneously. Record the final TSV via group allocation results, including data block IDs, a list of allocated via numbers, total bandwidth capacity, and estimated occupancy time. Update the resource allocation table of the TSV via array, marking the status of allocated vias to provide a basis for subsequent transmission scheduling and resource reclamation.
[0064] Step 1043: For each data block, based on the connection status of the TSV via group and each switch point corresponding to the data block, configure the switch path from NAND flash memory to the preset extended cache layer through the cross switch matrix to obtain the cache channel corresponding to the data block to be migrated.
[0065] A switching path is a signal transmission channel from the input port to the output port established by a crossbar switch matrix, defined by the connection states of a series of switch nodes. Path configuration refers to the process of establishing a specific transmission path by controlling the on / off states of each switch node in the crossbar switch matrix. A cache channel is a complete data transmission link formed by combining TSV via groups and switching paths, realizing an end-to-end connection between NAND flash memory and the extended cache layer.
[0066] Specifically, a complete data transmission channel is established through path planning and switch configuration algorithms. First, path planning is performed: the corresponding input and output ports in the cross-connect switch matrix are determined based on the physical location of the TSV via groups. A shortest path algorithm is used to search for the optimal path from the NAND flash port to the extended cache layer port in the cross-connect switch matrix, with optimization objectives including minimum hop count, minimum latency, and minimum collision probability. The current connection status of each switch node on the path is checked, identifying occupied nodes and available alternative nodes. A path collision detection mechanism is established to ensure that newly created paths do not interfere with existing active paths. When a path collision is detected, a rerouting algorithm is initiated to find an alternative path or adjust the existing path configuration. Then, switch configuration is performed: a sequence of switch node configuration instructions is generated based on the planned optimal path, including node coordinates, connection status, and priority information. Configuration instructions are issued through the control interface of the cross-connect switch matrix to sequentially set the connection status of each switch node on the path. The correctness of the path configuration is verified by sending test signals to check end-to-end connectivity and signal integrity. Complete cache channel information is recorded, including data block ID, TSV via group number, switch path description, total bandwidth capacity, transmission latency, and configuration timestamp. Establish a channel status monitoring mechanism to track the usage and performance indicators of the cached channel in real time, providing data support for subsequent transmission optimization and fault handling.
[0067] Step 105: When the transmission bandwidth utilization of any cache channel exceeds the bandwidth utilization threshold, adjust the cache optimization path according to the deviation rate between the transmission bandwidth utilization and the bandwidth utilization threshold.
[0068] Transmission bandwidth utilization refers to the ratio of the actual data transmission bandwidth currently used by a cached channel to the maximum available bandwidth of that channel, usually expressed as a percentage. The bandwidth utilization threshold is a system-preset upper limit for bandwidth utilization, used to trigger path adjustment mechanisms to prevent channel overload from affecting transmission performance. Deviation rate is the relative difference between the current transmission bandwidth utilization and the bandwidth utilization threshold, reflecting the severity of overload. For example, if a cached channel has a maximum bandwidth of 1GB / s and is currently using 850MB / s (85% utilization), and the threshold is set to 80%, then the deviation rate is 6.25%.
[0069] Specifically, intelligent optimization of the cache path is achieved through dynamic monitoring and adaptive adjustment algorithms. First, bandwidth utilization monitoring is performed: a bandwidth monitoring module is deployed on each cache channel, collecting real-time data transmission volume every 50 milliseconds, and calculating the instantaneous bandwidth utilization as the current bandwidth divided by the maximum bandwidth multiplied by 100%. A sliding window averaging algorithm is used to calculate a stable utilization rate, with a window size of 20 sampling points. The formula is: average utilization rate equals the sum of all instantaneous utilization rates divided by 20. A bandwidth utilization threshold of 75% is set; when the average utilization rate exceeds this threshold, a path adjustment mechanism is triggered. Then, the deviation rate and adjustment strategy are calculated: the deviation rate is calculated as the average utilization rate minus the threshold, then divided by the threshold and multiplied by 100%, quantifying the degree of overload. The adjustment intensity is determined based on the deviation rate, establishing a tiered adjustment mechanism: a slight adjustment is performed when the deviation rate is 5-15%, reallocating some data streams; a moderate adjustment is performed when the deviation rate is 15-30%, enabling backup transmission paths; and a severe adjustment is performed when the deviation rate exceeds 30%, reconfiguring the overall path configuration. A load-sharing algorithm is used to migrate some transmission tasks from overloaded channels to less loaded backup channels. Reconfigure path connections using a crossbar switch matrix to establish multi-path parallel transmission and distribute the load. Monitor the adjustment effect in real time to verify the stability and transmission efficiency of the new path configuration, ensuring that the adjusted bandwidth utilization drops below the threshold. Record path adjustment logs, including the reasons for the adjustment, a comparison of utilization before and after the adjustment, and the performance improvement effect.
[0070] Step 106: Based on the cache optimization path, perform cache migration operations for each data block to be migrated.
[0071] Cache migration refers to the complete data transfer process of moving data blocks from NAND flash memory storage location to extended cache layer storage location, including data reading, transmission, writing and status update operations.
[0072] Specifically, a multi-stage coordinated execution algorithm is used to achieve batch migration of data blocks. First, migration preparation and scheduling are performed: the transmission order of each data block to be migrated is determined based on the cache optimization path, and a migration queue is established according to cache priority and data dependencies. Corresponding cache channel resources, including TSV via groups, cross-connect paths, and target storage addresses, are allocated to each data block. Storage space is pre-allocated in the extended cache layer, and a mapping table from source address to target address is established. Then, batch migration and transmission are performed: a pipelined processing approach is adopted, dividing the migration process into three parallel stages: read, transmit, and write. Data read requests are initiated from the NAND flash controller, and reads are performed in 4KB block granularity to improve transmission efficiency. Data transmission is conducted through the configured cache channels, and bandwidth utilization and error rate metrics are monitored during transmission. Data write operations are performed in the extended cache layer, and a write confirmation mechanism is used to ensure data integrity. Finally, status updates and verification are performed: after data transmission is completed, the storage location information of the data blocks is updated, and the original data in the NAND flash memory is marked as reclaimable. The index table of the extended cache layer is updated, establishing a new mapping relationship from logical address to physical address. Perform data consistency checks, verifying the integrity of the migrated data using a checksum algorithm. Release cached channel resources used during the migration process and update the status information of the TSV via array and cross switch matrix. Generate a migration completion report, recording key metrics such as the number of successfully migrated data blocks, total transmission time, average transmission rate, and number of error recovery attempts.
[0073] In the above embodiments, a cache optimization framework based on high-bandwidth memory was implemented through access heat calculation, dynamic capacity allocation, and cache channel establishment. To further improve the reliability and adaptability of the data migration process and reduce the impact of transmission errors on cache performance, this application also provides another fault-aware adaptive migration execution control method. This method intelligently adjusts the migration strategy by real-time monitoring of the transmission error rate and dynamic reconfiguration of cache channels, enabling the system to more accurately handle cache migration needs in complex transmission environments and hardware failures. The following section combines... Figure 2 Another cache optimization method based on high-bandwidth memory is described in the embodiments of this application: Please see Figure 2 This is another flowchart illustrating a cache optimization method based on high-bandwidth memory in an embodiment of this application.
[0074] Step 201: Determine the migration order and corresponding cache channel for each data block to be migrated based on the cache optimization path.
[0075] Migration order refers to the sequence in which cache migration operations are performed on multiple data blocks to be migrated. It is determined based on the cache priority, data dependencies, and transmission resource allocation of the data blocks. The migration order directly impacts overall migration efficiency and system performance. The corresponding cache channel refers to the dedicated data transmission path allocated to each data block to be migrated, which includes specific TSV via groups and cross-connect switch path configurations.
[0076] Specifically, the optimal migration execution plan is determined through priority ranking and resource matching algorithms. First, migration order planning is performed: a multi-factor ranking model is established, with cache priority as the primary ranking factor (weight coefficient 0.6), data block capacity as the secondary ranking factor (weight coefficient 0.3), and data access frequency as the auxiliary ranking factor (weight coefficient 0.1). The comprehensive ranking score is calculated as: ranking score = cache priority multiplied by 0.6 + capacity normalized value multiplied by 0.3 + access frequency normalized value multiplied by 0.1. All data blocks to be migrated are arranged in descending order of ranking score, generating an initial migration order list. Logical dependencies between data blocks are checked to ensure that the migration order of dependent data blocks meets access constraints. Then, cache channel allocation is performed: based on the migration order and bandwidth requirements of the data blocks, a greedy allocation algorithm is used to allocate optimal cache channels to each data block. Priority is given to allocating cache channels with the best transmission performance to data blocks with high ranking scores, including channel combinations with the lowest latency, highest bandwidth, and lowest error rate. A channel conflict detection mechanism is established to ensure that multiple data blocks do not compete for the same cache channel resource within the same time period. Generate a migration plan table, recording the migration sequence number, allocated cache channel identifier, estimated start time, and estimated completion time for each data block. Establish a dynamic adjustment mechanism to reallocate backup cache channels and update the migration sequence when a channel failure or performance degradation is detected.
[0077] Step 202: Read the data blocks to be migrated from the NAND flash memory in sequence according to the migration order, and transfer the data blocks to be migrated to the preset extended cache layer through the corresponding cache channel.
[0078] Sequential reading refers to the process of retrieving the contents of data blocks to be migrated from the NAND flash memory one by one according to a predetermined migration order, using either a serial or parallel reading strategy. The transmission process refers to the complete flow of data blocks from the source storage location to the target storage location through a cache channel, including data encapsulation, path transmission, and integrity verification.
[0079] Specifically, efficient data migration is achieved through pipelined transmission and concurrency control algorithms. First, sequential read operations are performed: data read requests are sent to the NAND flash controller sequentially according to the migration order list. Each read request includes the logical address of the data block, data length, and read priority information. A pre-read buffer mechanism is employed, pre-reading the next data block into a temporary buffer while the current data block is being transmitted, reducing read latency. Large data blocks are divided into multiple 4KB transmission units, and a pipelined approach is used for fragmented reading and transmission, improving the continuity of data flow. NAND flash read performance metrics, including read latency, throughput, and error retries, are monitored, and read parameters are dynamically adjusted to optimize performance. Then, buffer channel transmission is performed: the read data blocks are encapsulated according to a predetermined format, adding control information such as data block identifiers, checksums, and transmission timestamps. Data transmission is initiated through the allocated buffer channel, utilizing the parallel transmission capability of TSV vias to increase transmission bandwidth. During transmission, the channel status is monitored in real time, including transmission rate, signal quality, and error rate metrics, to ensure transmission stability. Data integrity verification is performed at the extended buffer layer receiver, using a CRC check algorithm to verify the correctness of the transmitted data. Finally, transmission confirmation and status update are performed: After data is successfully written to the extended cache layer, a transmission completion confirmation signal is sent, and the storage location mapping table of the data block is updated. Currently used cache channel resources are released to prepare for the transmission of the next data block to be migrated. Transmission statistics, including actual transmission time, transmission rate, and resource utilization, are recorded for subsequent performance optimization and fault analysis.
[0080] Step 203: During data transmission, monitor the transmission error rate of each buffer channel; when the transmission error rate exceeds the preset error rate threshold, pause the transmission of the current data block to be migrated, reconfigure the switching path of the buffer channel through the cross switch matrix, and perform the transmission operation of the current data block to be migrated according to the reconfigured buffer channel.
[0081] The transmission error rate (ERR) is the ratio of the number of data packets that are erroneously transmitted during data transmission in a buffer channel to the total number of transmitted data packets. It is usually expressed as a percentage and reflects the channel's transmission quality and stability. The preset error rate threshold is a pre-set upper limit for the transmission error rate, used as a criterion for triggering channel reconfiguration to ensure the reliability of data transmission. Reconfiguration refers to the process of changing the data transmission path by adjusting the connection status of the switch nodes in the crossbar switch matrix, establishing a new connection relationship from the input port to the output port.
[0082] Specifically, high reliability of data transmission is ensured through real-time monitoring and dynamic reconfiguration algorithms. First, transmission error rate monitoring is implemented: an error detection module is deployed on each buffer channel, using a cyclic redundancy check (CRC) algorithm to perform real-time verification of transmitted data packets, recording the number of packets failing verification. A sliding window statistical mechanism is established, with 100 data packets per statistical window. The transmission error rate within the window is calculated as: error rate = number of erroneous data packets divided by the total number of data packets within the window multiplied by 100%. A preset error rate threshold of 0.3% is set; when the average error rate of three consecutive statistical windows exceeds this threshold, a reconfiguration mechanism is triggered. The timestamp of the error, error type, and affected data block identifier are recorded, establishing an error statistics database for fault analysis. Then, pause and reconfiguration operations are performed: upon detecting an error rate exceeding the threshold, a pause transmission command is immediately sent, stopping the transmission of the current data block on the buffer channel and protecting the untransmitted data block content. An alternative path search algorithm is initiated, searching for available alternative transmission paths in the cross-connector matrix, prioritizing the path combination with the lowest historical error rate. The connection status of the switch nodes is reconfigured through the cross-connector matrix control interface, disconnecting the switch connections on the faulty path and establishing a new transmission path. Perform connectivity tests on the new path, sending test packets to verify the path's transmission quality and stability. Once the error rate is confirmed to be below the threshold, resume data transmission. Restart the transmission operation of the currently migrated data block, continuing the data transmission task from the paused position to maintain the continuity and integrity of the transmission process.
[0083] In one possible implementation, the switching path of the buffer channel is reconfigured through a cross-switch matrix, specifically including steps 2031-2034, as follows: Step 2031: Obtain the current connection status and load status of each switch node in the cross switch matrix.
[0084] Load status refers to the current data transmission load carried by each switch node in the crossbar switch matrix, including indicators such as data traffic, processing latency, and power consumption levels through that node. Load status reflects the workload and remaining processing capacity of the switch nodes and is used to evaluate node availability and performance.
[0085] Specifically, complete operational status information of the crossbar switch matrix is obtained through status scanning and load assessment algorithms. First, connection status acquisition is performed: the connection configuration information of each switch node is read through the status register interface of the crossbar switch matrix, scanning the status bits of all M×N switch nodes in the matrix. A connection status matrix data structure is established, where a value of 1 indicates the corresponding switch node is connected, 0 indicates it is disconnected, and -1 indicates a node failure. The input and output port numbers of each connected node are recorded, constructing a complete connection topology. The validity of the connection status is checked to verify for conflicting connections or invalid configurations, ensuring that each input port is connected to at most one output port. Next, load status monitoring is performed: load monitoring circuits are deployed at each switch node to collect real-time data traffic, transmission delay, and power consumption parameters. The node load rate is calculated using the formula: load rate equals current data traffic divided by the node's maximum processing capacity multiplied by 100%, where the node's maximum processing capacity is typically 500MB / s. Signal transmission delay is measured by sending timestamped test signals to calculate the end-to-end delay time, including internal node processing delay and signal propagation delay. Monitor node power consumption status, read real-time power consumption data provided by the power management unit, and calculate heat dissipation density to assess node thermal stability. Establish a load status database to record historical load change trends for each node, used for load forecasting and capacity planning. Generate a cross-connector matrix status report, including a connection topology diagram, load distribution diagram for each node, and performance statistics, providing a basis for subsequent fault location and path reconfiguration decisions.
[0086] Step 2032: Based on the transmission error rate, locate the fault switch node corresponding to the target buffer channel where the transmission error occurred.
[0087] The target buffer channel refers to the specific buffer channel where a transmission error has occurred and fault location is required. It consists of TSV via groups and cross switch paths. The faulty switch node refers to the specific cross switch node in the transmission path of the target buffer channel that caused the transmission error, manifested as abnormal states such as degraded signal transmission quality, increased latency, or unstable connection.
[0088] Specifically, error correlation analysis and fault isolation algorithms are used to accurately locate faulty nodes in the buffer channel. First, error source localization is performed: based on the buffer channel identifier with an excessive transmission error rate, the complete transmission path configuration information of that channel is queried, including the TSV via group used and the sequence of cross-connection nodes traversed. An error propagation model is established to analyze the distribution characteristics of transmission errors in the channel path, and the fault area is determined by comparing the error occurrence frequency of different path segments. A binary search algorithm is used to narrow down the fault range, dividing the transmission path into multiple segments, testing the transmission quality of each segment separately, and gradually locating the specific faulty node. Diagnostic test signals are sent, transmitting known test data patterns through each switch node, and comparing the data differences between the receiving and transmitting ends to identify the fault location. Then, faulty node confirmation is performed: deep detection is performed on suspected faulty switch nodes, measuring the node's signal integrity parameters, including signal amplitude, rise time, and jitter level. The electrical characteristics of the node are checked, measuring parameters such as contact resistance, insulation impedance, and leakage current to determine if a hardware fault exists. Historical load data of the node is analyzed to calculate the performance change trend before and after the fault, confirming the correlation between performance degradation and transmission errors. Establish a fault severity assessment model to calculate a fault level score based on the increase in error rate, the degree of performance degradation, and the scope of impact. The formula is: Fault level = Error rate increment multiplied by 0.4 + Delay increment multiplied by 0.3 + Number of affected channels multiplied by 0.3. Record detailed information about fault switch nodes, including node coordinates, fault type, detection time, fault level, and a list of affected cached channels. Establish a fault database for fault prediction and preventative maintenance.
[0089] Step 2033: Based on the current connection status and load status of each switch node, select standby switch nodes from the cross switch matrix whose load status is lower than the preset load threshold and which have not experienced transmission errors.
[0090] The preset load threshold is a pre-defined upper limit for the load status of the switch nodes, used to determine whether a node has the capacity to handle additional transmission tasks. It is typically set to 70% of the node's maximum processing capacity. A backup switch node refers to a switch node in the cross-connect matrix that currently has a low load level and is operating normally. It serves as an alternative to a faulty node, ensuring redundancy and reliability of the data transmission path.
[0091] Specifically, suitable backup transmission resources are identified through load filtering and availability assessment algorithms. First, load status filtering is performed: all switch nodes in the cross-connect matrix are traversed, and real-time load status data for each node is read, including current data traffic, processing latency, and load rate. A preset load threshold of 70% is set, and switch nodes with load rates below this threshold are selected as candidate backup nodes. The remaining processing capacity of each candidate node is calculated using the formula: remaining capacity equals the node's maximum processing capacity multiplied by 1 minus the current load rate, where the node's maximum processing capacity is typically 500MB / s. A candidate node list is created, recording information such as node coordinates, current load rate, remaining processing capacity, and last update time. Next, transmission error filtering is performed: the transmission error history of each candidate node is queried, with the error filtering time window set to the past 2 hours, and the number and error rate of transmission errors within the time window are counted. Nodes with zero transmission errors and a zero error rate are selected to ensure good transmission reliability for backup nodes. The hardware status of the candidate nodes is checked, including electrical characteristics, signal integrity, and temperature status, eliminating nodes with potential failure risks. The connectivity reachability of the evaluation nodes is verified to ensure that candidate nodes can establish an effective transmission path from NAND flash memory to the extended cache layer. A priority ranking of backup nodes is established, taking into account remaining processing capacity, historical reliability, and connectivity convenience. The priority score is calculated using the formula: priority score = remaining capacity multiplied by 0.5 + reliability coefficient multiplied by 0.3 + connectivity convenience coefficient multiplied by 0.2. A final list of backup switch nodes is generated, sorted in descending order of priority score, providing a candidate resource pool for subsequent target node selection.
[0092] Step 2034: Based on the available bandwidth of the standby switch node, select the target standby switch node that meets the channel bandwidth requirements of the current data block to be migrated, and establish a new switch path from NAND flash memory to the preset extended cache layer according to the target standby switch node, thus completing the switch path reconfiguration of the cache channel.
[0093] Available bandwidth refers to the remaining data transmission bandwidth capacity that the standby switch node can currently provide, which is equal to the node's maximum bandwidth capacity minus the currently occupied bandwidth. The target standby switch node is a specific node selected from the standby switch nodes to ultimately establish a new transmission path, and it needs to meet the bandwidth requirements and transmission quality requirements of the data blocks to be migrated. The new switch path is an alternative transmission channel from NAND flash memory to the extended cache layer established by reconfiguring the target standby switch node in the cross-connect matrix.
[0094] Specifically, the reliability of the cache channel is restored through bandwidth matching and path reconstruction algorithms. First, target backup node selection is performed: the channel bandwidth requirement of the data block to be migrated is obtained, which has been determined in the previous bandwidth requirement calculation. The list of backup switch nodes is traversed, and the available bandwidth of each node is calculated using the formula: available bandwidth equals the node's maximum bandwidth capacity minus the currently occupied bandwidth, where the node's maximum bandwidth capacity is typically 500MB / s. Backup nodes with available bandwidth greater than or equal to the channel bandwidth requirement are selected to ensure that transmission performance requirements are met. An optimal matching algorithm is used to select the optimal target node from the backup nodes that meet the bandwidth requirements. Optimization objectives include minimizing bandwidth waste, minimizing transmission latency, and maximizing reliability. A comprehensive score for each candidate node is calculated using the formula: score equals bandwidth matching degree multiplied by 0.4 plus latency performance multiplied by 0.3 plus reliability index multiplied by 0.3. The node with the highest score is selected as the target backup switch node. Then, a new switch path is established: the location coordinates of the target backup switch node are analyzed to determine its input and output port connections in the cross-connect matrix. A complete transmission path from the NAND flash controller to the extended cache layer is designed, with the path passing through the target backup switch node to achieve data transmission. Configure the connection status of the target backup switch node through the crossbar switch matrix control interface, establishing the electrical connection from the input port to the output port. Perform path connectivity testing, sending test data to verify the transmission quality and stability of the new switch path, and measuring end-to-end transmission delay and error rate metrics. After completing the switch path reconfiguration, update the path information of the buffer channel, including the new switch node coordinates, path description, and performance parameters. Record detailed information about the reconfiguration operation, including the fault node identifier, target backup node identifier, configuration time, and performance improvement effect, establishing a configuration history database for subsequent fault analysis and optimization.
[0095] Step 204: After each data block to be migrated has been migrated, update the cache index table of the preset extended cache layer and mark the cache status of the migrated data blocks in the NAND flash memory.
[0096] The cache index table is a data structure in the extended cache layer used to record the storage location and status information of data blocks. It includes fields such as logical address, physical address, data length, storage time, and access permissions. Cache status refers to the current status of a data block in the storage system, indicating whether the data block has been migrated to the cache layer and its availability in its original storage location.
[0097] Specifically, the system state update after migration is completed through index management and state synchronization algorithms. First, the cache index table is updated: all migrated data blocks are traversed to obtain their new storage location information in the extended cache layer. A new index record is created for each migrated data block in the cache index table, with fields including data block identifier, logical address, physical address, data capacity, migration completion time, cache priority, and access counter. A hash index structure is established to improve query efficiency, using the data block logical address as the hash key to achieve fast location with O(1) time complexity. Global cache statistics are updated, including used cache capacity, remaining available capacity, cache hit rate, and data block distribution. An index backup mechanism is established to synchronize the updated cache index table to the backup storage location to prevent index data loss. Then, NAND flash status marking is performed: a cache status flag is set for each migrated data block in the NAND flash metadata area, with a flag value of 1 indicating a cached state. The data block management table in the NAND flash is updated, modifying the data block status field from uncached to cached, and recording the cache layer address information. Establish a bidirectional reference relationship, recording the storage address of the data block in the extended cache layer in the NAND flash memory, and recording the original address of the data block in the NAND flash memory in the extended cache layer. Perform a consistency check operation to verify whether the state information in the NAND flash memory and the extended cache layer remains synchronized, ensuring system state consistency. Generate a migration completion report, summarizing key performance indicators such as the number of successfully migrated data blocks, total migration time, average transfer rate, number of error recovery attempts, and final cache utilization.
[0098] Reference Figure 3 This application provides a cache optimization system based on high-bandwidth memory. The system includes: a data acquisition module, a data block determination module, a path determination module, and a cache optimization module, wherein: The data acquisition module is used to acquire access data of each data block in the NAND flash memory. The access data includes access frequency, data type and access time interval. The module for determining data blocks to be migrated is used to calculate the access heat value of each data block based on each access data, and classify each data block into cache priorities according to each access heat value to obtain the cache priority corresponding to each data block; according to the number of data blocks of each cache priority, the capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory is adjusted, and multiple data blocks to be migrated are determined according to the adjusted capacity allocation ratio and the cache priority of each data block; The path determination module is used to establish a cache channel corresponding to each data block to be migrated through the TSV via array and cross switch matrix integrated in the silicon interposer, and to determine the cache optimization path according to each cache channel. The cache channel is used for data transmission between the preset extended cache layer and NAND flash memory. The cache optimization module is used to adjust the cache optimization path based on the deviation rate between the transmission bandwidth utilization rate and the bandwidth utilization rate threshold when the transmission bandwidth utilization rate of any cache channel exceeds the bandwidth utilization rate threshold; and to perform cache migration operations for each data block to be migrated based on the cache optimization path.
[0099] Based on the above embodiments, the module for determining the data block to be migrated is further configured to: perform time window weighting on the fluctuation range of the access frequency of each data block within a preset transmission period to obtain a standardized access frequency; determine the data block type according to the data type, and determine the type weight factor corresponding to the data block type in a preset data block type mapping table, wherein the data block type includes sequential access data blocks, random access data blocks, and mixed access data blocks; calculate the access interval variance value of the data block based on each access time interval; and perform weighted fusion calculation on the standardized access frequency, access interval variance value, and type weight factor of the data block to obtain the access popularity value of the data block.
[0100] Based on the above embodiments, the module for determining the data blocks to be migrated is further used to collect the write rate of the NAND flash memory and calculate the deviation ratio between the write rate and the baseline write rate; based on the deviation ratio, the number of data blocks for each cache priority is weighted and corrected to obtain the corrected distribution of the number of data blocks; combined with the average write time and write power consumption of the data blocks corresponding to each cache priority, the write overhead coefficient corresponding to each cache priority is calculated; according to the distribution of the number of data blocks and the write overhead coefficient corresponding to each cache priority, the dynamic capacity distribution coefficient is calculated; the initial capacity allocation ratio between the extended cache area in the preset extended cache layer and the cache area in the NAND flash memory is obtained; the dynamic capacity distribution coefficient is calculated with the initial capacity allocation ratio to obtain the capacity allocation ratio between the extended cache area and the cache area.
[0101] Based on the above embodiments, the module for determining the data blocks to be migrated is further configured to determine the target available capacity of the extended cache area according to the capacity allocation ratio; extract the data blocks corresponding to each cache priority in descending order of each cache priority, and accumulate the data capacity of each data block; when the accumulated data capacity reaches the target available capacity of the extended cache area, mark each extracted data block as a data block to be migrated.
[0102] Based on the above embodiments, the path determination module is further configured to obtain the current load status of each TSV via in the TSV via array and the connection status of each switch node in the cross switch matrix; determine the channel bandwidth requirement corresponding to each data block to be migrated according to the data capacity and cache priority of each data block to be migrated; allocate corresponding TSV via groups to each data block to be migrated in the TSV via array based on the channel bandwidth requirements and the current load status; and for each data block, configure the switch path from NAND flash memory to the preset extended cache layer through the cross switch matrix according to the connection status of the TSV via group corresponding to the data block and each switch point, thereby obtaining the cache channel corresponding to the data block to be migrated.
[0103] Based on the above embodiments, the cache optimization module is further configured to determine the migration order and corresponding cache channel of each data block to be migrated according to the cache optimization path; read the data blocks to be migrated from the NAND flash memory in sequence according to the migration order, and transmit the data blocks to be migrated to the preset extended cache layer through the corresponding cache channel; monitor the transmission error rate of each cache channel during data transmission; when the transmission error rate exceeds the preset error rate threshold, pause the transmission of the current data block to be migrated, reconfigure the switching path of the cache channel through the cross switch matrix, and perform the transmission operation of the current data block to be migrated according to the reconfigured cache channel; after each data block to be migrated is migrated, update the cache index table of the preset extended cache layer, and mark the cache status of the migrated data blocks in the NAND flash memory.
[0104] Based on the above embodiments, the cache optimization module is also used to obtain the current connection status and load status of each switch node in the cross switch matrix; locate the faulty switch node corresponding to the target cache channel where the transmission error occurred based on the transmission error rate; select backup switch nodes with load status lower than a preset load threshold and no transmission error from the cross switch matrix according to the current connection status and load status of each switch node; select the target backup switch node that meets the channel bandwidth requirements of the current data block to be migrated according to the available bandwidth of the backup switch node; and establish a new switch path from NAND flash memory to the preset extended cache layer according to the target backup switch node, thereby completing the reconfiguration of the switch path of the cache channel.
[0105] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0106] This application also discloses an electronic device. (See reference...) Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 400 may include: at least one processor 401, at least one network interface 404, a user interface 403, a memory 405, and at least one communication bus 402.
[0107] The communication bus 402 is used to enable communication between these components.
[0108] The user interface 403 may include a display interface and a camera interface. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.
[0109] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0110] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface graphics, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 401 and may be implemented as a separate chip.
[0111] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory 405 may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. (Refer to...) Figure 3 The memory 405, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program based on a high-bandwidth memory caching optimization method.
[0112] exist Figure 3 In the illustrated electronic device 400, the user interface 403 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 401 can be used to call an application program stored in the memory 405 that uses a cache optimization method based on high-bandwidth memory. When executed by one or more processors 401, the electronic device 400 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0113] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0114] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0115] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0116] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0118] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practical disclosure.
[0119] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only.
Claims
1. A cache optimization method based on high bandwidth memory, characterized in that, The method comprises the following steps: acquiring access data of each data block in a NAND flash memory, the access data comprising access frequency, data type and access time interval; calculating access heat value of each data block based on the access data, and classifying each data block according to the access heat value to obtain the cache priority of each data block; adjusting the capacity allocation ratio of the extension cache area in the preset extension cache layer and the cache area in the NAND flash memory according to the number of data blocks of each cache priority, and determining a plurality of to-be-migrated data blocks according to the adjusted capacity allocation ratio and the cache priority of each data block; establishing a cache channel corresponding to each to-be-migrated data block through a TSV through-hole array and a crossbar matrix integrated in a silicon interposer, and determining a cache optimization path according to each cache channel, the cache channel being used for data transmission between the preset extension cache layer and the NAND flash memory; when the transmission bandwidth usage rate of any cache channel exceeds the bandwidth usage rate threshold, adjusting the cache optimization path according to the deviation rate of the transmission bandwidth usage rate and the bandwidth usage rate threshold; based on the cache optimization path, performing a cache migration operation on each to-be-migrated data block.
2. The method of claim 1, wherein, The method comprises the following steps: for each data block, performing time window weighting processing on the fluctuation amplitude of the access frequency of the data block within a preset transmission period to obtain a standardized access frequency; determining a data block type according to the data type, and determining a type weight factor corresponding to the data block type in a preset data block type mapping table, the data block type comprising sequential access data block, random access data block and mixed access data block; calculating the access interval variance value of the data block based on each access time interval; performing weighted fusion calculation on the standardized access frequency, the access interval variance value and the type weight factor of the data block to obtain the access heat value of the data block.
3. The method of claim 1, wherein, The method comprises the following steps: collecting the write rate of the NAND flash memory, and calculating the deviation ratio of the write rate and the reference write rate; based on the deviation ratio, performing weighted correction on the number of data blocks of each cache priority to obtain a corrected data block number distribution; combining the average write time and write power consumption of the data blocks corresponding to each cache priority, calculating a write overhead coefficient corresponding to each cache priority; calculating a dynamic capacity distribution coefficient according to the data block number distribution and the write overhead coefficient corresponding to each cache priority; obtaining the initial capacity allocation ratio of the extension cache area in the preset extension cache layer and the cache area in the NAND flash memory; calculating the dynamic capacity distribution coefficient and the initial capacity allocation ratio to obtain the capacity allocation ratio of the extension cache area and the cache area.
4. The method of claim 1, wherein, The method comprises the following steps of: determining a plurality of data blocks to be migrated according to the adjusted capacity allocation ratio and the cache priority of each data block; determining the target available capacity of the extended cache area according to the capacity allocation ratio; extracting the data block corresponding to each cache priority in turn according to the descending order of each cache priority, and accumulating the data capacity of each data block; 5. The method of claim 1, wherein, when the accumulated data capacity reaches the target available capacity of the extended cache area, marking each extracted data block as a data block to be migrated. The method comprises the following steps of: obtaining the current load state of each TSV via hole in the TSV via hole array and the connection state of each switch node in the crossbar matrix; determining the channel bandwidth requirement corresponding to each data block to be migrated according to the data capacity and cache priority of each data block to be migrated; allocating a corresponding TSV via hole group for each data block to be migrated in the TSV via hole array based on the channel bandwidth requirement and the current load state; 6. The method of claim 1, wherein, for each data block, configuring a switch path from the NAND flash memory to the preset extended cache layer through the crossbar matrix according to the connection state of the TSV via hole group corresponding to the data block and each switch node, to obtain the cache channel corresponding to the data block to be migrated. The method comprises the following steps of: determining the migration order and the corresponding cache channel of each data block to be migrated according to the cache optimization path; reading the data block to be migrated from the NAND flash memory in turn according to the migration order, and transmitting the data block to be migrated to the preset extended cache layer through the corresponding cache channel; monitoring the transmission error rate of each cache channel during data transmission; when the transmission error rate exceeds a preset error rate threshold, pausing the transmission of the current data block to be migrated, reconfiguring the switch path of the cache channel through the crossbar matrix, and performing the transmission operation of the current data block to be migrated according to the reconfigured cache channel; 7. The method of claim 6, wherein, after the migration of each data block to be migrated is completed, updating the cache index table of the preset extended cache layer, and marking the cache state of the migrated data block in the NAND flash memory. The method comprises the following steps of: obtaining the current connection state and load state of each switch node in the crossbar matrix; locating the fault switch node corresponding to the target cache channel where the transmission error occurs based on the transmission error rate; screening a standby switch node from the crossbar matrix according to the current connection state and load state of each switch node, wherein the standby switch node has a load state lower than a preset load threshold and has not occurred transmission error; According to the available bandwidth of the backup switch node, a target backup switch node satisfying the channel bandwidth requirement of the current data block to be migrated is selected, and a new switch path from the NAND flash memory to the preset extended cache layer is established according to the target backup switch node, so as to complete the switch path reconfiguration of the cache channel.
8. A high bandwidth memory based cache optimization system, comprising: The system comprises: a data acquisition module configured to acquire access data of each data block in the NAND flash memory, the access data including access frequency, data type and access time interval; a data block to be migrated determination module configured to calculate an access heat value of each data block based on the access data, and classify each data block according to the access heat value to obtain a cache priority corresponding to each data block; adjust a capacity allocation ratio of an extended cache area in a preset extended cache layer and a cache area in the NAND flash memory according to a number of data blocks of each cache priority, and determine a plurality of data blocks to be migrated according to the adjusted capacity allocation ratio and the cache priority of each data block; a path determination module configured to establish a cache channel corresponding to each data block to be migrated through a TSV via array and a crossbar matrix integrated in a silicon interposer, and determine a cache optimization path according to each cache channel, the cache channel being used for data transmission between the preset extended cache layer and the NAND flash memory; a cache optimization module configured to adjust the cache optimization path according to a deviation rate of a transmission bandwidth utilization rate of any cache channel and the bandwidth utilization rate threshold when the transmission bandwidth utilization rate exceeds the bandwidth utilization rate threshold, and perform a cache migration operation of each data block to be migrated based on the cache optimization path.
9. An electronic device, comprising: An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is configured to store instructions. The user interface and the network interface are configured to communicate with other devices. The processor is configured to execute the instructions stored in the memory to cause the electronic device to perform the cache optimization method based on high-bandwidth memory according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the cache optimization method based on high-bandwidth memory according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-level cache management method, system and device for solid-state storage device and medium
CN119127091A
Dynamic data distribution storage method and system for RAID (Redundant Array of Independent Disks)
CN119668506A
Method and equipment for optimizing high-bandwidth interface of intelligent storage high-computing power industrial control chip
CN119884003A
Coprocessing system for encrypting and quickly decrypting data of solid state disk
CN120611398A