A memory bank bit flip error real-time detection and correction method and system

CN122653900BActive Publication Date: 2026-09-29SHENZHEN ZHONGXIN HECHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611140284.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-29
Estimated Expiration
2046-07-30

AI Technical Summary

Technical Problem

由于后续读请求已在队列中紧密排队,若系统立即执行深度校验与校正,校正动作所消耗的时延将直接占用当前返回窗口,导致数据无法按时交付,并引发后续读请求在通道中堆积和延迟抖动;若为了保障时延而直接返回数据,又会使未校正的错误数据流出

Benefits of technology

[0016]有益效果:本申请提出的一种内存条位翻转错误实时检测校正方法及系统,通过获取待返回数据及其关联信息,并在当前返回窗口内对数据项进行即时风险判定,输出交付资格状态。当识别为风险项时,立即阻断原始数据的交付授权,并保留请求上下文信息,从而避免了错误数据流出。同时,在独立于当前返回窗口的处理流程中对风险项进行校正,校正完成后,基于请求上下文信息将校正结果与请求标识绑定,并根据链路调度策略将校正结果注入返回路径,替代原始数据完成交付。此外,根据校正处理状态对后续请求队列进行局部分流调度,维持与风险项无关的请求的连续返回,最大限度地保障了系统吞吐量。本申请具有在保障数据可靠性的同时,显著提升内存系统的并发处理能力和响应效率的有益效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653900B_ABST
    Figure CN122653900B_ABST
Patent Text Reader

Abstract

The application provides a memory bank bit flip error real-time detection and correction method and system, which is applied to the technical field of memory bank, obtains to-be-returned data and associated information, and performs instant risk judgment on the data items in the current return window to output a delivery qualification state. When the risk item is identified, the delivery authorization of the original data is blocked, the request context information is retained, and the risk item is corrected. Then, the correction result is bound to the request identifier based on the request context information, and the correction result is injected into the return path according to the link scheduling strategy to replace the original data and complete the delivery. In addition, the subsequent request queue is locally shunted according to the correction processing state, the continuous return of the request unrelated to the risk item is maintained, and the system throughput is maximally guaranteed. Therefore, the application has the beneficial effects of guaranteeing data reliability, significantly improving the concurrent processing capacity and response efficiency of the memory system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of memory module technology, and in particular to a method and system for real-time detection and correction of memory module bit flip errors. Background Technology

[0002] During computer system operation, to ensure the reliability of memory data, bit flip error detection and correction are typically performed during the data read phase. When the system initiates a read request, the memory controller retrieves data from the target storage unit. In the conventional process of returning the data through the data channel and bus interface, the data is simultaneously verified. If a bit flip error is detected, a correction mechanism is triggered, and the corrected data is delivered to the upper layer. However, in high-concurrency read access scenarios such as servers or high-performance computing, the memory controller's interaction port needs to continuously and frequently process a large number of dense read requests, and the data return time window is compressed to an extremely short period. Under these stringent boundary conditions, the real-time triggering of bit flip errors, the data receiving action of the memory controller's interaction port, and the determination of the data receiving status after correction are all highly compressed within the same microsecond or even nanosecond-level read request return window.

[0003] The problem begins to surface when a read data entry is flagged as having a bit-flip risk within the return window. Since subsequent read requests are already tightly queued, if the system immediately performs deep verification and correction, the latency consumed by the correction action will directly consume the current return window, causing data to fail to be delivered on time and leading to a backlog of subsequent read requests and latency jitter. Conversely, if data is returned directly to ensure latency, uncorrected erroneous data will leak out. Existing conventional methods typically employ fixed waiting delays or simple interrupt suspension mechanisms. These methods, under high concurrency, can easily cause a sharp drop in the throughput of the read / write link, failing to smoothly handle and distribute abnormal states within the extremely short return window. Therefore, in the critical data return window under high-concurrency read access, when the memory controller's interaction port receives read data with a bit-flip risk, the risk assessment, correction processing, and return output share the same short window, causing unclear data delivery status or waiting congestion in the return link, thus affecting the real-time detection and correction loop for bit-flip errors in the memory module.

[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0005] In view of the shortcomings of the prior art, this application provides a method and system for real-time detection and correction of memory module bit flip errors, which aims to solve the technical problem that in high-concurrency read access scenarios, the memory controller is unable to effectively detect and correct bit flip errors in real time and smoothly accept abnormal data within a very short data return window, resulting in data delivery delays or erroneous data leakage.

[0006] In a first aspect, this application discloses a real-time detection and correction method for memory module bit flip errors, used for error detection and correction during memory read access. The method includes the following steps: S1: Obtain the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; S2: Perform an immediate risk assessment on the data item to be returned within the current return window and output the corresponding delivery eligibility status; S3: When the delivery eligibility status identifies the data item to be returned as a risk item, block the delivery authorization of the original data corresponding to the risk item within the current return window, and retain its request context information; S4: In a processing flow independent of the current return window, the verification metadata is used to correct the risk item and generate a correction result; S5: Based on the request context information, bind the correction result to the request identifier of the risk item, and inject the correction result into the return path according to the scheduling policy of the return link determined by the link status information, so as to replace the original data to complete the delivery. S6: Based on the correction processing status of the risk item, the queue of requests to be returned after the current return window is partially split and scheduled to maintain the continuous return of requests unrelated to the risk item.

[0007] Furthermore, in step S1, the link status information includes at least the occupancy level of the current return window, the number of requests in the subsequent return request queue, the estimated time interval between two adjacent requests in the subsequent return request queue, the busy / idle status of the fast error correction unit that performs error correction within the current return window, and the bit flip error statistics of the physical address corresponding to the data to be returned in historical access.

[0008] Furthermore, step S2 includes: S21: Calculate the window load risk value of the current return window based on the data to be returned, the verification metadata, and the link status information; S22: Compare the window load risk value with multiple preset thresholds, and combine the judgment result of whether the verification metadata is complete and available, output the corresponding delivery qualification status; The delivery eligibility status includes the status of direct release, the status of release after minor corrections within the current return window, the status of needing to be postponed to the backend for processing, and the status of needing to be reliably reviewed due to the inability to determine the status.

[0009] Further, step S21 includes: S211: Based on the comparison result between the data to be returned and the verification metadata, determine the severity of bit flipping in the current data, and quantify the flipping score according to the severity. S212: Based on the number of requests in the queue of subsequent requests to be returned and the estimated time interval between two adjacent requests, determine the pressure of subsequent requests on the current return window, and quantify the pressure score according to the pressure level. S213: Based on the bit flip error statistics of the physical address corresponding to the data to be returned in the historical access, determine whether the address belongs to the hot spot area with frequent errors, and quantify the error score according to the error frequency. S214: Based on the busy / idle status of the fast error correction unit, determine whether there are currently any idle error correction resources available for use, and quantify the busy / idle score based on the idle or busy status. S215: The load risk value of the window is obtained by weighted summation of the flip score, the pressure score, the error score, and the busy / idle score.

[0010] Furthermore, step S22 includes: S221: When the load risk value of this window is lower than the first threshold value, it is determined to be in a state where it can be directly released; S222: When the window load risk value is between the first boundary value and the second boundary value, it is determined that the state can be slightly modified within the current window. S223: When the window's load risk value is higher than the second threshold value but the metadata is complete, it is determined that the processing needs to be postponed to the background. S224: When the verification metadata is incomplete or the window load risk value is higher than the third threshold and the judgment result is unstable, it is determined that the state cannot be reliably determined and needs to be reviewed. The preset thresholds include the first boundary value, the second boundary value, and the third boundary value.

[0011] Furthermore, when the delivery qualification status is one that can be released after minor corrections within the current return window, one that needs to be postponed to the backend for processing, or one that cannot be reliably determined to require review, the data item to be returned is the risk item, and step S3 includes: S31: When the delivery eligibility status identifies the data item to be returned as a risk item, the original data shall be prohibited from being output from the memory controller to the system bus or the upper layer requester within the current return window; S32: In the return tracking table, switch the request status corresponding to the risk item from normal pending return to abnormal pending acceptance, and retain the request context information of the risk item; The request context information includes at least: the request identifier, the physical address corresponding to the data to be returned, the original sequential position of the risk item in the current return window, the reserved output slot number corresponding to the risk item, and a copy of the original data and the verification metadata.

[0012] Furthermore, step S4 includes: S41: Send the risk item and its request context information into the exception handling buffer. The exception handling buffer is independent of the output queue of the main return link and does not occupy the time resources of the current return window. S42: If the risk item is a minor anomaly that can be corrected at present, then call the fast error correction unit within the current return window to complete the correction and generate the correction result; S43: If the risk item is a serious anomaly that needs to be postponed, then standard error correction and reconstruction will be performed in the anomaly acceptance buffer. The corrected data will be derived using the verification metadata, and the correction result will be generated. S44: If the risk item is in a state that cannot be reliably determined and needs to be reviewed, then a fast reread or redundant consistency review is triggered first. After the review result is stable, the correction result is generated. Alternatively, if the review fails, the risk item is marked as a persistently unstable anomaly and reported.

[0013] Further, step S5 includes: S51: Extract the request identifier of the risk item from the request context information, and bind the request identifier to the correction result one by one, so that the correction result can be identified as the formal return data of the original request corresponding to the risk item. S52: Based on the link status information, determine the return semantic type supported by the current return link, and thus determine the injection method of the correction result; S53: If the current return link supports out-of-order return, the correction result is directly inserted into the return path according to the injection priority, and the request identifier is carried so that the upper layer requester can reassemble the data according to the request identifier; S54: If the current return link supports a return semantic type that requires strict order return, then according to the original order position in the request context information, find the position originally occupied by the risk item in the return queue, accurately fill in the position, and after local alignment and waiting for the limited number of requests immediately following the position, inject the correction result. S55: Send the correction result as the official return data for this risk item to the requester, and mark the delivery final status field corresponding to the request identifier as delivered in the return tracking table.

[0014] Furthermore, the correction processing status includes at least the following states: separated and awaiting acceptance, correction in progress, result awaiting reinjection, reinjection completed, and correction failed; step S6 includes: S61: Iterate through each request in the queue of requests to be returned after the current return window, and determine whether the request has an address association and order dependency with the risk item; S62: When the correction processing status is separated and awaiting acceptance or in the correction process, for requests that are not associated with the risk item and have no sequential dependency, the risk item is allowed to return normally continuously; for requests that are associated with the risk item, the risk item is allowed to continue to return normally, but a temporary risk observation label is attached to the associated physical address at the same time. S63: When the correction processing status is "result pending injection", requests that are not associated with the address of the risk item and have no order dependency are allowed to return normally continuously, and the scheduler prepares to inject the correction result into the return path. S64: When the correction processing status is pending return of results and the system requires strict sequential return, for a preset number of requests in the pending return request queue that are located after the original sequence position of the risk item, a local short-term wait is performed until the correction result is filled in before they are released together; for requests after the preset number, normal return continues. S65: When the correction processing status is "correction failed", mark the risk item as failed and skip it. All requests in the pending return request queue that are after the risk item will not be affected and will continue to return normally.

[0015] Secondly, a real-time detection and correction system for memory module bit flip errors, used to implement the steps of any of the above methods, the system comprising: Acquisition module: Acquires the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; Judgment module: Performs real-time risk assessment on the data item to be returned within the current return window and outputs the corresponding delivery eligibility status; Blocking module: When the delivery eligibility status identifies the data item to be returned as a risk item, it blocks the delivery authorization of the original data corresponding to the risk item within the current return window, while retaining its request context information; Correction module: In a processing flow independent of the current return window, the verification metadata is used to correct the risk item and generate a correction result; Injection module: Based on the request context information, bind the correction result to the request identifier of the risk item, and inject the correction result into the return path according to the scheduling strategy of the return link determined by the link status information, so as to replace the original data to complete the delivery; Scheduling module: Based on the correction processing status of the risk item, the queue of requests to be returned after the current return window is partially distributed and scheduled to maintain the continuous return of requests unrelated to the risk item.

[0016] Beneficial Effects: This application proposes a real-time detection and correction method and system for memory module bit flip errors. It acquires the data to be returned and its associated information, performs real-time risk assessment on the data item within the current return window, and outputs the delivery eligibility status. When a risk item is identified, the delivery authorization of the original data is immediately blocked, while retaining the request context information, thereby preventing erroneous data from leaking out. Simultaneously, the risk item is corrected in a processing flow independent of the current return window. After correction, the correction result is bound to the request identifier based on the request context information, and the correction result is injected into the return path according to the link scheduling strategy, replacing the original data to complete delivery. Furthermore, subsequent request queues are partially distributed and scheduled according to the correction processing status, maintaining the continuous return of requests unrelated to risk items and maximizing system throughput. This application has the beneficial effect of significantly improving the concurrent processing capability and response efficiency of the memory system while ensuring data reliability. Attached Figure Description

[0017] Figure 1 This is a flowchart of a real-time detection and correction method for memory module bit flip errors proposed in this application.

[0018] Figure 2 This is a structural diagram of a real-time detection and correction system for memory module bit flip errors proposed in this application.

[0019] Labeling Explanation: 201. Acquisition Module; 202. Judgment Module; 203. Blocking Module; 204. Correction Module; 205. Injection Module; 206. Scheduling Module. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and marked in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0022] During computer system operation, to ensure the reliability of memory data, bit flip error detection and correction are typically performed during the data reading phase. In the conventional process where the memory controller retrieves data from the target storage unit when a read request is initiated, and then returns the data through the data channel and bus interface, the data is simultaneously verified. If a bit flip error is detected, a correction mechanism is triggered, delivering the corrected data to the upper layer. However, in high-concurrency read access scenarios such as servers or high-performance computing, the memory controller's interaction port needs to continuously and frequently process a large number of intensive read requests, compressing the data return time window to an extremely short duration. Under these stringent boundary conditions, the real-time triggering of bit flip errors, the data reception action of the memory controller's interaction port, and the determination of the data reception status after correction are all highly compressed within the same microsecond or even nanosecond-level read request return window.

[0023] When a piece of read data is determined to have a bit flip risk within the return window, since subsequent read requests are already closely queued in the queue, if the system immediately performs deep verification and correction, the latency consumed by the correction action will directly occupy the current return window, causing the data to fail to be delivered on time and causing subsequent read requests to accumulate in the channel and cause delay jitter; if the data is returned directly in order to ensure latency, uncorrected erroneous data will leak out.

[0024] Existing conventional processing methods typically employ fixed waiting delays or simple interruption suspension mechanisms. Under high concurrency conditions, this approach can easily lead to a sharp drop in the throughput of the read-write link, making it impossible to smoothly handle and distribute abnormal states within a very short return window.

[0025] Therefore, in the critical data return window under high-concurrency read access, when the memory controller interaction port receives read data with a risk of bit flip, the risk judgment, correction processing and return output share the same short window, which causes the data delivery status in the return link to be unclear or to be squeezed, thus affecting the real-time detection and real-time correction closed loop of the memory module for bit flip errors.

[0026] For this, please refer to Figure 1 This application discloses a real-time detection and correction method for memory module bit flip errors, used for error detection and correction during memory read access. The method includes the following steps: S1: Obtain the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; S2: Perform an immediate risk assessment on the data item to be returned within the current return window and output the corresponding delivery eligibility status; S3: When the delivery eligibility status identifies the data item to be returned as a risk item, block the delivery authorization of the original data corresponding to the risk item within the current return window, and retain its request context information; S4: In a processing flow independent of the current return window, the verification metadata is used to correct the risk item and generate a correction result; S5: Based on the request context information, bind the correction result to the request identifier of the risk item, and inject the correction result into the return path according to the scheduling policy of the return link determined by the link status information, so as to replace the original data to complete the delivery. S6: Based on the correction processing status of the risk item, perform partial diversion scheduling on the queue of requests to be returned after the current return window to maintain the continuous return of requests unrelated to the risk item.

[0027] This application effectively solves the contradiction between latency and throughput in memory bit flip error detection and correction in high-concurrency scenarios by introducing real-time risk assessment, blocking delivery authorization, independent correction processing, and intelligent traffic diversion scheduling mechanisms during memory read access, thus ensuring a balance between data reliability and system performance.

[0028] Among them, memory bit flip error refers to the phenomenon that the binary bits stored in the memory cell are unexpectedly flipped during transmission or storage (e.g., 0 becomes 1, or 1 becomes 0). This may be caused by factors such as hardware failure, electromagnetic interference, or cosmic rays.

[0029] Verification metadata is additional information stored or transmitted along with the original data to detect and / or correct errors in the original data, such as parity bits, cyclic redundancy check (CRC) codes, or error correction codes (ECC).

[0030] The request identifier is a unique identifier assigned by the system to each memory read request, used to track and manage the lifecycle of the request.

[0031] Link status information refers to the real-time operating status of the current memory data return path, including but not limited to the occupancy status of the return window, the length of the subsequent request queue, and the busy / idle status of the error correction unit.

[0032] The current return window refers to the limited time period allocated by the memory controller for returning data in response to a read request.

[0033] The core of this application lies in the real-time detection and correction of memory module flip-flop errors, and the optimization of data delivery processes in high-concurrency scenarios. The main technical features of this application will be described in detail below.

[0034] First, in step S1, it is necessary to obtain the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned. For example, after receiving a read request, the memory controller will read the corresponding data from the memory module. Simultaneously, verification metadata corresponding to the data can be generated from a specific area of ​​the memory module or through calculation. The request identifier is usually generated by the system when initiating the read request and appended to the request. Link status information can be collected in real time through sensors or counters within the memory controller; for example, the number of requests processed within the current return window can be simply counted, or the busy / idle status of the fast error correction unit can be obtained through a preset interface.

[0035] Secondly, in step S2, it is necessary to perform an immediate risk assessment on the data item to be returned within the current return window and output the corresponding delivery eligibility status. For example, a simple logic checker can be designed to directly determine a risk item if an inconsistency is found between the data to be returned and the verification metadata. Alternatively, a preset rule can be used, for example, if the verification metadata is missing, it can be directly determined as a risk item.

[0036] Next, in step S3, when the delivery eligibility status identifies the data item to be returned as a risk item, the delivery authorization of the original data corresponding to the risk item needs to be blocked within the current return window, while retaining its request context information. For example, when determined to be a risk item, the memory controller can simply stop sending the data to the system bus and store it in a temporary buffer. Simultaneously, information such as the request identifier and physical address associated with the risk item is recorded for subsequent processing.

[0037] Then, in step S4, in a processing flow independent of the current return window, the check metadata is used to correct the risk item, generating a correction result. For example, an independent error correction unit can be designed to detect and correct errors using its check metadata when a risk item is received. If the check metadata is a simple parity check bit, the error can be corrected by flipping the corresponding bit. If the check metadata is a more complex ECC code, multi-bit error correction can be performed using the ECC algorithm.

[0038] Subsequently, in step S5, based on the request context information, the correction result is bound to the request identifier of the risk item, and the correction result is injected into the return path according to the scheduling policy of the return link determined by the link status information, to replace the original data and complete the delivery. For example, after correction, the correction result can be associated with the previously retained request identifier. Then, based on the link status information, for example, if the return link is currently idle, the correction result can be sent directly. If the return link is processing other requests, the correction result can be placed in a priority queue, waiting for an appropriate time to be injected.

[0039] Finally, in step S6, based on the correction processing status of the risk item, a partial diversion scheduling is performed on the queue of requests to be returned after the current return window to maintain the continuous return of requests unrelated to the risk item. For example, when a risk item is undergoing correction processing, the scheduler can check the requests in the subsequent request queue. If these requests have no address or sequence dependency on the risk item, they can be allowed to continue returning normally without waiting for the correction of the risk item to be completed.

[0040] The real-time detection and correction method for bit flip errors in memory modules proposed in this application aims to solve the latency and throughput challenges faced by bit flip error detection and correction in high-concurrency memory read access scenarios by introducing a series of innovative mechanisms.

[0041] Specifically, during memory read access, the system first acquires the data to be returned, along with its associated verification metadata, request identifier, and link status information. This step provides the necessary foundational data for subsequent risk assessment and correction. Subsequently, the system performs an immediate risk assessment on the data item within the current return window and outputs the delivery eligibility status. This immediate assessment mechanism enables the system to quickly identify potential errors before data delivery, preventing the outflow of erroneous data.

[0042] When a data item is identified as a risk item, the key innovation of this application lies in blocking the delivery authorization of the original data corresponding to the risk item within the current return window, while retaining its request context information. This blocking action ensures that erroneous data is not delivered immediately, while the retained request context information provides a complete tracing basis for subsequent correction and back-injection. Unlike traditional methods that may directly return erroneous data or simply suspend the entire chain, this application selectively blocks risk items, thereby avoiding indiscriminate impact on the entire return chain.

[0043] Furthermore, this application separates the risk item correction process from the current return window. This means that the correction operation no longer consumes valuable real-time return window resources, thereby avoiding subsequent request backlog and latency jitter caused by correction delays. By utilizing verification metadata to correct risk items and generate correction results, the accuracy of the data is ensured.

[0044] After correction, based on the previously retained request context information, the correction result is bound to the request identifier of the risk item. This binding operation ensures that the corrected data can be correctly identified as a response to the original request. Subsequently, according to the return link scheduling strategy determined by the link state information, the correction result is injected into the return path to replace the original data and complete the delivery. This intelligent injection strategy can select the optimal injection method (e.g., out-of-order injection or sequential padding) based on the real-time status of the link, minimizing the impact on system performance.

[0045] Finally, based on the correction status of risk items, this application performs partial diversion scheduling on the queue of requests awaiting return after the current return window, maintaining the continuous return of requests unrelated to risk items. This diversion scheduling mechanism avoids the entire subsequent request queue from stalling due to the correction of a single risk item, thereby maintaining high throughput and low latency for memory read access in high-concurrency scenarios.

[0046] In summary, this application effectively addresses the latency and throughput contradiction faced by existing technologies in bit flip error detection and correction under high-concurrency memory read access scenarios by introducing a series of innovative technologies, including real-time risk assessment, selective blocking, independent correction processing, intelligent back injection, and localized flow scheduling. Compared to traditional fixed waiting delays or simple interrupt suspension mechanisms, the method in this application can handle abnormal situations more flexibly and efficiently, ensuring a balance between data reliability and system performance, and significantly improving the overall efficiency and stability of the memory system.

[0047] In some of the embodiments described above in this application, during memory read access, it is necessary to obtain the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned. Specifically, in order to more comprehensively assess the current operating status and potential risks of the memory system, the link status information can be further refined.

[0048] Furthermore, in step S1, the link status information includes at least the occupancy level of the current return window, the number of requests in the subsequent return request queue, the estimated time interval between two adjacent requests in the subsequent return request queue, the busy / idle status of the fast error correction unit that performs error correction within the current return window, and the bit flip error statistics of the physical address corresponding to the data to be returned in historical access.

[0049] The current return window occupancy rate refers to the proportion of resources or remaining capacity of the time window occupied by data items that have been allocated but not yet delivered in the current memory controller return path. Its purpose is to reflect the busy level of the current return path and provide a basis for subsequent data item delivery decisions.

[0050] Furthermore, the number of requests in the subsequent pending return request queue refers to the total number of memory requests waiting to be processed and returned after the current return window. The expected arrival time interval between two adjacent requests in the subsequent pending return request queue refers to the time difference between two consecutive requests in the queue arriving at the return path, predicted based on historical data or a pre-set model. This information is used to assess the pressure on future return paths for proactive scheduling and risk assessment.

[0051] Furthermore, the busy / idle status of the fast error correction unit performing error correction within the current return window refers to the current workload of the hardware unit used for fast error correction within the current return window. Its purpose is to determine whether there are available error correction resources to handle potential lightweight errors.

[0052] Specifically, the bit-flip error statistics for the physical address corresponding to the data to be returned refer to the cumulative data on the frequency, type, or severity of bit-flip errors occurring at a specific memory physical address over a period of time. This information helps identify hot memory regions, i.e., addresses that are more prone to bit-flip errors, thereby enabling a more rigorous risk assessment of data from these addresses.

[0053] Through the above technical solution, this application can significantly improve the intelligence and adaptability of the real-time detection and correction method for memory module flip-flop errors. Specifically, by acquiring multi-dimensional, fine-grained link status information, the system can more accurately assess the risk level of each data item to be returned and dynamically adjust the error handling strategy based on real-time system resource status and future load predictions. This not only helps to minimize the impact on normal data flow while ensuring data integrity and improving the overall throughput and response speed of the memory system, but also effectively identifies and manages potential hotspot error sources, thereby improving the long-term stability and reliability of the memory system.

[0054] In some of the embodiments described above in this application, real-time detection and correction of memory module flip-flop errors are required during memory read access. A key step is to perform immediate risk assessment on the data item to be returned within the current return window and output the corresponding delivery eligibility status. To this end, this application further proposes a specific implementation method for the aforementioned immediate risk assessment, making risk assessment more refined and intelligent.

[0055] Furthermore, step S2 includes: S21: Calculate the window load risk value of the current return window based on the data to be returned, the verification metadata, and the link status information; S22: Compare the window load risk value with multiple preset thresholds, and combine the judgment result of whether the verification metadata is complete and available, output the corresponding delivery qualification status; The delivery eligibility status includes the status of direct release, the status of release after minor corrections within the current return window, the status of needing to be postponed to the backend for processing, and the status of needing to be reliably reviewed due to the inability to determine the status.

[0056] Calculating the window load risk value of the current return window involves comprehensively considering multiple factors to quantify the overall risk level of the current memory return path. These factors include, but are not limited to, the error status of the data to be returned, verification metadata, the busyness of the return link, and the availability of error correction resources. By weighting or comprehensively evaluating these factors, a numerical value can be obtained that reflects the potential risk of processing and delivering the current data item within the current return window.

[0057] Furthermore, the window load risk value is compared with multiple preset thresholds, and combined with the judgment result of whether the metadata is complete and available, aiming to intelligently determine the data item delivery strategy based on the risk level and metadata status. The preset thresholds can be configured according to factors such as system performance requirements, error tolerance, and resource constraints to divide different risk ranges. The judgment result of whether the metadata is complete and available serves as an important auxiliary judgment basis, because missing or corrupted metadata will seriously affect subsequent error correction capabilities.

[0058] Therefore, the output delivery eligibility status is subdivided into several types to address risks of varying degrees and natures. Specifically: A "can be released directly" status indicates that the data item has extremely low risk and can be delivered directly to the requester via the return path without additional processing.

[0059] A status indicating that a data item has a minor risk, but can be corrected within the current return window using a fast error correction mechanism, and can then be released.

[0060] A status indicating that data items need to be deferred to the backend for processing indicates that the data items have a high risk or require complex error correction and are not suitable for immediate processing within the current return window. They need to be transferred to a separate backend processing flow.

[0061] The inability to reliably determine the status of a data item requiring review indicates that the risk situation is complex or uncertain, and a clear judgment cannot be made based on existing information. Further review mechanisms (such as rereading or redundancy consistency checks) are needed to confirm its status.

[0062] This application's solution refines the immediate risk assessment step S2 into calculating the window load risk value S21 and outputting the delivery eligibility status S22 based on the comparison of the risk value with a threshold and verification of metadata availability. This achieves a refined and multi-dimensional assessment of potential bit-flip errors during the return of memory data. Specifically, by calculating the window load risk value S21, the system can comprehensively consider various real-time and historical information such as the degree of data error, link pressure, historical error hotspots, and error correction resources, thereby obtaining a quantified risk indicator. It is precisely because of this comprehensive risk assessment that the subsequent delivery eligibility status S22 can intelligently select the most appropriate processing strategy based on different risk levels and the judgment results of verifying whether the metadata is complete and available. For example, low-risk data can be quickly released; medium-risk data can be attempted with lightweight correction; high-risk data can be delayed to avoid blocking the main return path; and for cases where reliable judgment is not possible, a review mechanism can be triggered to ensure data integrity. This hierarchical processing mechanism effectively balances the real-time nature and reliability of data delivery.

[0063] In some of the embodiments described above in this application, a window load risk value for the current return window is calculated based on the data to be returned, the verification metadata, and the link status information. However, in its implementation, if the risk value is calculated based on only a single or limited factor, it may not be able to fully and accurately reflect the actual risk situation of memory bit flip errors, thereby affecting the accuracy and reliability of subsequent delivery qualification determination.

[0064] Furthermore, step S21 includes: S211: Based on the comparison result between the data to be returned and the verification metadata, determine the severity of bit flipping in the current data, and quantify the flipping score according to the severity. S212: Based on the number of requests in the queue of subsequent requests to be returned and the estimated time interval between two adjacent requests, determine the pressure of subsequent requests on the current return window, and quantify the pressure score according to the pressure level. S213: Based on the bit flip error statistics of the physical address corresponding to the data to be returned in the historical access, determine whether the address belongs to the hot spot area with frequent errors, and quantify the error score according to the error frequency. S214: Based on the busy / idle status of the fast error correction unit, determine whether there are currently any idle error correction resources available for use, and quantify the busy / idle score based on the idle or busy status. S215: The load risk value of the window is obtained by weighted summation of the flip score, the pressure score, the error score, and the busy / idle score.

[0065] Specifically, in step S211, the flip score refers to the assessment of the number, location, or pattern of bit flip errors in the data by comparing the data to be returned with the verification metadata, thereby quantifying its severity. For example, the degree of flipping can be divided into different levels and assigned corresponding scores based on factors such as the number of flipped bits and whether it affects critical data areas, with the aim of intuitively reflecting the degree of error in the data itself.

[0066] In step S212, the pressure score can be understood as a predictive assessment of the future load on the current return window. Specifically, by analyzing the number of requests in the queue of subsequent requests and the estimated time interval between two adjacent requests, the pressure that subsequent requests may cause to the current return window can be determined. For example, the more requests and the shorter the interval, the greater the pressure and the higher the score. The purpose is to predict and avoid delivery delays or insufficient processing resources caused by congestion of subsequent requests.

[0067] In practical applications, in step S213, the error score is specifically determined based on the bit flip error statistics of the physical address corresponding to the data to be returned in historical accesses. For example, if a physical address has frequently experienced bit flip errors in the past, that address is considered a hotspot area, and its error score will be increased accordingly. The purpose is to identify and focus on high-risk memory areas and issue warnings.

[0068] Furthermore, in step S214, the busy / idle score refers to assessing whether there are sufficient error correction resources available for use based on the real-time busy / idle status of the fast error correction unit. For example, when the fast error correction unit is idle, the busy / idle score is low, indicating that resources are available for immediate error correction; conversely, if the unit is busy, the score is high. The purpose is to incorporate the availability of error correction resources into the risk assessment so as to allocate processing strategies reasonably.

[0069] Finally, in step S215, the window load risk value is obtained by weighted summation of the aforementioned flip-flop score, stress score, error score, and busy / idle score. The weights of each score can be configured according to actual system requirements and experience to reflect the relative importance of different factors in the overall risk assessment. The purpose is to comprehensively consider multi-dimensional risk factors and form a comprehensive risk assessment indicator.

[0070] Specifically, the weights of each score can be determined during the system design or manufacturing phase by injecting different types of bit-flip errors into the target memory platform and conducting extensive testing under varying queue loads and resource consumption conditions. This involves recording the overall system performance (e.g., average return latency, correction success rate, and the degree of anomaly disturbance to subsequent queues) under different weight combinations, and then selecting the set of weight values ​​that optimizes overall system performance as the default configuration. Subsequently, during system operation, the channel-level return jitter, the actual impact of a single anomaly handling on subsequent requests, the reliability of correction results, and the average busy / idle rate of the fast error correction unit are analyzed. When it is detected that the current weight combination leads to excessive window occupation by anomalies or a low correction success rate in a specific scenario, the default weights are fine-tuned based on experience. For example, the weight corresponding to the pressure score may be appropriately increased when queue pressure continues to increase, or the weight corresponding to the error score may be appropriately increased when errors occur frequently at a certain address.

[0071] Through the above technical solution, the real-time detection and correction method for memory bit flip errors achieves a more refined and accurate risk assessment capability. Compared to basic solutions, this application effectively avoids misjudgments or omissions caused by one-sided assessments by comprehensively considering the severity of the data error itself, the future load pressure of the system, the historical error frequency of memory addresses, and the real-time availability of error correction resources. This multi-dimensional and quantitative risk assessment mechanism enables the system to more intelligently identify risk items that truly require special handling, thereby optimizing resource allocation, improving the efficiency and reliability of error detection and correction, and ultimately ensuring the stability of memory data delivery and the overall performance of the system.

[0072] In some preferred embodiments, a specific example is given below. Suppose that during a memory read access, the system obtains data to be returned and needs to perform a risk assessment on it.

[0073] First, in step S211, by comparing the data to be returned with the associated verification metadata, three bit flip errors are found. According to the preset quantization rules, for example, each bit flip is worth 10 points, so the flip score is quantized to 30 points.

[0074] Secondly, in step S212, the system detects that there are 10 requests in the subsequent request queue, and the average expected arrival time between adjacent requests is 5 nanoseconds. According to the preset pressure quantification model, for example, the more requests there are and the shorter the interval, the higher the pressure score. In this case, the pressure score is quantified as 25 points.

[0075] Next, in step S213, the bit flip error statistics of the physical address corresponding to the data to be returned in the historical access are queried, and it is found that the address has had 5 bit flip errors in the past 24 hours. According to the preset error frequency quantization rules, for example, the higher the frequency, the higher the error score, the error score is quantized to 20 points.

[0076] Next, in step S214, the system detects that the fast error correction unit is currently in a busy state and there are no idle resources available for immediate use. Based on a preset busy / idle state quantification rule, for example, if the busy state score is high, the busy / idle score is quantized to 15 points.

[0077] Finally, in step S215, the four scores are weighted and summed. Assume the weight of the reversal score is 0.4, the weight of the pressure score is 0.2, the weight of the error score is 0.2, and the weight of the busy / idle score is 0.2.

[0078] The window load risk value is: (30*0.4)+(25*0.2)+(20*0.2)+(15*0.2)=12+5+4+3=24 points.

[0079] Through the above method, the system obtains a comprehensive window load risk value of 24 points. This score will be used for subsequent delivery eligibility status determination, thereby guiding the system to adopt the most appropriate error handling strategy.

[0080] In some embodiments described above, this application proposes to perform real-time risk assessment on the data items to be returned within the current return window, and output the corresponding delivery eligibility status based on the comparison between the window load risk value and a preset threshold, as well as the judgment result of verifying the completeness and usability of metadata. However, in practical applications, if the threshold setting is too simplistic or the judgment logic is not refined enough, it may fail to adequately distinguish between different levels of risk, leading to over-processing of minor errors or failure to take appropriate countermeasures for serious errors in a timely manner, thereby affecting system performance and data reliability.

[0081] Furthermore, step S22 includes: S221: When the load risk value of this window is lower than the first threshold value, it is determined to be in a state where it can be directly released; S222: When the window load risk value is between the first boundary value and the second boundary value, it is determined that the state can be slightly modified within the current window. S223: When the window's load risk value is higher than the second threshold value but the metadata is complete, it is determined that the processing needs to be postponed to the background. S224: When the verification metadata is incomplete or the window load risk value is higher than the third threshold and the judgment result is unstable, it is determined that the state cannot be reliably determined and needs to be reviewed. The preset thresholds include the first boundary value, the second boundary value, and the third boundary value.

[0082] Specifically, the first, second, and third boundary values ​​mentioned above are preset numerical thresholds used to distinguish different risk levels. They divide the window load risk value into multiple intervals, each corresponding to a specific delivery eligibility status. These boundary values ​​can be calibrated and optimized based on system design, performance requirements, error tolerance, and actual operating data. For example, the first boundary value can be set to a low risk value, indicating that the current memory access environment is very stable, the error probability is extremely low, and system resources are sufficient, allowing data to be directly passed. The second boundary value can be set to a medium risk value, indicating that there is some risk, but it can be quickly resolved through lightweight corrections within the current window. The third boundary value can be set to a high risk value, indicating a higher level of risk, which may require more complex processing or review.

[0083] Specifically, when the window load risk value is below the first threshold, it indicates that the risk of the current data item is extremely low, with negligible impact on system performance. Therefore, it is determined to be in a state where it can be directly released without any additional processing. When the window load risk value is between the first and second thresholds, it indicates that there is a slight risk, but the system still has the ability to perform quick and low-overhead corrections within the current return window. For example, it can complete a minor correction through a fast error correction unit, thus being determined to be in a state where minor corrections can be performed within the current window.

[0084] Furthermore, when the window load risk value is higher than the second threshold, but the verification metadata is complete, this indicates that although the risk is high, the data still has high recoverability due to the complete verification information. Therefore, it is determined that the processing should be postponed to the background to avoid occupying the valuable time resources of the current return window.

[0085] However, when the verification metadata is incomplete, or the window load risk value is even higher than the third threshold and the judgment result is unstable, it indicates that the data reliability is questionable and may not be corrected by conventional means. Therefore, it is judged as a state that cannot be reliably determined and needs to be reviewed, requiring a higher level of diagnosis and processing.

[0086] Through the above technical solution, this application enables more precise and flexible classification and processing of memory bit flip error risks. Compared to relying solely on a single threshold or simple judgment, the introduction of multi-level boundary values ​​and verification of metadata integrity allows the system to select the most appropriate delivery qualification status based on the actual risk level and error correction capability. This not only optimizes the use of system resources and avoids over-processing of low-risk data, but also ensures sufficient attention and proper handling of high-risk data, thereby significantly improving the reliability and efficiency of memory access. Specifically, for data that can be directly released, unnecessary processing overhead is reduced; for data that can be lightly corrected within the current window, rapid error correction is achieved, reducing latency; for data that needs to be delayed for processing in the background, blocking of the main return link is avoided; and for data that cannot be reliably determined and requires review, a final guarantee mechanism is provided. This refined risk assessment and status output effectively improves the intelligence level and overall performance of the memory error detection and correction method.

[0087] In some preferred embodiments, a specific example is given below. Assume that the system's preset first threshold value is set to 20, the second threshold value to 50, and the third threshold value to 80.

[0088] When the system receives data item A to be returned and calculates its window load risk value to be 15, and verifies that the metadata is complete, since 15 is lower than the first threshold value of 20, according to the above scheme, data item A will be determined to be in a state that can be directly released and immediately output from the memory controller to the system bus without any additional error correction processing, thereby achieving rapid delivery.

[0089] When the system receives data item B to be returned and calculates its window load risk value to be 35, and verifies that the metadata is complete, since 35 falls between the first threshold value of 20 and the second threshold value of 50, according to the above scheme, data item B will be determined to be in a state that can be lightly corrected within the current window. At this time, the system will immediately call the fast error correction unit to correct data item B. After the correction is completed, data item B is delivered as the correction result.

[0090] When the system receives data item C to be returned and calculates its window load risk value to be 60, and verifies that the metadata is complete, since 60 is higher than the second threshold of 50 but lower than the third threshold of 80, and the metadata is complete, according to the above scheme, data item C will be determined to be in a state that needs to be delayed for backend processing. At this time, the original data delivery authorization of data item C is blocked, its request context information is preserved, and it is sent to the exception handling buffer for standard error correction and reconstruction processing. It will be injected into the return path after the correction result is generated.

[0091] When the system receives data item D to be returned and calculates its window load risk value to be 90, and the verification metadata is incomplete, data item D will be determined to be in a state that cannot be reliably determined and requires verification, either because the verification metadata is incomplete or the window load risk value of 90 is higher than the third threshold value of 80, according to the above scheme. At this time, the system will trigger a fast reread or redundant consistency verification mechanism to further confirm the true state and recoverability of data item D, avoiding the delivery of uncertain data.

[0092] In some embodiments of this application described above, when a data item to be returned is identified as a risk item, the authorization to deliver its original data needs to be blocked, while retaining its request context information. Specifically, when the delivery eligibility status is a state where it can be released after minor correction within the current return window, a state where it needs to be postponed to the background for processing, or a state where it cannot be reliably determined that it needs to be reviewed, the data item to be returned is the risk item, and step S3 includes: S31: When the delivery eligibility status identifies the data item to be returned as a risk item, the original data shall be prohibited from being output from the memory controller to the system bus or the upper layer requester within the current return window; S32: In the return tracking table, switch the request status corresponding to the risk item from normal pending return to abnormal pending acceptance, and retain the request context information of the risk item; The request context information includes at least: the request identifier, the physical address corresponding to the data to be returned, the original sequential position of the risk item in the current return window, the reserved output slot number corresponding to the risk item, and a copy of the original data and the verification metadata.

[0093] Specifically, when the memory controller performs an immediate risk assessment on the data returned from memory read access within the current return window, if it identifies a data item to be returned as a risk item, and its delivery eligibility status is determined to require minor correction within the current return window before release, need to be delayed for background processing, or cannot be reliably determined to require review, a special handling procedure for that risk item will be triggered. Step S31 aims to ensure that problematic original data is not incorrectly delivered to the system bus or upper-layer requesters. This means that once data is marked as a risk item, its normal output path within the current return window will be immediately interrupted, effectively preventing the propagation of potentially erroneous data.

[0094] Furthermore, step S32 details how to manage the status and retain information of the risk item while blocking the authorization for original data delivery. In the return tracking table, the request status corresponding to the risk item will be updated from "normal pending return" to "abnormal pending acceptance," indicating that the request will leave the normal return process and enter the exception handling path. Simultaneously, for subsequent correction processing and correct data injection, all necessary information related to the risk item will be completely retained, forming request context information. Request context information is the key basis for subsequent correction and injection operations, and it includes at least: a request identifier, used to uniquely identify the memory read request; the physical address corresponding to the data to be returned, used to locate the memory location where the error occurred; the original sequential position of the risk item in the current return window, which is crucial for systems requiring strict sequential returns to ensure that the corrected data can be accurately injected back to its proper position; the reserved output slot number corresponding to the risk item, used to reserve space for the correction result in the return path; and copies of the original data and verification metadata, which are the basis for error correction processing.

[0095] This application's solution achieves timely isolation of potential bit-flip errors by immediately prohibiting the delivery of original data within the current return window upon identifying a risk item and switching its request status to "abnormal and pending," while fully preserving its request context information. This mechanism ensures that the erroneous data is not passed to the upper-layer system before it is corrected, preventing the spread of erroneous data. By preserving detailed request context information, including request identifier, physical address, original sequence position, reserved output slots, and copies of the original data and verification metadata, all necessary conditions are provided for subsequent correction of risk items in an independent processing flow. This allows the corrected data to be accurately bound, scheduled, and injected into the return path based on the original request context information, replacing the original data to complete delivery without affecting the continuous return of other normal requests.

[0096] In some embodiments described above, this application proposes to perform real-time risk assessment on the data item to be returned within the current return window. When the data item is identified as a risk item, the authorization to deliver its original data is blocked, while its request context information is preserved. Subsequently, in a processing flow independent of the current return window, the risk item is corrected using verification metadata to generate a correction result. However, in practical applications, the type and severity of memory module bit flip errors may vary. For example, some errors are minor and easily corrected quickly, while others are more serious or difficult to determine immediately. If a single correction process is used for all risk items, it may lead to inefficient handling of minor errors or inadequate and timely handling of complex errors, thereby affecting the overall system response speed and data reliability.

[0097] Furthermore, step S4 includes: S41: Send the risk item and its request context information into the exception handling buffer. The exception handling buffer is independent of the output queue of the main return link and does not occupy the time resources of the current return window. S42: If the risk item is a minor anomaly that can be corrected at present, then call the fast error correction unit within the current return window to complete the correction and generate the correction result; S43: If the risk item is a serious anomaly that needs to be postponed, then standard error correction and reconstruction will be performed in the anomaly acceptance buffer. The corrected data will be derived using the verification metadata, and the correction result will be generated. S44: If the risk item is in a state that cannot be reliably determined and needs to be reviewed, then a fast reread or redundant consistency review is triggered first. After the review result is stable, the correction result is generated. Alternatively, if the review fails, the risk item is marked as a persistently unstable anomaly and reported.

[0098] Specifically, in step S41, the exception handling buffer can be understood as a dedicated storage area for handling exception data items. Its design aims to separate complex or time-consuming correction processes from the main return path. This buffer is independent of the main return path's output queue, meaning its operation will not block or delay the return flow of normal data, thus ensuring the smoothness of the main data path. Request context information, such as request identifier, physical address, and original sequence position, is also sent to this buffer to ensure that the correction result can be accurately associated with the original request after subsequent corrections are completed.

[0099] Step S42 addresses risk items with minor anomalies that can be corrected immediately. This typically indicates a minor bit-flip error that can be quickly corrected within the current return window, for example, through simple parity checking or a lightweight ECC (Error Correcting Code) algorithm. In this case, the system invokes a fast error correction unit, usually implemented in hardware, which boasts extremely high processing speed and can complete the correction without significantly increasing return latency, generating the correction result.

[0100] In practical applications, step S43 handles risk items where "serious anomalies require delayed acceptance." Errors in these risk items may be more complex or severe, such as involving multiple bit flips, requiring more complex error correction algorithms, such as Reed-Solomon codes or stronger ECC algorithms. Since this type of correction processing can be time-consuming, standard error correction reconstruction is performed in the anomaly acceptance buffer, using the verification metadata to derive the corrected data and generate the correction result. This approach avoids time-consuming operations within the current return window, thus maintaining the efficiency of the main return link.

[0101] Furthermore, step S44 addresses risk items for which the status requiring review cannot be reliably determined. When the system cannot reliably determine the nature of the error or the result of the correction based on existing information, a further review mechanism is triggered. This may include a fast reread, i.e., rereading the data item from memory to verify its consistency; or a redundant consistency review, such as comparing it with copies stored in other redundant locations. A correction result is generated only after the review result is stable and the data is confirmed to be correctable. If the review fails, the risk item is marked as a persistently unstable anomaly and reported to the upper-level system or administrator for further diagnosis and handling.

[0102] Through the above technical solution, this application can flexibly select the most suitable correction strategy based on the specific type and severity of memory bit flip errors. This not only significantly improves the efficiency of error correction, especially enabling rapid response when handling minor errors, but also enhances the robustness of the system in handling complex or uncertain errors. By separating time-consuming operations from the main return path and providing multi-level correction and verification mechanisms, the negative impact of a single error on the overall system performance is effectively avoided, thereby improving the overall reliability and stability of memory access and ensuring the accuracy of data transmission.

[0103] Further, step S5 includes: S51: Extract the request identifier of the risk item from the request context information, and bind the request identifier to the correction result one by one, so that the correction result can be identified as the formal return data of the original request corresponding to the risk item. S52: Based on the link status information, determine the return semantic type supported by the current return link, and thus determine the injection method of the correction result; S53: If the current return link supports out-of-order return, the correction result is directly inserted into the return path according to the injection priority, and the request identifier is carried so that the upper layer requester can reassemble the data according to the request identifier; S54: If the current return link supports a return semantic type that requires strict order return, then according to the original order position in the request context information, find the position originally occupied by the risk item in the return queue, accurately fill in the position, and after local alignment and waiting for the limited number of requests immediately following the position, inject the correction result. S55: Send the correction result as the official return data for this risk item to the requester, and mark the delivery final status field corresponding to the request identifier as delivered in the return tracking table.

[0104] Specifically, in step S51, the request identifier is a tag used to uniquely identify the memory access request. Its binding with the correction result ensures that the upper-layer requester can accurately associate the corrected data with the original request. This binding mechanism is the foundation for achieving correct data injection.

[0105] In step S52, the link status information can be understood as a set of data describing the current operating mode and capabilities of the memory return path. The return semantic type refers to the order rules followed by the memory controller or system bus when returning data to the requester, mainly including out-of-order returns and strictly ordered returns. By determining this type, the injection strategy most suitable for the current system state can be flexibly selected.

[0106] In practical applications, step S53 addresses scenarios with out-of-order returns by assigning the correction result a back-injection priority and directly inserting it into the return path. Out-of-order returns mean that the system can accept data items arriving out of request order; as long as each data item carries the correct request identifier, the upper-layer system can reassemble it automatically. This approach maximizes the utilization of the return path's bandwidth and reduces waiting time.

[0107] Further, in step S54, the original sequence position refers to the expected position of the risk item in the return queue before it was blocked. Precise padding refers to accurately placing the correction result back to that position to maintain the continuity of the data flow. Local alignment waiting refers to briefly pausing the small number of requests immediately following the correction result after padding, to ensure they are released in the correct order after the correction result, thereby avoiding order confusion.

[0108] Furthermore, in step S55, the "Delivery Finality" field is a marker in the returned tracking table used to indicate whether the data for a specific request has been successfully delivered. Marking it as delivered means that the lifecycle of the request has been completed, and the system can release the associated resources.

[0109] Through the above technical solution, this application can significantly improve the delivery efficiency and reliability of memory module bit flip error correction results. Compared with the basic solution, this application avoids performance loss caused by waiting in out-of-order return scenarios and avoids data order disorder caused by blind injection in strict-order return scenarios by distinguishing and adapting to different return link semantic types.

[0110] In some preferred embodiments, a specific example is given below. Assume the memory controller receives a read request R1, whose corresponding data to be returned is identified as a risk item in step S1 and successfully corrected in step S4, generating a correction result C1. Specifically, in step S51, the request identifier of request R1 is extracted from the request context information, and the correction result C1 is bound to this request identifier. Subsequently, in step S52, the system determines the semantic type of the current return link based on the link status information.

[0111] If the semantic type is out-of-order return, in step S53, the correction result C1 is given a higher back-injection priority and directly inserted into the return path. For example, if other requests R2 and R3 are returning in the current return path, C1 can be inserted directly between or after them without waiting for R2 and R3 to complete, carrying the request identifier of R1. After receiving C1, the upper-layer requester identifies it as the correction data of R1 based on its request identifier and associates it with R1's original request.

[0112] If the semantic type requires maintaining a strict order, in step S54, the system will find the slot that R1 should have occupied in the return queue based on the original order position of R1 in the request context information. For example, if R1 was originally located after R_prev and before R_next in the queue, then the correction result C1 will be precisely padded between R_prev and R_next. To ensure a strict order, the system may briefly perform local alignment waiting for R_next and a small number of requests immediately following it (e.g., R_next+1, R_next+2), and then release these requests along with C1 in order after C1 is padded.

[0113] Finally, in step S55, regardless of the scenario, the correction result C1 is sent to the requester as the official return data of R1, and the delivery final state field corresponding to R1 is marked as delivered in the return tracking table. In this way, the system can flexibly and efficiently complete the delivery of correction data according to the actual return link characteristics.

[0114] Furthermore, the correction processing status includes at least the following states: separated and awaiting acceptance, correction in progress, result awaiting reinjection, reinjection completed, and correction failed; step S6 includes: S61: Iterate through each request in the queue of requests to be returned after the current return window, and determine whether the request has an address association and order dependency with the risk item; S62: When the correction processing status is separated and awaiting acceptance or in the correction process, for requests that are not associated with the risk item and have no sequential dependency, the risk item is allowed to return normally continuously; for requests that are associated with the risk item, the risk item is allowed to continue to return normally, but a temporary risk observation label is attached to the associated physical address at the same time. S63: When the correction processing status is "result pending injection", requests that are not associated with the address of the risk item and have no order dependency are allowed to return normally continuously, and the scheduler prepares to inject the correction result into the return path. S64: When the correction processing status is pending return of results and the system requires strict sequential return, for a preset number of requests in the pending return request queue that are located after the original sequence position of the risk item, a local short-term wait is performed until the correction result is filled in before they are released together; for requests after the preset number, normal return continues. S65: When the correction processing status is "correction failed", mark the risk item as failed and skip it. All requests in the pending return request queue that are after the risk item will not be affected and will continue to return normally.

[0115] Specifically, the correction processing status is a detailed division of different stages in the entire lifecycle of a risk item, from its identification to the completion or failure of correction.

[0116] The "separated and awaiting acceptance" status means that the risk item has been identified and blocked from the main return path, but the actual correction process has not yet started and is waiting to be taken over by the background correction unit.

[0117] The "correction in progress" status means that the correction process for the risk item is in progress, such as rapid error correction or standard error correction reconstruction.

[0118] The pending injection status means that the correction process for the risk item has been completed, the correction result has been generated, and it is waiting to be injected into the return path.

[0119] The "reinjection complete" status means that the correction results have been successfully injected into the return path and have replaced the original data for delivery.

[0120] A correction failure status means that the correction process failed to successfully fix the error, and this risk item is judged as uncorrectable.

[0121] Step S61 aims to identify potential correlations between subsequent requests and risk items in order to perform targeted scheduling. Address correlation means that the physical address accessed by a subsequent request is the same as or overlaps with the physical address of a risk item, which may mean that they are accessing the same memory region. Sequential dependency means that the correct processing of a subsequent request depends on the prior processing or return of a risk item, such as in a strictly sequential return chain.

[0122] In practical applications, in step S62, when a risk item is in a separated, pending acceptance or correction state, since the correction results have not yet been generated or injected, requests not directly related to the risk item can continue to return normally to maximize the throughput of the link. For requests with address association, although they are also allowed to return normally, a temporary risk observation tag will be attached to the associated physical address. The purpose is to continuously monitor the address to prevent it from becoming a hotspot for frequent errors.

[0123] Furthermore, in step S63, when the correction process enters the result-awaiting-injection state, it means that the correction result is about to be available. At this time, for unrelated requests, normal returns can still continue, and the scheduler will prepare to inject the correction result into the return path at an appropriate time.

[0124] In a preferred implementation, in step S64, when the system requires a strictly sequential return, to ensure the correctness of the returned data, the scheduler performs a local, short-term wait for a limited number of requests following the original sequence position of the risk item. This wait is local and time-limited, its purpose being to provide a window for the accurate filling of the correction result, preventing subsequent requests from returning before the correction result, thus disrupting the strict sequence. Once the correction result is filled, these waited requests will be allowed to proceed. Requests further back in time are unaffected by this local wait and continue to return normally.

[0125] Furthermore, in step S65, when the correction process ultimately determines that the correction has failed, in order to prevent system stagnation or error propagation, the risk item will be marked as failed and skipped. At this time, all subsequent requests, regardless of whether they are related to the risk item, will continue to return normally without being affected, ensuring the robustness and continuity of the system in the face of unrecoverable errors.

[0126] This application's solution effectively addresses the issues of low scheduling efficiency and impaired data flow continuity that may exist in the basic solution by finely defining the correction processing status of memory module bit flip errors and implementing differentiated local flow scheduling strategies based on these statuses. Specifically, when a data item to be returned is identified as a risk item and enters the correction process, its correction processing status is updated in real time. By judging the address association and sequence dependency between subsequent requests and risk items in step S61, the scheduler can accurately identify which requests can be processed independently and which requests require special handling.

[0127] For example, when a request is separated and awaiting acceptance or is undergoing correction, since the correction results have not yet been generated, the system allows immediate return for requests unrelated to risk items, avoiding unnecessary waiting and thus maintaining the continuity of the data flow. For requests with address associations, although return is also allowed, by attaching temporary risk observation tags, early warning and monitoring of potential hotspot areas are achieved, improving the system's predictability.

[0128] When the correction process enters the result-awaiting-injection state, the correction results are ready. At this point, unrelated requests can still return normally, while the scheduler actively prepares to inject the correction results into the return path. Especially in scenarios where the system requires strict sequential return, by performing a local short-term wait on a preset number of requests after the original order position of the risk item, the correction results are ensured to be accurately filled in, maintaining the strict order of data return, while avoiding long-term blocking of the entire request queue, achieving a balance between orderliness and efficiency.

[0129] Ultimately, even in the event of a correction failure, this approach can ensure that subsequent requests are returned normally without being affected by marking and skipping the failed items, thereby avoiding the impact of a single point of failure on the entire system's data flow and greatly enhancing the system's robustness and availability.

[0130] Through the above technical solution, this application enables more refined and intelligent scheduling management of the subsequent request queue during the memory bit flip error correction process. Compared with the relatively general scheduling method in the basic solution, this application significantly improves the efficiency and continuity of data return by clearly defining the correction processing status and implementing differentiated scheduling accordingly.

[0131] Please refer to Figure 2 A real-time detection and correction system for memory module bit flip errors, used to implement the steps of any of the above methods, the system comprising: Module 201: Acquires the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; Judgment module 202: Performs real-time risk assessment on the data item to be returned within the current return window and outputs the corresponding delivery qualification status; Blocking module 203: When the delivery eligibility status identifies the data item to be returned as a risk item, it blocks the delivery authorization of the original data corresponding to the risk item within the current return window, and retains its request context information; Correction module 204: In a processing flow independent of the current return window, the verification metadata is used to correct the risk item and generate a correction result; Injection module 205: Based on the request context information, binds the correction result to the request identifier of the risk item, and injects the correction result into the return path according to the scheduling strategy of the return link determined by the link status information, so as to replace the original data to complete the delivery; Scheduling module 206: Based on the correction processing status of the risk item, the queue of requests to be returned after the current return window is partially split and scheduled to maintain the continuous return of requests unrelated to the risk item.

[0132] This system modularizes error detection and correction during memory read access, enabling real-time and efficient handling of bit flip errors. Specifically, the acquisition module 201 collects necessary data and status information to lay the foundation for subsequent processing; the judgment module 202 quickly identifies potential risks within a very short return window to prevent erroneous data from flowing out; the blocking module 203 promptly blocks the delivery of risky data and properly preserves context information; the correction module 204 independently corrects erroneous data without occupying main return link resources; the injection module 205 ensures that the corrected data can be accurately and timely injected back into the return path; finally, the scheduling module 206 ensures the continuity and throughput of the system under high-concurrency scenarios through intelligent traffic routing. This modular design allows the system to flexibly cope with the challenges brought by high-concurrency read access, effectively solving the contradiction between latency and throughput in traditional solutions, and ensuring a balance between data reliability and system performance.

[0133] The core of this application lies in the real-time detection and correction of memory module flip-flop errors, and the optimization of data delivery processes in high-concurrency scenarios. The main technical features of this application will be described in detail below.

[0134] The acquisition module 201 of this application can be a hardware circuit unit, such as the data interface logic in a memory controller, configured to read data and metadata from the memory bus or data cache, and extract a request identifier from the request queue. Alternatively, the acquisition module can also be a software program segment running on the memory controller or a dedicated processor, acquiring the required information by accessing specific registers or memory regions.

[0135] The decision module 202 can be a hardware logic circuit, such as a state machine or a programmable logic array (FPGA), designed to complete data verification and risk assessment in a very short time. This decision module can also be implemented via firmware or software, running on a microcontroller inside the memory controller, executing a preset algorithm to calculate the risk value and output the delivery eligibility status.

[0136] The blocking module 203 can be a data path switch or multiplexer that immediately cuts off the transmission path of the original data to the system bus when the judgment module outputs a risk signal. Simultaneously, the blocking module can include a small buffer or register set for temporarily storing the request context information of the blocked data, such as the request identifier, physical address, and original sequence position.

[0137] The correction module 204 can be a dedicated hardware error correction unit (ECC engine) that operates independently of the main data path, receiving blocked risk items and their verification metadata, and executing error correction algorithms. Alternatively, the correction module can also be a software routine running on an auxiliary processor, asynchronously performing correction calculations on the risk data in the background.

[0138] The injection module 205 can be a data reassembler or a smart cache controller, responsible for re-associating the correction results with the original request identifier and determining the optimal injection timing and method based on link status information (such as link idle status, out-of-order return capability, etc.). For example, the injection module can be an output buffer with a priority queue to ensure that the correction results can be injected back in a timely and orderly manner.

[0139] The scheduling module 206 can be a hardware scheduler responsible for monitoring the correction status of risk items and the dependencies of subsequent request queues. This scheduling module can dynamically adjust the return order of subsequent requests. For example, requests unrelated to risk items are allowed to continue returning normally, while requests with dependencies are subject to partial waiting or reordering to minimize the impact on overall throughput.

[0140] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for real-time detection and correction of memory module bit flip errors, characterized in that, The method for error detection and correction during memory read access includes the following steps: S1: Obtain the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; S2: Perform real-time risk assessment on the data items to be returned within the current return window and output the corresponding delivery eligibility status; S3: When the delivery qualification status identifies the data item to be returned as a risk item, block the delivery authorization of the original data corresponding to the risk item within the current return window, and retain its request context information; S4: In a processing flow independent of the current return window, the risk item is corrected using the verification metadata to generate a correction result; S5: Based on the request context information, bind the correction result with the request identifier of the risk item, and inject the correction result into the return path according to the scheduling policy of the return link determined by the link status information, so as to replace the original data and complete the delivery. The request context information includes at least: the request identifier, the physical address corresponding to the data to be returned, the original sequential position of the risk item in the current return window, the reserved output slot number corresponding to the risk item, and a copy of the original data and the verification metadata; S6: Based on the correction processing status of the risk item, perform partial diversion scheduling on the queue of requests to be returned after the current return window to maintain the continuous return of requests unrelated to the risk item; The step S5 includes: S51: Extract the request identifier of the risk item from the request context information, and bind the request identifier to the correction result in a one-to-one correspondence, so that the correction result can be identified as the formal return data of the original request corresponding to the risk item. S52: Based on the link status information, determine the return semantic type supported by the current return link, thereby determining the injection method of the correction result; S53: If the current return link supports out-of-order return, the correction result is directly inserted into the return path according to the injection priority, and the request identifier is carried so that the upper layer requester can reassemble the data according to the request identifier; S54: If the return semantic type supported by the current return link requires strict order return, then according to the original order position in the request context information, find the position originally occupied by the risk item in the return queue, accurately fill in the position, and after local alignment and waiting for the limited number of requests immediately following the position, inject the correction result. S55: Send the correction result as the official return data of the risk item to the requester, and mark the delivery final state field corresponding to the request identifier as delivered in the return tracking table.

2. The method for real-time detection and correction of memory module bit flip errors according to claim 1, characterized in that, In step S1, the link status information includes at least the occupancy level of the current return window, the number of requests in the subsequent return request queue, the estimated arrival time interval between two adjacent requests in the subsequent return request queue, the busy / idle status of the fast error correction unit that performs error correction within the current return window, and the bit flip error statistics of the physical address corresponding to the data to be returned in historical access.

3. The method for real-time detection and correction of memory module bit flip errors according to claim 2, characterized in that, Step S2 includes: S21: Calculate the window load risk value of the current return window based on the data to be returned, the verification metadata, and the link status information; S22: Compare the window load risk value with multiple preset thresholds, and combine the judgment result of whether the verification metadata is complete and available, to output the corresponding delivery qualification status; The delivery eligibility status includes a status that can be released directly, a status that can be released after minor corrections within the current return window, a status that needs to be delayed for processing in the background, and a status that cannot be reliably determined and requires review.

4. The method for real-time detection and correction of memory module bit flip errors according to claim 3, characterized in that, Step S21 includes: S211: Based on the comparison result between the data to be returned and the verification metadata, determine the severity of bit flipping in the current data, and quantify the flipping score according to the severity. S212: Based on the number of requests in the subsequent request queue and the estimated time interval between two adjacent requests, determine the pressure of subsequent requests on the current return window, and quantify the pressure score according to the pressure level. S213: Based on the bit flip error statistics of the physical address corresponding to the data to be returned in the historical access, determine whether the address belongs to a hot spot area with frequent errors, and quantify the error score according to the error frequency. S214: Based on the busy / idle status of the fast error correction unit, determine whether there are currently any idle error correction resources available for use, and quantify the busy / idle score based on the idle or busy status. S215: The flip score, the pressure score, the error score, and the busy / idle score are weighted and summed to obtain the window load risk value.

5. The method for real-time detection and correction of memory module bit flip errors according to claim 4, characterized in that, Step S22 includes: S221: When the window load risk value is lower than the first threshold value, it is determined that the window can be directly released. S222: When the window load risk value is between the first boundary value and the second boundary value, it is determined that the state can be lightly corrected within the current window. S223: When the window load risk value is higher than the second threshold value but the verification metadata is complete, it is determined that the processing needs to be postponed to the background. S224: When the verification metadata is incomplete or the window load risk value is higher than the third threshold value and the judgment result is unstable, it is determined that the state that cannot be reliably determined and needs to be reviewed is not possible. The preset thresholds include the first boundary value, the second boundary value, and the third boundary value.

6. The method for real-time detection and correction of memory module bit flip errors according to claim 1, characterized in that, When the delivery qualification status is a state that can be released after minor correction within the current return window, a state that needs to be postponed to the background for processing, or a state that cannot be reliably determined to require review, the data item to be returned is the risk item, and step S3 includes: S31: When the delivery qualification status identifies the data item to be returned as a risk item, the original data shall be prohibited from being output from the memory controller to the system bus or the upper layer requester within the current return window; S32: In the return tracking table, switch the request status corresponding to the risk item from normal pending return to abnormal pending acceptance, and retain the request context information of the risk item.

7. The method for real-time detection and correction of memory module bit flip errors according to claim 6, characterized in that, Step S4 includes: S41: The risk item and its request context information are sent to the exception handling buffer. The exception handling buffer is independent of the output queue of the main return link and does not occupy the time resources of the current return window. S42: If the risk item is a minor anomaly that can be corrected at present, then call the fast error correction unit within the current return window to complete the correction and generate the correction result; S43: If the risk item is a serious anomaly that needs to be delayed in the acceptance state, then standard error correction and reconstruction are performed in the anomaly acceptance buffer, and the corrected data is derived using the verification metadata to generate the correction result. S44: If the risk item is a state that cannot be reliably determined and needs to be reviewed, then a fast reread or redundant consistency review is triggered first. After the review result is stable, the correction result is generated. Alternatively, if the review fails, the risk item is marked as a persistently unstable anomaly and reported.

8. The method for real-time detection and correction of memory module bit flip errors according to claim 1, characterized in that, The correction processing status includes at least the separated and awaiting acceptance status, the correction in progress status, the result awaiting reinjection status, the reinjection completed status, and the correction failure status; step S6 includes: S61: Traverse each request in the queue of requests to be returned after the current return window, and determine whether the request has an address association and order dependency with the risk item; S62: When the correction processing status is separated and awaiting acceptance or in the correction process, for requests that are not associated with the risk item and have no sequential dependency, the risk item is allowed to return normally continuously; for requests that are associated with the risk item, the risk item is allowed to continue to return normally, but a temporary risk observation label is attached to the associated physical address at the same time. S63: When the correction processing status is "results pending injection", for requests that have no address association with the risk item and no order dependency, they are allowed to return normally continuously, and the scheduler prepares to inject the correction result into the return path. S64: When the correction processing status is "results pending injection" and the system requires strict sequential return, for a preset number of requests in the pending request queue that are located after the original sequence position of the risk item, a local short-term wait is performed until the correction results are filled in before they are released together; for requests after the preset number, normal return continues. S65: When the correction processing status is "correction failed", the risk item is marked as failed and skipped. All requests in the pending return request queue that are after the risk item are unaffected and continue to return normally.

9. A real-time detection and correction system for memory module bit flip errors, characterized in that, The system, used to implement the method according to any one of claims 1-8, comprises: Acquisition module: Acquires the current data to be returned, as well as the verification metadata, request identifier, and link status information associated with the data to be returned; Judgment module: Performs real-time risk assessment on the data item to be returned within the current return window and outputs the corresponding delivery qualification status; Blocking module: When the delivery eligibility status identifies the data item to be returned as a risk item, the delivery authorization of the original data corresponding to the risk item is blocked within the current return window, while retaining its request context information; Correction module: In a processing flow independent of the current return window, the risk item is corrected using the verification metadata to generate a correction result; Injection module: Based on the request context information, bind the correction result to the request identifier of the risk item, and inject the correction result into the return path according to the scheduling strategy of the return link determined by the link status information, so as to replace the original data to complete the delivery; Scheduling module: Based on the correction processing status of the risk item, the queue of requests to be returned after the current return window is partially split and scheduled to maintain the continuous return of requests unrelated to the risk item.

Citation Information

Patent Citations

  • Log processing method and electronic equipment

    CN122309217A

  • System and method for prioritization of bit error correction attempts

    US20200381076A1