Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

33 results about "Memory faults" patented technology

Memory faults. A memory fault occurs when a process either uses memory in an incorrect way or uses memory that does not belong to it, according to the OS . On a Linux system, a memory fault is called a segmentation fault. When a memory fault occurs, the OS terminates the process immediately .

Server memory management system and cluster system

ActiveCN120560897ANon-redundant fault processingMemory faultsTerm memory
The invention discloses a server memory management system and a cluster system, and relates to the technical field of server memory management, the system identifies memory fault information of a server based on a fault reminding signal under the condition that a fault signal is received, and the current load state is determined based on operation data of the server. And matching the corresponding target repair strategy to perform fault repair on the server memory based on the target repair strategy. Therefore, the system can dynamically match different fault recovery strategies according to different fault scenes, so that the technical problem of poor flexibility of a memory fault recovery method in related technologies is solved, and the technical effects of improving the flexibility of memory fault recovery and the service operation stability of a server are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

CXL memory fault tolerance method, and server system, storage medium and electronic device

PCT designated stageWO2025227987A1TransmissionRedundant hardware error correctionMemory faultsMemory footprint
A CXL memory fault tolerance method, and a server system, a storage medium and an electronic device. The method comprises: acquiring parameter values of a group of operating parameters of CXL memory devices in a CXL memory device group, wherein the group of operating parameters are used for representing the operating states of the corresponding CXL memory devices; on the basis of the acquired parameter values of the group of operating parameters, predicting the operating states of the CXL memory devices in the CXL memory device group; and when it is predicted that there is an abnormal memory device operating abnormally in the CXL memory device group, performing controlling to execute a migration operation on memory data in the abnormal memory device, so as to migrate the memory data in the abnormal memory device to a target memory device operating normally in the CXL memory device group. By means of the present application, the problem of CXL memory fault tolerance methods in the prior art of the memory utilization rate of a server being low due to a hot standby memory occupying a server slot is solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

CXL memory fault processing method, computing device and computing system

The embodiment of the invention discloses a CXL memory fault processing method, computing equipment and a computing system, relates to the field of computing, and at least can realize memory fault synchronization between the computing equipment so as to ensure stable operation of the computing equipment. The CXL memory fault processing method comprises the steps of determining second computing equipment from a computing equipment cluster under the condition that access of first computing equipment to a first memory space is abnormal; wherein the first computing device is any computing device in a computing device cluster; the second computing device is a computing device sharing a first memory space with the first computing device in the computing device cluster; the first memory space is a memory space in the CXL memory device; sending the first message to the second computing device; the first message is used for indicating that the first memory space access is abnormal.
Owner:XFUSION DIGITAL TECH CO LTD

Memory diagnosis method and device, equipment and medium

The invention discloses a memory diagnosis method and device, equipment and a medium, which are applied to the technical field of servers, and the method comprises the following steps: acquiring a memory data sequence collected according to a preset time interval within a recent preset duration; memory data features are extracted based on the memory data sequence, and the memory data features comprise change trend features; inputting the memory data features into a target memory state prediction model to obtain a memory state prediction result of a future preset duration corresponding to the memory data sequence; wherein the target memory state prediction model is obtained by training a memory data sequence feature sample and a memory state label, and the memory state label represents a memory state of a future preset duration corresponding to the memory data sequence feature sample. Therefore, the memory fault can be predicted in real time, so that problems can be positioned and solved, and the server performance is improved.
Owner:INSPUR (SHANDONG) COMPUTER TECH CO LTD

Fault processing method and computing device

This application provides a fault processing method and a computing device. The method is applied to the computing device, and the computing device includes an out-of-band management module and processor firmware. The method includes: The out-of-band management module obtains memory information of a memory; the out-of-band management module determines fault information of the memory based on the memory information, where the fault information includes a fault location, and location precision included in the fault location is located at least at a row address of a fault; the out-of-band management module sends the fault information to the processor firmware; and the processor firmware isolates a memory at the fault location based on the fault information. In the foregoing method, a location that is in the memory and at which the fault occurs can be accurately determined, so that processing precision of the memory fault is improved.
Owner:XFUSION DIGITAL TECH CO LTD

A method, apparatus and storage medium for collecting memory fault information

ActiveCN115904773BFault responseMemory faultsTerm memory
This application provides a method, apparatus, and storage medium for collecting memory fault information, relating to the field of data storage, and capable of preventing missed fault information reports. The method is applied to a server, which includes memory, an out-of-band controller, and an in-band controller. The method includes: during server operation, the in-band controller performs fault detection on the memory and sends the detected fault information to the out-of-band controller; during server restart, if the server's startup mode is a cold restart, the in-band controller determines whether new fault information exists in the memory, wherein the new fault information is fault information not stored in the out-of-band controller; if new fault information exists in the memory, the in-band controller sends the new fault information to the out-of-band controller.
Owner:XFUSION DIGITAL TECH CO LTD

Fault memory locating method, system, device, computer device and storage medium

ActiveCN117234771BAffect bootingSolve the problem of being unable to locate faulty PMIC memoryWrite protectionMemory faults
The application relates to the technical field of servers and discloses a fault memory positioning method, system and device, computer equipment and a storage medium, the method comprising the following steps: in the case that a memory fault signal is received and a server is in a shutdown state, the write protection state of a power management chip register is released; candidate fault memories are determined according to the memory fault signal; a disable command is written into the power management chip register of the candidate fault memories, wherein the disable command is used to not power on the candidate fault memories when the server enters a startup state; in the case that the server enters the startup state, log information of the candidate fault memories is acquired; and according to the log information, a target fault memory and slot position information of the target fault memory are determined from the candidate fault memories. The application solves the problems that it is difficult to acquire abnormal information capable of accurately positioning the fault memories and the PMIC fault memories cannot be positioned.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

A fault injection method for a QEMU-based all-digital digital prototype

ActiveCN119883912BError detection/correctionFault toleranceMemory faults
The present application belongs to the technical field of embedded software virtualization testing, and particularly relates to a fault injection method for a QEMU-based all-digital digital prototype. The present application divides faults into three layers, i.e. an atomic fault injection layer, a composite fault injection layer and a fault injection simulation test layer, provides four types of atomic fault injection, i.e. register fault, memory fault, bus protocol fault and bus load fault, and provides composite fault injection and fault case set injection. The present application takes the QEMU-based all-digital simulation digital prototype as an injection target, improves the reliability of fault injection results, and realizes fault injection at different levels from the physical layer to the application layer through the provision of three levels and four dimensions of fault injection modes and rich fault injection methods, thereby improving the fault tolerance and reliability of embedded software and improving the testing efficiency.
Owner:XIAN AVIATION COMPUTING TECH RES INST OF AVIATION IND CORP OF CHINA

Memory fault tolerance using raid

Systems, methods, and computer-readable storage devices can include fabric networks for isolating and correcting failures in the fabric network and device failures. The fabric network connects to a group of devices and a host. The group of devices includes at least a target data device, other data devices, and a parity device. A redundant array of independent devices (RAID) engine, which is coupled to the one group of devices, performs an access operation. The fault tolerant engine is provided in a leaf switch of the fabric network. A routing processor determines a path for a request received from the host to the target data device. The routing processor is coupled to the fault tolerant engine, and the routing processor is provided in the leaf switch.
Owner:MICRON TECHNOLOGY INC

Npu processor architecture-based integrated circuit memory fault management method and system

The application belongs to the technical field of integrated circuit design and manufacturing, and particularly relates to an integrated circuit memory fault management method and system based on an NPU processor architecture. First, multi-dimensional detection is carried out specially for the memory size, storage area and handshake protocol effectiveness of a physical address, which changes the status of passive transmission of abnormal signals in the traditional scheme. Then, according to the fault detection result, a control signal is reversely output to intervene in the mapping logic, so as to effectively avoid fault diffusion caused by continuous transmission of abnormal addresses, and when an address abnormality is detected, subsequent physical addresses can be forcibly relocated to a preset safe address area, reliable safety backup is provided, and the running stability is improved.
Owner:SHANDONG UNIV

Computer operation fault rapid identification processing method

The invention discloses a computer operation fault rapid identification processing method, and relates to the technical field of data analysis, and the method comprises the following steps: S100, setting a first test parameter according to a mapping relation between a driving square wave and a write-in signal, and testing a memory based on an improved SDRAM test algorithm and the first test parameter to obtain a test result; a cross cubic network of memory addresses is constructed, a numerical matrix is constructed according to the cross cubic network, first calculation is performed based on the numerical matrix, and a fault address and a fault type are output. According to the method, accurate coordinates are provided for subsequent redundancy repair through automatic classification, the number of test steps is reduced by driving the mapping relation between square waves and write-in signals and exciting time sequence and amplitude coupling faults at a time, three-dimensional addresses are subjected to redundancy through the cross cubic network, and the storage fault recognition speed and accuracy are remarkably improved.
Owner:NANTONG INST OF TECH

Memory fault detection method and apparatus based on target fault type

ActiveCN119314542BStatic storageMemory faultsReliability engineering
The application provides a memory fault detection method and device based on a target fault type, which comprises the following steps: determining a fault type to be detected existing in a memory, and constructing a fault detection sequence library of the memory; determining a corresponding target detection sequence from the fault detection sequence library for any target fault type in the fault type to be detected; performing sequence recombination based on a write operation type for each target detection sequence to obtain an overall write operation detection sequence, inserting a read operation sequence element to obtain a fault detection sequence to be optimized, and finally performing screening processing based on the read operation sequence element to obtain a final fault detection sequence for detecting any target fault type in the fault type to be detected. Through the application, the defect that the test coverage rate for a specific fault type cannot reach 100% due to the fact that the memory fault detection method in the prior art is difficult to detect all fault types of the memory because of the numerous fault types of the memory is overcome.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Memory fault detection method and device, chip, medium and program product

PendingCN120636502AStatic storageDouble data rateMemory faults
The invention provides a memory fault detection method, which comprises the following steps of: generating expected data according to a preset mode; wherein the preset mode represents the test dimension of the memory, and the memory comprises a double-rate synchronous dynamic random access memory DDR subsystem and a dynamic random access memory DRAM subsystem; the expected data is written into the DRAM subsystem through the DDR subsystem; reading the read-back data from the DRAM subsystem; and under the condition that the read-back data is not matched with the expected data, determining that the memory has a fault. The invention further provides a fault detection device of the memory, a chip, a computer readable medium and a computer program product.
Owner:SANECHIPS TECH CO LTD

Local and global redundancy optimization repair method and system based on data mining

PendingCN121709004AStatic storageCluster algorithmMemory faults
The invention relates to the technical field of memory fault repair, and particularly discloses a local and global redundancy optimization repair method and system based on data mining, and the method comprises the steps: collecting historical fault data of a memory, cleaning the historical fault data, and extracting fault bit features; carrying out clustering analysis on the extracted fault bit features by adopting a clustering algorithm, and dividing a high-frequency fault region of the fault bit; constructing a distance model of the fault bit and the redundant unit, wherein the distance model fuses the physical distance and the logic distance of the fault bit and the redundant unit; when a new fault bit is detected, preferentially allocating a local redundancy unit with a closer logic distance and / or physical distance based on the distance model for repairing, and updating a resource occupation state of a local redundancy pool; and if the local redundant unit of the area where the new fault bit is located is exhausted, selecting an optimal resource from the global redundant unit according to the distance weight based on the distance model to repair.
Owner:BEIJING YUEXIN TECH CO LTD

Built-in computing macro-cell-oriented built-in self-test architecture, system and method

The invention relates to a built-in self-test method, architecture and system oriented to an in-memory computing macro-cell, and the method comprises the steps: executing a first test mode: inserting a computing activation vector and a computing enabling instruction in an MBIST test algorithm, and carrying out a computing test after a read operation to detect a coupling fault between a storage array and bit multiplication logic; and judging whether the storage reading result and the calculation result are correct or not, if so, entering a second test mode or judging that the self-test is correct and completing the detection, and otherwise, diagnosing the fault position based on the detection result. According to the test method provided by the invention, the calculation enabling and the calculation activation vector are inserted in the traditional MBIST algorithm, and the additionally executed multiply-accumulate calculation operation is verified after the read operation, so that the test efficiency is greatly improved on the premise of covering the fault type of the traditional memory. The specific fault of data coupling between the storage unit and the calculation unit can be further effectively detected, and the fault coverage rate is increased.
Owner:HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD

Memory fault processing method and apparatus, electronic device, and readable storage medium

This application provides a memory fault processing method: a faulty first storage cell is determined, the first storage cell is isolated, a pressure test is performed on the first storage cell to obtain a fault level of the first storage cell, and a corresponding operation is performed on the first storage cell based on the fault level of the first storage cell, where when the fault level of the first storage cell indicates that the first storage cell is at a first risk level, the operation performed on the first storage cell includes de-isolating the first storage cell. After a faulty storage cell is isolated, a pressure test may be performed on the storage cell to obtain a real fault level of the storage cell. When the fault level of the first storage cell is the first risk level, the first storage cell may continue to be used.
Owner:HUAWEI TECH CO LTD

Memory fault processing method, device, controller, system, medium and product

PendingCN120994448AFault responseMemory faultsControl store
The invention discloses a memory fault processing method and device, a controller, a system, a medium and a product, and relates to the field of fault processing, when a target memory has an IO overtime fault, the memory controller controls to close an IO retry function of the target memory; according to the method and the device, the first target requests of the target memory are received, the first target requests of the target memory are stopped, the first target requests of the current target memory are directly stopped, the IO requests which are not processed in the current target memory can be rapidly processed, and then an error processing thread can be called to carry out fault processing. By closing the IO retry function of the target memory, the time consumed by IO repeated retry is avoided, the IO request which is not processed completely is compressed for timeout fault detection, and the time consumed by retry is avoided, so that an error processing thread can be quickly awakened, the fault processing duration is greatly shortened, and the influence caused by overlong-time request blockage is avoided.
Owner:SANGFOR TECH INC

Runtime sparing for uncorrectable errors based on fault-aware analysis

A system can respond to detection or prediction of an uncorrectable error (UE) in memory based on fault-aware analysis. The fault-aware analysis enables the system to generate a determination of a specific hardware element of the memory that is faulty. In response to detection of an error, the system can correlate a hardware configuration of the memory device with historical data indicating memory faults for hardware elements of the hardware configuration. Based on a determination of the specific component that likely caused the UE, the system can identify a region of memory associated with the detected UE and mirror the faulty region to a reserved memory space of the memory device for access to data of the faulty region.
Owner:INTEL CORP

Apparatus and methods for memory fault detection within die architectures

ActiveUS12493514B2Redundant data error correctionMemory faultsData error
Methods and apparatuses directed to memory fault detection mechanisms within die architectures. In some examples, a die package includes decoder logic that receives multiple data words, and a first error correcting code for each of the data words. The decoder logic generates, for each of the data words, a second error correcting code based on a corresponding one of the data words. Further, the decoder logic generates, for each of the data words, an error status based on the first error correcting code and the second error correcting code that corresponds to each of the data words. The die package also includes error generation logic that receives the error status for the data words from the decoder logic, and generates error data based on a combination of the error statuses for the data words. The error generation logic can store the error data in a memory device.
Owner:QUALCOMM INC

File system, operating system, and computing device

A file system, comprising: a plurality of first file systems, wherein a plurality of first file systems are mounted in different partitions of a same storage medium or are mounted on different storage media; and a second file system, mounted on the plurality of first file systems and used for carrying out fault-tolerant verification on consistency data in memory metadata of the plurality of first file systems across the plurality of first file systems. In this way, one second file system is overlaid on the original first file systems, so that fault-tolerant verification can be carried out on the consistency data in the memory metadata of the first file systems across the first file systems, without requiring redundant backup of these consistency data in the first file systems, thereby reducing the memory overhead, and handling a memory fault with low memory overhead.
Owner:HUAWEI TECH CO LTD

Fault handling method, system, device and program product for distributed memory pool

The present disclosure provides a distributed memory pool fault processing method, system, device and program product, and relates to the technical field of computers. The distributed memory pool fault processing method comprises: in response to an uncorrectable memory fault triggered by a memory operation instruction, determining a physical address that has failed; based on the physical address that has failed and the correspondence relationship between the memory page and the physical address range, determining a first memory page containing the physical address; finding the copy position of the first memory page in the fault processing program table, and determining the corresponding second memory page in the copy position; in the case that the fault processing strategy is remapping, remapping the virtual address accessed by the memory operation instruction from the first memory page to the second memory page; and based on the memory operation instruction, accessing the stored data through the second memory page in the same execution thread.
Owner:XIAMEN UNIV

An adaptive memory testing method, system, device, product, and medium

PendingCN122637858AMemory faultsAlgorithm
The present application relates to the field of storage test, and provides a kind of adaptive memory test method, system, equipment, product and medium, including according to sampling ratio sampling, obtain edge phase;Obtain two-dimensional probability eye diagram, construct full convolution inference network, the resolution of two-dimensional probability eye diagram is promoted by full convolution inference network, obtain fine two-dimensional probability eye diagram;The error of fine two-dimensional probability eye diagram is zero probability and offset standard deviation, determine eye diagram adjustment strategy, to adjust fine two-dimensional probability eye diagram, obtain adjustment two-dimensional probability eye diagram;The timing margin and eye diagram offset of adjustment two-dimensional probability eye diagram are calculated, and the compensation strategy is determined according to timing margin and eye diagram offset, and the phase compensation amount is obtained according to the compensation strategy;According to the selection of compensation strategy, determine the area failure probability, generate fault area mark by area failure probability, generate memory fault report according to fault area mark and phase compensation amount, complete memory test.
Owner:CHINA STATE SHIPBUILDING CORP NO 707 RES INST

Memory fault diagnosis device using bloom filter, soc, bloom filter update method, and memory accessibility determination method

A memory fault diagnosis device, a System on a Chip (SoC), and a Bloom filter update method and a memory accessibility determination method can operate to diagnose memory faults in a minimum time by using a Bloom filter to detect failed memory areas and thereby determine memory accessibility in a minimum time.
Owner:SK HYNIX INC +1

Ecc fault-tolerant structure covering full communication path and memory fault-tolerant design method

This invention discloses an ECC fault-tolerant structure and memory fault-tolerant design method covering the entire communication path, relating to the field of on-chip memory technology. It addresses the technical problem that existing on-chip memory EDAC technology's protection scope is limited to the on-chip memory cell itself, and introduces additional errors during data transmission, making it difficult to guarantee data integrity. The ECC fault-tolerant structure covering the entire communication path includes a host module, an interconnect bus, and storage components. The host module is configured with an ECC codec. The host module is used to implement access control of the storage components. The interconnect bus is used to implement access routing between the host module and the storage components and establish a connection between them. The host module initiates a data write access, encodes the pre-written data using ECC, and then initiates a write access to the storage component. The interconnect bus routes the access to the corresponding storage component based on the access address.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Memory error processing method and apparatus, terminal, and storage medium

The embodiment of the application discloses a memory error processing method, device, terminal and storage medium, and belongs to the memory field, and comprises the following steps: when a memory fault occurs, it is judged whether a read-write action occurs in a CPU execution context; when the CPU execution context is executed, an interrupt context is acquired, and it is judged whether the memory fault is in a write stage by using the interrupt context; when the write stage is executed, a page table entry is acquired, and it is judged whether it is a file mapping according to the page table entry; when the file mapping is executed, a physical page frame corresponding to the memory fault is determined, all page table entries related to the physical page frame are determined through reverse mapping, a physical page frame entry in the page table entry is atomically deleted, and a physical memory page corresponding to the virtual address is deleted from a file page cache; a free physical memory page is allocated, added to the file page cache, and initialized for a read operation; and the virtual memory and the free physical memory page of the file page cache are associated through the page table entry.
Owner:KYLIN CORP

Memory fault information collection method and device and storage medium

PendingCN121349744AFault responseMemory faultsTerm memory
The embodiment of the invention provides a memory fault information collection method and device and a storage medium, relates to the field of data storage, and can avoid fault information report omission. The method is applied to the server, the server comprises a memory, an out-of-band controller and an in-band controller, and the method comprises the steps that in the running process of the server, the in-band controller conducts fault detection on the memory and sends detected fault information to the out-of-band controller; in the restarting process of the server, if the starting mode of the server is cold restarting, the in-band controller determines whether newly-added fault information is stored in a memory or not, and the newly-added fault information is fault information which is not stored in the out-of-band controller; and if the newly added fault information is stored in the memory, the in-band controller sends the newly added fault information to the out-of-band controller.
Owner:XFUSION DIGITAL TECH CO LTD

Memory error processing method and device, terminal and storage medium

The embodiment of the invention discloses a memory error processing method and device, a terminal and a storage medium, and belongs to the field of memorys.The memory error processing method comprises the steps that when a memory breaks down, whether a read-write action occurs when a CPU executes a context or not is judged, when the context is executed for the CPU, an interrupt context is obtained, and whether the memory breaks down or not is in a writing stage or not is judged through the interrupt context; in the writing stage, obtaining a page table item, and judging whether file mapping is performed or not according to the page table item; when file mapping is carried out, determining a physical page frame corresponding to a memory fault, determining all page table entries related to the physical page frame through reverse mapping, carrying out atomic deletion on the physical page frame entries in the page table entries, and deleting a physical memory page corresponding to the virtual address from a file page cache; distributing an idle physical memory page, adding a file page cache, and initializing a read operation; and associating the virtual memory with an idle physical memory page of the file page cache through a page table item.
Owner:KYLIN CORP

Fault processing method, system and device for distributed memory pool and program product

The invention provides a fault processing method, system and device for a distributed memory pool and a program product, and relates to the technical field of computers. The fault processing method for the distributed memory pool comprises the following steps: in response to an uncorrectable memory fault triggered by a memory operation instruction, determining a physical address with the fault; determining a first memory page containing the physical address based on the physical address with the fault and the corresponding relation between the memory page and the physical address range; searching a copy position of the first memory page in the fault processing program table, and determining a corresponding second memory page in the copy position; under the condition that the fault processing strategy is remapping, the virtual address accessed by the memory operation instruction is remapped to the second memory page from the first memory page; and accessing the stored data through the second memory page based on the memory operation instruction in the same execution thread.
Owner:XIAMEN UNIV