Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

48 results about "Failure domain" patented technology

In computer networking, a failure domain encompasses a section of a network that is negatively affected when a critical device or network service experiences problems. The size of a failure domain and its potential impact depends on the device or service that is malfunctioning. For example, a router potentially experiencing problems would generally create a more significant failure domain than a network switch would.

Cloud computing extension cluster high availability method based on dynamic fault domain and intelligent scheduling

The invention relates to a cloud computing extension cluster high availability method based on a dynamic fault domain and intelligent scheduling, and relates to the field of cloud computing extension cluster high availability. According to the method, hardware health data and network performance indexes of physical nodes are collected in real time, and the fault correlation degree between the nodes is calculated based on the hardware health similarity and a network topology attenuation factor; dynamically generating and updating a logic fault domain topological structure according to the fault correlation degree and a preset threshold value; based on the logic fault domain topological structure and a preset SLA strategy library, executing a virtual machine scheduling decision of multi-objective optimization so as to balance business service quality, fault domain risk and resource cost; and in response to the detected fault event, triggering a corresponding hierarchical migration process according to the service priority and the fault level. According to the method, closed-loop linkage of real-time reconstruction and intelligent scheduling of the dynamic fault domain in the cross-region cloud computing extension cluster can be realized, and the disaster tolerance capability and the resource utilization rate of the system are remarkably improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Distributed memory system based on erasure code and dynamic repair mechanism

The invention relates to the field of distributed storage, and discloses a distributed memory system based on erasure codes and a dynamic repair mechanism. Comprising a data node cluster, a metadata management cluster, a bandwidth token management module, a dynamic repair composer and a repair agent node. And the data node cluster divides the object into strip generation data fragments and verification fragments and dispersedly stores the data fragments and the verification fragments according to a placement rule. And the metadata management cluster maintains object stripe mapping, a fragment position set, erasure code protection configuration and a version epoch, and records a repair intention log. And the bandwidth token management module grants repair execution permission according to the fault domain budget. The dynamic repair composer generates availability repair and durability repair tasks, the temporary check fragments are written in when the number of the available check fragments is lower than a threshold value in the availability stage, and the configuration fragments are recovered and the temporary fragments are recovered in the durability stage. And the consistency of version epochs is verified before the agent node is repaired and written back, and if not, recalculation is carried out.
Owner:EASY POINT GEEK (BEIJING) TECHNOLOGY CO LTD

Distributed system network interface dynamic expansion method and device based on dependency perception

The invention relates to a distributed system network interface dynamic expansion method and device based on dependency perception, and relates to the field of distributed system network interface dynamic expansion. The method comprises the following steps: collecting topological information of each node in a distributed cluster, and calculating the dependency strength between node pairs based on the communication frequency between the nodes, a service call chain and a resource access mode; according to the number of copies and the minimum number of copies of each copy group, constructing a fault domain constraint model, and verifying whether the candidate expansion scheme meets the system availability requirement or not; a heuristic search method is adopted to generate a batch expansion scheduling sequence, and it is ensured that the node dependency relationship of each batch of expansion is weakest and meets constraint conditions; and executing configuration updating operation according to the scheduling sequence, and realizing atomized execution of configuration modification and service restart by adopting a two-stage submission protocol. According to the invention, the zero-shutdown dynamic expansion of the network interface of the distributed system is realized, the service continuity and the configuration consistency are ensured, and the operation and maintenance interruption time and the operation risk are obviously reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Virtual fault domain isolation and recovery method suitable for Lingqu interconnection system

The invention relates to the technical field of chip interconnection, in particular to a virtual fault domain isolation and recovery method suitable for a Lingqu interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; carrying out classification processing on the uncompleted transactions in the transaction shadow table, and monitoring the state of the access object; and after reconnection, task replay is carried out according to the transaction shadow table, and the state machine is switched back to the ACTIVE state. In order to solve the problem that in the prior art, a third-party node with a fault in a Lingqu interconnection system is prone to causing chain reaction, an isolation mechanism is added to an IO interconnection chip used for being connected to a third-party chip, a state machine of a fault domain is switched to a QUIESCE mode, a transaction shadow table of an uncompleted transaction is established, and the transaction shadow table of the uncompleted transaction is obtained. According to the method and the system, the IO interconnection chip performs proxy processing on part of transactions after the access object is recovered, so that faults of other equipment in the interconnection system are prevented from being caused in a fault reconnection stage, and replaying is performed after the access object is recovered, thereby realizing risk isolation and recovery processes.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Intelligently forming data stripes including multiple shards in a single failure domain

Redundant array of independent drives (RAID) sub-stripes are formed across one or more solid-state storage devices of storage nodes of a storage system. The RAID sub-stripes include corresponding shards of data to be stored at the solid-state storage devices, wherein at least one of the RAID sub-stripes has at least two of the corresponding shards of data on a same storage node. At least one global parity shard is generated for the RAID sub-stripes. The at least one global parity shard is to be stored on a first storage node that is different from each of the same storage nodes storing the at least two of the corresponding shards of data.
Owner:PURE STORAGE INC

OSD switching method, system and device in distributed storage pool and medium

The invention discloses an OSD switching method, system and device in a distributed storage pool and a medium, and relates to the technical field of distributed storage, and the method comprises the steps: determining a fault OSD in the distributed storage pool; determining a physical fault domain of a target storage pool where the fault OSD is located; acquiring a global allowable failure domain number of the physical fault domain; determining the number of failed fault domains of the physical fault domain; in response to the fact that the number of the failed fault domains is smaller than the number of the global allowable failure fault domains, marking the state of the fault OSD as down; the global allowable failure fault domain number comprises a minimum value of failure of the fault domain allowed by the target storage pool set, and the target storage pool set comprises a set of all storage pool combinations sharing the physical fault domain. According to the method, the number of the newest failed fault domains is at most equal to the number of the failure of the minimum fault domain, the situation that the number of the newest failed fault domains exceeds the number of the failure of the minimum fault domain is avoided, and the service reliability and the data security of distributed storage in a fault switching scene are improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Automated multi-region application recovery

A recovery orchestrator system receives a recovery plan, which may be used for a hosted-computing environment. The recovery plan includes multiple steps. The recovery orchestrator system receives monitoring metrics. The recovery orchestrator system executes the recovery plan based on the monitoring metrics. The recovery orchestrator system executes steps from the recovery plan in multiple fault domains. The recovery orchestrator system monitors the status of the execution of the recovery plan and provides the status update to a computing device.
Owner:AMAZON TECH INC

Distributing Certificate Bundles According To Fault Domains

Operations of a certificate bundle distribution service may include: detecting a trigger condition to distribute a certificate bundle that includes a set of certificate authority certificates; determining, for each of a plurality of network entities associated with a computer network, a fault domain representing at least one single point of failure; partitioning the plurality of network entities into a plurality of certificate distribution groups, based on a set of partitioning criteria that includes a fault domain of each particular network entity, in which each particular certificate distribution group includes a particular subset of network entities, and the particular subset of network entities are associated with a particular fault domain; selecting a particular certificate distribution group, of the plurality of certificate distribution groups, for distribution of the certificate bundle; and transmitting the certificate bundle to the particular subset of network entities in the particular certificate distribution group.
Owner:ORACLE INT CORP

Adaptive pc-kriging reliability analysis method and system based on active learning

This invention discloses an adaptive PC-Kriging reliability analysis method and system based on active learning. The method includes: using weighted clustering to obtain uniformly distributed first candidate samples from a pre-generated MC sample pool and constructing an initial PC-Kriging model; using interval reduction to select samples and construct a new sample pool, dividing the new sample pool into two sub-sample pools: a safe domain and a failure domain; for each sub-sample pool, weighted clustering is used again to select uniformly distributed second candidate samples; using the second candidate samples distributed in the failure domain and the safe domain, crossing points are constructed; the constructed crossing points are used as new experimental points to iteratively update the PC-Kriging model; convergence is determined and the results are output. The technical solution of this invention focuses on key regions through interval reduction and updates the model based on crossing points, which can more accurately approximate the limit state surface, reduce the number of experimental points, improve prediction accuracy, and is more efficient in utilizing computational resources.
Owner:CHANGZHOU INST OF TECH

Equipment fault processing method, host, computer program product and storage medium

The embodiment of the invention provides an equipment fault processing method, a host, a computer program product and a storage medium. The operating system of the target host can run an error report driving program after receiving an interrupt signal sent by any root interface and used for triggering an error report mechanism; based on an error report driving program, obtaining an address translation exception record from an event queue maintained by a memory management unit; if the address translation exception record points to any virtual device connected to the root interface, calling a kernel mode drive program corresponding to the virtual device through an error report drive program; and controlling the target virtual machine served by the virtual equipment to be paused based on the kernel mode driving program. Therefore, when a single virtual device has a fault, the fault domain can be accurately controlled on the virtual machine served by the faulted virtual device, so that only the virtual machine served by the faulted virtual device is suspended, and the influence on other virtual machines running on the target host can be effectively avoided.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Storage pool creation method and device, equipment, medium and product

ActiveCN121387205AInput/output to record carriersPoolRecursive computation
The invention discloses a storage pool creation method and device, equipment, a medium and a product, which are applied to the technical field of storage, and comprise the following steps: obtaining the number of fault domain units and the number of copies; when the number of the fault domain units is greater than the number of the copies, judging whether each fault domain unit meets a first preset capacity balance condition or not based on the minimum effective capacity and the expected effective capacity corresponding to each fault domain unit; if the first preset capacity balance condition is not met, grouping recursive calculation is carried out on each fault domain unit to obtain a target effective capacity, and whether a second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; and under the condition that a second preset capacity balance condition is not met, creating a main pool based on the target effective capacity and creating an auxiliary pool based on the residual capacity. Therefore, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method for analyzing gradual change reliability and global sensitivity of cabin door lifting mechanism

The invention provides a cabin door lifting mechanism gradual change reliability and global sensitivity analysis method, and relates to the technical field of reliability and sensitivity analysis. Dynamic modeling, gradual change failure domain description, Kriging proxy modeling based on active learning and global sensitivity analysis are organically fused; the performance degradation characteristics and the reliability evolution process of the cabin door lifting mechanism in the whole life cycle can be accurately described in a multi-level and multi-dimension mode. By introducing a gradual change failure membership function, continuous evolution of structural performance from safety to failure is completely reflected, the failure probability and the variation coefficient thereof are quantitatively calculated through Monte Carlo simulation, and clear redundancy margin is reserved for safety design. Meanwhile, by means of an iterative sampling strategy of active learning-Kriging and calculation of a global sensitivity index, the local and global prediction precision and convergence efficiency of the proxy model are remarkably improved, and the contribution degree and interaction effect of each influence factor to the overall reliability are revealed.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Network fault processing method and related device

The invention discloses a network fault processing method, which can reduce the calculation burden and message forwarding overhead of network equipment. In the method, when a link connected with network equipment fails, the network equipment sends a link fault message to neighbor equipment, and the link fault message is only diffused in a certain range taking the network equipment as a center, so that a fault domain with a fixed size is formed. Therefore, the network equipment connected with the fault link only needs to re-calculate the path in the fault domain range, and the re-calculated path is used for replacing the fault link, so that the routing convergence can be completed. By setting the fault domain with the fixed size, each network device with the link fault only needs to maintain the own local fault domain and calculate the corresponding path, so that the problem that the link fault message diffuses in the whole network to cause frequent path calculation of the network devices in the whole network is avoided; and the calculation burden and the message forwarding overhead of the network equipment are effectively reduced.
Owner:HUAWEI TECH CO LTD

Storage based on fault domains and storage classes

Examples include a computing device configured by executable instructions to categorize a plurality of storage components into a plurality of storage fault domains, each storage fault domain corresponding to at least one failure scenario not shared by the other storage fault domains. The computing device may receive a data object for storage. The computing device may store at least a portion of the data object to a first storage fault domain, and may store at least another portion of the data object to a second storage fault domain that is different from the first storage fault domain.
Owner:HITACHI VANTARA LLC

Methods and systems for chaos testing

PendingUS20260079825A1Error detection/correctionFailure domainLoad testing
Provided are systems for automated chaos including a processor and a memory having instructions stored thereon. The instructions, when executed, cause the processor to perform certain operations including connecting to an application infrastructure with one or more applications and inspecting a code of the one or more applications and configuring a chaos experiment. The configuring includes identifying fault domains of the applications. The operations also include enabling pre-execution tasks, including load testing and observability, executing the chaos experiment, and automatically subjecting the applications to features of the chaos experiment. The features may be configured to trigger a fault to occur from the applications. The operations collect information from the applications as a result of executing the chaos experiment and execute an AI / ML routine on the information to output a result. The result is representative of the resilience of the applications.
Owner:JPMORGAN CHASE BANK NA

Spanning tree protocol configuration automatic checking method, device, storage medium and system

The invention provides a spanning tree protocol configuration automatic checking method and device, a storage medium and a system, and the method comprises the steps: receiving and responding to a gateway down-moving operation instruction, and migrating the gateway function configuration of convergence layer equipment to access layer equipment; acquiring equipment information of an access layer, and generating a spanning tree protocol configuration information table according to the equipment information; and determining a difference configuration item according to the spanning tree protocol configuration information table and a preset spanning tree protocol configuration information table to complete verification of spanning tree protocol configuration, the difference configuration item being a configuration item in which the spanning tree protocol configuration information table is not consistent with the preset spanning tree protocol configuration information table. According to the method and the device, the problems that a data center gateway is generally deployed on convergence layer equipment, so that the coverage range of a single two-layer network segment is relatively large, and once a two-layer loop fault occurs, the fault domain diffusion range is wide, so that the operation risk is high are solved.
Owner:AGRICULTURAL BANK OF CHINA

Data consistency guarantee method, system and equipment under active-active architecture and medium

PendingCN121935080AImprove reliabilityAvoid synchronization blocking problemsHardware monitoringFailure domainEmbedded system
The invention relates to the technical field of data consistency guarantee. By providing the data consistency guarantee method, system, device and medium under the active-active architecture, the method comprises the following steps: detecting a data service state to generate a breakpoint event; fault domain positioning processing is carried out based on the breakpoint event, and a fault domain identifier is generated; performing packaging processing on the fault domain operation state to generate a packaging operation state, and performing formatting processing on the packaging operation state according to a standardized breakpoint context structure in the strategy template to generate a persistent breakpoint record; analyzing and processing the persistent breakpoint record through a predefined data service adapter to generate a cross-service compensation operation chain; executing the cross-service compensation operation chain in the isolation environment to generate an operation execution result; and performing consistency verification processing on the operation execution result to generate a consistency verification report so as to achieve the technical effects of improving the system throughput, reducing the fault recovery time and enhancing the data consistency verification reliability.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Isolation evaluation method and device for distributed storage system, equipment and medium

The invention provides an isolation evaluation method and device for a distributed storage system, equipment and a medium. In the technical scheme provided by the invention, a centralized isolation evaluation center is introduced and a cluster steady-state sensing mechanism is combined; isolation evaluation logics originally dispersed on OSDs are collected to an isolation evaluation center elected by a state monitoring assembly to be processed in a unified mode, all isolation requests can be executed after being subjected to serialization evaluation through the isolation evaluation center, and when the isolation evaluation center judges that the distributed storage system is in an unstable state according to the state of a global PG fragment, all the isolation requests can be executed; according to the method, whether the redundancy of each isolated PG still meets the lowest availability requirement of the redundancy strategy adopted by the PG to which the PG fragment belongs is strictly evaluated, the situation that part of PGs cannot normally provide read-write response due to cross-fault-domain and multi-point concurrent isolation caused by multi-node decentralized decision is avoided, and the data security and service continuity of the system are improved.
Owner:XINHUASAN INFORMATION TECH CO LTD

Virtual machine management methods, apparatus, media, devices and computer program products

A virtual machine management method, apparatus, medium, device, and computer program product are disclosed. The method includes: acquiring topology information of multiple virtual machines; determining fault domain information corresponding to each virtual machine based on the topology information; grouping the virtual machines based on the fault domain information to obtain multiple virtual machine groups, wherein virtual machines in the same virtual machine group correspond to the same fault domain information; in response to determining that the primary virtual machine corresponding to a target service is abnormal, determining candidate virtual machines from the multiple virtual machines corresponding to the target service based on the virtual machine group to which the primary virtual machine belongs, wherein the primary virtual machine is one of the multiple virtual machines corresponding to the target service, and the candidate virtual machines and the primary virtual machine belong to different virtual machine groups; and determining the updated primary virtual machine corresponding to the target service from the candidate virtual machines. This can improve the availability of the updated primary virtual machine and ensure the high availability of the target service.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Memory system and system construction method

Maintain adequate availability of storage systems in a cloud environment. [Solution] The system is provided with multiple storage nodes that constitute a first node group spanning multiple failure domains in a cloud environment. For each node, the domain ID of the failure domain in which the node was generated is obtained, and a second node group is formed from the first node group using the necessary number of nodes with domain IDs that do not overlap as much as possible. The number of member nodes in the second node group that exist in the same failure domain is less than or equal to the redundancy level. The redundancy level is the maximum number of member nodes in the second node group that are allowed to stop simultaneously. In the first node group, nodes other than those in the second node group are spare nodes that can be selected as failback destination nodes.
Owner:HITACHI VANTARA LTD

Dynamic storage resiliency

A computer system is configured to provision a plurality of storage volumes at a plurality of fault domains and thinly provision a plurality of cache volumes at the plurality of fault domains. The computer system is also configured to perform a write operation in a resilient manner that maintains a plurality of copies of data associated with the write operation. Performing the write operation in the resilient manner includes allocating a portion of storage in each of the plurality of cache volumes, and caching the data associated with the write operation in the portion of storage in each of the plurality of cache volumes. The cached data is then persistently stored in the plurality of storage volumes. After that, the portion of storage in each of the plurality of cache volumes is deallocated.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

System upgrading method and device, equipment, storage medium and computer program product

The invention discloses a system upgrading method and device, equipment, a storage medium and a computer program product. The method comprises the steps that a first node in a distributed storage cluster determines first information, second information and third information, the first information represents the state of the distributed storage cluster, the second information represents one or more second nodes to be subjected to system upgrading, and the third information represents a fault domain and a storage pool division result of the distributed storage cluster; the first information, the second information and the third information are utilized to determine one or more node groups and an upgrading sequence corresponding to each node group, each node group comprises one or more second nodes, the second nodes contained in different node groups are completely different, and different nodes contained in the same node group belong to the same fault domain or belong to different storage pools; and according to the upgrading sequence of each node group, performing system upgrading on all the second nodes contained in the node group in sequence.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

A storage pool creation method, apparatus, device, medium and product

ActiveCN121387205BInput/output to record carriersPoolRecursive computation
The application discloses a storage pool creation method and device, equipment, medium and product, applied to the storage technical field, including: obtaining the number of fault domain units and the number of copies; when the number of fault domain units is greater than the number of copies, whether each fault domain unit meets the first preset capacity balance condition is judged based on the minimum effective capacity corresponding to each fault domain unit and the expected effective capacity; if the first preset capacity balance condition is not met, the target effective capacity is obtained by grouping and recursively calculating each fault domain unit, and whether the second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; in the case of not meeting the second preset capacity balance condition, the main pool is created based on the target effective capacity, and the auxiliary pool is created based on the remaining capacity. In this way, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Virtual fault domain isolation and recovery method suitable for spirit and qi interconnection system

ActiveCN122019244BThird partyInterconnection
The present application relates to the technical field of chip interconnection, and particularly relates to a virtual fault domain isolation and recovery method suitable for a flexible interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; uncompleted transactions in the transaction shadow table are classified and processed, and the state of an access object is listened to; after reconnection, task replay is performed according to the transaction shadow table, and the state machine is switched back to an ACTIVE state. In view of the problem that a third-party node with a fault in the existing flexible interconnection system is prone to cause a chain reaction, an isolation mechanism is added to an IO interconnection chip used for accessing the third-party chip, the state machine of the fault domain is switched to the QUIESCE mode, and a transaction shadow table of uncompleted transactions is established, so that the IO interconnection chip performs proxy processing on part of the transactions, the fault of other devices in the interconnection system is avoided in the fault reconnection stage, and replay is performed after the access object is recovered, so that the risk isolation and recovery process is realized.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Data recovery method and device of storage system, electronic equipment and storage medium

The invention provides a data recovery method and device of a storage system, electronic equipment and a storage medium, and relates to the technical field of distributed storage. The free storage space is used for generating a locally optimized redundant data block which coexists with an original redundant data block of the data object and can recover data in a smaller fault domain range, and when a storage fault is detected, the corresponding redundant data block is selected and called according to the fault range to execute recovery operation. The problems that in the prior art, due to the fact that the locality of erasure codes is insufficient and the fault domain range of original redundant data block recovery data is large, data needs to be accessed across multiple nodes during fault recovery, the transmission amount is large, the recovery speed is low, and the system load is high can be solved, and the purposes of optimizing the locality of the erasure codes, flexibly adapting to different fault scenes and improving the recovery efficiency are achieved. And the cross-node data transmission overhead is reduced, the fault recovery efficiency is improved, and the overall load of the system is reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Fault tolerance of transaction mirroring

This paper describes systems and methods for promoting fault tolerance in transaction mirroring. The methods described herein include: receiving a commit command for a data transaction from an initiating node of the system, wherein the data transaction is associated with a first fault domain, and wherein the commit command is directed to a primary participant node and a secondary participant node of the system; in response to the receipt, determining whether a response to the commit command has been received at the primary participant node from the secondary participant node; and in response to determining that no response to the commit command has been received at the primary participant node, indicating that the secondary participant node is invalid in a data store associated with a second fault domain different from the first fault domain.
Owner:EMC IP HLDG CO LLC

Inter-satellite reliable routing method based on fault domain model

The application claims a kind of inter-satellite reliable routing method based on fault domain model, belong to satellite communication technical field.Aiming at the problems that inter-satellite link is influenced by space environment and human interference, leading to transmission failure and routing interruption, a centralized inter-satellite reliable routing method based on inter-satellite link attribute is proposed.According to the real-time state and failure reason of inter-satellite link, the inter-satellite link is classified, the initial topology model containing global network information is constructed, the boundary diffusion method is designed combined with the state of boundary node of fault link, the fault link is aggregated into block to form fault domain topology model to limit the influence range of fault.Analysis of the influence of buffer queue capacity, link signal-to-noise ratio and path length on routing, construct comprehensive link utility function to quantify link quality, propose multi-attribute bypass path selection method to solve the problem of load imbalance caused by single path selection criterion, reduce system packet loss rate and end-to-end delay, improve the fault response capability of satellite network.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Yield evaluation method for high-dimensional multi-failure domain of SRAM (Static Random Access Memory) and electronic equipment

The invention relates to the technical field of static random access memories, and provides a yield evaluation method and electronic equipment for high-dimensional multi-failure domains of an SRAM (Static Random Access Memory), and the method comprises the following steps: obtaining failure index parameters of the SRAM, which at least comprise a read access failure parameter, a read destruction failure parameter, a write failure parameter and a hold state failure parameter; performing Monte Carlo processing on the failure index parameters twice to respectively obtain sample point data and evaluation point data; training a preset initial model based on the sample point data to obtain a multi-classification logistic regression model, training adopting a high-dimensional space and a multi-failure domain, the high-dimensional space being determined by the type of the failure index parameter and the sample point data, and the failure domain being determined by dividing high-dimensional space data points based on the maximum radius and the minimum density of a region by a spatial clustering algorithm; and inputting evaluation point data into the model, and determining a yield evaluation result. The problem that high efficiency and high accuracy cannot be considered at the same time when yield evaluation is carried out on the SRAM in the prior art is solved.
Owner:BEIJING KUANWEN MICROELECTRONICS TECH CO LTD

Processor exception recovery method and apparatus, computing device, storage medium, and product

PendingCN122470422AObject basedGranularity
The present disclosure relates to a processor exception recovery method and device, a computing device, a storage medium and a product. The processor exception recovery method comprises: in response to an exception event of a processor, determining at least one exception object related to the exception event and obtaining state data of the at least one exception object; determining an association relationship between the at least one exception object based on the state data to construct an object relationship view of the at least one exception object; obtaining forward advancing state data of the processor; determining a fault domain of the processor based on the object relationship view and the forward advancing state data; determining a recovery granularity corresponding to the fault domain; and performing at least one exception recovery operation of the processor based on the recovery granularity, thereby identifying a real impact range of the exception and determining the exception recovery operation by using the association relationship view constructed around the exception object, improving the accuracy of the fault domain and the exception recovery operation, and further improving the effect of the exception recovery.
Owner:UNIONTECH SOFTWARE TECH CO LTD

Intelligently forming data stripes including multiple shards in a single failure domain

Redundant array of independent drives (RAID) sub-stripes are formed across one or more solid-state storage devices of storage nodes of a storage system. The RAID sub-stripes include corresponding shards of data to be stored at the solid-state storage devices, wherein at least one of the RAID sub-stripes has at least two of the corresponding shards of data on a same storage node. At least one global parity shard is generated for the RAID sub-stripes. The at least one global parity shard is to be stored on a first storage node that is different from each of the same storage nodes storing the at least two of the corresponding shards of data.
Owner:PURE STORAGE INC