Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Failure domain" patented technology

In computer networking, a failure domain encompasses a section of a network that is negatively affected when a critical device or network service experiences problems. The size of a failure domain and its potential impact depends on the device or service that is malfunctioning. For example, a router potentially experiencing problems would generally create a more significant failure domain than a network switch would.

Distributed memory system based on erasure code and dynamic repair mechanism

The invention relates to the field of distributed storage, and discloses a distributed memory system based on erasure codes and a dynamic repair mechanism. Comprising a data node cluster, a metadata management cluster, a bandwidth token management module, a dynamic repair composer and a repair agent node. And the data node cluster divides the object into strip generation data fragments and verification fragments and dispersedly stores the data fragments and the verification fragments according to a placement rule. And the metadata management cluster maintains object stripe mapping, a fragment position set, erasure code protection configuration and a version epoch, and records a repair intention log. And the bandwidth token management module grants repair execution permission according to the fault domain budget. The dynamic repair composer generates availability repair and durability repair tasks, the temporary check fragments are written in when the number of the available check fragments is lower than a threshold value in the availability stage, and the configuration fragments are recovered and the temporary fragments are recovered in the durability stage. And the consistency of version epochs is verified before the agent node is repaired and written back, and if not, recalculation is carried out.
Owner:EASY POINT GEEK (BEIJING) TECHNOLOGY CO LTD

Virtual fault domain isolation and recovery method suitable for Lingqu interconnection system

The invention relates to the technical field of chip interconnection, in particular to a virtual fault domain isolation and recovery method suitable for a Lingqu interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; carrying out classification processing on the uncompleted transactions in the transaction shadow table, and monitoring the state of the access object; and after reconnection, task replay is carried out according to the transaction shadow table, and the state machine is switched back to the ACTIVE state. In order to solve the problem that in the prior art, a third-party node with a fault in a Lingqu interconnection system is prone to causing chain reaction, an isolation mechanism is added to an IO interconnection chip used for being connected to a third-party chip, a state machine of a fault domain is switched to a QUIESCE mode, a transaction shadow table of an uncompleted transaction is established, and the transaction shadow table of the uncompleted transaction is obtained. According to the method and the system, the IO interconnection chip performs proxy processing on part of transactions after the access object is recovered, so that faults of other equipment in the interconnection system are prevented from being caused in a fault reconnection stage, and replaying is performed after the access object is recovered, thereby realizing risk isolation and recovery processes.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Intelligently forming data stripes including multiple shards in a single failure domain

Redundant array of independent drives (RAID) sub-stripes are formed across one or more solid-state storage devices of storage nodes of a storage system. The RAID sub-stripes include corresponding shards of data to be stored at the solid-state storage devices, wherein at least one of the RAID sub-stripes has at least two of the corresponding shards of data on a same storage node. At least one global parity shard is generated for the RAID sub-stripes. The at least one global parity shard is to be stored on a first storage node that is different from each of the same storage nodes storing the at least two of the corresponding shards of data.
Owner:PURE STORAGE INC

Adaptive pc-kriging reliability analysis method and system based on active learning

This invention discloses an adaptive PC-Kriging reliability analysis method and system based on active learning. The method includes: using weighted clustering to obtain uniformly distributed first candidate samples from a pre-generated MC sample pool and constructing an initial PC-Kriging model; using interval reduction to select samples and construct a new sample pool, dividing the new sample pool into two sub-sample pools: a safe domain and a failure domain; for each sub-sample pool, weighted clustering is used again to select uniformly distributed second candidate samples; using the second candidate samples distributed in the failure domain and the safe domain, crossing points are constructed; the constructed crossing points are used as new experimental points to iteratively update the PC-Kriging model; convergence is determined and the results are output. The technical solution of this invention focuses on key regions through interval reduction and updates the model based on crossing points, which can more accurately approximate the limit state surface, reduce the number of experimental points, improve prediction accuracy, and is more efficient in utilizing computational resources.
Owner:CHANGZHOU INST OF TECH

Equipment fault processing method, host, computer program product and storage medium

The embodiment of the invention provides an equipment fault processing method, a host, a computer program product and a storage medium. The operating system of the target host can run an error report driving program after receiving an interrupt signal sent by any root interface and used for triggering an error report mechanism; based on an error report driving program, obtaining an address translation exception record from an event queue maintained by a memory management unit; if the address translation exception record points to any virtual device connected to the root interface, calling a kernel mode drive program corresponding to the virtual device through an error report drive program; and controlling the target virtual machine served by the virtual equipment to be paused based on the kernel mode driving program. Therefore, when a single virtual device has a fault, the fault domain can be accurately controlled on the virtual machine served by the faulted virtual device, so that only the virtual machine served by the faulted virtual device is suspended, and the influence on other virtual machines running on the target host can be effectively avoided.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Storage pool creation method and device, equipment, medium and product

ActiveCN121387205AInput/output to record carriersPoolRecursive computation
The invention discloses a storage pool creation method and device, equipment, a medium and a product, which are applied to the technical field of storage, and comprise the following steps: obtaining the number of fault domain units and the number of copies; when the number of the fault domain units is greater than the number of the copies, judging whether each fault domain unit meets a first preset capacity balance condition or not based on the minimum effective capacity and the expected effective capacity corresponding to each fault domain unit; if the first preset capacity balance condition is not met, grouping recursive calculation is carried out on each fault domain unit to obtain a target effective capacity, and whether a second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; and under the condition that a second preset capacity balance condition is not met, creating a main pool based on the target effective capacity and creating an auxiliary pool based on the residual capacity. Therefore, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Storage based on fault domains and storage classes

Examples include a computing device configured by executable instructions to categorize a plurality of storage components into a plurality of storage fault domains, each storage fault domain corresponding to at least one failure scenario not shared by the other storage fault domains. The computing device may receive a data object for storage. The computing device may store at least a portion of the data object to a first storage fault domain, and may store at least another portion of the data object to a second storage fault domain that is different from the first storage fault domain.
Owner:HITACHI VANTARA LLC

Methods and systems for chaos testing

PendingUS20260079825A1Error detection/correctionFailure domainLoad testing
Provided are systems for automated chaos including a processor and a memory having instructions stored thereon. The instructions, when executed, cause the processor to perform certain operations including connecting to an application infrastructure with one or more applications and inspecting a code of the one or more applications and configuring a chaos experiment. The configuring includes identifying fault domains of the applications. The operations also include enabling pre-execution tasks, including load testing and observability, executing the chaos experiment, and automatically subjecting the applications to features of the chaos experiment. The features may be configured to trigger a fault to occur from the applications. The operations collect information from the applications as a result of executing the chaos experiment and execute an AI / ML routine on the information to output a result. The result is representative of the resilience of the applications.
Owner:JPMORGAN CHASE BANK NA

Spanning tree protocol configuration automatic checking method, device, storage medium and system

The invention provides a spanning tree protocol configuration automatic checking method and device, a storage medium and a system, and the method comprises the steps: receiving and responding to a gateway down-moving operation instruction, and migrating the gateway function configuration of convergence layer equipment to access layer equipment; acquiring equipment information of an access layer, and generating a spanning tree protocol configuration information table according to the equipment information; and determining a difference configuration item according to the spanning tree protocol configuration information table and a preset spanning tree protocol configuration information table to complete verification of spanning tree protocol configuration, the difference configuration item being a configuration item in which the spanning tree protocol configuration information table is not consistent with the preset spanning tree protocol configuration information table. According to the method and the device, the problems that a data center gateway is generally deployed on convergence layer equipment, so that the coverage range of a single two-layer network segment is relatively large, and once a two-layer loop fault occurs, the fault domain diffusion range is wide, so that the operation risk is high are solved.
Owner:AGRICULTURAL BANK OF CHINA

Data consistency guarantee method, system and equipment under active-active architecture and medium

PendingCN121935080AImprove reliabilityAvoid synchronization blocking problemsHardware monitoringFailure domainEmbedded system
The invention relates to the technical field of data consistency guarantee. By providing the data consistency guarantee method, system, device and medium under the active-active architecture, the method comprises the following steps: detecting a data service state to generate a breakpoint event; fault domain positioning processing is carried out based on the breakpoint event, and a fault domain identifier is generated; performing packaging processing on the fault domain operation state to generate a packaging operation state, and performing formatting processing on the packaging operation state according to a standardized breakpoint context structure in the strategy template to generate a persistent breakpoint record; analyzing and processing the persistent breakpoint record through a predefined data service adapter to generate a cross-service compensation operation chain; executing the cross-service compensation operation chain in the isolation environment to generate an operation execution result; and performing consistency verification processing on the operation execution result to generate a consistency verification report so as to achieve the technical effects of improving the system throughput, reducing the fault recovery time and enhancing the data consistency verification reliability.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Virtual machine management methods, apparatus, media, devices and computer program products

A virtual machine management method, apparatus, medium, device, and computer program product are disclosed. The method includes: acquiring topology information of multiple virtual machines; determining fault domain information corresponding to each virtual machine based on the topology information; grouping the virtual machines based on the fault domain information to obtain multiple virtual machine groups, wherein virtual machines in the same virtual machine group correspond to the same fault domain information; in response to determining that the primary virtual machine corresponding to a target service is abnormal, determining candidate virtual machines from the multiple virtual machines corresponding to the target service based on the virtual machine group to which the primary virtual machine belongs, wherein the primary virtual machine is one of the multiple virtual machines corresponding to the target service, and the candidate virtual machines and the primary virtual machine belong to different virtual machine groups; and determining the updated primary virtual machine corresponding to the target service from the candidate virtual machines. This can improve the availability of the updated primary virtual machine and ensure the high availability of the target service.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Memory system and system construction method

Maintain adequate availability of storage systems in a cloud environment. [Solution] The system is provided with multiple storage nodes that constitute a first node group spanning multiple failure domains in a cloud environment. For each node, the domain ID of the failure domain in which the node was generated is obtained, and a second node group is formed from the first node group using the necessary number of nodes with domain IDs that do not overlap as much as possible. The number of member nodes in the second node group that exist in the same failure domain is less than or equal to the redundancy level. The redundancy level is the maximum number of member nodes in the second node group that are allowed to stop simultaneously. In the first node group, nodes other than those in the second node group are spare nodes that can be selected as failback destination nodes.
Owner:HITACHI VANTARA LTD

A storage pool creation method, apparatus, device, medium and product

ActiveCN121387205BInput/output to record carriersPoolRecursive computation
The application discloses a storage pool creation method and device, equipment, medium and product, applied to the storage technical field, including: obtaining the number of fault domain units and the number of copies; when the number of fault domain units is greater than the number of copies, whether each fault domain unit meets the first preset capacity balance condition is judged based on the minimum effective capacity corresponding to each fault domain unit and the expected effective capacity; if the first preset capacity balance condition is not met, the target effective capacity is obtained by grouping and recursively calculating each fault domain unit, and whether the second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; in the case of not meeting the second preset capacity balance condition, the main pool is created based on the target effective capacity, and the auxiliary pool is created based on the remaining capacity. In this way, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Virtual fault domain isolation and recovery method suitable for spirit and qi interconnection system

ActiveCN122019244BThird partyInterconnection
The present application relates to the technical field of chip interconnection, and particularly relates to a virtual fault domain isolation and recovery method suitable for a flexible interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; uncompleted transactions in the transaction shadow table are classified and processed, and the state of an access object is listened to; after reconnection, task replay is performed according to the transaction shadow table, and the state machine is switched back to an ACTIVE state. In view of the problem that a third-party node with a fault in the existing flexible interconnection system is prone to cause a chain reaction, an isolation mechanism is added to an IO interconnection chip used for accessing the third-party chip, the state machine of the fault domain is switched to the QUIESCE mode, and a transaction shadow table of uncompleted transactions is established, so that the IO interconnection chip performs proxy processing on part of the transactions, the fault of other devices in the interconnection system is avoided in the fault reconnection stage, and replay is performed after the access object is recovered, so that the risk isolation and recovery process is realized.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Data recovery method and device of storage system, electronic equipment and storage medium

The invention provides a data recovery method and device of a storage system, electronic equipment and a storage medium, and relates to the technical field of distributed storage. The free storage space is used for generating a locally optimized redundant data block which coexists with an original redundant data block of the data object and can recover data in a smaller fault domain range, and when a storage fault is detected, the corresponding redundant data block is selected and called according to the fault range to execute recovery operation. The problems that in the prior art, due to the fact that the locality of erasure codes is insufficient and the fault domain range of original redundant data block recovery data is large, data needs to be accessed across multiple nodes during fault recovery, the transmission amount is large, the recovery speed is low, and the system load is high can be solved, and the purposes of optimizing the locality of the erasure codes, flexibly adapting to different fault scenes and improving the recovery efficiency are achieved. And the cross-node data transmission overhead is reduced, the fault recovery efficiency is improved, and the overall load of the system is reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Inter-satellite reliable routing method based on fault domain model

The application claims a kind of inter-satellite reliable routing method based on fault domain model, belong to satellite communication technical field.Aiming at the problems that inter-satellite link is influenced by space environment and human interference, leading to transmission failure and routing interruption, a centralized inter-satellite reliable routing method based on inter-satellite link attribute is proposed.According to the real-time state and failure reason of inter-satellite link, the inter-satellite link is classified, the initial topology model containing global network information is constructed, the boundary diffusion method is designed combined with the state of boundary node of fault link, the fault link is aggregated into block to form fault domain topology model to limit the influence range of fault.Analysis of the influence of buffer queue capacity, link signal-to-noise ratio and path length on routing, construct comprehensive link utility function to quantify link quality, propose multi-attribute bypass path selection method to solve the problem of load imbalance caused by single path selection criterion, reduce system packet loss rate and end-to-end delay, improve the fault response capability of satellite network.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Yield evaluation method for high-dimensional multi-failure domain of SRAM (Static Random Access Memory) and electronic equipment

The invention relates to the technical field of static random access memories, and provides a yield evaluation method and electronic equipment for high-dimensional multi-failure domains of an SRAM (Static Random Access Memory), and the method comprises the following steps: obtaining failure index parameters of the SRAM, which at least comprise a read access failure parameter, a read destruction failure parameter, a write failure parameter and a hold state failure parameter; performing Monte Carlo processing on the failure index parameters twice to respectively obtain sample point data and evaluation point data; training a preset initial model based on the sample point data to obtain a multi-classification logistic regression model, training adopting a high-dimensional space and a multi-failure domain, the high-dimensional space being determined by the type of the failure index parameter and the sample point data, and the failure domain being determined by dividing high-dimensional space data points based on the maximum radius and the minimum density of a region by a spatial clustering algorithm; and inputting evaluation point data into the model, and determining a yield evaluation result. The problem that high efficiency and high accuracy cannot be considered at the same time when yield evaluation is carried out on the SRAM in the prior art is solved.
Owner:BEIJING KUANWEN MICROELECTRONICS TECH CO LTD

Processor exception recovery method and apparatus, computing device, storage medium, and product

PendingCN122470422AObject basedGranularity
The present disclosure relates to a processor exception recovery method and device, a computing device, a storage medium and a product. The processor exception recovery method comprises: in response to an exception event of a processor, determining at least one exception object related to the exception event and obtaining state data of the at least one exception object; determining an association relationship between the at least one exception object based on the state data to construct an object relationship view of the at least one exception object; obtaining forward advancing state data of the processor; determining a fault domain of the processor based on the object relationship view and the forward advancing state data; determining a recovery granularity corresponding to the fault domain; and performing at least one exception recovery operation of the processor based on the recovery granularity, thereby identifying a real impact range of the exception and determining the exception recovery operation by using the association relationship view constructed around the exception object, improving the accuracy of the fault domain and the exception recovery operation, and further improving the effect of the exception recovery.
Owner:UNIONTECH SOFTWARE TECH CO LTD

Storage system and system construction method

A plurality of storage nodes constituting a first node group across a plurality of fault domains in a cloud environment are provided. For each node, a domain ID of a fault domain in which the node is generated is acquired, and a second node group is configured as a first node group from a necessary number of nodes whose domain IDs do not overlap as much as possible. The number of member nodes existing in the same fault domain in the second node group is equal to or less than the redundancy. The redundancy is the maximum number of member nodes allowed to stop simultaneously in the second node group. In the first node group, a node other than the second node group is a spare node that may be selected as a failback destination node.
Owner:HITACHI VANTARA LTD

Engineering design file multi-version collaborative management method based on block chain

PendingCN122045139ADigital data protectionFile system administrationRetrievabilityFailure domain
The invention discloses an engineering design file multi-version collaborative management method based on a block chain, and relates to the technical field of file version collaborative management, and the method comprises the steps: obtaining the chain state data of a version event, and calculating a settlement control quantity set containing a settlement confidence coefficient lower bound, a time threshold value and a rollback probability function; according to the set and the version level, the number of initial temporary storage copies of the under-chain object and a copy set are determined, and nodes of the copy set meet fault domain coverage constraints and are sorted according to availability scores; carrying out spot check on the copy set according to a spot check interval to obtain a retrievability confidence index, dynamically calculating a temporary storage state retrievability threshold value based on the residual exposed window length and the rollback probability, and triggering redundant copy adjustment; when a long-term storage conversion condition is met, judging whether erasure coding is allowed or not, dynamically determining the number of data pieces and the number of coding pieces, and generating erasure code fragments when the recovery success probability is not lower than a final state recoverability threshold value; retrieability and consistency in the collaboration process of the multi-version engineering design files are improved.
Owner:SHANGHAI YISHU CULTURE MEDIA CO LTD

Data processing method and device of storage system, electronic equipment and storage medium

The invention discloses a data processing method and device of a storage system, electronic equipment and a storage medium, and relates to the technical field of computers.According to the data processing method and device of the storage system, due to the fact that when a fault domain is changed, a communication suppression state is activated firstly to avoid disordered data interaction, and communication cost is accurately calculated in combination with a storage node topological relation; data updating pulling and scheduling reconstruction tasks are initiated according to needs, the limitation of undifferentiated PG broadcast is eliminated, the cross-cabinet communication flow during fault domain change of the distributed storage system is reduced, the communication efficiency of the data reconstruction tasks is optimized, network resource occupation is reduced, and the reliability of the distributed storage system is improved. And therefore, the technical effect of stable and efficient operation of the distributed storage system in scenes such as a data center and a hybrid cloud can be ensured.
Owner:JINAN INSPUR DATA TECH CO LTD

Metadata management system and method for a distributed storage system

The application discloses a metadata management system and method of a distributed storage system, and relates to the computer field, which comprises a cluster management module, a device management module and an object management module.The cluster management module is used for collecting cluster topology information, PG distribution information and fault domain information in the distributed storage system, generating cluster-level metadata according to the cluster topology information, the PG distribution information and the fault domain information, and managing the cluster-level metadata.The device management module is used for monitoring device states of storage devices, space usage of the storage devices and storage space allocation states of the storage devices in real time, generating device-level metadata according to the device states, the space usage and the storage space allocation states, and managing the device-level metadata.The object management module is used for generating object-level metadata of internal objects in the distributed storage system in the case that the internal objects are created, and managing the object-level metadata.
Owner:JINAN INSPUR DATA TECH CO LTD

Erasure Coding With Multiple Fragments On A Single Node

Techniques for efficiently and durably implementing erasure coding. Data objects are processed to extract metadata, which is used to identify fragments of the data object and the stripe to which the fragments belong. The techniques described herein evaluate failure domains at the drive-level rather than at the node-level, thereby greatly expanding the number of failure domains for fragment storage. To support object availability and durability, each drive-level failure may only hold one fragment from any given stripe. Further, the storage drives of each storage node are restricted such that the total number of fragments from any given stripe stored in the storage node does not exceed a limit. With this arrangement, not only can erasure coding can be implemented while using fewer computing resources when compared with node-level failure domains, but data objects can also be reconstructed without compromising data integrity under configurable levels of tolerance to unavailable fragments.
Owner:NETAPP INC

Method and system for positioning online anomaly based on fault domain

The invention discloses a method and a system for positioning an online anomaly based on a fault domain, and the method comprises the steps: S1, configuring a plurality of fault domains, and setting a threshold value of the anomaly of the fault domain; s2, receiving an abnormal alarm, automatically matching to a configured fault domain, and determining a possible fault range; and S3, according to the matched fault domain, calling a specific processing scheme to repair. The fault can be accurately positioned, manual intervention is reduced, the fault recovery speed is increased, and the operation and maintenance cost is reduced.
Owner:JIAXING QINIU INFORMATION TECHNOLOGY CO LTD