Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

57 results about "Failure domain" patented technology

In computer networking, a failure domain encompasses a section of a network that is negatively affected when a critical device or network service experiences problems. The size of a failure domain and its potential impact depends on the device or service that is malfunctioning. For example, a router potentially experiencing problems would generally create a more significant failure domain than a network switch would.

Cloud computing extension cluster high availability method based on dynamic fault domain and intelligent scheduling

The invention relates to a cloud computing extension cluster high availability method based on a dynamic fault domain and intelligent scheduling, and relates to the field of cloud computing extension cluster high availability. According to the method, hardware health data and network performance indexes of physical nodes are collected in real time, and the fault correlation degree between the nodes is calculated based on the hardware health similarity and a network topology attenuation factor; dynamically generating and updating a logic fault domain topological structure according to the fault correlation degree and a preset threshold value; based on the logic fault domain topological structure and a preset SLA strategy library, executing a virtual machine scheduling decision of multi-objective optimization so as to balance business service quality, fault domain risk and resource cost; and in response to the detected fault event, triggering a corresponding hierarchical migration process according to the service priority and the fault level. According to the method, closed-loop linkage of real-time reconstruction and intelligent scheduling of the dynamic fault domain in the cross-region cloud computing extension cluster can be realized, and the disaster tolerance capability and the resource utilization rate of the system are remarkably improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Distributed memory system based on erasure code and dynamic repair mechanism

The invention relates to the field of distributed storage, and discloses a distributed memory system based on erasure codes and a dynamic repair mechanism. Comprising a data node cluster, a metadata management cluster, a bandwidth token management module, a dynamic repair composer and a repair agent node. And the data node cluster divides the object into strip generation data fragments and verification fragments and dispersedly stores the data fragments and the verification fragments according to a placement rule. And the metadata management cluster maintains object stripe mapping, a fragment position set, erasure code protection configuration and a version epoch, and records a repair intention log. And the bandwidth token management module grants repair execution permission according to the fault domain budget. The dynamic repair composer generates availability repair and durability repair tasks, the temporary check fragments are written in when the number of the available check fragments is lower than a threshold value in the availability stage, and the configuration fragments are recovered and the temporary fragments are recovered in the durability stage. And the consistency of version epochs is verified before the agent node is repaired and written back, and if not, recalculation is carried out.
Owner:EASY POINT GEEK (BEIJING) TECHNOLOGY CO LTD

Distributing certificate bundles according to fault domains

Operations of a certificate bundle distribution service may include: detecting a trigger condition to distribute a certificate bundle that includes a set of certificate authority certificates; determining, for each of a plurality of network entities associated with a computer network, a fault domain representing at least one single point of failure; partitioning the plurality of network entities into a plurality of certificate distribution groups, based on a set of partitioning criteria that includes a fault domain of each particular network entity, in which each particular certificate distribution group includes a particular subset of network entities, and the particular subset of network entities are associated with a particular fault domain; selecting a particular certificate distribution group, of the plurality of certificate distribution groups, for distribution of the certificate bundle; and transmitting the certificate bundle to the particular subset of network entities in the particular certificate distribution group.
Owner:ORACLE INT CORP

Distributed system network interface dynamic expansion method and device based on dependency perception

The invention relates to a distributed system network interface dynamic expansion method and device based on dependency perception, and relates to the field of distributed system network interface dynamic expansion. The method comprises the following steps: collecting topological information of each node in a distributed cluster, and calculating the dependency strength between node pairs based on the communication frequency between the nodes, a service call chain and a resource access mode; according to the number of copies and the minimum number of copies of each copy group, constructing a fault domain constraint model, and verifying whether the candidate expansion scheme meets the system availability requirement or not; a heuristic search method is adopted to generate a batch expansion scheduling sequence, and it is ensured that the node dependency relationship of each batch of expansion is weakest and meets constraint conditions; and executing configuration updating operation according to the scheduling sequence, and realizing atomized execution of configuration modification and service restart by adopting a two-stage submission protocol. According to the invention, the zero-shutdown dynamic expansion of the network interface of the distributed system is realized, the service continuity and the configuration consistency are ensured, and the operation and maintenance interruption time and the operation risk are obviously reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Adaptive PC-Kriging reliability analysis method and system based on active learning

The invention discloses an adaptive PC-Kriging reliability analysis method and system based on active learning, and the method comprises the steps: obtaining uniformly distributed first candidate samples from a pre-generated MC sample pool through employing a weight clustering method, and constructing an initial PC-Kriging model; selecting samples by using an interval reduction method to construct a new sample pool, and dividing the new sample pool into two sub-sample pools of a security domain and an invalidation domain; aiming at the two sub-sample pools, adopting the weight clustering method again, and respectively selecting uniformly distributed second candidate samples from the two sub-sample pools; constructing a crossing point by using the second candidate samples distributed in the failure domain and the security domain; taking the constructed crossing point as a newly added experimental point, and carrying out iterative updating on the PC-Kriging model; and judging convergence and outputting a result. According to the technical scheme, through interval reduction of the focusing key area and model updating based on the crossing points, the limit state surface can be more accurately approached, the number of experimental points is reduced, the prediction precision is improved, and the calculation resource utilization is more efficient.
Owner:CHANGZHOU INST OF TECH

Virtual fault domain isolation and recovery method suitable for Lingqu interconnection system

The invention relates to the technical field of chip interconnection, in particular to a virtual fault domain isolation and recovery method suitable for a Lingqu interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; carrying out classification processing on the uncompleted transactions in the transaction shadow table, and monitoring the state of the access object; and after reconnection, task replay is carried out according to the transaction shadow table, and the state machine is switched back to the ACTIVE state. In order to solve the problem that in the prior art, a third-party node with a fault in a Lingqu interconnection system is prone to causing chain reaction, an isolation mechanism is added to an IO interconnection chip used for being connected to a third-party chip, a state machine of a fault domain is switched to a QUIESCE mode, a transaction shadow table of an uncompleted transaction is established, and the transaction shadow table of the uncompleted transaction is obtained. According to the method and the system, the IO interconnection chip performs proxy processing on part of transactions after the access object is recovered, so that faults of other equipment in the interconnection system are prevented from being caused in a fault reconnection stage, and replaying is performed after the access object is recovered, thereby realizing risk isolation and recovery processes.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Intelligently forming data stripes including multiple shards in a single failure domain

Redundant array of independent drives (RAID) sub-stripes are formed across one or more solid-state storage devices of storage nodes of a storage system. The RAID sub-stripes include corresponding shards of data to be stored at the solid-state storage devices, wherein at least one of the RAID sub-stripes has at least two of the corresponding shards of data on a same storage node. At least one global parity shard is generated for the RAID sub-stripes. The at least one global parity shard is to be stored on a first storage node that is different from each of the same storage nodes storing the at least two of the corresponding shards of data.
Owner:PURE STORAGE INC

OSD switching method, system and device in distributed storage pool and medium

The invention discloses an OSD switching method, system and device in a distributed storage pool and a medium, and relates to the technical field of distributed storage, and the method comprises the steps: determining a fault OSD in the distributed storage pool; determining a physical fault domain of a target storage pool where the fault OSD is located; acquiring a global allowable failure domain number of the physical fault domain; determining the number of failed fault domains of the physical fault domain; in response to the fact that the number of the failed fault domains is smaller than the number of the global allowable failure fault domains, marking the state of the fault OSD as down; the global allowable failure fault domain number comprises a minimum value of failure of the fault domain allowed by the target storage pool set, and the target storage pool set comprises a set of all storage pool combinations sharing the physical fault domain. According to the method, the number of the newest failed fault domains is at most equal to the number of the failure of the minimum fault domain, the situation that the number of the newest failed fault domains exceeds the number of the failure of the minimum fault domain is avoided, and the service reliability and the data security of distributed storage in a fault switching scene are improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Block-storage service supporting multi-attach and health check failover mechanism

A block-based storage system hosts logical volumes that are implemented via multiple replicas of volume data stored on multiple resource hosts in different failure domains. Also, the block-based storage service allows multiple client computing devices to attach to a same given logical volume at the same time. In order to prevent unnecessary failovers, a primary node storing a primary replica is configured with a health check application programmatic interface (API) and a secondary node storing a secondary replica determines whether or not to initiate a failover based on the health of the primary replica.
Owner:AMAZON TECH INC

Automated multi-region application recovery

A recovery orchestrator system receives a recovery plan, which may be used for a hosted-computing environment. The recovery plan includes multiple steps. The recovery orchestrator system receives monitoring metrics. The recovery orchestrator system executes the recovery plan based on the monitoring metrics. The recovery orchestrator system executes steps from the recovery plan in multiple fault domains. The recovery orchestrator system monitors the status of the execution of the recovery plan and provides the status update to a computing device.
Owner:AMAZON TECH INC

Distributing Certificate Bundles According To Fault Domains

Operations of a certificate bundle distribution service may include: detecting a trigger condition to distribute a certificate bundle that includes a set of certificate authority certificates; determining, for each of a plurality of network entities associated with a computer network, a fault domain representing at least one single point of failure; partitioning the plurality of network entities into a plurality of certificate distribution groups, based on a set of partitioning criteria that includes a fault domain of each particular network entity, in which each particular certificate distribution group includes a particular subset of network entities, and the particular subset of network entities are associated with a particular fault domain; selecting a particular certificate distribution group, of the plurality of certificate distribution groups, for distribution of the certificate bundle; and transmitting the certificate bundle to the particular subset of network entities in the particular certificate distribution group.
Owner:ORACLE INT CORP

Adaptive pc-kriging reliability analysis method and system based on active learning

This invention discloses an adaptive PC-Kriging reliability analysis method and system based on active learning. The method includes: using weighted clustering to obtain uniformly distributed first candidate samples from a pre-generated MC sample pool and constructing an initial PC-Kriging model; using interval reduction to select samples and construct a new sample pool, dividing the new sample pool into two sub-sample pools: a safe domain and a failure domain; for each sub-sample pool, weighted clustering is used again to select uniformly distributed second candidate samples; using the second candidate samples distributed in the failure domain and the safe domain, crossing points are constructed; the constructed crossing points are used as new experimental points to iteratively update the PC-Kriging model; convergence is determined and the results are output. The technical solution of this invention focuses on key regions through interval reduction and updates the model based on crossing points, which can more accurately approximate the limit state surface, reduce the number of experimental points, improve prediction accuracy, and is more efficient in utilizing computational resources.
Owner:CHANGZHOU INST OF TECH

Equipment fault processing method, host, computer program product and storage medium

The embodiment of the invention provides an equipment fault processing method, a host, a computer program product and a storage medium. The operating system of the target host can run an error report driving program after receiving an interrupt signal sent by any root interface and used for triggering an error report mechanism; based on an error report driving program, obtaining an address translation exception record from an event queue maintained by a memory management unit; if the address translation exception record points to any virtual device connected to the root interface, calling a kernel mode drive program corresponding to the virtual device through an error report drive program; and controlling the target virtual machine served by the virtual equipment to be paused based on the kernel mode driving program. Therefore, when a single virtual device has a fault, the fault domain can be accurately controlled on the virtual machine served by the faulted virtual device, so that only the virtual machine served by the faulted virtual device is suspended, and the influence on other virtual machines running on the target host can be effectively avoided.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Storage pool creation method and device, equipment, medium and product

ActiveCN121387205AInput/output to record carriersPoolRecursive computation
The invention discloses a storage pool creation method and device, equipment, a medium and a product, which are applied to the technical field of storage, and comprise the following steps: obtaining the number of fault domain units and the number of copies; when the number of the fault domain units is greater than the number of the copies, judging whether each fault domain unit meets a first preset capacity balance condition or not based on the minimum effective capacity and the expected effective capacity corresponding to each fault domain unit; if the first preset capacity balance condition is not met, grouping recursive calculation is carried out on each fault domain unit to obtain a target effective capacity, and whether a second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; and under the condition that a second preset capacity balance condition is not met, creating a main pool based on the target effective capacity and creating an auxiliary pool based on the residual capacity. Therefore, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method for analyzing gradual change reliability and global sensitivity of cabin door lifting mechanism

The invention provides a cabin door lifting mechanism gradual change reliability and global sensitivity analysis method, and relates to the technical field of reliability and sensitivity analysis. Dynamic modeling, gradual change failure domain description, Kriging proxy modeling based on active learning and global sensitivity analysis are organically fused; the performance degradation characteristics and the reliability evolution process of the cabin door lifting mechanism in the whole life cycle can be accurately described in a multi-level and multi-dimension mode. By introducing a gradual change failure membership function, continuous evolution of structural performance from safety to failure is completely reflected, the failure probability and the variation coefficient thereof are quantitatively calculated through Monte Carlo simulation, and clear redundancy margin is reserved for safety design. Meanwhile, by means of an iterative sampling strategy of active learning-Kriging and calculation of a global sensitivity index, the local and global prediction precision and convergence efficiency of the proxy model are remarkably improved, and the contribution degree and interaction effect of each influence factor to the overall reliability are revealed.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Network fault processing method and related device

The invention discloses a network fault processing method, which can reduce the calculation burden and message forwarding overhead of network equipment. In the method, when a link connected with network equipment fails, the network equipment sends a link fault message to neighbor equipment, and the link fault message is only diffused in a certain range taking the network equipment as a center, so that a fault domain with a fixed size is formed. Therefore, the network equipment connected with the fault link only needs to re-calculate the path in the fault domain range, and the re-calculated path is used for replacing the fault link, so that the routing convergence can be completed. By setting the fault domain with the fixed size, each network device with the link fault only needs to maintain the own local fault domain and calculate the corresponding path, so that the problem that the link fault message diffuses in the whole network to cause frequent path calculation of the network devices in the whole network is avoided; and the calculation burden and the message forwarding overhead of the network equipment are effectively reduced.
Owner:HUAWEI TECH CO LTD

Storage based on fault domains and storage classes

Examples include a computing device configured by executable instructions to categorize a plurality of storage components into a plurality of storage fault domains, each storage fault domain corresponding to at least one failure scenario not shared by the other storage fault domains. The computing device may receive a data object for storage. The computing device may store at least a portion of the data object to a first storage fault domain, and may store at least another portion of the data object to a second storage fault domain that is different from the first storage fault domain.
Owner:HITACHI VANTARA LLC

Methods and systems for chaos testing

Provided are systems for automated chaos including a processor and a memory having instructions stored thereon. The instructions, when executed, cause the processor to perform certain operations including connecting to an application infrastructure with one or more applications and inspecting a code of the one or more applications and configuring a chaos experiment. The configuring includes identifying fault domains of the applications. The operations also include enabling pre-execution tasks, including load testing and observability, executing the chaos experiment, and automatically subjecting the applications to features of the chaos experiment. The features may be configured to trigger a fault to occur from the applications. The operations collect information from the applications as a result of executing the chaos experiment and execute an AI / ML routine on the information to output a result. The result is representative of the resilience of the applications.
Owner:JPMORGAN CHASE BANK NA

Techniques for generating a configuration for electrically isolating fault domains in a data center

A computer system may receive a layout of a data center, the layout of the data center identifying physical locations of a plurality of server racks, electrical distribution feeds, and uninterruptible power supplies. The computer system may receive a fault domain configuration for the datacenter, the fault domain configuration identifying virtual locations of a plurality of logical fault domains for distributing one or more instances so that the instances are stored on independent physical hardware devices within a single availability fault domain. The computer system may determine the configuration for the data center by assigning the plurality of fault domains to a plurality of electrical zones, wherein each electrical zone provides a redundant electrical power supply across the plurality of logical fault domains in an event of a failure of one or more electrical distribution feeds. The computer system may display the configuration for the data center on a display.
Owner:ORACLE INT CORP

Spanning tree protocol configuration automatic checking method, device, storage medium and system

The invention provides a spanning tree protocol configuration automatic checking method and device, a storage medium and a system, and the method comprises the steps: receiving and responding to a gateway down-moving operation instruction, and migrating the gateway function configuration of convergence layer equipment to access layer equipment; acquiring equipment information of an access layer, and generating a spanning tree protocol configuration information table according to the equipment information; and determining a difference configuration item according to the spanning tree protocol configuration information table and a preset spanning tree protocol configuration information table to complete verification of spanning tree protocol configuration, the difference configuration item being a configuration item in which the spanning tree protocol configuration information table is not consistent with the preset spanning tree protocol configuration information table. According to the method and the device, the problems that a data center gateway is generally deployed on convergence layer equipment, so that the coverage range of a single two-layer network segment is relatively large, and once a two-layer loop fault occurs, the fault domain diffusion range is wide, so that the operation risk is high are solved.
Owner:AGRICULTURAL BANK OF CHINA

Data consistency guarantee method, system and equipment under active-active architecture and medium

PendingCN121935080AImprove reliabilityAvoid synchronization blocking problemsHardware monitoringFailure domainEmbedded system
The invention relates to the technical field of data consistency guarantee. By providing the data consistency guarantee method, system, device and medium under the active-active architecture, the method comprises the following steps: detecting a data service state to generate a breakpoint event; fault domain positioning processing is carried out based on the breakpoint event, and a fault domain identifier is generated; performing packaging processing on the fault domain operation state to generate a packaging operation state, and performing formatting processing on the packaging operation state according to a standardized breakpoint context structure in the strategy template to generate a persistent breakpoint record; analyzing and processing the persistent breakpoint record through a predefined data service adapter to generate a cross-service compensation operation chain; executing the cross-service compensation operation chain in the isolation environment to generate an operation execution result; and performing consistency verification processing on the operation execution result to generate a consistency verification report so as to achieve the technical effects of improving the system throughput, reducing the fault recovery time and enhancing the data consistency verification reliability.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Isolation evaluation method and device for distributed storage system, equipment and medium

The invention provides an isolation evaluation method and device for a distributed storage system, equipment and a medium. In the technical scheme provided by the invention, a centralized isolation evaluation center is introduced and a cluster steady-state sensing mechanism is combined; isolation evaluation logics originally dispersed on OSDs are collected to an isolation evaluation center elected by a state monitoring assembly to be processed in a unified mode, all isolation requests can be executed after being subjected to serialization evaluation through the isolation evaluation center, and when the isolation evaluation center judges that the distributed storage system is in an unstable state according to the state of a global PG fragment, all the isolation requests can be executed; according to the method, whether the redundancy of each isolated PG still meets the lowest availability requirement of the redundancy strategy adopted by the PG to which the PG fragment belongs is strictly evaluated, the situation that part of PGs cannot normally provide read-write response due to cross-fault-domain and multi-point concurrent isolation caused by multi-node decentralized decision is avoided, and the data security and service continuity of the system are improved.
Owner:XINHUASAN INFORMATION TECH CO LTD

Virtual machine management methods, apparatus, media, devices and computer program products

A virtual machine management method, apparatus, medium, device, and computer program product are disclosed. The method includes: acquiring topology information of multiple virtual machines; determining fault domain information corresponding to each virtual machine based on the topology information; grouping the virtual machines based on the fault domain information to obtain multiple virtual machine groups, wherein virtual machines in the same virtual machine group correspond to the same fault domain information; in response to determining that the primary virtual machine corresponding to a target service is abnormal, determining candidate virtual machines from the multiple virtual machines corresponding to the target service based on the virtual machine group to which the primary virtual machine belongs, wherein the primary virtual machine is one of the multiple virtual machines corresponding to the target service, and the candidate virtual machines and the primary virtual machine belong to different virtual machine groups; and determining the updated primary virtual machine corresponding to the target service from the candidate virtual machines. This can improve the availability of the updated primary virtual machine and ensure the high availability of the target service.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Memory system and system construction method

Maintain adequate availability of storage systems in a cloud environment. [Solution] The system is provided with multiple storage nodes that constitute a first node group spanning multiple failure domains in a cloud environment. For each node, the domain ID of the failure domain in which the node was generated is obtained, and a second node group is formed from the first node group using the necessary number of nodes with domain IDs that do not overlap as much as possible. The number of member nodes in the second node group that exist in the same failure domain is less than or equal to the redundancy level. The redundancy level is the maximum number of member nodes in the second node group that are allowed to stop simultaneously. In the first node group, nodes other than those in the second node group are spare nodes that can be selected as failback destination nodes.
Owner:HITACHI VANTARA LTD

Dynamic storage resiliency

A computer system is configured to provision a plurality of storage volumes at a plurality of fault domains and thinly provision a plurality of cache volumes at the plurality of fault domains. The computer system is also configured to perform a write operation in a resilient manner that maintains a plurality of copies of data associated with the write operation. Performing the write operation in the resilient manner includes allocating a portion of storage in each of the plurality of cache volumes, and caching the data associated with the write operation in the portion of storage in each of the plurality of cache volumes. The cached data is then persistently stored in the plurality of storage volumes. After that, the portion of storage in each of the plurality of cache volumes is deallocated.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-failure domain structure reliability optimization method and system based on adaptive clustering

The invention discloses a multi-failure domain structure reliability optimization method and system based on adaptive clustering, and the method comprises the steps: carrying out the first-stage training of a Kriging agent model through an active learning method based on engineering structure data, and constructing a first-stage Kriging agent model; performing failure domain identification on the first-stage Kriging agent model based on an adaptive clustering algorithm, and constructing an important sampling density; based on the important sampling density, performing second-stage training on the first-stage Kriging agent model through an active learning method, and outputting a second-stage Kriging agent model; and performing failure probability calculation on the engineering structure based on the second-stage Kriging agent model to obtain a reliability calculation result of the engineering structure. According to the method, the potential sub-failure domain of the engineering structure can be accurately identified, so that the calculation efficiency of the failure probability of the engineering structure is improved. The multi-failure domain structure reliability optimization method and system based on adaptive clustering can be widely applied to the technical field of structure reliability optimization.
Owner:FOSHAN UNIVERSITY

System upgrading method and device, equipment, storage medium and computer program product

The invention discloses a system upgrading method and device, equipment, a storage medium and a computer program product. The method comprises the steps that a first node in a distributed storage cluster determines first information, second information and third information, the first information represents the state of the distributed storage cluster, the second information represents one or more second nodes to be subjected to system upgrading, and the third information represents a fault domain and a storage pool division result of the distributed storage cluster; the first information, the second information and the third information are utilized to determine one or more node groups and an upgrading sequence corresponding to each node group, each node group comprises one or more second nodes, the second nodes contained in different node groups are completely different, and different nodes contained in the same node group belong to the same fault domain or belong to different storage pools; and according to the upgrading sequence of each node group, performing system upgrading on all the second nodes contained in the node group in sequence.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

A storage pool creation method, apparatus, device, medium and product

ActiveCN121387205BInput/output to record carriersPoolRecursive computation
The application discloses a storage pool creation method and device, equipment, medium and product, applied to the storage technical field, including: obtaining the number of fault domain units and the number of copies; when the number of fault domain units is greater than the number of copies, whether each fault domain unit meets the first preset capacity balance condition is judged based on the minimum effective capacity corresponding to each fault domain unit and the expected effective capacity; if the first preset capacity balance condition is not met, the target effective capacity is obtained by grouping and recursively calculating each fault domain unit, and whether the second preset capacity balance condition is met is judged based on the target effective capacity and the expected effective capacity; in the case of not meeting the second preset capacity balance condition, the main pool is created based on the target effective capacity, and the auxiliary pool is created based on the remaining capacity. In this way, the utilization rate of storage resources can be improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Virtual fault domain isolation and recovery method suitable for spirit and qi interconnection system

ActiveCN122019244BThird partyInterconnection
The present application relates to the technical field of chip interconnection, and particularly relates to a virtual fault domain isolation and recovery method suitable for a flexible interconnection system, which comprises the following steps: when a fault event occurs, a state machine is switched to a QUIESCE mode, and a transaction shadow table is maintained; uncompleted transactions in the transaction shadow table are classified and processed, and the state of an access object is listened to; after reconnection, task replay is performed according to the transaction shadow table, and the state machine is switched back to an ACTIVE state. In view of the problem that a third-party node with a fault in the existing flexible interconnection system is prone to cause a chain reaction, an isolation mechanism is added to an IO interconnection chip used for accessing the third-party chip, the state machine of the fault domain is switched to the QUIESCE mode, and a transaction shadow table of uncompleted transactions is established, so that the IO interconnection chip performs proxy processing on part of the transactions, the fault of other devices in the interconnection system is avoided in the fault reconnection stage, and replay is performed after the access object is recovered, so that the risk isolation and recovery process is realized.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Data storage method, storage system, storage device and storage equipment

The embodiment of the invention discloses a data storage method, a storage system, a storage device and storage equipment, which are used for improving the storage space utilization rate of the storage system. The data storage method comprises the steps that each first-level fault domain is divided into a plurality of second-level fault domains, each second-level fault domain comprises part of storage devices in a plurality of storage devices, and the part of the storage devices included in each second-level fault domain are different; to-be-written data is divided into a plurality of data block groups, each data block group comprises N data blocks, and N is an integer greater than zero; generating at least one primary verification block according to the N data blocks in each data block group to obtain multiple groups of primary verification blocks; and after different data blocks in each data block group are stored in storage devices contained in different second fault domains, a second-level verification block is generated according to the multiple groups of first-level verification blocks, and the second-level verification block is a verification block of the multiple groups of data blocks.
Owner:CHENGDU HUAWEI TECH CO LTD