Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2419results about "Redundant hardware error correction" patented technology

Multi-cluster distributed training-oriented disaster recovery drill and performance evaluation method and system

The invention relates to the technical field of cloud computing and artificial intelligence, in particular to a disaster recovery drill and performance evaluation system and method for multi-cluster distributed training. In order to solve the problem of disaster recovery and performance evaluation in multi-cluster distributed training, the method comprises the steps of environment analysis and modeling, fault injection and drill control, cross-cluster scheduling and resource monitoring, performance data acquisition and index aggregation, training result evaluation and tuning closed loop and the like. Multi-cluster environment elements are identified, a fault and index mapping relation is constructed, faults are injected by using a fault injection tool, and task scheduling, resource monitoring and performance data acquisition and analysis are realized by means of a multi-cluster management and monitoring tool, so that an adjusting and optimizing strategy is formed, and closed-loop management is realized. According to the method, faults can be dynamically simulated in a multi-cluster environment, the training performance is monitored, the disaster recovery and recovery efficiency is evaluated, the training efficiency and robustness of the system are improved, and the requirements for high availability and high performance in the fields of finance, automatic driving and the like are met.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

IT asset fault propagation prediction method and system based on dynamic evolution of knowledge graph

The invention discloses an IT asset fault propagation prediction method and system based on dynamic evolution of a knowledge graph, and relates to the technical field of cloud computing and large-scale IT operation and maintenance management. Through an asynchronous message bus and a logic clock, the knowledge graph is updated immediately when resources are abnormal and a scheduling event occurs; the knowledge graph uniformly integrates physical connection, logic dependence and multi-copy redundancy, so that the cross-machine-room asset relationship is clear at a glance. And then, based on a weighted logistic regression model, node features and relation weights in the knowledge graph are fused, the node fault probability is accurately calculated, the limitation of traditional single-dimensional analysis is solved, self-healing operation is supported, end-to-end intelligent operation and maintenance from fault detection to prediction and early warning to closed-loop self-healing are realized, and the fault detection efficiency is improved. The problems that in a cross-machine-room and multi-live-site environment, resource topology is split, real-time state and alarm information cannot be fused with an asset dependence model, and large-scale real-time deployment of a traditional single-dimensional fault analysis and high-complexity prediction algorithm is difficult are effectively solved.
Owner:GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU

Heterogeneous computing multi-target adaptive task scheduling method based on deep reinforcement learning

The invention discloses a heterogeneous computing multi-target adaptive task scheduling method based on deep reinforcement learning, and the method comprises the following steps: S1, constructing a multi-dimensional dynamic perception model of a heterogeneous computing environment, and collecting and computing node performance indexes, task feature parameters and network states in real time; s2, defining a reward function as a multi-target weighted combination, fusing task completion time, energy consumption, resource utilization rate and cost, and dynamically adjusting the weight by a fuzzy comprehensive evaluation algorithm; s3, establishing a dual-channel deep reinforcement learning network architecture based on an attention mechanism; s4, establishing an adaptive exploration mechanism, combining an epsilon-greedy strategy and entropy regularization, and balancing exploration and utilization; the method has the beneficial effects that dynamic balance of multiple indexes such as task completion time, energy consumption and resource utilization rate is realized through combination of deep reinforcement learning and multi-objective optimization, a dual-channel network and a cross attention mechanism are adopted, and a task time sequence characteristic and a topological dependency relationship are modeled at the same time, so that a scheduling strategy is more accurate.
Owner:王立强

Distributed simulation method and system based on containerized deployment and elastic expansion

The invention relates to the technical field of distributed simulation based on containerized deployment and elastic expansion and contraction, and discloses a distributed simulation method and system based on containerized deployment and elastic expansion and contraction. According to the distributed simulation method and system based on containerized deployment and elastic expansion, real-time resource monitoring and historical load trend data of a simulation task are collected, and unified load evaluation is carried out in combination with a simulation calculation complexity parameter and an I / O density parameter; the refined modeling and resource demand pre-judgment of the simulation task are realized, and the accuracy of task scheduling and allocation is improved; modularized deployment of simulation tasks is achieved through subtask segmentation based on minimum executable units and a standard containerization packaging mechanism in cooperation with a container arrangement platform, and then concurrent execution and elastic scheduling in a multi-node environment are supported.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Multi-source heterogeneous data-based industry return on investment real-time acquisition method

The invention discloses an industry return on investment real-time acquisition method based on multi-source heterogeneous data, and relates to the technical field of data analysis and processing, and the method comprises the following steps: normalizing a multi-source timestamp and retaining metadata information based on a UTC reference, an adjustable gain and a power exponent; in the streaming processing, reordering of out-of-order data and marking of expired data are realized by setting a buffer queue and a watermark threshold value; when null value fields or cross-source conflicts are detected, interpolation and conflict processing are carried out, and complete data subjected to consistency correction are output; then, an external high-precision reference or a cross-correlation function is used for further fine tuning the timestamp at a millisecond level, and if the adjustment amplitude exceeds a safety boundary, the timestamp is marked as suspicious; and finally, the corrected data is distributed to multiple nodes according to a fragment mapping strategy, and fault-tolerant consensus and difference repair are triggered under fault duration judgment, so that high-precision time sequence alignment and high availability in a massive concurrent scene are kept, data loss or precision attenuation caused by time sequence inconsistency and node faults is avoided, and the method can be widely applied to high-frequency analysis.
Owner:JIANGXI LAYOUT DIGITAL TECHNOLOGY CO LTD

Reservoir dam operation safety sky-ground work intelligent sensing system and operation method

The invention relates to a reservoir dam operation safety sky-land project intelligent sensing system and an operation method, and relates to the technical field of hydraulic engineering safety monitoring. The system is composed of a sky-land water conservancy project integrated monitoring and sensing system, a self-adaptive sampling module, a layered distributed architecture and a software and hardware integrated module, and multi-source data such as deformation, seepage, stress strain, vibration and environmental quantity are cooperatively collected through five dimensions of sky domain, airspace, territory, water domain and work domain. The monitoring frequency is dynamically adjusted by using an adaptive sampling strategy, and data cleaning, standardization, space-time registration and fusion processing are completed through a distributed architecture to generate a high-quality comprehensive data set. The system can realize total-factor and whole-process refined monitoring, effectively eliminates data islands, improves data quality and monitoring efficiency, has high reliability, real-time performance and expandability, and provides powerful data support and decision basis for dam safety assessment and intelligent early warning.
Owner:CHANGJIANG SPATIAL INFORMATION TECH ENG CO LTD (WUHAN) +1

Full-stack NPU system supporting multi-stage fault mitigation mechanism

The invention discloses a full-stack NPU system supporting a multistage fault mitigation mechanism, and the system comprises a plurality of NPU functional blocks which are used for executing a neural network calculation task; the plurality of hardware protection modules are used for providing error detection and correction capabilities for at least one NPU functional block in the NPU or a data path in the NPU; and a control mechanism for coordinating operation of the plurality of hardware protection modules. By means of the scheme, NPU functional block calculation errors or data path transmission errors caused by hardware faults and the like can be detected and corrected in time, wrong calculation results are prevented from being continuously used or output, and the accuracy of final results is guaranteed. Meanwhile, even if part of hardware breaks down, an error correction mechanism can also cover errors to a large extent, normal operation of the system is maintained, the overall availability and task success rate of the system are improved, and therefore the reliability and robustness of the NPU are remarkably improved.
Owner:DALIAN UNIV OF TECH

Techniques for efficient replication and recovery

Techniques are described for efficient replication and maintaining snapshot data consistency during file storage replication between file systems in different cloud infrastructure regions. In certain embodiments, provenance IDs are used to efficiently identify a starting point (e.g., a base snapshot) for a cross-region replication process, conserve cloud resources while reducing network and IO traffic.
Owner:ORACLE INT CORP

Techniques for maintaining snapshot data consistency during file system cross-region replication

Techniques are described for efficient replication and maintaining snapshot data consistency during file storage replication between file systems in different cloud infrastructure regions. In certain embodiments, snapshot creation and deletion requests that occur during cross-region replications may be temporarily withheld until appropriate times to execute such requests safely, depending on the timing relationship between such requests and cross-region replication cycles.
Owner:ORACLE INT CORP

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

AI server data processing optimization system and method based on distributed heterogeneous computing

The invention discloses an AI server data processing optimization system and method based on distributed heterogeneous computing, and relates to the technical field of artificial intelligence and distributed heterogeneous computing, the AI server data processing optimization system comprises a distributed heterogeneous computing cluster, a dynamic collaborative scheduling subsystem and a double-layer closed loop verification subsystem; the distributed heterogeneous computing cluster is characterized in that the distributed heterogeneous computing cluster comprises a plurality of computing nodes which are interconnected, and each node integrates at least two hardware units. According to the AI server data processing optimization system and method based on distributed heterogeneous computing, deep collaborative optimization of computing resource scheduling and data stream transmission is achieved by constructing a hardware-data-task ternary graph model and a space-time collaborative propagation algorithm, so that end-to-end processing delay is shortened, and the heterogeneous resource utilization rate is maximized.
Owner:LOGOSDATA

Vehicle control device

A control execution unit is connected both a main bus and a sub bus, and includes an execution unit and a selection unit. The execution unit performs vehicle control according to a selected manipulated variable being either a main manipulated variable from a main processing unit connected to the main bus or a sub manipulated variable from a sub processing unit connected to the sub bus. The selected manipulated variable is set to the main manipulated variable in an initial state, and the selection unit switches over the selected manipulated variable from the main manipulated variable to the sub manipulated variable when communication performed via the main bus satisfies a preset switchover condition.
Owner:DENSO CORP

Large model batch reasoning and data flow optimization system oriented to MOE architecture

The invention relates to the technical field of project management, in particular to a large-model batch reasoning and data flow optimization system oriented to an MOE architecture. According to the method, a collaborative architecture of the request access module, the environment sensing module, the expert routing engine, the resource scheduling module and the dynamic optimization control module is set, the text length and the subject type are extracted by using the request access module, a basis is provided for accurate routing, and the GPU video memory, the I / O bandwidth and the request queue depth are acquired in real time through the environment sensing module, so that the real-time routing is realized. The system load is comprehensively monitored, meanwhile, an expert sub-network is activated through an expert routing engine according to request features, invalid calculation is avoided, weight loading and resource allocation are managed through a resource scheduling module, the I / O bottleneck is reduced, and finally an optimization strategy is intelligently triggered through a dynamic optimization control module based on routing conflict factors. The problems of large reasoning delay fluctuation and unbalanced resource utilization rate mentioned in the background technology are solved, and stable low-delay response and resource collaborative optimization in a high-concurrency scene is realized.
Owner:VIRTAI TECH BEIJING CO LTD

Managing resource constraints in a cloud environment

Techniques for managing resource constraints of a cloud environment are disclosed. A system receives a request to initiate a provisioning process for provisioning a first service in the cloud environment. The system determines a resource constraint associated with a resource that the first service utilizes. Based on the resource constraint, the system determines a set of candidate services that also utilize the resource as candidates for deprovisioning from the cloud environment. The system identifies respective service features of the set of candidate services and generates a ranking of the set of candidate services based on weighting metrics associated with the respective service features. Based on the ranking, the system selects a second service of the set of candidate services for deprovisioning from the cloud environment. The system deprovisions the second service to alleviate the resource constraint and then provisions the first service by executing the provisioning process.
Owner:ORACLE INT CORP

Fast database recovery in a multi-volume database environment via transactional awareness

Techniques for fast database recovery in a multi-volume database environment via transactional awareness are described. In the event of a failure associated with a first volume storing database page data, the first volume can be restored to a point in time and transactional metadata from a second volume storing logical change data can be obtained for a limited number of transactions occurring at / after that point in time, as opposed to analyzing extremely large change log files. These transactions can be checked to ensure that they have all been persisted, and if not, change data for those transactions can be obtained from the second volume and used to replay these transactions on the restored first volume.
Owner:AMAZON TECH INC

Satellite-borne equipment single event upset resisting system based on FPGA (Field Programmable Gate Array) and implementation method

The invention relates to an FPGA (Field Programmable Gate Array)-based satellite-borne equipment single event upset resisting system and an implementation method. The system comprises an anti-fuse FPGA chip, an SRAM (Static Random Access Memory) FPGA chip and an FPGA chip, the SRAM type FPGA chip is used for completing signal processing; the three FLASH chips are used for solidifying three identical SRAM (Static Random Access Memory) FPGA programs; the anti-fuse FPGA chip reads backup programs from the three FLASH chips respectively at the same time after the equipment is normally powered on, compares the three backup programs, starts to vote when at least two kinds of data are consistent through comparison, brushes the same data into the SRAM type FPGA chip to complete configuration of the FPGA chip, and refreshes the FPGA chip at regular time at the same time. By adopting the method, the single event upset resistance of the FPGA in the space environment can be enhanced, and the method has the advantages of high stability, low cost and the like.
Owner:HUNAN ZHONGDIAN HUARONG ENTERPRISE MANAGEMENT CO LTD

End-to-end restartability of cross-region replication using a new replication

Techniques are described for performing different types of restart operations for a file storage replication between a source file system and a target file system in different cloud infrastructure regions. In certain embodiments, the disclosed techniques perform a restart operation to terminate a current cross-region replication by synchronizing resource cleanup operations in the source file system and the target file system, respectively. In other embodiments, disclosed techniques perform a restart operation to allow a customer to reuse the source file system by identifying a restartable base snapshot in the source file system without dependency on the target file system.
Owner:ORACLE INT CORP

Dynamically selecting artificial intelligence models and hardware environments to execute tasks

The present disclosure relates to systems, non-transitory computer-readable media, and methods for selecting machine-learning models and hardware environments for executing a task. In particular, in one or more embodiments, the disclosed systems select a designated machine-learning model for executing a task based on workload features of the task and task routing metrics for a plurality of machine-learning models. In addition, in one or more embodiments, the disclosed systems select a designated hardware environment for executing the task based on workload features for the task and task routing metrics for a plurality of hardware environments. In some embodiments, the disclosed systems select a fallback machine-learning model and a fallback hardware environment for executing the task if the designated machine-learning model or designated hardware environment are unavailable. Moreover, in one or more embodiments, the disclosed systems can pause and initiate tasks based on bandwidth availability.
Owner:DROPBOX INC

Extension of network control system into public cloud

Some embodiments provide a method for a first data compute node (DCN) operating in a public datacenter. The method receives an encryption rule from a centralized network controller. The method determines that the network encryption rule requires encryption of packets between second and third DCNs operating in the public datacenter. The method requests a first key from a secure key storage. Upon receipt of the first key, the method uses the first key and additional parameters to generate second and third keys. The method distributes the second key to the second DCN and the third key to the third DCN in the public datacenter.
Owner:VMWARE INC

Data migration method, access processing method, data migration system and electronic equipment

The embodiment of the invention provides a data migration method, an access processing method, a data migration system and electronic equipment. The data migration method is applied to a switch, the switch communicates with a plurality of memory devices in a memory pool based on an interconnection protocol, and each memory device is provided with a plurality of memory areas. When data migration needs to be carried out on a first memory area in a memory pool, a switch sequentially migrates a plurality of data pages obtained by dividing data in the first memory area into available second memory areas according to pages, and in the migration process, a page bitmap determined for the first memory region is utilized to record migration states of the plurality of data pages. By adopting the scheme, the problem of large-area fault of the server caused by the memory problem can be effectively solved, the page-level granularity is adopted to migrate the data in the first memory region page by page in the migration process, the access of the server to the first memory region is continuously processed in the migration process, and the influence on the performance of the server is reduced.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Flash memory switching starting method, system and device, electronic equipment and medium

The invention discloses a flash memory switching starting method, system and device, electronic equipment and a medium, and relates to the technical field of servers, and the method comprises the steps of obtaining address information in a memory in response to completion of power-on initialization of a controller; historical starting information is determined according to the mapping relation between the address information and the flash memories, and the historical starting information comprises at least one of non-starting, first flash memory starting and second flash memory starting; determining a target flash memory corresponding to the address information based on the historical startup information; the target flash memory is used for starting the basic input / output system, so that the technical problem that firmware of the basic input / output system is started from the faulty flash memory at first and then switched to the non-faulty flash memory after the controller receives the switching mark every time when the basic input / output system is started is solved, and the aim of preventing the firmware from being started from the faulty flash memory every time when the basic input / output system is started is achieved. And the starting time is reduced.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Totally-enclosed line subway signal fault diagnosis system

The invention discloses a fully-enclosed line subway signal fault diagnosis system, and relates to the technical field of subway signals, the system comprises a communication monitoring module, a fault identification module, a self-healing control unit, a redundancy communication management module and a data recording module, and the communication monitoring module further comprises a temperature adaptation logic module and a feedback unit. Through the collection, classification and feature extraction technology of historical train operation data, the difference and generality between different signal transmission nodes are analyzed, an abnormal feature library is formed and stored and recorded in a data recording module, then when the system operates, abnormal features are matched according to actual data, corresponding node regulation and control instructions are dynamically generated, and the corresponding node regulation and control instructions are sent to the system. The instructions can actively optimize the signal transmission quality and reduce or completely solve the risk of instant communication interruption, information is fed back in real time through the fault recognition module, and the redundancy control system is switched to achieve smooth transition when the main system fails.
Owner:SHENYANG METRO CO LTD

Fault module restart system and power supply system

The invention discloses a fault module restarting system and a power supply system, and relates to the technical field of servers, the fault module restarting system comprises a first controller, a power supply module, a serial port fuse module, a fan fuse module and a short circuit detection circuit, performing short circuit detection on a serial port device corresponding to the serial port fuse module and a fan corresponding to the fan fuse module, and sending a short circuit detection signal to a first controller through a first short circuit detection output end or a second short circuit detection output end; the first controller is used for responding to the short circuit detection signal to determine that the system breaks down, determining the fault type of the system, and controlling the power module or the first fuse module to be restarted according to the fault type. The reliability of the service provided by the server can be improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Intelligent database switching method and device, computer equipment and storage medium

The invention discloses an intelligent database switching method and device, computer equipment and a storage medium, belongs to the technical field of big data, and is applied to a financial database downtime processing scene. According to the method, the node state of the database is monitored in real time, multi-dimensional load evaluation is combined, the optimal main and standby database combination is automatically selected, second-level fault sensing and rapid switching are achieved, and a minute-level time window needed by traditional database recovery is remarkably shortened. The database identifier is embedded in the service main key, so that the request can accurately position the corresponding database node, complex cross-database query and routing overhead are avoided, and the system response efficiency is improved. In addition, through data seamless transmission and automatic abnormal switching between the main database and the standby database, the risk of service interruption caused by database faults is reduced. According to the scheme, the high availability and the fault recovery speed of the database are improved, repeated construction of a multi-product-line database cluster is reduced through resource integration, and a large amount of hardware and operation and maintenance cost are saved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Cloud edge collaborative edge end device operator hot update method, system and device, and medium

The invention relates to the technical field of edge computing and artificial intelligence model updating, and provides a cloud edge collaborative edge end equipment operator hot updating method, system and device and a medium. The method comprises the steps that a cloud detects a new version of an operator and generates an incremental update package between the new version and the old version; the edge node pulls the incremental update package and performs multiple security verification on the incremental update package; a double-instance inference engine is deployed in the edge node, operators passing verification are preloaded to a standby engine, and hot switching from an operation engine to the standby engine is achieved through a state synchronization mechanism; monitoring the running state of the new engine after switching, if the running state is abnormal, triggering a rollback mechanism, and switching back to the original engine; and optimizing edge node resources, cleaning old version operators, and reporting an update state and a resource use condition to the cloud. Core mechanisms such as differential updating, double-instance engine hot switching, multiple safety verification and dynamic resource scheduling are fused, and therefore efficient, safe and non-perceptual operator updating is achieved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Fault prediction and self-repairing method and device, electronic equipment and storage medium

The invention discloses a fault prediction and self-repairing method and device, electronic equipment and a storage medium, and relates to the technical field of computers. According to the method, the function of dynamically monitoring various data of processor hardware for the target node can be realized, the dynamically monitored hardware state data is input, the fault probability is predicted by a method of weighting and combining key indexes through a bidirectional long-short-term memory model and an attention mechanism, the possible faults are intelligently predicted, and the fault prediction efficiency is improved. And when the target node has a test fault in advance, the target node is repaired in a gradual load reduction mode, so that the effects of real-time monitoring, accurate prediction and rapid self-regulation are achieved. Manual intervention is reduced, the intelligent decision-making capability is achieved, operation and maintenance automation and intelligentization are achieved, resource self-adaptive repairing can be integrated, the self-adaptive capability is improved, and the complex scene fault sensing capability is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Increased replication for managed volumes

A scale-out computing cluster may include a large number of computing servers and storage devices. In order to provide high reliability, the computing cluster must be able to handle failures of individual devices. Reliability of the computing cluster may be improved by providing a standby server for each active server in the computing cluster. If any active server fails, the corresponding standby server is activated. The failed server may be brought back online or replaced, at which time the restored server becomes the standby server for the now-active original standby server. During the restoration period, if any other active server fails, the standby server for that active server is immediately activated. As a result, the recovery ability of the computing cluster is only challenged if both servers of an active / standby pair fail during the restoration period, substantially improving reliability.
Owner:SAP SE

Techniques for maintaining data consistency during disaster recovery

Techniques are described for maintaining data consistency when failure events occur during file storage replications between file systems in different cloud infrastructure regions. In certain embodiments, two generation numbers (or different identifications) are assigned to two groups of processed B-tree key-value pairs, one before and one after a failure event, within a key range. In some embodiments, the two generation numbers are assigned to a group of B-tree key-value pairs processed by a failed thread and another group of B-tree key-value pairs processed by a substitute thread taking over the failed thread to avoid potential data corruption.
Owner:ORACLE INT CORP

High-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) coprocessing architecture of audio frequency integrated signal processor

The invention discloses a high-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) cooperative processing architecture of an audio frequency integrated signal processor, belonging to the technical field of audio frequency integrated signal processing, the architecture is based on a dynamic priority scheduling model and realizes efficient collaboration of a CPU and a GPU through hierarchical resource management and protocol level optimization, a hardware layer adopts a multi-GPU cluster and a distributed storage node, and the CPU-GPU cooperative processing architecture has the advantages that the multi-GPU cluster and the distributed storage node are integrated; high-concurrency task processing is supported; the transmission layer fuses RapidIO and an Ethernet protocol, and adapts to a short frame control signal and a long packet data stream respectively; and the application layer calculates task resource demands through dynamic priority weights, and ensures low time delay of key tasks in combination with a normalized allocation algorithm. According to the architecture, in an audio signal processing scene, the task scheduling efficiency is improved by 40%, the short frame transmission delay is as low as 0.5, the throughput of a long data stream reaches 100 Gbps, and the requirements for real-time performance and calculation precision in a complex acoustic environment can be met.
Owner:CHINA SHIP DEV & DESIGN CENT +1

Server fault processing method and system, electronic equipment, computer storage medium and computer program product

The embodiment of the invention provides a server fault processing method and system, electronic equipment, a computer storage medium and a computer program product. The server fault processing method comprises the following steps: monitoring the working states of at least two processor systems in a target server when accessing respective mounted local storage devices, each processor system comprising computing resources deployed to a virtual machine of the processor system, the local storage device mounted on each processor system comprises storage resources deployed to the virtual machine of the processor system; if the working state of the first processor system in the at least two processor systems indicates that the first processor system has a system fault, migrating a virtual machine in the first processor system to a second processor system in the at least two processor systems, and switching the local storage device mounted on the first processor system to be mounted on the second processor system.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD