Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1551results about "Redundant hardware error correction" patented technology

IT asset fault propagation prediction method and system based on dynamic evolution of knowledge graph

The invention discloses an IT asset fault propagation prediction method and system based on dynamic evolution of a knowledge graph, and relates to the technical field of cloud computing and large-scale IT operation and maintenance management. Through an asynchronous message bus and a logic clock, the knowledge graph is updated immediately when resources are abnormal and a scheduling event occurs; the knowledge graph uniformly integrates physical connection, logic dependence and multi-copy redundancy, so that the cross-machine-room asset relationship is clear at a glance. And then, based on a weighted logistic regression model, node features and relation weights in the knowledge graph are fused, the node fault probability is accurately calculated, the limitation of traditional single-dimensional analysis is solved, self-healing operation is supported, end-to-end intelligent operation and maintenance from fault detection to prediction and early warning to closed-loop self-healing are realized, and the fault detection efficiency is improved. The problems that in a cross-machine-room and multi-live-site environment, resource topology is split, real-time state and alarm information cannot be fused with an asset dependence model, and large-scale real-time deployment of a traditional single-dimensional fault analysis and high-complexity prediction algorithm is difficult are effectively solved.
Owner:GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU

Distributed simulation method and system based on containerized deployment and elastic expansion

The invention relates to the technical field of distributed simulation based on containerized deployment and elastic expansion and contraction, and discloses a distributed simulation method and system based on containerized deployment and elastic expansion and contraction. According to the distributed simulation method and system based on containerized deployment and elastic expansion, real-time resource monitoring and historical load trend data of a simulation task are collected, and unified load evaluation is carried out in combination with a simulation calculation complexity parameter and an I / O density parameter; the refined modeling and resource demand pre-judgment of the simulation task are realized, and the accuracy of task scheduling and allocation is improved; modularized deployment of simulation tasks is achieved through subtask segmentation based on minimum executable units and a standard containerization packaging mechanism in cooperation with a container arrangement platform, and then concurrent execution and elastic scheduling in a multi-node environment are supported.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Reservoir dam operation safety sky-ground work intelligent sensing system and operation method

The invention relates to a reservoir dam operation safety sky-land project intelligent sensing system and an operation method, and relates to the technical field of hydraulic engineering safety monitoring. The system is composed of a sky-land water conservancy project integrated monitoring and sensing system, a self-adaptive sampling module, a layered distributed architecture and a software and hardware integrated module, and multi-source data such as deformation, seepage, stress strain, vibration and environmental quantity are cooperatively collected through five dimensions of sky domain, airspace, territory, water domain and work domain. The monitoring frequency is dynamically adjusted by using an adaptive sampling strategy, and data cleaning, standardization, space-time registration and fusion processing are completed through a distributed architecture to generate a high-quality comprehensive data set. The system can realize total-factor and whole-process refined monitoring, effectively eliminates data islands, improves data quality and monitoring efficiency, has high reliability, real-time performance and expandability, and provides powerful data support and decision basis for dam safety assessment and intelligent early warning.
Owner:CHANGJIANG SPATIAL INFORMATION TECH ENG CO LTD (WUHAN) +1

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Large model batch reasoning and data flow optimization system oriented to MOE architecture

The invention relates to the technical field of project management, in particular to a large-model batch reasoning and data flow optimization system oriented to an MOE architecture. According to the method, a collaborative architecture of the request access module, the environment sensing module, the expert routing engine, the resource scheduling module and the dynamic optimization control module is set, the text length and the subject type are extracted by using the request access module, a basis is provided for accurate routing, and the GPU video memory, the I / O bandwidth and the request queue depth are acquired in real time through the environment sensing module, so that the real-time routing is realized. The system load is comprehensively monitored, meanwhile, an expert sub-network is activated through an expert routing engine according to request features, invalid calculation is avoided, weight loading and resource allocation are managed through a resource scheduling module, the I / O bottleneck is reduced, and finally an optimization strategy is intelligently triggered through a dynamic optimization control module based on routing conflict factors. The problems of large reasoning delay fluctuation and unbalanced resource utilization rate mentioned in the background technology are solved, and stable low-delay response and resource collaborative optimization in a high-concurrency scene is realized.
Owner:VIRTAI TECH BEIJING CO LTD

Managing resource constraints in a cloud environment

Techniques for managing resource constraints of a cloud environment are disclosed. A system receives a request to initiate a provisioning process for provisioning a first service in the cloud environment. The system determines a resource constraint associated with a resource that the first service utilizes. Based on the resource constraint, the system determines a set of candidate services that also utilize the resource as candidates for deprovisioning from the cloud environment. The system identifies respective service features of the set of candidate services and generates a ranking of the set of candidate services based on weighting metrics associated with the respective service features. Based on the ranking, the system selects a second service of the set of candidate services for deprovisioning from the cloud environment. The system deprovisions the second service to alleviate the resource constraint and then provisions the first service by executing the provisioning process.
Owner:ORACLE INT CORP

End-to-end restartability of cross-region replication using a new replication

Techniques are described for performing different types of restart operations for a file storage replication between a source file system and a target file system in different cloud infrastructure regions. In certain embodiments, the disclosed techniques perform a restart operation to terminate a current cross-region replication by synchronizing resource cleanup operations in the source file system and the target file system, respectively. In other embodiments, disclosed techniques perform a restart operation to allow a customer to reuse the source file system by identifying a restartable base snapshot in the source file system without dependency on the target file system.
Owner:ORACLE INT CORP

Dynamically selecting artificial intelligence models and hardware environments to execute tasks

The present disclosure relates to systems, non-transitory computer-readable media, and methods for selecting machine-learning models and hardware environments for executing a task. In particular, in one or more embodiments, the disclosed systems select a designated machine-learning model for executing a task based on workload features of the task and task routing metrics for a plurality of machine-learning models. In addition, in one or more embodiments, the disclosed systems select a designated hardware environment for executing the task based on workload features for the task and task routing metrics for a plurality of hardware environments. In some embodiments, the disclosed systems select a fallback machine-learning model and a fallback hardware environment for executing the task if the designated machine-learning model or designated hardware environment are unavailable. Moreover, in one or more embodiments, the disclosed systems can pause and initiate tasks based on bandwidth availability.
Owner:DROPBOX INC

Extension of network control system into public cloud

Some embodiments provide a method for a first data compute node (DCN) operating in a public datacenter. The method receives an encryption rule from a centralized network controller. The method determines that the network encryption rule requires encryption of packets between second and third DCNs operating in the public datacenter. The method requests a first key from a secure key storage. Upon receipt of the first key, the method uses the first key and additional parameters to generate second and third keys. The method distributes the second key to the second DCN and the third key to the third DCN in the public datacenter.
Owner:VMWARE INC

Intelligent database switching method and device, computer equipment and storage medium

The invention discloses an intelligent database switching method and device, computer equipment and a storage medium, belongs to the technical field of big data, and is applied to a financial database downtime processing scene. According to the method, the node state of the database is monitored in real time, multi-dimensional load evaluation is combined, the optimal main and standby database combination is automatically selected, second-level fault sensing and rapid switching are achieved, and a minute-level time window needed by traditional database recovery is remarkably shortened. The database identifier is embedded in the service main key, so that the request can accurately position the corresponding database node, complex cross-database query and routing overhead are avoided, and the system response efficiency is improved. In addition, through data seamless transmission and automatic abnormal switching between the main database and the standby database, the risk of service interruption caused by database faults is reduced. According to the scheme, the high availability and the fault recovery speed of the database are improved, repeated construction of a multi-product-line database cluster is reduced through resource integration, and a large amount of hardware and operation and maintenance cost are saved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Cloud edge collaborative edge end device operator hot update method, system and device, and medium

The invention relates to the technical field of edge computing and artificial intelligence model updating, and provides a cloud edge collaborative edge end equipment operator hot updating method, system and device and a medium. The method comprises the steps that a cloud detects a new version of an operator and generates an incremental update package between the new version and the old version; the edge node pulls the incremental update package and performs multiple security verification on the incremental update package; a double-instance inference engine is deployed in the edge node, operators passing verification are preloaded to a standby engine, and hot switching from an operation engine to the standby engine is achieved through a state synchronization mechanism; monitoring the running state of the new engine after switching, if the running state is abnormal, triggering a rollback mechanism, and switching back to the original engine; and optimizing edge node resources, cleaning old version operators, and reporting an update state and a resource use condition to the cloud. Core mechanisms such as differential updating, double-instance engine hot switching, multiple safety verification and dynamic resource scheduling are fused, and therefore efficient, safe and non-perceptual operator updating is achieved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Techniques for maintaining data consistency during disaster recovery

Techniques are described for maintaining data consistency when failure events occur during file storage replications between file systems in different cloud infrastructure regions. In certain embodiments, two generation numbers (or different identifications) are assigned to two groups of processed B-tree key-value pairs, one before and one after a failure event, within a key range. In some embodiments, the two generation numbers are assigned to a group of B-tree key-value pairs processed by a failed thread and another group of B-tree key-value pairs processed by a substitute thread taking over the failed thread to avoid potential data corruption.
Owner:ORACLE INT CORP

High-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) coprocessing architecture of audio frequency integrated signal processor

The invention discloses a high-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) cooperative processing architecture of an audio frequency integrated signal processor, belonging to the technical field of audio frequency integrated signal processing, the architecture is based on a dynamic priority scheduling model and realizes efficient collaboration of a CPU and a GPU through hierarchical resource management and protocol level optimization, a hardware layer adopts a multi-GPU cluster and a distributed storage node, and the CPU-GPU cooperative processing architecture has the advantages that the multi-GPU cluster and the distributed storage node are integrated; high-concurrency task processing is supported; the transmission layer fuses RapidIO and an Ethernet protocol, and adapts to a short frame control signal and a long packet data stream respectively; and the application layer calculates task resource demands through dynamic priority weights, and ensures low time delay of key tasks in combination with a normalized allocation algorithm. According to the architecture, in an audio signal processing scene, the task scheduling efficiency is improved by 40%, the short frame transmission delay is as low as 0.5, the throughput of a long data stream reaches 100 Gbps, and the requirements for real-time performance and calculation precision in a complex acoustic environment can be met.
Owner:CHINA SHIP DEV & DESIGN CENT +1

GPU cluster, redundancy optimization method of GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a GPU cluster, a redundancy optimization method of the GPU cluster, electronic equipment, a storage medium and a computer program product, the GPU cluster comprises a plurality of GPU cabinets, and each GPU cabinet comprises a plurality of GPU nodes, a GPU extension frame and an interconnection module; for any GPU cabinet, each GPU node in the GPU cabinet comprises a plurality of physical GPUs, and a GPU expansion frame in the GPU cabinet comprises a plurality of standby virtual GPUs; the interconnection module in the GPU cabinet is used for carrying out GPU interconnection on a plurality of GPU nodes and a plurality of GPU expansion frames in the GPU cabinet; and the plurality of standby virtual GPUs in the GPU extension frame in the GPU cabinet are used for providing redundant computing resources for the plurality of physical GPUs in each GPU node in the GPU cabinet. According to the embodiment of the invention, the GPU card-level redundancy capability can be effectively realized.
Owner:MOORE THREADS TECH CO LTD

Digital twinning-oriented multi-modal man-machine interaction method

The invention discloses a digital twinning-oriented multi-modal man-machine interaction method. The method comprises the following steps of: obtaining interactive input of a user on a client for man-machine interaction; when abnormal operation is detected, the human-computer interaction management program adaptively executes corresponding control logic according to a current abnormal mode, an abnormal mode database established by historical data and a network bandwidth; comprising the steps of dynamically adjusting a client interface layout and / or operation logic to replace a standby client, automatically matching an optimal interactive input mode, adjusting an interactive priority and recalculating an optimal interactive input mode. According to the method and the device, the optimal interactive input mode can be switched by adjusting the priority matching of the current interactive input according to the self-adaptive dynamic control client interface layout and / or operation logic, so that the interactive response time is remarkably shortened; therefore, efficient interaction under different working conditions is ensured, unique requirements of different users under complex scenes are met, the satisfaction degree and the interaction efficiency of interaction are improved, and the overall experience and acceptability of the users are improved.
Owner:CHINA RAILWAY CONSTR HEAVY IND

Data storage method and device based on partition identifier mapping, equipment and medium

The invention relates to the technical field of distributed storage, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a data storage method, device, equipment and medium based on partition identifier mapping, and the method comprises the following steps: receiving a file and dividing the file into data blocks to generate block identifiers; generating a partition identifier based on the file identifier and the block identifier through hash mapping; creating a partition disk pack mapping table and writing an initial relationship; querying the mapping table to obtain a disk group, and writing the data block into a physical disk to generate a copy; monitoring the health of the disk, keeping the partition identifier and the block identifier unchanged when a fault occurs, and updating the mapping relation to a new disk group; and receiving a read-write request of the target partition identifier, querying the mapping table to obtain the target disk pack, and executing access. According to the method, the partition identification and the block identification are kept unchanged, fast fault tolerance is achieved only by updating the mapping relation, metadata updating expenditure is reduced, bottom layer change is shielded through partition identification routing, and access transparency and high availability are achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

CXL memory fault tolerance method, and server system, storage medium and electronic device

PCT designated stageWO2025227987A1TransmissionRedundant hardware error correctionMemory faultsMemory footprint
A CXL memory fault tolerance method, and a server system, a storage medium and an electronic device. The method comprises: acquiring parameter values of a group of operating parameters of CXL memory devices in a CXL memory device group, wherein the group of operating parameters are used for representing the operating states of the corresponding CXL memory devices; on the basis of the acquired parameter values of the group of operating parameters, predicting the operating states of the CXL memory devices in the CXL memory device group; and when it is predicted that there is an abnormal memory device operating abnormally in the CXL memory device group, performing controlling to execute a migration operation on memory data in the abnormal memory device, so as to migrate the memory data in the abnormal memory device to a target memory device operating normally in the CXL memory device group. By means of the present application, the problem of CXL memory fault tolerance methods in the prior art of the memory utilization rate of a server being low due to a hot standby memory occupying a server slot is solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Fault switching method and device, electronic equipment and storage medium

The invention discloses a failover method and device, electronic equipment and a storage medium, and relates to the technical field of distributed storage, and the failover method comprises the following steps: determining a first disk resource group corresponding to a first storage node and a second disk resource group corresponding to a second storage node; and establishing a double-control relationship between the first disk resource group and the second disk resource group. And after the establishment of the double-control relationship is completed, carrying out fault switching on different disk resource groups under the double-control relationship based on the first mechanism and the second mechanism. The technical problem that in a distributed storage system, efficient and safe disk resource management and data access control of a fault controller cannot be effectively achieved under a double-controller architecture is solved, and the technical effects of improving rapidness, smoothness and automatic back-switching of fault switching and remarkably improving the high availability of the distributed storage system are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Fault identification and recovery for distributed training

Example embodiments of the present disclosure relate to a method, a device and a non-transitory computer-readable medium for distributed training. The method comprises obtaining, during a distributed training task performed across a plurality of computing nodes, at least one heartbeat message from the plurality of computing nodes, each computing node including multiple GPU workers; detecting, based on the at least one heartbeat message, an abnormal status of the distributed training task; commanding the plurality of computing nodes to run at least one self-check diagnostics test; identifying, based on results of the at least one self-check diagnostics test, at least one faulty node from the plurality of computing nodes; and replacing the at least one faulty node with an equivalent number of heathy computing nodes that have passed the at least one self-check diagnostics test.
Owner:LEMON INC(GB)

Data Processing Method, Switching Board, Data Processing System and Data Processing Apparatus

A data processing method, a switching board, a data processing system and a data processing apparatus are provided. The method includes: receiving a first fault instruction sent by a controller; on the basis of the first fault instruction, acquiring first target data; and transmitting the first target data to a second host system, so as to instruct the second host system to control a second processor to continue to process first data according to the first target data. wherein the second host system is a host system in a normal operating state among a plurality of host systems connected to a CXL switching device, and the second processor is a processor allocated to the second host system by the CXL switching device.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Fault processing method of multi-control storage system and electronic equipment

The invention discloses a fault processing method of a multi-control storage system and electronic equipment, and relates to the technical field of data storage, and the fault processing method comprises the following steps: controlling a cache state machine to execute a first fault processing flow under the condition of identifying that a specified mirror image pair exists in at least one mirror image pair, the specified mirror image pair is a mirror image pair containing two fault controllers; and under the condition that the first fault processing flow is not executed completely and a specified controller is identified to exist, executing a second fault processing flow of the cache state machine based on an execution stage of the first fault processing flow so as to control the multi-control storage system to continue to perform service processing, the problem of long-time service interruption of the multi-control storage system in a fault scene in the prior art is solved, and the service continuity and the system stability are remarkably enhanced.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Dynamically selecting artificial intelligence models and hardware environments to execute tasks

The present disclosure relates to systems, non-transitory computer-readable media, and methods for selecting machine-learning models and hardware environments for executing a task. In particular, in one or more embodiments, the disclosed systems select a designated machine-learning model for executing a task based on workload features of the task and task routing metrics for a plurality of machine-learning models. In addition, in one or more embodiments, the disclosed systems select a designated hardware environment for executing the task based on workload features for the task and task routing metrics for a plurality of hardware environments. In some embodiments, the disclosed systems select a fallback machine-learning model and a fallback hardware environment for executing the task if the designated machine-learning model or designated hardware environment are unavailable. Moreover, in one or more embodiments, the disclosed systems can pause and initiate tasks based on bandwidth availability.
Owner:DROPBOX INC

Proactive input-output failover based on predicting optical transceiver module hardware failures in host bus adapter

Method and apparatus for performing proactive input-output (IO) failover based on predicting optical transceiver module hardware failures in a host bus adapter (HBA) are described. An example method includes registering with an HBA to receive a notification that a first path via the HBA is predicted to fail. The notification is based on a set of parameters of an optical transceiver module coupled to the first path. A set of parameters of the optical transceiver module are monitored in accordance with the registration. At least one of the set of parameter is determined to satisfy a predetermined criteria. Responsive to the determination, the notification that the first path via the HBA is predicted to fail is sent, a second path via the HBA for issuing an IO operation is determined, and the IO operation is issued via the second path.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Methods to maintain read-write consistency and dependent write order consistency within a cross-site storage system

The present storage solution provides an order of operations of a computer-implemented method for performing transient failure handling with an improved application I / O resumption time for a symmetric distributed storage system; an order of operations of a computer-implemented method for performing persistent failure handling with an improved application I / O resumption time for a symmetric distributed storage system; an order of operations of a computer-implemented method for performing transient failure handling with an improved application I / O resumption time to maintain dependent write order consistency for a symmetric distributed storage system; an order of operations of a computer-implemented method for performing secondary side write Op handling to maintain dependent write order consistency for a symmetric distributed storage system; and an order of operations of a computer-implemented method for performing secondary side read Op handling to maintain dependent write order consistency for a symmetric distributed storage system in accordance with some embodiments.
Owner:NETAPP INC

Model distributed training automatic fault tolerance method in large-scale cloud native scene

The invention provides an automatic fault tolerance method for distributed training of a model in a large-scale cloud native scene, and relates to the technical field of data management, and the method comprises the steps: collecting and based on hardware monitoring data of each node in a cluster, scheduling a distributed training task to each node, and starting check point storage; monitoring the running state of each training task, collecting a training log and a chip acceleration platform log, and collecting CPU, memory and acceleration card resource index data of each node and each training task; building a model training fault detection classification model by combining a supervision and machine learning method on the basis of the collected logs and hardware index data; and based on the model training fault detection classification model, judging whether each running training task has a fault and the fault type, and if the node equipment has a fault, rescheduling and loading the latest check point data to finish training task breakpoint continuous training and automatic fault tolerance. According to the invention, model breakpoint continuous training is realized, and the stability and efficiency of model training are improved.
Owner:HANGZHOU HARMONYCLOUD TECH CO LTD

Methods and systems of an all purpose broadband network with publish subscribe broker network

A system including a server and a wireless RF access node connected to a communication network is provided. The server provides a first publish-subscribe broker of one or more publish-subscribe brokers forming part of a publish-subscribe broker network. The server connects to the wireless RF access node. A first entity connects via the wireless RF access node to the first publish-subscribe broker using a unicast IP address and thereafter publishes data packets. A second entity connects to the publish-subscribe broker network via any of the one or more publish-subscribe brokers. The server provides packet distribution services via the first publish-subscribe broker for the first entity, the publish-subscribe broker network routes communications from the first entity to the second entity when the second entity is subscribed; and the data packets published by the first entity are routed through the publish-subscribe broker to which the second entity is connected.
Owner:ALL PURPOSE NETWORKS INC

Vehicle-mounted distributed data storage and management method for intelligent vehicle

The invention discloses a vehicle-mounted distributed data storage and management method for an intelligent vehicle, which comprises the following steps: in a dynamic election mode, acquiring performance parameters of each node by a default main storage node, calculating a comprehensive score, electing a main storage node and an alternative storage node according to the comprehensive score, and taking the rest as slave storage nodes; the main storage node formulates a differential fragmentation storage strategy and a network QoS strategy according to the data multi-dimensional label and distributes the strategies; each node stores collected data in a local designated area according to a strategy, and data aggregation and preprocessing in the nodes are completed; the alternative storage node and the slave storage node converge the preprocessed data to the main storage node at different priorities according to a preset rule; and the main and alternative storage nodes periodically synchronize data. According to the method, the distributed architecture is combined with the intelligent fragmentation, the single-point bottleneck is avoided, the data is stored and processed nearby, the network delay is effectively reduced, and the storage and network resource utilization efficiency is improved.
Owner:WUHAN JIANGXIA CHUNENG AUTOMOBILE TECHNOLOGY R&D CO LTD

Techniques for resource utilization in replication pipeline processing

Techniques are described for ensuring end-to-end fair-share resource utilization during cross-region replication. In certain embodiments, a fair-share architecture is used for communication among pipeline stages performing a cross-region replication between different cloud infrastructure regions. Cross-region replication-related jobs are distributed evenly from a pipeline stage into a temporary buffer in the fair-share architecture, and then further distributed evenly form the fair-share architecture to parallel running threads of next pipeline stage for execute. Techniques for static and dynamic resource allocations are also disclosed.
Owner:ORACLE INT CORP