Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

939 results about "Fault tolerance" patented technology

Fault tolerance is the property that enables a system to continue operating properly in the event of the failure of (or one or more faults within) some of its components. If its operating quality decreases at all, the decrease is proportional to the severity of the failure, as compared to a naively designed system, in which even a small failure can cause total breakdown. Fault tolerance is particularly sought after in high-availability or life-critical systems. The ability of maintaining functionality when portions of a system break down is referred to as graceful degradation.

Power grid equipment state sensing driving dynamic response method based on Internet of Things technology

The invention discloses a power grid equipment state sensing driving dynamic response method based on the Internet of Things technology, and relates to the technical field of power system automation and informatization, and the method comprises the following steps: S001, collecting original multi-dimensional state signals of a plurality of sensors deployed in a power grid equipment state sensing channel, constructing an electromagnetic disturbance recognition model, and carrying out the recognition of the original multi-dimensional state signals; frequency domain and time domain feature extraction is carried out on the signals, and a feature comparison parameter set used for distinguishing electromagnetic interference and real faults is generated. Frequency domain and time domain features are extracted through an electromagnetic disturbance recognition model, dynamic threshold judgment and response strategy adjustment are achieved in combination with multi-source sensing data, interference and faults can be accurately distinguished, early warning and protection logic can be corrected in real time, protection actions can be accurately triggered, closed-loop control is achieved, the delayed fault tolerance and multi-source verification capacity is achieved, and the method is suitable for large-scale popularization and application. The false operation rate and the false stop risk are effectively reduced, and the intelligence, the safety and the stability of power grid operation are improved.
Owner:GUANGDONG POWER GRID CO LTD INFORMATION CENT

Cloud-edge collaborative production and manufacturing management system based on artificial intelligence

The invention relates to the technical field of artificial intelligence, in particular to a cloud edge collaborative production manufacturing management system based on artificial intelligence, which comprises a task scheduling optimization module, a resource scheduling and load balancing module, a production efficiency evaluation module, a dynamic load adjustment module and a fault tracing analysis module. According to the invention, by accurately analyzing the task dependency relationship, optimizing the task scheduling strategy and improving the response sequence and delay control of task execution, the processing bottleneck problem caused by data transmission delay is avoided, the task execution period is flexibly adjusted, the high-frequency tasks in the production process are effectively managed, and the load unevenness caused by fluctuation is reduced; the method can accurately identify and regulate the production bottleneck, improve the adaptive capability of the production process, accurately identify the potential fault source, shorten the fault diagnosis time, reduce the complexity, more rapidly position the problem, improve the stability, efficiency and fault tolerance of the production system, and integrally enhance the self-optimization capability of the production system.
Owner:HUICHENG DAGONG TECH HENAN CO LTD

Intelligent agent collaborative question-answering method and system based on context awareness and authority control

The invention provides an agent collaborative question-answering method and system based on context awareness and authority control, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring user query, and triggering a user context manager to generate a session state user context object containing a permission label and preference; the central scheduler performs task decomposition based on the query intention and the context, dynamically selects an intelligent agent through a multi-dimensional scoring model and binds a knowledge source; the authority-aware agent execution engine executes triple authority verification and intermediate result desensitization processing in the whole process, and supports agent fault tolerance and audit log recording. The method realizes unification of knowledge security and access flexibility, supports multi-scene adaptive question and answer, and has good expandability and fault-tolerant capability.
Owner:NANJING SHENYE INTELLIGENT SYST ENG

Computing power resource dynamic scheduling system integrating environmental perception and power self-balancing

The invention discloses a computing power resource dynamic scheduling system integrating environmental perception and power self-balancing, and relates to the technical field of computing power resource scheduling. According to the system, multi-dimensional data inside and outside a specified computing power facility operated by a server cluster are collected in real time through a data collection and environment perception fusion module, and dynamic / static fusion processing is carried out on the multi-dimensional data and current environment perception parameters, so that comprehensive perception of the environment and the operation state is realized; then, resource configuration is optimized by combining a combined scheduling and self-power balance control module with a fusion processing result; finally, a hierarchical scheduling scheme is generated through a hierarchical scheduling and multi-scene adaptation verification module, and whether the generated hierarchical scheduling scheme adapts to the current scene or not is judged by combining progressive derating judgment and error-tolerant rate verification, so that dynamic linkage regulation and control of the load, the environment and the power are realized, the problem of unbalanced power distribution in computing power scheduling is effectively solved, and the service life of the computing power scheduling system is prolonged. And the energy efficiency ratio and the environment adaptive capability of the server cluster are improved.
Owner:BEIJING AEROSPACE STAR BRIDGE TECH CO LTD

Charging pile liquid cooling module intelligent cooperative control system based on industrial Ethernet

The invention discloses a charging pile liquid cooling module intelligent cooperative control system based on an industrial Ethernet, and relates to the technical field of liquid cooling control. The charging pile liquid cooling module intelligent cooperative control system based on the industrial Ethernet comprises a liquid cooling data acquisition and preprocessing unit used for acquiring and preprocessing liquid cooling cooperative operation data; the thermal load sensing and intelligent liquid cooling regulation and control unit is used for dividing independent cooling branches, evaluating the thermal load degree of each cooling branch, judging the real-time thermal load state and generating and executing a control instruction; the anomaly detection and cooling redundancy control unit is used for evaluating the thermal response deviation condition of a cooling branch and judging whether a double-ring redundancy pipeline is started or not; and the fault root cause judgment and recovery enhancement unit is used for evaluating the component fault risk degree in the abnormal branch and executing fault-tolerant control and branch recovery strategies. The problem that a whole set of liquid cooling system is shut down due to the fact that local fault tolerance cannot be achieved when a local branch breaks down in a traditional pipeline structure is solved.
Owner:TIANJIN TIER TECHNOLOGY CO LTD

Anti-quantum fuzzy keyword processing method and system and electronic equipment

The invention provides an anti-quantum fuzzy keyword processing method and system and electronic equipment, and relates to the technical field of networks and security. The method comprises the following steps: a client performs wildcard character extension on a keyword set based on a target shared key, generates an encryption index, performs authentication encryption on file identifiers by using the target shared key, forms an encrypted file identifier set, constructs an index table, and uploads the index table and the encrypted file set to a cloud server. The server generates a second wildcard character set according to the query keyword and a fault-tolerant threshold value, generates a trap door set based on the same target shared key and a pseudo-random function and sends the trap door set to the cloud server, and the cloud server traverses the index table, compares the index table with the trap door set and sends the trap door set to the server; and finding out the matched target encryption index and the associated target encryption file identifier and returning the matched target encryption index and the associated target encryption file identifier to the server, and decrypting and acquiring the target file from the cloud server by the server. In this way, high-safety, high-efficiency and extensible privacy protection search service is achieved.
Owner:中电信量子信息科技集团有限公司

Intelligent monitoring system for bridge rotation

The invention provides a bridge swivel intelligent monitoring system, belongs to the field of bridge engineering monitoring, and is used for solving the problems of poor multi-form adaptation, low calculation precision, insufficient extreme pre-judgment, weak fault tolerance and no long-term evolution ability of a traditional bridge swivel monitoring system in related technologies. The system comprises a multi-form identification module, a multi-field coupling calculation module, a digital twinborn mapping module, a fault-tolerant cooperation module and a full-cycle evolution module, all the modules cooperate to achieve accurate identification of a rotation form, nonlinear calculation of core parameters, generation of extreme scene samples, multi-fault fault tolerance and full-life-cycle parameter optimization, and horizontal rotation, vertical rotation and combined rotation scenes can be covered. The monitoring precision and the system robustness are improved, the bridge swivel construction safety is guaranteed, and the intelligent level is improved.
Owner:CHINA CONSTRUCTION SIXTH ENGINEERING DIVISION CO LTD +1

FTU-based power distribution network fault positioning method and system

The invention discloses an FTU-based power distribution network fault positioning method and system, and belongs to the technical field of power distribution automation, and the method comprises the steps: constructing a space-time correlation feature matrix according to a transient current sequence and a voltage drop sequence during a fault period, and extracting the convolution features of a graph to obtain a fault feature graph containing the fault correlation degree between nodes; according to a static topological structure in the power distribution information model, virtual impedance is calculated based on the fault feature graph, and network equivalent topology is dynamically identified to obtain a dynamic virtual topological structure graph; and based on the dynamic virtual topological structure diagram structure and the fault feature diagram, performing fault section confidence competing decision through the intelligent agent unit corresponding to each FTU based on an incomplete information game, and outputting a fault section positioning result. The power distribution network fault positioning method solves the problems that a traditional power distribution network fault positioning method is insufficient in positioning accuracy and poor in fault tolerance and excessively depends on centralized processing and global information synchronization when information is incomplete, fault features are complex and network topology dynamically changes.
Owner:HONGHE POWER SUPPLY BUREAU OF YUNNAN POWER GRID

Database access parameter real-time cooperative processing method and device based on multi-level cache

The invention discloses a database access parameter real-time cooperative processing method and device based on multi-level cache, and belongs to the technical field of computer data caching. The method comprises the steps that a version number management table is configured; establishing a four-level cache architecture of a transaction-level cache layer, a process-level cache layer, a distributed cache layer and a database cache layer; receiving a transaction request, creating a transaction level cache layer and loading current version number information; comparing the version number in the transaction level cache layer with the version number in the process level cache layer; according to the version number comparison result, obtaining parameter data from the corresponding cache level according to a preset cache access priority strategy; and when the parameter change is detected, updating the corresponding version number in the version number management table, and updating the cache layer data as required. According to the method, real-time global effectiveness of parameter change is realized, parameter consistency of in-transit transaction is guaranteed, system processing performance and throughput are improved, system reliability and fault-tolerant capability are enhanced, and system resource utilization rate is optimized.
Owner:SHANDONG CITY COMMERCIAL BANK COOP ALLIANCE CO LTD

Variable constraint control method for stage equipment based on risk perception and dynamic security domain

The invention belongs to the technical field of stage equipment boundary safety protection control, and particularly relates to a variable constraint control method for stage equipment based on risk perception and a dynamic safety domain. The characteristic that stage equipment is usually in a fixed application scene in the performance process is utilized, the inherent safety level and active protection capacity of a stage equipment system are greatly improved through tight combination of real-time collection and variable constraint, and therefore safer, more accurate and more smooth control is achieved. According to the method, a construction algorithm is embedded in software, a dynamic region division and segmentation threshold adjustment strategy is utilized, and a multi-stage braking redundancy cooperation mechanism is triggered, so that the anti-interference capability and fault tolerance performance of the control system are effectively enhanced, the implementation cost is low, a large amount of manpower and material debugging can be saved, and the system is suitable for large-scale popularization and application. And a high-precision and high-robustness safety control scheme is provided for large stage machinery such as a seat vehicle platform and a rotating stage.
Owner:BEIJING BEITE SHENGDI TECH DEV CO LTD

Heavy-load train automatic driving curve planning method, electronic equipment and storage medium

The invention discloses a heavy haul train automatic driving curve planning method, electronic equipment and a storage medium. The method comprises the steps that a DP state transition table is loaded, and a plurality of initial states, a plurality of actions capable of being adopted in each initial state and a transition state capable of being reached by each action are defined in the DP state transition table; acquiring a protection curve as a speed limit line, and drawing according to the protection curve to obtain a guide line; performing curve planning by taking N DP step lengths as a cycle, for each DP step length, sequentially traversing each migration state corresponding to the initial state of the DP step length from top to bottom in the DP state migration table, and performing judgment based on a speed limit line and a guide line to obtain respective target migration states of the N DP step lengths; and generating a speed-mileage planning curve of the train according to the final speed of a differential interval obtained by performing dynamic differential calculation by adopting the action corresponding to the target migration state of each DP step length. According to the invention, the planning success rate and fault tolerance can be improved.
Owner:CASCO SIGNAL LTD

Data storage method and device based on partition identifier mapping, equipment and medium

The invention relates to the technical field of distributed storage, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a data storage method, device, equipment and medium based on partition identifier mapping, and the method comprises the following steps: receiving a file and dividing the file into data blocks to generate block identifiers; generating a partition identifier based on the file identifier and the block identifier through hash mapping; creating a partition disk pack mapping table and writing an initial relationship; querying the mapping table to obtain a disk group, and writing the data block into a physical disk to generate a copy; monitoring the health of the disk, keeping the partition identifier and the block identifier unchanged when a fault occurs, and updating the mapping relation to a new disk group; and receiving a read-write request of the target partition identifier, querying the mapping table to obtain the target disk pack, and executing access. According to the method, the partition identification and the block identification are kept unchanged, fast fault tolerance is achieved only by updating the mapping relation, metadata updating expenditure is reduced, bottom layer change is shielded through partition identification routing, and access transparency and high availability are achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Reusable rocket attitude control system redundancy verification and fault tolerance evaluation method

The invention relates to the technical field of verification and evaluation, in particular to a reusable rocket attitude control system redundancy verification and fault tolerance evaluation method, which comprises the following steps: performing multi-dimensional combination on disturbance units according to disturbance types, action time periods and intensity grades to form a disturbance combination set, constructing a group of induction task profiles covering different attitude control abnormal situations; recording a switching state of an internal control path of the attitude control system, a redundant component intervention time sequence and a function takeover sequence in real time, and constructing a dynamic response path map of a redundant link; and generating a response integrity index and a fault tolerance conflict degree criterion of redundancy configuration as a fault tolerance capability evaluation result of the attitude control system. Compared with an existing method which only verifies individual static disturbance scenes, the method has the advantages that systematic design and composite excitation of disturbance conditions can be realized, scene complexity and coverage of redundancy verification are enhanced, and authenticity and representativeness of fault-tolerant testing are improved.
Owner:BEIJING YIZHUANG REUSABLE ROCKET TECHNOLOGY INNOVATION CENTER CO LTD

Multi-target collaborative liquid cooling system fault prediction and self-healing control method and energy storage heat management system

The invention relates to the technical field of cooling system fault prediction, in particular to a multi-target collaborative liquid cooling system fault prediction and self-healing control method and an energy storage thermal management system.The method comprises the steps that 1, a system health state database is constructed, and data in the database is dynamically updated; step 2, quantizing a parameter change rule based on calculus, constructing a mathematical model of core parameters of the liquid cooling system through a heat balance equation, and quantizing correlation between parameter change and a system state; and step 3, based on the mathematical model in the step 2 and a machine learning algorithm, analyzing the fused data set to realize fault prediction and graded early warning. And 4, according to the fault type in the step 3, by solving a dynamic control equation, actuator adjustment parameters are automatically calculated, and self-healing operation is achieved. According to the method, multi-source data fusion and fault prediction are deeply combined, a self-healing strategy of stage processing and redundant fault tolerance is designed, and high-precision early warning and continuous operation of the system in a fault state are achieved.
Owner:NEW UNITED RAIL TRANSIT TECH

Data sharing and storage model based on double-layer block chain

The invention provides an Internet of Vehicles data sharing and storage model based on a double-layer block chain, and aims to solve the problems of low data interaction efficiency, poor expansibility, insufficient security and the like in the traditional Internet of Vehicles. According to the model, road side units are grouped geographically, an optimized Raft protocol is adopted in each group to realize rapid consensus, and an improved PBFT protocol is adopted among the groups to ensure global consistency. Self-adaptive leader election, batch processing and assembly line mechanisms are introduced into the Raft protocol, so that the throughput and response speed of the system are improved; a reputation value-based VRF main node election mechanism and a BLS aggregation signature technology are introduced into the PBFT protocol, so that the communication complexity is reduced, and the system security and fault-tolerant capability are enhanced. According to the method, the fault-tolerant rate and the message complexity of the system are analyzed theoretically, the advantages of the system in the aspects of throughput, time delay and fault-tolerant performance are verified through simulation experiments, and the method is suitable for large-scale and high-concurrency car networking application scenes.
Owner:BEIJING TECH & BUSINESS UNIV

Distributed big data real-time processing and analysis system

The invention discloses a distributed big data real-time processing and analysis system, which belongs to the technical field of big data real-time processing, and comprises a multi-dimensional data flow acquisition module, an elastic resource scheduling module, an increment state management module, an intelligent fault-tolerant coordination module and a real-time analysis output module, refined dynamic scheduling of computing resources is realized through an adaptive weight adjustment strategy, and the resource utilization rate is improved by more than 40%. Through a fine-grained increment state management technology, only change data is recalculated, the processing delay is reduced by 60%, the throughput is improved by 150%, and through an intelligent fault-tolerant coordination mechanism, the fault recovery time is controlled within 200ms, and the availability reaches 99.99%. The problems of inflexible resource scheduling, low data processing efficiency and weak fault-tolerant capability in the prior art are solved, efficient, low-delay and high-reliability processing of mass real-time data streams is realized, and the method is suitable for application scenes such as financial risk control, telecommunication monitoring and Internet of Things analysis.
Owner:HOHAI UNIV

BMS (Battery Management System) double-partition collaborative flashing method and system and electronic equipment

The invention discloses a BMS (Battery Management System) double-partition collaborative flashing method and system and electronic equipment. Relates to the field of program flashing. A BMS storage area is divided into a first bootstrap program area, two independent bootstrap program partitions and two independent application program partitions, an effective flag bit and a program version number are set for each partition, and state judgment of a partition type arbitration mechanism and a program flashing flag bit is combined, so that the application program is flashed and written, and the application program is flashed and written. A complete process from system starting, partition selection and flash control is constructed. According to the scheme, independent dual-backup management of the bootstrap program and the application program is realized, and when any flash fails or the program is abnormal, a runnable stable version is kept all the time; the technical problems of system paralysis, rollback mechanism deficiency, irreversible upgrading process and the like caused by flashing failure in a traditional single-partition or single-side dual-backup architecture are effectively solved, and the reliability, the safety and the fault-tolerant capability of BMS software updating are remarkably improved.
Owner:FARASIS TECH (GANZHOU) CO LTD

Distributed file management system and method

The invention discloses a distributed file management system and method, and belongs to the technical field of computer software, the system adopts a tree hierarchical structure to organize files, and the system comprises an application layer used for receiving file operation requests from a plurality of independent clients, and each request carries a unique application identifier; the file system layer comprises a configuration module, an operation module, a metadata management module, a day module, a timed task module and a gateway module, and the request is routed to a corresponding logic storage partition in the data layer according to an application identifier; the data layer comprises a database cluster and an object storage cluster, and performs persistent storage and management of logic isolation on metadata and file data according to application identifiers; and the infrastructure layer provides an extensible server and network resources. Based on a distributed architecture, load balancing and efficient access are realized through data storage partitioning, a metadata management and index technology and a client load balancing technology, and the expandability, fault tolerance and performance of the system are improved.
Owner:JIANGSU SECURITIES

PCIe device function number border crossing prevention method supporting ARI

The invention relates to the technical field of chip packaging, in particular to a PCIe (peripheral component interface express) equipment function number border crossing prevention method supporting ARI (automatic repeat interface). The hardware layer protection process is as follows: S1-1, hot plugging; s1-2, if it is detected that the equipment declarates to support ARI, the PCIe controller hardware circuit immediately sets an equipment number register to zero; s1-3, only allowing the PCIe controller of the system to read the configuration space unidirectionally during the low level period of the reset signal PERST #; the interception process of the driving layer is as follows: S2-1, in an equipment enumeration stage, an operating system drives to dynamically read an ARI capability register; s2-2, if the equipment declarates to support ARI, forcing an equipment number register to be equal to 0, and releasing an original equipment number bit space for function number expansion; s2-3, border crossing access is blocked; and S2-4, synchronizing the state register. Compared with the prior art, correctness of device function numbers is guaranteed through two dimensions of hardware register state solidification and driver layer access truncation, and'perception-free 'fault tolerance of the PCIe device under ARI / traditional mode switching is achieved for the first time.
Owner:CHIPMOS TECHNOLOGIES (SHANGHAI) LTD

Control method of double-arm robot, double-arm robot and storage medium

The embodiment of the invention discloses a control method of a double-arm robot, the double-arm robot and a storage medium. The method comprises the steps that the first pose deviation of the tail end of a first mechanical arm relative to a target object and the second pose deviation of the tail end of a second mechanical arm relative to the target object are determined, and a first cooperation parameter and a second cooperation parameter which are opposite in number are generated based on the first pose deviation and the second pose deviation, and the first cooperative parameter and the second cooperative parameter are input into the virtual positive dynamic model, a first joint acceleration and a second joint acceleration are obtained, and the first mechanical arm and the second mechanical arm are controlled based on the first joint acceleration and the second joint acceleration. The first cooperative parameter and the second cooperative parameter provided by the embodiment of the invention are equivalent to a virtual spring, and the virtual spring has an elastic characteristic and a force buffering capability, can be matched with the requirements of a double-arm cooperative task on compliant fault tolerance and robust interference resistance, and is beneficial to improving the double-arm cooperative capability, compliance and robustness of the double-arm robot.
Owner:PAXINI TECHNOLOGY (SHENZHEN) CO LTD

KV-cache streaming for improved performance and fault tolerance in generative model serving

A method of serving a generative transformer model includes determining a batch size to use in processing inference requests and allocating at least one prompt pipeline and at least on token pipeline to the generative transformer model to process the batch of inference requests. The number of prompt pipelines and the number of token pipelines, and the depths of the pipelines are determined based on the batch size, an average prompt length, a cache requirement per stage, and a memory footprint of model weights for the generative model per stage using a resource allocator component of the model serving system. Cache streaming is used to stream prompt cache from prompt pipelines to token pipelines to generate tokens. Cache streaming involves gather-copy operations which may be performed using compute kernels.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Large model-based data security risk automatic research and judgment processing system and method

The large-model-based data security risk automatic research and judgment processing system comprises a fusion research and judgment module which is used for carrying out comprehensive research and judgment on the risk level, attack intention and potential influence of a security event based on security data in combination with deep learning and a large-scale language model, and transmitting a research and judgment conclusion to an automatic processing module; the business process disassembling and analyzing module is used for automatically discovering and modeling based on the security data through process mining and flow analysis, perceiving a business process and a dependency graph associated with a security event in real time, evaluating potential influences of different disposal measures on business continuity, and transmitting business influence evaluation to the automatic disposal module; and the automatic disposal module is used for selecting and executing an optimal risk disposal strategy from the security strategy library by adopting a distributed architecture of an AI intelligent agent according to the research and judgment conclusion and the business influence evaluation. The accuracy of alarm study and judgment is improved, 'alarm fatigue 'is relieved, and the method has high elasticity, high fault tolerance and strong cooperative capability.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Equipment networking method and system based on distributed soft bus

The invention provides an equipment networking method and system based on a distributed soft bus, and the method comprises the steps: starting a distributed soft bus module and carrying out the discovery of adjacent equipment, and obtaining a trusted equipment set after identity verification; a point-to-point communication link between the devices is established according to the trusted device set, and a distributed mesh interconnection topology is formed; performing routing calculation and data forwarding on data transmission in the mesh topology to realize end-to-end communication; meanwhile, dynamic changes of equipment in the network are monitored, topology updating is carried out, and the network stability is maintained. The scheme does not depend on the dependence of the traditional star network on the central node, realizes direct interconnection between devices through the distributed soft bus technology, improves the fault-tolerant capability and transmission efficiency of the network, and solves the problems of poor network reliability and limited expansibility in the prior art.
Owner:SHENZHEN HONGYUAN ZHITONG TECH CO LTD

Iterative task execution method and system based on dynamic feedback and causal fault tolerance

The invention discloses an iterative task execution method and system based on dynamic feedback and causal fault tolerance, belongs to the technical field of artificial intelligence and data analysis, and remarkably improves the execution success rate and system robustness of a complex data analysis task by fusing multi-modal perception, dynamic feedback and causal reasoning. According to the method, the disassembling deviation caused by information splitting is avoided, the task complexity is quantified through the graph neural network, self-adaptive task disassembling is achieved, a dynamic feedback mechanism is combined, the task structure is corrected in real time in the execution process, and insertion, combination and sequence adjustment of subtasks are supported. A causal fault-tolerant strategy based on a historical failure trajectory is introduced, a high-risk path can be pre-judged, defensive operation can be automatically embedded, and error propagation is blocked. The whole process execution track is stored in a memory bank and is used for continuously optimizing the model and the strategy to form the closed-loop learning ability.
Owner:BEIJING SILICON MOBILE TECHNOLOGY CO LTD

LoRA training fault tolerance method based on state awareness

The invention relates to the technical field of machine learning training fault tolerance, and discloses a LoRA training fault tolerance method based on state awareness. The method comprises the steps that training indexes are collected through a sensor network, a time sequence knowledge base is constructed, and sampling frequency is adjusted in a self-adaptive mode; historical data is utilized to generate a state change trend baseline, and abnormity is preliminarily identified through deviation degree comparison. And for the abnormal behavior, performing verification and risk assessment by fusing the gradient frequency domain characteristics and the loss curve form of the abnormal behavior. Encoding the anomalies into a high-dimensional point set, analyzing and identifying a continuous homologous structure of the high-dimensional point set by adopting topological data, and judging a systematic gradient anomaly mode according to topological characteristics; and deducing a compensation coefficient according to geometric attributes of the topological characteristics, calibrating a risk estimation value, and dynamically reconstructing a monitoring strategy. According to the method, systematic training faults can be accurately identified from a data structure level, adaptive intelligent fault-tolerant control is realized, and the reliability and efficiency of a training process are improved.
Owner:HUNAN PAN CLOUD DATA CO LTD

Power business anti-quantum cryptography migration progressive control method based on risk perception

The invention provides a power business anti-quantum cryptography migration progressive control method based on risk awareness, and belongs to the technical field of migration control, and the method comprises the steps: collecting key parameters of all to-be-migrated businesses in a power system, carrying out the standardization preprocessing of the collected key parameters, and constructing a business feature matrix; acquiring relevant parameters of quantum attacks in real time, and calculating a real-time quantum attack risk value of each service to be migrated; based on the service feature matrix and the real-time quantum attack risk value, constructing a dynamic migration rhythm adaptation model, and calculating a migration priority, a migration rate and a migration interval of each service to be migrated; and constructing a migration fault-tolerant and recovery model, monitoring the migration process in real time, calculating a fault influence range and recovery cost based on fault information, and executing corresponding adjustment operation. Accurate, dynamic and high-reliability control of anti-quantum password migration of the power business is realized, safe and stable operation of the power business in the migration process is guaranteed, and the migration efficiency and the anti-quantum attack capability are improved.
Owner:NANJING NANZI DIGITAL SECURITY TECH CO LTD +1

Model distributed training automatic fault tolerance method in large-scale cloud native scene

The invention provides an automatic fault tolerance method for distributed training of a model in a large-scale cloud native scene, and relates to the technical field of data management, and the method comprises the steps: collecting and based on hardware monitoring data of each node in a cluster, scheduling a distributed training task to each node, and starting check point storage; monitoring the running state of each training task, collecting a training log and a chip acceleration platform log, and collecting CPU, memory and acceleration card resource index data of each node and each training task; building a model training fault detection classification model by combining a supervision and machine learning method on the basis of the collected logs and hardware index data; and based on the model training fault detection classification model, judging whether each running training task has a fault and the fault type, and if the node equipment has a fault, rescheduling and loading the latest check point data to finish training task breakpoint continuous training and automatic fault tolerance. According to the invention, model breakpoint continuous training is realized, and the stability and efficiency of model training are improved.
Owner:HANGZHOU HARMONYCLOUD TECH CO LTD

Aviation permanent magnet synchronous motor control architecture based on master-slave redundancy and model prediction

The invention provides an aviation permanent magnet synchronous motor control architecture based on master-slave redundancy and model prediction, belongs to the field of aviation electric propulsion control systems, and aims to solve the problems of extreme environment parameter drift, single-fault shutdown, multi-target control conflict, fault diagnosis lag and the like of an existing system. The control architecture comprises a master-slave dual-redundancy control unit, a wide-temperature-range adaptive model prediction control unit, an electromagnetic interference adaptive suppression unit and a multi-dimensional fault tolerance unit. According to the invention, a DSP + FPGA heterogeneous redundant architecture is matched with two-out-of-three voting logic, so that fault non-perception switching is realized; an extended Kalman filter observer is embedded to update motor parameters on line, and take-off / cruise / landing three modes are preset; interference at the frequency band of 100kHz-10MHz is controlled and suppressed through adaptive notch filtering and a sliding mode variable structure; a long short-term memory network is fused to realize accurate diagnosis of early faults, and the system is adaptive to aviation scenes such as electric general aviation aircrafts and unmanned aerial vehicles.
Owner:TAIHANG NATIONAL LABORATORY

AI service intelligent fault tolerance and degradation method and device based on Java bytecode enhancement technology, equipment, medium and product

The invention discloses an AI service intelligent fault tolerance and degradation method, device and equipment based on a Java bytecode enhancement technology, a medium and a product, and relates to the technical field of AI service integration. According to the method, in a class loading period of Java application starting, a loaded class is scanned to identify target methods marked with declarative fault-tolerant annotations, a bytecode operation framework is utilized to enhance the target methods so as to implant a unified fault-tolerant logic interception entry, and then in a running period, when the target methods are called, the target methods are subjected to fault-tolerant annotation processing. And triggering a fault-tolerant interceptor through the interception entrance to analyze the annotation to obtain a fault-tolerant strategy, allocating an independent fault-tolerant component instance to the fault-tolerant interceptor based on the unique identifier of the method, constructing and executing an enhanced call chain containing a fusing check and retry mechanism, and finally, if the execution of the call chain fails, executing the fault-tolerant interceptor. And if yes, a preset degradation method in the fault-tolerant strategy is automatically executed, so that complete non-intrusive decoupling of AI service fault-tolerant management and service logic can be realized, and the reliability of the system is remarkably improved.
Owner:BEIJING INSIGHT NETWORK CO LTD

CAE workflow control subsystem and method based on cloud computing

The invention discloses a CAE (Computer Aided Engineering) workflow control subsystem and method based on cloud computing, and belongs to the technical field of crossing of computer simulation and cloud computing, and the subsystem comprises a workflow configuration analysis module which is used for converting a CAE workflow configuration file possibly containing cyclic dependence into a directed acyclic graph state machine through strong connectivity component analysis; the workflow state machine operation module is used for dynamically deriving task groups which can be executed in parallel in the next stage in the state machine according to the current state; and the workflow scheduling module is used for submitting the tasks to a cloud solving platform in parallel and realizing fault recovery through state persistence. Through deep fusion of a graph theory method and a cloud computing architecture, automatic scheduling, efficient parallel execution and reliable fault tolerance of the CAE workflow are realized, and the computing resource utilization rate and the system reliability are remarkably improved.
Owner:ANHUI ZHONGAN ZHIQING TECHNOLOGY CO LTD