Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1344 results about "System failure" patented technology

A system failure can occur because of a hardware failure or a severe software issue, causing the system to freeze, reboot, or stop functioning altogether. A system failure may or may not result in an error being displayed on the screen.

Electromechanical system fault pre-diagnosis method and system based on digital twinning

The invention discloses an electromechanical system fault pre-diagnosis method and system based on digital twinning. The method comprises the following steps of obtaining multi-source data in an electromechanical system operation process; preprocessing the acquired multi-source data, wherein the preprocessing comprises data cleaning, normalization processing and feature extraction; and on the basis of the preprocessed multi-source data, an electromechanical system design drawing, a three-dimensional geometric model, material attributes and a kinetic equation are fused, and a digital twin model is constructed. According to the invention, through a digital twin model dynamic calibration and prediction algorithm, early abnormity of the equipment is identified in advance, the fault probability and the residual life are output, and non-planned shutdown is reduced; by constructing a cross-physical domain fault feature system and fusing model simulation and actual measurement data, the potential fault identification accuracy is improved, and the missed diagnosis rate is reduced; by calibrating parameters of the digital twin model in real time, the method adapts to nonlinear changes of equipment, ensures high-fidelity mapping of the model, and improves fault prediction precision.
Owner:CHENGDU TECHNICIAN COLLEGE (CHENGDU VOCATIONAL & TECH COLLEGE OF IND & TRADE CHENGDU ADVANCED TECH SCHOOL CHENGDU RAILWAY ENG SCHOOL)

Integrated test scene prediction method and system based on multi-dimensional data

The invention provides an integrated test scene prediction method and system based on multi-dimensional data, and the method comprises the steps: carrying out the multi-source data collection of an integrated test scene based on a dynamic sampling strategy, obtaining integrated test multi-dimensional data, carrying out the multi-dimensional data preprocessing of the integrated test multi-dimensional data, and obtaining a pre-trained multi-dimensional data analysis model; inputting the pre-processed integrated test multi-dimensional data into the multi-dimensional data analysis model to carry out sub-module data analysis, obtaining a pre-trained intelligent regulation and control model, importing a system diagnosis report and a key index deviation degree report into the intelligent regulation and control model as inputs, generating an intelligent regulation and control decision, and carrying out intelligent regulation and control. By means of the method, system faults, performance bottlenecks and potential problems occurring in the integrated test scene can be accurately predicted, intelligent regulation and control decisions are provided for solving the problems, and accurate system performance evaluation and optimization services can be provided for users.
Owner:SHANGHAI ZEZHONG SOFTWARE TECH CO LTD

Processing environment switching and recovering method and device, equipment and medium

PendingCN121092357AFault responseRecovery methodMulti source data
The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a processing environment switching and recovery method, device, equipment and medium. The method comprises the steps that multi-source heterogeneous data in a main processing environment and a standby processing environment are acquired, and the system fault probability is obtained through multi-model collaborative prediction; a dynamic threshold value is generated in combination with a historical service period mode and a real-time service load, when the fault probability exceeds the threshold value, a switching strategy is generated based on the fault scene knowledge base and the service priority, and flow scheduling between the main processing environment and the standby processing environment is executed; and monitoring the business index of the standby processing environment during the scheduling period, and triggering the fusing rollback when the business index is lower than the health standard. According to the method, the fault identification precision is improved through multi-source data fusion and multi-model prediction, adaptive scheduling is realized in combination with a dynamic threshold and a switching strategy, and fusing rollback is triggered to guarantee high availability and data consistency, so that the continuity and stability of key services are enhanced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Telecommunication system failures prediction through machine learning and artificial intelligence

In an embodiment, a method may be implemented in a computer system comprising a processor, memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor, the computer system interconnected with a telecommunications system, the method comprising: receiving, at the computer system, data relating to operation of the telecommunication system, obtaining, at the computer system, at least one machine learning model trained to detect and predict faults in the operation of the telecommunication system, selecting, at the computer system, computing infrastructure upon which to execute the at least one machine learning model, wherein the selected computing infrastructure comprises a mesh of interconnected micro-applications;, executing, at the computer system, the at least one machine learning model using the selected computing infrastructure to detect and predict faults in the operation of the telecommunication system, and automatically correcting at least some of the detected faults.
Owner:GENESIS INTELLIGENCE LLC

Train control method and train control system for entering check based on trusted interval

The invention discloses a train control method and a train control system based on credible interval entry check. The train control method comprises the following steps: starting a backup command; the temporary speed limit server starts a sending module used for sending authorization to a train self-protection system; performing interval credible entry check; determining the existence condition of the train in the interval; the temporary speed limit server updates a train existence condition list in the interval; the train obtains a front driving permission; the train runs according to the running permission. According to the train control method and the train control system based on the credible interval entry check, the axle counting equipment is arranged at the interval entrance, software of the temporary speed limit server only needs to adjust individual modules, and the software of the temporary speed limit server can be only used for recording the train condition in the interval and is not used for authorized calculation, so that the train control method and the train control system are convenient to use. The deployment cost of the temporary speed limit server is reduced, the transportation efficiency under the condition of a CTCS-2 system fault is improved, and the toughness of the CTCS-2 system is improved.
Owner:CASCO SIGNAL LTD

Cloud native system fault root cause positioning method and device

The invention provides a cloud native system fault root cause positioning method and device, and the method comprises the steps: determining the abnormal performance index data of each instance and a server based on the micro-service instance of each target micro-service and the performance index abnormal score of the server, so as to construct the nodes corresponding to each micro-service instance and the server, according to a fault propagation direction between the abnormal performance index data, constructing an edge between the nodes to obtain an index-level cause and effect graph; normalizing the performance index anomaly score to obtain a target anomaly score of each node, and constructing a transition probability matrix; and obtaining the access frequency of each node in the causal graph by using a random walk algorithm so as to obtain fault root cause positioning result data of the cloud native system. According to the invention, the automation degree, efficiency and accuracy of fault root cause positioning of the cloud native system can be effectively improved, the efficiency and reliability of fault early warning and recovery of the cloud native system can be effectively improved, and the operation stability of the cloud native system can be improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Health evaluation method, electronic equipment, storage medium and product

The invention discloses a health evaluation method, electronic equipment, a storage medium and a product, and relates to the technical field of computers, and the method comprises the steps: based on time sequence index data and log data in a to-be-evaluated system, respectively determining a monitoring index deviation degree corresponding to the time sequence index data and a log abnormal probability corresponding to the log data; based on the monitoring index deviation degree and the log abnormal probability, determining an index fusion weight corresponding to the monitoring index deviation degree and a log fusion weight corresponding to the log abnormal probability; and according to the monitoring index deviation degree, the log abnormal probability, the index fusion weight and the log fusion weight, in combination with a health degree scoring function, determining a health degree score of the to-be-evaluated system, so as to execute a recovery operation on the to-be-evaluated system according to the health degree score. The system fault can be detected and evaluated more accurately, the recovery strategy can be dynamically adjusted according to different fault severity degrees and actual conditions of the system, and the recovery efficiency and fault-tolerant capability of the system are remarkably improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

H-bridge key equipment service life and system reliability evaluation method and system for cascade networking type energy storage system

The invention discloses an H-bridge key equipment service life and system reliability evaluation method and system for a cascade network construction type energy storage system, and belongs to the technical field of power system automation. The method comprises the following steps: firstly, extracting task profile parameters under multiple time scales, and constructing a time sequence feature model; secondly, estimating a hot spot temperature sequence of the IGBT device and the capacitor based on a multilayer feedforward neural network; then, in combination with a continuous extreme point paired temperature cycle extraction method and a Miner linear cumulative damage criterion, the damage factor and the residual life of the device are evaluated; then, task profile samples are expanded based on a generative adversarial network with gradient penalty, and life distribution and reliability indexes of key devices under different profiles are calculated; and finally, based on H-bridge series structure mapping device level information, constructing a system level reliability model, obtaining system failure rate, average fault-free operation time and a reliability function, and realizing health state perception and reliability quantitative evaluation of the energy storage system.
Owner:SOUTHEAST UNIV

Context-driven fault solution suggestion generation method

The invention relates to a context-driven fault solution suggestion generation method, which specifically comprises the following steps of: firstly, collecting various log data generated in system operation, and preprocessing the data as a basis for subsequent analysis; then, a BERTomaly model based on a multi-head cross attention mechanism is constructed, feature extraction, feature fusion and template matching can be performed on the log data, anomaly detection of system faults can be realized, and then, based on the system operation log data, the system fault detection efficiency is improved. The method comprises the following steps of: firstly, constructing a fault solution library which is of a context-driven type and has a dynamic updating capability by using a large language model in combination with the output of a BERTomally model and the output of an FP-Growth association analysis algorithm, and carrying out real-time detection on each fault, so that the fault solution library has a dynamic updating capability. And a targeted solution is output.
Owner:SOUTHEAST UNIV +1

Energy storage system fault processing method, device, equipment, medium and product

The embodiment of the invention provides an energy storage system fault processing method and device, equipment, a medium and a product, and relates to the field of photovoltaic power generation, the method solves the problem of low fault processing efficiency in the prior art by constructing a cloud-side-end cooperative fault processing architecture, changes the traditional static threshold judgment logic, and improves the fault processing efficiency. A decision-making mode combining multi-source information fusion and dynamic prediction is adopted, firstly, preliminary diagnosis of real-time data is achieved through edge calculation, the instantaneity of response is ensured, and then historical trend analysis, equipment health prediction and multi-dimensional real-time state evaluation are fused through a cloud platform. According to the method, the dynamic fault evaluation capable of accurately reflecting the actual severity and development trend of the fault is generated, so that the system can be adaptive to equipment aging and environment change, and a grading response strategy accurately matched with the fault grade is triggered, thereby realizing the crossing from passive alarm to active and accurate fault management and control, and improving the fault processing efficiency.
Owner:QINGDAO NAHUI ENERGY TECH CO LTD

New energy wind power system fault detection method, system and device based on multi-parameter fusion and medium

The invention relates to the technical field of new energy power generation monitoring, in particular to a new energy wind power system fault detection method, system and device based on multi-parameter fusion and a medium. The method comprises the following steps: acquiring various operating parameters such as wind speed, wind direction, generator speed, temperature, voltage and current of a wind power system in real time, and performing feature extraction on the operating parameters to obtain key features related to a fault; analyzing the key features based on a deep learning algorithm, and judging the fault type of the wind power system; and outputting a fault diagnosis result in a visual mode, and sending the fault diagnosis result to a control center for remote monitoring and fault early warning. The technical problems that fault detection of a traditional wind power system is poor in real-time performance, not high in accuracy, low in intelligent degree and the like are solved, and reliable guarantee is provided for safe and stable operation of the wind power system.
Owner:HUANENG LIAOCHENG THERMAL POWER CO LTD

Non-volatile memory rapid recovery method based on metadata priority and on-demand loading

The invention relates to the technical field of computer system structures and storage, in particular to a non-volatile memory quick recovery method based on metadata priority and on-demand loading. The method comprises the following steps: in response to a system fault signal detected by a voltage monitoring unit, freezing a processor context and traversing a page table structure to extract system configuration information; writing the system configuration information and business data codes in the volatile memory into a nonvolatile medium to generate a persistent state mirror image; analyzing the persistent state mirror image, extracting address conversion metadata, and reconstructing a mapping relation from a virtual address to a physical page frame in a volatile memory; and generating an address mapping table, wherein the physical page frame pointed by the address mapping table is set to be in an existing state but is not associated with the effective service data. According to the method, decoupling of the control flow and the data flow is realized by constructing a virtual ready state, and quick starting of the system and immediate response of key services are realized on the premise of not depending on the total capacity of a memory.
Owner:CHENGDU FUYUNXUN TECHNOLOGY CO LTD +1

Electric power system optimization scheduling method based on artificial intelligence

The invention discloses a power system optimization scheduling method based on artificial intelligence, and relates to the technical field of power system fault testing. The method comprises the steps of obtaining production monitoring data, operation data and management data under a target power system as a power data set, and performing preprocessing; constructing a load prediction model, inputting historical load data, user power consumption behaviors, meteorological data and calendar characteristics, and outputting a predicted power grid load curve; the technical key points are as follows: through combination with an equipment health state evaluation model, dynamic adjustment of health scores and subsequent real-time scheduling optimization, an optimal management scheme with extremely high linkage is formed, the effect of accurately identifying potential fault risks is realized, operation and maintenance personnel can take prevention measures in advance, and the risk of potential faults is prevented. The situation of misjudgment or missing report caused by neglecting external factors is avoided, and the pre-judgment and response capability of the whole scheduling scheme to the equipment health risk is enhanced.
Owner:GANSU SHINING SCI & TECH

PLC scheduling algorithm based on multi-core soft real-time operating system

The invention discloses a PLC scheduling algorithm based on a multi-core soft real-time operating system, and relates to the technical field of voltage control, and the algorithm comprises the following four steps: 1, starting a target industrial control system, completing hardware self-inspection, memory allocation and communication module starting, and sending an initialization ready signal after confirming that an assembly is normal; 2, according to PLC task functions and real-time performance, four types of standardized tasks are divided, priority levels are defined, sub-priorities are configured, a dependency relationship is identified, and a task chain is established; 3, counting the number of tasks and task chains, distributing resources by using multiple binding cores, and constructing a global ready queue to adapt to dynamic scheduling; and 4, monitoring an interrupt event through an interrupt manager, realizing high-priority task preemption by semaphores, and allocating tasks to corresponding CPU cores, so that the problems of soft real-time multi-core task preemption and data conflict are solved, the scheduling real-time performance and the data security are improved, the system failure rate is reduced, and the industrial control efficiency and stability are improved.
Owner:NANJING AOTUO AUTOMATION TECH CO LTD

Automatic fault injection and recovery test method and system based on cloud native

The invention belongs to the technical field of cloud computing and software testing, and particularly discloses an automatic fault injection and recovery testing method based on cloud native, which comprises the following steps: acquiring test basic data of a to-be-tested cloud native system; constructing a fault test execution model according to the test basic data; according to the fault test execution model, executing an automatic fault injection operation; generating an abnormal alarm log according to the system running state after the fault injection operation, and synchronously recording a fault injection operation log; according to the abnormal alarm log, executing automatic fault recovery operation; and generating a fault test report according to the fault injection operation log, the system monitoring data and the recovery operation process log. The invention aims to solve the problems of low automation degree, narrow scene coverage, weak monitoring evaluation and passive recovery in cloud native system fault test in the prior art, and provides a comprehensive, automatic, monitorable and evaluable fault injection and recovery test scheme.
Owner:HUANLE ENTERTAINMENT SHANGHAI TECH CO LTD

Combustion system fault correlation analysis and diagnosis method and system based on knowledge graph

The invention relates to the technical field of industrial intelligent control, in particular to a combustion system fault correlation analysis and diagnosis method and system based on a knowledge graph. The method comprises the following steps: constructing a mechanism topological graph of a combustion system, obtaining real-time load data of the combustion system, calculating dynamic conduction lag time of a parent node influencing a child node in the mechanism topological graph based on a fluid mechanics load correction mechanism in combination with a flow resistance correction factor and reference transmission time, performing time sequence alignment based on dynamic conduction delay time, calculating causal association strength of a father node pointing to a child node and performing confidence attenuation processing, performing reverse search on a mechanism topological graph when an abnormal trigger point is monitored, calculating accumulated abnormal energy of each father node based on the causal association strength, and determining the abnormal trigger point according to the accumulated abnormal energy. And determining the father node with the highest accumulated abnormal energy as a fault root cause. According to the scheme of the invention, accurate time sequence alignment under working condition fluctuation can be realized, diagnosis failure can be prevented, and fault root causes can be accurately positioned.
Owner:YIXING HOTTEEN ENVIRONMENTAL PROTECTION ENG

Fault self-checking method and system for circulating ball type electric power steering system and readable medium

The invention discloses a recirculating ball type electric power steering system fault self-checking method and system and a readable medium, and relates to the technical field of vehicles, and the method comprises the steps: storing a signal transmission path and each function node, determining the type of an influence factor which influences the instruction execution accuracy of the function node, and determining the instruction execution accuracy of the function node; analyzing and acquiring the relationship between the fluctuation amplitude of the data signal at each function node and the numerical value of the influence factor, storing the relationship as a correction model, and storing a first relationship model for reflecting the data association relationship among the function nodes on the same signal transmission path; acquiring real-time data and source data of various impact factors, and generating theoretical response data according to the first relation model, the real-time impact factor data and the correction model; the actual response data is collected, the deviation amplitude between the actual response data and the theoretical response data is calculated, whether the function node breaks down or has a potential fault is judged according to the deviation amplitude, early warning is output, and the fault detection accuracy and the potential fault pre-judgment capability are greatly improved through the technical scheme.
Owner:WUHAN PUXIXIN ELECTRONIC TECH CO LTD

MMC (Modular Multilevel Converter) fault processing method and device based on sub-module, equipment and storage medium

The invention discloses an MMC sub-module fault processing method, device and equipment and a storage medium, and the method comprises the steps: obtaining the capacitor voltage value of a sub-module of an MMC and the bridge arm current value of the sub-module, the MMC comprises six bridge arms of three phases, each bridge arm comprises a plurality of sub-modules, the capacitor of each sub-module is connected with an equivalent series resistor in series, and further, the equivalent series resistor is connected with the capacitor voltage value of the sub-module of the MMC; the resistance value of an equivalent series resistor in the sub-module is monitored through the capacitor voltage value and the bridge arm current value, if the resistance value of the equivalent series resistor is higher than a preset resistance value, it is determined that the capacitor of the sub-module fails, fault-tolerant control is carried out on the sub-module through a voltage-sharing and circulating current suppression algorithm, a three-phase switching signal is output, and the three-phase switching signal acts on the MMC. Therefore, the circulating current fundamental frequency component and the double frequency component caused by fault-tolerant control after the fault can be eliminated through the circulating current suppression strategy, the effect of circulating current suppression is achieved, and the stability of system fault operation is improved.
Owner:ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID +1

System abnormity intelligent diagnosis and recovery method and system fusing time sequence logs

The invention provides a system abnormity intelligent diagnosis and recovery method and system fusing time sequence logs, and relates to the field of distributed system operation and maintaining.The method comprises the steps that operation logs are collected from a plurality of service nodes of a distributed system, time sequence alignment is conducted, and calling relation and performance measurement data are extracted; time sequence change features are calculated in the sliding time window, dynamic feature vectors are generated through fusion coding, and a cross-service time sequence association graph is constructed; identifying an abnormal service node and matching the abnormal service node with historical fault data to obtain a historical propagation path and a feature vector; calculating a propagation convergence coefficient and a characteristic deviation degree based on a historical propagation path to obtain a causal intensity score; and identifying a fault influence degree according to the score, and adaptively adjusting a resource isolation and flow scheduling strategy. According to the invention, accurate diagnosis and efficient recovery of distributed system faults are realized.
Owner:SMIC WANYE TECHNOLOGY CO LTD

HPLC system fault early warning self-calibration method

The invention relates to the technical field of intelligent control of HPLC (High Performance Liquid Chromatography) analytical instruments, in particular to an HPLC system fault early warning self-calibration method. Comprising the following steps: S1, acquiring a high-frequency pressure signal behind a pump and a chromatographic signal of a detector in real time; calculating a relative physical disturbance quantity based on the post-pumping high-frequency pressure signal; calculating a chemical state deviation degree based on the detector chromatographic signal; s2, fusing the relative physical disturbance quantity and the chemical state deviation degree, and constructing an instantaneous health index; and S3, comparing the instantaneous health index with a preset first early warning threshold value and a preset second early warning threshold value, and executing a closed-loop control strategy corresponding to a comparison result. According to the method, accurate fault early warning is realized by constructing a physical-chemical two-dimensional monitoring system, the defects of offline and lagging applicability test of a traditional system are overcome through real-time monitoring and fusion evaluation of the two dimensions, and the accuracy and timeliness of early warning are remarkably improved.
Owner:NINGBO SMART PHARMA

Micromotor fault prediction and health management system

The invention belongs to the crossing field of artificial intelligence and mechanical engineering, particularly relates to a micro-motor fault prediction and health management system, and aims to solve the problems that early faults of a micro-motor are difficult to recognize, degradation modeling is inaccurate and maintenance lags. The system collects multi-source data through high-density sensing, combines denoising reconstruction, composite feature extraction and time-varying weighted fusion to generate health indexes, identifies health stages by using a segmented hidden Markov model, iteratively updates residual life prediction based on a Wiener process, outputs an estimation result with a confidence interval, and links a hierarchical maintenance strategy. And continuous optimization of the model is realized through federal learning. The system improves the fault early warning accuracy and prediction reliability, and reduces the operation and maintenance cost.
Owner:SHANGHAI SIDAPU IND CO LTD

Deicing system real-time state evaluation and fault prediction method and system

The invention discloses a deicing system real-time state evaluation and fault prediction method and system, and the method comprises the steps: simulating the fault state of a deicing system through a ground state monitoring test of the deicing system, and obtaining sensor data in different fault states; constructing a data set by using sensor data in different fault states to train a state evaluation model and a fault prediction model; the method comprises the following steps: installing a sensor on an aircraft deployed with a deicing system to collect working data of the deicing system; and inputting data acquired by the sensor into the trained state evaluation model and the fault prediction model, evaluating the real-time state of the deicing system, and predicting possible faults. According to the invention, the current working state of the anti-icing and deicing system can be evaluated according to the current state monitoring data, and possible faults of the system can be predicted.
Owner:成都流体动力创新中心

Neutral point non-effective grounding power distribution network secondary overline fault detection method

The invention relates to the technical field of power system fault detection and diagnosis, in particular to a neutral point non-effective grounding power distribution network secondary overline fault detection method, device and equipment and a computer storage medium. According to the method for detecting the secondary overline fault of the neutral point non-effectively grounded power distribution network, an overline different-phase secondary fault equivalent circuit model is established, a fault electrical quantity analytical expression is solved, the characteristics of each stage of the secondary grounding fault under different fault phase sequences are analyzed, and the fault detection accuracy is improved. Corresponding different-phase secondary overline grounding fault detection logic is provided, the accuracy and reliability of the secondary grounding fault detection method are improved, and the method has positive significance in improving the power supply safety of the power distribution network.
Owner:TSINGHUA UNIVERSITY +1

Dynamic event trigger fault detection method under DoS network attack

The invention relates to the technical field of network detection, and provides a dynamic event trigger fault detection method under DoS network attack, which comprises the following steps: establishing a linear state space model of a network control system; defining a non-attack interval and an attack interval of system operation, and constraining attack frequency and duration; the observer gain is switched according to the current non-attack interval or attack interval of the system; a dynamic event triggering mechanism is constructed, and the triggering condition depends on the output state and the internal dynamic variable of the full-order switching observer and is used for dynamically adjusting the data transmission frequency; establishing a closed-loop switching system model; and analyzing system index stability according to the closed-loop switching system model, and cooperatively designing observer gain, controller gain and event triggering parameters. According to the invention, through quantification of the DoS attack model, the dynamic event triggering mechanism and collaborative optimization design, the effects of effectively detecting the system fault and improving the utilization rate of the network channel are achieved.
Owner:GUANGZHOU UNIVERSITY

Back-end system fault self-recovery method and device, storage medium and computer equipment

According to the back-end system fault self-recovery method and device, the storage medium and the computer equipment provided by the invention, in the system operation process, multi-dimensional state analysis is performed on the operation data collected by the system micro-service in real time to obtain the operation state; when the operation state is abnormal, the fault analysis model is used for performing fault analysis on the operation data, so that the fault positioning efficiency is improved; after fault data output by the model are obtained, a self-healing strategy corresponding to the fault data is obtained through matching in the dynamic knowledge base, then self-healing operation is executed on the micro-service, and quick response of fault processing is achieved. If the self-healing result is failure, self-healing data in the self-healing operation process is obtained, the self-healing strategy of the micro-service is adjusted to be more adaptive to the actual fault condition, then the self-healing operation is executed on the micro-service again based on the new self-healing strategy until the self-healing result meets the self-healing ending condition, and therefore the flexibility of system fault self-healing is improved; and the influence of the fault on the system stability is reduced to the greatest extent.
Owner:创优数字科技(广东)有限公司

A method, device, computer-readable storage medium and electronic device for intelligently processing distributed system faults

The present application relates to a method and device for intelligently processing faults in a distributed system. The method includes: starting an application service and generating a heartbeat packet; sending a heartbeat packet to a heartbeat queue at a set time interval; a monitoring service receiving and saving the data in the heartbeat queue; continuously collecting the data in the heartbeat queue from each node and preprocessing it, extracting features related to fault detection, and using an isolation forest model for model training; using an adaptive threshold adjustment method based on an isolation forest algorithm to analyze the data in the heartbeat packet; comparing the heartbeat according to the threshold, and when an abnormal heartbeat is found in a node, determining the fault location according to the feature code; generating a compensation message, and the application service executing the compensation task according to the compensation message. The method detects faults by combining a heartbeat mechanism with an AI algorithm, and recovers unfinished tasks by a compensation mechanism. The method is adaptable to complex distributed environments, has high real-time performance and sensitivity, and significantly improves the stability and reliability of the system.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Intelligent operation and maintenance and fault handling method based on AI large model

The invention discloses an intelligent operation and maintenance and fault handling method based on an AI large model, and relates to the technical field of artificial intelligence and information, and the intelligent operation and maintenance method comprises the following steps: multi-source heterogeneous data fusion, cross-level fault analysis, handling strategy generation and verification, and resource optimization scheduling. The method has the advantages that multi-source heterogeneous data of the equipment layer, the software layer and the service layer are synchronously collected through the sensing layer, the space-time aligned feature matrix is constructed, the limitation that traditional operation and maintenance only depend on single-level data is solved, the time dependence and level incidence relation of the data can be mined at the same time in combination with a multi-mode AI large model of the analysis layer, and the data processing efficiency is improved. The fault root cause probability distribution and the influence hierarchy are accurately output, misjudgment or missed judgment caused by single hierarchy data deviation is avoided, and the accuracy of complex system fault positioning is remarkably improved.
Owner:ZHEJIANG SHUOANG TECHNOLOGY CO LTD

Automatic operating system fault repairing method based on artificial intelligence

The invention discloses an automatic fault repairing method for an operating system based on artificial intelligence, and relates to the technical field of automatic fault repairing, and the method comprises the steps: carrying out the feature analysis of system operation data through a pre-trained fault feature extraction model, generating a fault feature vector, inputting the fault feature vector into a fault classifier, and obtaining a fault feature vector; a current fault type is identified through a multi-classification algorithm, fault cause primary tracing is performed according to the fault type to obtain a fault generation factor, secondary tracing is performed on the fault generation factor to obtain a fault influence factor, positioning is performed based on the fault generation factor, and a corresponding repair strategy is matched from a knowledge base. The method comprises the following steps: acquiring historical system operation data with relevance on the basis of a fault influence factor, acquiring updated real-time system operation data after executing a repair operation to calculate a system optimization coefficient, judging a forward trend of a repair strategy according to a preset optimization threshold value, and updating the forward trend into a knowledge base to realize rapid and efficient automatic repair.
Owner:SICHUAN CHANGFU INFORMATION TECHNOLOGY SERVICE CO LTD

Intelligent core file debugger for cluster file system serviceability

Providing issue resolution in a cluster system by monitoring system operation to detect occurrence of a system error, and automatically generating, upon detection of the error condition, a core file for a user node. The core file captures a current memory state of a respective node, where the current memory state comprises system statistics, system information, and logs. An intelligent core debugger extracts information from a core file to generate a core file report that is sent to a vendor for a quick determination of whether sufficient information is in the report to allow the vendor to recommend a fix, or whether further information from is required, including the core file itself, if necessary. This prevents the need to send an entire core file to a vendor in every instance of a system fault.
Owner:DELL PROD LP

Fault analysis method of energy storage system

The invention relates to the technical field of energy storage system fault analysis, in particular to an energy storage system fault analysis method, which comprises the following steps: arranging a plurality of sensors, and collecting fault associated data of an energy storage system; constructing a fault analysis model, and generating a fault analysis result according to the fault associated data; detecting whether abnormal data exists in the fault associated data or not, if yes, performing fault multi-stage analysis, and obtaining fault analysis results, including obtaining a first fault analysis result and a second fault analysis result by inputting the fault associated data and the corrected fault associated data into a fault analysis model, and comparing and analyzing whether the two results are the same or not, and analyzing reasons for abnormal data in combination with sensor fault detection and historical fault analysis results to obtain a fault analysis result. According to the scheme, when the sensor data is abnormal, reason analysis can be carried out, the output of a fault analysis result is not influenced, the accuracy of fault analysis is improved, and the safety of an energy storage system is ensured.
Owner:CHONGQING HITEN ENERGY CO LTD