Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4359results about "Fault response" patented technology

Log aggregation fault diagnosis method and system based on artificial intelligence

The invention relates to the field of log fault analysis, in particular to a log aggregation fault diagnosis method and system based on artificial intelligence. The method comprises the following steps: collecting a multi-modal heterogeneous log, carrying out sliding time sequence slicing processing, carrying out time sequence association sequence reconstruction, and constructing a time sequence reconstruction log data stream; log event deep semantic analysis is carried out on the time sequence reconstruction log data stream, event semantic topological evolution is carried out, and a multi-dimensional event topological representation matrix is constructed; performing routine event behavior analysis and abnormal fault mode inference based on the multi-dimensional event topology representation matrix, and marking abnormal fault points; and the occurrence timestamp and the abnormal propagation rate of the abnormal fault point are calculated, fault space-time diffusion evolution is carried out, and a dynamic fault propagation path map is constructed. Through efficient and accurate fault traceability analysis, the fault diagnosis efficiency is greatly improved, and the stability and reliability of log data are improved.
Owner:SHANGHAI FEIWEI INFORMATION TECH CO LTD +2

Industrial system automatic fault diagnosis method based on large language model

The invention discloses an industrial system automatic fault diagnosis method based on a large language model. According to the method, a three-layer mapping system of industrial data, natural language description and knowledge reasoning is constructed, field multi-source sensor data are subjected to semantic conversion, and a quantitative calculation model based on a large language model is constructed based on historical data and logs. And a fault case is matched in real time with the help of a retrieval-enhancement generation technology to serve as a reference, the fault case and abnormal information are input into a knowledge reasoning model based on a large language model together, a structured logical reasoning chain is generated, and a diagnosis conclusion containing candidate faults, cause analysis and disposal suggestions is further output. Meanwhile, through user feedback and a reinforcement learning mechanism, the model and the knowledge base are adaptively updated, the defects of traditional static rules and expert experience are effectively overcome, the accuracy, interpretability and robustness of fault detection are remarkably improved, and the method adapts to complex and changeable working condition requirements.
Owner:ZHEJIANG UNIV

Monitoring fault analysis method fused with multi-modal knowledge base

The invention relates to the technical field of fault analysis, and particularly provides a monitoring fault analysis method fused with a multi-modal knowledge base, which comprises the following steps: collecting original data of a monitoring fault log, and preprocessing and storing the original data; performing data cleaning and feature extraction on the obtained original data of the monitoring fault log; constructing a searchable knowledge base based on the cleaned data; when the system triggers an alarm, mixed retrieval is executed through a dynamic routing mechanism; aggregating the plurality of retrieval results to generate an executable repair scheme; iteratively optimizing the decision process through manual feedback; and continuously optimizing the knowledge base and the diagnosis model to form a closed loop iteration mechanism. According to the scheme, the accuracy and response efficiency of fault diagnosis are improved.
Owner:ADVANCED OPERATING SYST INNOVATION CENT (TIANJIN) CO LTD

Fault root cause positioning method and system driven by dynamic knowledge graph

The invention discloses a fault root cause positioning method and system driven by a dynamic knowledge graph, and relates to the technical field of fault root cause localization, and the method comprises the steps: collecting and obtaining a multi-source fault associated data set, carrying out the entity association extraction of the multi-source fault associated data set, and obtaining a fault entity set and an entity relationship set; performing graph node cascading and incremental learning updating, and constructing a fault updating knowledge graph; monitoring and acquiring target fault data, performing mode matching reasoning, and generating a fault mode candidate root cause set; and performing similarity matching on the fault mode candidate root cause set in combination with a historical fault case library, and determining a target fault root cause positioning result. The technical problem of low fault diagnosis efficiency caused by inaccurate fault root cause positioning and knowledge graph updating lagging in the prior art is solved, and the technical effects of realizing accurate positioning of the fault root cause and dynamic improvement of the knowledge graph and improving the fault diagnosis efficiency and accuracy are achieved.
Owner:BEIJING JIANXING TECHNOLOGY CO LTD

Intelligent agent system optimization method and device based on intelligent fault analysis and cross-generation knowledge inheritance

The invention relates to an intelligent agent system optimization method and device based on intelligent fault analysis and cross-generation knowledge inheritance, and belongs to the technical field of artificial intelligence. According to the method, interaction abnormal signals are captured in real time by deploying a lightweight log probe, and a tool benefit prediction model based on reinforcement learning is constructed to automatically generate an improvement proposal when the failure rate exceeds a threshold value; an agent genealogy map is established to realize automatic inheritance of a new agent on core memory and abandonment of failure knowledge, and a disastrous forgetting blocker is deployed to dynamically extract a functional module from a genealogy to deal with key capability degradation. Aiming at the problems of fault response lag, knowledge inheritance fracture, key capability degradation and the like in an intelligent agent system iteration process, the invention creatively provides a cooperation mechanism of an intelligent fault analysis layer and a cross-generation knowledge inheritance network, and the fault self-healing capability, version stability and service continuity guarantee level of the system are remarkably improved.
Owner:KUNLUN YUAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (SHANGHAI) CO LTD

Operation maintenance management method of integrated management system

The invention discloses an operation and maintenance management method of an integrated management system, and belongs to the technical field of operation and maintenance of systems. The invention discloses an operation and maintenance management method of an integrated management system, and aims to solve the problems of data islands, slow fault positioning, experience dependence on strategies and the like in traditional operation and maintenance. The method comprises the following nine core processes: dynamically accessing multi-source heterogeneous data and carrying out standardization processing; constructing a hierarchical time series data storage structure; generating a modeling dependency and fault path of the equipment knowledge graph; adopting a three-layer anomaly detection model to identify anomaly; fault root causes are positioned through causal reasoning and a Bayesian network; generating an energy efficiency strategy based on reinforcement learning and multi-objective optimization; triggering the self-healing workflow to execute operation; testing the robustness of the system in a sandbox environment; and iteratively updating the knowledge graph and the AI model to form a closed loop. According to the method, automation and intelligentization of the whole operation and maintenance process are realized, and the system availability and the energy efficiency management level are improved.
Owner:TIBET SHENGMEIJIA NETWORK TECHNOLOGY CO LTD

Data center operation and maintenance fault prediction system and method based on deep learning

The invention discloses a data center operation and maintenance fault prediction system and method based on deep learning. The system comprises a multi-source heterogeneous data acquisition module, a data preprocessing module, a deep learning prediction model module and the like. The method comprises the following steps: acquiring multi-dimensional operation data of a data center through full-quantity acquisition of multi-source data, and inputting a CNN-LSTM-Attention hybrid model to realize fault prediction after preprocessing and feature enhancement; fault grades are divided in combination with fault grading, early warning is pushed in multiple channels, a coping strategy is intelligently generated, the effect is verified in a closed loop mode, and finally the model is iteratively optimized. According to the scheme, the fault prediction precision and real-time performance are improved, the operation and maintenance response time is shortened, the service interruption risk caused by faults is reduced, and the method is suitable for efficient operation and maintenance of large-scale data centers.
Owner:SHANGHAI DIPU XINCHENG INTELLIGENT TECH CO LTD

Fault diagnosis method for server

The invention relates to a fault diagnosis method for a server. Efficient fault positioning and processing are achieved by constructing a multi-dimensional intelligent diagnosis system. Constructing a server virtual model at each node and establishing a connection relationship, and collecting various data and mapping the data to the virtual model; performing weighted fusion on the data through correlation analysis, and extracting fault features including time and space correlation by using a bidirectional long-short-term memory network; fault probability distribution is obtained by means of a dynamic fault knowledge graph and graph neural network reasoning, and a fault evaluation result is visualized in combination with a virtual model; and finally analyzing the behavior deviation through a neural network and generating a processing strategy. Real-time processing of multi-source data, intelligent extraction of fault features and dynamic deduction of fault propagation are achieved, the accuracy, the real-time performance and the automation level of server fault diagnosis are effectively improved, and predictive maintenance and intelligent decision making of a complex server system are achieved.
Owner:BEIJING HUAKUN ZHENYU INTELLIGENT TECH CO LTD

Automatically generating reports of incident events

ActiveUS12487874B1Fault responseSpecial data processing applicationsData setIncident management (ITSM)
A computer-implemented method executed using one or more processors of an incident management system, the computer-implemented method comprising accessing one or more data sets of information associated with an incident event corresponding to an incident associated with a computer system; generating a prompt based on the one or more data sets of information, wherein generating the prompt comprises generating a plurality of sub-prompts to be provided to a machine-learning model for generating a report of the incident event in accordance with a predetermined criteria; inputting the prompt into a machine-learning model that has been trained to generate a report of the incident event based on the prompt; outputting, by the machine-learning model, the report of the incident event, wherein the report comprises an analysis of the incident event; transmitting the report to one or more computing devices associated with the computer system.
Owner:PAGERDUTY INC

Intelligent operation and maintenance method and device based on knowledge graph and large model and electronic equipment

The invention relates to an intelligent operation and maintenance method and apparatus based on a knowledge graph and a large model, and an electronic device. The method comprises the steps of collecting multi-source runtime data of a Kubernetes cluster; constructing a knowledge graph with time dimension based on the resource change event, recording termination time in response to graph relationship failure, recording starting time in response to a newly added relationship and not setting the termination time, and associating the entity with the performance index and the log data; responding to the diagnosis request, scheduling a specialized agent by a coordination agent through multi-agent cooperation to retrieve associated information from multi-source data, and iteratively integrating to generate a structured context; and inputting the generated structured context information into a large model reasoning service to output a fault root cause diagnosis and solution. The technical problems that information dispersion and relevance are weak, root cause positioning is difficult and time-consuming, comprehensive context sensing ability is lacked and expert experience is excessively relied on are solved.
Owner:INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2

Switch equipment fault early warning method based on multi-source data fusion

The invention relates to the technical field of fault prediction, in particular to a switch equipment fault early warning method based on multi-source data fusion, which comprises the following steps of: acquiring a channel transition time sequence to generate a time interval, marking a behavior label to generate a time sequence, constructing a sliding window to form a combined alarm, analyzing a cross matrix to extract an abnormal time period, and performing early warning. And identifying the leading channel and sending out early warning. According to the method, alignment processing is carried out through the time intervals of multi-channel state transition responses, dynamic linkage recognition among multiple signals is achieved, time sequence correlation among abnormal behaviors is enhanced by means of behavior label combination and a sliding statistical strategy, and the sensitivity and accuracy of abnormal recognition are enhanced through aggregation expression of composite alarm events. In combination with cross-channel behavior label matrix analysis, the continuity of abnormal labels is clear and key channels are identified, so that the accuracy and timeliness of early warning of the switch equipment are improved, the risk of false alarm and missing alarm is reduced, and the actual demand of multi-source sensing data fusion analysis of the switch equipment is met.
Owner:BEIJING YANENG ELECTRIC EQUIP CO LTD

Processing environment switching and recovering method and device, equipment and medium

PendingCN121092357AFault responseRecovery methodMulti source data
The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a processing environment switching and recovery method, device, equipment and medium. The method comprises the steps that multi-source heterogeneous data in a main processing environment and a standby processing environment are acquired, and the system fault probability is obtained through multi-model collaborative prediction; a dynamic threshold value is generated in combination with a historical service period mode and a real-time service load, when the fault probability exceeds the threshold value, a switching strategy is generated based on the fault scene knowledge base and the service priority, and flow scheduling between the main processing environment and the standby processing environment is executed; and monitoring the business index of the standby processing environment during the scheduling period, and triggering the fusing rollback when the business index is lower than the health standard. According to the method, the fault identification precision is improved through multi-source data fusion and multi-model prediction, adaptive scheduling is realized in combination with a dynamic threshold and a switching strategy, and fusing rollback is triggered to guarantee high availability and data consistency, so that the continuity and stability of key services are enhanced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Intelligent data monitoring method and monitoring platform

The invention relates to the technical field of data monitoring, and discloses an intelligent data monitoring method and a monitoring platform, and the intelligent data monitoring method comprises the steps: constructing a dynamic heterogeneous hypergraph model, and representing the complex relation and time sequence evolution characteristics among multi-source data nodes; based on a dynamic heterogeneous hypergraph model, realizing high-order relationship representation and calculation, and obtaining embedded representation of nodes and relationships; performing heterogeneous graph information transmission and aggregation by utilizing embedded representation of nodes and relationships, and capturing complex interaction characteristics among multiple nodes; according to the complex interaction characteristics, an abnormal propagation rate model is established, and accurate prediction of an abnormal diffusion path is realized; based on the abnormal propagation rate model and the abnormal diffusion path, multi-level collaborative anomaly detection is implemented, and collaborative faults across subsystems are identified; according to the invention, the abnormal propagation path can be predicted in advance, the abnormal prediction accuracy is improved, and the system fault response time is advanced.
Owner:ZHANGJIAGANG BIG DATA CO LTD

AIOps anomaly detection and root cause positioning method

The invention discloses an AIOps anomaly detection and root cause positioning method, and relates to the technical field of anomaly detection. The method comprises the following steps: 1, processing a preset time window according to services and instances under a unified timeline, generating monitoring type abnormal fragments for monitoring indexes, and extracting a client and server span in distributed tracking for fragment pairing; step 2, executing stitching by taking distributed tracking as guidance to obtain a candidate evidence chain set, and taking a segment at the tail end of each evidence chain unit as a candidate root cause direction; and step 3, outputting a root cause list for the candidate evidence chain set according to a deterministic rule, and giving a time range of a related template text and an adjacent monitoring type abnormal fragment. According to the method, the abnormal propagation path can be accurately identified in the multi-source heterogeneous data, association verification is carried out on the upstream representation and the downstream resource failure, and a clear root cause target and evidence explanation are provided.
Owner:NINGBO SANYANG INFORMATION TECH CO LTD

Fully automated log analysis and fault handling system and method based on NLP large model

PCT designated stageWO2025156166A1Fault responseFault analysisBusiness process
Disclosed in the present invention is a fully automated log analysis and fault handling system and method based on an NLP large model. The method comprises: collecting a service state, an operating environment, and log files; extracting and generating log summaries from the log files, extracting key information from the log summaries to obtain state information of a service process; if an error is fed back, the NLP large model performing fault analysis on the error by using the service state and the operating environment as prior knowledge, and recording upstream and downstream key information involved in the error; on the basis of the upstream and downstream key information involved in the error, finding a fault cause and providing a fault handling action; and then executing a plan on the basis of the fault handling action. Log collection of the present invention is not limited to a prescribed format, and an NLP large model using a service state and an operating environment as prior knowledge is introduced for fault analysis.
Owner:ZHEJIANG LAB

Fault rapid positioning method and device in MGX system

The invention provides a fault rapid positioning method and device in an MGX system, and belongs to the field of server system management and fault diagnosis. The method comprises the following steps: determining an idle universal serial bus interface on a mainboard, and selecting a target USB interface and a fixed IP network segment configuration corresponding to the target USB interface from the idle USB interface; the BMC establishes a physical connection with a device where the HMC is located through the target USB interface, and constructs a virtual local area network channel; in the virtual local area network channel, performing out-of-band communication between the BMC and the HMC; the BMC obtains list information of each component in the MGX system from the HMC, and classifies the list information according to component types; the HMC pre-collects and stores the state information of each category according to the classification result; and the BMC periodically reads the state information based on the Redfish protocol, and identifies the fault in the MGX system according to the state information. According to the method and the device provided by the invention, the fault can be quickly and accurately positioned.
Owner:ENGINETECH COMPUTER CO LTD

Thermal power plant equipment defect intelligent monitoring and early warning method and system based on multi-modal large model

The invention relates to a thermal power plant equipment defect intelligent monitoring and early warning method based on a multi-modal large model. The method comprises the following steps: collecting operation data, namely multi-modal data, of equipment in real time; the method comprises the following steps of: constructing a multi-mode Transform model; carrying out cross-modal data fusion and analysis; calculating the abnormal degree of the equipment; future equipment state prediction is carried out, and the probability of potential fault occurrence is predicted; and carrying out optimization on the multi-mode Transform model. The invention further discloses a thermal power plant equipment defect intelligent monitoring and early warning system based on the multi-mode large model. According to the method, the multi-modal Transform model is adopted, cross-modal feature fusion is carried out in combination with the image, sound, sensor and text data, the accuracy is higher, and the false alarm rate and the missing report rate are lower; trend analysis is carried out on the equipment state, the fault occurrence time can be predicted, the prediction advance is greatly improved, and the probability of sudden faults is reduced through combination of anomaly detection and prediction. In different thermal power plants and different devices, the migration adaptability is high.
Owner:ANHUI ELECTRIC POWER DESIGN INST CEEC

Root cause analysis method and device based on large language model, equipment and medium

The invention discloses a root cause analysis method and device based on a large language model, equipment and a medium. The method comprises the steps that a target entity and a topological relation are extracted in response to a root cause analysis request; searching matched historical root cause cases in a vector database, inputting the topological relation, the historical root cause cases and cue words into a large language model to generate a doubtful point list, and extracting downstream entities in the doubtful point list; calling a detection tool to collect entity diagnosis data and identify abnormal downstream entities; returning to execute the operation of retrieving the historical root cause case until a preset iteration ending condition is met; and inputting the final doubtful point list, the abnormal diagnosis data, the topological relation and the historical root cause case into the large language model again to obtain a root cause reasoning result and generate a root cause report. According to the embodiment of the invention, through the full-chain design of natural language understanding, topological constraint, historical cases, dynamic detection and iterative reasoning, the fault positioning efficiency and the root cause accuracy are improved, and the operation and maintenance labor cost and the service fault time consumption are remarkably reduced.
Owner:BEIJING YOUTEJIE INFORMATION TECH

Hardware detection process exception handling method and device and electronic equipment

The invention discloses a hardware detection process exception handling method, a hardware detection process exception handling device and electronic equipment, and relates to the technical field of computers, a hardware detection process is split into a plurality of independent execution units according to various detection functions supported by the hardware detection process, and state monitoring is carried out on each execution unit in the running process, so that the hardware detection process is more accurate. When a certain execution unit is abnormal, the abnormal execution unit is processed in a targeted manner, so that the abnormality of the process is processed fundamentally, the hardware detection process can be recovered to a normal operation state in time, and the situation that hardware faults in a system are diffused due to the fact that the hardware faults cannot be detected in time is avoided; and a foundation is laid for improving the hardware safety of the to-be-detected computing equipment.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Fault reason determination method and device, storage medium, electronic equipment and computer program product

The embodiment of the invention provides a fault cause determination method and device, a storage medium, electronic equipment and a computer program product, and relates to the field of data analysis, and the method comprises the steps: receiving a fault analysis request sent by a target object, and determining a preliminary fault cause list corresponding to the fault analysis request; an analysis task corresponding to the preliminary fault reason in the preliminary fault reason list is determined, the analysis task is executed to obtain abnormal data corresponding to the analysis task, and the abnormal data at least comprises one of abnormal log data, abnormal index data and abnormal service data; and determining the fault contribution evaluation information corresponding to the plurality of candidate fault reasons according to the abnormal data, and determining the target fault reason of the target fault in the plurality of candidate fault reasons according to the fault contribution evaluation information, thereby solving the problem of relatively long time consumption caused by determining the fault reason depending on personal experience in related technologies.
Owner:JINAN INSPUR DATA TECH CO LTD

Application state dynamic diagnosis and automatic repair method and system

The invention relates to the technical field of application state diagnosis, in particular to an application state dynamic diagnosis and automatic repair method and system.The method comprises the steps that multi-source heterogeneous logs are collected and standardized in real time, key fields are extracted, context information is injected, and structured log data are generated; inputting the structured log data into a dynamic anomaly detection model, constructing a dual-path detection mechanism based on LSTM time sequence analysis and a graph neural network, identifying an abnormal mode and positioning a fault root cause; according to the output of the anomaly detection model, a repair action is triggered in a grading manner through an intelligent repair strategy engine; in the repairing process, system state changes are stored and recorded through a pre-writing type redundancy log, automatic rollback during abnormity is achieved on the basis of check point information, and data consistency and system stability are guaranteed. The problem of service interruption or data inconsistency possibly caused by traditional automatic repair is avoided, and the reliability of automatic operation and maintenance is improved.
Owner:SHANDONG ARTAPLAY INTELLIGENT TECH CO LTD

System fault diagnosis method, device, medium and program product

The invention discloses a system fault diagnosis method and device, a medium and a program product, and relates to the technical field of computers, and the method comprises the steps: extracting fault data fragments for system fault diagnosis from a system database based on fault description information, reducing the influence of huge and redundant information amount on useful information extraction, and improving the system fault diagnosis efficiency. According to the system fault diagnosis method and device, the fault knowledge fragments used for system fault diagnosis are extracted from the system knowledge base based on the fault description information, fault diagnosis is conducted according to the fault data fragments and the fault knowledge fragments, experience dependence is reduced, the technical problems that the system fault diagnosis process is tedious and low in efficiency are solved, and the technical effect of accurate fault diagnosis is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Server cluster operation and maintenance method based on multi-source heterogeneous data fusion and dynamic knowledge graph

The invention provides a server cluster operation and maintenance method based on multi-source heterogeneous data fusion and a dynamic knowledge graph, and the method comprises the following steps: collecting a performance index, a log text and topological structure data of a server cluster, splicing the performance data and the log data based on a unified time window, and generating a multi-modal feature sequence; and analyzing the sequence by using an unsupervised deep learning model, constructing a dynamic health baseline, and generating a health degree portrait through the deviation with real-time data. When an exception is detected, mapping an exception event into a dynamic topological graph constructed based on a topological structure; analyzing a fault propagation probability between nodes by using a graph neural network algorithm, positioning a root cause node, and generating a disposal strategy to execute disposal operation; and collecting the processed recovery data as a feedback signal, and updating the deep learning model by using incremental learning. The method has the beneficial effects that the fault discovery accuracy is improved, the alarm storm is effectively inhibited, the root cause is directly positioned, and the model self-iteration adaptability is higher.
Owner:金品计算机科技(天津)有限公司 +1

Centralized log visualization for analysis debugging in cluster networks

A multi-node, multi-container cluster system that generates, aggregates, and manages log files from services and components to be used for audit logs and to debug and perform other serviceability tasks provided by a vendor of the cluster system. Logs are collected from all components of the system and aggregated into a consistent format for user analysis and debugging. Embodiments provide a comprehensive way to parse and index vast numbers of log files that can then be packaged and displayed to a user in a way that facilitates analysis and debugging and / or efficient input to appropriate debugging programs.
Owner:DELL PROD LP

Big data platform storage data isolation method in SaaS mode

The invention discloses a big data platform storage data isolation method in a SaaS mode, and relates to the technical field of big data, and the method comprises the steps: inputting a storage situation data set into a causal graph neural network, capturing a data time sequence mode and causal association features through a feature extraction layer, carrying out the multi-hop neighborhood feature aggregation through a feature fusion layer, and generating a storage anomaly detection vector; inputting the stored anomaly detection vector into a causal inference engine, executing risk quantification by using an improved causal inference tree algorithm, obtaining a causal effect score, carrying out risk division through a three-level threshold, generating an anomaly risk level, carrying out entropy calculation on the anomaly risk level by using a Shannon entropy formula, obtaining an anomaly entropy value, and obtaining an anomaly result. Carrying out interval classification on the abnormal entropy value to form a sensitivity level; according to the invention, through the constructed causal graph neural network, the improved causal inference tree algorithm and the analytic hierarchy process, dynamic identification of abnormal risks is realized, and redundancy isolation of low-risk data is also avoided.
Owner:ANHUI VALLEY DATA TECHNOLOGY CO LTD

Systems and methods for using multi-tiered guardrail architecture to generate dynamic conversational responses in sparse data environments

Systems and methods for uses and / or improvements to artificial intelligence applications, particularly in the area of generating conversational dynamic responses. As one example, systems and methods are for generating conversational dynamic responses using a multi-tiered guardrail architecture. As one example, systems and methods are for generating conversational dynamic responses using a multi-tiered guardrail architecture in data sparse environments.
Owner:CAPITAL ONE SERVICES LLC

Intelligent operation and maintenance management method based on big data algorithm

The invention relates to the technical field of big data, in particular to an intelligent operation and maintenance management method based on a big data algorithm, and the method comprises the steps: constructing and continuously updating a dynamic fault association graph through inputting multi-source heterogeneous operation and maintenance data; starting full-graph scanning based on a predefined period, detecting an abnormal topological structure through a graph pattern recognition algorithm, and marking potential risk nodes; executing dynamic influence diffusion simulation on the potential risk nodes, calculating a business influence severity quantized value after the fault, and marking fault propagation vulnerabilities according to the quantized value; taking the potential risk node as a starting point, executing a reverse traceability algorithm for preferentially exploring a path pointing to a fault propagation vulnerable point, and outputting a fault propagation path and a source fault node identifier; and finally generating and executing a fault processing strategy. The process solves the problem that traditional operation and maintenance cannot quantitatively evaluate and discriminate the highest priority disposal object from numerous potential risks, and realizes accurate positioning and active prevention and control of weak links of fault propagation.
Owner:HANGZHOU FOCUS TECHNOLOGY CO LTD

Server fault diagnosis method and device, storage medium and program product

The invention discloses a server fault diagnosis method and device, a storage medium and a program product, and relates to the technical field of server fault localization, and the method comprises the steps: carrying out the preprocessing of original log data, determining a state space according to the obtained preprocessed log data, a fault classification model containing a value network and an action network is obtained through reinforcement learning in advance, the value network in the fault classification model is used for carrying out value evaluation on the state space, and the action network in the fault classification model is used for carrying out action probability distribution calculation on the state space according to a target value obtained through value evaluation; and determining the target fault type according to the calculated probability distribution result, thereby solving the technical problem that new features or anomalies cannot be effectively judged, effectively improving the quality and integrity of log data, improving the server fault diagnosis efficiency and accuracy, realizing intelligent identification of the fault type, and improving the fault diagnosis efficiency and accuracy. The method has the technical effects of good training stability and generalization ability.
Owner:ZHENGZHOU YUNHAI INFORMATION TECH CO LTD

Software fault repair method and system fused with intelligent analysis

The invention belongs to the technical field of computers, and particularly relates to a software fault repairing method and system fused with intelligent analysis, which comprises the steps of collecting a multi-level running log and performing structured preprocessing, constructing a dynamic calling graph through a time sequence encoder and a graph neural network, inferring a fault root cause in combination with a Bayesian causal inference model, and repairing a fault fault according to the fault root cause. And matching the repair strategy to generate an atomization instruction sequence, and deploying the atomization instruction sequence to a production system after sandbox environment verification. The system comprises a log acquisition module, a feature coding module, a graph construction module, a causal reasoning module, a strategy matching module, an instruction generation module, a sandbox verification module, a deployment feedback module and the like. Through end-to-end intelligent analysis and a closed loop verification mechanism, the fault positioning precision and the repair safety are remarkably improved, system self-evolution is supported, and operation and maintenance are promoted to be transformed from passive response to active autonomy.
Owner:HARBIN BLACK ANT TECHNOLOGY CO LTD

System and method for ai based incident impact and root cause analysis

ActiveUS20250245091A1Fault responseMachine learningLinguistic modelIncident report
A system and method for reducing data record processing in incident report generation is provided. The method includes: accessing a plurality of event records, each event record generated based on an event in a computing environment; parsing each event record based on a predetermined data field; extracting from each predetermined data field a data value; correlating a group of event records of the plurality of event records based on at least an extracted data value; generating an incident data record based on the extracted data values of the correlated group of event records; generating a prompt based on the incident data record; and generating an incident report by configuring a large language model (LLM) to execute the generated prompt.
Owner:BIGPANDA INC