Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4977results about "Non-redundant fault processing" patented technology

Language model assisted error analysis system

Computer-implemented systems and methods including language models for explaining and resolving code errors. A computer-implemented method may include: receiving or accessing a log comprising an error message, the error message indicating an error in code; determining the error message from the log; determining a context associated with the error; generating a prompt for a large language model (“LLM”), the prompt comprising at least: the error message, and the context associated with the error; transmitting the prompt to the LLM; and receiving an output from the LLM in response to the prompt, the output comprising at least: an explanation of the error message, and a suggested fix for the error.
Owner:PALANTIR TECHNOLOGIES INC

Distributed accounting asynchronous data processing method and system based on MEMO state machine

The invention relates to the technical field of computer communication, provides a distributed accounting asynchronous data processing method and system based on an MEMO state machine, and is used for improving the intelligent adaptability of a distributed accounting network to a dynamic service scene. The method comprises the following steps: acquiring a transaction request set in a distributed accounting network, extracting an asynchronous processing trigger event of the transaction request set, analyzing the asynchronous processing trigger event, generating a state transition instruction sequence containing M state transition paths associated with K transaction data blocks, a parallel prediction module based on an MEMO state machine performs asynchronous processing on the M state transition paths to generate N prediction state vectors and corresponding conflict detection marks; performing feature fusion on the N prediction state vectors according to the conflict detection mark to generate a global state updating matrix; and writing the global state updating matrix into a consensus node set of the distributed accounting network, and generating an asynchronous processing state log matched with the transaction request set.
Owner:CHINA RONGXIN CLOUD TECH CO LTD

Distributed computing power scheduling intelligent optimization method and system

The invention relates to the technical field of distributed computing, and discloses a distributed computing power scheduling intelligent optimization method and system.The distributed computing power scheduling intelligent optimization method comprises the steps that a multi-level system dependency graph and a causal relationship model are constructed; a multi-dimensional anomaly detection and fault prediction technology is utilized to monitor the system state in real time and predict potential faults; when an anomaly is detected or a potential fault is predicted, performing key path analysis by combining a multi-level system dependency graph and a causal relationship model, identifying a fault propagation path and a key influence node, executing a fault isolation strategy, blocking fault propagation and implementing preventive resource scheduling; implementing intelligent repair and resource recombination; according to the method, early prediction, accurate positioning and efficient processing of faults in the distributed system are realized, allocation and scheduling of computing power resources are optimized, and the reliability, availability and performance of the system are improved.
Owner:ZHONGTONG INFORMATION SERVICE CO LTD +1

Integrated AI-Driven System for Automating IT and Cybersecurity Operations

A system for automating information technology software actions using advanced AI techniques including a processing system, storage medium, a communications interface, a user interface, a natural language processing model or neural network, operable to interface with at least one of an embedded prompt or chatbot prompt and hosted on a control node which communicates with one or more nodes comprised by a node network via the communications interface, and program instructions on storage medium that direct the processing system to receive an instruction from the user interface, process the instruction using the one or more natural language processing model or neural network, along with one or more AI agents, and execute the instruction one of locally or on a node of the node network via the communications interface.
Owner:AUGSTRA LLC

Self-adaptive cloud management platform system based on intelligent resource scheduling and container arrangement

The invention discloses a self-adaptive cloud management platform system based on intelligent resource scheduling and container arrangement, and relates to the field of computer information management. The system comprises a refined resource scheduling and adaptive optimization module, a containerized application life cycle management and dynamic container arrangement module, a high-precision operation and maintenance monitoring and self-healing mechanism module based on big data analysis, and a dynamic resource allocation and elastic scaling strategy module of an intelligent scheduling engine. The system takes a containerization technology as a core, realizes centralized management and monitoring of cloud computing resources, can realize dynamic intelligent resource scheduling and optimization, accelerates application deployment, improves system flexibility, strengthens operation and maintenance monitoring and platform safety guarantee, and comprehensively improves resource optimization and cost effectiveness. The resource scheduling efficiency is improved, the application deployment is simplified, the operation and maintenance monitoring is enhanced, and efficient resource management is realized through an adaptive optimization technology.
Owner:CHINA IND INTERNET RES INST

Method and system for intelligent capacity planning in hybrid cloud environment

The invention relates to the technical field of cloud computing, in particular to an intelligent capacity planning method and system in a hybrid cloud environment. According to the method and the system for intelligent capacity planning in the hybrid cloud environment, heterogeneous resources in a hybrid cloud are modeled by using a declarative description language, a dependency relationship among the resources is identified, and a directed acyclic graph (DAG) is constructed; the resource operation tasks are divided in batches, and the tasks without the dependency relationship are executed in parallel; synchronizing the running state of each resource in real time, and performing state consistency verification and abnormity marking; triggering a retry mechanism when an exception occurs; and after the resource arrangement process is completed, summarizing data in the whole process, and performing performance evaluation and strategy optimization on resource operation. According to the method and the system for intelligent capacity planning in the hybrid cloud environment, the use condition of resources can be monitored and analyzed in real time, the resource demand can be accurately predicted, a reasonable resource scheduling strategy can be generated, manual intervention is reduced, and elastic expansion and optimal configuration of cloud resources are realized.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Intelligent data quality monitoring method and system

The invention provides an intelligent data quality monitoring method and system. The method comprises the following steps: performing data processing on multi-source heterogeneous data by utilizing a unified data model; inputting historical data as a training set into the pre-training model for training optimization to obtain an optimized pre-training model, and performing threshold determination on the standardized distribution data stream by using the optimized pre-training model in combination with an adaptive threshold module; inputting the processed abnormal event list into a context awareness and dynamic priority management module to carry out de-duplication optimization transmission; carrying out root analysis on the abnormal event list after the de-weighting optimization by utilizing a root analysis tool; generating a repair strategy for the root analysis result by using the assistance of a repair suggestion engine; and inputting the repair strategy and the visual report into a cross-platform monitoring interface, and generating a repair suggestion and an early warning notification. According to the method, the data consanguinity map is constructed based on the graph database, minute-level traceability of abnormal events is achieved, and the troubleshooting time is shortened by 50%.
Owner:JIANGXI TONGRUI INFORMATION TECH CO LTD

Memory fault repairing method and device, equipment, medium and computer program product

The invention discloses a memory fault repairing method and device, equipment, a medium and a computer program product, and relates to the technical field of computers.The memory fault repairing method includes the steps that memory error information sent by a memory controller is obtained, a fault target memory page can be positioned, and a standby memory page is obtained from a standby memory pool; the memory address mapping table is updated, the physical address of the target memory page with the fault is mapped to the physical address of the standby memory page, the repairing process does not depend on triggering of a system management interrupt mechanism, memory repairing in the system running stage is achieved, the system downtime caused by memory fault repairing is shortened, and therefore the memory repairing efficiency is improved. The problem of high delay in memory fault processing can be solved, and the technical effect of improving the memory fault repairing efficiency is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method and system for automatically testing reliability of solid state disk based on multiple threads

The invention relates to the technical field of hard disk testing and verification, in particular to a multi-thread-based solid state disk reliability automatic testing method and system.The method comprises the steps that firstly, SMART information is deeply analyzed through microsecond-level high-granularity continuous performance monitoring, and multi-thread parallel processing is assisted; according to the method, fine performance fluctuation of the solid state disk under the concurrent load can be quickly captured, a fault mode can be identified, then early warning is realized by utilizing the extracted multi-dimensional features and a machine learning model, and a detailed fault diagnosis report is generated; and through dynamic error correction code strength verification and data integrity verification under pressure, an internal error correction mechanism of the solid state disk is actively detected and optimized. And finally, in combination with prediction reliability modeling, the system can estimate the remaining service life and predict faults, and provides product optimization suggestions for design, manufacturing and firmware optimization of the solid state disk, so that automation, intelligence and full life cycle management of the fault detection reliability of the solid state disk are realized.
Owner:GUIZHOU SHUSUAN INTERNET TECHNOLOGY CO LTD

Multi-node database disaster recovery backup system

The invention relates to the technical field of database disaster recovery backup, and discloses a multi-node database disaster recovery backup system. The system comprises a data fragmentation module, a real-time synchronization module, a fault detection module, a data recovery module and a strategy optimization module. The data fragmentation module performs logic fragmentation based on a consistent Hash algorithm to realize load balancing and redundant backup; the real-time synchronization module captures data change of the main node by using an incremental log synchronization technology; the fault detection module evaluates the health condition of the node through a self-adaptive heartbeat detection mechanism; the data recovery module recovers data by means of a multi-version snapshot recovery algorithm; and the strategy optimization module generates an optimal disaster recovery backup strategy by applying a dynamic weight adjustment algorithm. The system solves the problems of non-uniform data distribution, low synchronization efficiency, untimely fault detection, incomplete data recovery, lack of strategy optimization and the like in multi-node database disaster recovery backup, improves system performance, stability and data security, and meets diversified disaster recovery backup requirements.
Owner:QUANZHOU INST OF INFORMATION ENG

Data center intelligent operation and maintenance and fault prediction system and method

The invention provides an intelligent operation and maintenance and fault prediction system and method for a data center, and the system comprises a self-healing decision module, a closed-loop verification module, an automatic execution module, a monitoring feedback module, and a strategy optimization module, and is characterized in that the self-healing decision module is used for determining the type and severity of a fault. The self-healing decision tree is generated according to the fault type and severity in combination with expert knowledge and historical data, the self-healing accuracy and safety are improved, the self-healing decision is executed through the automatic execution module, manual intervention is reduced, the operation and maintenance efficiency is improved, the self-healing process is monitored through the monitoring feedback module, feedback information is collected, and the operation and maintenance efficiency is improved. The self-healing effect is evaluated, the correctness and effectiveness of self-healing operation are ensured, a self-healing strategy is continuously optimized through a strategy optimization module according to feedback information and new data, the self-adaptive capacity and long-term performance of the system are improved, the system can continuously learn and adapt to a new fault mode, and therefore the self-healing capacity is continuously improved.
Owner:HUAZHANG DATA (SHENZHEN) CO LTD

System and method for automated identification and assisted repair of can-related faults across multiple vehicle ecus

A diagnostic system and method for identifying and resolving vehicle Controller Area Network (CAN) bus faults is disclosed. The system connects to a vehicle's data link connector (DLC) and automatically retrieves diagnostic trouble codes (DTCs) from multiple electronic control units (ECUs). It filters the DTCs to identify those related to CAN communication errors, groups repeated DTCs across ECUs to highlight systemic faults, and assigns repair priority based on status, scope, and criticality. A user interface displays prioritized DTCs with definitions, affected systems, possible causes, and repair guidance. Upon technician input, the system re-scans the network to confirm resolution and updates the display accordingly. The invention supports semantic and conceptual recognition of communication faults, enabling robust, language-agnostic analysis. This system reduces diagnostic time, improves repair accuracy, and supports multilingual or manufacturer-specific DTC definitions without requiring external infrastructure or advanced user expertise.
Owner:INNOVA ELECTRONICS CORP

Electric power information system inspection and resource allocation method based on knowledge base and intelligent agent

The invention relates to the field of intelligent inspection and resource allocation, in particular to an electric power information system inspection and resource allocation method based on a knowledge base and an intelligent agent, and the method comprises the steps: collecting the data of an electric power information system, carrying out the heterogeneous data integration, and constructing the knowledge base of the electric power information system; constructing an intelligent agent oriented to the electric power information system based on a large language model, and performing RAG on the intelligent agent based on a knowledge base and open source knowledge to enhance the performance of the intelligent agent; acquiring a routing inspection demand in a natural language form, automatically generating a routing inspection script by the intelligent agent based on the routing inspection demand, and compiling the routing inspection script to generate a routing inspection task; the intelligent agent automatically executes the inspection task and compares the inspection result with the knowledge base to generate an inspection report; and the intelligent agent performs intelligent analysis on the inspection report, and performs resource allocation on abnormal items in the inspection report in combination with the knowledge base. According to the method, the operation and maintenance efficiency of the electric power information system is effectively improved, and the operation and maintenance cost is reduced.
Owner:INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD

Conversational automated event response and remediation

ActiveUS12487875B1Biological modelsHardware monitoringIncident management (ITSM)Software engineering
In one embodiment, a computer-implemented method executed using one or more processors of an incident management system comprises receiving a notification of an incident associated with a computer system, and in response to receiving the notification: extracting an error message from the notification; reading a set of computer program code changes that have been implemented in the computer system; matching the error message to the set of computer program code changes and outputting a set of one or more candidate code changes that may correspond to the error message; based on the set of candidate code changes, generating an automatic remediation for the incident; and using the one or more processors of the incident management system, executing the automatic remediation on the computer system.
Owner:PAGERDUTY INC

Storage hard disk remote diagnosis system and method based on Internet of Things

The invention discloses a storage hard disk remote diagnosis system and method based on the Internet of Things, and relates to the technical field of health management of the Internet of Things and storage equipment. The system comprises a data fusion module, a causal analysis module, a risk analysis module and a map construction module. The data fusion module collects SMART parameters, IO operation time sequence data and environment data through Internet of Things equipment, dynamically distributes multi-source data weights by using an attention mechanism, and extracts hard disk health state features. And the causal analysis module is combined with the multi-scale time sequence convolutional network and the Bayesian causal network to identify periodic abnormal fluctuation and generate a fault root cause analysis report. The risk analysis module matches historical cases through federal incremental learning, calculates a hard disk fault risk index and generates an early warning signal. And the atlas construction module optimizes resource isolation, data migration and response paths according to the risk indexes, generates a visual operation and maintenance atlas, provides fault positioning, risk links and repair priorities, and improves the operation and maintenance management efficiency.
Owner:SHENZHEN SANSHANG SCIENCE & TECHNOLOGY CO LTD

Fault diagnosis and recovery verification method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a fault diagnosis and recovery verification method and device, equipment and a medium. Inputting the candidate event and the context state into an intelligent diagnosis model to obtain a diagnosis result and a fault type, selecting a recovery path in a recovery strategy library based on the diagnosis result, generating a recovery execution instruction, creating a verification experiment based on the instruction, and injecting drill data to obtain verification data, and forming a verification conclusion according to the verification data, associating a recovery execution instruction, and issuing the instruction in the execution environment to complete recovery operation. According to the method, through multi-source monitoring data processing, intelligent diagnosis model reasoning, strategy library path selection, verification experiment verification and execution environment automatic recovery, a full-link automatic process from detection, diagnosis, recovery to verification is constructed, and accurate diagnosis and dynamic recovery in a complex environment are achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Database abnormal root cause positioning method and system based on knowledge graph

The invention discloses a database abnormal root cause positioning method and system based on a knowledge graph, and belongs to the technical field of database management and intelligent operation and maintenance, and the method is implemented by the following steps: constructing a database domain knowledge graph, integrating multi-source data, and storing by adopting a graph database; carrying out anomaly detection and feature extraction, finding anomalies through time sequence index detection, query mode analysis and log anomaly recognition, and extracting numerical features, semantic features and graph features; performing root cause reasoning and positioning based on the knowledge graph, and determining abnormal root causes by utilizing rule reasoning, graph traversal reasoning and probability graph model reasoning and combining with multi-evidence fusion; generating and executing a repair strategy, matching a solution from the repair strategy library according to the root cause, and automatically executing and monitoring a repair effect; and dynamically updating the knowledge graph. According to the method, automatic detection, accurate root cause positioning and targeted repair strategy generation of database exceptions can be realized, the manual operation and maintenance cost is reduced, and the exception handling efficiency and accuracy are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Supercomputing center emergency response method and system based on data fusion analysis

The invention provides a supercomputing center emergency response method and system based on data fusion analysis, and the method comprises the steps: obtaining a multi-source real-time monitoring data set of a supercomputing center, carrying out the data fusion processing of the multi-source real-time monitoring data set, generating a multi-dimensional fusion feature set corresponding to the supercomputing center, and carrying out the emergency response of the supercomputing center based on a preset anomaly detection strategy. Performing real-time anomaly detection analysis on the multi-dimensional fusion feature set, generating an abnormal event set of the supercomputing center, calling an emergency response strategy template matched with an abnormal type identifier according to the abnormal type identifier of each abnormal event in the abnormal event set, and generating an emergency response strategy set of the supercomputing center; and executing a target emergency response strategy in the emergency response strategy set, obtaining execution feedback data of the target emergency response strategy in real time, and updating matching parameters of the emergency response strategy template. According to the invention, the self-adaptive recovery capability and system stability of the supercomputing center for coping with complex abnormal events can be enhanced.
Owner:POWERCHINA RAILWAY CONSTR +2

Comprehensive financial IT operation and maintenance management system and method based on artificial intelligence

The invention provides a comprehensive financial IT operation and maintenance management system and method based on artificial intelligence, and the system comprises an intelligent monitoring and predictive maintenance module which is used for achieving the real-time monitoring, abnormality recognition and performance prediction of an IT system through collection of multi-dimensional time series data, and when an abnormal condition is monitored or a prediction result does not meet a preset requirement, the intelligent monitoring and predictive maintenance module carries out the operation and maintenance of the IT system; if so, triggering intelligent early warning; the AI-driven fault diagnosis and self-healing module is used for carrying out intelligent fault analysis and repair in combination with IT system logs and historical fault data based on the detected abnormal condition; the self-adaptive resource management module is used for predicting a future load trend and a future performance trend, and performing self-adaptive resource management based on the predicted future load trend and the future performance trend in combination with cost and energy consumption factors; and the security operation and maintenance automation module is used for monitoring security threats of the IT system in real time and automatically executing security strategies.
Owner:SHANGHAI GREAT WISDOM

Method and system for detecting anomalous sub-sequences in metadata

A method for managing an anomaly in a client includes: obtaining, by an analyzer, historical metadata (HM); obtaining, by the analyzer, an error description that is associated with the HM; analyzing, by the analyzer, the HM to generate a first data frame (DF); generating, by the analyzer, a second DF and a third DF based on the first DF, in which the second DF and the third DF are sent to an engine; generating, by the analyzer, a fourth DF based on the first DF and error description, in which the fourth DF is sent to the engine; tuning, by the engine, an anomaly detection model (ADM) to obtain a tuned ADM using: a first target parameter (TP) and the second DF; a second TP and the third DF; a third TP and the fourth DF; and initiating, by the engine, notification of an administrator about the tuned ADM.
Owner:DELL PROD LP

Alarm event root cause analysis method, device and equipment based on large language model and service topology

The invention relates to an alarm event root cause analysis method, device and equipment based on a large language model and business topology, and the method comprises the steps: based on an analysis engine, according to an alarm event sent by an alarm system, obtaining an association node set, collecting multi-source data, generating a preliminary context packet, carrying out the extraction of abnormal information according to the preliminary context packet, and carrying out the analysis of the abnormal information. And generating a refined context package by removing irrelevant node information, and generating an interpretable root cause analysis report through large language model constraint and reasoning based on the constructed large language cue word template. According to the method, the initial context packet covering the fault propagation link is constructed based on the service topological graph, rapid modeling of the context is achieved, the high-frequency error mode template is generated through clustering, irrelevant nodes are dynamically pruned to generate the refined context packet, accurate focusing and data filtering are achieved, and the fault propagation efficiency is improved. And embedding the refined evidence chain through a structured cue word template, and driving the large language model to output an interpretable root cause report.
Owner:BEIJING ZHIWEI YINGXUN NETWORK TECH CO LTD

Prioritized fault remediation

Methods, computer program products, and systems are presented. The method computer program products, and systems can include, for instance: iteratively examining logging data; detecting multiple faults in a computer environment in dependence on the examining of the logging data; generating for respective ones of the detected multiple faults one or more candidate remediation to provide a set of candidate remediations for the computer environment; prioritizing remediations defining the set of candidate remediations from the generating and ordering the remediations in a remediation queue according to an order of the prioritizing; and deploying remediations according to the ordering of remediations in the remediation queue.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Software defect information fusion method and system based on multi-source data

The invention provides a software defect information fusion method and system based on multi-source data. The method comprises the following steps: acquiring static code characteristics, a runtime log sequence and user feedback from cross-platform software; constructing a code dependency graph based on static code features and a historical defect library, defining a log structure through code logic association, and performing space-time alignment with defect trigger nodes fed back by a user to generate a context matrix; synchronously acquiring physical parameters of hardware nodes, and dynamically coupling the physical parameters with the matrix to generate a state evolution graph; fusing the code dependency graph, the state evolution graph and user feedback, and extracting cross-modal defect features; and generating defect positioning probability distribution based on the features, and determining a repair scheme by combining physical parameter anomaly analysis. Through integration of codes, operation states and user feedback, accurate positioning and dynamic repair of defects are realized. According to the technical scheme provided by the invention, the precision and reliability of software defect information fusion can be improved.
Owner:BEIJING ZHONGKE CHANGFENG TECHNOLOGY CO LTD

Computing power sharing system and method in multi-tenant environment

The invention relates to the technical field of computing power sharing, in particular to a computing power sharing system and method in a multi-tenant environment, and the system comprises a computing power monitoring module, a load analysis and prediction module, a task priority judgment module, a resource optimization configuration module and a fault management module. According to the invention, by implementing real-time resource monitoring and data analysis and accurately capturing the speed and mode of resource consumption, the resource allocation is more accurate, the waste of computing power resources is avoided, the response capability is improved, the load prediction is generated by using historical and real-time data, the computing power resources can be ensured to be fully utilized in the peak period of demand, overload is avoided, and the service life of the system is prolonged. Through intelligent analysis of the resource utilization rate and the predicted completion time and optimization of priority ranking of the tasks, the emergency tasks can be rapidly responded, implementation of a real-time computing power distribution strategy is achieved, high-efficiency computing power resource use is guaranteed, the stability and reliability of the system are improved, and the cost problem caused by fault response delay is reduced.
Owner:YUNJU DATA TECH (SHANGHAI) CO LTD

System and method for root cause change detection

A system and method for root cause analysis in incident processing, including root cause change detection, is presented. The method includes processing a plurality of event records, including a plurality of alert records and a plurality of change records, each event record generated based on an event in a computing environment; parsing each alert record based on a predetermined data field; extracting from each predetermined data field a data value; correlating a group of alert records of the plurality of alert records based on at least an extracted data value; generating an incident data record based on the extracted data values of the correlated group of alert records; detecting a change record of the plurality of change records related to the incident data record; determining that the change record is a root cause change of the incident data record; and initiating a mitigation action based on the root cause change.
Owner:BIGPANDA INC

Cross-system fault diagnosis method and system combined with multi-dimensional anomaly detection

The invention discloses a cross-system fault diagnosis method and system combined with multi-dimensional anomaly detection, and relates to the technical field of fault diagnos.The method comprises the steps that a graph neural network with a liquid time constant network unit as a node is constructed through a dynamic topology dependency relationship and a multi-dimensional key performance index flow; each unit describes state evolution through a coupled ordinary differential equation system, and a liquid state time constant can be adaptively adjusted. A time back propagation algorithm is adopted to train a model to learn a normal behavior track contour reference, and anomaly is detected through a dynamic time warping distance. And determining a fault propagation path and a root cause through anti-fact intervention and forward integral solution. And generating an optimal diagnosis action sequence in a liquid graph neural network simulation environment, and calculating a reward value based on execution efficiency, accuracy and a repair effect to carry out strategy optimization. The abnormal detection accuracy and the root cause positioning precision are improved, the fault repair time is shortened, the operation and maintenance cost is reduced, and an intelligent fault diagnosis solution is provided for a complex information technology system.
Owner:SHANGHAI QINGCHUANG INFORMATION TECH CO LTD