Fault monitoring method and device based on artificial intelligence, equipment and medium
By collecting, structuring, and standardizing multi-dimensional data, and utilizing artificial intelligence analysis modules for fault detection and visualization, the system addresses the shortcomings of existing systems in data processing and diagnostic accuracy. This enables efficient and accurate fault monitoring and diagnosis, adapting to complex web environments and multi-source heterogeneous data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN HEALTH INSURANCE CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-15
AI Technical Summary
Existing AI-based fault monitoring and diagnosis systems are inadequate in terms of data processing capabilities, real-time response speed, diagnostic accuracy, and user interface interactivity. They are ill-suited to complex web environments and multi-source heterogeneous data, and cannot meet the efficient and accurate fault monitoring needs of the fintech and healthcare sectors.
Collect multi-dimensional data and store it in the cloud or on a local server. Perform structured analysis and standardization processing, extract key features using an artificial intelligence analysis module, and infer the cause of the fault through a fault association network. Then, generate fault cause information and display it on a visual operation interface.
It enables efficient processing of multi-source heterogeneous data, quickly identifies potential faults and provides repair suggestions, improves the real-time performance and accuracy of fault monitoring and diagnosis, meets the operational needs of users at different levels, and ensures system stability and user experience.
Smart Images

Figure CN122053344A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a fault monitoring method, apparatus, device, and storage medium based on artificial intelligence. Background Technology
[0002] In the operational system of modern e-commerce websites, website stability and performance are directly related to user experience and business growth, forming the core foundation for e-commerce platform development. Current traditional fault monitoring and diagnosis methods rely heavily on human experience or simple strategy systems, generally suffering from insufficient real-time performance, low diagnostic accuracy, and limited fault analysis capabilities, making them ill-suited to the complex operational needs of e-commerce websites. With the rapid development of artificial intelligence technology, AI-based fault monitoring and diagnosis systems have become a research hotspot in this field, but existing systems still have many areas for improvement. First, data processing capabilities are limited, making it difficult to efficiently handle multi-source heterogeneous data in complex web environments; second, the real-time response speed and diagnostic accuracy of fault diagnosis models still have room for improvement; third, there is a lack of comprehensive analysis and diagnosis capabilities for complex faults, making it difficult to handle compound faults caused by multiple factors; and fourth, the user interface lacks interactivity, failing to meet the operational and usage needs of users at different levels.
[0003] In the operation of online platforms in the fintech sector, the stability and performance of the platform are core elements for ensuring a good financial service experience and supporting continuous business growth. They are also a crucial foundation for compliant fintech operations and risk control. Currently, traditional fault monitoring and diagnosis methods in the industry largely rely on manual experience or simplistic strategy systems, which generally suffer from insufficient real-time performance, low diagnostic accuracy, and limited fault analysis capabilities. These methods are ill-suited to the complex operating environment and high security and reliability requirements of online fintech platforms. While AI-based fault monitoring and diagnosis systems currently in use in the industry have become a hot topic in technology research and implementation, they still face many pressing technical challenges: Data processing capabilities are significantly limited, making it difficult to efficiently handle multi-source heterogeneous data in the complex web environment of fintech businesses, and unable to integrate multi-dimensional data such as transaction data, system operation data, and user behavior data; the real-time response speed and diagnostic accuracy of fault diagnosis models do not meet industry requirements, making it difficult to match the high-frequency trading and time-sensitive nature of fintech businesses; there is a lack of comprehensive analysis and diagnosis capabilities for complex faults, making it impossible to accurately trace the source and conduct comprehensive analysis of complex faults caused by multiple factors such as systems, networks, and data in fintech businesses; and the user interface lacks interactivity, failing to meet the operational and usage needs of different levels of users in the fintech field, including technical maintenance, business operations, and risk management, hindering rapid and collaborative fault handling.
[0004] In the healthcare sector, the stability and performance of medical systems directly impact patient safety, efficiency, and the overall patient experience, serving as a core guarantee for healthcare services. Current traditional fault monitoring and diagnosis methods in the medical field largely rely on the experience and judgment of medical staff or simplistic maintenance strategies. These methods generally suffer from a lack of real-time performance, insufficient diagnostic accuracy, and limited depth of fault analysis, making them ill-suited to the complex system operation requirements of smart healthcare. With the gradual application of artificial intelligence technology in the medical field, AI-based medical system fault monitoring and diagnosis solutions have begun to attract attention. However, existing related systems still face many unresolved challenges: First, data processing capabilities are significantly limited. Medical scenarios involve massive amounts of heterogeneous data from multiple sources (such as electronic medical records, medical equipment operation data, treatment process data, and patient vital sign monitoring data), diverse formats, and high privacy requirements, making it difficult for existing systems to efficiently integrate and process the data. Second, fault diagnosis models lack real-time response and accuracy. Issues such as medical equipment failures, abnormal system data transmission, and treatment process delays all require immediate responses, but existing models often experience diagnostic delays or misjudgments, potentially hindering treatment. Third, there is a lack of comprehensive diagnostic capabilities for complex faults. Medical system faults are often caused by a complex interplay of factors such as equipment, data, and processes, making it difficult for existing systems to disentangle complex relationships and pinpoint the root cause. Fourth, user interface interactivity is lacking. Users in the medical field include professional maintenance personnel, medical staff, and other diverse groups. Existing interfaces often focus on technical presentation and do not adequately adapt to the operational needs of non-technical personnel, hindering rapid fault reporting and collaborative processing. Summary of the Invention
[0005] The main objective of this invention is to provide a fault monitoring method, apparatus, device, and storage medium based on artificial intelligence, aiming to solve the problems of low data processing efficiency and low accuracy in existing fault monitoring and diagnosis based on artificial intelligence.
[0006] To achieve the above objectives, the present invention provides a fault monitoring method based on artificial intelligence, comprising: Collect multi-dimensional data and store the multi-dimensional data on a cloud server and / or a local server; The multi-dimensional data is subjected to structured parsing and standardization to extract key features; An artificial intelligence analysis module is used to perform fault detection on the multi-dimensional data based on the key features, and fault data is generated. The fault data is input into a fault association network for inference, and fault cause information is output. Repair suggestions are generated based on the fault cause information. The fault cause information and repair suggestions are displayed on the visual operation interface.
[0007] Furthermore, to achieve the above objectives, the present invention provides an artificial intelligence-based fault monitoring device, comprising: The data acquisition module is used to collect multi-dimensional data and store the multi-dimensional data to a cloud server and / or a local server; The data processing module is used to perform structured parsing and standardization processing on the multi-dimensional data and extract key features; The fault detection module is used to perform fault detection on the multi-dimensional data based on the key features using the artificial intelligence analysis module, and generate fault data. The fault diagnosis module is used to input the fault data into the fault association network for reasoning, output fault cause information, and generate repair suggestions based on the fault cause information. The user interface module is used to display the fault cause information and repair suggestions on a visual operation interface.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an artificial intelligence-based fault monitoring program stored in the memory and executable on the processor, wherein when the artificial intelligence-based fault monitoring program is executed by the processor, it implements the steps of the artificial intelligence-based fault monitoring method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an artificial intelligence-based fault monitoring program, wherein the artificial intelligence-based fault monitoring program, when executed by a processor, implements the steps of the artificial intelligence-based fault monitoring method as described above.
[0010] Beneficial Effects: This invention relates to the field of artificial intelligence technology and discloses an AI-based fault monitoring method, comprising: collecting multi-dimensional data and storing the multi-dimensional data in a cloud server and / or a local server; performing structured parsing and standardization processing on the multi-dimensional data to extract key features; using an AI analysis module to perform fault detection on the multi-dimensional data based on the key features, generating fault data; inputting the fault data into a fault association network for inference, outputting fault cause information, and generating repair suggestions based on the fault cause information; and displaying the fault cause information and repair suggestions on a visual operation interface. This invention can be applied to network fault detection scenarios such as fintech and healthcare. By monitoring multi-dimensional data during website operation in real time, it can quickly identify potential faults and provide repair suggestions through artificial intelligence, achieving efficient fault early warning and handling, and realizing efficient and accurate fault monitoring and diagnosis. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based fault monitoring method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the fault monitoring method based on artificial intelligence according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the fault monitoring device based on artificial intelligence of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The fault monitoring method based on artificial intelligence provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server terminal via a network. The server terminal can collect multi-dimensional data from the user terminal and store the multi-dimensional data on a cloud server and / or a local server; it performs structured parsing and standardization processing on the multi-dimensional data to extract key features; it uses an artificial intelligence analysis module to perform fault detection on the multi-dimensional data based on the key features, generating fault data; it inputs the fault data into a fault association network for inference, outputs fault cause information, and generates repair suggestions based on the fault cause information; it displays the fault cause information and repair suggestions on a visual operation interface. This invention can be applied to network fault detection scenarios such as fintech and healthcare. By monitoring multi-dimensional data during website operation in real time, it quickly identifies potential faults and provides repair opinions through artificial intelligence, achieving efficient fault early warning and handling, and realizing efficient and accurate fault monitoring and diagnosis. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server terminal can be implemented using a separate server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the artificial intelligence-based fault monitoring method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2As shown, the fault monitoring method based on artificial intelligence proposed in this invention includes the following steps: S100. Collect multi-dimensional data and store the multi-dimensional data in a cloud server and / or a local server; S200. Perform structured parsing and standardization on the multi-dimensional data to extract key features; S300: Using an artificial intelligence analysis module, fault detection is performed on the multi-dimensional data based on the key features to generate fault data; S400: Input the fault data into the fault association network for inference, output fault cause information, and generate repair suggestions based on the fault cause information; S500: Display the fault cause information and repair suggestions on the visual operation interface.
[0016] In this embodiment, multi-dimensional data collection covers server operation logs, application operation logs, network performance data, and user behavior data. After collection, the data is securely and stably stored in cloud servers or local servers through various network transmission protocols such as HTTP and WebSocket, providing a complete and comprehensive data foundation for subsequent analysis and processing.
[0017] In the data processing and analysis stage, the collected multi-dimensional raw data is first cleaned by noise reduction, standardization and other operations to remove invalid, redundant or abnormal interference data to ensure data quality. Then, through professional feature extraction algorithms, key features that can reflect the system's operating status and potential fault signs are screened from the cleaned data. These key features are the core basis for subsequent fault detection and diagnosis, which can greatly improve the efficiency and accuracy of subsequent analysis.
[0018] The AI analysis module leverages machine learning models, deep learning algorithms, and natural language processing (NLP, a technology that uses computers to understand and analyze human language) to conduct comprehensive analysis of multi-dimensional data based on extracted key features. By monitoring data changes in real time, identifying abnormal behavior, performing time-series analysis to predict resource bottlenecks, and locating potential code problems based on code context and runtime data, it accurately detects various faults in the system, generating fault data that includes information such as fault type, occurrence time, and scope of impact.
[0019] The fault correlation network is built based on the relationships between devices, logs, resources, and code. After the generated fault data is input into the network, logical deduction is performed through a preset reasoning strategy to quickly trace the root cause of the fault and output detailed fault cause information, including fault triggering conditions and related factors. Then, based on the fault cause information, combined with system operation rules and past handling experience, targeted and actionable repair suggestions are automatically generated to provide clear guidance for fault handling.
[0020] The visual operation interface displays fault cause information, repair suggestions, system operation status, resource usage, etc. in an intuitive and easy-to-understand way. It supports users to view fault warning information in real time, and also has the function of generating and exporting fault analysis reports, which makes it convenient for users of different levels to quickly grasp the system status and carry out fault handling work in a timely manner.
[0021] In the fintech business sector, fintech platforms (such as online banking and securities trading platforms) have extremely high requirements for system stability and security. Multi-dimensional data collection can cover transaction logs, user account operation records, server load data, network transmission security data, etc. After storage, key features are extracted through processing. Artificial intelligence analysis modules can detect transaction anomalies, system vulnerabilities, payment failures, and other issues in real time. By correlating faults with the network, the causes can be quickly located (such as transaction interruptions triggered by network attacks, or abnormal account data caused by code vulnerabilities). Repair suggestions can be generated (such as strengthening firewall policies, fixing code vulnerabilities, and optimizing transaction verification processes). The visual interface allows maintenance and risk control personnel to monitor in real time, handle faults in a timely manner, ensure the safety and smooth operation of financial transactions, and maintain the safety of user funds and the reputation of the platform.
[0022] In the field of medical technology, medical technology systems (such as telemedicine platforms, electronic medical record systems, and medical device network management systems) are directly related to the quality of medical services and patient safety. Multi-dimensional data can include medical device operation logs, patient diagnosis and treatment data transmission records, system response time, and data storage security information. After collection, storage, and feature extraction, artificial intelligence analysis modules can detect problems such as medical device malfunctions (e.g., abnormal transmission of imaging equipment), system data leakage risks, and delays in diagnosis and treatment data processing. Fault correlation networks can quickly identify the causes (e.g., aging equipment hardware, network transmission protocol compatibility issues, and system permission setting vulnerabilities) and generate repair suggestions (e.g., replacing equipment components, optimizing transmission protocols, and adjusting permission management strategies). The visual interface allows medical staff and technical maintenance personnel to monitor the system status in real time, quickly resolve faults, ensure the smooth operation of telemedicine, secure storage and efficient retrieval of electronic medical records, and improve the continuity and reliability of medical services.
[0023] In one embodiment, S100 includes: S101. Pre-construct a full-link dynamic perception and acquisition mechanism; S102. Through the full-link dynamic perception and acquisition mechanism, adopt a priority dynamic scheduling strategy to collect multi-dimensional data including server operation logs, application program operation logs, network performance data, and user access behavior data; S103. Pre-construct a hierarchical intelligent storage system; S104. The hierarchical intelligent storage system adopts a distributed storage mechanism, combines a data heat dynamic grading mechanism, and transmits the multi-dimensional data to a cloud server and / or a local server for storage through a transmission protocol, and establishes an association index for each dimension of data.
[0024] In this embodiment, the full-link dynamic perception and acquisition mechanism is a core acquisition system pre-built to achieve comprehensive monitoring of the entire process of the operation of a Web e-commerce website. It covers key links such as servers, application programs, network transmissions, and user interactions involved in website operation, and can achieve the联动感知 and synchronous capture of data in each link, ensuring that key data dimensions affecting system stability and performance are not missed. Supported by this mechanism, a priority dynamic scheduling strategy is adopted to carry out multi-dimensional data acquisition work: for the core logs generated during the operation of the server (such as the operation records of mainstream Web servers), the operation logs of e-commerce website application programs, the key performance data during network transmission (including latency, bandwidth occupancy, data packet loss rate, etc.), and the behavior data generated during user access (such as page loading time, click stream trajectory, etc.), the acquisition frequency and acquisition accuracy of different dimension data are dynamically adjusted according to the real-time load status and failure risk level of the system. For example, during high-risk periods with high system load or frequent failures, focus on high-density acquisition of data related to core services first to ensure that key information is not missing; during stable periods with low system load, appropriately increase the acquisition breadth of non-core dimension data to achieve full data coverage, which not only ensures acquisition efficiency but also meets the requirements for data integrity in subsequent analysis.
[0025] The tiered intelligent storage system is a pre-built storage architecture that combines data storage needs with subsequent application scenarios. Its core adopts a distributed storage model, adaptable to storage requirements of different data scales, ensuring data storage stability and scalability. During data transmission and storage, a dynamic data popularity grading mechanism is used, employing adaptable transmission protocols such as HTTP and WebSocket to flexibly transmit collected multi-dimensional data to cloud servers or local servers for categorized storage: high-frequency access real-time monitoring data (such as system real-time operating status data, fault instantaneous trigger data, etc.) is stored in cloud cache or local core storage nodes, ensuring efficient response for subsequent queries and analysis; low-frequency access data such as historical log data and non-critical redundant data are archived using cloud cold storage, saving storage resources while ensuring data traceability. Simultaneously, during the data storage phase, a correlation index is established between data dimensions, logically linking server logs, application logs, network performance data, and user behavior data. This provides efficient data retrieval and correlation analysis support for subsequent AI-based fault detection, log analysis, resource bottleneck location, and fault cause reasoning, improving the efficiency and accuracy of overall fault monitoring and diagnosis.
[0026] In the fintech business, the stable operation of fintech platforms is directly related to user fund security and smooth transactions. Multi-dimensional data collection needs to capture core data in a targeted manner: server operation logs include the operating status records of various dedicated servers such as trading servers and risk control servers, to promptly detect abnormal server response and excessive load; application operation logs cover the execution logs of core applications such as transaction processing programs and account management programs, tracking the program operation status of the entire process of transaction initiation, verification, and completion, and troubleshooting issues such as transaction failure and data synchronization errors; network performance data focuses on monitoring the latency of cross-regional transaction transmission, the bandwidth usage of encrypted financial data transmission, and packet loss during data transmission to avoid transaction timeouts and data loss due to network problems; user access behavior data records the user's login, transfer, query, and transaction order placement operation trajectory, combined with operation time, device information, etc., to help identify risky behaviors such as stolen login and abnormal transactions. The data transmission module can select an encrypted transmission protocol to transmit data to a cloud server that complies with financial regulatory requirements (to achieve centralized management and real-time sharing of data across multiple branches) or a local server (to meet the compliance requirements for local storage of highly sensitive data) based on the security and real-time requirements of financial data. The stored data not only provides a basis for subsequent transaction risk analysis and system fault investigation, but also supports the compliance audit work of financial regulatory authorities, ensuring the safe and stable operation of financial technology businesses.
[0027] In the field of medical technology, medical technology systems directly impact the quality of medical services and patient safety. Multi-dimensional data collection must be tailored to the specific characteristics of medical scenarios: server operation logs include those of diagnostic data servers, electronic medical record storage servers, and remote consultation servers, ensuring stable storage and retrieval of medical data and preventing interruptions in treatment due to server failures; application operation logs cover the operation records of medical record entry programs, treatment plan recommendation programs, and medical data exchange programs, identifying issues such as program crashes, data entry errors, and failures in cross-hospital data sharing; network performance data monitors video transmission latency in remote diagnosis and treatment, transmission bandwidth of medical images (such as CT and MRI images), and packet loss rate of real-time uploads of medical data (such as patient vital signs data), ensuring smooth video playback, clear and complete transmission of medical images, and real-time synchronization of vital signs data during remote consultations; user access behavior data records medical staff's medical record queries, diagnostic and treatment operations, and medical order issuance, as well as patients' remote consultation appointments and health data uploads, assisting in standardizing medical practices and identifying medical data errors caused by improper operation. During data transmission, considering the high sensitivity and privacy requirements of medical data, a highly secure transmission protocol can be selected to transmit data to a local server (meeting the privacy requirements of local storage of medical data within the hospital) or a cloud server that complies with medical data security standards (supporting data sharing during cross-hospital collaborative diagnosis and treatment). The stored data provides reliable support for rapid troubleshooting of medical system faults, traceability of medical quality, and statistical analysis of medical data, helping to improve the stability of medical technology systems and the efficiency and security of medical services.
[0028] This embodiment's multi-dimensional data collection covers core scenarios such as server operation, application operation, network performance, and user access behavior, comprehensively capturing key information across the entire operational chain of a web e-commerce website. It breaks through the limitations of single data dimensions, providing complete data support for subsequent fault analysis. Through multiple transmission protocols such as HTTP and WebSocket, it can flexibly achieve stable transmission of multi-dimensional data to the cloud or local servers, adapting to different deployment needs and ensuring the real-time performance and compatibility of data transmission. Data is stored on a dedicated server, achieving centralized management of multi-source data, avoiding data dispersion and loss, and providing high-quality data input for modules such as AI analysis and knowledge graph reasoning. This helps to accurately identify faults and locate the root causes of problems, laying a solid data foundation for efficient system monitoring and intelligent diagnosis, thereby improving website operational stability and user experience.
[0029] In one embodiment, S200 includes: S201. Obtain multi-dimensional data and associated indexes of each dimension of data stored on a cloud server or local server; S202. Transform the structured log data of multi-dimensional data into standardized structured data by using a preset field mapping strategy. S203. Through regular expression matching and adaptive template learning modules, identify the data field logic of semi-normalized structured data of multi-dimensional data and generate standardized structured data. S204. Extract entities, attributes, and relationships between entities and attributes from non-standardized structured data of multi-dimensional data through the semantic parsing module; S205. Convert non-standardized structured data into standardized structured data based on entities, attributes, and the relationships between entities and attributes; S206. Standardize the standardized structural data and extract key features from it.
[0030] In this embodiment, the focus is on the comprehensive retrieval and correlation of data. When acquiring multi-dimensional data stored on cloud servers or local servers, it not only covers various logs generated during server operation, the running records of e-commerce website applications, performance parameters in network transmission (such as latency, bandwidth, packet loss rate, etc.), and behavioral data formed during user access (such as page loading time, click flow trajectory, etc.), but also simultaneously retrieves pre-established data correlation indexes for each dimension. Through the index, the rapid correlation of data from different sources is achieved, ensuring the integrity and correlation of data in subsequent processing and laying the foundation for cross-dimensional analysis.
[0031] For structured log data in multi-dimensional data, a pre-defined field mapping strategy is used to perform structure transformation. Based on pre-defined standardized field mapping rules that reflect the characteristics of data in Web e-commerce scenarios, well-formatted, clearly defined structured log data is directly converted into standardized structured data. This ensures that such data can quickly adapt to subsequent standardized processing procedures, improving the overall efficiency of data processing.
[0032] For semi-standardized structured data, a collaborative approach using regular expression matching and adaptive template learning modules is employed. First, regular expression matching is used to initially identify fields with fixed format characteristics. Then, the adaptive template learning module learns autonomously based on a large number of semi-standardized data samples, continuously optimizing the field identification logic to accurately capture the underlying field relationships. Ultimately, this transforms the originally inconsistent and unclear semi-standardized structured data into standardized structured data, solving the problem of messy and difficult-to-process semi-standardized data.
[0033] This module performs deep analysis on non-standardized structured data, extracting core information through a semantic parsing module. It accurately extracts entities (such as fault-related equipment names and operational behaviors), their corresponding attributes (such as equipment operating parameters and operation execution times), and the inherent relationships between entities and attributes. This breaks down barriers to the effective identification and utilization of non-standardized data, providing crucial support for subsequent data structuring transformation.
[0034] Based on the entities, attributes, and their relationships extracted from S204, non-standardized structured data is systematically transformed according to a unified standard data structure specification. By establishing the correspondence between entities, attributes, and standard fields, scattered and unformatted non-standardized data is integrated into standardized structured data with a unified format and clear logic. This achieves structural unification of different types of raw data, creating conditions for subsequent centralized processing and analysis.
[0035] After standardizing and transforming various data structures, further refined data processing and key feature extraction are carried out. First, the integrated, standardized data structure is denoised, completed, and formatted to eliminate potential deviations and redundant information that may have occurred during data collection and transmission, ensuring data accuracy and consistency. Then, considering the core needs of fault monitoring and diagnosis in the Web e-commerce scenario, key features are extracted from the data through a combination of statistical analysis and deep learning algorithms. These features include fault precursor features, resource usage correlation features, and correlation features between user behavior and system anomalies, providing accurate and efficient data support for subsequent fault detection, cause reasoning, and resource optimization.
[0036] In the fintech business, the multi-dimensional data generated by fintech platforms includes transaction logs, account operation records, server load data, network transmission security data, and user transaction behavior data. After the data processing module retrieves this data from the server, the preprocessing stage removes fraudulent transaction records, corrects account information entry errors, and filters duplicate transaction logs. Standardization processing unifies and standardizes transaction data in different formats and server resource data with different statistical units, ensuring data accuracy and consistency. Subsequently, key features are extracted through parsing, such as the amount range of abnormal transactions, the characteristics of devices with high-frequency logins, the time patterns corresponding to peak server load, and abnormal access ports in network security data. These key features support subsequent risk prevention and fault diagnosis; for example, abnormal transaction features can quickly identify fraud risks, and server load characteristics can predict resource bottlenecks, ensuring the security and stability of financial transactions and the smooth operation of the platform.
[0037] In the field of medical technology, multi-dimensional data from medical technology systems includes medical equipment operation logs, patient treatment data, medical records, network transmission data, and data on the operational behavior of medical staff. After the data processing module retrieves the data, the preprocessing stage cleans up invalid error messages in the medical equipment logs, corrects incorrectly entered treatment data in the medical records, and removes duplicate patient health records. Standardization processing unifies and standardizes heterogeneous data generated by different medical devices and medical record formats from different hospitals, solving the problem of data interoperability. Key features extracted during the analysis process include changes in operating parameters of medical equipment before malfunctions, key physiological indicators corresponding to patient symptoms, critical thresholds for network latency in remote diagnosis and treatment, and high-frequency scenarios of operational errors by medical staff. These key features can help quickly locate the causes of medical equipment malfunctions, assist in disease diagnosis and analysis, optimize the network environment for remote diagnosis and treatment, standardize the operating procedures of medical staff, improve the efficiency and safety of medical services, and ensure the patient's treatment experience.
[0038] This embodiment achieves centralized access to website operation data across all scenarios by acquiring multi-dimensional data stored in the cloud or on local servers. This breaks down the barriers of scattered data storage and provides a comprehensive data source for subsequent analysis. Structured parsing of the multi-dimensional data transforms irregular logs, performance, and behavioral data into standardized structured data, eliminating analytical obstacles caused by data format differences and improving data usability. Standardization removes data noise and unifies data formats, ensuring data quality. Key feature extraction then accurately identifies core information related to fault detection and resource monitoring, significantly reducing the computational cost of subsequent AI analysis. This series of operations builds a high-quality data foundation, enabling the AI analysis module to quickly and accurately identify faults and locate bottlenecks, significantly improving the real-time performance and accuracy of fault diagnosis. Simultaneously, it provides reliable data support for knowledge graph reasoning, helping the system achieve efficient and intelligent monitoring and diagnosis.
[0039] In one embodiment, S300 includes: S301. The application's runtime logs are analyzed in real time using an anomaly detection model, and fault information is output. S302. The server operation log and application operation log are parsed through the natural language processing module to generate abnormal behavior information; S303. Use the time series analysis module to predict the trend of network performance data and generate resource bottleneck early warning information; S304. Use a behavior sequence analysis model to identify anomalies in user access behavior data and generate user experience-related fault information. S305. Based on the code context and runtime data, locate potential faults in the code through a data correlation analysis model and generate code problem warning information; S306. The fault information, abnormal behavior information, resource bottleneck warning information, code problem warning information and user experience related fault information are fused and processed to output fault data.
[0040] In this embodiment, the artificial intelligence analysis module performs fault detection and generates fault data based on key features. This process involves using various professional technical models to perform targeted analysis on different types of multi-dimensional data, and finally integrating various results to form comprehensive fault data. Anomaly detection models, as a type of machine learning model, monitor and analyze application runtime logs in real time, accurately identifying abnormal situations that deviate from the normal program trajectory and directly outputting fault information including the fault occurrence node and fault manifestation. Natural language processing modules parse the text content in server and application runtime logs, extracting valuable clues from unstructured log information, identifying abnormal behaviors in system operation, and generating detailed abnormal behavior information. Time series analysis modules focus on mining the patterns of network performance data (such as latency, bandwidth, and packet loss rate) over time, predicting potential network congestion and resource shortages through trend forecasting, and generating resource bottleneck warnings. Behavioral sequence analysis models deeply analyze user access behavior data (such as page load time and clickstream data), identifying abnormal interactive behaviors during user access, and generating user experience-related fault information such as slow page response and malfunctioning functions. Simultaneously, by combining the contextual logic of the code and real-time data during program runtime, data correlation analysis models can accurately locate potential problems hidden in the code, such as logical vulnerabilities and syntax errors, generating code problem warnings. Finally, the fault information, abnormal behavior information, resource bottleneck warning information, code problem warning information, and user experience-related fault information obtained above are systematically integrated to form complete fault data covering various fault types, causes, and impact ranges, providing a comprehensive basis for subsequent fault diagnosis and repair.
[0041] Specifically, the necessary components of the anomaly detection model rely on multi-level system collaboration. The core includes a data input interface, a feature processing submodule, an algorithm operation submodule, and a result output interface. It also needs to establish bidirectional connections with the data processing module and the knowledge graph layer. The data input interface receives multi-source data (server logs, application logs, network performance data, user behavior data, etc.) after being cleaned and standardized by the data processing module. The feature processing submodule reuses the key features extracted by the data processing module and supplements them with fault-related specific features. After the algorithm operation submodule completes the anomaly identification, it synchronizes the fault information to other modules of the AI analysis module and the knowledge graph layer through the result output interface, providing a foundation for fault correlation reasoning. The anomaly detection model first collects massive amounts of historical normal operation data and various fault data (including code bugs, resource bottlenecks, network anomalies, etc.), dividing the training and test sets in a 7:3 ratio. Preprocessing the data, such as deduplication and missing value imputation, ensures data integrity. Based on the feature extraction results from the data processing module, it further filters strongly correlated features related to faults (such as server CPU utilization, application response time, network packet loss rate, etc.), standardizing the feature vectors through normalization and feature encoding. Finally, it selects anomaly detection algorithms adapted to multi-source data (such as isolated forest, autoencoder, etc.), initializing key parameters (the number of trees in the isolated forest is set to 100-200, sample anomalies, etc.). The constant threshold was set to 0.8-0.9, the hidden layer dimension of the autoencoder was set to 1 / 2-2 / 3 of the input dimension, and the learning rate was set to 0.001-0.01. The training was iteratively conducted with the goal of "minimizing the anomaly identification error" and the parameters were continuously adjusted to optimize the model performance. The accuracy and recall of the model were verified with the test set. The model's judgment logic was corrected by combining the fault association strategy in the knowledge graph and the feature interference corresponding to the misjudged cases was eliminated to ensure that the model can accurately identify different types of production faults. The trained model was embedded into the AI analysis module to receive new operating data in real time and to update the training dataset regularly (incorporating newly emerging fault cases) to continuously optimize the model parameters to adapt to the dynamic changes in the system's operating status.
[0042] The core components of the behavioral sequence analysis model need to connect multiple levels of the system. Essential modules include a data access module, a sequence preprocessing module, a feature extraction module, a model inference module, and a result output module. It also needs to be closely connected with the data processing module and the AI analysis module. The data access module receives user access behavior data (such as page load time, clickstream data, etc.) after it has been cleaned by the data processing module. After the preprocessing module completes the sequence normalization, the feature extraction module mines the temporal and correlation features in the behavioral sequence. After the model inference module completes the anomaly identification, the result output module synchronizes the user experience-related fault information to the fusion processing unit and knowledge graph layer of the AI analysis module to provide support for fault correlation analysis. The behavior sequence analysis model collects massive amounts of normal and abnormal user access behavior sequence data (such as repeated refreshes, page freezes followed by redirects, and process interruptions). The training and test sets are divided in an 8:2 ratio. The data undergoes sequence completion, noise reduction, and standardization to ensure the integrity of the behavior time sequence. User access behaviors are sorted by timestamps and converted into fixed-length behavior sequence vectors. Variable-length sequences are truncated or padded to standardize the format, with a sequence length threshold of 30-50 behavior nodes. Temporal features (such as behavior intervals and access durations), correlation features (such as page jump logic and operation path relevance), and statistical features (such as the frequency of a certain behavior) are extracted from the behavior sequences. Discrete behaviors are transformed into continuous feature vectors through one-hot encoding and embedding layers. Algorithms adapted to the sequence data (such as L...) are selected. (LSTM, GRU), initialize key parameters: LSTM hidden layer dimension set to 64-128, learning rate set to 0.001-0.005, number of iterations set to 50-100 rounds, batch size set to 32-64. With the goal of "minimizing the error rate of abnormal behavior recognition", iterative training is performed using the training set, dynamically adjusting the number of network layers and neurons. The anomaly recognition accuracy and recall rate of the model are verified using the test set. The judgment threshold is adjusted in combination with the actual scenarios of user experience-related faults (usually the anomaly confidence threshold is set to 0.85-0.9), and interference from misjudged normal behavior sequences is eliminated. The trained model is embedded into the AI analysis module to receive new user behavior data in real time, and new abnormal behavior cases are regularly added to update the training set to continuously optimize model parameters to adapt to the dynamic changes in user access behavior.
[0043] In the fintech business, the stable operation of fintech platforms is directly related to user fund security and smooth transactions. Anomaly detection models can analyze the runtime logs of core applications such as transaction processing programs and account management applications in real time, quickly detect faults such as failed transaction submissions and abnormal account balance calculations, and output fault information. Natural language processing modules can parse text records in financial server and application logs to identify abnormal behavior information such as unauthorized access attempts and transaction data tampering. Time series analysis modules predict trends in network performance data for financial data transmission, providing early warnings of resource bottlenecks such as cross-border transaction delays and batch clearing processing lags caused by insufficient network bandwidth. Behavioral sequence analysis models can monitor user access behavior data such as login, transfer, and payment, identifying potentially risky behaviors such as abnormal logins from different locations and high-frequency small-amount transfers, while also detecting faults that affect user experience, such as slow payment page loading and unresponsive transaction confirmation buttons. Data correlation analysis models combine the context of core financial code and runtime data to locate potential faults in the code that may lead to transaction reconciliation errors and risk control strategy failures, generating code problem warning information. The integrated fault data can help fintech platforms respond quickly to transaction failures, prevent security risks, optimize user transaction experience, and ensure the stable and compliant operation of financial businesses.
[0044] In the field of medical technology, the reliability of medical technology systems directly affects the quality of medical services and patient safety. Anomaly detection models can analyze the operational logs of electronic medical record entry programs and remote consultation applications in real time, promptly identifying faults such as failed medical record data saving and interrupted transmission of treatment instructions, and outputting fault information. Natural language processing modules can parse the text content in medical server and application logs, identifying abnormal behaviors such as medical data leaks and unauthorized access to patient medical records. Time series analysis modules predict trends in network performance data for video transmission networks and medical image data transmission in remote diagnosis and treatment, providing early warnings of excessively high network latency. This leads to resource bottlenecks such as lag in remote consultations and incomplete transmission of CT / MRI images. Behavioral sequence analysis models monitor data from medical staff's actions, including medical record inquiries and prescription issuance, as well as patients' remote appointment bookings and health data uploads. These models identify faults affecting medical service efficiency and user experience, such as abnormal operational processes, failed health data uploads, and slow response times on consultation pages. Data association analysis models combine the context and runtime data of core medical system code (such as medical record data interaction code and medical device control code) to pinpoint potential faults in the code that may cause errors in diagnostic data or malfunctions in medical device control commands. The integrated fault data helps medical technology platforms quickly troubleshoot system faults, ensure medical data security, improve the stability and reliability of remote diagnosis and treatment, medical record management, and other services, providing technical support for the smooth operation of medical services.
[0045] This embodiment achieves comprehensive and accurate fault detection by collaboratively analyzing multi-dimensional data using multiple models. The anomaly detection model captures application runtime faults in real time, the natural language processing module deeply analyzes logs to uncover abnormal behavior, time series analysis accurately predicts network resource bottlenecks, the behavioral sequence analysis model focuses on user experience-related faults, and the data association analysis model precisely locates potential code problems. After fusion processing, the multi-dimensional analysis results overcome the limitations of single data sources and analysis models, covering all fault types across various scenarios, including program execution, network resources, code quality, and user experience. Furthermore, data fusion eliminates information silos, enabling comprehensive and in-depth fault identification. This process significantly improves the real-time performance and comprehensiveness of fault detection, providing complete and accurate fault data support for subsequent fault diagnosis and repair, effectively reducing missed and false fault detections, facilitating rapid fault response, and ensuring stable website operation.
[0046] In one embodiment, S400 includes: S401, Build a knowledge graph of the relationships and dependencies between devices, logs, resources, and code; S402. Construct a fault association network based on the aforementioned association relationships and dependency logic; S403. Based on a preset reasoning strategy, and combined with a fault association network, reason about the fault data and output the fault cause. S404. Analyze the cause of the fault and obtain the fault type, scope of impact and system operating parameters; S405. Generate repair suggestions based on the fault type, scope of impact, and system operating parameters.
[0047] In this embodiment, fault data is input into a fault association network for reasoning and to generate repair suggestions. The core is to use a knowledge graph-based association system to accurately locate the cause of the fault and output targeted solutions. First, through knowledge graph technology, the system comprehensively sorts out the complex relationships and dependencies between devices (such as hardware facilities like servers and network devices), logs (server operation logs, application operation logs, etc.), resources (system resources such as network bandwidth and server computing power), and code (business logic code, underlying support code, etc.). For example, it clarifies the association between a certain type of application log and the operating status of a specific server hardware, or the execution of a certain piece of code depends on specific network resource support, forming a structured association graph. Based on this, a fault association network covering the entire system is built using these relationships and dependencies as the framework, allowing various fault-related elements to form a traceable and deductive logical network. Subsequently, based on a pre-defined inference strategy (such as policy-based inference or case-based inference), the previously generated fault data is imported into the network. By traversing related nodes and analyzing dependency chains, the root cause of the fault is traced. For example, it is determined whether a server hardware failure is causing abnormal application logs, or whether a code vulnerability is causing excessive resource consumption and thus triggering a network bottleneck. Finally, clear and accurate fault cause information is output. Finally, combining the specific type of fault cause, its scope of impact, the optimal parameter configuration for system operation, and past fault handling experience, actionable repair suggestions are generated, clearly indicating the resource configurations that need adjustment, the code vulnerabilities that need fixing, the device operating parameters that need optimization, or the network problems that need to be investigated.
[0048] In the fintech business, the knowledge graph of the fintech platform focuses on constructing the relationships and dependencies between transaction servers, risk control equipment, core business application logs, fund clearing resources, and financial transaction code. For example, it clarifies that the execution of cross-border transaction code depends on international network bandwidth resources, and user account change logs are directly related to the monitoring data of risk control equipment. Based on this fault correlation network, when the imported fault data shows "user transfer transaction failed," the system will traverse the network through reasoning strategies: first, it checks whether the transaction code execution is abnormal, then it checks the corresponding server operating status, whether the network bandwidth meets the cross-border transmission requirements, and whether the risk control equipment has misjudged and blocked transactions, etc., and finally accurately locates the cause of the fault, which may be "temporary congestion of international network bandwidth" or "the abnormal transaction identification strategy in the risk control code is too strict." Corresponding repair suggestions are generated for different causes, such as "temporarily expand international network bandwidth resources" and "optimize the transaction identification threshold parameters in the risk control code," to help technical personnel quickly handle faults, reduce transaction losses, and ensure user fund security and transaction experience.
[0049] In the field of medical technology, the knowledge graph of medical technology systems focuses on constructing the relationships and dependencies between medical devices (such as imaging equipment and vital signs monitors), electronic medical record logs, medical data transmission resources, and diagnostic and treatment procedure code. For example, it clarifies that the operation of CT image transmission programs depends on the hospital's internal high-speed network resources, and that patient vital signs data logs are closely related to the operating status of monitor hardware. Upon receiving fault data indicating that "CT images cannot be uploaded to the electronic medical record system," the constructed fault correlation network analyzes the data through reasoning strategies: checking for anomalies in the CT equipment's operating logs, ensuring sufficient bandwidth in the data transmission network, and identifying execution errors in the image upload-related code. Ultimately, it determines the cause of the fault, which may be "incompatibility between the image transmission code and the newly upgraded electronic medical record system interface" or "a partial interruption in the medical data transmission network." It then generates repair suggestions, such as "updating the image upload code to adapt to the electronic medical record system interface" and "investigating and repairing local network fault nodes," ensuring smooth medical data transmission and guaranteeing the smooth operation of medical services such as remote diagnosis and treatment and multi-departmental collaboration, thus preventing system failures from affecting the patient's treatment process.
[0050] This embodiment constructs a knowledge graph connecting devices, logs, resources, and code, clearly outlining the relationships and dependencies among these elements and establishing a systematic framework for fault analysis. The fault correlation network built upon this framework breaks down the isolated barriers between different data dimensions, allowing diverse fault-related information to form an organic whole. Combined with preset inference strategies, deep reasoning on fault data can penetrate surface phenomena to pinpoint the root cause, avoiding one-sided judgments and significantly improving the accuracy of fault diagnosis. The resulting repair suggestions closely align with the root cause of the fault and the system's relational logic, possessing strong pertinence and operability. The entire process achieves an intelligent closed loop from fault data to cause location and repair solutions, significantly reducing manual troubleshooting costs, accelerating fault handling, and effectively ensuring stable website operation and a good user experience.
[0051] In one embodiment, S500 includes: S501. The system's operating status, fault warning information, resource usage, and code fault location results are displayed in real time through the user interface module's visual interface. S502. When any of the system's operating status, fault warning information, and resource usage is abnormal, based on the knowledge graph of the relationships and dependencies between devices, logs, resources, and code, the associated data dimensions are highlighted through a visual interface, and combined with the fault detection results, resource bottleneck location data, and code fault location results from the AI analysis module, a fault link time-series trajectory diagram is generated. S503. Based on the fault cause information and repair suggestions output by the intelligent diagnostic module, generate a fault analysis report and simultaneously display the fault cause information and repair suggestions on the visualization interface.
[0052] In this embodiment, the user interface module's visualization interface possesses full-dimensional real-time display capabilities, dynamically presenting the core operating status of the system, various fault warning information, real-time usage of hardware and network resources, and specific location results of potential faults in the code. This allows users to intuitively grasp the overall operating status and key details of the system. When system operating status fluctuates, fault warning information is triggered, or resource usage exhibits abnormalities, the interface relies on a pre-built knowledge graph—which deeply associates the inherent connections and dependencies between devices, logs, resources, and code—to automatically identify and highlight all data dimensions directly related to the abnormal fault, clearly outlining the correlation links of the abnormal data and helping users quickly focus on key information. Simultaneously, the interface integrates accurate fault detection results, resource bottleneck location data, and code fault location results output by the AI analysis module, generating a fault link time-series trajectory diagram based on timeline logic. This intuitively reconstructs the complete timeline from the fault's inception to its spread, clearly presenting the business modules involved, data flow paths, and the correlation relationships between each stage, providing visual support for fault root cause location. The system receives detailed fault cause information and targeted repair suggestions from the intelligent diagnostic module in real time. Based on standardized specifications, it automatically generates a fault analysis report containing core content such as fault occurrence time, scope of impact, core causes, troubleshooting process, and repair solutions. The report also displays a core summary of the fault cause and key repair suggestions in a prominent position on the visual interface. Users can view the report details with one click, and the repair suggestions can be directly linked to the corresponding code adjustment entry and resource configuration interface, achieving seamless connection between fault diagnosis and repair operations and greatly improving fault handling efficiency.
[0053] In the fintech business, this visualization function has significant practical value. The fintech platform's visualization interface presents the real-time operational status of the core trading system (such as transaction volume per second, transaction success rate, etc.), fault warning information (such as abnormal account login warnings, transaction payment failure alerts, risk control strategy trigger reminders, etc.), resource usage (such as transaction server load, cross-regional data transmission bandwidth usage, database storage capacity, etc.), and fault location results for the core financial code (such as the location of logical vulnerabilities in the fund clearing code, security risks in the user authentication code, etc.). Users (such as operations and maintenance personnel, risk control specialists) can filter cross-border transaction faults and account security-related faults within specific time periods according to business needs, generating standardized fault analysis reports. The reports clearly state the causes of the faults, such as "cross-border transaction failures at a certain time were due to temporary insufficient international network bandwidth causing data transmission timeouts" or "abnormal account logins were due to inadequate verification strategies for remote login scenarios in the user authentication code," and simultaneously provide remediation suggestions, such as "temporarily expanding international network bandwidth resources" or "optimizing the user authentication code and adding multi-factor authentication steps for remote logins." Through a visual interface, fintech platform personnel can quickly respond to faults, accurately handle problems, ensure the security and smooth operation of financial transactions, and reduce business losses.
[0054] In the field of medical technology operations, this visualization function also plays a crucial role. The visualization interface of the medical technology system displays the real-time operational status of the medical system (such as electronic medical record data read / write speed, and the clarity of remote diagnosis and treatment video transmission), fault warning information (such as alarms for abnormal medical equipment operation, reminders of failed patient medical record data uploads, and warnings of interrupted remote consultation connections), resource usage (such as the capacity of medical image storage servers, hospital network bandwidth usage, and computing power allocation for diagnostic and treatment equipment), and fault location results for medical-related codes (such as compatibility issues in image transmission codes and data verification vulnerabilities in medical record entry programs). Medical staff and technical maintenance personnel can filter specific departmental medical equipment faults and specific types of diagnostic and treatment data transmission faults according to the needs of their work, generating standardized fault analysis reports. The report will clearly identify the cause of the malfunction, such as "CT images cannot be uploaded because the image transmission code is incompatible with the newly upgraded electronic medical record system interface" or "remote consultation lag is due to excessive bandwidth consumption by other services on the hospital's network." It will also provide repair suggestions, such as "update the image transmission code to adapt to the electronic medical record system interface" and "optimize the allocation of hospital network resources to reserve dedicated bandwidth for remote diagnosis and treatment services." With the help of a visual interface, relevant personnel can quickly troubleshoot and resolve malfunctions, ensuring the secure transmission and efficient access to medical data, guaranteeing the smooth operation of medical services such as remote diagnosis and treatment and medical record management, and improving the quality of medical services and the patient's treatment experience.
[0055] This embodiment presents the system's operating status, fault warnings, resource usage, and code fault location results in real time through a visual interface, making various key information intuitive and perceptible. It breaks down the barriers of fragmented data presentation, allowing users at different levels to quickly grasp the core system situation. Standardized fault analysis reports, combined with fault causes and repair suggestions, are displayed simultaneously, providing technical personnel with precise fault handling guidelines and enabling non-professional users to clearly understand the essence of the problem and the direction of solution. This intuitive and integrated information display method significantly reduces the cost of information acquisition and understanding, and shortens the fault response and handling cycle. At the same time, the user-friendly display format meets the needs of different users, helping them efficiently conduct system monitoring, fault handling, and optimization decisions, further ensuring the stable operation of the website and improving overall operational efficiency and user experience.
[0056] In one embodiment, after inputting the fault data into a fault association network for inference, outputting fault cause information, and generating repair suggestions based on the fault cause information, the method further includes: Real-time acquisition of resource bottleneck warning information and system resource configuration data output by the artificial intelligence analysis module; Based on resource bottleneck early warning information and system resource configuration data, a resource optimization plan is generated through a dynamic resource scheduling algorithm according to the characteristics of real-time website traffic and peak business periods. The resource optimization scheme will be synchronized to the visual operation interface.
[0057] In this embodiment, after completing the fault cause reasoning and repair suggestion generation, the system will further carry out dynamic adaptation work related to resource optimization, forming a closed loop of "fault handling + resource optimization". The system will continuously capture resource bottleneck warning information output by the artificial intelligence analysis module (such as warning prompts that may affect system operating efficiency, such as insufficient network bandwidth, server computing power shortage, and storage resource shortage) and current system resource configuration data (including core configuration parameters such as server CPU utilization, memory usage, network bandwidth allocation, and storage capacity usage status). Based on this, combined with the fluctuation of real-time website traffic (such as the number of current online users and page request frequency) and the characteristics of peak business periods (such as the business traffic patterns of specific periods such as e-commerce platform promotion activities, financial platform fund clearing, and medical platform remote diagnosis peaks), the system will perform precise calculation and analysis through dynamic resource scheduling algorithms to generate targeted resource optimization solutions. These solutions may include specific measures such as temporarily expanding the computing power of high-load servers, adjusting the network bandwidth allocation ratio to ensure core business data transmission, optimizing the partition management of storage resources, and scheduling idle resources to high-demand business modules. Finally, the generated resource optimization plan will be synchronized to the visual operation interface and presented in the form of clear charts, text descriptions and other forms, so that users can intuitively understand the direction, specific measures and expected effects of resource optimization, and provide a direct basis for resource adjustment decisions.
[0058] In the fintech business, the fintech platform receives real-time resource bottleneck warnings (such as insufficient server computing power during fund clearing periods and tight cross-border transaction network bandwidth) and system resource configuration data (such as trading server CPU load, database storage utilization, and payment channel bandwidth allocation) from the AI analysis module. Combining real-time access volume (such as peak access times for securities trading during market opening hours and third-party payment transfer requests during holidays) and peak business period characteristics (such as peak repayment periods on monthly paydays and peak payment settlement periods during e-commerce promotions), the dynamic resource scheduling algorithm generates precise resource optimization solutions: for example, before the peak of securities trading opening hours, it automatically expands the computing power and memory resources of the trading system server to ensure real-time push of market data and rapid execution of trading instructions; during busy periods of cross-border payment business, it adjusts the network bandwidth allocation strategy to prioritize the transmission channels for cross-border transfer data; and for situations where database storage pressure is too high, it initiates a data tiered storage solution, storing frequently accessed transaction data in high-speed storage media and migrating infrequently accessed historical data to high-capacity storage devices. After these resource optimization plans are synchronized to the visual interface, operations and maintenance personnel can monitor resource adjustments in real time and flexibly confirm or fine-tune them according to actual business needs. This ensures that financial services can maintain a smooth and stable operation in high-concurrency scenarios and avoid problems such as transaction delays and failures due to insufficient resources.
[0059] In the medical technology business, the medical technology system captures real-time resource bottleneck warnings (such as insufficient network bandwidth warnings during peak remote diagnosis and treatment periods, and critical capacity alerts for medical image storage servers) and system resource configuration data (such as image transmission server load, electronic medical record database response speed, and hospital network bandwidth usage distribution) output by the artificial intelligence analysis module. Combining real-time access volume (such as the number of remote consultations conducted during a specific time period, and the number of medical staff simultaneously querying electronic medical records) and peak business period characteristics (such as peak outpatient visits on weekday mornings, and peak medical equipment data transmission during surgery), the dynamic resource scheduling algorithm generates resource optimization solutions tailored to the needs of the medical scenario: for example, during peak remote diagnosis and treatment periods, dedicated network bandwidth is automatically allocated to consultation services to ensure smooth video playback and real-time synchronization of medical data; to address medical image storage pressure, an image compression and hierarchical access mechanism is activated to compress image file sizes without affecting diagnostic accuracy, and frequently used image data is prioritized for storage in high-speed access areas; when multiple departments simultaneously and frequently access the electronic medical record system, database resource allocation is dynamically adjusted to improve the response speed of medical record queries, modifications, and saves. Once these resource optimization plans are synchronized to the visual interface, the hospital's technical maintenance personnel and medical management personnel can clearly understand the resource allocation status and optimization direction, quickly respond to clinical business needs, ensure the normal operation of medical equipment and the efficient flow of medical data, provide stable resource support for precise diagnosis and treatment and efficient collaboration, and improve the quality and efficiency of medical services.
[0060] This embodiment accurately identifies the core pain points of resource supply and demand imbalance by acquiring real-time resource bottleneck early warning information and system resource configuration data output by the AI analysis module, providing real-time and reliable data support for resource optimization. Combining real-time website traffic and peak business period characteristics, the optimization scheme generated by the dynamic resource scheduling algorithm breaks through the limitations of traditional static configuration, achieving precise matching between resource allocation and business needs—automatic scaling up during peak periods ensures service stability, while reasonable scaling down during off-peak periods avoids resource waste. The optimization scheme is synchronized to a visual operation interface, making resource adjustment strategies intuitive and easy for users to understand and implement. This process not only improves resource utilization efficiency and reduces operation and maintenance costs, but also ensures smooth website operation under different business scenarios by proactively avoiding resource bottlenecks, further enhancing system stability and user experience, and adapting to the dynamic business needs of e-commerce platforms of different sizes.
[0061] In one embodiment, an AI-based fault monitoring device is provided, which corresponds one-to-one with the AI-based fault monitoring method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the artificial intelligence-based fault monitoring device of the present invention. The modules include a data acquisition module 10, a data processing module 20, a fault detection module 30, a fault diagnosis module 40, and a user interface module 50. Detailed descriptions of each functional module are as follows: Data acquisition module 10 is used to collect multi-dimensional data and store the multi-dimensional data to a cloud server and / or a local server; Data processing module 20 is used to perform structured parsing and standardization processing on the multi-dimensional data and extract key features; Fault detection module 30 is used to perform fault detection on the multi-dimensional data based on the key features using an artificial intelligence analysis module, and generate fault data; The fault diagnosis module 40 is used to input the fault data into the fault association network for reasoning, output fault cause information, and generate repair suggestions based on the fault cause information. User interface module 50 is used to display the fault cause information and repair suggestions on a visual operation interface.
[0062] In one embodiment, the data acquisition module 10 includes: The data acquisition module collects multi-dimensional data, including server operation logs, application operation logs, network performance data, and user access behavior data. The data transmission module uses a transmission protocol to transmit the multi-dimensional data to a cloud server and / or a local server for storage.
[0063] In one embodiment, the data processing module 20 includes: The data processing module acquires multi-dimensional data stored on a cloud server or local server. The multi-dimensional data is subjected to structured parsing to generate standardized structured data; The standardized structural data is standardized, and key features are extracted from it.
[0064] In one embodiment, the fault detection module 30 includes: The application's runtime logs are analyzed in real time using an anomaly detection model, and fault information is output. The server operation logs and application operation logs are parsed using the natural language processing module to generate abnormal behavior information; The time series analysis module is used to predict trends in network performance data and generate early warning information about resource bottlenecks. Anomalies in user access behavior data are identified through a behavior sequence analysis model, generating user experience-related fault information. Based on code context and runtime data, potential faults in the code are located through a data correlation analysis model, and code problem warning information is generated. The fault information, abnormal behavior information, resource bottleneck warning information, code problem warning information, and user experience-related fault information are fused and processed to output fault data.
[0065] In one embodiment, the fault diagnosis module 40 includes: Build a knowledge graph of the relationships and dependencies between devices, logs, resources, and code; Construct a fault association network based on the aforementioned relationships and dependency logic; Based on a preset reasoning strategy, the fault data is inferred using a fault association network to output the cause of the fault. The causes of the faults are analyzed to obtain the fault type, scope of impact, and system operating parameters; Repair suggestions are generated based on the fault type, scope of impact, and system operating parameters.
[0066] In one embodiment, the user interface module 50 includes: The system's operating status, fault warning information, resource usage, and code fault location results are displayed in real time through the user interface module's visual interface. A standardized fault analysis report is generated based on the fault cause information and repair suggestions, and the fault cause information and repair suggestions are displayed synchronously on the visualization interface.
[0067] In one embodiment, the resource optimization module further includes: Real-time acquisition of resource bottleneck warning information and system resource configuration data output by the artificial intelligence analysis module; Based on resource bottleneck early warning information and system resource configuration data, a resource optimization plan is generated through a dynamic resource scheduling algorithm according to the characteristics of real-time website traffic and peak business periods. The resource optimization scheme will be synchronized to the visual operation interface.
[0068] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side fault monitoring method based on artificial intelligence.
[0069] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the user-side functions or steps of an artificial intelligence-based fault monitoring method. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Collect multi-dimensional data and store the multi-dimensional data on a cloud server and / or a local server; The multi-dimensional data is subjected to structured parsing and standardization to extract key features; An artificial intelligence analysis module is used to perform fault detection on the multi-dimensional data based on the key features, and fault data is generated. The fault data is input into a fault association network for inference, and fault cause information is output. Repair suggestions are generated based on the fault cause information. The fault cause information and repair suggestions are displayed on the visual operation interface.
[0070] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Collect multi-dimensional data and store the multi-dimensional data on a cloud server and / or a local server; The multi-dimensional data is subjected to structured parsing and standardization to extract key features; An artificial intelligence analysis module is used to perform fault detection on the multi-dimensional data based on the key features, and fault data is generated. The fault data is input into a fault association network for inference, and fault cause information is output. Repair suggestions are generated based on the fault cause information. The fault cause information and repair suggestions are displayed on the visual operation interface.
[0071] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0072] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0073] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0074] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0075] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A fault monitoring method based on artificial intelligence, characterized in that, Includes the following steps: Collect multi-dimensional data and store the multi-dimensional data on a cloud server and / or a local server; The multi-dimensional data is subjected to structured parsing and standardization to extract key features; An artificial intelligence analysis module is used to perform fault detection on the multi-dimensional data based on the key features, and fault data is generated. The fault data is input into a fault association network for inference, and fault cause information is output. Repair suggestions are generated based on the fault cause information. The fault cause information and repair suggestions are displayed on the visual operation interface.
2. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, The process of collecting multi-dimensional data and storing the multi-dimensional data on a cloud server and / or a local server includes: Pre-build a dynamic sensing and acquisition mechanism across the entire chain; Through the aforementioned end-to-end dynamic sensing and acquisition mechanism, a priority dynamic scheduling strategy is adopted to collect multi-dimensional data including server operation logs, application operation logs, network performance data, and user access behavior data. Pre-construct a tiered intelligent storage system; The hierarchical intelligent storage system adopts a distributed storage mechanism and combines a dynamic data popularity grading mechanism to transmit the multi-dimensional data to cloud servers and / or local servers for storage through a transmission protocol, and establishes an associated index for each dimension of data.
3. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, The process of performing structured parsing and standardization on the multi-dimensional data to extract key features includes: Retrieve multi-dimensional data and associated indexes of each dimension of data stored on cloud servers or local servers; By using a preset field mapping strategy, the structured log data of multi-dimensional data is transformed to generate standardized structured data; The module uses regular expression matching and adaptive template learning to identify the logic of data fields in the semi-normalized structured data of multi-dimensional data and generate standardized structured data. The semantic parsing module extracts entities, attributes, and relationships between entities and attributes from non-standardized structured data of multi-dimensional data. Transform non-standardized structured data into standardized structured data based on entities, attributes, and the relationships between entities and attributes; The standardized structural data is standardized, and key features are extracted from it.
4. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, The method employs an artificial intelligence analysis module to perform fault detection on the multi-dimensional data based on the key features, generating fault data, including: The application's runtime logs are analyzed in real time using an anomaly detection model, and fault information is output. The server operation logs and application operation logs are parsed using the natural language processing module to generate abnormal behavior information; The time series analysis module is used to predict trends in network performance data and generate early warning information about resource bottlenecks. Anomalies in user access behavior data are identified through a behavior sequence analysis model, generating user experience-related fault information. Based on code context and runtime data, potential faults in the code are located through a data correlation analysis model, and code problem warning information is generated. The fault information, abnormal behavior information, resource bottleneck warning information, code problem warning information, and user experience-related fault information are fused and processed to output fault data.
5. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, The step of inputting the fault data into a fault association network for inference, outputting fault cause information, and generating repair suggestions based on the fault cause information includes: Build a knowledge graph of the relationships and dependencies between devices, logs, resources, and code; Construct a fault association network based on the aforementioned relationships and dependency logic; Based on a preset reasoning strategy, the fault data is inferred using a fault association network to output the cause of the fault. The causes of the faults are analyzed to obtain the fault type, scope of impact, and system operating parameters; Repair suggestions are generated based on the fault type, scope of impact, and system operating parameters.
6. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, The step of displaying the fault cause information and repair suggestions on a visual operation interface includes: The system's operating status, fault warning information, resource usage, and code fault location results are displayed in real time through the user interface module's visual interface. When any of the system's operating status, fault warning information, and resource usage experiences an abnormal fault, a fault link time-series trajectory diagram is generated based on the knowledge graph of the relationships and dependencies between devices, logs, resources, and code. This is achieved by highlighting the associated data dimensions through a visual interface and combining the fault detection results, resource bottleneck location data, and code fault location results from the AI analysis module. Based on the fault cause information and repair suggestions output by the intelligent diagnostic module, a fault analysis report is generated and the fault cause information and repair suggestions are displayed synchronously on the visualization interface.
7. The fault monitoring method based on artificial intelligence as described in claim 1, characterized in that, After inputting the fault data into a fault association network for inference, outputting fault cause information, and generating repair suggestions based on the fault cause information, the process further includes: Real-time acquisition of resource bottleneck warning information and system resource configuration data output by the artificial intelligence analysis module; Based on resource bottleneck early warning information and system resource configuration data, a resource optimization plan is generated through a dynamic resource scheduling algorithm according to the characteristics of real-time website traffic and peak business periods. The resource optimization scheme will be synchronized to the visual operation interface.
8. A fault monitoring device based on artificial intelligence, characterized in that, The artificial intelligence-based fault monitoring device includes: The data acquisition module is used to collect multi-dimensional data and store the multi-dimensional data to a cloud server and / or a local server; The data processing module is used to perform structured parsing and standardization processing on the multi-dimensional data and extract key features; The fault detection module is used to perform fault detection on the multi-dimensional data based on the key features using the artificial intelligence analysis module, and generate fault data. The fault diagnosis module is used to input the fault data into the fault association network for reasoning, output fault cause information, and generate repair suggestions based on the fault cause information. The user interface module is used to display the fault cause information and repair suggestions on a visual operation interface.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and an AI-based fault monitoring program stored in the memory and executable on the processor. When executed by the processor, the AI-based fault monitoring program implements the steps of the AI-based fault monitoring method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an AI-based fault monitoring program, which, when executed by a processor, implements the steps of the AI-based fault monitoring method as described in any one of claims 1-7.