Comprehensive financial IT operation and maintenance management system and method based on artificial intelligence

By introducing artificial intelligence technology into the financial IT operation and maintenance management system, intelligent monitoring, intelligent fault diagnosis, adaptive resource management and automated safe operation and maintenance have been solved, and the operation and maintenance efficiency and system reliability have been significantly improved.

CN120029858AActive Publication Date: 2025-05-23SHANGHAI GREAT WISDOM

Patent Information

Application Number
CN202510510795.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing financial IT operation and maintenance solutions have shortcomings in terms of intelligence, security and compliance, adaptability, knowledge management, cross-system collaboration and business relevance, and it is difficult to meet the needs of complex financial IT environments.

Method used

The comprehensive financial IT operation and maintenance management system based on artificial intelligence is adopted, including intelligent monitoring and predictive maintenance modules, AI-driven fault diagnosis and self-healing modules, adaptive resource management modules and safe operation and maintenance automation modules. Through technical means such as multi-dimensional timing data analysis, knowledge graphs, reinforcement learning and multi-objective optimization, real-time monitoring, intelligent fault diagnosis, adaptive resource management and automated safe operation and maintenance are achieved.

Benefits of technology

It significantly improves the accuracy of fault prediction and average repair time, enhances the adaptability and security of IT operations and maintenance, improves resource utilization and system reliability, reduces operation and maintenance costs and risks, and improves customer satisfaction and operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029858A_ABST
    Figure CN120029858A_ABST
Patent Text Reader

Abstract

The invention provides a comprehensive financial IT operation and maintenance management system and method based on artificial intelligence, and the system comprises an intelligent monitoring and predictive maintenance module which is used for achieving the real-time monitoring, abnormality recognition and performance prediction of an IT system through collection of multi-dimensional time series data, and when an abnormal condition is monitored or a prediction result does not meet a preset requirement, the intelligent monitoring and predictive maintenance module carries out the operation and maintenance of the IT system; if so, triggering intelligent early warning; the AI-driven fault diagnosis and self-healing module is used for carrying out intelligent fault analysis and repair in combination with IT system logs and historical fault data based on the detected abnormal condition; the self-adaptive resource management module is used for predicting a future load trend and a future performance trend, and performing self-adaptive resource management based on the predicted future load trend and the future performance trend in combination with cost and energy consumption factors; and the security operation and maintenance automation module is used for monitoring security threats of the IT system in real time and automatically executing security strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of IT operation and maintenance automation technology in the field of financial technology, and specifically, to a comprehensive financial IT operation and maintenance management system and method based on artificial intelligence. Background Art

[0002] With the rapid development of financial technology, the IT systems of financial institutions are becoming increasingly complex, and the requirements for IT operation and maintenance are becoming higher and higher. However, the existing financial IT operation and maintenance solutions have the following main limitations: Insufficient intelligence: Existing systems have some capabilities in anomaly detection and basic alarms, but their intelligence levels in predictive maintenance, automatic fault diagnosis, and self-healing are still limited.

[0003] Incomplete consideration of security and compliance: Insufficient consideration of the strict security and compliance requirements unique to the financial industry, and a lack of comprehensive security risk assessment and automated security operation and maintenance capabilities.

[0004] Lack of adaptive capabilities and dynamic optimization: Most existing solutions are relatively static and lack the ability to automatically adjust according to environmental changes and business needs.

[0005] Insufficient operation and maintenance knowledge management and continuous learning: The lack of systematic knowledge management and continuous learning mechanisms makes it difficult to effectively accumulate and utilize long-term operation and maintenance experience.

[0006] Insufficient cross-system collaboration and business relevance: There is a lack of effective support for cross-team and cross-system collaboration, as well as the ability to closely link IT operations and maintenance with business impact.

[0007] According to IDC's forecast, by 2025, global financial institutions' spending on IT operations will reach $200 billion, with a compound annual growth rate of about 8%. However, traditional operations methods have been unable to cope with the increasingly complex financial IT environment. Therefore, developing an intelligent financial IT operations management system that can overcome the above limitations has important practical significance and market value. Summary of the invention

[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a comprehensive financial IT operation and maintenance management system and method based on artificial intelligence.

[0009] According to the present invention, a comprehensive financial IT operation and maintenance management system based on artificial intelligence includes: The intelligent monitoring and predictive maintenance module is used to collect multi-dimensional time series data to achieve real-time monitoring, anomaly identification and performance prediction of IT systems. When an abnormal situation is monitored or the prediction result does not meet the preset requirements, an intelligent warning is triggered; An AI-driven fault diagnosis and self-healing module for performing intelligent fault analysis and repair based on detected anomalies, combined with IT system logs and historical fault data; An adaptive resource management module for predicting future load trends and future performance trends, and performing adaptive resource management based on the predicted future load trends and future performance trends, combined with cost and energy consumption factors; A security operation and maintenance automation module for real-time monitoring of IT system security threats and automatic execution of security policies.

[0010] Preferably, the intelligent monitoring and predictive maintenance module includes: A data collection sub-module: collecting multi-dimensional time series data including performance metrics, logs, and events through a distributed collection module; A data preprocessing sub-module: preprocessing the collected multi-dimensional time series data, including data cleaning, normalization, and feature extraction; An anomaly detection sub-module: identifying anomaly points in the preprocessed multi-dimensional time series data using the Isolation Forest algorithm; constructing a network behavior model using a graph neural network and detecting abnormal activities using the network behavior model; A trend prediction sub-module: analyzing multi-dimensional time series data using an LSTM network to predict data change trends and obtaining trend prediction results; A dynamic threshold adjustment sub-module: dynamically adjusting the alarm threshold based on statistical process control methods, comparing the obtained trend prediction results with the alarm threshold, and considering it abnormal when the trend prediction results exceed the alarm threshold; An early warning generation sub-module: triggering an intelligent early warning based on detected anomalies.

[0011] Preferably, the AI-driven fault diagnosis and self-healing module includes: A knowledge graph engine sub-module: constructing a knowledge graph of IT system topology, component relationships, and fault modes; Based on rule-based regular matching and a BERT entity recognition model, extracting device entities, logical components, and fault modes from IT system configuration items and logs; constructing static dependency relationships through dependency syntactic analysis, and generating dynamic association relationships using Prometheus monitoring data; storing triples in a Neo4j graph database and defining topology level constraints; An inference engine sub-module: performing fault propagation analysis and root cause localization based on the constructed knowledge graph; Integrating the Apache Jena rule engine to implement fault propagation path reasoning; aligning the ITIL fault library with real-time log events based on TF-IDF and GloVe word vectors, and performing knowledge embedding using the TransE algorithm; NLP engine submodule: uses a multimodal deep learning model based on the Transformer architecture to analyze unstructured logs and unstructured text information related to IT systems or business processes to extract key diagnostic information; obtains comprehensive analysis results based on the extracted key diagnostic information, fault propagation analysis, and root cause location; Define the learnable parameter matrix , which is used to learn the adaptation weights in the space after modal concatenation and align the features of log text, performance indicators and topology maps:

[0012] Among them, softmax is used to normalize the output and obtain the attention weights of each modality , indicating the importance of each mode in the final decision at the current moment; Represents the feature vector of the i-th log text, represents the performance indicator feature vector related to the log, A node graph vector representing the event in the system topology; Reinforcement learning agent submodule: selects the optimal repair strategy based on historical experience based on comprehensive analysis results; Among them, a dynamic attention mechanism is introduced into the multimodal deep learning model based on the Transformer architecture, and the importance weights of different data sources are adaptively adjusted according to actual conditions based on the dynamic attention mechanism.

[0013] Preferably, the adaptive resource management module includes: Load forecasting submodule: Use the Prophet model to predict future load and performance trends based on historical load data; Multi-objective optimization submodule: Based on the predicted future load and performance trends, the NSGA-II algorithm is used to balance multiple objectives including resource performance, cost, and energy consumption; Resource scheduling submodule: implements dynamic scaling and migration at the container level through Kubernetes.

[0014] Preferably, the security operation and maintenance automation module includes: DevSecOps integration submodule: embeds automated security scanning and compliance checks into the CI / CD process to obtain inspection results; Network behavior analysis submodule: The inspection results use graph neural networks to build network behavior models and detect abnormal activities; Threat intelligence analysis submodule: abnormal activities are analyzed based on the MITRE ATT&CK framework and associated with security events; Automatic response submodule: Automatically triggers security defense and repair according to predefined policies based on associated security events.

[0015] Preferably, the system further comprises: an intelligent customer service module; The intelligent customer service module is used to provide intelligent IT support services, including: intelligent question and answer, automatic classification and routing of work orders, and intelligent knowledge recommendation.

[0016] Preferably, the intelligent customer service module includes: Intent recognition submodule: Use the BERT model to recognize the intent and entities of user queries; Intelligent question-answering submodule: Based on the identified intent and entities of user queries, the knowledge graph is used for semantic understanding and answer generation; Ticket Routing Submodule: Use deep reinforcement learning to optimize ticket assignment strategies during customer service. Knowledge recommendation submodule: During the customer service process, relevant knowledge and solution strategies are intelligently recommended based on the context.

[0017] Preferably, the system further comprises: an open architecture module; The open architecture module is used to provide standardized APIs and integration interfaces to achieve flexible integration, expansion and interoperability of IT systems.

[0018] Preferably, the open architecture module includes: a RESTful API and a GraphQL interface; The RESTful API is used to provide a standardized HTTP interface and support CRUD operations of resources; The GraphQL interface is used to support data query and aggregation; The RESTful API and the GraphQL interface set up message queues, WebSockets, and security authentication; The message queue is used to implement asynchronous communication using message middleware; The WebSocket is used to support real-time data push and two-way communication; The security authentication is used to implement an API authentication and authorization mechanism based on OAuth 2.0 and JWT.

[0019] According to the present invention, a comprehensive financial IT operation and maintenance management method based on artificial intelligence includes: Step S1: The intelligent monitoring and predictive maintenance module collects multi-dimensional time series data to achieve real-time monitoring, anomaly identification and performance prediction of the IT system. When an abnormal situation is monitored or the prediction result does not meet the preset requirements, an intelligent warning is triggered; Step S2: The AI-driven fault diagnosis and self-healing module performs intelligent fault analysis and repair based on the detected abnormal conditions, combined with IT system logs and historical fault data; Step S3: using the adaptive resource management module to predict future load trends and future performance trends, and based on the predicted future load trends and future performance trends, performing adaptive resource management in combination with cost and energy consumption factors; Step S4: Use the security operation and maintenance automation module to monitor IT system security threats in real time and automatically execute security policies.

[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention has significantly improved the accuracy of fault prediction and shortened the mean repair time from passive response to active prediction. Compared with traditional methods, the accuracy of fault prediction has been improved and the mean repair time has been shortened. 2. The present invention can better adapt to the complex and ever-changing IT operation and maintenance environment and provide more accurate, real-time and explainable diagnostic results; 3. The present invention realizes continuous optimization of system configuration, resource allocation and security strategy through reinforcement learning and multi-objective optimization algorithm, adapting to dynamically changing business needs and technical environment; compared with static rules, resource utilization is improved and energy consumption is reduced; 4. The present invention combines knowledge graphs with machine learning technology to build an intelligent fault diagnosis and problem-solving system, accelerating knowledge accumulation and experience inheritance; compared with traditional knowledge bases, the problem-solving efficiency is improved by 50%.

[0021] 5. This invention deeply embeds security protection and compliance management into the DevOps process to achieve "security left shift"; through advanced algorithms such as graph neural networks, it improves the ability to detect advanced threats in complex network environments; compared with traditional security operations, the threat detection accuracy is improved and the response time is shortened; 6. The present invention is based on microservices and API design, supports flexible integration with existing IT systems and tools, and realizes end-to-end process automation; through open interfaces, it supports financial institutions to customize and expand according to their own needs; 7. The present invention significantly reduces manual intervention through AI-driven automation, and improves the efficiency of operation and maintenance personnel by 200%; through predictive maintenance and intelligent fault diagnosis, the system mean time between failures (MTBF) is increased by 50%; through adaptive resource management, the utilization rate of IT infrastructure is increased by 25%, and the total cost of ownership (TCO) is reduced; through intelligent security analysis and automatic response, the average processing time of security incidents is shortened by 70%; through intelligent customer service and knowledge management, customer satisfaction is improved by 30%; through open architecture, it supports the rapid integration of new technologies and services, shortening the innovation cycle by 40%; 8. The present invention can effectively cope with the growing complexity and uncertainty of financial IT systems, significantly improve operation and maintenance efficiency and system reliability, and provide strong support for the digital transformation of financial institutions; 9. Through machine learning and deep learning technologies, the present invention can continuously learn from operation and maintenance data, optimize decision-making strategies, and adapt to the ever-changing IT environment and business needs; at the same time, its open architecture design enables the platform to have good scalability and integration capabilities, and can be seamlessly connected with existing IT systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 It is the overall system architecture diagram of the present invention; Figure 2 This is a schematic diagram of the intelligent monitoring and predictive maintenance module; Figure 3 This is a schematic diagram of the AI-driven fault diagnosis and self-healing module; Figure 4 It is a schematic diagram of the adaptive resource management module; Figure 5 This is a schematic diagram of the security operation and maintenance automation module; Figure 6 This is a schematic diagram of the smart customer service module. DETAILED DESCRIPTION

[0023] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0024] Example 1 According to the present invention, a comprehensive financial IT operation and maintenance management system based on artificial intelligence is provided, such as Figure 1 As shown, including: Intelligent monitoring and predictive maintenance module: By collecting multi-dimensional time series data, it can realize real-time monitoring, abnormality identification and performance prediction of IT systems, so as to detect potential problems in advance and perform preventive maintenance; Specific functions include but are not limited to: real-time collection of system performance indicators, logs, and event data; use of machine learning algorithms (such as isolation forests) for anomaly detection; use of deep learning models (such as LSTM networks) to predict system performance trends; AI-driven fault diagnosis and self-healing module: Based on detected anomalies, combined with system logs and historical fault data, intelligent fault analysis and repair are performed to reduce system downtime and manual intervention; Specific functions include but are not limited to: using knowledge graphs and other technologies to analyze fault propagation, applying machine learning algorithms to locate root causes, automatically generating repair strategies based on analysis results, and triggering and executing automatic repair operations when appropriate; Adaptive resource management module: responsible for predicting the future load and performance trends of the system, and performing intelligent resource management based on cost and energy consumption factors to improve resource utilization and application performance; Specific functions include but are not limited to: using time series prediction models (such as Prophet) to predict future load and performance trends, using multi-objective optimization algorithms (such as NSGA-II) to balance factors such as performance, cost and energy consumption, and realizing dynamic allocation and optimal scheduling of resources.

[0025] Security operation and maintenance automation module: responsible for real-time monitoring of system security threats and automatically executing corresponding security policies to enhance system security and compliance; Specific functions include but are not limited to: using machine learning algorithms to detect abnormal behavior, automatically implementing security defense measures, and integrating automated security checks into DevOps processes.

[0026] Intelligent customer service module: provides intelligent IT support services to improve customer service efficiency and satisfaction; Specific functions include but are not limited to: intelligent question-answering system based on natural language processing technology, automatic classification and routing of work orders using machine learning algorithms, and intelligent knowledge recommendation based on user context and historical data.

[0027] Open architecture module: By providing standardized APIs and integration interfaces, the system can be flexibly expanded and seamlessly connected with other systems to support business innovation and technological evolution.

[0028] Specific functions include but are not limited to: providing RESTful API and GraphQL interface, supporting microservice architecture, and integrating with third-party systems and tools.

[0029] The present invention is based on the above modules to jointly form a comprehensive, intelligent and adaptive financial IT operation and maintenance management system, which can effectively cope with the challenges of complex financial IT environment.

[0030] Specifically, Figure 2 As shown, the intelligent monitoring and predictive maintenance module includes: Data collection submodule: collects multi-dimensional time series data such as system performance indicators, logs, and events through distributed collection agents; Data preprocessing submodule: perform data cleaning, normalization and feature extraction.

[0031] Anomaly detection submodule: Use the isolation forest algorithm to identify anomalies in multidimensional data; use graph neural networks to build network behavior models, and use network behavior models to detect abnormal activities; Trend prediction submodule: LSTM network is used to analyze time series data to predict performance trends. This embodiment predicts future trends by analyzing historical data, which can detect possible performance bottlenecks or system failures early and reduce unexpected downtime. At the same time, computing resources can be allocated more reasonably to avoid resource waste or shortage, which helps maintain stable system performance and improve user experience. This embodiment analyzes these multi-dimensional time series data through the LSTM network to capture complex nonlinear patterns and long-term dependencies, thereby providing more accurate performance trend predictions and supporting smarter IT operation and maintenance decisions.

[0032] Dynamic threshold adjustment submodule: dynamically adjusts the alarm threshold based on statistical process control methods.

[0033] In this embodiment, the dynamically adjusting the alarm threshold based on the statistical process control method includes: Collect historical data and calculate the mean μ and standard deviation σ; Set initial control limits: Upper control limit (UCL) = μ + 3σ; Lower control limit (LCL) = μ-3σ; Real-time monitoring of new data: If the data points fall within the control limits, the process is considered to be in control; If multiple points in a row (usually 7-8) show a monotonic trend, or if one point is outside the control limit, the process is considered out of control; When the process is detected to be out of control, μ and σ are recalculated and the control limits are updated; Use the new control limits as dynamic alarm thresholds.

[0034] Warning generation submodule: Generate intelligent warnings based on the results of anomaly detection and trend prediction.

[0035] In this embodiment, the integrated abnormality detection and trend prediction results generate intelligent warnings, including: If the current point is detected as abnormal and the forecast trend shows that the problem may persist or worsen, a high priority warning is generated; If the current point is normal, but the forecast trend shows that there may be problems in the future, a low priority warning is generated; If the current point is normal and the forecast trend is good, no warning is generated; This embodiment provides early warning based on both anomaly detection and trend prediction. Anomaly detection focuses on the current state, while trend prediction helps assess future risks. The combination of the two can provide more comprehensive early warning information.

[0036] Specifically, Figure 3 As shown, the AI-driven fault diagnosis and self-healing module includes: Knowledge graph engine submodule: Builds knowledge graphs of IT system topology, component relationships, and failure modes.

[0037] First, model the topology and component relationships: Use rule-based regular matching and BERT entity recognition models to extract device entities (servers, switches), logical components (microservices, database instances), and failure modes (CPU overload, network packet loss) from IT system configuration items (CMDB) and logs; Static dependencies (such as "server A hosts service B") are constructed through dependency syntax analysis, while dynamic associations (such as "service call latency is positively correlated with database query volume") are generated using Prometheus monitoring data.

[0038] The Neo4j graph database is used to store triples and define topology-level constraints (such as “a switch must be connected to at least two servers”).

[0039] Integrate Apache Jena rule engine to implement fault propagation path reasoning (example rule: `If service S depends on database D and D response time > 1s, then the failure probability of S increases by 30%`).

[0040] Reasoning engine submodule: performs fault propagation analysis and root cause location based on knowledge graph; The ITIL fault database and real-time log events are aligned based on TF-IDF and GloVe word vectors, and the TransE algorithm is used for knowledge embedding.

[0041] NLP engine submodule: uses a multimodal deep learning model based on the Transformer architecture to analyze unstructured logs and text information and extract key diagnostic information; integrates the extracted key diagnostic information with fault propagation analysis and root cause location results; Reinforcement learning agent submodule: Based on the integrated results, the optimal repair solution is selected according to historical experience.

[0042] The multimodal deep learning model based on the Transformer architecture includes: Multimodal fusion learning: A multimodal deep learning model based on the Transformer architecture is designed, which can simultaneously process multiple data types such as system logs, performance indicators, network traffic, etc. A dynamic attention mechanism is introduced to adaptively adjust the importance weights of different data sources according to the current situation, thereby improving the adaptability and accuracy of the model.

[0043] Online incremental learning: We developed a streaming data processing module based on Reservoir Sampling to achieve consistency in data distribution. We introduced elastic window technology to dynamically balance the impact of historical data and new data, and improve the stability and real-time performance of the model. We used Elastic Weight Consolidation (EWC) technology to prevent the model from forgetting important information during the update process.

[0044] Explainable AI: Integrate an explanation framework based on SHAP (SHapley Additive exPlanations) to provide understandable explanations for each prediction result. Develop an interactive visualization module based on D3.js to intuitively display feature importance and decision paths, and enhance the interpretability and credibility of the model.

[0045] Transfer learning and few-sample learning: Implement a meta-learning framework based on MAML (Model-Agnostic Meta-Learning) to improve the model's ability to quickly adapt to new fault types. Introduce the SimCLR (Simple Framework for Contrastive Learning of Visual Representations) contrastive learning method to enhance feature representation capabilities and improve the model's generalization ability in small sample situations.

[0046] Through the above improvements, the AI-driven fault diagnosis and self-healing module of the present invention can better adapt to the complex and changeable IT operation and maintenance environment, provide more accurate, real-time and explainable prediction and diagnosis results, significantly improve the accuracy of fault prediction and shorten the mean repair time.

[0047] Specifically, Figure 4 As shown, the adaptive resource management module includes: Load prediction submodule: Use the Prophet model based on historical load data to predict future load trends and performance trends; in this embodiment, the historical load data includes: CPU usage, memory usage, network traffic, number of requests, transaction volume and other time series data reflecting the system load; these data are collected over a long period of time through the monitoring system.

[0048] Multi-objective optimization submodule: uses the NSGA-II algorithm to balance multiple objectives such as performance, cost and energy consumption.

[0049] Resource scheduling submodule: realizes dynamic scaling and migration at the container level through Kubernetes; In this embodiment, the Horizontal Pod Autoscaler (HPA) is used to monitor the usage of container resources; the number of Pods is automatically adjusted according to a predefined policy; and the scheduler is used to migrate Pods between nodes to balance the load.

[0050] Feedback optimization submodule: collects actual operation data and continuously optimizes prediction models and scheduling strategies. Scheduling strategies include container deployment strategy, load balancing strategy, resource allocation strategy, failover strategy, etc.

[0051] Specifically, Figure 5 As shown, the security operation and maintenance automation module includes: Network behavior analysis submodule: Use graph neural networks to build network behavior models and detect abnormal activities; Threat intelligence analysis submodule: Analyze and correlate security events based on the MITRE ATT&CK framework.

[0052] In this embodiment, the security events analyzed and associated based on the MITRE ATT&CK framework include: Event collection and standardization: Collect raw security event logs from various security devices and systems; standardize logs in different formats and extract key fields such as timestamp, source IP, destination IP, user, operation, etc.

[0053] Event mapping: Map standardized events to corresponding techniques in the MITRE ATT&CK framework; Context enrichment: Adding additional contextual information to events, such as asset information, threat intelligence, etc., helps to better understand the impact and severity of the event.

[0054] Pattern recognition: Using machine learning or rules engines to identify attack patterns in a sequence of events; for example, identifying a series of events that may constitute a complete "lateral movement" tactic.

[0055] Correlation analysis: Correlate related events based on attributes such as time, IP, and user to identify events that may belong to different stages of the same attack chain.

[0056] Tactical reconstruction: Map the associated event sequence to MITRE ATT&CK tactics to reconstruct the attacker's possible actions.

[0057] Threat Score: Score the entire attack chain based on the number of techniques, tactics, and severity observed.

[0058] Visualization: Use the MITRE ATT&CK matrix to visualize the detected techniques and tactics; generate an attack chain timeline to show the evolution of the attack.

[0059] Automated response: Based on the identified ATT&CK techniques, trigger the corresponding automated response actions.

[0060] Continuous learning: Feedback analysis results into detection rules to continuously optimize and update detection capabilities.

[0061] Automatic response submodule: automatically performs security defense and repair operations according to predefined strategies.

[0062] In this embodiment, when a high-risk vulnerability is detected, patches are automatically deployed or temporary mitigation measures are implemented; when malware is found, the infected system is isolated and the malware removal program is started; when abnormal login behavior is detected, the account is temporarily locked and multi-factor authentication is required; when a data leakage attempt is identified, suspicious connections are blocked and sensitive data is encrypted; when unauthorized configuration changes are found, the security configuration is rolled back, and the changes are recorded and reported; when a DDoS attack is detected, the traffic cleaning service is activated and the firewall rules are adjusted; when lateral movement behavior is identified, the affected network segment is isolated and network segmentation is enhanced; when abnormal access to sensitive data is found, user access rights are revoked and data access audits are started; when a brute force attack is detected, the attack source IP is temporarily blocked and the password policy is enhanced; when a SQL injection attempt is identified, the Web application firewall rules are updated and the vulnerability is patched; when abnormal process behavior is found, the suspicious process is terminated and the process behavior is analyzed; when unauthorized API access is detected, the API key is revoked and the API authentication mechanism is enhanced. These automatic response strategies are usually graded according to the severity and potential impact of the threat to ensure the appropriateness of the response. At the same time, the system will also retain the interface for manual intervention, allowing the security team to manually adjust or cancel the automatic response operation when necessary. By implementing such an automatic response mechanism, the system can take action as soon as a threat is detected, greatly reducing the response time and improving the efficiency and effectiveness of security defense.

[0063] DevSecOps integration submodule: Embed automated security scanning and compliance checks in CI / CD processes.

[0064] Specifically, Figure 6 As shown, the intelligent customer service module includes: Intent identification submodule: Use the BERT model to identify the intent and entities of user queries.

[0065] Intelligent question-answering submodule: semantic understanding and answer generation based on knowledge graph.

[0066] In this embodiment, the semantic understanding and answer generation based on the knowledge graph includes: mapping user queries to entities and relationships in the knowledge graph, reasoning in the knowledge graph to find relevant information nodes, scoring and ranking candidate answers, and converting the information in the graph into natural language answers.

[0067] Work order routing submodule: Use deep reinforcement learning to optimize the work order assignment strategy.

[0068] In this embodiment, the use of deep reinforcement learning to optimize the work order allocation strategy includes: encoding information such as the current work order queue and customer service status into a state vector; defining possible allocation actions (such as allocating the work order to a specific customer service); designing a reward function based on factors such as processing time and customer satisfaction; training a deep neural network to select the optimal allocation action based on the current state; storing and reusing historical allocation experience to improve learning efficiency; and continuously optimizing the allocation strategy by interacting with the environment; Knowledge recommendation submodule: intelligently recommends relevant knowledge and solutions based on the context.

[0069] Specifically, the open architecture module includes: RESTful API: Provides a standardized HTTP interface and supports CRUD operations on resources.

[0070] GraphQL interface: supports flexible data query and aggregation.

[0071] Message Queue: Use message middleware such as Kafka to implement asynchronous communication.

[0072] WebSocket: Supports real-time data push and two-way communication.

[0073] Security Authentication: Implement API authentication and authorization mechanisms based on OAuth 2.0 and JWT.

[0074] The present invention also provides an artificial intelligence-based comprehensive financial IT operation and maintenance management system, which can be implemented by executing the process steps of the artificial intelligence-based comprehensive financial IT operation and maintenance management method, that is, those skilled in the art can understand the artificial intelligence-based comprehensive financial IT operation and maintenance management method as a preferred implementation of the artificial intelligence-based comprehensive financial IT operation and maintenance management system.

[0075] The present invention realizes the intelligence, automation and coordination of financial IT operation and maintenance, can effectively cope with the growing complexity and uncertainty of financial IT systems, significantly improves operation and maintenance efficiency and system reliability, and provides strong support for the digital transformation of financial institutions.

[0076] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0077] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A comprehensive financial IT operation and maintenance management system based on artificial intelligence, characterized in that: include: The intelligent monitoring and predictive maintenance module is used to collect multi-dimensional time series data to achieve real-time monitoring, anomaly identification and performance prediction of IT systems. When an abnormal situation is monitored or the prediction result does not meet the preset requirements, an intelligent warning is triggered; AI-driven fault diagnosis and self-healing module, which is used to perform intelligent fault analysis and repair based on detected anomalies combined with IT system logs and historical fault data; An adaptive resource management module is used to predict future load trends and future performance trends, and to perform adaptive resource management based on the predicted future load trends and future performance trends combined with cost and energy consumption factors; The security operation and maintenance automation module is used to monitor IT system security threats in real time and automatically execute security policies.

2. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The intelligent monitoring and predictive maintenance module includes: Data collection submodule: collects multi-dimensional time series data including performance indicators, logs, and events through distributed collection modules; Data preprocessing submodule: preprocesses the collected multi-dimensional time series data, including data cleaning, normalization and feature extraction; Anomaly detection submodule: Use the isolation forest algorithm to identify anomalies in preprocessed multi-dimensional time series data; use graph neural networks to build network behavior models, and use network behavior models to detect abnormal activities; Trend prediction submodule: Use LSTM network to analyze multi-dimensional time series data to predict data change trends and obtain trend prediction results; Dynamic threshold adjustment submodule: dynamically adjusts the alarm threshold based on the statistical process control method, and compares the obtained trend prediction result with the alarm threshold. When the trend prediction result exceeds the alarm threshold, it is considered to be abnormal; Warning generation submodule: triggers intelligent warnings based on detected abnormal situations.

3. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The AI-driven fault diagnosis and self-healing module includes: Knowledge graph engine submodule: builds knowledge graphs of IT system topology, component relationships, and failure modes; Based on rule-based regular matching and BERT entity recognition models, we extract equipment entities, logical components, and failure modes from IT system configuration items and logs. We build static dependencies through dependency syntax analysis and generate dynamic associations using Prometheus monitoring data. We use Neo4j graph database to store triples and define topology constraints. Reasoning engine submodule: performs fault propagation analysis and root cause location based on the constructed knowledge graph; Integrate Apache Jena rule engine to implement fault propagation path reasoning; align ITIL fault database with real-time log events based on TF-IDF and GloVe word vectors, and use TransE algorithm for knowledge embedding; NLP engine submodule: uses a multimodal deep learning model based on the Transformer architecture to analyze unstructured logs and unstructured text information related to IT systems or business processes to extract key diagnostic information; obtains comprehensive analysis results based on the extracted key diagnostic information, fault propagation analysis, and root cause location; Define the learnable parameter matrix , which is used to learn the adaptation weights in the space after modal concatenation and align the features of log text, performance indicators and topology maps: Among them, softmax is used to normalize the output and obtain the attention weights of each modality , indicating the importance of each mode in the final decision at the current moment; Represents the feature vector of the i-th log text, represents the performance indicator feature vector related to the log, A node graph vector representing the event in the system topology; Reinforcement learning agent submodule: selects the optimal repair strategy based on historical experience based on comprehensive analysis results; Among them, a dynamic attention mechanism is introduced into the multimodal deep learning model based on the Transformer architecture, and the importance weights of different data sources are adaptively adjusted according to actual conditions based on the dynamic attention mechanism.

4. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The adaptive resource management module includes: Load forecasting submodule: Use the Prophet model to predict future load and performance trends based on historical load data; Multi-objective optimization submodule: Based on the predicted future load and performance trends, the NSGA-II algorithm is used to balance multiple objectives including resource performance, cost, and energy consumption; Resource scheduling submodule: implements dynamic scaling and migration at the container level through Kubernetes.

5. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The security operation and maintenance automation module includes: DevSecOps integration submodule: embeds automated security scanning and compliance checks into the CI / CD process to obtain inspection results; Network behavior analysis submodule: The inspection results use graph neural networks to build network behavior models and detect abnormal activities; Threat intelligence analysis submodule: abnormal activities are analyzed based on the MITRE ATT&CK framework and associated with security events; Automatic response submodule: Automatically triggers security defense and repair according to predefined policies based on associated security events.

6. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The system also includes: an intelligent customer service module; The intelligent customer service module is used to provide intelligent IT support services, including: intelligent question and answer, automatic classification and routing of work orders, and intelligent knowledge recommendation.

7. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 6 is characterized in that: The intelligent customer service module includes: Intent recognition submodule: Use the BERT model to recognize the intent and entities of user queries; Intelligent question-answering submodule: Based on the identified intent and entities of user queries, the knowledge graph is used for semantic understanding and answer generation; Ticket Routing Submodule: Use deep reinforcement learning to optimize ticket assignment strategies during customer service. Knowledge recommendation submodule: During the customer service process, relevant knowledge and solution strategies are intelligently recommended based on the context.

8. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The system further comprises: an open architecture module; The open architecture module is used to provide standardized APIs and integration interfaces to achieve flexible integration, expansion and interoperability of IT systems.

9. The comprehensive financial IT operation and maintenance management system based on artificial intelligence according to claim 8 is characterized in that: The open architecture module includes: RESTful API and GraphQL interface; The RESTful API is used to provide a standardized HTTP interface and support CRUD operations of resources; The GraphQL interface is used to support data query and aggregation; The RESTful API and the GraphQL interface set up message queues, WebSockets, and security authentication; The message queue is used to implement asynchronous communication using message middleware; The WebSocket is used to support real-time data push and two-way communication; The security authentication is used to implement an API authentication and authorization mechanism based on OAuth 2.0 and JWT.

10. A comprehensive financial IT operation and maintenance management method based on artificial intelligence, characterized in that: include: Step S1: The intelligent monitoring and predictive maintenance module collects multi-dimensional time series data to achieve real-time monitoring, anomaly identification and performance prediction of the IT system. When an abnormal situation is monitored or the prediction result does not meet the preset requirements, an intelligent warning is triggered; Step S2: The AI-driven fault diagnosis and self-healing module performs intelligent fault analysis and repair based on the detected abnormal conditions, combined with IT system logs and historical fault data; Step S3: using the adaptive resource management module to predict future load trends and future performance trends, and based on the predicted future load trends and future performance trends, adaptive resource management is performed in combination with cost and energy consumption factors; Step S4: Use the security operation and maintenance automation module to monitor IT system security threats in real time and automatically execute security policies.

Citation Information

Patent Citations

  • Intelligent correction method for sea wave forecast

    CN118504779A

  • Operation and maintenance decision driving method and system based on multi-modal data knowledge graph

    CN119722037A

  • Robot online teaching method and system based on artificial intelligence

    CN119741172A

  • Intelligent operation and maintenance management and alarm system based on large model agent

    CN119847802A

  • Preventative diagnosis prediction and solution determination of future event using internet of things and artificial intelligence

    US20200019893A1

Cited By

  • AI-based meter reading data management and resource scheduling optimization method and system

    CN120317639A

  • Self-healing operation and maintenance method, device and equipment of big data component and storage medium

    CN120560702A

  • Operation maintenance management method of integrated management system

    CN120610842A

  • Intelligent early warning system and method based on log collection

    CN120913371A

  • Intelligent operation and maintenance method and system based on CMDB

    CN120934999A