Big data platform operation and maintenance management system based on artificial intelligence

Through the operation and maintenance management system of the big data platform based on artificial intelligence, the intelligent operation and maintenance management of the big data platform is realized, solving the problems of untimely fault response, unreasonable resource allocation and insufficient security protection in the traditional operation and maintenance methods, and improving operation and maintenance efficiency, stability and security.

CN120448168APending Publication Date: 2025-08-08NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510519630.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The operation and maintenance management of traditional big data platforms relies on manual experience, making it difficult to detect faults in real time, dynamically optimize resource allocation, and insufficient security protection, resulting in low operation and maintenance efficiency, poor stability and weak security.

Method used

The operation and maintenance management system of big data platform based on artificial intelligence is adopted, including data collection and preprocessing, intelligent fault detection and early warning, intelligent resource allocation, security protection and intelligent decision support modules, and uses machine learning, deep learning and reinforcement learning algorithms for automated operation and maintenance and security monitoring to provide real-time decision support.

Benefits of technology

It improves operation and maintenance efficiency, enhances system stability and security, reduces the probability of failure, provides scientific decision-making basis, and improves resource utilization efficiency and security protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448168A_ABST
    Figure CN120448168A_ABST
Patent Text Reader

Abstract

The invention discloses a big data platform operation and maintenance management system based on artificial intelligence, and relates to the field of big data and artificial intelligence. Through automatic data acquisition, processing and analysis and an intelligent fault detection and early warning mechanism, problems can be found and solved in time, manual intervention is reduced, and the operation and maintenance efficiency is improved. Resource allocation is monitored in real time and dynamically adjusted, stable operation of the big data platform under different load conditions can be ensured, and the probability of fault occurrence is reduced. Security threats are monitored and prevented in real time through the artificial intelligence technology, network attacks and data leakage can be effectively resisted, and the security of the big data platform is guaranteed. Through deep mining and analysis of operation and maintenance data, a scientific decision basis is provided for operation and maintenance personnel, and the operation and maintenance personnel are helped to make a more reasonable operation and maintenance strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of big data and artificial intelligence, and specifically to an artificial intelligence-based big data platform operation and maintenance management system. Background Art

[0002] With the rapid development of big data technology, big data platforms have been widely adopted in various fields. However, the operation and maintenance management of big data platforms faces numerous challenges. Traditional operation and maintenance management methods rely primarily on manual experience and are unable to cope with the challenges posed by the massive data processing, complex system architecture, and high concurrent access of big data platforms. For example, in terms of fault detection, manual inspections fail to detect potential faults in real time, resulting in a delay in responding to faults and disrupting the normal operation of the platform. In terms of resource allocation, manual deployment often fails to dynamically optimize according to the platform's real-time load, resulting in wasted or insufficient resources. Furthermore, the security of big data platforms faces severe challenges, as traditional security strategies are unable to effectively defend against increasingly sophisticated cyberattacks. Therefore, there is an urgent need for an intelligent operation and maintenance management system to improve the efficiency, stability, and security of big data platforms. Summary of the Invention

[0003] In response to the problems of low efficiency, untimely fault response, unreasonable resource allocation and insufficient security protection in the existing technology, the purpose of the present invention is to provide an artificial intelligence-based big data platform operation and maintenance management system to realize intelligent and automated operation and maintenance management of the big data platform.

[0004] In order to achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: an artificial intelligence-based big data platform operation and maintenance management system, including a data acquisition and preprocessing module, an intelligent fault detection and early warning module, an intelligent resource allocation module, a security protection module, and an intelligent decision support module; the data acquisition and preprocessing module is used to collect various types of operating data of the big data platform in real time, and clean, convert and normalize the collected data; the intelligent fault detection and early warning module uses machine learning and deep learning algorithms to model and analyze the preprocessed data, monitors the operating status of the big data platform in real time, and issues early warnings in time when abnormal data or potential faults are detected. The intelligent resource allocation module, based on artificial intelligence algorithms, dynamically adjusts resource allocation according to the real-time load and business needs of the big data platform, and monitors and evaluates resource usage in real time; the security protection module uses artificial intelligence technology to monitor and analyze the network traffic and user behavior of the big data platform in real time, identify potential security threats, and use intelligent firewalls and intrusion detection systems to automatically intercept and prevent security threats; the intelligent decision support module, based on big data analysis and artificial intelligence algorithms, conducts in-depth mining and analysis of various data in the operation and maintenance management process, provides decision support for operation and maintenance personnel, and provides a visual interface to display the analysis results.

[0005] The data acquisition and preprocessing module adopts distributed data acquisition technology, deploys data acquisition agents at each node of the big data platform, and transmits the collected data to the data processing center through the message queue.

[0006] The intelligent fault detection and early warning module uses a convolutional neural network (CNN) to extract and analyze server performance data and build a fault prediction model.

[0007] The intelligent resource allocation module adopts a reinforcement learning algorithm to abstract the resource allocation problem of the big data platform into a Markov decision process and learn the optimal resource allocation strategy.

[0008] The security protection module uses a recurrent neural network (RNN) to analyze user login behavior and determine whether there is any abnormal login.

[0009] The intelligent decision support module uses an association rule mining algorithm to analyze the association relationship between fault data and resource usage data.

[0010] The present invention has the following beneficial effects:

[0011] 1. Improve operation and maintenance efficiency: Through automated data collection, processing and analysis, as well as intelligent fault detection and early warning mechanisms, problems can be discovered and solved in a timely manner, reducing manual intervention and improving operation and maintenance efficiency.

[0012] 2. Enhance system stability: Real-time monitoring and dynamic adjustment of resource allocation can ensure the stable operation of the big data platform under different load conditions and reduce the probability of failure.

[0013] 3. Improve security protection capabilities: Use artificial intelligence technology to monitor and prevent security threats in real time, which can effectively resist network attacks and data leaks and ensure the security of big data platforms.

[0014] 4. Provide decision support: Through in-depth mining and analysis of operation and maintenance data, provide operation and maintenance personnel with scientific decision-making basis and help them formulate more reasonable operation and maintenance strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments;

[0016] Figure 1 This is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0017] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.

[0018] Reference Figure 1 ,This specific implementation adopts the following technical solutions: an artificial intelligence-based big data platform operation and maintenance management system, including a data acquisition and preprocessing module 1, an intelligent fault detection and early warning module 2, an intelligent resource allocation module 3, a security protection module 4, and an intelligent decision support module 5;

[0019] Among them, the data collection and preprocessing module is responsible for real-time collection of various operating data of the big data platform, including server performance indicators (such as CPU usage, memory usage, disk I / O, etc.), network traffic data, application logs, etc.

[0020] The collected data is cleaned, converted and normalized to remove noise data and outliers, and data in different formats are converted into a format that is convenient for subsequent analysis.

[0021] Intelligent Fault Detection and Early Warning Module: This module uses machine learning and deep learning algorithms to model and analyze preprocessed data to build a fault prediction model. For example, it uses a convolutional neural network (CNN) to extract and analyze features from server performance data to identify potential failure modes.

[0022] Real-time monitoring of the operating status of the big data platform. When abnormal data or potential failures are detected, timely warning information is issued. Warning information can be notified to operation and maintenance personnel via SMS, email, or internal system messages.

[0023] Intelligent Resource Allocation Module: Based on artificial intelligence algorithms, it dynamically adjusts resource allocation based on the real-time load of the big data platform and business needs. For example, it uses reinforcement learning algorithms to continuously interact with the environment to learn optimal resource allocation strategies and achieve efficient resource utilization.

[0024] Monitor and evaluate resource usage in real time, and automatically adjust and optimize resources when insufficient or wasted resources are found.

[0025] Security protection module: Uses artificial intelligence technology to monitor and analyze the network traffic and user behavior of the big data platform in real time, and identify potential security threats such as network attacks and data leaks.

[0026] Intelligent firewalls and intrusion detection systems are used to automatically intercept and prevent security threats. At the same time, encryption technology is used to protect sensitive data to ensure data security and integrity.

[0027] Intelligent Decision Support Module: Based on big data analysis and artificial intelligence algorithms, it conducts in-depth mining and analysis of various data in the operation and maintenance management process to provide decision support for operation and maintenance personnel. For example, by analyzing historical failure data, it can predict the type and time of possible failures and formulate response strategies in advance.

[0028] Provides a visual interface to present analysis results to operation and maintenance personnel in the form of intuitive charts and reports, facilitating their decision-making and management.

[0029] This specific embodiment utilizes distributed data collection technology for data collection. Data collection agents are deployed at each node of the big data platform to collect server performance data, network traffic data, and application logs in real time. The collected data is transmitted to the data processing center via a message queue. During the data preprocessing phase, a data cleaning algorithm is used to remove noise and outliers. A data conversion algorithm is used to convert data in different formats into a standard format. Finally, the data is normalized to facilitate subsequent analysis and modeling. Historical data is used to train machine learning and deep learning models. For example, a convolutional neural network (CNN) is trained on server performance data to learn data characteristics under normal operating conditions. During real-time monitoring, the collected real-time data is input into the trained model. When the model identifies data anomalies, an early warning mechanism is triggered, promptly notifying operations and maintenance personnel. A reinforcement learning algorithm is used to abstract the resource allocation problem of the big data platform into a Markov decision process. Through continuous interaction with the environment, the optimal resource allocation strategy is learned. During real-time operation, resource allocation is dynamically adjusted based on the platform's real-time load and business needs, achieving efficient resource utilization.

[0030] This specific implementation method uses deep learning algorithms to model and analyze network traffic and user behavior to identify potential security threats. For example, a recurrent neural network (RNN) is used to analyze user login behavior to determine whether there are any abnormal logins. At the same time, intelligent firewalls and intrusion detection systems are deployed to automatically intercept and prevent security threats. Big data analysis technology is used to deeply mine and analyze various data in the operation and maintenance management process, such as using association rule mining algorithms to analyze the correlation between fault data and resource usage data. The analysis results are presented to operation and maintenance personnel in the form of visual charts and reports to provide them with decision support.

[0031] Example 1: Dynamic resource scheduling and fault prevention during e-commerce promotions

[0032] An e-commerce platform faced a surge in traffic during the "Double 11" shopping festival and needed to ensure high system availability. Traditional static resource allocation could easily lead to server overload and resource waste.

[0033] Therefore, the system operation process of the present invention is as follows:

[0034] 1. Data collection and preprocessing

[0035] The collection agent deployed in the server cluster collects data such as CPU usage (peak value up to 90%) and network latency (sudden increase of 200ms) in real time, and transmits it to the data processing center through the Kafka message queue.

[0036] The preprocessing module uses the box plot algorithm to remove sensor anomalies in disk I / O (such as dirty data that exceeds 100% instantaneously) and converts the log data into JSON format.

[0037] 2. Intelligent fault detection

[0038] The CNN model detects that the memory usage of a batch processing node continuously exceeds the threshold (>95%), and the log reports the error "GC overhead limit exceeded".

[0039] The system triggers a level 3 warning (email + SMS), prompting "Node A may crash due to a memory leak" and automatically generates a snapshot for analysis.

[0040] 3. Dynamic adjustment of resources

[0041] The reinforcement learning model (based on Q-learning) adds 10 idle computing nodes to the cluster based on the real-time load, and delays the execution of non-core businesses (such as data analysis tasks).

[0042] Resource utilization dropped from 80% to 65%, avoiding service degradation.

[0043] 4. Safety protection

[0044] RNN detected that a certain IP attempted to log in 500 times within 1 minute, and the behavioral characteristics matched the brute force cracking pattern. The intelligent firewall immediately blocked the IP and isolated the relevant virtual machines.

[0045] 5. Decision support

[0046] The visualization panel displays a "historical promotion failure heat map." The operation and maintenance team expanded the database connection pool in advance, and the failure rate dropped by 70% year-on-year.

[0047] This embodiment

[0048] Example 2: Abnormal transaction identification and self-healing in financial systems

[0049] A bank's big data platform needs to monitor transaction fraud risks in real time while ensuring the integrity of compliance audit data. The specific system operation process is as follows:

[0050] 1. Data Collection

[0051] The collection module obtains transaction logs from Kafka streams (TPS 100,000+), simultaneously collects Oracle database audit logs, and implements unified stream and batch processing through Flink.

[0052] 2. Fault detection and early warning

[0053] The LSTM model found that the standard deviation of transaction response time suddenly increased by 3 times during a certain period of time. Correlation analysis suggested that it was related to the timeout of the third-party payment interface.

[0054] The system automatically triggers the degradation strategy, switches to the backup channel, and marks the abnormal link in red in the topology diagram.

[0055] 3. Safety protection

[0056] The GAN-based anomaly detection model identified that an internal account was exporting customer data in batches during non-working hours, immediately terminating the session and alerting the security team.

[0057] The data encryption module desensitizes sensitive fields (such as ID card number) in export operations in real time.

[0058] 4. Intelligent Decision-Making

[0059] Association rule mining shows that "when disk write latency is >50ms, the probability of a database deadlock occurring the next day increases by 40%."

[0060] Based on this, the operation and maintenance personnel optimized the storage strategy and pre-allocated table space, reducing the number of deadlocks by 60%.

[0061] 5. Resource optimization

[0062] The reinforcement learning model automatically scales down container instances by 50% during the nighttime off-peak period, saving approximately $15,000 per month in cloud computing costs.

[0063] The above-mentioned embodiments simultaneously process structured metrics (CPU), unstructured logs (StackTrace), and streaming data (transaction records). CNNs are used for spatial features (e.g., visualization of server metric time series), RNNs / LSTMs handle sequence dependencies (e.g., log streams), and reinforcement learning enables dynamic optimization. This creates an automated pipeline from detection (model) → decision-making (strategy library) → execution (API calls), reducing manual intervention.

[0064] These two cases verify the flexibility and practicality of the system in complex scenarios.

[0065] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An artificial intelligence-based big data platform operation and maintenance management system, characterized in that: The system includes a data acquisition and preprocessing module, an intelligent fault detection and early warning module, an intelligent resource allocation module, a security protection module, and an intelligent decision support module. The data acquisition and preprocessing module is used to collect various operating data of the big data platform in real time and clean, convert and normalize the collected data. The intelligent fault detection and early warning module uses machine learning and deep learning algorithms to model and analyze the preprocessed data, monitor the operating status of the big data platform in real time, and issue early warning information in a timely manner when abnormal data or potential faults are detected. The intelligent resource allocation module, based on artificial intelligence algorithms, dynamically adjusts resource allocation according to the real-time load and business needs of the big data platform, and monitors and evaluates resource usage in real time. The security protection module uses artificial intelligence technology to monitor and analyze the network traffic and user behavior of the big data platform in real time, identify potential security threats, and automatically intercept and prevent security threats using intelligent firewalls and intrusion detection systems. The intelligent decision support module, based on big data analysis and artificial intelligence algorithms, conducts in-depth mining and analysis of various data in the operation and maintenance management process, provides decision support for operation and maintenance personnel, and provides a visual interface to display the analysis results.

2. The big data platform operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The data acquisition and preprocessing module adopts distributed data acquisition technology, deploys data acquisition agents at each node of the big data platform, and transmits the collected data to the data processing center through the message queue.

3. The big data platform operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The intelligent fault detection and early warning module uses a convolutional neural network (CNN) to extract and analyze server performance data and build a fault prediction model.

4. The big data platform operation and maintenance management system based on artificial intelligence according to claim 1, characterized in that: The intelligent resource allocation module adopts a reinforcement learning algorithm to abstract the resource allocation problem of the big data platform into a Markov decision process and learn the optimal resource allocation strategy.

5. The big data platform operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The security protection module uses a recurrent neural network (RNN) to analyze user login behavior and determine whether there is any abnormal login.

6. The big data platform operation and maintenance management system based on artificial intelligence according to claim 1 is characterized in that: The intelligent decision support module uses an association rule mining algorithm to analyze the association relationship between fault data and resource usage data.