An operation and maintenance management method, device, electronic device and storage medium
By setting up standardized interfaces and monitoring operation and maintenance data in the domestic IT innovation environment, as well as correlating and analyzing anomaly information and root cause information, the problems of low efficiency and poor accuracy of anomaly detection in the domestic IT innovation environment are solved, and efficient and accurate operation and maintenance management is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PING AN LIFE INSURANCE CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
In the context of domestic IT innovation, existing technologies cannot effectively correlate multi-dimensional anomaly information, resulting in low detection efficiency and poor accuracy of IT equipment, making it difficult to quickly locate the root cause of faults and affecting the efficiency and accuracy of operation and maintenance management.
By setting up standardized interfaces in the domestic IT innovation environment, data deployment and abnormal information acquisition are realized, operation and maintenance data monitoring and anomaly detection are carried out, the first and second abnormal information are correlated and detected, the root cause information is analyzed, and the repair information is processed. Seamless deployment and operation and maintenance are achieved by using standardized interfaces and automated tools.
It improves the efficiency and accuracy of abnormal state detection in the domestic IT innovation environment, shortens the troubleshooting time, enhances the efficiency and accuracy of operation and maintenance management, and ensures the safety and compliance of operation and maintenance.
Smart Images

Figure CN122489326A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the general technical field, and in particular to an operation and maintenance management method, apparatus, electronic device and storage medium. Background Technology
[0002] In business scenarios such as fintech and healthcare, with the continuous increase in the number and functionality of computers, IT equipment faces multiple threats. Therefore, monitoring and maintenance management of IT equipment is crucial.
[0003] Existing IT equipment anomaly detection technologies are isolated and cannot be linked to multi-dimensional anomaly information. To achieve multi-dimensional status information for operation and maintenance management in a domestic IT innovation environment, it is usually necessary to rely on human experience, which leads to low detection efficiency and poor accuracy. Summary of the Invention
[0004] The main objective of this application is to propose an operation and maintenance management method, device, electronic device, and storage medium, which aims to improve the detection efficiency and accuracy of abnormal states in the context of information technology innovation.
[0005] To achieve the above objectives, a first aspect of this application proposes an operation and maintenance management method, comprising: responding to the detection of first abnormal information in a domestic IT innovation environment, performing correlation detection on the abnormal information, determining second abnormal information and root cause information associated with the first abnormal information; analyzing the root cause information, determining repair information for the first abnormal information and the second abnormal information; and processing the first abnormal information and the second abnormal information according to the repair information.
[0006] In some embodiments, before performing correlation detection on the anomaly information detected in the domestic IT innovation environment, the operation and maintenance management method further includes: setting a standardized interface in the domestic IT innovation environment to deploy or update data of at least one system in the domestic IT innovation environment through the standardized interface; and / or obtaining the anomaly information in the domestic IT innovation environment through the standardized interface; the deployment or update of data of at least one system in the domestic IT innovation environment includes: obtaining data to be deployed in the target system, the data to be deployed including standardized deployment steps, environment configuration, script data and application release; and configuring the data to be deployed in the target system.
[0007] In some embodiments, before performing correlation detection on the abnormal information detected in the domestic IT innovation environment, the operation and maintenance management method further includes: monitoring the operation and maintenance data in the domestic IT innovation environment; the operation and maintenance data includes at least one of performance indicators, log information, event information, network traffic, and user behavior information; performing abnormal detection on the operation and maintenance status of the domestic IT innovation environment based on the acquired operation and maintenance data and corresponding set standard parameters; and acquiring the first abnormal information when an abnormality is detected in the operation and maintenance data.
[0008] In some embodiments, the process of performing correlation detection on the first abnormal information to determine the second abnormal information associated with the first abnormal information includes: determining the abnormal type of the first abnormal information; determining the associated data associated with the first abnormal information based on preset correlation information of the abnormal type and historical data; comparing the associated data with preset standard parameters; and obtaining the second abnormal information of the abnormal associated data when an abnormality is detected in the associated data.
[0009] In some embodiments, the process of performing correlation detection on the first abnormal information to determine the root cause information of the first abnormal information includes: analyzing the causal relationship between the first abnormal information and the second abnormal information, and determining the root cause information based on the analysis results; or, determining the correlation data between the first abnormal information and the second abnormal information, and determining the root cause information based on the correlation data and a preset fault database.
[0010] In some embodiments, the process of analyzing the root cause data to determine the repair information for the first anomaly information and the second anomaly information includes: analyzing the type and data indicators of the root cause data, matching the analysis results with a preset strategy database and historical processing data, determining the matching target repair path, target repair data, and repair operation; and using the target repair path, the target repair data, and the target repair operation to determine the repair information.
[0011] In some embodiments, the method further includes: during the transmission of the data to be deployed, the first abnormal information, the second abnormal data, the root cause data, and the repair information of the target system, identifying the transmitted data and obtaining sensitive information; and encrypting the sensitive information based on its category.
[0012] To achieve the above objectives, a second aspect of this application proposes an operation and maintenance management device, including a response module, an analysis module, and a processing module: the response module is used to respond to the detection of first abnormal information in the information technology innovation environment, perform correlation detection on the abnormal information, and determine second abnormal information and root cause information associated with the first abnormal information; the analysis module is used to analyze the root cause information and determine repair information for the first abnormal information and the second abnormal information; the processing module is used to process the first abnormal information and the second abnormal information according to the repair information.
[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0015] This application proposes an operation and maintenance management method, apparatus, electronic device, and storage medium, belonging to the field of financial technology technology. The operation and maintenance management method includes: responding to the detection of first abnormal information in a domestic IT innovation environment, performing correlation detection on the abnormal information to determine second abnormal information and root cause information associated with the first abnormal information; analyzing the root cause information to determine repair information for the first and second abnormal information; and processing the first and second abnormal information according to the repair information, which can improve the detection efficiency and accuracy of abnormal states in a domestic IT innovation environment. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the operation and maintenance management method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the process for obtaining the first abnormality information provided in an embodiment of this application; Figure 3 This is a schematic diagram of the process for obtaining the second abnormality information provided in an embodiment of this application; Figure 4 This is a schematic diagram of the encryption process provided in an embodiment of this application; Figure 5 This is a schematic diagram of the operation and maintenance management device provided in the embodiments of this application; Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] In the context of domestic IT innovation in business scenarios such as fintech and healthcare, numerous domestically produced software and hardware products have been introduced. These include domestically produced servers running operating systems like Kylin and Tongxin, domestic databases like DM and Kingbase, and middleware and security devices based on domestic IT innovation architectures. While these products have laid the foundation for the independent and controllable development of fintech, they have also brought new technical challenges.
[0021] First, data from business scenarios such as fintech and healthcare are characterized by high frequency and complexity. The data format and processing logic in the information technology innovation system are not entirely consistent with those in the traditional system, which causes anomaly detection algorithms trained on historical data to produce false positives and false negatives in the new environment.
[0022] Secondly, the relationships between components in the domestic IT innovation environment are complex and lack mature research and practical experience. Differences in interface specifications and communication protocols between different domestic software and hardware products make it extremely difficult to determine the second anomaly associated with the first. When core business systems in scenarios such as fintech and healthcare experience performance issues, the unclear relationships between domestic servers, operating systems, and middleware make it difficult to quickly pinpoint which component's anomaly triggered the chain reaction, increasing the time and difficulty of troubleshooting.
[0023] Finally, business processes in scenarios such as fintech and healthcare involve multiple stages and various technical components. In the context of information technology innovation, due to the diversity and autonomy of technologies, the root cause of a fault may be hidden at multiple levels. Accurately identifying the root cause requires a deep understanding of various technologies in the information technology innovation environment. However, the current reserves of relevant technical personnel and technical information are relatively insufficient, and different operation and maintenance personnel may have different analysis results for the same anomaly, which affects the efficiency and accuracy of problem solving.
[0024] Based on this, embodiments of this application provide an operation and maintenance management method, apparatus, electronic device, and storage medium, which can be applied in business scenarios such as fintech and healthcare to improve the detection efficiency and accuracy of abnormal states in the context of information technology innovation in these business scenarios. The specific implementation is illustrated in the following embodiments, first describing the operation and maintenance management method in this application.
[0025] The operation and maintenance management method provided in this application can be applied to electronic devices, which can be terminals or servers. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Furthermore, the operation and maintenance management method provided in this application can also be applied to the software of electronic devices, which can be applications that implement the operation and maintenance management method, but is not limited to the above forms.
[0026] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0027] Figure 1 This is an optional flowchart of the operation and maintenance management method provided in the embodiments of this application. Figure 1 The operation and maintenance management method described herein is applied in the information technology innovation environment. The information technology innovation environment is an IT infrastructure, system, and application ecosystem built upon domestically developed software and hardware systems. Its core objective is to achieve independent control, security, and reliability in the field of information technology, and to break free from dependence on foreign software and hardware products. The operation and maintenance management method provided in this embodiment may include, but is not limited to, steps S101 to S103.
[0028] Step S101: In response to the detection of the first abnormal information in the information technology innovation environment, perform correlation detection on the abnormal information to determine the second abnormal information and the root cause information associated with the first abnormal information.
[0029] Optionally, the first anomaly information is the data information obtained corresponding to the initial anomaly signal that is first detected in the domestic IT innovation environment and triggers subsequent related detection processes. Here, the initial anomaly signal can be obtained by real-time monitoring of operation and maintenance data, or it can be determined in response to alarm prompts.
[0030] Optionally, the second anomaly information is data information corresponding to other anomaly signals that have a direct or indirect causal relationship or synergistic influence with the first anomaly information, obtained after performing correlation detection on the "first anomaly information". Here, the second anomaly information can be second anomaly information that is causally related to the first anomaly information, such as second anomaly information caused by the first anomaly information, or first anomaly information caused by the second anomaly information. Alternatively, the second anomaly information can also be second anomaly information that is related to the first anomaly information from the same source, that is, both the first anomaly information and the second anomaly information are anomalies caused by the same root source information.
[0031] Optionally, the root cause information is the fundamental reason or core triggering factor that leads to the generation of the first abnormal information and the second abnormal information. It can be understood as the "core target" for the occurrence of the first abnormal information and / or the occurrence of the second abnormal information.
[0032] For example, when a hard disk I / O anomaly is detected, the first anomaly information of the hard disk I / O anomaly is obtained, and the first anomaly information of the hard disk I / O anomaly is correlated with detection, including detecting performance indicators (such as CPU and memory usage), log events, network traffic data, and user behavior. Here, if anomaly information such as error logs or latency peaks is detected, the second anomaly information is determined based on the corresponding log time and network traffic data. When hardware aging or configuration errors are detected, the root cause information is determined based on the data information of hardware aging or configuration errors.
[0033] In one embodiment, before performing correlation detection on the anomaly information detected in the domestic IT innovation environment, the operation and maintenance management method further includes: In a domestically developed information technology (IT) environment, standardized interfaces are set up to deploy or update data in at least one system within that environment; and / or, Obtain anomaly information in the domestic IT innovation environment through standardized interfaces; Optionally, considering the significant challenges that the complexity and diversity of the domestic IT innovation environment poses to automated operation and maintenance management, an automated adaptation layer needs to be set up as a core component to abstract and encapsulate the poor heterogeneity of the underlying domestic software and hardware. For example, the current domestic CPU architectures in the domestic IT innovation environment include Kunpeng, Phytium, and Loongson; operating systems include Kylin and UnionTech UOS; databases include DM and Kingbase; and there are middleware, etc. Optionally, standardized interfaces are set up in the automatic adaptation layer for upper-layer automated tools (such as deployment scripts, configuration management, and monitoring probes) to call, while internally implementing compatibility logic for specific domestic chip instruction sets, operating system APIs, database drivers, dependency libraries, and other details. Optionally, through the set standardized interfaces, the upper-layer automated deployment processes (such as installation, configuration, and application startup) and operation and maintenance management (such as monitoring collection and log parsing) do not need to care about the underlying differences and can be executed seamlessly in diverse domestic IT innovation environments, ensuring deployment consistency and operation and maintenance effectiveness.
[0034] Optionally, the standardized interface is a uniform calling specification provided by the automatic adaptation layer. Its core function is to enable upper-layer automation tools (deployment scripts, configuration management tools, monitoring probes) to call according to unified rules without having to care about the underlying components.
[0035] In practice, standardized interfaces typically exist as APIs (such as REST APIs or RPC interfaces) or tool plugins (such as Ansible modules or Jenkins plugins), with completely unified interface parameters and return formats. For example, when upper-layer tools call standardized commands through the adaptation layer, such as when a deployment script calls the "application installation interface," only two parameters need to be passed: "application package path and target component type (such as 'DaMeng Database')," without needing to pass "whether it's a Kunpeng or Feiteng CPU" or "whether it's a Kylin or Tongxin OS." This way, setting up standardized interfaces in the domestic IT innovation environment makes deployment scripts and configuration management more convenient, achieving seamless deployment (such as installing and configuring applications).
[0036] For example, in a domestic IT innovation environment, when the exception information of different components is "severely fragmented in format," the system log of Kylin OS is in text format ( / var / log / messages), while the error log of DM database is in XML format (dmdbms / log), and the error code definitions of the two are inconsistent. Therefore, the raw exception information of each component can be received through a standardized interface (such as through the built-in "monitoring probe adaptation module" of the interface, receiving the journalctl command of Kylin OS and the dm_log_viewer tool of DM). Here, the raw exception information is mapped to "unified structured data," such as JSON format including: {"Exception Type":"Database Connection Exception","Occurrence Time":"2024-10-01 15:30:00","Component":"DM8","Exception Description":"Connection pool maximum number of connections 50 exhausted","Error Code":"DM-6002"}. In this way, the upper-layer anomaly detection module only needs to obtain the JSON format data from the interface, without having to parse the log formats and error code rules of different components. This solves the problem of inconsistent analysis caused by the messy format of anomaly information in traditional operation and maintenance. It enables correlation detection based directly on the unified data output by the interface, ensuring the effectiveness of operation and maintenance and improving detection efficiency and accuracy.
[0037] In one implementation, deploying or updating data in at least one system within the domestic IT innovation environment includes: Obtain the data to be deployed from the target system. This data includes standardized deployment steps, environment configuration, script data, and application deployment. Configure the data to be deployed in the target system.
[0038] Environment configuration is a set of underlying parameter settings and resource configurations that ensure the data to be deployed can run normally in the target system (such as development, testing, or production environments). Its core function is to provide a consistent operating foundation for applications and scripts, avoiding deployment failures or functional abnormalities due to environmental differences. Optionally, standardized deployment steps form the "process skeleton" of the data to be deployed. The core is to break down complex deployment actions into reusable, unambiguous, and automatically executable ordered steps, avoiding deployment failures or environmental inconsistencies due to differences in manual operation. Optionally, application release is the core material of the data to be deployed, referring to the business-related content that needs to be updated in the target system. It is the "final deliverable" of the deployment process; all "deployment steps" and "script data" are for the successful operation of this content on the target system. Optionally, script data is the "executive muscle" of standardized deployment steps. By converting the above "process steps" into directly executable code scripts, automation tools can complete deployment without manual intervention, solving problems such as batch execution of complex operations.
[0039] Optionally, when deploying or updating data for at least one system in the domestic IT innovation environment, the standardized deployment steps, environment configuration, application release, and other operations of the target system are encapsulated into reproducible deployment data. Tools such as Jenkins, GitLab CI / CD, and Ansible can be used for this. This standardized deployment method ensures environmental consistency, allowing for rapid and accurate reproduction across development, testing, and production environments. Here, by deploying or updating data for at least one system in the domestic IT innovation environment based on the deployment data, complex operations such as version rollback are implemented, significantly shortening the application deployment cycle, reducing the risk of human error, improving the predictability and reliability of deployment, enabling frequent iterations and continuous delivery, and greatly improving operational efficiency.
[0040] In one implementation, in response to detecting first abnormal information in the domestic IT innovation environment, before performing correlation detection on the abnormal information, the operation and maintenance management method further includes monitoring operation and maintenance data in the domestic IT innovation environment, such as... Figure 2 As shown, the specific steps include: Step S201: Monitor the operation and maintenance data in the domestic IT innovation environment.
[0041] Optionally, the operation and maintenance data in the domestic IT innovation environment can be monitored in real time or based on a preset time period to determine if any anomalies exist. Here, for operation and maintenance data that has experienced faults or requires special attention, a higher frequency of detection can be set accordingly.
[0042] In practice, key data affecting core business operations (such as transaction response time in financial information technology innovation systems and lock wait counts in domestic databases) are collected in real time to ensure immediate detection of anomalies. For non-core data (such as daily server disk usage and ordinary user login logs), data can be collected according to preset time periods to reduce resource consumption. For data that has "experienced failures" or is "high-risk" (such as domestic middleware that has historically experienced downtime or servers carrying sensitive business operations), a higher frequency can be set for "focused monitoring."
[0043] Optionally, monitoring can be dynamically adjusted based on the characteristics of the domestic IT innovation environment. In practice, the frequency of log information collection from nodes involved in the upgrade can be temporarily increased to promptly identify compatibility issues.
[0044] Optionally, operational data includes at least one of the following: performance metrics, log information, event information, network traffic, and user behavior information. Performance metrics reflect the "operational load, resource consumption, and responsiveness" of the hardware and software, and are core data for assessing system health. Log information refers to text data automatically recorded during operation, including operation traces, error records, and status changes. Event information refers to specific key behaviors or status changes defined in operations and maintenance. Network traffic refers to network-level data such as data transmission volume, connection count, and protocol types within and outside the domestic IT innovation environment. User behavior information is the behavioral trajectory of users operating within the domestic IT innovation environment.
[0045] Step S202: Based on the acquired operation and maintenance data and the corresponding set standard parameters, perform anomaly detection on the operation and maintenance status of the domestic IT innovation environment; Optionally, standard parameters can be set for different types of operation and maintenance data. Then, by comparing the acquired operation and maintenance data with the corresponding standard parameters, it can be determined whether the data is abnormal. For example, real-time collected CPU utilization (e.g., "92%, lasting 3 minutes") can be compared with standard parameters. If it is found that "it exceeds the peak value and the duration is too long," then the operation and maintenance data is determined to be abnormal.
[0046] Step S203: Determine if there are any anomalies in the operation and maintenance data. If yes, proceed to step S204; otherwise, continue to step S201.
[0047] Step S204: Obtain the first abnormality information based on the detected abnormality.
[0048] Optionally, when obtaining the first abnormal information, the first abnormal information may include identification information, such as abnormal ID, abnormal time, abnormal source identifier, etc.; the first abnormal information may include abnormal characteristics, such as abnormal type, abnormal data, abnormal performance characteristics, etc.; the first abnormal information may include related information, such as related business, initial scope of impact, and domestic IT innovation adaptation characteristics, etc.
[0049] In this way, by acquiring the first abnormal information, the ambiguous abnormal signals can be transformed into structured, analyzable, and accurate data, which helps to ensure the efficiency, security, and compliance of the operation and maintenance of the information technology innovation environment.
[0050] In one implementation, such as Figure 3 As shown, correlation detection is performed on the first abnormal information to determine the second abnormal information associated with the first abnormal information, including: Step S301: Determine the anomaly type of the first anomaly information; Optionally, the types of anomalies in the first anomaly information include performance indicator anomalies, log information anomalies, event information anomalies, network traffic anomalies, and user behavior anomalies. Among these, performance indicators are the "quantitative reflection of the health" of the domestic IT innovation environment. Performance indicator anomalies directly reflect component resource shortages or performance degradation, and are the most basic anomaly type that easily triggers alarms. Examples of performance indicator anomalies include excessive CPU / memory / disk usage and response delays in domestic databases. Logs are the "operational trajectory recorder" of the domestic IT innovation environment. Log information anomalies expose hidden component faults (such as configuration errors and compatibility issues) through text information. They need to be identified in conjunction with the specific log characteristics of domestic components. Examples of log information anomalies include error codes in domestic operating system logs (such as the "ERROR" level logs in the Kylin system) and exception stacks in application logs. Events are the "critical node markers" in domestic IT innovation operations and maintenance. Event information anomalies focus on "active operation results" or "component state mutations," directly related to the effectiveness of operations and maintenance actions (such as configuration changes and security defenses). Examples of event information anomalies include "intrusion alarm events" triggered by domestic security devices and configuration change failure events. The network serves as the "data transmission channel" in a domestically developed IT environment. Abnormal network traffic focuses on anomalies in traffic volume, source, and protocol, potentially linked to network equipment malfunctions or security attacks (such as DDoS attacks). Examples of abnormal network traffic include sudden increases in port traffic on domestically produced switches and unusual IP connections. User behavior serves as the "human-computer interaction record" in a domestically developed IT environment. Abnormal user behavior focuses on unauthorized operations by "human or system accounts," directly related to data security (such as sensitive data leaks) or operational errors. Examples of abnormal user behavior include unauthorized users operating domestically produced core systems and bulk downloading of sensitive data.
[0051] Step S302: Based on the preset association information of the anomaly type and historical data, determine the association data associated with the first anomaly information.
[0052] Optionally, historical data includes past anomaly cases and related records occurring in the domestic IT innovation environment. For example, if historical data determines that three past instances of abnormal Kunpeng CPU usage were accompanied by the exhaustion of the thread pool of the domestic middleware (TongWeb), then when a new "CPU anomaly" occurs, the system will prioritize checking the "TongWeb thread pool status." Similarly, if "multiple login failures by users in a certain department are often followed by a sudden increase in data download volume," then when that department experiences another "login failure anomaly," the system will proactively check the "data download traffic."
[0053] In some embodiments, when an abnormal login (such as unauthorized IP access) is detected, network traffic and user behavior data can be analyzed, and the IP’s historical attack records (from security logs), server vulnerability status (from vulnerability scan data), and recent configuration changes (from configuration management database) can be correlated to determine a second abnormal information.
[0054] Step S303: Compare the associated data with the preset standard parameters to determine if there is any anomaly in the associated data. If yes, proceed to step S304; otherwise, continue to step S302.
[0055] Optionally, the preset standard parameters include preset static thresholds, dynamic baselines, and domestically developed innovation-specific standards. The preset static thresholds are fixed values set based on component performance limits or industry standards. The dynamic baseline is a "normal fluctuation range" calculated based on historical data, which can be dynamically adjusted according to time and business scenarios. Domestically developed innovation-specific standards are unique standards set for the characteristics of domestically produced components. For example, based on key indicators of multi-core collaboration under domestic architecture, the inter-core communication latency of Phytium CPUs is set to a standard range of less than or equal to 50ns.
[0056] Step S304: When an anomaly is detected in the associated data, obtain the second anomaly information of the abnormal associated data.
[0057] Optionally, after obtaining multiple associated data related to the first abnormal data, each associated data is compared with preset standard parameters. Thus, when an anomaly is detected in the associated data, corresponding second anomaly information is obtained. By obtaining this second anomaly information, the root cause can be effectively avoided, helping to improve the efficiency and accuracy of anomaly detection.
[0058] In one embodiment, correlation detection is performed on the first abnormal information to determine the root source information related to the first abnormal information, including: Analyze the causal relationship between the first and second abnormal information, and determine the root cause information based on the analysis results; or, Determine the correlation data between the first and second abnormal information, and determine the root cause information based on the correlation data and the preset fault database.
[0059] Optionally, the first anomaly corresponding to the first anomaly information and the second anomaly corresponding to the second anomaly information are arranged in chronological order or logical influence relationship, and it is checked whether the "hypothetical cause" can really lead to the "effect" (e.g., assuming "high disk I / O" is the cause and "slow database response" is the effect, it is necessary to verify whether "high disk I / O directly leads to database read / write latency"). Optionally, based on the causal relationship between the first and second anomaly information, one of them is determined as the source, and the corresponding root cause information is determined.
[0060] In practice, the first abnormal information can be the root cause of the second abnormal information, or the second abnormal information can be the root cause of the first abnormal information.
[0061] For example, when the first anomaly is detected as an abnormal login (such as unauthorized IP access), network traffic and user behavior data can be analyzed, including historical attack records of that IP (from security logs), server vulnerability status (from vulnerability scan data), and recent configuration changes (from the configuration management database). Upon detecting an anomaly, a second anomaly is identified. Here, a causal analysis is performed on the first and second anomalies to determine that the abnormal login was caused by the second anomaly (such as the existence of a vulnerability). Thus, the second anomaly is identified as the root cause.
[0062] Optionally, key feature data between the first and second anomaly information can be collected, such as anomaly type, occurrence time, involved components, and indicator values. This associated data is then compared with a pre-defined "fault database" (which records the characteristics, root causes, and solutions of historical faults) to find the most similar cases. Optionally, if a match is successful, the root cause of the historical case is directly reused; if there is a partial match, a weighted similarity judgment is applied.
[0063] For example, real-time monitoring data (such as response time and throughput) and historical performance trends can be analyzed. Bottleneck detection algorithms (such as queue theory analysis or resource utilization models) can be used to identify bottlenecks (such as insufficient database connection pool) when metrics exceed normal ranges (such as CPU utilization consistently above 90%).
[0064] Optionally, in actual IT innovation operation and maintenance, the two methods can often be used in combination, such as first using causal analysis to sort out the direction, and then using a fault database to verify, so as to ensure the accuracy of root cause location.
[0065] Step S102: Analyze the root cause information to determine the repair information for the first and second abnormal information.
[0066] In one embodiment, the root cause data is analyzed to determine repair information for the first and second anomalies, including: The types and metrics of the root cause data are analyzed, and the analysis results are matched with the preset strategy database and historical processing data to determine the matching target repair path, target repair data and repair operations. The target repair path, target repair data, and target repair operations are used to determine the repair information.
[0067] Optionally, the preset strategy database includes a built-in automated rule base, which is determined based on Ansible scripts and elastic scaling strategies. Optionally, it can analyze the type and metrics of root source data based on big data, extract quantitative features, and match these quantitative features with the preset strategy database to determine the most suitable target strategy for the problem, and correspondingly determine the target repair path, target repair data, and repair operations.
[0068] For example, if disk I / O utilization is detected to be ≥90% for 5 minutes, the elastic scaling rules based on the domestic cloud platform will determine that the storage volume needs to be expanded for repair; if the Kylin system's "vm.swappiness" parameter is not 10 (vendor recommended value), the repair information that can be executed is sysctl -w vm.swappiness=10; if a hardware failure of a domestic server is detected, the domestic high availability solution based on Keepalived can provide repair information that triggers a master-slave switchover script.
[0069] Optionally, historical processing data includes historical cases learned through AI / ML models. In this way, by analyzing the root cause data types and data indicators through AI / ML models, the target repair path, target repair data, and repair operations can be determined.
[0070] For example, if historical processing data records similar disk I / O bottlenecks, the issue can be resolved by adjusting the report task execution time and increasing IOPS. In this way, based on the current ML model and historical processing data, it predicts that if the current I / O load is not addressed, it may lead to database crashes after a certain period, and provides corresponding quick scaling solutions and related repair information.
[0071] In this way, by monitoring the operation and maintenance data in the domestic IT innovation environment, a closed loop of analysis, decision-making, and execution can be achieved. This is especially suitable for the characteristics of the domestic IT innovation environment, which has a large number of domestic software and hardware and high operation and maintenance complexity, and greatly improves the efficiency and stability of anomaly handling.
[0072] In one embodiment, the method further includes: During the transmission of data to be deployed, first anomaly information, second anomaly data, root cause data and repair information of the target system, the transmitted data is identified and sensitive information is obtained. Sensitive information is encrypted based on its category.
[0073] like Figure 4 As shown, the categories of sensitive information include statically stored sensitive configurations (such as passwords in configuration files), dynamically transmitted sensitive instructions (such as SQL commands in repair scripts), and identity credential information (such as server login passwords and API keys).
[0074] Optionally, for sensitive configurations stored statically, Advanced Encryption Standard - 256-bit encryption (i.e., AES-256 encryption, Vault management key) can be configured; for dynamic sensitive instructions in transit, Transport Layer Security / Secure Sockets Layer Protocol (i.e., TLS / SSL, temporary key encryption) can be configured; and for identity credential information, dynamic key / credential generation (i.e., Vault dynamic generation) can be configured.
[0075] In actual implementation, when storing sensitive configurations in static storage, the password (e.g., "dm_password=123456") in the configuration file (e.g., dm_svc.conf for DM database) is automatically recognized before deployment. Here, Vault's API can be called to obtain the AES encryption key (Vault automatically generates and stores the key, requiring no manual intervention). Optionally, the password field can be encrypted using the AES-256 algorithm to generate "dm_password=Enc (123456, AES key)". Optionally, the encrypted configuration file is transmitted via a TLS link, and the decryption key is only temporarily injected by Vault when the target server performs deployment (and is destroyed after use).
[0076] In practical implementation, when encrypting dynamically sensitive instructions during transmission, and when identifying sensitive operation instructions in the repair information, a one-time temporary symmetric key can be generated before transmission, based on the domestic SM4 algorithm and adapted to the domestically developed encryption standard. Optionally, the temporary key can be used to encrypt SQL instructions before transmission via a TLS link. Upon receiving the instruction, the target server obtains the temporary key through a pre-shared key exchange mechanism, decrypts it, and executes the instruction. Here, the temporary key transmission process can also be encrypted using TLS.
[0077] In practice, when encrypting identity credential information, after the "server root password" in the data to be deployed is recognized, a request is made to Vault for "temporary login credentials for the server." Correspondingly, Vault generates a temporary password valid for 10 minutes based on RBAC permissions (accessible only to the deployment role on the target server). This temporary password is transmitted to the target server via TLS encryption for deployment verification, automatically expires upon timeout, and is never stored or transmitted in plaintext.
[0078] In this way, encryption technology is used to ensure the security of critical data. Strong encryption (such as TLS / SSL, AES, and the Secrets management tool Vault) is applied to deployment scripts, configuration files, images / packages in transit, and stored sensitive information (such as passwords and keys). This ensures that only authorized entities (people / systems) can trigger deployments, modify configurations, or access sensitive operational data. This effectively prevents risks such as unauthorized access, man-in-the-middle attacks, and sensitive information leakage, building a solid security barrier for automated processes.
[0079] Step S103: Process the first and second abnormal information based on the repair information.
[0080] Optionally, based on the determined repair information applicable to standardized operations (such as parameter modification or service restart), the script can be executed by calling the "standardized interface" of the domestic IT innovation environment. For example, a Python script adapted to the Kunpeng architecture can be called, and the MAX_SESSIONS parameter can be modified through the dm_svc.conf configuration file of the DM database. The script will automatically record operation logs to a domestic logging system.
[0081] Optionally, when repairing the first and second anomalies based on the repair information, the validity of the repair information can also be verified. Verified and valid repair operations (such as scripts for adjusting memory parameters of domestic servers and scripts for optimizing database connection pools) can be included in the "automatic repair library" of the CI / CD pipeline. When a similar anomaly occurs again, the script can be directly called to handle it automatically, reducing manual intervention.
[0082] In one embodiment, the method further includes: acquiring time-series data; the time-series data includes historical traffic logs and business metrics; inputting the time-series data into a preset target prediction model for analysis to determine the prediction result of the business traffic trend; and outputting the prediction result.
[0083] Historical traffic logs are the original records of user visits (such as hourly UV / PV, page navigation paths, and dwell time for e-commerce platforms; and minutely form submissions for government service platforms). These need to be aggregated by time granularity to form a "time-traffic" sequence. For example, it can be represented as: "2024-10-01 08:00, UV=1200; 09:00, UV=2500". Business performance data that are strongly correlated with traffic (such as "add-to-cart / orders" for e-commerce platforms, and "number of bullet comments / gifts" for live streaming platforms) are also important. These metrics have a linkage with traffic (for example, an increase in orders is usually accompanied by a traffic peak), which can help the model capture the business-driven logic behind the traffic.
[0084] Alternatively, external influencing factors can also be considered, such as marketing campaign data, environmental factors, and unforeseen events.
[0085] Alternatively, an ARIMA or LSTM neural network can be selected as the prediction model. In this way, the prediction model is learned based on historical traffic logs, business metrics, and external influencing factors to predict future traffic trends (such as predicting peak periods). In one implementation, current traffic data can be accessed in real time, and new external factors (such as a sudden trending topic event expected to bring an additional 1 million traffic) can be added to ensure the timeliness of the model input. An "online learning" mechanism (such as incremental training of LSTM) is used to fine-tune the model parameters hourly with the latest data. Optionally, when the predicted value deviates from the actual traffic collected in real time by more than a threshold, it can be identified as a first / second anomaly. Based on the increased business traffic, real-time calibration logic is triggered accordingly (e.g., increasing the prediction by 10 points to 6 million), and repair information is simultaneously determined, such as expanding server resources in advance.
[0086] In summary, the above-mentioned operation and maintenance management methods can improve the efficiency and accuracy of abnormal state detection in the context of information technology innovation.
[0087] This application embodiment monitors the domestic IT innovation environment in real time and automatically performs repair processing when abnormal information is acquired. In response to the detection of the first abnormal information in the domestic IT innovation environment, it performs correlation detection on the abnormal information to determine the second abnormal information and root cause information associated with the first abnormal information; it analyzes the root cause information to determine the repair information for the first and second abnormal information; and it processes the first and second abnormal information according to the repair information, thereby improving the detection efficiency and accuracy of abnormal states in the domestic IT innovation environment.
[0088] Please see Figure 5This application provides an operation and maintenance management device 500, which can implement the above-mentioned operation and maintenance management method, including a response module 501, an analysis module 502 and a processing module 503.
[0089] Among them, the response module 501 is used to respond to the detection of the first abnormal information in the information innovation environment, perform correlation detection on the abnormal information, and determine the second abnormal information and root cause information associated with the first abnormal information; Analysis module 502 is used to analyze the root cause information and determine the repair information for the first and second anomalies. The processing module 503 is used to process the first abnormal information and the second abnormal information based on the repair information.
[0090] The specific implementation method of this operation and maintenance management device is basically the same as the specific implementation method of the above-mentioned operation and maintenance management method, and will not be described again here.
[0091] Please see Figure 6 This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described operation and maintenance management method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0092] Please see Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 601 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 602 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 602 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called and executed by the processor 601 using the operation and maintenance management method of the embodiments of this application. The input / output interface 603 is used to implement information input and output; The communication interface 604 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 605 transmits information between various components of the device (e.g., processor 601, memory 602, input / output interface 603, and communication interface 604); The processor 601, memory 602, input / output interface 603, and communication interface 604 are connected to each other within the device via bus 605.
[0093] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described operation and maintenance management method.
[0094] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0095] This application also provides a computer program product, which is stored in a storage medium and implements the above-described operation and maintenance management method when executed by at least one processor.
[0096] The operation and maintenance management method, device, electronic device, and storage medium provided in this application belong to the field of financial technology technology. In response to the detection of first abnormal information in the information technology innovation environment, the method performs correlation detection on the abnormal information to determine second abnormal information and root cause information associated with the first abnormal information; analyzes the root cause information to determine repair information for the first and second abnormal information; and processes the first and second abnormal information based on the repair information. Through deep learning and real-time analysis of massive system operation monitoring data (including performance indicators, logs, events, network traffic, user behavior, etc.), the model can automatically identify complex patterns, correlations, and abnormal situations, far exceeding traditional threshold-based alarms. This enables operation and maintenance teams to shift from reactive firefighting to proactive prevention and intervention, achieving a leap from "monitoring" to "insight," significantly improving system stability and user experience, while optimizing resource utilization.
[0097] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0098] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0100] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0101] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0102] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0103] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0104] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0106] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0107] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An operation and maintenance management method, characterized in that, include: In response to the detection of first abnormal information in the information technology innovation environment, the abnormal information is correlated and detected to determine second abnormal information and root cause information associated with the first abnormal information; The root cause information is analyzed to determine the repair information for the first anomaly information and the second anomaly information; The first and second anomaly information are processed based on the repair information.
2. The method according to claim 1, characterized in that, Before performing correlation detection on the anomaly information detected in the domestic IT innovation environment, the operation and maintenance management method further includes: A standardized interface is set up in the domestic IT innovation environment to deploy or update data of at least one system within the environment; and / or, The abnormal information in the domestic IT innovation environment is obtained through the standardized interface. The deployment or updating of data in at least one system within the domestic IT innovation environment includes: Obtain the data to be deployed from the target system, including standardized deployment steps, environment configuration, script data, and application deployment. Configure the data to be deployed in the target system.
3. The method according to claim 1, characterized in that, Before performing correlation detection on the anomaly information detected in the domestic IT innovation environment, the operation and maintenance management method further includes: The operation and maintenance data in the aforementioned domestic IT innovation environment are monitored; the operation and maintenance data includes at least one of performance indicators, log information, event information, network traffic, and user behavior information. Based on the acquired operation and maintenance data and the corresponding set standard parameters, anomaly detection is performed on the operation and maintenance status of the domestic IT innovation environment. When an anomaly is detected in the operation and maintenance data, the corresponding first anomaly information is obtained.
4. The method according to claim 1, characterized in that, The step of performing correlation detection on the first abnormal information to determine the second abnormal information associated with the first abnormal information includes: Determine the anomaly type of the first anomaly information; Based on the preset association information of the anomaly type and historical data, determine the association data associated with the first anomaly information; The associated data is compared with preset standard parameters. When an anomaly is detected in the associated data, the second anomaly information of the abnormal associated data is obtained.
5. The method according to claim 1, characterized in that, The step of performing correlation detection on the first abnormal information to determine the root source information related to the first abnormal information includes: The causal relationship between the first abnormal information and the second abnormal information is analyzed, and the root cause information is determined based on the analysis results; or, Determine the correlation data between the first abnormal information and the second abnormal information, and determine the root cause information based on the correlation data and a preset fault database.
6. The method according to claim 1, characterized in that, The step of analyzing the root cause data to determine the repair information for the first anomaly information and the second anomaly information includes: The types and data indicators of the root cause data are analyzed, and the analysis results are matched with the preset strategy database and historical processing data to determine the matching target repair path, target repair data and repair operation. The target repair path, the target repair data, and the target repair operation are used to determine the repair information.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: During the transmission of the target system's data to be deployed, the first abnormal information, the second abnormal data, the root cause data, and the repair information, the transmitted data is identified to obtain sensitive information. The sensitive information is encrypted based on its category.
8. An operation and maintenance management device, characterized in that, Includes a response module, an analysis module, and a processing module: The response module is used to respond to the detection of first abnormal information in the information technology innovation environment, perform correlation detection on the abnormal information, and determine second abnormal information and root cause information associated with the first abnormal information. The analysis module is used to analyze the root cause information and determine the repair information for the first abnormal information and the second abnormal information; The processing module is used to process the first abnormal information and the second abnormal information according to the repair information.
9. An electronic device, characterized in that, The electronic device includes a memory and a synchronizer. The memory stores a computer program, and the synchronizer executes the computer program to implement the operation and maintenance management method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the synchronizer, it implements the operation and maintenance management method as described in any one of claims 1 to 7.