AIOps Multitask Scheduling Method and System Driven by Collaborative Large and Small Models

By performing consistency calibration and causal information analysis on heterogeneous operation and maintenance timing data sets, combined with the coordinated scheduling strategy of size and model, the problem of over-optimization traps in AIOps technology is solved, and efficient operation and maintenance response and system stability are achieved.

CN120085999BActive Publication Date: 2025-07-25QINGDAO HAIBO TECH INFORMATION SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510587616.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-25
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing AIOps technology is prone to ignore potential overoptimization traps during operation and maintenance, and it is difficult to make effective real-time responses to high-dynamic scenarios, resulting in frequent personnel intervention and adjustments, and the response speed is difficult to meet the needs of high-dynamic operation and maintenance.

Method used

By obtaining heterogeneous operation and maintenance timing data sets, performing data consistency calibration, determining multi-source standardized operation and maintenance data flow, analyzing the dynamic causal information of operation and maintenance indicators, determining over-optimized target event information, and real-time operation and maintenance adjustments are made according to the coordinated scheduling strategy of the size and model, and outputting collaborative operation and maintenance reports.

Benefits of technology

It effectively solves the over-optimization trap problem, improves the response speed of the operation and maintenance process, ensures the system's high dynamic response needs, and ensures system stability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085999B_ABST
    Figure CN120085999B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence operation and maintenance technology, and particularly to an AIOps multi-task scheduling method and system driven by the collaboration of large and small models. The method includes: obtaining a heterogeneous operation and maintenance time series data set, calibrating the data consistency of the heterogeneous operation and maintenance time series data set to determine a multi-source standardized operation and maintenance data stream; analyzing the multi-source standardized operation and maintenance data stream to determine the dynamic causal information of operation and maintenance metrics, and determining the over-optimized target event information according to the dynamic causal information of operation and maintenance metrics; determining a real-time operation and maintenance task chain according to the over-optimized target event information, and determining a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain; making real-time operation and maintenance adjustments to the system state according to the collaborative scheduling strategy for large and small models, and determining and outputting a collaborative operation and maintenance report. This application effectively solves the problem of over-optimization traps in the operation and maintenance process using AIOps technology, making the operation and maintenance process highly match the high-dynamic response requirements of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of artificial intelligence operation and maintenance, and particularly to an AIOps multi-task scheduling method and system driven by the collaboration of large and small models. Background Art

[0002] With the wide application of technologies such as cloud computing, containerization, and microservices, operation and maintenance work has become increasingly complex. Traditional operation and maintenance methods are difficult to meet the increasingly complex operation and maintenance requirements. AIOps (Artificial Intelligence for IT Operations), as a technology that improves operation and maintenance efficiency through data analysis and machine learning methods, has become the focus of attention in the operation and maintenance industry.

[0003] However, existing AIOps technologies usually use a single model and preset optimization goals to uniformly process different operation and maintenance tasks, which easily ignores potential over-optimization traps and is difficult to make effective real-time operation and maintenance responses to high-dynamic scenarios. As a result, during the operation and maintenance process based on AIOps technology, frequent human intervention and adjustment are still required, and the response speed is difficult to meet the high-dynamic operation and maintenance requirements. Summary of the Invention

[0004] This application provides an AIOps multi-task scheduling method and system driven by the collaboration of large and small models to solve the above technical problems.

[0005] In a first aspect, this application provides an AIOps multi-task scheduling method driven by the collaboration of large and small models, and the method includes:

[0006] Obtain a heterogeneous operation and maintenance time series data set, perform data consistency calibration on the heterogeneous operation and maintenance time series data set, and determine a multi-source standardized operation and maintenance data stream;

[0007] Analyze the multi-source standardized operation and maintenance data stream, determine the dynamic causal information of operation and maintenance metrics, and determine the over-optimization target event information according to the dynamic causal information of operation and maintenance metrics;

[0008] Determine a real-time operation and maintenance task chain according to the over-optimization target event information, and determine a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain;

[0009] Perform real-time operation and maintenance adjustment on the system state according to the collaborative scheduling strategy for large and small models, and determine and output a collaborative operation and maintenance report.

[0010] Through this solution, data consistency calibration is performed on the heterogeneous operation and maintenance time series dataset to determine the multi-source standardized operation and maintenance data stream. On this basis, the operation and maintenance index dynamic causal information used to characterize the dynamic causal relationship of each operation and maintenance index is analyzed, so as to determine the over-optimization target event information. According to the over-optimization target event information, a real-time operation and maintenance task chain required to solve the over-optimization event is formed. Based on the principle of collaborative operation of large and small models, a collaborative scheduling strategy for large and small models is determined, so as to perform real-time operation and maintenance adjustment on the system state, and provide the corresponding collaborative operation and maintenance report to the operation and maintenance personnel, effectively solving the over-optimization trap problem existing in the operation and maintenance process using AIOps technology, making the operation and maintenance process highly match the high-dynamic response requirements of the system, significantly improving the response speed of the operation and maintenance process while ensuring the operation and maintenance effect, and ensuring the system stability and user experience in real time.

[0011] Optionally, the data consistency calibration of the heterogeneous operation and maintenance time series dataset to determine the multi-source standardized operation and maintenance data stream includes:

[0012] According to the heterogeneous operation and maintenance time series dataset, extract the operation and maintenance time series data corresponding to different types of operation and maintenance indicators;

[0013] According to several pieces of the operation and maintenance time series data, determine the service node to which each piece of the operation and maintenance time series data belongs, and thereby determine the network communication hop count between any two service nodes;

[0014] According to the network communication hop count, determine the topological weights of different data nodes in any two pieces of the operation and maintenance time series data at the corresponding time nodes;

[0015] Analyze several pieces of the operation and maintenance time series data, extract the timestamps corresponding to different data nodes in each piece of the operation and maintenance time series data, and determine the time difference impact factor according to the timestamps;

[0016] Based on several pieces of the operation and maintenance time series data, determine the minimum alignment cost between any two pieces of the operation and maintenance time series data according to the topological weights and the time difference impact factor;

[0017] According to the minimum alignment cost, perform unified alignment processing on different data nodes in several pieces of the operation and maintenance time series data to construct the multi-source standardized operation and maintenance data stream.

[0018] Through this solution, starting from two dimensions of network structure and time, according to the network communication hops between different service nodes and the timestamps of different data nodes reflected in the operation and maintenance time-series data, the topological weight used to characterize the impact of the communication distance between service nodes on data consistency and the time difference impact factor used to characterize the impact of the timestamp differences of different operation and maintenance time-series data on data alignment are respectively analyzed. Furthermore, the minimum alignment cost required to achieve data time-series alignment is comprehensively evaluated from the two dimensions of network structure and time. According to the minimum alignment cost, different data nodes in several operation and maintenance time-series data are uniformly aligned to construct a multi-source standardized operation and maintenance data stream, significantly improving the data alignment speed while ensuring the alignment accuracy.

[0019] Optionally, based on the several operation and maintenance time-series data, according to the topological weight and the time difference impact factor, determining the minimum alignment cost between any two of the operation and maintenance time-series data is specifically the following formula: ;

[0020] Wherein, is the minimum alignment cost, is the topological weight, is the operation and maintenance time-series data at the time node the observed value of the data node at, is the operation and maintenance time-series data at the time node the observed value of the data node at, is the time difference impact factor.

[0021] Through this solution, using mathematical analysis means, based on several operation and maintenance time-series data, according to the topological weight and the time difference impact factor, the minimum alignment cost between any two operation and maintenance time-series data is quantified. Through the topological weight mechanism, the interference of multi-hop node data deviation is reduced, so that the aligned data truly reflects the state correlation of adjacent service nodes. Through the time difference impact factor, the matching of pseudo-event sequences caused by clock asynchrony is effectively suppressed, improving the accuracy of subsequent operation and maintenance event chain reconstruction.

[0022] Optionally, according to the network communication hops, determining the topological weight of different data nodes in any two of the operation and maintenance time-series data at the corresponding time nodes is specifically the following formula:

[0023] ;

[0024] Wherein, is the topological weight, is the service node and the service node the network communication hops between, is a preset attenuation coefficient;

[0025] Based on the time stamp, determine the time difference impact factor, specifically as the following formula:

[0026] ;

[0027] wherein, is the time difference impact factor, is the operation and maintenance time series data at the time node the time stamp of the data node at, the operation and maintenance time series data at the time node the time stamp of the data node at, is a preset time offset threshold.

[0028] Through this solution, by means of mathematical analysis, according to the network communication hop count and time stamp, the topological weight and the time difference impact factor are respectively quantified to clarify the mathematical characteristics of the topological weight and the time difference impact factor, improve the scientificity and accuracy of the topological weight and the time difference impact factor, and further improve the accuracy of the minimum alignment cost.

[0029] Optionally, analyzing the multi-source standardized operation and maintenance data stream, determining the dynamic causal information of operation and maintenance indicators, and according to the dynamic causal information of operation and maintenance indicators, determining the over-optimized target event information, including:

[0030] Analyze the multi-source standardized operation and maintenance data stream, determine the change amount of different operation and maintenance indicators at different time points, thereby determining the change rate corresponding to each operation and maintenance indicator, and according to the change rate, determine the dynamic lag time window corresponding to each operation and maintenance indicator;

[0031] Based on the change amount of each operation and maintenance indicator at different time points and the dynamic lag time window, determine the lag impact factor between any two operation and maintenance indicators;

[0032] According to the preset index level attribution information, analyze the multi-source standardized operation and maintenance data stream, determine the level correlation index between any two operation and maintenance indicators and the collaborative impact weight of each operation and maintenance indicator;

[0033] Based on the level correlation index and the collaborative impact weight, according to the dynamic lag time window and the lag impact factor, determine the causal correlation index between any two operation and maintenance indicators;

[0034] Analyze the multi-source standardized operation and maintenance data stream, and determine several obvious abnormal operation and maintenance indicators;

[0035] Based on the causal association index, analyze a number of the explicit abnormal operation and maintenance indicators, determine a number of implicit root cause indicators, and judge whether there is an over-optimization event according to the number of the explicit abnormal operation and maintenance indicators and the corresponding number of the implicit root cause indicators;

[0036] If there is the over-optimization event, use the corresponding number of the explicit abnormal operation and maintenance indicators and the implicit root cause indicators as the over-optimization target event information.

[0037] Through this solution, analyze the dynamic causal relationship of different operation and maintenance indicators from multiple dimensions of time series change, hierarchical structure, and collaborative influence. By evaluating the causal association index between different operation and maintenance indicators and combining the current explicit abnormal operation and maintenance indicators, determine the root cause operation and maintenance indicators corresponding to the explicit abnormal operation and maintenance indicators, thereby judging whether there is an over-optimization event, and use the corresponding number of explicit abnormal operation and maintenance indicators and implicit root cause indicators as the over-optimization target event information, making the judgment of the over-optimization event more comprehensive and accurate, and at the same time providing a scientific data basis for the subsequent operation and maintenance task analysis based on the optimization target event information.

[0038] Optionally, based on the hierarchical association index and the collaborative influence weight, determine the causal association index between any two operation and maintenance indicators according to the dynamic lag time window and the lag influence factor, specifically as the following formula:

[0039] ;

[0040] Among them, is the current time point for the operation and maintenance indicator and the operation and maintenance indicator the causal association index between them, is the preset data dimension conversion coefficient, is the operation and maintenance indicator and the operation and maintenance indicator the hierarchical association index between them, is the collaborative influence weight of the operation and maintenance indicator , is the dynamic lag time window, is the operation and maintenance indicator at the time point the change amount, is the current time point for the operation and maintenance indicator and the operation and maintenance indicator the lag influence factor between them.

[0041] Through this solution, by means of mathematical analysis, based on the hierarchical correlation index and the collaborative influence weight, and according to the dynamic lag time window and the lag influence factor, the causal correlation index used to characterize the dynamic causal association strength between operation and maintenance indicators is comprehensively quantified from three dimensions: hierarchical influence, collaborative influence, and lag influence, so as to improve the comprehensiveness and accuracy of the causal correlation index, and further improve the judgment accuracy of over-optimization events.

[0042] Optionally, the lag influence factor between any two operation and maintenance indicators is determined based on the change amount of each operation and maintenance indicator at different time points, specifically as the following formula:

[0043] ;

[0044] Wherein, is the current time point under the operation and maintenance indicator and the operation and maintenance indicator the lag influence factor therebetween, is the preset maximum lag time, is the covariance calculation function, is the dynamic lag time window, is the operation and maintenance indicator at the time point the change amount thereof, is the operation and maintenance indicator at the time point the change amount thereof.

[0045] Through this solution, by means of mathematical analysis, based on the change amount of each operation and maintenance indicator at different time points, the overall lag effect shown by the current operation and maintenance indicator under the influence of different moments within the dynamic lag time window is described, and then the lag time window is normalized to quantify the lag influence factor, so as to improve the scientificity and accuracy of the lag influence factor.

[0046] Optionally, the determining of the real-time operation and maintenance task chain according to the over-optimization target event information and the determining of the collaborative scheduling strategy of the large and small models according to the real-time operation and maintenance task chain include:

[0047] Analyze a number of the explicit abnormal operation and maintenance indicators and the implicit root cause indicators, and taking each of the implicit root cause indicators as the starting task node, and taking the corresponding number of the explicit abnormal operation and maintenance indicators as the chain nodes, construct a number of operation and maintenance target data chains;

[0048] Based on the starting task node, analyze the operation and maintenance target data chain to determine the path depth of each chain node;

[0049] Analyze several of the operation and maintenance target data chains according to the preset operation and maintenance knowledge graph, extract the operation and maintenance tasks corresponding to each starting task node or chain node and the dependency index between each node, construct several target task chains, and determine the event complexity of each target task chain and the task importance index of each target task in each target task chain;

[0050] Determine the dynamic scheduling mapping index of each target task in each target task chain according to the path depth, the dependency index, the event complexity, and the task importance index;

[0051] Determine the large model target task set and the small model target task set according to the dynamic scheduling mapping index of each target task, and construct the collaborative scheduling strategy for the large and small models.

[0052] Through this solution, starting from two dynamic dimensions of depth and breadth, and at the same time combining two static dimensions of the absolute complexity and importance of the operation and maintenance tasks themselves, comprehensively evaluate the complexity of the operation and maintenance tasks according to the path depth, the dependency index, the event complexity, and the task importance index, obtain the dynamic scheduling mapping index corresponding to different target tasks, so as to accurately measure the actual complexity of the operation and maintenance tasks. On this basis, integrate to obtain the model target task set and the small model target task set respectively, and then construct the collaborative scheduling strategy for the large and small models, so that different target tasks are highly matched with the corresponding execution model scale.

[0053] Optionally, the determining the dynamic scheduling mapping index of each target task in each target task chain according to the path depth, the dependency index, the event complexity, and the task importance index is specifically the following formula:

[0054] ;

[0055] Where is the dynamic scheduling mapping index of the th target task in the current target task chain, is the preset depth influence weight, is the path depth corresponding to the current target task, is the preset dependency influence weight, is the set of target tasks that have a dependency relationship with the current target task in the current target task chain, is the th target task in the target task set, is the th target task and the th target task, is the preset resource influence weight, is the event complexity of the target task chain corresponding to the current target task, is the task importance index corresponding to the current target task.

[0056] Through this solution, by using mathematical analysis means, based on the path depth, dependency index, event complexity, and task importance index, clarify the impact of each parameter on the evaluation of the target task complexity, and quantitatively obtain the dynamic scheduling mapping index reflecting the overall complexity of the target task, so as to accurately reflect the complexity of the target task and provide an accurate scientific data basis for the collaborative scheduling of large and small models.

[0057] In a second aspect, the present application provides an AIOps multi-task scheduling system driven by collaborative large and small models, and the system includes:

[0058] A data processing module, configured to obtain a heterogeneous operation and maintenance time series data set, perform data consistency calibration on the heterogeneous operation and maintenance time series data set, and determine a multi-source standardized operation and maintenance data stream; an over-optimization analysis module, configured to analyze the multi-source standardized operation and maintenance data stream, determine dynamic causal information of operation and maintenance indicators, and determine over-optimization target event information according to the dynamic causal information of the operation and maintenance indicators; a scheduling analysis module, configured to determine a real-time operation and maintenance task chain according to the over-optimization target event information, and determine a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain; an operation and maintenance output module, configured to perform real-time operation and maintenance adjustment on the system state according to the collaborative scheduling strategy for large and small models, and determine and output a collaborative operation and maintenance report. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0060] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0061] Figure 2 is a flowchart of a method for collaborative AIOps multi-task scheduling driven by large and small models provided by an embodiment of the present application;

[0062] Figure 3 is a schematic structural diagram of an AIOps multi-task scheduling system driven by collaborative large and small models provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.

[0064] In addition, the term "and / or" in this document is merely a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0065] The embodiments of this application will be further described in detail below with reference to the accompanying drawings of the specification.

[0066] Existing AIOps technologies usually use a single model and a preset optimization objective to uniformly process different operation and maintenance tasks, which easily ignores potential over-optimization traps and makes it difficult to make effective real-time operation and maintenance responses to highly dynamic scenarios. As a result, during the operation and maintenance process based on AIOps technology, frequent human intervention and adjustment are still required, and the response speed is difficult to meet the requirements of highly dynamic operation and maintenance.

[0067] Based on this, this application provides an AIOps multi-task scheduling method and system driven by collaborative large and small models. It calibrates the data consistency of heterogeneous operation and maintenance time series datasets to determine multi-source standardized operation and maintenance data streams. On this basis, it analyzes and obtains operation and maintenance metric dynamic causal information used to characterize the dynamic causal relationships of various operation and maintenance metrics, thereby determining over-optimization target event information. According to the over-optimization target event information, it forms a real-time operation and maintenance task chain required to solve over-optimization events. Based on the principle of collaborative operation of large and small models, it determines the collaborative scheduling strategy of large and small models, thereby making real-time operation and maintenance adjustments to the system state and providing the corresponding collaborative operation and maintenance report to operation and maintenance personnel, effectively solving the over-optimization trap problem existing in the operation and maintenance process using AIOps technology, making the operation and maintenance process highly match the high-dynamic response requirements of the system, significantly improving the response speed of the operation and maintenance process while ensuring the operation and maintenance effect, and real-time guaranteeing system stability and user experience.

[0068] Figure 1 This is a schematic diagram of an application scenario provided by this application. During the automated operation and maintenance process, by applying the method provided by this application, the over-optimization trap problem existing in the operation and maintenance process using AIOps technology is effectively solved, and the operation and maintenance process highly matches the high-dynamic response requirements of the system.

[0069] Specifically, the method of the present application is applied to any server, which communicates with a log monitoring component. Through this server, it obtains and analyzes the heterogeneous operation and maintenance time series data set provided by the log monitoring component, calibrates the data consistency of the heterogeneous operation and maintenance time series data set, determines the multi-source standardized operation and maintenance data stream. On this basis, it analyzes and obtains the operation and maintenance index dynamic causal information used to characterize the dynamic causal relationship of each operation and maintenance index, thereby determining the over-optimization target event information. And according to the over-optimization target event information, it forms a real-time operation and maintenance task chain required to solve the over-optimization event. Based on the principle of cooperation between large and small models, it determines the cooperation scheduling strategy of large and small models, thereby making real-time operation and maintenance adjustments to the system state, and providing the corresponding cooperative operation and maintenance report to the operation and maintenance personnel, effectively solving the over-optimization trap problem existing in the operation and maintenance process using AIOps technology, making the operation and maintenance process highly match the high-dynamic response requirements of the system, and significantly improving the response speed of the operation and maintenance process while ensuring the operation and maintenance effect, so as to ensure the system stability and user experience in real time.

[0070] Specific implementation manners can refer to the following embodiments.

[0071] Figure 2 The figure is a flowchart of an AIOps multi-task scheduling method driven by cooperation between large and small models provided by an embodiment of the present application. The method of this embodiment can be applied to the server in the above scenario. As Figure 2 shown, the method includes:

[0072] S201. Obtain a heterogeneous operation and maintenance time series data set, calibrate the data consistency of the heterogeneous operation and maintenance time series data set, and determine a multi-source standardized operation and maintenance data stream.

[0073] The heterogeneous operation and maintenance time series data set can be a set of operation and maintenance monitoring index data containing different sources and different formats, such as time series data of server CPU utilization rate, network throughput, database connection number, etc. The heterogeneous operation and maintenance time series data set can be obtained through the log monitoring component in the target operation and maintenance system.

[0074] The multi-source standardized operation and maintenance data stream can be a standardized operation and maintenance data stream that has been processed by time alignment, format unification, and numerical normalization.

[0075] Specifically, in modern distributed IT systems, there are problems such as sampling frequency differences, timestamp offsets, and inconsistent measurement units in the operation and maintenance data collected by various monitoring tools. Using automated mathematical analysis means, the spatio-temporal deviation between cross-source data is eliminated through data consistency calibration, and a standardized data stream with a unified time reference and numerical dimension is constructed to provide high-quality input for subsequent causal reasoning.

[0076] S202, analyzing multi-source standardized operation and maintenance data streams, determining dynamic causal information of operation and maintenance indicators, and determining over-optimization target event information based on the dynamic causal information of operation and maintenance indicators.

[0077] The dynamic causal information of operation and maintenance indicators can be information reflecting the time-varying causal relationship and intensity between indicators.

[0078] The over-optimization target event information may be currently existing operation and maintenance event information that has a negative impact on some operation and maintenance indicators due to over-optimization of a certain operation and maintenance indicator.

[0079] Specifically, in the process of using AIOps technology to automate the operation and maintenance of distributed systems, since the distributed system is composed of multiple microservices and these microservices are in a highly dynamic process, and the operation and maintenance goals set by the operation and maintenance team are often static, some work stages of some microservices do not match the operation and maintenance goals, resulting in over-optimization. That is, the current microservice status does not require high-level indicators in the operation and maintenance goals. Affected by the operation and maintenance goals, the automated operation and maintenance process excessively pursues the optimization of the current operation and maintenance indicators, resulting in uneven performance of some system components or the whole, loss of redundancy or reduced reliability, thereby causing failures, delays or reduced user experience. Such events are called "over-optimization events". The determination of over-optimization events cannot be concluded simply by measuring whether there are abnormalities in the operation and maintenance indicators. This is because when an over-optimization event occurs, the "over-optimization indicator" that triggers the over-optimization event is usually normal under the measurement system of the operation and maintenance target, while other abnormal operation and maintenance indicators are often not caused by abnormalities in the components to which they belong. It is necessary to clarify the causal relationship between different operation and maintenance indicators under the current system working state in order to quickly determine whether there is an over-optimization event. Through mathematical analysis, the causal relationship of each operation and maintenance indicator represented in the current multi-source standardized operation and maintenance data stream is automatically sorted out, and the dynamic causal information of the operation and maintenance indicators is integrated to determine whether there is an over-optimization event and determine the over-optimization target event information.

[0080] S203: Determine the real-time operation and maintenance task chain according to the optimized target event information, and determine the large and small model collaborative scheduling strategy according to the real-time operation and maintenance task chain.

[0081] The real-time operation and maintenance task chain can be an operation and maintenance operation sequence organized according to the causal path between operation and maintenance indicators.

[0082] The large and small model collaborative scheduling strategy can be a task allocation solution that combines the semantic understanding capability of the large language model with the real-time reasoning capability of the lightweight model.

[0083] Specifically, after determining the over-optimization target event information, since the over-optimization target event usually involves multiple operation and maintenance metrics, and the optimization and stability of different operation and maintenance metrics correspond to different operation and maintenance tasks, the solution of the over-optimization event requires the execution of multiple different operation and maintenance tasks. By extracting the operation and maintenance tasks to which several different operation and maintenance metrics corresponding in the over-optimization target event information, a real-time operation and maintenance task chain is constructed. Existing AIOps technologies usually uniformly execute different operation and maintenance tasks through an overall single model. Although this method reduces the overall complexity of the system, due to the lack of strong pertinence to different operation and maintenance tasks, the overall operation and maintenance efficiency is not high, and it is difficult to meet the system with high response requirements. Therefore, after obtaining the real-time operation and maintenance task chain, through mathematical analysis means, the complexity of different tasks in the real-time operation and maintenance task chain is quantified respectively. The operation and maintenance tasks with high complexity are assigned to the large model, and the operation and maintenance tasks with low complexity are assigned to the small model, so as to construct a collaborative scheduling strategy for the large and small models. Through the collaborative operation and maintenance scheduling of the large and small models, while ensuring the operation and maintenance effect, rapid operation and maintenance response is achieved.

[0084] S204. According to the collaborative scheduling strategy of the large and small models, perform real-time operation and maintenance adjustment on the system state, and determine and output a collaborative operation and maintenance report.

[0085] The collaborative operation and maintenance report can be report information including the change trends of various operation and maintenance metrics during the collaborative operation and maintenance of the large and small models.

[0086] Specifically, according to the operation and maintenance tasks responsible for by different models in the large and small model system scheduling strategy, through the collaborative execution of the large model and the small model for different operation and maintenance tasks, perform real-time operation and maintenance adjustment on the system state, record the change data of different operation and maintenance metrics during the execution of different operation and maintenance tasks, and through data visualization technology, perform visual integration processing on the change data of different operation and maintenance metrics, construct a collaborative operation and maintenance report, and through a human-computer interaction device, such as a high-definition display screen, provide the corresponding collaborative operation and maintenance report to the operation and maintenance personnel, so that the operation and maintenance personnel can effectively master the current operation and maintenance state.

[0087] Through this solution, data consistency calibration is performed on heterogeneous operation and maintenance time series datasets to determine multi-source standardized operation and maintenance data streams. On this basis, operation and maintenance index dynamic causal information used to characterize the dynamic causal relationships of various operation and maintenance indicators is analyzed, so as to determine over-optimization target event information. According to the over-optimization target event information, a real-time operation and maintenance task chain required to solve over-optimization events is formed. Based on the principle of collaboration between large and small models, a collaborative scheduling strategy for large and small models is determined, so as to perform real-time operation and maintenance adjustment on the system state and provide the corresponding collaborative operation and maintenance report to operation and maintenance personnel, effectively solving the over-optimization trap problem existing in the operation and maintenance process using AIOps technology, making the operation and maintenance process highly match the high-dynamic response requirements of the system, significantly improving the response speed of the operation and maintenance process while ensuring the operation and maintenance effect, so as to ensure system stability and user experience in real time.

[0088] In some embodiments, according to the heterogeneous operation and maintenance time series datasets, the operation and maintenance time series data corresponding to different types of operation and maintenance indicators are extracted; according to a number of operation and maintenance time series data, the service node to which each operation and maintenance time series data belongs is determined, and based on this, the network communication hop count between any two service nodes is determined; according to the network communication hop count, the topological weights of different data nodes in any two operation and maintenance time series data at the corresponding time nodes are determined; a number of operation and maintenance time series data are analyzed, the timestamps corresponding to different data nodes in each operation and maintenance time series data are extracted, and based on the timestamps, the time difference impact factor is determined; based on a number of operation and maintenance time series data, according to the topological weights and the time difference impact factor, the minimum alignment cost between any two operation and maintenance time series data is determined; according to the minimum alignment cost, unified alignment processing is performed on different data nodes in a number of operation and maintenance time series data to construct a multi-source standardized operation and maintenance data stream.

[0089] The operation and maintenance time series data can be the values corresponding to different operation and maintenance indicators at different time points.

[0090] The service node can be the microservice node to which the operation and maintenance indicator belongs.

[0091] The network communication hop count can be the number of intermediate nodes required for data transmission between two service nodes

[0092] The topological weight can be a quantitative indicator reflecting the impact of the communication distance between service nodes on data consistency.

[0093] The time difference impact factor can be a quantitative indicator used to characterize the impact of the timestamp difference of different operation and maintenance time series data on data alignment.

[0094] The minimum alignment cost can be the minimum adjustment cost required to achieve the time series alignment processing of different data nodes.

[0095] Specifically, in the operation and maintenance scenario of a distributed system, due to the network architecture differences and clock synchronization errors among different service nodes, there are two core problems in the heterogeneous operation and maintenance time series data set collected: one is the mismatch in the spatial dimension: traditional data calibration methods only focus on timestamp alignment but ignore the impact of network topology on data transmission delay. For example, for the metric data between service A and service B that span multiple levels of nodes, due to the long network transmission path, the collected time series data actually reflects the states of different service node positions, and direct alignment will introduce spatial deviation; the other is the distortion in the time dimension: existing timestamp alignment algorithms usually assume that the clocks of each node are completely synchronized, while in the actual system, there are clock offsets at the millisecond or even second level between nodes. If only aligned according to the original timestamps, it will misassociate operation and maintenance events with different time bases, resulting in the failure of subsequent causal reasoning. Therefore, this solution adopts a "network-time two-dimensional calibration mechanism" for the time series calibration of the heterogeneous operation and maintenance time series data set. Through mathematical analysis, based on the network communication hops between different service nodes and the timestamps of different data nodes reflected in the operation and maintenance time series data, the topological weight used to characterize the impact of the communication distance between service nodes on data consistency and the time difference impact factor used to characterize the impact of the timestamp differences of different operation and maintenance time series data on data alignment are respectively quantified. Furthermore, the minimum alignment cost required to achieve data time series alignment is comprehensively evaluated from the two dimensions of network structure and time, and according to the minimum alignment cost, different data nodes in several operation and maintenance time series data are uniformly aligned to construct a multi-source standardized operation and maintenance data stream, significantly improving the data alignment speed while ensuring the alignment accuracy.

[0096] Through this solution, starting from the two dimensions of network structure and time, based on the network communication hops between different service nodes and the timestamps of different data nodes reflected in the operation and maintenance time series data, the topological weight used to characterize the impact of the communication distance between service nodes on data consistency and the time difference impact factor used to characterize the impact of the timestamp differences of different operation and maintenance time series data on data alignment are respectively analyzed. Furthermore, the minimum alignment cost required to achieve data time series alignment is comprehensively evaluated from the two dimensions of network structure and time, and according to the minimum alignment cost, different data nodes in several operation and maintenance time series data are uniformly aligned to construct a multi-source standardized operation and maintenance data stream, significantly improving the data alignment speed while ensuring the alignment accuracy.

[0097] In some embodiments, based on several operation and maintenance time series data, the minimum alignment cost between any two operation and maintenance time series data is determined according to the topological weight and the time difference impact factor, specifically as the following formula (1):

[0098] (1)

[0099] Wherein, is the minimum alignment cost, is the topological weight, is the operation and maintenance time series data at the time node the observed value of the data node at this point, is the operation and maintenance time series data at the time node the observed value of the data node at this point, is the time difference impact factor.

[0100] Specifically, through the in formula (1) to describe the single-step alignment cost between two data nodes. This cost is determined by the topological weight and the distance between nodes. The distance measurement standard can use the Euclidean distance. At the same time, through the to describe the cumulative minimum alignment cost of two data nodes, and introduce the time difference impact factor in the process of measuring the minimum alignment cost to clarify the impact of timestamp differences, and then comprehensively quantify the minimum alignment cost between two data nodes.

[0101] Through this solution, using mathematical analysis methods, based on a number of operation and maintenance time series data, according to the topological weight and the time difference impact factor, quantify the minimum alignment cost between any two operation and maintenance time series data, reduce the interference of multi-hop node data deviation through the topological weight mechanism, so that the aligned data truly reflects the state association of adjacent service nodes, and through the time difference impact factor, effectively suppress the matching of pseudo-event sequences caused by clock asynchrony, and improve the accuracy of subsequent operation and maintenance event chain reconstruction.

[0102] In some embodiments, according to the number of network communication hops, determine the topological weight of different data nodes in any two operation and maintenance time series data at the corresponding time nodes. Specifically, it is the following formula (2):

[0103] (2)

[0104] Wherein, the topological weight, is the service node and the service node the number of network communication hops between them, is the preset attenuation coefficient; according to the timestamp, determine the time difference impact factor. Specifically, it is the following formula (3):

[0105] (3)

[0106] Wherein, is the time difference impact factor, is the operation and maintenance time series data at the time node the timestamp of the data node at this point, operation and maintenance time series data At the time node the timestamp of the data node, is the preset time offset threshold.

[0107] Specifically, the topological weight is used to reflect the communication influence strength between two service nodes in the network at a certain time node. Its quantization principle is based on the number of network communication hops. The larger the number of hops, the higher the delay and cost of data communication, and the smaller the weight. On the contrary, the smaller the number of hops, the more direct the connection between nodes, and the larger the weight. Through the exponential decay function in formula (3), it is realized that the larger the number of hops, the faster its weight needs to decrease, so as to show the weakening of the influence of distant nodes. At the same time, the attenuation degree needs to be adjustable to adapt to the objectively existing non-linear dependence relationship between different nodes in the actual network topology; the time difference influence factor is used to quantify the influence of the timestamp difference between two different data nodes at different time points on the matching cost. If the timestamps of two time-series data are exactly the same at two time nodes, the time difference influence factor should be 0 because they are synchronized in time. If the timestamps are inconsistent, that is, there is a difference, the greater the difference, the higher the time difference influence factor, thus imposing a greater penalty on the cost of the data alignment process. Through the measures the absolute time difference between two data nodes, and then through normalizes the time difference to quantitatively obtain the time difference influence factor.

[0108] Through this solution, by using mathematical analysis methods, according to the number of network communication hops and timestamps, the topological weight and the time difference influence factor are quantitatively obtained respectively, so as to clarify the mathematical characteristics of the topological weight and the time difference influence factor, improve the scientificity and accuracy of the topological weight and the time difference influence factor, and further improve the accuracy of the minimum alignment cost.

[0109] In some embodiments, analyze the multi-source standardized operation and maintenance data stream, determine the change amount of different operation and maintenance metrics at different time points, thereby determine the change rate corresponding to each operation and maintenance metric, and according to the change rate, determine the dynamic lag time window corresponding to each operation and maintenance metric; based on the change amount of each operation and maintenance metric at different time points and the dynamic lag time window, determine the lag impact factor between any two operation and maintenance metrics; according to the preset index hierarchy attribution information, analyze the multi-source standardized operation and maintenance data stream, determine the hierarchy correlation index between any two operation and maintenance metrics and the collaborative impact weight of each operation and maintenance metric; based on the hierarchy correlation index and the collaborative impact weight, according to the dynamic lag time window and the lag impact factor, determine the causal correlation index between any two operation and maintenance metrics; analyze the multi-source standardized operation and maintenance data stream, determine a number of obvious abnormal operation and maintenance metrics; based on the causal correlation index, analyze the number of obvious abnormal operation and maintenance metrics, determine a number of hidden root cause metrics, and according to the number of obvious abnormal operation and maintenance metrics and the corresponding number of hidden root cause metrics, determine whether there is an over-optimization event; if there is an over-optimization event, then use the corresponding number of obvious abnormal operation and maintenance metrics and hidden root cause metrics as the over-optimization target event information.

[0110] The dynamic lag time window can be a quantitative time range reflecting the impact of the change rate of the operation and maintenance metric on the timeliness of causal correlation. The time window corresponding to the operation and maintenance metric with a faster change rate is shorter.

[0111] The lag impact factor can be a quantitative metric used to characterize the cross-time dimension impact relationship between two operation and maintenance metrics.

[0112] The change amount can be the numerical difference amount of the operation and maintenance metric between the current time point and the previous sampling point.

[0113] The preset index hierarchy attribution information can be the preset information used to describe the hierarchical structure relationship between different operation and maintenance metrics.

[0114] The hierarchy correlation index can be the quantitative weight of the subordination relationship between operation and maintenance metrics determined based on the preset index hierarchy attribution information.

[0115] The collaborative impact weight can be a quantitative parameter used to characterize the influence of a single operation and maintenance metric in the collaborative operation and maintenance scenario.

[0116] The obvious abnormal operation and maintenance metric can be an operation and maintenance metric with an obviously abnormal value (such as CPU usage > 90%).

[0117] The hidden root cause metric can be a potential operation and maintenance metric that does not directly show abnormalities but has a significant impact on the obvious abnormal operation and maintenance metric through causal correlation.

[0118] Specifically, during the real-time operation and maintenance of a distributed system, since most operation and maintenance metrics are in a state of real-time change and there are differences in the change rates of different operation and maintenance metrics, in the process of analyzing the causal relationships of operation and maintenance metrics, if a fixed time window is adopted, the accuracy and efficiency of analyzing the causal relationships of operation and maintenance metrics with different change rates will be greatly reduced. It is necessary to determine the dynamic lag time window corresponding to each operation and maintenance metric according to the change rate corresponding to each operation and maintenance metric, following the principle that the faster the change rate, the shorter the time window. Then, based on the dynamic lag time window and the change amount corresponding to each operation and maintenance metric, through mathematical analysis means, analyze the cross-time dimension influence relationships between different operation and maintenance metrics, and quantitatively obtain the lag influence factor. At the same time, since the hierarchical relationships and collaborative influence weights defined during the system architecture design stage between different operation and maintenance metrics directly affect the degree of dynamic causal relationships shown by different operation and maintenance metrics during the system operation process, through mathematical analysis means, according to the dynamic lag time window and the lag influence factor, combined with the hierarchical correlation index and the collaborative influence weight, comprehensively analyze the causal relationships between different operation and maintenance metrics to quantitatively obtain the causal correlation index reflecting the causal correlation degree between different operation and maintenance metrics. On this basis, analyze the multi-source standardized operation and maintenance data stream, and use numerical verification methods such as threshold verification method to screen out a number of obvious abnormal operation and maintenance metrics. These obvious abnormal operation and maintenance metrics may be caused by "over-optimization events". Further, according to the causal correlation index, extract the root cause operation and maintenance metrics corresponding to the obvious abnormal operation and maintenance metrics. If there are root cause operation and maintenance metrics at a high level within their corresponding ranges among the extracted root cause operation and maintenance metrics, it indicates that there is an over-optimization event, and the corresponding obvious abnormal operation and maintenance metrics and hidden root cause metrics are used as over-optimization target event information.

[0119] Through this solution, analyze the dynamic causal relationships of different operation and maintenance metrics from multiple dimensions of time series change, hierarchical structure, and collaborative influence. By evaluating the causal correlation index between different operation and maintenance metrics and combining the current obvious abnormal operation and maintenance metrics, determine the root cause operation and maintenance metrics corresponding to the obvious abnormal operation and maintenance metrics, so as to judge whether there is an over-optimization event, and use the corresponding obvious abnormal operation and maintenance metrics and hidden root cause metrics as over-optimization target event information, making the judgment of over-optimization events more comprehensive and accurate, and at the same time providing a scientific data basis for subsequent operation and maintenance task analysis based on the optimization target event information.

[0120] In some embodiments, based on the hierarchical correlation index and the collaborative influence weight, according to the dynamic lag time window and the lag influence factor, determine the causal correlation index between any two operation and maintenance metrics, specifically as the following formula (4):

[0121] (4)

[0122] Wherein, is the current time point the following operation and maintenance metrics and the operation and maintenance metrics the causal association index between them is the preset data dimension conversion coefficient is the operation and maintenance metric and the operation and maintenance metric the hierarchical association index between them is the operation and maintenance metric the collaborative impact weight is the dynamic lag time window is the operation and maintenance metric at the time point the change amount is the current time point the following operation and maintenance metrics and the operation and maintenance metric the lag impact factor between them

[0123] The preset data dimension conversion coefficient can be a quantitative value used to uniformly scale the causal association impact of each parameter

[0124] Specifically, through in formula (4) to describe the collaborative impact of other operation and maintenance metrics on the two currently measured operation and maintenance metrics, through to describe the contribution of the fluctuations of different operation and maintenance metrics in the dynamic time lag window to the current collaborative impact. Through the collaborative impact, combined with the hierarchical association index used to describe the hierarchical impact and the lag impact factor to comprehensively quantify the causal association index

[0125] Through this solution, using mathematical analysis means, based on the hierarchical association index and the collaborative impact weight, according to the dynamic lag time window and the lag impact factor, comprehensively quantify the causal association index used to characterize the dynamic causal association strength between operation and maintenance metrics from three dimensions of hierarchical impact, collaborative impact, and lag impact, improve the comprehensiveness and accuracy of the causal association index, and further improve the judgment accuracy of over-optimization events

[0126] In some embodiments, based on the change amount of each operation and maintenance metric at different time points, determine the lag impact factor between any two operation and maintenance metrics, specifically as the following formula (5):

[0127] (5)

[0128] where is the current time point the following operation and maintenance metrics and the operation and maintenance metric The lag impact factor between is the preset maximum lag time, is the covariance calculation function, is the dynamic lag time window, is the operation and maintenance metric At the time point The change amount at is the operation and maintenance metric At the time point The change amount at

[0129] The preset maximum lag time is the maximum data delay time allowed for the service corresponding to the current operation and maintenance metric.

[0130] Specifically, during the operation and maintenance process, some operation and maintenance metrics are usually not completely independent. The change of the target metric may be caused by a certain fluctuation of the metric within the lag time point (such as a previous period of time). Therefore, the lag impact between different metrics is evaluated by measuring the contribution of the change at the historical time point to the current change. Through in formula (5), it describes the overall lag effect shown by the current operation and maintenance metric under the influence of different moments within the dynamic lag time window. The lag impact will cause large numerical fluctuations due to different selections of the dynamic lag time window. Therefore, in order to ensure the comparability of the results of the lag impact factor in different scenarios, it is necessary to normalize the lag time window. Through data normalization is achieved, and then the lag impact factor is accurately quantified.

[0131] Through this solution, by using mathematical analysis means, based on the change amount of each operation and maintenance metric at different time points, it describes the overall lag effect shown by the current operation and maintenance metric under the influence of different moments within the dynamic lag time window. Then, the lag time window is normalized to quantify the lag impact factor, improving the scientificity and accuracy of the lag impact factor.

[0132] In some embodiments, a number of explicit abnormal operation and maintenance metrics and implicit root cause metrics are analyzed. Taking each implicit root cause metric as the starting task node, a number of corresponding explicit abnormal operation and maintenance metrics are used as chain nodes to construct a number of operation and maintenance target data chains; based on the starting task node, the operation and maintenance target data chains are analyzed to determine the path depth of each chain node; according to the preset operation and maintenance knowledge graph, a number of operation and maintenance target data chains are analyzed to extract the operation and maintenance tasks corresponding to each starting task node or chain node and the dependency index between each node, construct a number of target task chains, and determine the event complexity of each target task chain and the task importance index of each target task in each target task chain; according to the path depth, dependency index, event complexity and task importance index, determine the dynamic scheduling mapping index of each target task in each target task chain; according to the dynamic scheduling mapping index of each target task, determine the large model target task set and the small model target task set, and construct a collaborative scheduling strategy for the large and small models.

[0133] The operation and maintenance target data chain can be a series of operation and maintenance metrics composed of implicit root cause metrics and explicit abnormal metrics, which are used to represent the current over-optimization problem.

[0134] The path depth can be the number of levels between the current task node and the root cause node.

[0135] The preset operation and maintenance knowledge graph can be the knowledge graph information used to represent the corresponding relationship between operation and maintenance metrics and operation and maintenance work. The operation and maintenance knowledge graph can be obtained through the analysis of historical operation and maintenance data and the evaluation of professional operation and maintenance personnel.

[0136] The operation and maintenance task can be the operation and maintenance work required to optimize the corresponding operation and maintenance metric.

[0137] The dependency index can be the quantitative data used to represent the mutual dependency relationship between operation and maintenance tasks.

[0138] The event complexity can be the quantitative index reflecting the overall resources required to process the current task chain. The event complexity can be obtained through the estimation of the execution time scale of different operation and maintenance tasks.

[0139] The task importance index can be the index used to represent the importance degree of the current operation and maintenance task in the overall operation and maintenance work. The task importance index can be obtained through the estimation of the influence range of the operation and maintenance task.

[0140] The dynamic scheduling mapping index can be the core decision parameter used to determine whether the task is assigned to the large model or the small model.

[0141] The large model target task set can be the set of operation and maintenance tasks responsible for execution by the large model.

[0142] The small model target task set can be a set of operation and maintenance tasks that the small model is responsible for executing.

[0143] Specifically, in the process of collaborative scheduling analysis of large and small models, the basic principle is that high-complexity operation and maintenance tasks are handled by the large model, and low-complexity operation and maintenance tasks are handled by the small model. This requires an accurate assessment of the complexity of operation and maintenance tasks. First, the operation and maintenance tasks closer to the root cause node (which can be regarded as the main node) have higher complexity because the scope of the impact of their operation and maintenance process spreads outward is larger. On the contrary, the operation and maintenance tasks farther from the root cause node have relatively lower complexity (which can be regarded as the branch nodes) because their operation and maintenance processes are relatively independent. In addition, the more the number of other operation and maintenance tasks associated with an operation and maintenance task, the higher the complexity of this operation and maintenance task because more other operation and maintenance indicators need to be concerned simultaneously when maintaining it. Therefore, it is necessary to comprehensively evaluate the complexity of operation and maintenance tasks from two dynamic dimensions of depth and breadth, and at the same time combine two static dimensions of the absolute complexity and importance of the operation and maintenance tasks themselves, so as to accurately measure the actual complexity of operation and maintenance tasks, and further provide accurate data basis for the collaborative scheduling analysis of large and small models. Through mathematical analysis means, according to the path depth, dependency index, event complexity, and task importance index, scientifically quantify the dynamic scheduling mapping index used to reflect the comprehensive complexity of operation and maintenance tasks. The larger the dynamic scheduling mapping index, the more the operation and maintenance task corresponding to this dynamic scheduling mapping index tends to use the large model. On the contrary, it means that the operation and maintenance task corresponding to this dynamic scheduling mapping index tends to use the small model. According to the comparison results between different dynamic scheduling mapping indexes and the mapping index threshold obtained by fitting historical data, determine the delivery model object of the corresponding operation and maintenance task, so as to separately integrate to obtain the model target task set and the small model target task set, and then construct the collaborative scheduling strategy of large and small models.

[0144] Through this solution, starting from two dynamic dimensions of depth and breadth, and at the same time combining two static dimensions of the absolute complexity and importance of the operation and maintenance tasks themselves, according to the path depth, dependency index, event complexity, and task importance index, comprehensively evaluate the complexity of operation and maintenance tasks, obtain the dynamic scheduling mapping indexes corresponding to different target tasks, so as to accurately measure the actual complexity of operation and maintenance tasks. On this basis, separately integrate to obtain the model target task set and the small model target task set, and then construct the collaborative scheduling strategy of large and small models, so that different target tasks are highly matched with their corresponding execution model scales.

[0145] In some embodiments, according to the path depth, dependency index, event complexity, and task importance index, determine the dynamic scheduling mapping index of each target task in each target task chain, specifically as the following formula (6):

[0146] (6)

[0147] Among them, is the dynamic scheduling mapping index of the th target task in the current target task chain, is the preset depth influence weight, is the path depth corresponding to the current target task, is the preset dependency influence weight, is the set of target tasks in the current target task chain that have a dependency relationship with the current target task, is the th target task in the set of target tasks, is the rd target task and the th target task, is the preset resource influence weight, is the event complexity of the target task chain corresponding to the current target task, is the task importance index corresponding to the current target task.

[0148] The preset depth influence weight can be a quantization weight used to represent the proportion of the influence of the path depth of the current target task on the dynamic scheduling mapping index.

[0149] The preset dependency influence weight is a quantization weight used to represent the proportion of the influence of the dependency relationship of the current target task on the dynamic scheduling mapping index.

[0150] The preset resource influence weight is a quantization weight used to standardize the proportion of the influence of the complexity and importance of the current target task itself on the dynamic scheduling mapping index.

[0151] The above influence weights can all be set based on the experience of professional operation and maintenance personnel, or can be obtained by fitting experimental data.

[0152] Specifically, through in formula (6) to describe the influence of the path depth of the target task on the dynamic scheduling mapping index, and using is to reflect the inverse relationship between the dynamic scheduling mapping index and the path depth. The smaller the path depth, the closer the current target task is to the root cause node, and the larger its corresponding dynamic scheduling mapping index; through to describe the restraint situation of the current target task by other operation and maintenance tasks that have a dependency relationship with it. The more restrained it is, the higher the overall complexity of the current target task, and the larger its corresponding dynamic scheduling mapping index; through to describe the linear relationship between the current target task and the event complexity and task importance index. The larger the event complexity and task importance index, the larger the corresponding dynamic scheduling mapping index; comprehensively considering the above influence relationships, the dynamic scheduling mapping index is quantitatively obtained.

[0153] Through this solution, by means of mathematical analysis, according to the path depth, dependence index, event complexity, and task importance index, the impact of each parameter on the complexity evaluation of the target task is clarified, and a dynamic scheduling mapping index reflecting the overall complexity of the target task is quantitatively obtained to accurately reflect the complexity of the target task, providing an accurate scientific data basis for the collaborative scheduling of large and small models.

[0154] Figure 3 It is a schematic structural diagram of an AIOps multi-task scheduling system driven by the collaboration of large and small models provided by an embodiment of the present application. As Figure 3 shown, an AIOps multi-task scheduling system 300 driven by the collaboration of large and small models in this embodiment includes: a data processing module 301, an over-optimization analysis module 302, a scheduling analysis module 303, and an operation and maintenance output module 304.

[0155] The data processing module 301 is used to obtain a heterogeneous operation and maintenance time series data set, perform data consistency calibration on the heterogeneous operation and maintenance time series data set, and determine a multi-source standardized operation and maintenance data stream;

[0156] The over-optimization analysis module 302 is used to analyze the multi-source standardized operation and maintenance data stream, determine the dynamic causal information of operation and maintenance indicators, and determine the over-optimization target event information according to the dynamic causal information of operation and maintenance indicators;

[0157] The scheduling analysis module 303 is used to determine a real-time operation and maintenance task chain according to the over-optimization target event information, and determine a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain;

[0158] The operation and maintenance output module 304 is used to perform real-time operation and maintenance adjustment on the system state according to the collaborative scheduling strategy for large and small models, and determine and output a collaborative operation and maintenance report.

[0159] Optionally, the data processing module 301 is specifically configured to: extract the operation and maintenance time series data corresponding to different types of operation and maintenance metrics according to the heterogeneous operation and maintenance time series data set; determine the service nodes to which each operation and maintenance time series data belongs according to a plurality of the operation and maintenance time series data, and thereby determine the network communication hop count between any two service nodes; determine the topological weights of different data nodes in any two operation and maintenance time series data at corresponding time nodes according to the network communication hop count; analyze a plurality of the operation and maintenance time series data, extract the timestamps corresponding to different data nodes in each operation and maintenance time series data, and determine the time difference impact factor according to the timestamps; based on a plurality of the operation and maintenance time series data, determine the minimum alignment cost between any two operation and maintenance time series data according to the topological weights and the time difference impact factor; and perform unified alignment processing on different data nodes in a plurality of the operation and maintenance time series data according to the minimum alignment cost to construct the multi-source standardized operation and maintenance data stream.

[0160] Optionally, when the data processing module 301 determines the minimum alignment cost between any two operation and maintenance time series data based on a plurality of the operation and maintenance time series data according to the topological weights and the time difference impact factor, it is specifically the following formula:

[0161] ;

[0162] where is the minimum alignment cost, is the topological weight, is the operation and maintenance time series data at the time node of the data node's observation value, is the operation and maintenance time series data at the time node of the data node's observation value, is the time difference impact factor.

[0163] Optionally, when the data processing module 301 determines the topological weights of different data nodes in any two operation and maintenance time series data at corresponding time nodes according to the network communication hop count, it is specifically the following formula:

[0164] ;

[0165] where is the topological weight, is the service node and the service node between the network communication hop count, is a preset attenuation coefficient; the determining the time difference impact factor according to the timestamps is specifically the following formula:

[0166] ;

[0167] wherein, is the time difference impact factor, is the operation and maintenance time series data at the data node at the time node the time stamp of the data node, the operation and maintenance time series data at the data node at the time node the time stamp of the data node, is the preset time offset threshold.

[0168] Optionally, the over-optimization analysis module 302 is specifically configured to: analyze the multi-source standardized operation and maintenance data stream, determine the change amount of different operation and maintenance metrics at different time points, thereby determining the change rate corresponding to each operation and maintenance metric, and determine the dynamic lag time window corresponding to each operation and maintenance metric according to the change rate; based on the change amount of each operation and maintenance metric at different time points and the dynamic lag time window, determine the lag impact factor between any two operation and maintenance metrics; according to the preset metric level attribution information, analyze the multi-source standardized operation and maintenance data stream, determine the level correlation index between any two operation and maintenance metrics and the co-influence weight of each operation and maintenance metric; based on the level correlation index and the co-influence weight, according to the dynamic lag time window and the lag impact factor, determine the causal correlation index between any two operation and maintenance metrics; analyze the multi-source standardized operation and maintenance data stream, determine several obvious abnormal operation and maintenance metrics; based on the causal correlation index, analyze several of the obvious abnormal operation and maintenance metrics, determine several hidden root cause metrics, and judge whether there is an over-optimization event according to several of the obvious abnormal operation and maintenance metrics and the corresponding several hidden root cause metrics; if there is the over-optimization event, then use the corresponding several obvious abnormal operation and maintenance metrics and the hidden root cause metrics as the over-optimization target event information.

[0169] Optionally, when the over-optimization analysis module 302 determines the causal correlation index between any two operation and maintenance metrics based on the level correlation index and the co-influence weight, according to the dynamic lag time window and the lag impact factor, it is specifically the following formula:

[0170] ;

[0171] wherein, is the current time point under the operation and maintenance metric and the operation and maintenance metric the causal correlation index between them, is the preset data dimension conversion coefficient, is the operation and maintenance metric and the operation and maintenance metric the hierarchical association index therebetween is the operation and maintenance metric the collaborative influence weight thereof is the dynamic lag time window is the operation and maintenance metric at the time point the change amount is the current time point under which the operation and maintenance metric and the operation and maintenance metric the lag influence factor therebetween

[0172] Optionally, when determining the lag influence factor between any two operation and maintenance metrics based on the change amounts of each operation and maintenance metric at different time points, the over-optimization analysis module 302 is specifically the following formula:

[0173] ;

[0174] wherein is the current time point under which the operation and maintenance metric and the operation and maintenance metric the lag influence factor therebetween is the preset maximum lag time is the covariance calculation function is the dynamic lag time window is the operation and maintenance metric at the time point the change amount is the operation and maintenance metric at the time point the change amount

[0175] Optionally, the scheduling analysis module 303 is specifically configured to: analyze a number of the explicit abnormal operation and maintenance metrics and the implicit root cause metrics, use each of the implicit root cause metrics as a starting task node, and use a number of the explicit abnormal operation and maintenance metrics corresponding thereto as chain nodes to construct a number of operation and maintenance target data chains; based on the starting task nodes, analyze the operation and maintenance target data chains to determine the path depth of each chain node; according to a preset operation and maintenance knowledge graph, analyze a number of the operation and maintenance target data chains, extract the operation and maintenance tasks corresponding to each starting task node or chain node and the dependency index between each node, construct a number of target task chains, and determine the event complexity of each target task chain and the task importance index of each target task in each target task chain; according to the path depth, the dependency index, the event complexity, and the task importance index, determine the dynamic scheduling mapping index of each target task in each target task chain; according to the dynamic scheduling mapping index of each target task, determine a large model target task set and a small model target task set, and construct the collaborative scheduling strategy for the large and small models.

[0176] Optionally, when the scheduling analysis module 303 determines the dynamic scheduling mapping index of each target task in each target task chain according to the path depth, the dependency index, the event complexity, and the task importance index, it is specifically the following formula:

[0177] ;

[0178] Wherein, is the dynamic scheduling mapping index of the th target task in the current target task chain, is a preset depth influence weight, is the path depth corresponding to the current target task, is a preset dependency influence weight, is the set of target tasks in the current target task chain that have a dependency relationship with the current target task, is the th target task in the target task set, is the th target task and the th target task, is a preset resource influence weight, is the event complexity of the target task chain corresponding to the current target task, is the task importance index corresponding to the current target task.

[0179] The system of this embodiment can be used to execute the method of any of the above embodiments. The implementation principles and technical effects are similar, and will not be elaborated here.

Claims

1. An AIOps multi-task scheduling method co-driven by a large model and a small model, characterized in that, Including: Obtain a heterogeneous operation and maintenance time series dataset, calibrate the data consistency of the heterogeneous operation and maintenance time series dataset, and determine a multi-source standardized operation and maintenance data stream; Analyze the multi-source standardized operation and maintenance data stream, determine the dynamic causal information of operation and maintenance metrics, and determine the over-optimization target event information according to the dynamic causal information of operation and maintenance metrics; The dynamic causal information of operation and maintenance metrics is information reflecting the time-varying causal relationship and intensity between operation and maintenance metrics in the multi-source standardized operation and maintenance data stream, that is, the causal association index between different operation and maintenance metrics; The analyzing the multi-source standardized operation and maintenance data stream, determining the dynamic causal information of operation and maintenance metrics, and determining the over-optimization target event information according to the dynamic causal information of operation and maintenance metrics includes: Analyze the multi-source standardized operation and maintenance data stream, determine the change amount of different operation and maintenance metrics at different time points, thereby determine the change rate corresponding to each operation and maintenance metric, and determine the dynamic lag time window corresponding to each operation and maintenance metric according to the change rate; Based on the change amount of each operation and maintenance metric at different time points and the dynamic lag time window, determine the lag impact factor between any two operation and maintenance metrics; According to the preset index hierarchy attribution information, analyze the multi-source standardized operation and maintenance data stream, and determine the hierarchy association index between any two operation and maintenance metrics and the collaborative impact weight of each operation and maintenance metric; Based on the hierarchy association index and the collaborative impact weight, determine the causal association index between any two operation and maintenance metrics according to the dynamic lag time window and the lag impact factor; Analyze the multi-source standardized operation and maintenance data stream to determine several obvious abnormal operation and maintenance metrics; Based on the causal association index, analyze several of the obvious abnormal operation and maintenance metrics to determine several implicit root cause metrics, and judge whether there is an over-optimization event according to several of the obvious abnormal operation and maintenance metrics and the corresponding several implicit root cause metrics; If there is the over-optimization event, then use the corresponding several obvious abnormal operation and maintenance metrics and the implicit root cause metrics as the over-optimization target event information; Determine a real-time operation and maintenance task chain according to the over-optimization target event information, and determine a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain; The determining a real-time operation and maintenance task chain according to the over-optimization target event information, and determining a collaborative scheduling strategy for large and small models according to the real-time operation and maintenance task chain includes: Analyze several of the obvious abnormal operation and maintenance metrics and the implicit root cause metrics, use each implicit root cause metric as the starting task node, and use the corresponding several obvious abnormal operation and maintenance metrics as chain nodes to construct several operation and maintenance target data chains; Based on the starting task node, analyze the operation and maintenance target data chain to determine the path depth of each chain node; According to the preset operation and maintenance knowledge graph, analyze several of the operation and maintenance target data chains, extract the operation and maintenance tasks corresponding to each starting task node or chain node and the dependency index between each node, construct several target task chains, and determine the event complexity of each target task chain and the task importance index of each target task in each target task chain; Determine the dynamic scheduling mapping index of each target task in each target task chain according to the path depth, the dependency index, the event complexity, and the task importance index; Determine the large model target task set and the small model target task set according to the dynamic scheduling mapping index of each target task, and construct the collaborative scheduling strategy for the large and small models; According to the collaborative scheduling strategy for the large and small models, perform real-time operation and maintenance adjustment on the system state, and determine and output a collaborative operation and maintenance report.

2. The method according to claim 1, wherein The data consistency calibration of the heterogeneous operation and maintenance time series data set to determine the multi-source standardized operation and maintenance data stream includes: Extract the operation and maintenance time series data corresponding to different types of operation and maintenance indicators according to the heterogeneous operation and maintenance time series data set; Determine the service node to which each operation and maintenance time series data belongs according to several pieces of the operation and maintenance time series data, and thereby determine the network communication hop count between any two service nodes; Determine the topological weight of different data nodes in any two operation and maintenance time series data at the corresponding time node according to the network communication hop count; Analyze several pieces of the operation and maintenance time series data, extract the time stamps corresponding to different data nodes in each operation and maintenance time series data, and determine the time difference impact factor according to the time stamps; Based on several pieces of the operation and maintenance time series data, determine the minimum alignment cost between any two operation and maintenance time series data according to the topological weight and the time difference impact factor; Perform unified alignment processing on different data nodes in several pieces of the operation and maintenance time series data according to the minimum alignment cost, and construct the multi-source standardized operation and maintenance data stream.

3. The method according to claim 2, characterized in that, The determination of the minimum alignment cost between any two operation and maintenance time series data based on several pieces of the operation and maintenance time series data according to the topological weight and the time difference impact factor is specifically the following formula: ; wherein, is the minimum alignment cost, is the topological weight, is the operation and maintenance time series data at the time node the observed value of the data node at, is the operation and maintenance time series data at the time node the observed value of the data node at, is the time difference impact factor.

4. The method according to claim 2, characterized in that, The determination of the topological weight of different data nodes in any two operation and maintenance time series data at the corresponding time node according to the network communication hop count is specifically the following formula: ; Among them, the topological weight, is the number of network communication hops between the service node and the service node, and is a preset attenuation coefficient; The determination of the time difference impact factor according to the time stamp is specifically the following formula: ; wherein, is the time difference impact factor, is the operation and maintenance time series data at the time node the time stamp of the data node at that point, the operation and maintenance time series data at the time node the time stamp of the data node at that point, is the preset time offset threshold.

5. The method according to claim 1, characterized in that Based on the hierarchical association index and the collaborative influence weight, determine the causal association index between any two operation and maintenance indicators according to the dynamic lag time window and the lag influence factor, specifically the following formula: ; Among them, is the current time point the following operation and maintenance metric and the operation and maintenance metric the causal correlation index therebetween is the preset data dimension conversion coefficient is the operation and maintenance metric and the operation and maintenance metric the hierarchical correlation index therebetween is the operation and maintenance metric the collaborative influence weight thereof is the dynamic lag time window is the operation and maintenance metric at the time point the change amount thereat is the current time point the following operation and maintenance metric and the operation and maintenance metric the lag influence factor therebetween 6. The method according to claim 5, wherein Based on the change amount of each operation and maintenance indicator at different time points, determine the lag influence factor between any two operation and maintenance indicators, specifically the following formula: ; Wherein, is the current time point of the operation and maintenance metric and the operation and maintenance metric of the lag impact factor therebetween, is the preset maximum lag time, is the covariance calculation function, is the dynamic lag time window, is the operation and maintenance metric at the time point of the change amount, is the operation and maintenance metric at the time point of the change amount.

7. The method according to claim 6, wherein The determination of the dynamic scheduling mapping index of each target task in each target task chain according to the path depth, the dependency index, the event complexity, and the task importance index is specifically the following formula: ; Among them, is the dynamic scheduling mapping index of the th target task in the current target task chain, is the preset depth influence weight, is the path depth corresponding to the current target task, is the preset dependency influence weight, is the set of target tasks in the current target task chain that have a dependency relationship with the current target task, is the th target task in the set of target tasks, is the th target task and the th target task of the is the preset resource influence weight, is the event complexity of the target task chain corresponding to the current target task, is the task importance index corresponding to the current target task.

8. An AIOps multi-task scheduling system co-driven by a large model and a small model, characterized in that, Including: A data processing module, configured to obtain a heterogeneous operation and maintenance time series data set, perform data consistency calibration on the heterogeneous operation and maintenance time series data set, and determine a multi-source standardized operation and maintenance data stream; An over-optimization analysis module is used to analyze the multi-source standardized operation and maintenance data stream, determine the dynamic causal information of operation and maintenance metrics, and determine the over-optimization target event information based on the dynamic causal information of operation and maintenance metrics; the dynamic causal information of operation and maintenance metrics is information reflecting the time-varying causal relationship and intensity among various operation and maintenance metrics in the multi-source standardized operation and maintenance data stream, that is, the causal association index between different operation and maintenance metrics. The over-optimization analysis module is specifically used for: Analyze the multi-source standardized operation and maintenance data stream, determine the change amount of different operation and maintenance metrics at different time points, thereby determine the change rate corresponding to each operation and maintenance metric, and determine the dynamic lag time window corresponding to each operation and maintenance metric according to the change rate. Based on the change amount of each operation and maintenance metric at different time points and the dynamic lag time window, determine the lag impact factor between any two operation and maintenance metrics. According to the preset index level attribution information, analyze the multi-source standardized operation and maintenance data stream, and determine the level association index between any two operation and maintenance metrics and the collaborative impact weight of each operation and maintenance metric. Based on the level association index and the collaborative impact weight, determine the causal association index between any two operation and maintenance metrics according to the dynamic lag time window and the lag impact factor. Analyze the multi-source standardized operation and maintenance data stream to determine several obvious abnormal operation and maintenance metrics. Based on the causal association index, analyze several of the obvious abnormal operation and maintenance metrics, determine several hidden root cause metrics, and judge whether there is an over-optimization event according to the several obvious abnormal operation and maintenance metrics and the corresponding several hidden root cause metrics. If there is the over-optimization event, then use the corresponding several obvious abnormal operation and maintenance metrics and the hidden root cause metrics as the over-optimization target event information. A scheduling analysis module is used to determine the real-time operation and maintenance task chain according to the over-optimization target event information, and determine the collaborative scheduling strategy of the large and small models according to the real-time operation and maintenance task chain. The scheduling analysis module is specifically used for: Analyze several of the obvious abnormal operation and maintenance metrics and the hidden root cause metrics, take each hidden root cause metric as the starting task node, and use the corresponding several obvious abnormal operation and maintenance metrics as chain nodes to construct several operation and maintenance target data chains. Based on the starting task node, analyze the operation and maintenance target data chain to determine the path depth of each chain node. According to the preset operation and maintenance knowledge graph, analyze several of the operation and maintenance target data chains, extract the operation and maintenance tasks corresponding to each starting task node or chain node and the dependency index between each node, construct several target task chains, and determine the event complexity of each target task chain and the task importance index of each target task in each target task chain. According to the path depth, the dependency index, the event complexity, and the task importance index, determine the dynamic scheduling mapping index of each target task in each target task chain. According to the dynamic scheduling mapping index of each target task, determine the large model target task set and the small model target task set, and construct the collaborative scheduling strategy of the large and small models. An operation and maintenance output module, which is used to perform real-time operation and maintenance adjustment on the system status according to the collaborative scheduling strategy of the large and small models, and determine and output a collaborative operation and maintenance report.

Citation Information

Patent Citations

  • Industrial fault diagnosis model generation method based on combination of large and small models

    CN118468120A

  • Internet of vehicles platform operation and maintenance method, system and device based on multi-modal large model

    CN119449576A