A service system running state monitoring and early warning method and system

By using artificial intelligence to monitor and warn about the operational status of business systems, the problem of relying on manual monitoring in traditional data center monitoring has been solved. This enables real-time and accurate identification and warning of abnormal states, reduces labor costs, and improves the stability and efficiency of data center operations.

CN114168409BActive Publication Date: 2025-10-28SHANGHAI MOONPAC INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111424577.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-10-28
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

Traditional data center monitoring relies on manual on-duty personnel, which has problems such as high professional requirements, high labor costs, low response efficiency and poor real-time performance, making it difficult to achieve efficient operation and maintenance management.

Method used

An AI-based business system operation status monitoring and early warning method is adopted. By collecting, analyzing and preprocessing data, a combined model is generated for real-time monitoring and early warning. This includes data collection, quality exploration, feature analysis, similarity measurement, pattern mining and association rule generation, to achieve real-time monitoring and early warning of the data center.

Benefits of technology

It enables 24/7 uninterrupted identification and early warning of abnormal states, reduces labor costs, improves the accuracy and timeliness of fault finding, and enhances the stability and efficiency of data center operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168409B_ABST
    Figure CN114168409B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for monitoring and early warning of the operational status of a business system. The method includes: collecting operational status information of the business system to be monitored and storing it as a first operational status dataset; performing quality exploration and feature analysis on the first operational status dataset to obtain a second operational status dataset; preprocessing the second operational status dataset to obtain a third operational status dataset; performing similarity measurement on the third operational status dataset to obtain the similarity distance of the operational status data; performing pattern mining processing on the third operational status dataset to obtain data association rules; generating a first combined model based on the similarity distance and data association rules; and predicting data change trends based on the first combined model for real-time monitoring and early warning. This invention can reflect the operational indicators of various businesses in the data center in real time and intuitively, ensuring system availability, preventing emergencies, and improving system operational efficiency and safe operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method and system for monitoring and early warning of the operational status of a business system. Background Technology

[0002] As the central location for storing and processing business data, the secure operation of data centers, including their power, environment, security, network, and equipment, is crucial to the overall system's operation. Therefore, data center security monitoring, as a vital component of operation, plays a significant role in ensuring the secure operation of company business and improving daily operational efficiency.

[0003] Traditional data center maintenance relies primarily on on-duty staff, but manual monitoring has significant limitations. First, data center maintenance demands a high level of expertise from management personnel, and manual operations have blind spots, failing to guarantee immediate detection of equipment malfunctions. Second, data center equipment operates 24 / 7, requiring substantial manpower for monitoring, and achieving continuous manual monitoring is challenging. Finally, traditional manual monitoring suffers from poor real-time alarm response and low efficiency, hindering efficient data center operation and maintenance. The economic losses from system or equipment failures are incalculable; therefore, real-time intelligent monitoring and management of data centers are crucial.

[0004] With the development of technologies such as big data and artificial intelligence, building highly intelligent "unmanned" data centers will be a new development in data center environmental control systems. Summary of the Invention

[0005] In view of the deficiencies in the existing technology, the present invention provides a method and system for monitoring and early warning of the operation status of a business system, which can reflect the various business operation indicators of the data center in real time and intuitively, ensure the availability of the system, prevent the occurrence of emergencies, and improve the system's operating efficiency and safe operation.

[0006] In a first aspect, the present invention provides a method for monitoring and early warning of the operating status of a business system, comprising the following steps:

[0007] Step S101: Collect the operating status information of the business system to be tested and store it as the first operating status dataset;

[0008] Step S103: Perform quality exploration and feature analysis on the first running state dataset to obtain the second running state dataset;

[0009] Step S105: Preprocess the second running state dataset to obtain the third running state dataset;

[0010] Step S107: Perform a similarity measurement on the third running state dataset to obtain the similarity distance of the running state data;

[0011] Step S109: Perform pattern mining processing on the third running state dataset to obtain data association rules;

[0012] Step S111: Generate a first combined model based on the similarity distance and data association rules;

[0013] Step S113: Based on the first combined model, predict the data change trend and perform real-time monitoring and early warning.

[0014] The quality exploration and feature analysis in step S103 include at least data missingness analysis, outlier analysis, overall data distribution, statistical analysis, and correlation analysis.

[0015] The preprocessing in step S105 includes: data cleaning, data standardization, data reduction, and data discretization.

[0016] The data standardization includes:

[0017]

[0018] Among them, X j The values ​​of the data objects in the original dataset. X is the standardized value of this data object. max X is the maximum value in this data object. min It is the minimum value in this data object.

[0019] Specifically, the data reduction includes:

[0020]

[0021] in, It refers to the value of a data object after data reduction, ultimately transforming the original dataset of length n into a dataset of length N by segmenting and aggregating the segments.

[0022] The data discretization includes:

[0023]

[0024] in, It is the value of the data object after data discretization. Finally, based on its data change trend, the time series dataset is transformed into a string set {u,l,d}, where t is the data fluctuation threshold.

[0025] Specifically, step S107 includes:

[0026] Based on the third operating status dataset and for different indicator data of operating conditions, a template subsequence is obtained with a period of hour, day, week, and month, respectively. The similarity distance between the current time subsequence of the same indicator data and the template subsequence is calculated to predict data changes in the future.

[0027] Specifically, step S109 includes:

[0028] Step S1091: Based on the third running status dataset, mine the frequent patterns of each indicator data in the business system;

[0029] Step S1093: Based on the frequent patterns of each indicator data, mine the frequent patterns between different indicator data in the business system.

[0030] Step S1095: Generate association rules for the third running status dataset based on the frequent patterns between the different indicator data.

[0031] Specifically, step S111 includes:

[0032] The first combined model is generated by weighting the prediction results based on similarity distance and association rules.

[0033] Secondly, the present invention also provides a business system operation status monitoring and early warning system, comprising:

[0034] The data acquisition module collects operational status information of the business system under test and stores it as the first operational status dataset.

[0035] The data exploration module performs quality exploration and feature analysis on the first running state dataset to obtain the second running state dataset.

[0036] The data preprocessing module preprocesses the second running state dataset to obtain the third running state dataset;

[0037] The sequence matching module performs similarity measurement on the third running state dataset and obtains the similarity distance between the running state data.

[0038] The pattern mining module performs pattern mining on the third running state dataset to obtain data association rules.

[0039] The model optimization module generates the first combined model based on similarity distance and data association rules;

[0040] The monitoring and early warning module predicts data change trends based on the first combined model and performs real-time monitoring and early warning.

[0041] This invention proposes a business system operation status monitoring and early warning method and system based on artificial intelligence, which can identify and warn of abnormal states 24 / 7, reduce the workload of data center operators, and greatly improve the accuracy and timeliness of fault finding, the stability and real-time performance of data center operation and maintenance, thereby further improving the efficiency of the company's internal system operation. Attached Figure Description

[0042] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the present disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0043] Figure 1 This is a flowchart illustrating a business system operation status monitoring and early warning method according to an embodiment of the present invention;

[0044] Figure 2 This is a schematic diagram illustrating the preprocessing according to an embodiment of the present invention;

[0045] Figure 3 This is a flowchart illustrating the pattern mining process according to an embodiment of the present invention;

[0046] Figure 4 This is a schematic diagram illustrating a business system operation status monitoring and early warning system according to an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.

[0048] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "an," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.

[0049] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0050] The optional embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0051] Example 1

[0052] like Figure 1 As shown, this invention discloses a method for monitoring and early warning of the operational status of a business system, comprising the following steps:

[0053] Step S101: Collect the operating status information of the business system to be tested and store it as the first operating status dataset;

[0054] Step S103: Perform quality exploration and feature analysis on the first running state dataset to obtain the second running state dataset;

[0055] Step S105: Preprocess the second running state dataset to obtain the third running state dataset;

[0056] Step S107: Perform a similarity measurement on the third running state dataset to obtain the similarity distance of the running state data;

[0057] Step S109: Perform pattern mining processing on the third running state dataset to obtain data association rules;

[0058] Step S111: Generate the first combined model based on similarity distance and data association rules;

[0059] Step S113: Based on the first combined model, predict the trend of data change and perform real-time monitoring and early warning.

[0060] Example 2

[0061] like Figure 1 As shown, this invention proposes a method for monitoring and early warning of the operational status of a business system, comprising the following steps:

[0062] Step S101: Collect the operating status information of the business system to be tested and store it as the first operating status dataset; wherein, the operating status information includes the monitoring index data information of the operating status of each machine in this business system, specifically including CPU utilization, memory utilization, disk utilization, CPU idle rate, buffer usage, total file system utilization, disk IO wait time, response time, average network outflow / inflow per second, etc.

[0063] Step S103: Perform quality exploration and feature analysis on the first running status dataset to obtain the second running status dataset. This step involves data exploration of the collected business system running information data set to understand the general quality, size, features, sample quantity, data type, and probability distribution of the dataset, so as to better understand its data characteristics.

[0064] Step S105: Preprocess the second running state dataset to obtain the third running state dataset;

[0065] Step S107: Perform similarity measurement on the third operating status dataset to obtain the similarity distance of the operating status data. This step involves sequence matching of relevant information within a single monitoring indicator of the monitored system. On the third operating status dataset, for different indicator data, from a time perspective, a template subsequence is obtained with a period of hour, day, week, and month, respectively, and the similarity distance between the current subsequence and the template subsequence is calculated.

[0066] Step S109: Perform pattern mining on the third operating status dataset to obtain data association rules; this step of pattern mining is the process of searching for hidden information in a large amount of equipment operating data through algorithms.

[0067] Step S111: Generate the first combined model based on similarity distance and data association rules;

[0068] Step S113: Based on the first combined model, predict the trend of data change and perform real-time monitoring and early warning.

[0069] To facilitate understanding of the quality exploration and feature analysis in step S103, this embodiment of the invention provides a further description. In this embodiment, the quality exploration aims to check the correctness and validity of the data. Its main task is to check whether there is dirty data in the original data. Dirty data generally refers to data that does not meet the requirements or cannot be directly analyzed.

[0070] In practical application scenarios, the quality exploration and feature analysis mentioned in step S103 of this embodiment include at least data missingness, outlier analysis, overall data distribution, statistical analysis, and correlation analysis.

[0071] To further understand data missingness, outlier analysis, overall data distribution, statistical analysis, and correlation analysis, these are described in more detail below.

[0072] Data missing generally refers to missing observations and missing values ​​of variables within those observations.

[0073] Outliers refer to data entry errors or the presence of unreasonable data. These are usually isolated values ​​in the sample that deviate significantly from the other observations. In most cases, variables are not allowed to have negative values; that is, negative values ​​are outliers.

[0074] For outlier analysis, certain business background knowledge can be used to identify anomalies from the range of variable values, thereby determining whether there are data processing errors. The 3-sigma principle (if the data follows a normal distribution, under the 3σ principle, outliers are defined as values ​​in a set of measurements that deviate from the mean by more than three standard deviations) and IQR (outliers are typically defined as values ​​less than Q) can also be used. L -l.5IQR or greater than Q u +1.5IQR, Q L This is called the lower quartile, Q. u The upper quartile is called the interquartile range, and the IQR is called the interquartile range, which is the Q value. u Upper quartiles and Q L The difference in the lower quartiles includes half of all observations. Statistical methods such as these are used to analyze whether the data are outliers.

[0075] The overall distribution of data can be observed through plotting methods such as histograms and frequency graphs to reveal the shape and trend of the data distribution.

[0076] Statistical analysis is a descriptive statistical method for analyzing data, including the range of variable values ​​and the degree of deviation.

[0077] The most commonly used statistics are the maximum and minimum values, which are used to determine whether the values ​​of a variable exceed a reasonable range. Other statistics that can be calculated include the mean, mode, quantiles, standard deviation, and variance.

[0078] Correlation analysis is used to analyze the strength of the linear correlation between two continuous variables, and it can be performed by calculating the Pearson correlation coefficient. The formula is as follows:

[0079]

[0080] In the formula, E() is the expectation, μ X Let E(X) represent the expectation of X, μ YLet E(Y) represent the expected value of Y, and ρ represent the Pearson correlation coefficient between two continuous variables (X,Y). X,Y It equals their covariance cov(X,Y) divided by the product of their respective standard deviations (σ). X ,σ Y ).

[0081] Example 3:

[0082] Based on the above embodiments, this embodiment may further include the following:

[0083] like Figure 2 As shown, the preprocessing of the second running state dataset in step S105 may include: data cleaning, data standardization, data reduction, and data discretization.

[0084] To help those skilled in the art better understand the preprocessing procedure in this embodiment, it is further described below.

[0085] Data standardization in this embodiment may include:

[0086]

[0087] where X j The values ​​of the data objects in the original dataset. X is the standardized value of this data object. max X is the maximum value in this data object. min It is the minimum value.

[0088] Data reduction processing is applied to standardized data, which may specifically include:

[0089]

[0090] in It refers to the value of a data object after data reduction, ultimately transforming the original dataset of length n into a dataset of length N by segmenting and aggregating the segments.

[0091] Discretizing the reduced data can specifically include:

[0092]

[0093] in, It is the value of the data object after data discretization. Finally, based on its data change trend, the time series dataset is transformed into a string set {u,l,d}, where t is the data fluctuation threshold.

[0094] This embodiment employs methods such as filling in missing values, replacing invalid data, and removing noise from the data for data cleaning. To clarify the difference between data cleaning and quality exploration and feature analysis in this embodiment, the differences are further described. The data quality exploration and analysis in this embodiment primarily focuses on discovering dirty data. Specifically, its main task is to check whether dirty data exists in the original data. Dirty data generally refers to data that does not meet requirements and cannot be directly analyzed. Data cleaning, on the other hand, involves correcting or discarding this dirty data.

[0095] Example 4

[0096] Based on the above embodiments, this embodiment may further include the following:

[0097] This embodiment performs a similarity measurement on the preprocessed third running state dataset. Specifically, step S107 may include:

[0098] Based on the third operational status dataset and for different indicator data of operational information, a template subsequence is obtained with a period of hour, day, week, and month respectively. The similarity distance between the current time subsequence of the same indicator data and the template subsequence is calculated to predict the data changes in the future.

[0099] To facilitate understanding of step S107 in this embodiment, some of its contents will be described in detail. Specifically, the metrics data for the operational status of the business system in this embodiment include CPU utilization, memory utilization, disk utilization, access volume, etc.

[0100] In this embodiment, when obtaining a template subsequence, a random selection is made based on the template sequence. A subsequence with a period of hour, day, week, or month is selected and used as the template subsequence.

[0101] Example 5

[0102] Based on the above embodiments, this embodiment may further include the following:

[0103] In addition to performing similarity measurement on the third running state dataset, this embodiment also requires pattern mining processing on the third running state dataset. Specifically, such as... Figure 3 As shown, step S109 in this embodiment may include:

[0104] Step S1091: Based on the third running status dataset, mine the frequent patterns of each indicator data in the business system;

[0105] Step S1093: Based on the frequent patterns of each indicator data, mine the frequent patterns between different indicator data in the business system.

[0106] Step S1095: Generate association rules for the third running status dataset based on the frequent patterns between different indicator data.

[0107] Algorithms for pattern mining mainly include Apriori and FP-Trees. In this implementation, the Apriori algorithm is preferred for pattern mining. Specifically, step S1091 can include the following methods:

[0108] First, set the maximum length of the pattern to k, and the minimum support threshold for frequent patterns to min.

[0109] The Apriori association analysis algorithm is used to mine frequent patterns in the preprocessed third-running-state dataset. First, based on frequent patterns of size 1 (consisting of symbols u, l, d in the string set {u, l, d}), pairwise concatenation yields frequent patterns of length 2. A complete scan of the discretized dataset is then performed to obtain the frequency of each frequent pattern, called its support count, and a list of its occurrence positions in the dataset is recorded, including the initial and final positions of the pattern. Frequent patterns are then filtered, retaining those with a support value greater than the minimum support threshold (min), resulting in a frequent pattern set of length 2. Repeating this process yields frequent pattern sets of length 3, their corresponding position lists, and so on, until a pattern set of length k is obtained, along with its corresponding position list.

[0110] Regarding step S1093, this embodiment utilizes the Apriori association analysis algorithm to perform pattern mining between different indicator data of the business system operation. The specific method may include:

[0111] First, set the maximum length of the pattern to k2, the minimum support threshold for frequent patterns to min2, and the maximum support threshold for frequent patterns to max2.

[0112] Randomly select a business system running a certain indicator data, starting with frequent pattern 1 within the sequence. Connect each frequent pattern within the sequence to other frequent patterns in the sequence, generating a candidate frequent pattern between sequences of length 2. Generate a position list of candidate frequent patterns between sequences of length 2 from the position list of candidate frequent patterns within the sequence, and determine whether the length of this position list is between the minimum support count min2 and the maximum support count max2. If it is, add the candidate frequent pattern between sequences of length 2 to the set of frequent patterns between sequences; otherwise, delete the candidate frequent pattern between sequences of length 2. Continue in this manner, using two frequent patterns of length k2-1 to generate candidate frequent patterns between sequences of length k2. Generate a new position list from the position lists of the two candidate frequent patterns within the sequence, and determine whether the length of the new position list is between the minimum support count min2 and the maximum support count max2. If it is, add the candidate frequent pattern between sequences of length k2 to the set of frequent patterns between sequences; otherwise, delete the candidate frequent pattern between sequences of length k2, until the length of the frequent patterns in the set of frequent patterns between sequences reaches the set maximum length k2.

[0113] For step S1095, the specific method may include:

[0114] First, set the minimum confidence threshold of the rules to min3. Generate association rules for the business system operation status dataset based on the frequent patterns between different index data. Then, filter the generated association rules according to the set minimum confidence threshold min3 to obtain the final association rules.

[0115] Example 6

[0116] Based on the above embodiments, this embodiment may further include the following:

[0117] In this embodiment, step S111, generating the first combined model based on similarity distance and data association rules, may specifically include:

[0118] The prediction results based on similarity distance and association rules are weighted to generate the first combined model;

[0119] Based on the experimental results, the parameters of the first combined model were adjusted and optimized, and the weights of the similarity distance and data association rule output results were adjusted to optimize the first combined model.

[0120] After obtaining the optimized first combined model, this embodiment uses step S113 for real-time monitoring and early warning, which may specifically include:

[0121] The similarity distance between the template subsequence and the current time series is calculated. Using real sequence data as experimental data, the trend of future time series changes is predicted. If the trend exceeds a certain threshold, it is judged to be abnormal.

[0122] Based on the correlation rules between the mined data, we can obtain the changing patterns of one data point from another. When the actual data trend does not match the predicted data changing patterns, we can determine that an anomaly has occurred.

[0123] The results of single sequences and data sequences are combined to comprehensively assess the likelihood of anomalies, thereby enabling monitoring and early warning.

[0124] This embodiment can also display various indicators of the business system's operating status data on a single graph after data analysis and processing, enabling comprehensive dynamic monitoring and facilitating managers' understanding of the data center equipment's operating status.

[0125] Example 7

[0126] like Figure 4 As shown in the figure, this embodiment of the invention also proposes a business system operation status monitoring and early warning system, which includes:

[0127] The data acquisition module collects operational status information of the business system under test and stores it as the first operational status dataset.

[0128] The data exploration module performs quality exploration and feature analysis on the first running state dataset to obtain the second running state dataset.

[0129] The data preprocessing module preprocesses the second running state dataset to obtain the third running state dataset;

[0130] The sequence matching module performs similarity measurement on the third running state dataset and obtains the similarity distance between the running state data.

[0131] The pattern mining module performs pattern mining on the third running state dataset to obtain data association rules.

[0132] The model optimization module generates the first combined model based on similarity distance and data association rules;

[0133] The monitoring and early warning module predicts data change trends based on the first combined model and performs real-time monitoring and early warning.

[0134] Example 8

[0135] This disclosure provides a non-volatile computer storage medium storing computer-executable instructions that can perform the method steps described in the above embodiments.

[0136] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0137] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0138] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (AN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0140] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0141] The preferred embodiments of the present invention have been described above to make the spirit of the present invention clearer and easier to understand, and are not intended to limit the present invention. All modifications, substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope summarized by the appended claims.

Claims

1. A method for monitoring and early warning of the operational status of a business system, characterized in that, Includes the following steps: Step S101: Collect the operating status information of the business system to be tested and store it as the first operating status dataset; Step S103: Perform quality exploration and feature analysis on the first running status dataset to obtain a second running status dataset; the quality exploration and feature analysis includes at least data missingness analysis, outlier analysis, overall data distribution, statistical analysis, and correlation analysis; Step S105: Preprocess the second running state dataset to obtain the third running state dataset; Step S107: Perform similarity measurement on the third operating status dataset and obtain the similarity distance of the operating status data, including: for different indicator data of operating condition information, obtain template subsequences with hours, days, weeks and months as periods respectively, and calculate the similarity distance between the current time subsequence of the same indicator data and the template subsequence; Step S109: Perform pattern mining processing on the third running state dataset to obtain data association rules; Step S111: Based on the similarity distance and data association rules, generate a first combined model, specifically including: generating a first combined model by weighting the prediction results of similarity distance and association rules; adjusting and optimizing the parameters of the first combined model according to experimental results, adjusting the weights of the output results of similarity distance and data association rules, and optimizing the first combined model. Step S113: Based on the first combined model, predict the data change trend and perform real-time monitoring and early warning. Specifically, this includes: calculating the similarity distance between the template subsequence and the current time series, using real sequence data as experimental data, predicting the change trend of data at a set future time, and judging that an anomaly has occurred if the change trend exceeds a certain threshold; based on the association rules between the mined data, the change pattern of one data can be obtained from one data, and when the actual data trend and the predicted data change pattern do not match, it is judged that an anomaly has occurred; combining the anomaly results between the real sequence data of a single monitoring indicator and the real sequence data of multiple monitoring indicators, comprehensively judging the probability of an anomaly, and thus performing monitoring and early warning.

2. The method as described in claim 1, characterized in that, The preprocessing in step S105 includes: data cleaning, data standardization, data reduction, and data discretization.

3. The method as described in claim 2, characterized in that, The data standardization includes: Among them, X j The values ​​of the data objects in the original dataset. X is the standardized value of this data object. max X is the maximum value in this data object. min It is the minimum value in this data object.

4. The method as described in claim 3, characterized in that, The data reduction specifically includes: in, It refers to the value of a data object after data reduction, ultimately transforming the original dataset of length n into a dataset of length N by segmenting and aggregating the segments.

5. The method as described in claim 4, characterized in that, The data discretization includes: in, It is the value of the data object after data discretization. Finally, based on its data change trend, the time series dataset is transformed into a string set {u,l,d}, where t is the data fluctuation threshold.

6. The method as described in claim 5, characterized in that, Step S109 specifically includes: Step S1091: Based on the third running status dataset, mine the frequent patterns of each indicator data in the business system; Step S1093: Based on the frequent patterns of each indicator data, mine the frequent patterns between different indicator data in the business system. Step S1095: Generate association rules for the third running status dataset based on the frequent patterns between the different indicator data.

7. A business system operation status monitoring and early warning system, characterized in that, include: The data acquisition module collects operational status information of the business system under test and stores it as the first operational status dataset. The data exploration module performs quality exploration and feature analysis on the first running state dataset to obtain the second running state dataset; the quality exploration and feature analysis includes at least data missingness analysis, outlier analysis, overall data distribution, statistical analysis, and correlation analysis. The data preprocessing module preprocesses the second running state dataset to obtain the third running state dataset; The sequence matching module performs similarity measurement on the third operating status dataset and obtains the similarity distance of the operating status data. This includes: obtaining template subsequences for different indicator data of operating condition information with periods of hours, days, weeks, and months, respectively, and calculating the similarity distance between the current time subsequence of the same indicator data and the template subsequence. The pattern mining module performs pattern mining on the third running state dataset to obtain data association rules. The model optimization module generates a first combined model based on similarity distance and data association rules. Specifically, it includes: generating a first combined model by weighting the prediction results based on similarity distance and association rules; adjusting and optimizing the parameters of the first combined model based on experimental results, adjusting the weights of the output results of similarity distance and data association rules, and optimizing the first combined model. The monitoring and early warning module predicts data change trends based on the first combined model and performs real-time monitoring and early warning. Specifically, it includes: calculating the similarity distance between the template subsequence and the current time series, using real sequence data as experimental data, predicting the change trend of data at a set future time; if the change trend exceeds a certain threshold, it is judged to be abnormal; based on the association rules between the mined data, it can obtain the change pattern of one data from another data; when the actual data trend does not match the predicted data change pattern, it is judged to be abnormal; combining the abnormal results of the real sequence data of a single monitoring indicator and the real sequence data of multiple monitoring indicators, comprehensively judging the probability of abnormality, and thus performing monitoring and early warning.