Data availability observation and alarm method and device, equipment and medium
By using intelligent data availability observation and alerting methods, combined with automated indicator analysis and deep learning technology, the problems of false alarms, missed alarms and insufficient multi-dimensional observation in traditional observation methods have been solved, achieving efficient and accurate data observation and alerting.
Patent Information
- Application Number
- CN202510989271.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional Argus system instance observation and alarm methods suffer from false alarms, missed alarms, insufficient multi-dimensional observation capabilities, and a lack of automated response mechanisms, resulting in low data observation efficiency and difficulty in adapting to dynamic business scenarios.
By integrating automated indicator analysis, time feature analysis, dynamic strategy optimization, feature dimensionality reduction and weighted regression, time series alignment and deep learning technologies, intelligent data availability observation and alerting are achieved, generating target strategies that dynamically adapt to different business scenarios, and enabling real-time observation and automated classification of anomaly levels.
It significantly improves the efficiency and accuracy of data observation, reduces false alarm and false negative rates, enhances the intelligence level and automated response capability of computer systems, and provides efficient and intelligent alarms that adapt to complex environments.
Smart Images

Figure CN120811933A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data analysis, and in particular to a data availability observation and alarm method, device, equipment and medium. BACKGROUND
[0002] In the Argus-based China system instance availability observation and alarm scenario, the current traditional observation method has many disadvantages. The alarm rules in the traditional observation tool are not intelligent enough, the fixed threshold is difficult to adapt to the dynamic business scenario, and false positives and false negatives are prone to occur; the multi-dimensional observation capability is lacking, it is difficult to consider the network, storage, application and other aspects at the same time; and there is a lack of automatic response mechanism, so after fault detection, it can only be handled manually, which is low in efficiency.
[0003] For example, in the medical health scenario, a hospital based on Argus builds a China system instance observation system, and in the traditional way, the registration system is observed, only the server CPU and memory are concerned, the alarm threshold is fixed, and when the registration peak period comes, the CPU usage rate exceeds the threshold but the business is normal, but a false alarm is triggered, and when the network is congested, the registration page loads slowly, and because the network index is not observed, a false negative occurs, and the network, storage (registration record storage) and application (registration process logic) states of the registration system cannot be monitored at the same time, so it is difficult to fully understand the complex operation of the medical system, and the patient registration is affected.
[0004] For example, in the financial technology business scenario, a securities trading platform based on Argus builds an observation system, and in the traditional observation, only the CPU and memory usage of the transaction server are observed, and a fixed alarm threshold is set. The fixed threshold cannot adapt to the characteristics of the financial technology business that changes in an instant, and false positives and false negatives affect business decisions; and during the transaction peak, the CPU temporarily exceeds the threshold but the business is normal, triggering a false alarm, and when the transaction system fails, such as database connection interruption, only manual troubleshooting and service restart can be relied on, and the efficiency of manual fault handling is low, which easily causes business interruption and brings economic losses and reputation risks to financial institutions.
[0005] Therefore, how to improve the data observation efficiency and real-time alarm has become a problem to be solved. SUMMARY
[0006] The present application provides a data availability observation and alarm method, device, equipment and medium, which mainly aims to solve the problem of low data observation efficiency.
[0007] In the first aspect, to achieve the above-mentioned purpose, the present application provides a data availability observation and alarm method, comprising: Obtaining observation data of a target observation object, performing index analysis on the observation data to obtain observation index data; perform time feature analysis on the observation index data to obtain observation data features; perform strategy optimization on a predefined initial observation alarm strategy according to the observation data features to obtain a target observation alarm strategy; perform feature anomaly detection on the observation data features to obtain target abnormal data; determine an abnormal level corresponding to the target abnormal data based on a time sequence fluctuation risk score of the target abnormal data; push the target abnormal data to a preset target observation object end according to the abnormal level and the target observation alarm strategy, and perform real-time observation on the pushed target abnormal data to obtain target observation data.
[0008] In a second aspect, the present application further provides a data availability observation and alarm device, comprising: an index analysis module configured to obtain observation data of a target observation object, perform index analysis on the observation data, and obtain observation index data; a feature extraction module configured to perform time feature analysis on the observation index data to obtain observation data features; a strategy optimization module configured to perform strategy optimization on a predefined initial observation alarm strategy according to the observation data features to obtain a target observation alarm strategy; an anomaly detection module configured to perform feature anomaly detection on the observation data features to obtain target abnormal data; a level determination module configured to determine an abnormal level corresponding to the target abnormal data based on a time sequence fluctuation risk score of the target abnormal data; a real-time observation module configured to push the target abnormal data to a preset target observation object end according to the abnormal level and the target observation alarm strategy, and perform real-time observation on the pushed target abnormal data to obtain target observation data.
[0009] In a third aspect, the present application further provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data availability observation and alarm method described above.
[0010] In a fourth aspect, the present application also provides a computer readable storage medium, wherein at least one computer program is stored in the computer readable storage medium, and the at least one computer program is executed by a processor in an electronic device to implement the data availability observation and alarm method.
[0011] In the embodiment of the present application, the efficiency and reliability of the computer system in processing observation data are significantly improved through the automatic index analysis process, the regular expression analysis calculation expression realizes the flexible expansion of index definition, without modifying the underlying code to adapt to new business requirements, avoiding the risk of manual SQL writing errors, effectively solving the problems of poor data quality, rigid analysis and delayed response in traditional observation systems; the time alignment technology adopts high-precision timestamp standardization and dynamic resampling algorithm to ensure that the time reference error of multi-source heterogeneous data is small and reduced, solving the analysis deviation problem caused by different clock synchronization in traditional observation; through the dynamic strategy optimization based on the characteristics of observation data, the intelligent level of the computer observation system is significantly improved, and the generated target strategy can dynamically adapt to different business scenarios, providing an efficient and accurate intelligent alarm solution for complex environments such as cloud computing and big data; through the collaborative optimization of feature dimension reduction and weighted regression, the detection efficiency of abnormal data in computer observation is significantly improved, and the feature dimension reduction technology compresses high-dimensional observation data into a low-dimensional space to reduce the computational overhead; through the fusion of time alignment and deep learning technology, the accuracy and automation level of abnormal level determination in observation are significantly improved, realizing the automatic classification (low / medium / high) of abnormal level, reducing manual intervention, and improving the efficiency of data observation and alarm; through the collaborative mechanism of probe collection, multi-dimensional threshold marking and intelligent merging, the efficiency and accuracy of real-time observation are significantly improved, and the lightweight data collection probe realizes millisecond-level index grabbing, covering system, application and network data in all dimensions to ensure that no abnormal data is missed; the dynamic threshold and multi-index correlation marking technology improves the data observation efficiency and accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0013] Figure 1 An application environment schematic diagram of a data availability observation and alarm method in an embodiment of the present application; Figure 2 A flowchart of a data availability observation and alarm method provided by an embodiment of the present application; Figure 3 A flowchart for determining an abnormality level corresponding to target abnormality data according to the target abnormality data is provided for an embodiment of the present application. Figure 4 A module diagram of a data availability observation and alarm device is provided for an embodiment of the present application. Figure 5 A structural diagram of an electronic device for implementing a data availability observation and alarm method is provided for an embodiment of the present application. Figure 6 Another structural diagram of an electronic device for implementing a data availability observation and alarm method is provided for an embodiment of the present application.
[0014] The purposes, functional features and advantages of the present application will be further described with reference to the accompanying drawings in conjunction with embodiments. DETAILED DESCRIPTION
[0015] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to achieve the implementation process of the corresponding technical effects by applying technical means to solve technical problems, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all embodiments. The embodiments of the present disclosure and various features in the embodiments can be combined with each other without conflict, and the technical solutions formed thereby are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should be within the protection scope of the present disclosure.
[0016] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or equipment including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0017] The embodiment of the present application provides a data availability observation and alarm method. The execution subject of the data availability observation and alarm method includes but is not limited to at least one of electronic devices such as a server, a terminal and the like which can be configured to execute the device provided by the embodiment of the present application. In other words, the data availability observation and alarm method can be executed by software or hardware installed in a terminal device or a server device. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster and the like. The server can be a stand-alone server, or a cloud server providing cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platform.
[0018] The data availability observation and alarm method can be applied to, for example, Figure 1The application environment is shown in the figure. The client communicates with the server through the network. The server can obtain observation data through the client, and the automatic index analysis process significantly improves the efficiency and reliability of the computer system in processing observation data. The regular expression analysis calculation expression realizes the flexible expansion of index definition, without modifying the underlying code to adapt to new business needs, avoiding the risk of manual SQL writing errors, and effectively solving the problems of poor data quality, rigid analysis, and delayed response in traditional observation systems. Through multi-dimensional time feature analysis, the computer system's ability to observe data is significantly improved. The time alignment technology uses high-precision timestamp standardization and dynamic resampling algorithm to ensure that the time reference error of multi-source heterogeneous data is small, and solves the analysis deviation problem caused by different clock synchronization in traditional observation. Through dynamic strategy optimization based on observation data characteristics, the intelligent level of the computer observation system is significantly improved. The generated target strategy can dynamically adapt to different business scenarios, providing an efficient and accurate intelligent alarm solution for complex environments such as cloud computing and big data. Through the collaborative optimization of feature dimension reduction and weighted regression, the detection efficiency of abnormal data in computer observation is significantly improved. The feature dimension reduction technology compresses high-dimensional observation data into low-dimensional space, reducing the computational overhead. Through the integration of time series alignment and deep learning technology, the accuracy and automation level of abnormal level determination in observation are significantly improved, realizing automatic classification of abnormal levels (low / medium / high), reducing manual intervention, and improving data observation and alarm efficiency. Through the collaborative mechanism of probe collection, multi-dimensional threshold marking, and intelligent merging, the efficiency and accuracy of real-time observation are significantly improved. The lightweight data collection probe realizes millisecond-level index grabbing, covering system, application, and network data in all dimensions, ensuring that no abnormal data is missed. The dynamic threshold and multi-index correlation marking technology improves data observation efficiency and accuracy. Finally, the target observation data is output and fed back to the client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The application will be described in detail through specific embodiments.
[0019] Referring to Figure 2 Fig. 1 is a flowchart of a data availability observation and alarm method according to an embodiment of the application. In this embodiment, the data availability observation and alarm method comprises: S1, obtaining observation data of a target observation object, and performing index analysis on the observation data to obtain observation index data.
[0020] In the embodiment of the application, the observation data refers to a set of key indicators reflecting the running state and performance of network devices, links, applications, etc., and the observation index data includes but is not limited to network interface traffic, error rate, packet loss rate, device CPU / memory utilization, link bandwidth utilization, QoS indicators (delay, jitter), BGP routing information, and security events (DDoS attack, abnormal traffic), etc., which are organized by a standardized model (such as YANG) and encoded and transmitted in a structured format (GPB, JSON, XML) to provide real-time state perception and fault location basis for network operation and maintenance.
[0021] In detail, local client collection can be performed, including local data collection of server hardware state, process resource occupation, log information, etc. Real-time indicators of CPU, memory, disk, network, etc. can be obtained by reading the / proc file system, or the survival state and resource consumption (CPU, memory, FD usage) of the process can be observed. The data is derived from the statistical files under the / proc / ${pid} / directory.
[0022] Specifically, whether the target port is open can be detected through a communication protocol such as TCP / UDP. When the target port is open, an HTTP / HTTPS request is sent and the returned content is verified (such as requiring a status code of 200 and a response body containing a specific string).
[0023] Exemplarily, in the hospital observation and patient data security guarantee in the medical health scene, a certain first-class hospital deploys an electronic medical record system (EMR), a remote diagnosis and treatment platform, a medical image storage (PACS), and Internet of Things devices (such as intelligent infusion pumps and vital sign monitors). In order to ensure stable operation of the system and safety of patient data, network performance, device state, and security events need to be observed in real time.
[0024] Specifically, an observation client can be deployed on the EMR and PACS servers to collect CPU, memory, and disk I / O usage every 5 seconds, observe the number of database connections and query response time, and parse server logs to extract key events (such as database login failure and abnormal access request) and identify potential security threats (such as SQL injection attempts) in combination with regular expressions.
[0025] In the embodiment of the application, in combination with the locally collected server logs and Telemetry traffic data, abnormal behaviors are identified. For example, if a certain PACS server has a large number of external IP accesses during non-working hours, and at the same time, the traffic of the network segment where the server is located is detected to surge, it may be a data leakage attack, so the server is immediately isolated and the security team is notified, and the security event detection accuracy is improved.
[0026] Exemplarily, in the field of financial technology, in the observation and anti-fraud of the bank core transaction system, a commercial bank deploys a core transaction system (processing deposit, transfer, payment and other businesses), an online banking platform and a risk control system, in order to meet the regulatory requirements (such as the second version of the network security protection) and protect the user experience, it is necessary to observe the transaction link performance, system resources and fraud behavior in real time.
[0027] In detail, the observation client is deployed on the core transaction server, and the transaction response time, TPS (transactions per second) and database lock waiting time are collected every second, if the TPS drops by 50% or the response time exceeds 2 seconds, the expansion process (such as starting the standby node) is automatically triggered, and the backlog message number and consumer delay of the message queue (such as Kafka) are observed, if the backlog of a certain queue exceeds 100,000 and lasts for 5 minutes, the number of consumer threads is adjusted to avoid transaction blocking.
[0028] In the embodiment of the application, the observation data is subjected to index analysis to obtain observation index data, comprising: The observation data is subjected to data cleaning to obtain cleaned observation data; The dimension attribute and the calculation expression corresponding to the cleaned observation data are obtained, the calculation expression is parsed according to a pre-defined regular expression specification to obtain a corresponding index name and a calculation function; The dimension attribute, the index name and the calculation function are spliced to obtain a SQL expression; The observation data is subjected to index calculation according to the SQL expression to obtain observation index data.
[0029] In detail, the data cleaning is the basis for ensuring the accuracy of the analysis result, mainly solving the problems of missing values, abnormal values, repeated values and inconsistent formats in the original data; if the key field (such as device IP, timestamp) of a certain observation record is missing, the record is directly discarded (applicable to the scene with a missing rate of less than 5%); for the missing values of non-key fields, forward filling (using the same field value of the previous record to supplement) or mean filling (such as using the average value of the historical traffic of the device to supplement the current missing value) is adopted.
[0030] Among them, the reasonable range of the index can be calculated based on the historical data (such as the CPU utilization rate being normally in the range of 0%-100%), and the values exceeding the range are marked as abnormal and corrected to the boundary value (such as 100% or 0%), providing high-quality input for subsequent analysis.
[0031] Specifically, extract the dimensions (such as device type, region) and calculation logic (such as sum, average) required for analysis from the cleaning observation data, and parse complex expressions through regular expressions; define the correspondence between data fields and dimensions (such as field device_type mapped to dimension device type) in advance, automatically extract dimensions by querying the mapping table, and use NLP models (such as BERT) to extract key entities (such as database in database connection failure) for unstructured fields (such as error description in logs).
[0032] Among them, define rule matching common calculation function (such as SUM$(.*)$ matching sum expression, AVG$(.*)$ matching average expression), extract the field in the parentheses as the calculation object (such as SUM(interface_traffic) in "interface_traffic" as the calculation field), for complex expression (such as (SUM(in_bytes) + SUM(out_bytes)) / 1024), build abstract syntax tree (AST) to decompose operation sequence and function call, identify index name (such as "total traffic") and calculation function (such as sum divided by 1024").
[0033] In detail, convert the parsed dimensions, indicators and functions into executable SQL queries, which need to handle field aliases, aggregation functions and condition filtering, use templates (such as SELECT {dimensions}, {function}({metric}) AS{alias} FROM {table} WHERE {conditions}) to dynamically fill parameters; for example, the dimension is "device type", the indicator is "interface traffic", and the calculation function is "sum", generate SQL.
[0034] Among them, execute SQL through database engine or distributed computing framework to get observation index data, use MySQL, PostgreSQL and other relational databases in single machine scene to directly execute SQL, if for large-scale data, use SparkSQL or Presto distributed computing, convert SQL to MapReduce task parallel processing (such as calculating the average traffic of all network devices), convert the calculated observation index data to unit, convert byte to MB / GB, or convert millisecond to second, etc.
[0035] In the embodiment of the present application, the efficiency and reliability of the computer system in processing observation data are significantly improved through an automatic index analysis process; in the data cleaning stage, the rule engine and machine learning model are used to efficiently filter missing values and abnormal values, reduce subsequent calculation noise, and realize flexible expansion of index definition through regular expression analysis calculation expression, without modifying the underlying code to adapt to new business needs, avoiding the risk of errors caused by manual SQL writing, and effectively solving the problems of poor data quality, rigid analysis, and delayed response in traditional observation systems.
[0036] S2, time characteristic analysis is performed on the observation index data to obtain observation data characteristics.
[0037] In the embodiment of the present application, the observation index data is time-aligned, segmented by a sliding window, and the time domain and frequency domain features are extracted and fused, and finally the comprehensive observation data characteristics are obtained.
[0038] In the embodiment of the present application, the time characteristic analysis on the observation index data to obtain observation data characteristics comprises: The timestamp of the observation index data is obtained, and the observation index data is time-aligned according to the timestamp to obtain standard index data; A time window is obtained, and the standard index data is segmented by sliding according to the time window to obtain a plurality of time segment index data; The time domain characteristics of a plurality of time segment index data are calculated; The frequency domain transformation is performed on a plurality of time segment index data, and the frequency domain characteristics of the time segment index data after frequency domain transformation are extracted; The time domain characteristics and the frequency domain characteristics are fused to obtain observation data characteristics.
[0039] The time domain characteristics include mean, variance, maximum and minimum, and the frequency domain characteristics include main frequency component and energy distribution.
[0040] In detail, the time alignment is a key step to ensure that observation data from different sources or with different sampling frequencies can be analyzed on the same time reference, mainly solving the problems of inconsistent timestamp format and non-uniform sampling interval, and can be realized through timestamp standardization, that is, the timestamp in the original data may exist in the form of string, Unix timestamp or local time zone time, which is recognized and converted into uniform UTC time format through regular expression, avoiding time zone confusion, and if the timestamp of a record is missing, it is inferred according to the timestamp of adjacent records (such as linear interpolation or directly copying the timestamp of the previous record).
[0041] Specifically, if the sampling frequencies of different indicators are different (such as CPU utilization once per second and network traffic once every 5 seconds), they can be unified to the same frequency (such as once every 5 seconds) through aggregation (such as taking the average of high-frequency data) or interpolation (such as using linear interpolation to complete low-frequency data). After alignment, the timestamp error of the data is reduced, the sampling frequency is unified, and a consistent time base is provided for subsequent analysis.
[0042] Furthermore, the sliding time window is used to divide continuous data into multiple segments, which is convenient for analyzing behavioral patterns within a local time range (such as burst traffic, periodic fluctuations). The time window can be set to a fixed window length (such as a window every 10 minutes), which is suitable for analyzing periodic characteristics (such as daily traffic peaks), or set to a sliding window, where the window slides at a fixed step size (such as every 1 minute) to generate overlapping segments (such as a window length of 10 minutes and a step size of 1 minute), which is suitable for detecting rapidly changing events (such as network attacks).
[0043] Among them, the window length can be adaptively adjusted according to the data characteristics (for example, shorten the window to capture details when traffic bursts, and extend the window to reduce the amount of calculation when traffic is stable).
[0044] Specifically, the time domain features directly reflect the statistical characteristics of the data in the time dimension, and are used to describe trends, volatility, and extreme values, calculate the mean and median of data within a segment, reflect typical values (such as average CPU utilization), calculate the standard deviation and range (maximum value - minimum value), measure data volatility (such as the jitter degree of network delay), calculate skewness (asymmetry of data distribution) and kurtosis (the degree of data sharpness), and identify abnormal distributions (for example, a bimodal distribution may indicate a mixed load).
[0045] Specifically, frequency domain analysis converts time domain signals into frequency domain representation, revealing the periodic components and frequency distribution characteristics in the data. It can decompose time domain data into a combination of sine waves of different frequencies, or perform FFT on the data in a sliding window to generate a time-frequency spectrum to calculate the energy (squared amplitude) of each frequency component, identify the dominant frequency (such as the main period of traffic fluctuations), and calculate the energy proportion of each frequency band to distinguish different types of behaviors (such as low frequency may correspond to business load, high frequency may correspond to noise or attack).
[0046] Further, the feature fusion combines the complementarity of time domain and frequency domain information, improves the richness and discriminability of feature expression, can directly splice the time domain feature vector (such as [mean, standard deviation, skewness]) and the frequency domain feature vector (such as [dominant frequency, low frequency energy proportion]) into a longer vector (such as [mean, standard deviation, skewness, dominant frequency, low frequency energy proportion]), and is suitable for linear models (such as logistic regression); the feature importance can also be used to assign weights (such as using information gain or mutual information to calculate the weight), and the time domain and frequency domain features are weighted and summed to highlight key features (such as giving higher weight to attack features in the frequency domain).
[0047] In the embodiment of the application, the multi-dimensional time feature analysis significantly improves the computer system's ability to analyze observation data, the time alignment technology uses high-precision timestamp standardization and dynamic resampling algorithm to ensure that the time reference error of multi-source heterogeneous data is small and reduced, and solves the analysis deviation problem caused by different clock synchronization in traditional observation; the sliding time window combined with the adaptive segmentation strategy can flexibly capture various time patterns from second-level burst to hour-level cycle, improve event detection coverage, and reduce computing resource consumption, thereby providing an efficient and accurate real-time feature extraction framework for intelligent observation alarm.
[0048] S3, performing strategy optimization on the initial observation alarm strategy according to the observation data features to obtain a target observation alarm strategy.
[0049] In the embodiment of the application, the matching degree analysis, strategy loss calculation and iterative optimization are used, that is, the strategy matching degree and loss value are calculated, inefficient strategies are screened and cross-optimized, and the optimal observation alarm strategy is output after multiple iterations.
[0050] In the embodiment of the application, the strategy optimization on the initial observation alarm strategy according to the observation data features to obtain a target observation alarm strategy comprises: Obtaining a plurality of initial observation alarm strategies, analyzing the plurality of initial observation alarm strategies to obtain a plurality of alarm targets; Performing matching degree analysis on the plurality of alarm targets and the observation data features to obtain data alarm matching degrees; Constructing a target loss function according to the alarm targets, the data alarm matching degrees and the observation data features, and determining a strategy loss value of each initial observation alarm strategy according to the target loss function; Screening a first observation alarm strategy greater than a preset strategy loss threshold according to the strategy loss value; Performing cross processing on the first observation alarm strategy to obtain a second observation alarm strategy; Returning the second observation alarm strategy to the step of determining the strategy loss value of each of the initial observation alarm strategies according to the target loss function, and counting the number of returns; When the number of returns reaches a preset number, the return is stopped and the final second observation alarm strategy is used as the target observation alarm strategy.
[0051] In detail, the initial observation alarm strategy usually exists in the form of a configuration file, database record or API parameter, and may include threshold rules (such as CPU utilization > 90% triggering an alarm), combination conditions (such as a sudden increase in traffic and an increase in error rate) or timing patterns (such as delay exceeding the limit is detected three times in a row).
[0052] Among them, the initial observation alarm strategies from different sources (such as XML configuration, JSON API, text rules) are uniformly converted into internal intermediate representations (such as key-value pairs or tree structures) to facilitate subsequent processing; when the alarm target is a single conditional target, the observation indicators (such as cpu_usage), comparison operators (such as >) and thresholds (such as 90) in the strategy are directly extracted as targets. If it is a combined conditional target, the logically combined strategy is decomposed into multiple sub-targets and the logical relationships are recorded.
[0053] Specifically, matching analysis is used to quantify the degree of fit between the alarm target and the actual observed data characteristics, solving the problem that traditional fixed threshold alarms cannot adapt to dynamic network environments; for each alarm target, the characteristic value of the corresponding indicator is extracted from the observed data characteristics, and the degree of fit between the characteristic value and the threshold is calculated.
[0054] For example, if the average CPU utilization is 85% and the threshold is 90%, the match can be defined as 1 - (85 / 90) (the smaller the value, the closer it is to the threshold, the higher the match). To detect trends (such as a continuous increase in traffic), use linear regression to calculate the slope of the feature sequence. If the slope is positive and significant (for example, p-value < 0.05), the match is high.
[0055] Furthermore, the objective loss function is used to quantify the optimization direction of the strategy and balance key indicators such as matching degree, false alarm rate and missed alarm rate. The loss function includes a matching degree penalty: the lower the matching degree, the higher the loss; a false alarm penalty: if the strategy has triggered too many false alarms in the past (through historical log statistics), the loss will be increased; a missed alarm penalty: if the strategy does not cover known problem scenarios (such as historical missed alarm events), the loss will be increased; and each penalty item is weighted and summed.
[0056] Specifically, high-loss strategies are screened based on the summed loss values, that is, the mean and standard deviation of the loss values of all strategies are calculated to screen out abnormally high-loss strategies. The screened strategies are marked as requiring optimization, and their original loss values and alarm targets are recorded for reference in subsequent cross-processing.
[0057] Specifically, the crossover process refers to generating a new strategy by combining the advantages of different strategies, similar to the crossover operation in the genetic algorithm, but optimized for the particularity of the alarm strategy; taking the average or random value of the same indicator target (such as the CPU threshold) of the two strategies.
[0058] For example, the CPU threshold for policy A is 90, and that for policy B is 85. After crossing, the value can be 87.5 or a randomly selected value between 85 and 90.
[0059] Among them, each crossover generates multiple candidate strategies (such as 3-5) to avoid falling into the local optimum. The iterative optimization gradually approaches the optimal strategy through multiple crossovers and loss evaluations, which is similar to the iterative process of gradient descent or genetic algorithms in machine learning. The number of iterations is preset (such as 10 times) and stops after it is reached. The current optimal strategy (lowest loss) is recorded in each iteration, and the historical optimal result is finally selected instead of the last iteration result to avoid overfitting.
[0060] In the embodiment of the present invention, the intelligence level of the computer observation system is significantly improved through dynamic strategy optimization based on the characteristics of observation data. Data-driven matching analysis is used to replace the traditional fixed threshold, so that the alarm strategy can adapt to network fluctuations and reduce the false alarm rate. By constructing a multi-dimensional target loss function, the accuracy, stability and complexity of the strategy are comprehensively evaluated to improve the coverage of the optimized strategy. The target strategy finally generated can dynamically adapt to different business scenarios, providing an efficient and accurate intelligent alarm solution for complex environments such as cloud computing and big data.
[0061] S4. Perform feature anomaly detection on the observed data features to obtain target anomaly data.
[0062] In the embodiment of the present invention, weighted regression analysis is performed on the observed features after dimensionality reduction, the anomaly probability is calculated, and feature vectors below a threshold are screened, thereby identifying target abnormal data.
[0063] In the embodiment of the present invention, the step of performing feature anomaly detection on the observed data features to obtain target anomaly data includes: Performing feature dimensionality reduction on the observation data features, and constructing a dimensionality reduction feature matrix based on the feature dimensionality reduced observation data features; Calculating the allocation weight of each reduced-dimensionality feature vector in the reduced-dimensionality feature matrix; Performing regression calculation on the reduced-dimensionality feature vector according to the assigned weight to obtain an abnormal probability of the reduced-dimensionality feature vector; The reduced-dimensional feature vector whose abnormal probability is greater than a preset abnormal threshold is used as target abnormal data.
[0064] In detail, the observation data features usually contain a large number of redundant or highly correlated indicators (such as CPU utilization, memory occupation, disk I / O, etc.), and direct analysis will lead to high computational complexity and high noise interference, and feature dimension reduction can improve the efficiency of anomaly detection by retaining key information and removing redundancy.
[0065] In the formula, the original features are projected to the direction with the maximum variance (principal component) through linear transformation, and the first one or several principal components (contribution rate > 95%) are retained, and if there is a class label (such as normal / abnormal), the linear transformation can maximize the inter-class distance and minimize the intra-class distance to generate more discriminative dimension-reduced features; the dimension-reduced features are arranged in time sequence to form a matrix, for example, each row represents a dimension-reduced feature vector (5 dimensions) at a time point, each column represents a feature dimension, and the matrix size is time point number x 5.
[0066] Further, weights are assigned to quantify the importance of each feature vector in the overall data, highlight abnormal sensitive features, and suppress noise interference, the variance of each feature dimension is calculated, and the greater the variance, the more intense the fluctuation, which may contain more abnormal information, and the weight is set as variance / total variance; wherein, the weight can be dynamically adjusted according to the historical abnormal detection results, for example, if a certain feature dimension frequently participates in abnormal detection in the past, the weight thereof is increased.
[0067] Specifically, the regression calculation is used to quantify the degree of deviation of the feature vector from the normal mode, and the key abnormal signal is highlighted by combining the weight, the dimension-reduced feature vector is multiplied by the weight and input into a logistic regression model, and an abnormal probability is output; the weight can also be introduced into the loss function of a weighted support vector machine (SVM) to make the abnormal samples (high weight) have a greater influence on the decision boundary.
[0068] Further, the abnormal threshold is used to filter high-confidence anomalies to avoid misjudgment of low-probability noise as an anomaly, and a fixed threshold (such as a vector with an abnormal probability > 0.8 is regarded as an anomaly) can be set according to business requirements, a vector with an abnormal probability p satisfying p greater than the abnormal threshold is marked as an anomaly, and otherwise, it is marked as normal.
[0069] In the embodiment of the application, the feature dimension reduction and weighted regression are cooperatively optimized to significantly improve the detection efficiency of abnormal data in computer observation, the feature dimension reduction technology compresses high-dimensional observation data into a low-dimensional space to reduce the computational overhead, and the dynamic weight distribution mechanism based on statistics or distance highlights the abnormal sensitive features and reduces the false positive rate, which can process TB-level observation data in real time and provide efficient and accurate anomaly warning support for cloud computing, data centers and other scenarios.
[0070] S5, determining an abnormal level corresponding to the target abnormal data based on the time sequence fluctuation risk score of the target abnormal data.
[0071] In the embodiment of the present application, the time sequence dependency and prediction deviation of the abnormal data are analyzed by the long short-term memory network, the risk score is generated by combining the volatility evaluation, and finally the abnormal level of the target abnormal data is determined.
[0072] As shown in Figure 3 In the embodiment of the present application, the time sequence volatility risk score based on the target abnormal data is used to determine the abnormal level corresponding to the target abnormal data, which includes: The time dependency of the target abnormal data is captured by using a preset long short-term memory layer; The time dependency is subjected to multi-layer full connection to obtain a data prediction value of the target abnormal data; The target abnormal data and the data prediction value are subjected to visual analysis to obtain an abnormal data trend chart; The abnormal data trend chart is subjected to volatility analysis, and the time sequence volatility risk score is generated according to the analysis result; The abnormal level corresponding to the target abnormal data is determined according to the time sequence volatility risk score.
[0073] In detail, the long short-term memory network (LSTM) has a gating mechanism to remember key historical information, the input gate controls the update degree of the input information at the current time point to the abnormal data, the forgetting gate decides to retain or discard the information in the historical cell state, and the output gate generates the output at the current time point according to the state of the abnormal data; for example, combining the long-term CPU rising trend and the short-term memory sudden increase, the output high abnormal probability.
[0074] Further, the time dependency output by the LSTM needs to be mapped to a specific prediction value through a full connection layer (FC), and a multi-layer structure can enhance the nonlinear expression ability, and high-level features are extracted step by step by stacking multiple full connection layers (such as 128→64→32→1); for example, the first layer extracts the short-term trend, the second layer extracts the long-term mode, and the output layer uses linear activation to directly generate a continuous prediction value (such as the utilization percentage).
[0075] Among them, visualization can intuitively show the deviation between actual value and prediction value, and assist artificial rapid positioning of abnormal mode (such as peak, continuous deviation), including line chart, horizontal axis for time, vertical axis for observation index value, actual value (blue) and prediction value (orange) are distinguished by different colors, heat map is used for multi-index abnormality, and deviation degree is represented by color depth, for example, red represents high deviation, and blue represents low deviation.
[0076] In detail, the volatility analysis quantifies the severity and persistence of the anomaly, providing a quantitative basis for the risk score, and calculates volatility indicators such as standard deviation and maximum deviation rate, the standard deviation measures the degree of data deviation from the mean, for example, if the CPU utilization standard deviation is 20%, and the current volatility reaches 30%, it means that the anomaly is severe.
[0077] The maximum deviation rate refers to the maximum deviation of the actual value from the predicted value, for example, the predicted value is 50%, the actual value is up to 90%, and the deviation rate is 80%.
[0078] Wherein, the standard deviation, deviation rate and the like are assigned weights (such as 0.4, 0.4) to calculate the comprehensive score, namely the time sequence volatility risk score, the anomaly level includes low risk, medium risk and high risk; low risk (0-60 points): temporary fluctuation, no immediate treatment is needed, for example, the CPU utilization temporarily rises to 70% and then recovers; medium risk (61-80 points): attention is needed, which may cause subsequent problems, for example, memory occupancy is above 80% for a long time; high risk (81-100 points): emergency disposal is needed to avoid system crash, for example, disk I / O is 100% for more than 10 minutes.
[0079] In the embodiment of the present application, through the fusion of time sequence alignment and deep learning technology, the accuracy and automation level of anomaly level determination in observation are significantly improved, the LSTM network effectively captures the long-term dependence relationship in abnormal data, improves the accuracy of anomaly pattern recognition, and the dynamic scoring mechanism based on volatility analysis quantifies the abnormal risk combined with the visual trend chart, realizes the automatic classification (low / medium / high) of the anomaly level, reduces the manual intervention, and improves the data observation alarm efficiency.
[0080] S6, according to the anomaly level and the target observation alarm strategy, the target abnormal data is pushed to the preset target observation object end, and the target abnormal data after pushing is observed in real time to obtain target observation data.
[0081] In the embodiment of the present application, the target abnormal data is observed in real time by the data acquisition probe, the multi-dimensional performance indicators are extracted and the abnormal marking and merging processing are performed, and finally the target observation data is generated.
[0082] In the embodiment of the present application, through the multi-channel redundant design, it is ensured that the alarm is accurately reached, and the false alarm caused by single channel failure is avoided, the push channel type includes instant messaging type, supports text message, link jump (such as directly jumping to the abnormal detail page), short message / voice call, or triggering a third party system (such as Jira work order, PagerDuty) to automatically create a task, requiring the receiver to confirm the alarm (such as clicking the "processed" button), and when not confirmed, upgrading the push (such as from group chat to single phone call), so as to ensure that the alarm is effectively processed.
[0083] In the embodiment of the present application, the target abnormal data after pushing is observed in real time to obtain target observation data, comprising: Obtaining a data collection probe of the target observation object end, performing data observation on the target abnormal data according to the data collection probe to obtain an original observation data stream; Extracting a multi-dimensional performance index of the original observation data stream, and performing abnormal alarm on the performance index deviating from a preset index threshold in the multi-dimensional performance index to obtain marked observation data; Merging the marked observation data, and taking the merged marked observation data as target observation data.
[0084] In detail, the data collection probe is a lightweight software module deployed on the observation object (such as a server, a network device), which is responsible for real-time collection of original indexes related to abnormal data; wherein, the host-level probe is installed at the operating system level (such as Node Exporter of Linux), which collects system-level indexes such as CPU utilization, memory occupation, and disk I / O; the application-level probe is embedded in the application program (such as JMX probe of Java application), which collects application layer indexes such as business transaction volume, response time, and error rate.
[0085] Further, the probe periodically sends a request to the observation object (such as requesting QPS of MySQL once every 10 seconds), obtains the latest index value, the observation object actively pushes the index change event (such as Kafka message) to the probe, and the probe receives and forwards in real time to obtain the original observation data stream.
[0086] Specifically, the multi-dimensional performance index reflects the running state of the observation object in different dimensions, the abnormal marker quickly locates the indexes deviating from the normal range through threshold comparison, the system dimension includes CPU (user state / kernel state / idle rate), memory (usage / caching / swap partition), disk (read / write rate / IOPS / utilization), application dimension includes business transaction volume (order number / login times), response time (P50 / P90 / P99), error rate (HTTP 5xx error ratio), network dimension includes bandwidth utilization (in / out direction), packet loss rate (TCP / UDP), delay (RTT average / max).
[0087] Among them, according to historical experience or equipment specification, a fixed value is set, any index exceeding the threshold is marked, the same index in the same time window (such as 1 minute) is merged into one record, for example, CPU utilization exceeds 85% for 3 times in 1 minute, which is merged into 1 "CPU high load" record, or the observation object with dependent relationship is aggregated, for example, the sudden increase of database server delay leads to the lengthening of application server response time, which is merged into 1 "database-application link performance decline" record.
[0088] In the embodiment of the present application, through the cooperative mechanism of probe collection, multi-dimensional threshold marking and intelligent merging, the efficiency and accuracy of real-time observation are significantly improved. The lightweight data collection probe realizes millisecond-level index grabbing, covers system, application and network full-dimensional data, and ensures that no abnormal data is missed. The dynamic threshold and multi-index correlation marking technology improve the data observation efficiency and accuracy.
[0089] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0090] As shown in Figure 4 is a functional module diagram of a data availability observation and alarm device provided by an embodiment of the present application.
[0091] In the embodiment of the present application, a data availability observation and alarm device is provided, which corresponds one-to-one to the above-mentioned data availability observation and alarm method. As shown in Figure 4 , the data availability observation and alarm device 100 can be installed in an electronic device. According to the functions implemented, the data availability observation and alarm device 100 includes an index analysis module 101, a feature extraction module 102, a strategy optimization module 103, an anomaly detection module 104, a level determination module 105, and a real-time observation module 106. The detailed descriptions of each functional module are as follows: The index analysis module 101 is configured to obtain observation data of a target observation object, perform index analysis on the observation data, and obtain observation index data. The feature extraction module 102 is configured to perform time feature analysis on the observation index data, and obtain observation data features. The strategy optimization module 103 is configured to perform strategy optimization on a predefined initial observation alarm strategy according to the observation data features, and obtain a target observation alarm strategy. The anomaly detection module 104 is configured to perform feature anomaly detection on the observation data features, and obtain target abnormal data. The level determination module 105 is configured to determine an abnormal level corresponding to the target abnormal data based on a time sequence fluctuation risk score of the target abnormal data. The real-time observation module 106 is configured to push the target abnormal data to a preset target observation object end according to the abnormal level and the target observation alarm strategy, and perform real-time observation on the pushed target abnormal data, and obtain target observation data.
[0092] In an embodiment, the index analysis module 101, when performing index analysis on the observation data to obtain observation index data, is configured to: perform data cleaning on the observation data to obtain cleaned observation data; obtain dimension attributes and a calculation expression corresponding to the cleaned observation data, parse the calculation expression according to a predefined regular expression specification to obtain a corresponding index name and a calculation function; splice the dimension attributes, the index name, and the calculation function to obtain an SQL expression; perform index calculation on the observation data according to the SQL expression to obtain observation index data.
[0093] In an embodiment, the feature extraction module 102, when performing time feature analysis on the observation index data to obtain observation data features, is configured to: obtain a timestamp of the observation index data, perform time alignment processing on the observation index data according to the timestamp to obtain standard index data; obtain a time window, perform sliding segmentation on the standard index data according to the time window to obtain a plurality of time segment index data; calculate time domain features of the plurality of time segment index data; perform frequency domain transformation on the plurality of time segment index data, and extract frequency domain features of the time segment index data after frequency domain transformation; perform feature fusion on the time domain features and the frequency domain features to obtain observation data features.
[0094] In an embodiment, the strategy optimization module 103, when performing strategy optimization on a predefined initial observation alarm strategy according to the observation data features to obtain a target observation alarm strategy, is configured to: obtain a plurality of initial observation alarm strategies, parse the plurality of initial observation alarm strategies to obtain a plurality of alarm targets; perform matching degree analysis on the plurality of alarm targets and the observation data features to obtain data alarm matching degrees; construct a target loss function according to the alarm targets, the data alarm matching degrees, and the observation data features, and determine a strategy loss value of each initial observation alarm strategy according to the target loss function; select a first observation alarm strategy greater than a preset strategy loss threshold according to the strategy loss value; perform cross processing on the first observation alarm strategy to obtain a second observation alarm strategy; returning the second observation alarm strategy according to a preset number of times, and stopping returning when the number of times reaches the preset number of times, and taking the final second observation alarm strategy as a target observation alarm strategy.
[0095] In an embodiment, the anomaly detection module 104, when performing feature anomaly detection on the observation data features to obtain target abnormal data, is configured to: perform feature dimension reduction on the observation data features, and construct a dimension-reduced feature matrix according to the dimension-reduced observation data features; calculate an assigned weight of each dimension-reduced feature vector in the dimension-reduced feature matrix; perform regression calculation on the dimension-reduced feature vector according to the assigned weight to obtain an anomaly probability of the dimension-reduced feature vector; take the dimension-reduced feature vector with an anomaly probability greater than a preset anomaly threshold as the target abnormal data.
[0096] In an embodiment, the level determination module 105, when performing time series fluctuation risk scoring based on the target abnormal data to determine an abnormal level corresponding to the target abnormal data, is configured to: capture time dependence of the target abnormal data by using a preset long short-term memory layer; perform multi-layer full connection on the time dependence to obtain a data predicted value of the target abnormal data; perform visual analysis on the target abnormal data and the data predicted value to obtain an abnormal data trend graph; perform fluctuation analysis on the abnormal data trend graph, and generate a time series fluctuation risk score according to an analysis result; determine the abnormal level corresponding to the target abnormal data according to the time series fluctuation risk score.
[0097] In an embodiment, the real-time observation module 106, when performing real-time observation on the target abnormal data after pushing to obtain target observation data, is configured to: obtain a data collection probe of the target observation object end, perform data observation on the target abnormal data according to the data collection probe to obtain an original observation data stream; extract a multi-dimensional performance index of the original observation data stream, and perform anomaly alarm on a performance index deviating from a preset index threshold in the multi-dimensional performance index to obtain labeled observation data; perform data merging on the labeled observation data, and take the merged labeled observation data as the target observation data.
[0098] In the present application, the specific definition of the data availability observation and alarm device can refer to the definition of the data availability observation and alarm method in the foregoing, which will not be described herein again. Each module in the data availability observation and alarm device described above can be realized by software, hardware, and combinations thereof, in whole or in part. The modules described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0099] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program, when executed by the processor, implements the functions or steps of the server side of the data availability observation and alarm method.
[0100] In an embodiment, a computer device is provided, which can be a client, and an internal structure diagram thereof can be as shown in Figure 6 The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program, when executed by the processor, implements the functions or steps of the client side of the data availability observation and alarm method.
[0101] In an embodiment, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program: Obtaining observation data of a target observation object, performing index analysis on the observation data to obtain observation index data; Performing time characteristic analysis on the observation index data to obtain observation data characteristics; According to the observation data characteristics, a pre-defined initial observation alarm strategy is optimized to obtain a target observation alarm strategy; Feature anomaly detection is performed on the observation data characteristics to obtain target abnormal data; Based on the time sequence fluctuation risk score of the target abnormal data, an abnormal level corresponding to the target abnormal data is determined; According to the abnormal level and the target observation alarm strategy, the target abnormal data is pushed to a preset target observation object end, and real-time observation is performed on the pushed target abnormal data to obtain target observation data.
[0102] In several embodiments provided in the present application, it should be understood that the disclosed devices and apparatuses can be implemented in other manners. For example, the above-described system embodiments are merely illustrative, and the division of the modules is merely a logical function division, and there can be another division manner in actual implementation.
[0103] In addition, each function module in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware, or in the form of hardware plus software function module.
[0104] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0105] In some embodiments of the present embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium, wherein the computer program is executed by a processor to implement the steps of the method described in the above embodiments.
[0106] The readable storage medium of the present application stores a computer program, and the computer program can realize the following when executed by a processor of an electronic device: Obtaining observation data of a target observation object, performing index analysis on the observation data to obtain observation index data; Performing time feature analysis on the observation index data to obtain observation data characteristics; According to the observation data characteristics, a pre-defined initial observation alarm strategy is optimized to obtain a target observation alarm strategy; Feature anomaly detection is performed on the observation data characteristics to obtain target abnormal data; determine an abnormality level corresponding to the target abnormal data based on a time sequence fluctuation risk score of the target abnormal data; push the target abnormal data to a preset target observation object end according to the abnormality level and the target observation alarm strategy, and perform real-time observation on the pushed target abnormal data to obtain target observation data.
[0107] It should be noted that the functions or steps described above with respect to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0108] The computer-readable storage medium can also store at least one computer executable program / instruction, such as computer readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may, for example, include read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then when the computing device runs the computer readable instructions stored on the computer readable storage medium, the various methods described above can be performed.
[0109] In addition, the computer device can also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and an input / output device (for example, a keyboard, a mouse, a speaker, etc.), etc.
[0110] In one embodiment, the at least one computer executable instruction can also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the steps of various functions and / or methods in the embodiments described in the present technology.
[0111] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium and can include the processes of the above-mentioned embodiments when executed. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory.
[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit, module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units or modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0113] In the embodiments provided by the present disclosure, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are merely illustrative, for example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, program segment or part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the block can occur in different order from that noted in the drawings. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for implementing the specified function or action, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0114] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
[0115] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction, and do not represent actual use.
Claims
1. A data availability observation and alarm method, characterized in that: The method comprises: Acquiring observation data of a target observation object, performing index analysis on the observation data, and obtaining observation index data; Performing time characteristic analysis on the observation indicator data to obtain observation data characteristics; Optimizing a predefined initial observation alarm strategy according to the observation data characteristics to obtain a target observation alarm strategy; Performing feature anomaly detection on the observed data features to obtain target anomaly data; Determining an abnormality level corresponding to the target abnormal data based on the time series fluctuation risk score of the target abnormal data; The target abnormal data is pushed to a preset target observation object end according to the abnormal level and the target observation alarm strategy, and the pushed target abnormal data is observed in real time to obtain target observation data.
2. The data availability observation and alarm method according to claim 1, characterized in that: The performing of indicator analysis on the observation data to obtain observation indicator data includes: performing data cleaning on the observation data to obtain cleaned observation data; Obtain the dimension attributes and calculation expressions corresponding to the cleaned observation data, parse the calculation expressions according to predefined regular expression specifications, and obtain the corresponding indicator names and calculation functions; Concatenate the dimension attribute, the indicator name, and the calculation function to obtain an SQL expression; The observed data is subjected to an index calculation according to the SQL expression to obtain observed index data.
3. The data availability observation and alarm method according to claim 1, characterized in that: The performing of time characteristic analysis on the observation indicator data to obtain observation data characteristics includes: Obtaining a timestamp of the observed indicator data, and performing time alignment processing on the observed indicator data according to the timestamp to obtain standard indicator data; Obtain a time window, and perform sliding segmentation on the standard indicator data according to the time window to obtain multiple time segment indicator data; Calculating time domain features of a plurality of the time segment indicator data; Performing frequency domain transformation on the plurality of time segment indicator data, and extracting frequency domain features of the time segment indicator data after the frequency domain transformation; The time domain features and the frequency domain features are fused to obtain observation data features.
4. The data availability observation and alarm method according to claim 1, wherein: The optimizing the predefined initial observation alarm strategy according to the observation data characteristics to obtain the target observation alarm strategy includes: Acquire multiple initial observation alarm strategies, and analyze the multiple initial observation alarm strategies to obtain multiple alarm targets; Performing matching analysis on the plurality of alarm targets and the observation data features to obtain a data alarm matching degree; Constructing a target loss function according to the alarm target, the data alarm matching degree, and the observed data characteristics, and determining a strategy loss value of each of the initial observation alarm strategies according to the target loss function; Filtering out a first observation alarm strategy that is greater than a preset strategy loss threshold according to the strategy loss value; Cross-processing the first observation alarm strategy to obtain a second observation alarm strategy; Returning the second observation alarm strategy to the step of determining the strategy loss value of each of the initial observation alarm strategies according to the target loss function, and counting the number of returns; When the number of returns reaches a preset number, the return is stopped and the final second observation alarm strategy is used as the target observation alarm strategy.
5. The data availability observation and alarm method according to claim 1, wherein: The performing feature anomaly detection on the observed data features to obtain target anomaly data includes: Performing feature dimensionality reduction on the observation data features, and constructing a dimensionality reduction feature matrix based on the feature dimensionality reduced observation data features; Calculating the allocation weight of each reduced-dimensionality feature vector in the reduced-dimensionality feature matrix; Performing regression calculation on the reduced-dimensionality feature vector according to the assigned weight to obtain an abnormal probability of the reduced-dimensionality feature vector; The reduced-dimensional feature vector whose abnormal probability is greater than a preset abnormal threshold is used as target abnormal data.
6. The data availability observation and alarm method according to claim 1, wherein: The determining, based on the time series fluctuation risk score of the target abnormal data, the abnormality level corresponding to the target abnormal data includes: Using a preset long short-term memory layer to capture the temporal dependency of the target abnormal data; Performing multi-layer full connection on the time dependency to obtain a data prediction value of the target abnormal data; Performing visual analysis on the target abnormal data and the data prediction value to obtain an abnormal data trend graph; Performing a volatility analysis on the abnormal data trend graph, and generating a time series volatility risk score based on the volatility obtained from the analysis results; The abnormality level corresponding to the target abnormal data is determined according to the time series fluctuation risk score.
7. The data availability observation and alarm method according to claim 1, characterized in that: The real-time observation of the pushed target abnormal data to obtain target observation data includes: Acquire a data acquisition probe at the target observation object end, and observe the target abnormal data according to the data acquisition probe to obtain an original observation data stream; Extracting multidimensional performance indicators of the original observation data stream, and issuing abnormal alarms for performance indicators in the multidimensional performance indicators that deviate from preset indicator thresholds to obtain marked observation data; The labeled observation data are merged, and the merged labeled observation data are used as target observation data.
8. A data availability observation and alarm device, characterized in that: The device comprises: An index analysis module is used to obtain observation data of a target observation object, perform index analysis on the observation data, and obtain observation index data; A feature extraction module is used to perform time feature analysis on the observation indicator data to obtain observation data features; A strategy optimization module is used to optimize the predefined initial observation alarm strategy according to the observation data characteristics to obtain a target observation alarm strategy; An anomaly detection module, used to perform feature anomaly detection on the observed data features to obtain target anomaly data; A level determination module, configured to determine an abnormality level corresponding to the target abnormal data based on the time series fluctuation risk score of the target abnormal data; The real-time observation module is used to push the target abnormal data to the preset target observation object end according to the abnormal level and the target observation alarm strategy, and perform real-time observation on the pushed target abnormal data to obtain target observation data.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute a data availability observation and alarm method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data availability observation and alarm method according to any one of claims 1 to 7 is implemented.