A software running risk early warning system based on machine learning
Patent Information
- Application Number
- CN202610998785.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
(1)现有软件运行预警系统多依赖固定阈值、单项性能指标、人工规则和普通日志告警,对日志模板、接口调用链、资源占用数据和异常事件之间的关联利用不足,难以将软件运行过程表达为状态节点、状态转移边和风险转移路径,导致系统只能识别单点异常,不能准确判断低频状态转移、异常路径演化和风险传播过程
(1)利用日志模板解析和软件运行状态机生成,通过日志模板解析模块对日志数据进行日志模板提取、变量字段识别、日志事件编码和日志序列构建,再通过软件运行状态机生成模块根据日志模板序列、接口调用链数据、资源占用数据和异常事件数据识别软件运行状态节点、状态转移边和风险转移路径,生成能够表达软件运行过程结构的软件运行状态转移图,解决了现有预警系统只依赖孤立日志和单项指标,难以表达软件运行状态演化关系和风险传播路径的问题。
Smart Images

Figure CN122817032A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of software operation monitoring technology, and in particular relates to a software operation risk early warning system based on machine learning. Background Technology
[0002] As enterprise software systems evolve towards microservice architecture, cloud-native deployment, distributed calls, and continuous version iteration, the software continuously generates application logs, system logs, API call chains, resource usage data, exception events, version release data, configuration change data, and user access data during operation. This operational data reflects changes in the software's operational status during request processing, API calls, database access, cache access, external API calls, exception recovery, and after version releases. Failure to promptly identify low-frequency anomalies, performance spikes, state transition anomalies, and risk propagation paths from this data can easily lead to delayed software fault detection, an abnormal expansion of business interfaces, continuously increasing resource consumption, and system service unavailability.
[0003] The existing technology has at least the following problems that need to be improved: (1) Existing software operation early warning systems mostly rely on fixed thresholds, single performance indicators, manual rules and ordinary log alarms. They do not make sufficient use of the correlation between log templates, interface call chains, resource usage data and abnormal events. It is difficult to express the software operation process as state nodes, state transition edges and risk transition paths, resulting in the system being able to identify single-point anomalies and not accurately judge low-frequency state transitions, abnormal path evolution and risk propagation processes.
[0004] (2) Existing machine learning-based software risk identification schemes often directly use historical log samples and operational indicator samples to train models. The number of low-frequency abnormal samples, boundary risk samples and unknown risk path samples is insufficient. The models are unstable in identifying data drift, performance regression after version release, sudden abnormal events and continuous deterioration trends. At the same time, false alarm samples, missed alarm samples and manual confirmation results are difficult to write back to the model training, state transition diagram and warning threshold in a timely manner, resulting in the warning model being fixed for a long time and difficult to adapt to software version iteration and changes in business operation mode.
[0005] Therefore, a machine learning-based software operation risk early warning system is needed to transform software operation data into structured risk features and combine state machine-enhanced samples, multi-model risk identification, change point fusion early warning, and feedback update mechanisms to achieve early identification, level output, and continuous optimization of software operation risks. Summary of the Invention
[0006] To address the above issues and overcome the shortcomings of existing technologies, this invention provides a machine learning-based software operation risk early warning system. Utilizing log template parsing, software operation state machine generation, rare risk sample enhancement, multi-model risk identification, change point fusion early warning, and feedback update principles, it transforms logs, call chains, resource indicators, and abnormal events into software operation state transition diagrams, rare risk sample sets, machine learning risk scores, and fusion early warning scores. This enables low-frequency risk identification, performance mutation early warning, risk level output, and continuous model correction, solving the problems of traditional software early warning systems, such as single-point judgment, insufficient samples, high false positives and false negatives, and difficulty in adapting to version iterations.
[0007] The technical solution adopted in this invention is as follows: a software operation risk early warning system based on machine learning, including an operation data acquisition module, a log template parsing module, a software operation state machine generation module, a rare risk sample enhancement module, a multi-model risk identification module, a change point fusion early warning module, a risk level output module, and a feedback update module; The runtime data acquisition module collects log data, interface call chain data, resource usage data, abnormal event data, version release data, configuration change data, and user access data during the operation of the target software. It then performs time alignment, field cleaning, source tagging, and runtime instance binding on the above data to generate the original dataset of the software operation. The log template parsing module extracts log templates, identifies variable fields, encodes log events, and constructs log sequences from the log data in the original dataset of software operation, generating a log template sequence. The software runtime state machine generation module identifies software runtime state nodes, state transition edges, and risk transition paths based on log template sequences, interface call chain data, resource usage data, and abnormal event data, and generates a software runtime state transition diagram. The rare risk sample enhancement module generates enhanced training samples based on the normal state transitions, abnormal state transitions, low-frequency state transitions, and boundary state transitions in the software operation state transition graph, forming a rare risk sample set. The multi-model risk identification module inputs the original software running dataset, log template sequence, software running state transition diagram and rare risk sample set into multiple machine learning models to generate log risk score, performance risk score, call chain risk score, abnormal event risk score and state transition risk score respectively; The change point fusion early warning module generates a fusion early warning score based on the time series of operating indicators, machine learning risk scores, performance change point detection results, and abnormal event change results, and determines whether the software operation risk has changed abruptly. The risk level output module generates a software operation risk level based on the fusion early warning score, risk duration, risk impact scope, and business importance. The software operation risk level includes low risk level, medium risk level, high risk level, and severe risk level. The feedback update module receives manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results, and version rollback results, and updates the log template, software operation status transition diagram, rare risk sample set, machine learning model parameters, and early warning thresholds based on the feedback results.
[0008] Furthermore, the runtime data acquisition module includes a data source access unit, a time alignment unit, a runtime instance binding unit, a field cleaning unit, and a data quality marking unit. The data source access unit receives application logs, system logs, interface call chains, CPU utilization, memory utilization, disk I / O rate, network throughput, interface response time, interface error rate, exception stack, alarm records, version release time, configuration change records, and user access volume. The time alignment unit aligns data with different sampling frequencies to the same time axis according to a unified time window. The runtime instance binding unit binds runtime data to the corresponding software runtime instance based on the service name, instance number, container number, host number, version number, and interface name. The field cleaning unit performs unified processing on the time field, thread field, exception field, interface field, user field, resource field, and version field in the log data. The data quality marking unit generates data quality marks based on the data missing ratio, acquisition latency, number of field anomalies, and source stability.
[0009] Furthermore, the log template parsing module includes a log segmentation unit, a template candidate generation unit, a variable field identification unit, a template matching scoring unit, and a log event encoding unit. The log segmentation unit performs structured segmentation of log data according to timestamp, log level, thread number, service name, and log text to obtain a log lexical sequence. The template candidate generation unit distinguishes between fixed and variable lexicals in the log lexical sequence and extracts log template candidates. The variable field identification unit identifies numeric variables, address variables, interface variables, error code variables, thread variables, user variables, and path variables in the log text. The template matching scoring unit calculates the log template matching score based on the degree of lexical matching, variable type matching, time context matching, and call chain context matching. The log event encoding unit converts the successfully matched log templates into log event codes and generates a log template sequence in chronological order.
[0010] Furthermore, the software runtime state machine generation module includes a state node identification unit, a state transition edge construction unit, a transition frequency statistics unit, a risk path identification unit, and a state transition graph output unit. The state node identification unit generates software runtime state nodes based on log event encoding, interface call phase, resource usage interval, and abnormal event type. The state transition edge construction unit establishes state transition edges according to the order of software runtime state nodes within adjacent time windows and records the starting state node, target state node, transition time, transition frequency, average transition time, transition error rate, and transition resource consumption value. The transition frequency statistics unit counts the state transition frequency in the normal operation cycle, abnormal operation cycle, and post-release observation cycle, and identifies low-frequency state transitions and newly added state transitions. The risk path identification unit generates the state transition risk intensity based on the error rate, time consumption growth rate, resource consumption growth rate, and number of abnormal event triggers of the state transition edges. The state transition graph output unit combines the software runtime state nodes, state transition edges, state transition risk intensity, low-frequency state transitions, and risk paths into a software runtime state transition graph.
[0011] Furthermore, the rare risk sample enhancement module includes a low-frequency transition extraction unit, an abnormal path extraction unit, a boundary sample construction unit, an enhanced sample weight calculation unit, and an enhanced training set output unit. The low-frequency transition extraction unit extracts state transition edges with a frequency lower than the low-frequency threshold from the software operation state transition graph. The abnormal path extraction unit extracts transition paths from normal state nodes to abnormal state nodes within abnormal operation cycles. The boundary sample construction unit constructs boundary risk samples close to the risk level boundary line based on the risk score boundary between historical normal samples and abnormal samples. The enhanced sample weight calculation unit generates rare risk sample enhancement weights based on state transition rarity, state transition risk intensity, abnormal path length, boundary distance, and data quality label. The enhanced training set output unit filters candidate enhanced samples according to the rare risk sample enhancement weights and merges the filtered candidate enhanced samples with real operation samples to generate an enhanced training set.
[0012] Furthermore, the multi-model risk identification module includes a feature fusion unit, a basic model training unit, a sequence model training unit, a graph model training unit, a model voting unit, and a machine learning risk score output unit. The feature fusion unit concatenates runtime data features, log template sequence features, state transition graph features, rare risk sample features, version release features, and configuration change features to generate a software runtime risk feature vector. The basic model training unit trains a tree model and a linear classification model using an augmented training set and outputs the basic model risk probability. The sequence model training unit trains a time-series risk identification model using log template sequences and runtime indicator time series and outputs the sequence model risk probability. The graph model training unit trains a state transition risk identification model using a software runtime state transition graph and outputs the graph model risk probability. The model voting unit performs weighted fusion of the basic model risk probability, sequence model risk probability, graph model risk probability, abnormal event risk probability, and resource indicator risk probability to generate a machine learning risk score. The machine learning risk score output unit sends the machine learning risk score to the change point fusion early warning module.
[0013] Furthermore, the change point fusion early warning module includes an indicator window construction unit, a performance change point detection unit, an event surge detection unit, a model risk trend detection unit, a change point fusion score calculation unit, and an early warning triggering unit. The indicator window construction unit constructs a current window and a historical baseline window according to a preset time length. The performance change point detection unit calculates the degree of change in performance indicators between the current window and the historical baseline window, and generates a performance change point score. The event surge detection unit calculates the degree of surge in the number of abnormal events, the number of abnormal logs, and the number of error interfaces, and generates an abnormal event surge score. The model risk trend detection unit calculates the rate of increase and duration of the machine learning risk score within the continuous early warning window, and generates a model risk trend score. The change point fusion score calculation unit generates a change point fusion score based on the performance change point score, the abnormal event surge score, and the model risk trend score. The early warning triggering unit generates a software operation risk early warning signal when the change point fusion score reaches the change point early warning threshold, and sends the software operation risk early warning signal to the risk level output module.
[0014] Furthermore, the risk level output module includes a comprehensive risk score calculation unit, a risk level classification unit, a risk cause identification unit, an early warning information generation unit, and a handling suggestion output unit. The comprehensive risk score calculation unit generates a comprehensive risk early warning score based on machine learning risk scores, change point fusion scores, state transition risk intensity, abnormal event severity, business importance, and feedback correction coefficients. The risk level classification unit compares the comprehensive risk early warning score with low-risk, medium-risk, high-risk, and severe-risk thresholds to generate low-risk, medium-risk, high-risk, and severe-risk levels. The risk cause identification unit identifies the risk cause based on the highest-performing score source in the comprehensive risk early warning score. The early warning information generation unit generates software operation risk early warning information based on the risk level, risk cause, affected services, affected interfaces, abnormal state nodes, risk transfer paths, and occurrence time. The handling suggestion output unit outputs handling suggestions based on the risk cause.
[0015] Furthermore, the feedback update module includes a feedback data receiving unit, a false alarm sample marking unit, a missed alarm sample supplementation unit, a state transition graph update unit, a model parameter update unit, and a threshold update unit. The feedback data receiving unit receives manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results, version rollback results, and recovery confirmation results. The false alarm sample marking unit marks warning windows that are manually confirmed to have no risk as false alarm samples and records the reasons for the false alarms. The missed alarm sample supplementation unit marks running windows that were not warned by the system in a timely manner but were later confirmed to have malfunctioned as missed alarm samples and supplements the actual risk level, risk cause, and handling results. The state transition graph update unit updates the software running state transition graph based on newly added log templates, newly added state nodes, newly added state transition edges, false alarm samples, and missed alarm samples. The model parameter update unit adds false alarm samples, missed alarm samples, and rare risk samples to the incremental training sample pool and incrementally updates the model parameters in the multi-model risk identification module. The threshold update unit adjusts the template matching threshold, low-frequency threshold, risk observation threshold, change point warning threshold, and risk level threshold based on the warning accuracy, false alarm rate, missed alarm rate, and average early warning time.
[0016] The beneficial effects of this invention are as follows: (1) By using log template parsing and software running state machine generation, the log template parsing module extracts log templates, identifies variable fields, encodes log events, and constructs log sequences. Then, the software running state machine generation module identifies software running state nodes, state transition edges, and risk transition paths based on log template sequences, interface call chain data, resource usage data, and abnormal event data, and generates a software running state transition diagram that can express the structure of the software running process. This solves the problem that existing early warning systems rely only on isolated logs and single indicators, making it difficult to express the evolution relationship of software running state and risk propagation paths.
[0017] (2) By using rare risk samples under state machine constraints, the rare risk sample enhancement module extracts normal state transitions, abnormal state transitions, low-frequency state transitions and boundary state transitions from the software operation state transition graph. Based on the rarity of state transitions, the intensity of state transition risks, the length of abnormal paths, the boundary distance and data quality labels, a rare risk sample set is generated to increase the coverage of low-frequency abnormal samples, boundary risk samples and risk path samples, and improve the machine learning model's ability to identify low-frequency risks and unknown anomalies. This solves the problems of insufficient abnormal samples, insufficient boundary risk learning and weak generalization ability of traditional machine learning models.
[0018] (3) By using multi-model risk identification, change point fusion early warning and feedback update, the multi-model risk identification module generates machine learning risk scores, the change point fusion early warning module integrates performance change point scores, abnormal event surge scores and model risk trend scores, the risk level output module generates software operation risk level and early warning information, and the feedback update module writes back the manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results and version rollback results to the log template, software operation state transition diagram, rare risk sample set, machine learning model parameters and early warning threshold, forming a closed-loop early warning process from risk identification, risk mutation judgment, risk level output to continuous model correction, which solves the problems of high false alarm rate, high missed alarm rate, unclear risk level and difficulty in continuous updating after the model is launched. Attached Figure Description
[0019] The accompanying drawings are provided to further understand the present invention and form part of the specification. They are used together with the embodiments of the present invention to explain the invention and do not constitute a limitation thereof.
[0020] Figure 1 This is a schematic diagram of the overall structure of the machine learning-based software operation risk early warning system proposed in this invention. Figure 2 This is a schematic diagram of the software operation state transition diagram proposed in this invention; Figure 3 This is a schematic diagram of the system data flow proposed in this invention. Detailed Implementation
[0021] Example 1, see Figures 1-3 The present invention provides a software operation risk early warning system based on machine learning, including an operation data acquisition module, a log template parsing module, a software operation state machine generation module, a rare risk sample enhancement module, a multi-model risk identification module, a change point fusion early warning module, a risk level output module, and a feedback update module.
[0022] The runtime data acquisition module collects log data, interface call chain data, resource usage data, abnormal event data, version release data, configuration change data, and user access data of the target software during its operation. It then performs time alignment, field cleaning, source tagging, and runtime instance binding on the above data to generate the original dataset of the software operation.
[0023] The log template parsing module extracts log templates, identifies variable fields, encodes log events, and constructs log sequences from the log data in the original dataset of software operation, generating a log template sequence.
[0024] The software runtime state machine generation module identifies software runtime state nodes, state transition edges, and risk transition paths based on log template sequences, interface call chain data, resource usage data, and abnormal event data, and generates a software runtime state transition diagram.
[0025] The rare risk sample enhancement module generates enhanced training samples based on the normal state transition, abnormal state transition, low-frequency state transition and boundary state transition in the software operation state transition graph, forming a rare risk sample set.
[0026] The multi-model risk identification module inputs the original software running dataset, log template sequence, software running state transition diagram and rare risk sample set into multiple machine learning models to generate log risk score, performance risk score, call chain risk score, abnormal event risk score and state transition risk score respectively.
[0027] The change point fusion early warning module generates a fusion early warning score based on the time series of operating indicators, machine learning risk score, performance change point detection results, and abnormal event change results, and determines whether the software operation risk has changed abruptly.
[0028] The risk level output module generates a software operation risk level based on the fusion early warning score, risk duration, risk impact scope, and business importance. The software operation risk level includes low risk, medium risk, high risk, and severe risk.
[0029] The feedback update module receives manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results, and version rollback results, and updates the log template, state transition diagram, rare risk sample set, machine learning model parameters, and early warning thresholds based on the feedback results.
[0030] Through the modules described above, this system can uniformly convert software logs, operational metrics, interface call chains, abnormal events, and version changes into risk data that can be trained, identified, warned, and updated with feedback. This elevates software operation risk warning from a single log alarm to a closed-loop warning system that enhances state machines, identifies multiple models, fuses change points, and provides feedback correction.
[0031] Example 2: Based on all the above examples, the running data acquisition module specifically includes a data source access unit, a time alignment unit, a running instance binding unit, a field cleaning unit, and a data quality marking unit.
[0032] The data source access unit receives application logs, system logs, interface call chains, CPU utilization, memory utilization, disk I / O rate, network throughput, interface response time, interface error rate, exception stack, alarm records, version release time, configuration change records, and user access volume.
[0033] The time alignment unit uses a unified time window as a reference to align data with different sampling frequencies to the same time axis. The unified time window includes a minute-level window, an hour-level window, and a post-release observation window.
[0034] The instance binding unit binds runtime data to the corresponding software runtime instance based on the service name, instance number, container number, host number, version number, and interface name.
[0035] The field cleaning unit performs unified processing on time, thread, exception, interface, user, resource, and version fields in log data, removing invalid and duplicate fields.
[0036] The data quality tagging unit generates data quality tags based on the data missing ratio, collection delay time, number of field anomalies, and source stability, and writes the data quality tags into the original dataset used by the software.
[0037] Through the above processing, the runtime data acquisition module can unify runtime data scattered in the log system, monitoring system, call chain system and version management system into a single data base, solving the problems of scattered sources, inconsistent time standards and difficulty in matching runtime instances in traditional software risk warning data.
[0038] Example 3: Based on all the above examples, the log template parsing module specifically includes a log segmentation unit, a template candidate generation unit, a variable field identification unit, a template matching and scoring unit, and a log event encoding unit.
[0039] The log segmentation unit performs structured segmentation of log data according to timestamp, log level, thread number, service name, and log text to obtain a log term sequence.
[0040] The template candidate generation unit distinguishes between fixed and variable terms in the log term sequence and extracts log template candidates.
[0041] The variable field identification unit identifies numeric variables, address variables, interface variables, error code variables, thread variables, user variables, and path variables in the log text and replaces these variable fields with a unified variable identifier.
[0042] The template matching scoring unit calculates the log template matching score based on the degree of word matching, variable type matching, time context matching, and call chain context matching. The log template matching score is calculated according to Formula 1:
[0043] In Formula 1, For the lth log entry and the lth log entry Log template matching scores between log templates The number of fixed-term matches between the l-th log entry and the k-th log template. The total number of terms in the l-th log entry. Match the number of variables by field type. This represents the total number of variable fields. Match values for time context. To match values for the call chain context, For noise field interference values, , , , and These are the weight coefficients for the corresponding items.
[0044] The log event encoding unit converts successfully matched log templates into log event codes and generates a log template sequence in chronological order. When the log template matching score is lower than the template matching threshold, the log template parsing module generates new template candidates and sends the new template candidates to the feedback update module.
[0045] Regarding parameter adjustments: The first step is to increase the weight of variable type matching when there are many variable fields in the software log.
[0046] The second step is to increase the weight corresponding to the degree of context matching of the call chain when the completeness of the interface call chain is high.
[0047] The third step is to increase the weight of noise field interference values when there are a large number of debugging and random fields in the logs.
[0048] The fourth step is to add the new template candidate to the log template library when the number of consecutive occurrences of the new template candidate reaches the new template confirmation threshold.
[0049] Through the above processing, the log template parsing module can transform unstructured logs into stable log event encoding sequences, solving the problems of feature dimension confusion, inability to handle template drift, and unstable identification of newly added abnormal logs caused by traditional machine learning early warning systems that directly use raw logs.
[0050] Example 4: Based on all the above examples, the software running state machine generation module specifically includes a state node identification unit, a state transition edge construction unit, a transition frequency statistics unit, a risk path identification unit, and a state transition graph output unit.
[0051] The status node identification unit generates software running status nodes based on log event encoding, interface call stage, resource usage range, and abnormal event type. The software running status nodes include startup status node, authentication status node, request processing status node, database access status node, cache access status node, external interface call status node, degradation processing status node, exception recovery status node, and termination status node.
[0052] The state transition edge construction unit establishes state transition edges according to the order of software running state nodes within adjacent time windows. The state transition edge records the starting state node, the target state node, the transition time, the transition frequency, the average transition time, the transition error rate, and the transition resource consumption value.
[0053] The state transition frequency statistics unit counts the state transition frequency during normal operation cycles, abnormal operation cycles, and post-release observation cycles, and identifies low-frequency state transitions and newly added state transitions.
[0054] The risk path identification unit calculates the state transition risk intensity based on the error rate of the state transition edge, the growth rate of time consumption, the growth rate of resource consumption, and the number of abnormal event triggers. The state transition risk intensity is calculated according to Formula 2:
[0055] In Formula 2, The strength of the state transition risk from the p-th state node to the q-th state node. This represents the number of errors on the corresponding state transition edge. This represents the total number of transitions on the corresponding state transition edge. The average time taken for the current state transition. This represents the average time taken for state transitions within a historically stable period. This represents the time fluctuation value within a historically stable period. The resource consumption value is transferred to the current state. This represents the average resource consumption over a historically stable period. This represents the fluctuation value of resource consumption within a historically stable period. This represents the number of times an abnormal event is triggered on the corresponding state transition edge. This serves as the baseline value for triggering abnormal events. , , and These are the weight coefficients for the corresponding items.
[0056] The state transition diagram output unit combines software running state nodes, state transition edges, state transition risk intensity, low-frequency state transitions, and risky paths into a software running state transition diagram.
[0057] Regarding parameter adjustments: The first step is to increase the weight of the percentage of errors when the main risk to the software operation is an increase in the interface error rate.
[0058] The second step is to increase the weight of the item that increases the time consumption when the main risk of software operation is slow interface response.
[0059] The third step is to increase the weight of the resource consumption growth item when the main risks to software operation are memory leaks and abnormal CPU spikes.
[0060] Fourth, when a risk event occurs on a newly added state transition edge after the version is released, increase the risk intensity maintenance coefficient of the newly added state transition edge.
[0061] Through the above processing, the software runtime state machine generation module can reconstruct the software execution state and state transition relationship from log sequences and runtime data, and identify risk transfer paths, solving the problem that traditional risk warning models only look at isolated indicators and are difficult to express the structure of the software operation process and the risk propagation path.
[0062] Example 5: Based on all the above examples, the rare risk sample enhancement module specifically includes a low-frequency transfer extraction unit, an abnormal path extraction unit, a boundary sample construction unit, an enhanced sample weight calculation unit, and an enhanced training set output unit.
[0063] The low-frequency transition extraction unit extracts state transition edges that occur less frequently than the low-frequency threshold from the software running state transition graph and marks them as low-frequency state transitions.
[0064] The abnormal path extraction unit extracts the transition paths from normal state nodes to abnormal state nodes within the abnormal operation cycle and marks them as abnormal state transition paths.
[0065] The boundary sample construction unit constructs boundary risk samples that are close to the risk level boundary line based on the risk scoring boundary between historical normal samples and abnormal samples.
[0066] The augmentation weight calculation unit calculates the augmentation weight of rare-risk samples based on state transition rarity, state transition risk intensity, outlier path length, boundary distance, and data quality label. The augmentation weight of rare-risk samples is calculated according to Formula 3:
[0067] In Formula 3, Increase the weight of the rare risk sample for the s-th candidate boosting sample. Let be the frequency of state transitions corresponding to the s-th candidate augmentation sample. Let the state transition risk intensity be the s-th candidate enhancement sample. Let be the length of the abnormal path corresponding to the s-th candidate augmentation sample. This represents the maximum length of the abnormal path. Let be the distance between the s-th candidate augmented sample and the risk level boundary line. Let be the data quality penalty value for the s-th candidate augmentation sample. , , , and These are the weight coefficients for the corresponding items.
[0068] The augmented training set output unit filters candidate augmented samples according to the augmentation weight of rare risk samples, and merges the filtered candidate augmented samples with the real running samples to generate an augmented training set.
[0069] Regarding parameter adjustments: The first step is to increase the weight corresponding to the rarity of state transitions when the number of historical abnormal samples is insufficient, so that low-frequency risk paths can obtain a higher probability of enhancement.
[0070] The second step is to increase the weight corresponding to the abnormal path length when the abnormal path is long and spans multiple services, so that cross-service risk propagation samples can be included in the enhanced training set.
[0071] The third step is to increase the weight corresponding to the boundary distance when the model has a lot of false alarms near the risk level boundary, so that the model focuses on learning boundary risk samples.
[0072] The fourth step is to increase the data quality penalty value when the candidate augmentation sample comes from the collection window with low data quality, thereby reducing the probability of the corresponding sample entering the augmentation training set.
[0073] Through the above processing, the rare risk sample enhancement module can generate low-frequency risk samples, abnormal path samples, and boundary risk samples based on the software operation state transition diagram, solving the problems of few abnormal samples, insufficient coverage of rare risks, and weak generalization ability of the model to unknown anomalies in software operation risk warning.
[0074] Example 6: Based on all the above examples, the multi-model risk identification module specifically includes a feature fusion unit, a basic model training unit, a sequence model training unit, a graph model training unit, a model voting unit, and a machine learning risk scoring output unit.
[0075] The feature fusion unit combines runtime data features, log template sequence features, state transition diagram features, rare risk sample features, version release features, and configuration change features to generate a software runtime risk feature vector.
[0076] The base model training unit uses the augmented training set to train the tree model and the linear classification model, and outputs the base model risk probability.
[0077] The sequence model training unit uses log template sequences and time series of operational metrics to train a time series risk identification model and outputs the risk probability of the sequence model.
[0078] The graphical model training unit uses the software runtime state transition graph to train the state transition risk identification model and outputs the graphical model risk probability.
[0079] The model voting unit performs a weighted fusion of the risk probabilities of the basic model, sequence model, graph model, anomaly event, and resource indicator. The machine learning risk score is calculated according to Formula 4:
[0080] In Formula 4, For the first Machine learning risk scoring for each early warning window. Based on the risk probability of the basic model, For the risk probability of the sequence model, For the graph model risk probability, This represents the probability of abnormal event risk. This represents the probability of risk for resource indicators. , , , and These are the weight coefficients for the corresponding items.
[0081] The machine learning risk score output unit sends the machine learning risk score to the change point fusion early warning module.
[0082] Regarding parameter adjustments: The first step is to increase the weight of the risk probability corresponding to the sequence model when the software risk is mainly triggered by log anomalies.
[0083] The second step is to increase the weights corresponding to the risk probabilities in the graph model when software risks are mainly triggered by call chain anomalies and state transition anomalies.
[0084] The third step is to increase the weight of the probability of abnormal events when software risks are mainly triggered by a concentration of abnormal events.
[0085] The fourth step is to increase the weight of resource indicator risk probabilities when software risks are mainly triggered by abnormal CPU utilization, memory utilization, and interface response time.
[0086] Through the above processing, the multi-model risk identification module can comprehensively utilize ordinary machine learning models, time series models, and state transition diagram models to improve the stability of software operation risk identification and solve the problems of high false alarm rate and high false negative rate when traditional single-model early warning faces data drift, scarce abnormal samples, and low-frequency risks.
[0087] Example 7: This example is based on all the above examples. The change point fusion early warning module specifically includes an indicator window construction unit, a performance change point detection unit, an event surge detection unit, a model risk trend detection unit, a change point fusion scoring calculation unit, and an early warning triggering unit.
[0088] The indicator window construction unit constructs a current window and a historical benchmark window according to a preset time length. The current window includes the current interface response time, current error rate, current CPU utilization, current memory utilization, current network throughput, and current number of abnormal events. The historical benchmark window includes the corresponding indicators within the historical stable period.
[0089] The performance change detection unit calculates the degree of change in performance metrics between the current window and the historical baseline window, and generates a performance change score.
[0090] The event surge detection unit calculates the degree of surge in the number of abnormal events, the number of abnormal logs, and the number of error interfaces, and generates an abnormal event surge score.
[0091] The model risk trend detection unit calculates the rate of increase and duration of the machine learning risk score within a continuous warning window, and generates the model risk trend score.
[0092] The variable point fusion scoring unit generates a variable point fusion score based on the performance variable point score, the abnormal event surge score, and the model risk trend score. The variable point fusion score is calculated according to Formula 5:
[0093] In Formula 5, For the first The fusion score of change points in each early warning window This is the overall performance metric value for the current window. This is the comprehensive average of performance metrics from a historical baseline window. This represents the performance index fluctuation value of the historical benchmark window. This is the combined value of all exceptions for the current window. This is the comprehensive average of outliers over a historical baseline window. This represents the fluctuation value of abnormal events within the historical baseline window. Assess the machine learning risk score for the current window. For the machine learning risk score in the previous window, This refers to the number of times the risk observation threshold is exceeded within a consecutive window. The number of consecutive windows. , , and These are the weight coefficients for the corresponding items.
[0094] When the change point fusion score reaches the change point warning threshold, the warning triggering unit generates a software operation risk warning signal and sends the software operation risk warning signal to the risk level output module.
[0095] Regarding parameter adjustments: The first step is to increase the weight of performance metric changes when the software enters the observation period after a version release, so that performance regressions after release can be identified in advance.
[0096] The second step is to increase the weight of the abnormal event surge score when the number of abnormal events suddenly increases in a short period of time.
[0097] The third step is to increase the weight of the number of consecutive window exceedances when the machine learning risk score rises continuously but a single window does not reach the alarm threshold.
[0098] The fourth step is to increase the suppression effect of historical fluctuation values on the change point fusion score when the historical benchmark window fluctuates significantly, thereby reducing false alarms caused by natural fluctuations.
[0099] Through the above processing, the change point fusion early warning module can integrate runtime performance mutations, sudden increases in abnormal events, and machine learning risk trends into a unified early warning judgment, solving the problem that traditional software risk early warning only looks at single-point anomalies and cannot identify continuous degradation, performance regression, and sudden changes in risk trends.
[0100] Example 8: This example is based on all the above examples. The risk level output module specifically includes a comprehensive risk score calculation unit, a risk level classification unit, a risk cause location unit, an early warning information generation unit, and a handling suggestion output unit.
[0101] The comprehensive risk score calculation unit calculates the comprehensive risk warning score based on the machine learning risk score, change point fusion score, state transition risk intensity, anomaly event severity, business importance, and feedback correction coefficient. The comprehensive risk warning score is calculated according to Formula Six:
[0102] In Formula Six, Let the comprehensive risk warning score for the t-th warning window be... For machine learning risk scoring, For variable point fusion scoring, The intensity of state transition risk, The severity of the abnormal event. Based on business importance, For feedback correction coefficient, , , , , and These are the weight coefficients for the corresponding items.
[0103] The risk level classification unit compares the comprehensive risk warning score with the low-risk threshold, medium-risk threshold, high-risk threshold, and severe-risk threshold to generate low-risk level, medium-risk level, high-risk level, and severe-risk level.
[0104] The risk cause identification unit identifies the risk cause based on the highest-scoring source in the comprehensive risk warning score. Risk causes include log template anomalies, state transition anomalies, resource indicator anomalies, performance change point anomalies, interface call chain anomalies, sudden increase in abnormal events, and version change anomalies.
[0105] The early warning information generation unit generates software operation risk early warning information based on risk level, risk cause, affected services, affected interfaces, abnormal status nodes, risk transfer path, and occurrence time.
[0106] The handling suggestion output unit outputs corresponding handling suggestions based on the cause of the risk. The handling suggestions include log investigation suggestions, interface rate limiting suggestions, cache clearing suggestions, version rollback suggestions, configuration restoration suggestions, service restart suggestions, and manual review suggestions.
[0107] Through the above processing, the risk level output module can convert the machine learning model output results, change point detection results, state transition risks, and severity of abnormal events into executable software operation risk warning information, solving the problem that traditional warning systems only output alarm scores and lack risk levels, risk causes, and handling suggestions.
[0108] Example 9: This example is based on all the above examples. The feedback update module specifically includes a feedback data receiving unit, a false alarm sample marking unit, a missed alarm sample supplementation unit, a state transition graph update unit, a model parameter update unit, and a threshold update unit.
[0109] The feedback data receiving unit receives manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results, version rollback results, and recovery confirmation results.
[0110] The false alarm sample marking unit marks warning windows that are manually confirmed to have no risk as false alarm samples and records the reasons for the false alarms. The reasons for the false alarms include natural traffic fluctuations, planned version releases, temporary stress tests, short-term resource jitter, and non-business abnormal logs.
[0111] The missed sample supplementation unit marks the operation window that was not promptly alerted by the system but was later confirmed to have malfunctioned as a missed sample, and supplements the actual risk level, risk cause and handling result.
[0112] The state transition graph update unit updates the software operation state transition graph based on newly added log templates, newly added state nodes, newly added state transition edges, false alarm samples, and missed alarm samples.
[0113] The model parameter update unit adds false positive samples, false negative samples, and rare risk samples to the incremental training sample pool to incrementally update the model parameters in the multi-model risk identification module.
[0114] The threshold update unit adjusts the template matching threshold, low-frequency threshold, risk observation threshold, change point warning threshold, and risk level threshold based on the warning accuracy, false alarm rate, missed alarm rate, and average early warning time.
[0115] Through the above processing, the feedback update module can write back the results of manual confirmation, fault handling and operation recovery to the log template, state transition diagram, rare risk sample set, machine learning model and warning threshold, forming a continuous learning closed loop, which solves the problem that traditional software risk warning models are fixed for a long time after going online and cannot adapt to version iteration and changes in business operation mode.
Claims
1. A software operation risk early warning system based on machine learning, characterized in that: It includes a data acquisition module, a log template parsing module, a software state machine generation module, a rare risk sample enhancement module, a multi-model risk identification module, a change point fusion early warning module, a risk level output module, and a feedback update module; The runtime data acquisition module collects log data, interface call chain data, resource usage data, abnormal event data, version release data, configuration change data, and user access data during the operation of the target software. It then performs time alignment, field cleaning, source tagging, and runtime instance binding on the above data to generate the original dataset of the software operation. The log template parsing module extracts log templates, identifies variable fields, encodes log events, and constructs log sequences from the log data in the original dataset of software operation, generating a log template sequence. The software runtime state machine generation module identifies software runtime state nodes, state transition edges, and risk transition paths based on log template sequences, interface call chain data, resource usage data, and abnormal event data, and generates a software runtime state transition diagram. The rare risk sample enhancement module generates enhanced training samples based on the normal state transitions, abnormal state transitions, low-frequency state transitions, and boundary state transitions in the software operation state transition graph, forming a rare risk sample set. The multi-model risk identification module inputs the original software running dataset, log template sequence, software running state transition diagram and rare risk sample set into multiple machine learning models to generate log risk score, performance risk score, call chain risk score, abnormal event risk score and state transition risk score respectively; The change point fusion early warning module generates a fusion early warning score based on the time series of operating indicators, machine learning risk scores, performance change point detection results, and abnormal event change results, and determines whether the software operation risk has changed abruptly. The risk level output module generates a software operation risk level based on the integrated early warning score, risk duration, risk impact scope, and business importance. The feedback update module receives manual confirmation results, fault handling results, false alarm marking results, missed alarm supplementation results, and version rollback results, and updates the log template, software operation status transition diagram, rare risk sample set, machine learning model parameters, and early warning thresholds based on the feedback results.
2. The software operation risk early warning system based on machine learning according to claim 1, characterized in that: The runtime data acquisition module includes a data source access unit, a time alignment unit, a runtime instance binding unit, a field cleaning unit, and a data quality marking unit. The data source access unit receives application logs, system logs, interface call chains, CPU utilization, memory utilization, disk I / O rate, network throughput, interface response time, interface error rate, exception stack traces, alarm records, version release time, configuration change records, and user access volume. The time alignment unit aligns data from different sampling frequencies to the same time axis according to a unified time window. The runtime instance binding unit binds runtime data to the corresponding software runtime instance based on the service name, instance number, container number, host number, version number, and interface name. The field cleaning unit performs unified processing on the time field, thread field, exception field, interface field, user field, resource field, and version field in the log data. The data quality marking unit generates data quality marks based on the data missing ratio, acquisition latency, number of field anomalies, and source stability.
3. The software operation risk early warning system based on machine learning according to claim 2, characterized in that: The log template parsing module includes a log segmentation unit, a template candidate generation unit, a variable field identification unit, a template matching scoring unit, and a log event encoding unit. The log segmentation unit performs structured segmentation of log data according to timestamp, log level, thread number, service name, and log text to obtain a log lexical sequence. The template candidate generation unit distinguishes between fixed and variable lexicals in the log lexical sequence and extracts log template candidates. The variable field identification unit identifies numeric variables, address variables, interface variables, error code variables, thread variables, user variables, and path variables in the log text. The template matching scoring unit calculates the log template matching score based on the degree of lexical matching, variable type matching, time context matching, and call chain context matching. The log event encoding unit converts successfully matched log templates into log event codes and generates a log template sequence in chronological order.
4. The software operation risk early warning system based on machine learning according to claim 3, characterized in that: The software runtime state machine generation module includes a state node identification unit, a state transition edge construction unit, a transition frequency statistics unit, a risk path identification unit, and a state transition graph output unit. The state node identification unit generates software runtime state nodes based on log event codes, interface call stages, resource usage intervals, and abnormal event types. The state transition edge construction unit establishes state transition edges according to the order of software runtime state nodes within adjacent time windows and records the starting state node, target state node, transition time, transition frequency, average transition time, transition error rate, and transition resource consumption value. The transition frequency statistics unit counts the state transition frequency in the normal operation cycle, abnormal operation cycle, and post-release observation cycle, and identifies low-frequency state transitions and newly added state transitions. The risk path identification unit generates the state transition risk intensity based on the error rate of state transition edges, the growth rate of time consumption, the growth rate of resource consumption, and the number of abnormal event triggers; the state transition graph output unit combines software running state nodes, state transition edges, state transition risk intensity, low-frequency state transitions, and risk paths into a software running state transition graph.
5. The software operation risk early warning system based on machine learning according to claim 4, characterized in that: The rare risk sample enhancement module includes a low-frequency transition extraction unit, an abnormal path extraction unit, a boundary sample construction unit, an enhanced sample weight calculation unit, and an enhanced training set output unit. The low-frequency transition extraction unit extracts state transition edges with a frequency lower than the low-frequency threshold from the software operation state transition graph. The abnormal path extraction unit extracts transition paths from normal state nodes to abnormal state nodes within abnormal operation cycles. The boundary sample construction unit constructs boundary risk samples close to the risk level boundary line based on the risk score boundary between historical normal samples and abnormal samples. The enhanced sample weight calculation unit generates enhanced weights for rare risk samples based on state transition rarity, state transition risk intensity, outlier path length, boundary distance, and data quality label. The augmented training set output unit filters candidate augmented samples according to the augmentation weight of rare risk samples, and merges the filtered candidate augmented samples with the real running samples to generate an augmented training set.
6. The software operation risk early warning system based on machine learning according to claim 5, characterized in that: The multi-model risk identification module includes a feature fusion unit, a basic model training unit, a sequence model training unit, a graph model training unit, a model voting unit, and a machine learning risk score output unit. The feature fusion unit concatenates runtime data features, log template sequence features, state transition diagram features, rare risk sample features, version release features, and configuration change features to generate a software runtime risk feature vector; the basic model training unit trains a tree model and a linear classification model using an enhanced training set and outputs the basic model risk probability; the sequence model training unit trains a time-series risk identification model using log template sequences and runtime indicator time series and outputs the sequence model risk probability. The graphical model training unit uses the software runtime state transition graph to train the state transition risk identification model and outputs the graphical model risk probability. The model voting unit performs weighted fusion of the risk probabilities of the basic model, sequence model, graph model, abnormal event, and resource indicator to generate a machine learning risk score; The machine learning risk score output unit sends the machine learning risk score to the change point fusion early warning module.
7. A machine learning-based software operation risk early warning system according to claim 6, characterized in that: The change point fusion early warning module includes an indicator window construction unit, a performance change point detection unit, an event surge detection unit, a model risk trend detection unit, a change point fusion score calculation unit, and an early warning triggering unit; The indicator window construction unit constructs the current window and the historical benchmark window according to a preset time length; The performance change point detection unit calculates the degree of change in performance metrics between the current window and the historical baseline window, and generates a performance change point score. The event surge detection unit calculates the degree of surge in the number of abnormal events, the number of abnormal logs, and the number of error interfaces, and generates an abnormal event surge score. The model risk trend detection unit calculates the rate of increase and duration of the machine learning risk score within a continuous warning window, and generates the model risk trend score. The variable point fusion scoring calculation unit generates a variable point fusion score based on the performance variable point score, the abnormal event surge score, and the model risk trend score. When the change point fusion score reaches the change point warning threshold, the warning triggering unit generates a software operation risk warning signal and sends the software operation risk warning signal to the risk level output module.
8. The software operation risk early warning system based on machine learning according to claim 7, characterized in that: The risk level output module includes a comprehensive risk score calculation unit, a risk level classification unit, a risk cause location unit, an early warning information generation unit, and a handling suggestion output unit. The comprehensive risk score calculation unit generates a comprehensive risk warning score based on machine learning risk score, change point fusion score, state transition risk intensity, abnormal event severity, business importance, and feedback correction coefficient. The risk level classification unit compares the comprehensive risk warning score with low-risk thresholds, medium-risk thresholds, high-risk thresholds, and severe-risk thresholds to generate low-risk, medium-risk, high-risk, and severe-risk levels. The risk cause identification unit identifies the risk cause based on the highest-scoring source in the comprehensive risk warning score. The early warning information generation unit generates software operation risk early warning information based on risk level, risk cause, affected services, affected interfaces, abnormal status nodes, risk transfer path, and occurrence time; the handling suggestion output unit outputs handling suggestions based on the risk cause. The feedback update module includes a feedback data receiving unit, a false alarm sample marking unit, a missed alarm sample supplementation unit, a state transition diagram update unit, a model parameter update unit, and a threshold update unit. The feedback update module updates the software running state transition diagram, machine learning model parameters, and early warning thresholds based on the results of manual confirmation, fault handling, false alarm marking, missed alarm supplementation, version rollback, and recovery confirmation.