A Method and System for Operational Anomaly Detection and Root Cause Analysis Based on a Large Model

CN122412207BActive Publication Date: 2026-09-18LIAONING RONGKE ZHIWEIYUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610896281.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-18
Estimated Expiration
2046-06-22

AI Technical Summary

Technical Problem

一方面,静态阈值和规则库的构建高度依赖运维专家的领域知识,难以适应系统架构的动态演化与未知异常类型

Benefits of technology

[0054]Anomaly detection and root cause analysis based on a large language model significantly improves the accuracy and automation of fault identification in operational scenarios. By converting historical operational data into semantic text and mapping it to the state space, the system can automatically label key seed samples from a large amount of data, effectively reducing the cost of manual annotation. Combining reverse reasoning and sensitivity analysis, a high-coverage test dataset is dynamically generated, allowing the model to expose potential edge scenarios during the training phase, significantly reducing the probability of missed and false alarms in actual operations. By calculating the sensitivity threshold of state nodes, the key links leading to anomalies are accurately identified. Combined with the directional and amplitude components of the difference matrix, the prompt word template and perturbation intensity can be adaptively adjusted, ensuring high robustness of the model in complex and dynamic operational environments. This avoids the excessive reliance on expert experience in traditional rule-based methods, significantly shortening the model iteration cycle and improving the generalization effect of anomaly detection. The collaborative optimization of updating the prompt word template and perturbation intensity allows the large language model to continuously adapt to the distribution drift of operational data. The directional component ensures that the prompt word instructions always focus on the most relevant anomaly features, while the amplitude component controls the reasonable boundary of perturbation injection, preventing excessive perturbation from causing model failure. This dynamic adjustment mechanism effectively solves the pain point of performance degradation caused by changes in data distribution in traditional methods, ensuring that the anomaly detection rate remains stable and high during long-term deployment, while reducing the frequency of manual parameter tuning and realizing the continuous evolution of intelligent operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122412207B_ABST
    Figure CN122412207B_ABST
Patent Text Reader

Abstract

This invention relates to the field of operation and maintenance technology, and in particular to a method and system for operation and maintenance anomaly detection and root cause analysis based on a large model. By constructing a large language model and configuring an initial prompt word template, historical operation and maintenance data is converted into training semantic text input to the model to obtain judgment text, which is then encoded and mapped to a state space to label seed samples. The seed samples are then used for reverse reasoning to generate inverse state sequences, and perturbations are injected based on sensitivity to generate a test dataset. The test dataset is input into the model to obtain verification text, and the difference matrix is ​​decomposed to update the prompt word template and adjust the perturbation intensity. Finally, the operation and maintenance data to be detected is input into the model based on the updated prompt word template to obtain judgment text and locate the root cause, thus achieving accurate and efficient operation and maintenance anomaly detection and root cause localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of operation and maintenance technology, and in particular to a method and system for operation and maintenance anomaly detection and root cause analysis based on a large model. Background Technology

[0002] With the rapid development of information technology, the complexity of operation and maintenance systems is increasing daily, making anomaly detection and root cause analysis crucial for ensuring service stability. Current technologies primarily rely on two conventional approaches for anomaly detection: one is based on static threshold-based monitoring rules, where operations personnel manually set warning lines for various indicators based on historical experience and business characteristics; the other uses traditional machine learning models, such as isolated forests, support vector machines, or deep learning-based autoencoders, to perform pattern recognition on time-series indicators to identify data points deviating from normal behavior. Root cause analysis typically relies on expert systems or association rule mining, using predefined causal relationship graphs or frequent itemset extraction to locate the source of the fault. Some solutions also incorporate manual troubleshooting processes, inferring based on the timestamps and topological relationships of alarm events.

[0003] These conventional practices generally suffer from two major drawbacks. Firstly, the construction of static thresholds and rule bases heavily relies on the domain knowledge of operations and maintenance experts, making it difficult to adapt to the dynamic evolution of system architecture and unknown anomaly types. During peak business periods or when configuration changes cause baseline drift, fixed thresholds are highly susceptible to false positives or false negatives, leading to a sharp increase in rule maintenance costs. Secondly, traditional machine learning methods require a large number of labeled anomaly samples during the training phase, while fault data is sparse and labeling is costly in operations and maintenance scenarios, limiting the model's generalization ability. Root cause analysis modules often operate independently of anomaly detection modules, lacking a deep understanding of the semantic context of anomaly events. This results in location results relying on the experience coverage of association rules, leading to inefficient reasoning for emerging or combined faults and failing to achieve automated, interpretable end-to-end root cause tracing.

[0004] Furthermore, existing technologies for detection and localization primarily output discrete indicator points or alarm lists, failing to effectively utilize the semantic relationships and causal logic inherent in operational data. This makes it difficult for the system to distinguish between the underlying causes and surface symptoms of anomalies when faced with multimodal, high-dimensional logs and indicators, further exacerbating the burden on operational personnel in investigating massive amounts of alarms. Summary of the Invention

[0005] The embodiments of the present invention provide a method and system for detecting operational anomalies and performing root cause analysis based on a large model, which can solve the problems in the prior art.

[0006] A first aspect of this invention provides a method for detecting operational anomalies and performing root cause analysis based on a large model, comprising:

[0007] Construct a large language model including encoding and decoding layers and configure initial prompt word templates;

[0008] Acquire historical operation and maintenance data, and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template;

[0009] The training semantic text is input into the large language model to obtain the decision text. The decision text is then encoded into semantic vectors and mapped to the state space to obtain the projection coordinates. Seed samples are labeled according to the distribution and dispersion of the projection coordinates.

[0010] Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset.

[0011] The test dataset is input into the large language model to obtain the verification text. The difference matrix between the verification text and the expected result is calculated. The difference matrix is ​​decomposed into directional components and amplitude components. The initial prompt word template is updated according to the directional components to obtain the updated prompt word template. The perturbation intensity is adjusted according to the amplitude components.

[0012] The system acquires the operation and maintenance data to be detected, converts the data into semantic text based on the update prompt word template, and inputs it into a large language model to obtain the judgment text. The system then locates the state node with the highest matching degree in the reverse state sequence based on the judgment text as the root cause localization result.

[0013] Obtaining historical operation and maintenance data, and converting the historical operation and maintenance data into training semantic text based on the initial prompt word template includes:

[0014] Collect log information, indicator information, alarm information, and topology information as historical operation and maintenance data;

[0015] Extract the trigger timestamp of the alarm information as the time alignment reference point, calculate the time offset of the log information and indicator information relative to the time alignment reference point, determine the time window boundary based on the distribution statistics of the time offset, extract the log information and indicator information within the time window boundary and combine them with the alarm information to form a time synchronization dataset.

[0016] Based on the topology information, a propagation path is constructed starting from the service node corresponding to the alarm information. Log information and indicator information of each node are extracted from the time-series synchronization dataset along the propagation path. The extracted log information, indicator information, alarm information from the time-series synchronization dataset, and propagation path are organized into a causal event chain according to the propagation path order.

[0017] Named placeholders are extracted from the initial prompt word template, and log information, indicator information, alarm information and propagation path in the causal event chain are filled into the corresponding named placeholder positions to generate training semantic text.

[0018] The training semantic text is input into the large language model to obtain the decision text. The decision text is then semantically encoded and mapped to the state space to obtain the projected coordinates. Seed samples are labeled according to the distribution and dispersion of the projected coordinates, including:

[0019] The training semantic text is input into the large language model. The probability distribution of multiple candidate outputs is extracted from the decoding layer of the large language model. The probability distribution is sampled to generate multiple judgment texts. The multiple judgment texts are semantically encoded to obtain multiple judgment semantic vectors. The similarity distribution between the multiple judgment semantic vectors is calculated and statistically aggregated to obtain the judgment consistency score. The judgment semantic vector corresponding to the judgment text with the highest judgment consistency score is selected as the final judgment semantic vector.

[0020] The training semantic text is semantically encoded to obtain the input semantic vector. The difference between the final decision semantic vector and the input semantic vector is calculated as the decision offset vector. The decision offset vector is then mapped to the state space through dimensionality reduction transformation to obtain the projection coordinates.

[0021] Cluster analysis is performed on the projected coordinates to obtain cluster centers and cluster boundaries. The minimum distance from the projected coordinates to the cluster boundaries is calculated as the boundary distance, and the distance from the projected coordinates to the cluster centers is calculated as the center distance. The ratio of the boundary distance to the center distance is calculated as the distribution dispersion. The training semantic text corresponding to the projected coordinates with a distribution dispersion exceeding a preset dispersion threshold is marked as seed samples.

[0022] Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes with sensitivity exceeding a preset sensitivity threshold to generate a test dataset, including:

[0023] The seed sample is input into the large language model and forward propagation is performed to obtain the decision output. A computation graph from the decision output to the seed sample input is constructed. The decision output is fixed and backpropagation is performed along the computation graph. The gradient activation state of each layer during the backpropagation process is recorded as a state node. The state nodes are organized according to the propagation order from the output layer to the input layer to form a reverse state sequence.

[0024] The gradient magnitude of each state node in the reverse state sequence is calculated and normalized to obtain the sensitivity of each state node. State nodes with sensitivity exceeding the preset sensitivity threshold are selected. The corresponding input feature dimension of the state nodes with sensitivity exceeding the preset sensitivity threshold is located by tracing back along the computation graph. The gradient direction of the state nodes with sensitivity exceeding the preset sensitivity threshold is extracted and the gradient direction is reversed to generate the adversarial direction.

[0025] Construct a perturbation vector along the adversarial direction according to a preset perturbation intensity, inject the perturbation vector at the perturbation target dimension to generate test data, and add the test data to the test dataset.

[0026] A perturbation vector is constructed along the adversarial direction according to a preset perturbation intensity. The perturbation vector is then injected at the target dimension to generate test data, including:

[0027] Calculate the vector magnitude of the adversarial direction, and normalize the adversarial direction based on the vector magnitude to obtain a unit adversarial vector;

[0028] The basic perturbation vector is obtained by adjusting the amplitude of the unit adversarial vector based on the preset perturbation intensity. The original feature values ​​of the perturbation target dimension position in the seed sample are extracted. The upper limit and lower limit of the perturbation boundary are determined according to the value range of the original feature values. The basic perturbation vector is truncated based on the upper limit and lower limit of the perturbation boundary to obtain the perturbation vector.

[0029] Construct a zero vector with the same number of dimensions as the total number of feature dimensions of the seed sample as the injection vector template. Traverse the perturbation target dimension index and fill the perturbation component in the perturbation vector corresponding to the perturbation target dimension index into the corresponding index position in the injection vector template. After completing the traversal, the injection vector is obtained.

[0030] The injected vector is superimposed on the feature vector of the seed sample to obtain the perturbed feature vector. The perturbed feature vector is then used to replace the original feature vector of the seed sample to generate test data.

[0031] The test dataset is input into the large language model to obtain the validation text. The difference matrix between the validation text and the expected result is calculated. The difference matrix is ​​decomposed into directional and amplitude components. The initial prompt word template is updated based on the directional component to obtain the updated prompt word template. The perturbation intensity is adjusted based on the amplitude component, including:

[0032] The test data in the test dataset is concatenated with the initial prompt word template and then lexical encoded. The lexical encoding is then input into the encoding layer of the large language model to obtain the context representation. The context representation is then input into the decoding layer to generate the output lexical sequence. The output lexical sequence is then decoded to obtain the verification text.

[0033] The verification text and the expected result are lexicalized and semantic vectors are extracted to obtain the verification semantic vector and the expected semantic vector. The element difference between the verification semantic vector and the expected semantic vector is calculated to construct the difference matrix. The difference matrix is ​​decomposed into singular value to obtain the left singular vector matrix and the singular value vector. The first column vector of the left singular vector matrix is ​​extracted as the direction component, and the singular value vector is extracted as the amplitude component.

[0034] Calculate the cosine similarity between the directional component and the semantic vector of each word in the initial prompt word template, and select the word position corresponding to the maximum value as the insertion position. Take the opposite direction of the directional component to obtain the compensation vector. Calculate the cosine similarity between the compensation vector and the semantic vector of each word in the vocabulary, and select the word corresponding to the maximum value as the correction word. Insert the correction word at the insertion position to generate the updated prompt word template.

[0035] The average value of the amplitude components is calculated as the offset strength. The ratio of the offset strength to the preset reference strength is used as the adjustment coefficient to adjust the disturbance strength, thus obtaining the adjusted disturbance strength.

[0036] The process involves acquiring the operational data to be detected, converting it into semantic text based on an update prompt word template, and inputting it into a large language model to obtain anomaly detection results. Based on these results, the process locates the state node with the highest matching degree in the reverse state sequence as the root cause localization result, including:

[0037] The system acquires the operation and maintenance data to be detected and performs field parsing on the data to obtain a structured field set. The structured field set is then mapped and filled with named placeholders in the update prompt word template to form semantic text for detection.

[0038] The semantically encoded text is input into a large language model to generate an output word sequence. The output word sequence is then decoded to obtain the judgment text.

[0039] Semantic parsing is performed on the judgment text to extract anomaly identification information and anomaly feature description. Based on the anomaly identification information, it is determined whether there is anomaly in the operation and maintenance data to be detected and anomaly detection results are generated.

[0040] An anomaly semantic vector is obtained by semantic vector encoding of the anomaly feature description, and the state nodes in the reverse state sequence are traversed and the state description information of the state nodes is extracted.

[0041] Semantic vector encoding is performed on the state description information to obtain the state semantic vector. The cosine similarity between the anomaly semantic vector and each state semantic vector is calculated to obtain the similarity sequence. The state node with the largest value in the similarity sequence is selected as the root cause localization result.

[0042] A second aspect of this invention provides a system for detecting and analyzing operational anomalies based on a large model, comprising:

[0043] Build a configuration unit to construct a large language model including an encoding layer and a decoding layer and configure the initial prompt word template;

[0044] The data conversion unit is used to acquire historical operation and maintenance data and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template.

[0045] The seed labeling unit is used to input the training semantic text into the large language model to obtain the decision text, encode the decision text into semantic vectors and map it to the state space to obtain the projection coordinates, and label the seed samples according to the distribution and dispersion of the projection coordinates.

[0046] The perturbation generation unit is used to perform reverse reasoning on the seed sample to generate a reverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset.

[0047] The validation and adjustment unit is used to input the test dataset into the large language model to obtain the validation text, calculate the difference matrix between the validation text and the expected result, decompose the difference matrix into directional and amplitude components, update the initial prompt word template according to the directional component to obtain the updated prompt word template, and adjust the perturbation intensity according to the amplitude component.

[0048] The root cause localization unit is used to acquire the operation and maintenance data to be detected, convert the operation and maintenance data to be detected into semantic text based on the updated prompt word template, and input it into the large language model to obtain the judgment text. The state node with the highest matching degree in the reverse state sequence is located based on the judgment text as the root cause localization result.

[0049] A third aspect of the present invention provides an electronic device, comprising:

[0050] processor;

[0051] Memory used to store processor-executable instructions;

[0052] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0053] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0054] Anomaly detection and root cause analysis based on a large language model significantly improves the accuracy and automation of fault identification in operational scenarios. By converting historical operational data into semantic text and mapping it to the state space, the system can automatically label key seed samples from a large amount of data, effectively reducing the cost of manual annotation. Combining reverse reasoning and sensitivity analysis, a high-coverage test dataset is dynamically generated, allowing the model to expose potential edge scenarios during the training phase, significantly reducing the probability of missed and false alarms in actual operations. By calculating the sensitivity threshold of state nodes, the key links leading to anomalies are accurately identified. Combined with the directional and amplitude components of the difference matrix, the prompt word template and perturbation intensity can be adaptively adjusted, ensuring high robustness of the model in complex and dynamic operational environments. This avoids the excessive reliance on expert experience in traditional rule-based methods, significantly shortening the model iteration cycle and improving the generalization effect of anomaly detection. The collaborative optimization of updating the prompt word template and perturbation intensity allows the large language model to continuously adapt to the distribution drift of operational data. The directional component ensures that the prompt word instructions always focus on the most relevant anomaly features, while the amplitude component controls the reasonable boundary of perturbation injection, preventing excessive perturbation from causing model failure. This dynamic adjustment mechanism effectively solves the pain point of performance degradation caused by changes in data distribution in traditional methods, ensuring that the anomaly detection rate remains stable and high during long-term deployment, while reducing the frequency of manual parameter tuning and realizing the continuous evolution of intelligent operation and maintenance.

[0055] In the root cause localization phase, the operational data to be detected is matched with the reverse state sequence, directly identifying the state node with the highest matching degree as the root cause of the fault. This method understands the anomaly correlation at the semantic level, avoiding the limitations of traditional methods that rely on statistical correlation. It can accurately distinguish between causality and correlation, improving the root cause localization accuracy to an industrially usable level. Ultimately, it achieves end-to-end automation from anomaly discovery to root cause tracing, significantly shortening the mean time to repair faults and reducing operational costs and the risk of business interruption. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating the operation and maintenance anomaly detection and root cause analysis method based on a large model.

[0057] Figure 2 This is a flowchart of the operation process for anomaly detection and root cause localization of operation and maintenance data based on a large language model. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0060] Figure 1 This is a flowchart illustrating the operation and maintenance anomaly detection and root cause analysis method based on a large model according to an embodiment of the present invention.

[0061] Methods for operational anomaly detection and root cause analysis based on large models include:

[0062] Construct a large language model including encoding and decoding layers and configure initial prompt word templates;

[0063] Acquire historical operation and maintenance data, and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template;

[0064] The training semantic text is input into the large language model to obtain the decision text. The decision text is then encoded into semantic vectors and mapped to the state space to obtain the projection coordinates. Seed samples are labeled according to the distribution and dispersion of the projection coordinates.

[0065] Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset.

[0066] The test dataset is input into the large language model to obtain the verification text. The difference matrix between the verification text and the expected result is calculated. The difference matrix is ​​decomposed into directional components and amplitude components. The initial prompt word template is updated according to the directional components to obtain the updated prompt word template. The perturbation intensity is adjusted according to the amplitude components.

[0067] The system acquires the operation and maintenance data to be detected, converts the data into semantic text based on the update prompt word template, and inputs it into a large language model to obtain the judgment text. The system then locates the state node with the highest matching degree in the reverse state sequence based on the judgment text as the root cause localization result.

[0068] Obtaining historical operation and maintenance data, and converting the historical operation and maintenance data into training semantic text based on the initial prompt word template includes:

[0069] Collect log information, indicator information, alarm information, and topology information as historical operation and maintenance data;

[0070] Extract the trigger timestamp of the alarm information as the time alignment reference point, calculate the time offset of the log information and indicator information relative to the time alignment reference point, determine the time window boundary based on the distribution statistics of the time offset, extract the log information and indicator information within the time window boundary and combine them with the alarm information to form a time synchronization dataset.

[0071] Based on the topology information, a propagation path is constructed starting from the service node corresponding to the alarm information. Log information and indicator information of each node are extracted from the time-series synchronization dataset along the propagation path. The extracted log information, indicator information, alarm information from the time-series synchronization dataset, and propagation path are organized into a causal event chain according to the propagation path order.

[0072] Named placeholders are extracted from the initial prompt word template, and log information, indicator information, alarm information and propagation path in the causal event chain are filled into the corresponding named placeholder positions to generate training semantic text.

[0073] In real-world operation and maintenance scenarios, historical operation and maintenance data comes from a wide range of sources and varies in format. Before being directly used for training large language models, it requires systematic collection, alignment, and structuring. The collection phase focuses on four core data sources: log information from the runtime log files of each service node, including error stacks, request records, and state change events; metric information from the monitoring and collection system, covering time-series metrics such as CPU utilization, memory usage, network throughput, and response latency; alarm information from the alarm management platform, recording alarm levels, alarm rule names, and trigger times; and topology information from the service dependency graph of the configuration management database or service mesh, describing the call relationships and dependency levels between service nodes. These four types of data differ significantly in time granularity and structural form; therefore, a unified time-series alignment mechanism needs to be established before proceeding to subsequent processing.

[0074] Time-series alignment uses the trigger timestamp of the alarm information as the reference point. Alarm events in operational scenarios have clear time markers and are usually observable results in the anomaly propagation chain. Using their trigger timestamp as the time-series alignment reference point can anchor scattered logs and metric data to the same reference axis. Let the alarm trigger timestamp be... For a specific log record or a specific metric sampling point, its time offset Defined as the timestamp of the record and The difference, i.e. .when When, it indicates that the record occurred before the alarm was triggered, belonging to a precursor event; when When the alarm is triggered, it indicates that the record occurred after the alarm was triggered and belongs to a subsequent response event.

[0075] After collecting a large number of historical alarm samples, statistics were compiled from all samples. The distribution of values, for example, in calculation The percentile distribution is used, with the quantiles covering more than 95% of the effective information as the left boundary of the time window. and right boundary Specifically, take The 2.5 percentile of the distribution is used as Take the 97.5th percentile as This determines the boundaries of the time window. Within this window, the corresponding log and metric information is extracted and combined with the alarm information that triggered the window to form a time-series synchronization dataset. This process ensures the consistency of different types of data in the time dimension and avoids data misalignment problems caused by different sampling frequencies or clock drift.

[0076] After time-series alignment, topological information is used to further construct causal propagation paths. Topological information describes the dependencies between service nodes in the form of a directed graph, where nodes represent service instances and directed edges represent call or dependency directions. Starting from the service node corresponding to the alarm message, a path traversal is performed along the edges of the dependency graph to obtain a set of propagation paths reachable from that node. The traversal strategy for propagation paths can employ breadth-first search, expanding layer by layer according to the call hierarchy until a leaf node is reached or a preset maximum propagation depth is achieved. In cases where multiple propagation paths exist, all paths are retained to cover different fault propagation modes.

[0077] Along each propagation path, corresponding log and metric information is extracted node-by-node from the time-series synchronized dataset. Each node may have multiple log records and multiple metric sample values ​​within a time window; the complete time-series sequence is preserved during extraction rather than a single snapshot to capture dynamic trends in subsequent analysis. The extracted log information, metric information, alarm information, and the propagation path itself are arranged sequentially according to the node order of the propagation path, forming a causal event chain. The structure of the causal event chain reflects the spatiotemporal propagation logic of anomalies in the service topology: the starting end of the chain corresponds to the state of the alarm triggering node, the middle segment of the chain corresponds to the behavior records of each node along the propagation path within the time window, and the end of the chain corresponds to the state of the furthest propagation node. This ordered causal structure provides a clear contextual framework for the semantic understanding of large language models.

[0078] When converting the causal event chain into training semantic text, named placeholders are extracted from the initial prompt word template. Named placeholders are predefined named tags in the template, such as '{log_content}', '{metric_snapshot}', '{alert_description}', '{propagation_path}', etc., each placeholder corresponding to a type of information in the causal event chain. The filling process replaces placeholder names one by one according to the correspondence between them and the data categories in the causal event chain: the text content of the log information is filled into the position corresponding to '{log_content}', the numerical sequence of key metrics is filled into the position corresponding to '{metric_snapshot}' in structured text, the description field of the alarm information is filled into the position corresponding to '{alert_description}', and the node sequence of the propagation path is filled into the position corresponding to '{propagation_path}' in readable text.

[0079] For text-based processing of log information, it is necessary to clean the raw logs, remove irrelevant debugging information and duplicate entries, retain key log lines related to abnormal states, and append node identifiers and relative time offsets when filling in the data. This enables the model to perceive the location and temporal relationship of log records within the propagation path. For metric information, the time-series numerical sequences are converted into natural language descriptions, such as describing the peak, mean, and trend of a metric within a time window, helping the model understand the dynamic characteristics of the metric rather than just focusing on single-point values. The textualization of the propagation path converts the node sequence into a service call chain description connected by arrows or hierarchical descriptions, enabling the model to understand the direction and scope of fault propagation.

[0080] After the above filling steps, each causal event chain generates a corresponding training semantic text. This text fully includes the alarm triggering background, the running status of related services, abnormal indicators, and fault propagation path, forming a semantically coherent and structurally complete natural language description. It can be directly used as the training input of a large language model, enabling the model to learn the semantic patterns and causal reasoning capabilities of abnormal operation and maintenance scenarios during the training process.

[0081] The training semantic text is input into the large language model to obtain the decision text. The decision text is then semantically encoded and mapped to the state space to obtain the projected coordinates. Seed samples are labeled according to the distribution and dispersion of the projected coordinates, including:

[0082] The training semantic text is input into the large language model. The probability distribution of multiple candidate outputs is extracted from the decoding layer of the large language model. The probability distribution is sampled to generate multiple judgment texts. The multiple judgment texts are semantically encoded to obtain multiple judgment semantic vectors. The similarity distribution between the multiple judgment semantic vectors is calculated and statistically aggregated to obtain the judgment consistency score. The judgment semantic vector corresponding to the judgment text with the highest judgment consistency score is selected as the final judgment semantic vector.

[0083] The training semantic text is semantically encoded to obtain the input semantic vector. The difference between the final decision semantic vector and the input semantic vector is calculated as the decision offset vector. The decision offset vector is then mapped to the state space through dimensionality reduction transformation to obtain the projection coordinates.

[0084] Cluster analysis is performed on the projected coordinates to obtain cluster centers and cluster boundaries. The minimum distance from the projected coordinates to the cluster boundaries is calculated as the boundary distance, and the distance from the projected coordinates to the cluster centers is calculated as the center distance. The ratio of the boundary distance to the center distance is calculated as the distribution dispersion. The training semantic text corresponding to the projected coordinates with a distribution dispersion exceeding a preset dispersion threshold is marked as seed samples.

[0085] After inputting the trained semantically encoded text into the large language model, instead of directly taking a single output as the judgment result, the probability distribution of multiple candidate outputs is extracted from the decoding layer of the large language model. Specifically, the decoding layer maintains a probability distribution on the vocabulary at each generation step. By independently sampling this probability distribution multiple times, multiple judgment texts with different semantic content can be obtained. A temperature parameter can be introduced during the sampling process to soften the probability distribution, making the sampling results more diverse and thus covering the various semantic tendencies that the model may output under the input conditions. This multiple sampling strategy can effectively capture the output uncertainty of the large language model when facing the same input, providing a basis for subsequent consistency evaluation.

[0086] Semantic vector encoding is performed on each of the plurality of judgment texts described above to obtain a corresponding plurality of judgment semantic vectors. In the encoding process, a semantic encoder matched with a large language model is used to map each judgment text into a dense vector of fixed dimensions, thereby retaining the semantic information of the text. Subsequently, the cosine similarity between every two of the plurality of judgment semantic vectors is calculated to form a similarity matrix, and all similarity values in the matrix are statistically aggregated, for example, by taking the mean or median, to obtain a judgment consistency score. The judgment consistency score reflects the concentration degree of output semantics of the model under the current input: if the judgment texts obtained from multiple samplings are highly consistent semantically, the consistency score will be high, indicating that the model's judgment on this input is relatively certain; conversely, if the semantics of the sampling results are scattered, the consistency score will be low, indicating that the model has great uncertainty. The judgment semantic vector corresponding to the judgment text with the highest judgment consistency score is selected as the final judgment semantic vector, which is used for subsequent state space mapping.

[0087] Semantic vector encoding is also performed on the training semantic text per se to obtain an input semantic vector. The input semantic vector and the final judgment semantic vector are in the same semantic space, and both are dense vectors with equal dimensions. The vector difference between the final judgment semantic vector and the input semantic vector is calculated to obtain a judgment offset vector. The judgment offset vector characterizes the directional offset between the input semantics and output semantics of the large language model when processing the training sample, and its magnitude and direction jointly reflect the model's understanding and inference degree for the sample. If the magnitude of the judgment offset vector is large, it indicates that there is a large semantic gap between the output and input of the model for this sample, which may correspond to the model identifying anomaly or root cause information; if the magnitude is small, it indicates that the semantics of the output of the model are close to that of the input, corresponding to a normal sample or a sample that the model fails to fully distinguish.

[0088] The dimension of the judgment offset vector is usually high, and direct analysis in the original high-dimensional space has the problem of dimension disaster. Therefore, the judgment offset vector is mapped to a low-dimensional state space through dimensionality reduction transformation to obtain projection coordinates. Dimensionality reduction transformation can adopt methods such as principal component analysis or manifold learning to compress a high-dimensional offset vector into two-dimensional or three-dimensional coordinates, while retaining the geometric structure and distribution characteristics in the original vector space as much as possible. The distribution of projection coordinates in the state space directly reflects the differences of different training samples at the level of model semantic inference. Samples with similar operation and maintenance states tend to cluster in the state space, while samples with special states (such as anomalies or boundary states) tend to deviate from the main clustering area.

[0089] Cluster analysis is performed on the projected coordinates in the state space to obtain several cluster centers and corresponding cluster boundaries. Density-based clustering methods can be used to adapt to the sparse distribution of abnormal samples and dense distribution of normal samples in operational data. The cluster boundary is defined as the outer contour of the clustered region and can be determined by kernel density estimation or convex hull methods. For each projected coordinate, the minimum distance to its corresponding cluster boundary is calculated and denoted as the boundary distance. Simultaneously, calculate the distance from the node to its cluster center, denoted as the center distance. Boundary distance and center distance together describe the relative position of the projected coordinates within the cluster structure.

[0090] Distribution Dispersion Defined as the ratio of the boundary distance to the center distance, i.e. When the projected coordinates are located inside the cluster and close to the center, Smaller and Larger A larger value indicates that the sample is in the core region of the cluster, its semantic state is stable, and the model's judgment of it is highly consistent; when the projected coordinates are located at the cluster edge or in the transition region between clusters... Larger and Smaller A smaller value indicates that the sample is in a semantically ambiguous region, and the model's judgment of it is subject to greater uncertainty, thus possessing higher informational value. The distribution dispersion... Exceeding the preset discrete threshold The training semantic text labels corresponding to the projected coordinates are used as seed samples.

[0091] The selection logic for seed samples lies in selecting samples located in the core cluster region ( Larger samples correspond to typical operational states that the model can stably handle, while samples at the cluster edge or across cluster regions ( Smaller samples correspond to boundary states or anomalous states that the model has not yet fully learned. It is important to note that in the above determination... Exceeding the preset discrete threshold The meaning needs to be understood in conjunction with the specific definition of dispersion: if Defined as the ratio of center distance to boundary distance (i.e. ),but A larger threshold indicates that the sample is further away from the center and closer to the boundary. Samples exceeding this threshold are considered boundary region samples and are suitable as seed samples for subsequent perturbation testing. Regardless of the definition approach, the core principle is to select samples with representative boundary characteristics in the state space distribution to ensure that subsequent back-inference and perturbation injection can cover the weak areas of the model, thereby improving the robustness of operational anomaly detection and root cause analysis. Preset discrete threshold. The threshold can be adaptively set according to the overall distribution statistical characteristics of the training dataset. For example, a certain percentile of the dispersion of all sample distributions can be used as the threshold to ensure that the number of seed samples is within a reasonable range, while taking into account the coverage and computational efficiency of the subsequent test dataset generation.

[0092] Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes with sensitivity exceeding a preset sensitivity threshold to generate a test dataset, including:

[0093] The seed sample is input into the large language model and forward propagation is performed to obtain the decision output. A computation graph from the decision output to the seed sample input is constructed. The decision output is fixed and backpropagation is performed along the computation graph. The gradient activation state of each layer during the backpropagation process is recorded as a state node. The state nodes are organized according to the propagation order from the output layer to the input layer to form a reverse state sequence.

[0094] The gradient magnitude of each state node in the reverse state sequence is calculated and normalized to obtain the sensitivity of each state node. State nodes with sensitivity exceeding the preset sensitivity threshold are selected. The corresponding input feature dimension of the state nodes with sensitivity exceeding the preset sensitivity threshold is located by tracing back along the computation graph. The gradient direction of the state nodes with sensitivity exceeding the preset sensitivity threshold is extracted and the gradient direction is reversed to generate the adversarial direction.

[0095] Construct a perturbation vector along the adversarial direction according to a preset perturbation intensity, inject the perturbation vector at the perturbation target dimension to generate test data, and add the test data to the test dataset.

[0096] Performing backpropagation on seed samples to generate inverse state sequences is the core step in the entire test dataset construction process. The selected seed samples are fed into the large language model, and a complete forward propagation process is executed. The model sequentially passes through each sub-layer of the encoding and decoding layers, ultimately producing the corresponding decision output. After the forward propagation is complete, the intermediate computational processes are not immediately discarded; instead, the complete computational graph from seed sample input to decision output is retained. This computational graph records the dependencies between the input tensor, weight matrix, activation function, and output tensor of each layer, forming the basic data structure for subsequent backpropagation.

[0097] After fixing the decision output, the backpropagation process is initiated along the computation graph. Unlike the backpropagation in the standard training phase, the purpose of backpropagation here is not to update model parameters, but to track the gradient activation states exhibited by each layer as the signal propagates from the output layer to the input layer. At each step of backpropagation, the gradient tensor and its activation mode of the current layer are recorded and saved as a state node. The state node contains attribute information such as the layer number, the numerical distribution of the gradient tensor, and the sparsity of activation. All state nodes are arranged sequentially according to the propagation order from the output layer to the input layer, forming an inverse state sequence. In the inverse state sequence, the beginning of the sequence corresponds to the gradient state of the last layer of the model, and the end of the sequence corresponds to the gradient state of the embedding layer or feature extraction layer closest to the input. The entire sequence completely depicts the propagation path of the sensitivity of the decision output to the input.

[0098] After establishing the inverse state sequence, the gradient magnitude is calculated for each state node in the sequence. Let the gradient tensor at a certain state node be... ,in Let $\mathbf{ ... Norm, i.e. To ensure comparability of gradient magnitudes across different layers and orders of magnitude, the gradient magnitudes of all state nodes are normalized. The sensitivity of each state node is then obtained after normalization. The calculation method is as follows:

[0099]

[0100] in Traverse the indexes of all state nodes in the reverse state sequence. Normalized sensitivity. Reflects the first The relative contribution of each state node to the final output decision is indicated by the value. A larger value indicates a stronger gradient signal in that layer and a more significant impact on the output.

[0101] Sensitivity of each state node With preset sensitivity threshold Compare and filter those that meet the requirements. A set of state nodes. Preset sensitivity threshold. It can be configured according to the needs of the disturbance coverage in the actual operation and maintenance scenario. It is usually set to the upper quartile of the normalized sensitivity distribution to ensure that the number of highly sensitive nodes selected is moderate, which can cover the key feature dimensions without introducing too many redundant disturbance targets.

[0102] For each sensitivity exceeding the preset sensitivity threshold For each state node, a backtracking operation is performed along the computation graph to trace the gradient signal of that node back to its corresponding input feature dimension. The backtracking process utilizes the dependencies between layers in the computation graph to propagate the source information of the gradient upwards layer by layer, ultimately locating the dimensions in the seed sample input vector that contribute the most to the gradient of that state node. These dimensions are recorded as the perturbation target dimensions. The accuracy of locating the perturbation target dimensions directly determines the targeting of subsequent perturbations. By accurately locating the input features in highly sensitive regions, the generated test data can produce the greatest perturbation effect on the model's decision output with minimal modification, thereby effectively verifying the model's robustness to boundary samples.

[0103] After determining the dimensions of the perturbation target, extract each sensitivity exceeding a preset sensitivity threshold. The gradient direction at the state node. The gradient direction is a unit vector. This indicates that the gradient tensor is about to be used. Divide it The norm is obtained. Since the gradient direction points in the direction that increases the decision output, in order to generate adversarial test data, the gradient direction needs to be reversed to obtain the adversarial direction. The adversarial direction points in the direction that reduces the decision output. Applying perturbation along this direction can maximally change the model's decision result, enabling the test data to challenge the model's decision boundary. When there are multiple state nodes in the reverse state sequence with sensitivity exceeding the threshold, the adversarial direction corresponding to each node is extracted, and perturbation is applied to the corresponding perturbation target dimension to cover multiple potential root cause feature dimensions.

[0104] According to the preset disturbance intensity along the confrontation direction Construct the perturbation vector. The construction method is to perturb the target dimension along the adversarial direction. move Units, i.e. Preset disturbance intensity The deviation between the test data and the original seed sample was controlled. Its initial value was pre-configured before the method ran and dynamically adjusted subsequently based on the magnitude components of the difference matrix. The perturbation vector... A test data point is obtained by injecting the perturbation target dimension position of the seed sample input vector. If there are multiple state nodes with sensitivity exceeding the threshold, a test data point is generated for each node. Alternatively, multiple perturbation vectors can be superimposed and injected on their respective target dimensions to generate test data with combined perturbation, in order to simulate multi-dimensional abnormal concurrent operation and maintenance scenarios.

[0105] The generated test data is organized in the form of semantic text, maintaining the same format and specifications as the training semantic text, ensuring it can be directly input into the large language model for subsequent validation and inference. All generated test data is added to the test dataset, which includes perturbation samples from different seed samples and covers the feature dimensions corresponding to high-sensitivity nodes at different levels in the inverse state sequence. This allows for a comprehensive multi-granularity and multi-level evaluation of the large language model's judgment capabilities. The test dataset constructed in this way can effectively simulate various boundary anomaly scenarios in real-world operational environments, providing sufficient and targeted validation basis for subsequent updates and optimizations of prompt word templates.

[0106] A perturbation vector is constructed along the adversarial direction according to a preset perturbation intensity. The perturbation vector is then injected at the target dimension to generate test data, including:

[0107] Calculate the vector magnitude of the adversarial direction, and normalize the adversarial direction based on the vector magnitude to obtain a unit adversarial vector;

[0108] The basic perturbation vector is obtained by adjusting the amplitude of the unit adversarial vector based on the preset perturbation intensity. The original feature values ​​of the perturbation target dimension position in the seed sample are extracted. The upper limit and lower limit of the perturbation boundary are determined according to the value range of the original feature values. The basic perturbation vector is truncated based on the upper limit and lower limit of the perturbation boundary to obtain the perturbation vector.

[0109] Construct a zero vector with the same number of dimensions as the total number of feature dimensions of the seed sample as the injection vector template. Traverse the perturbation target dimension index and fill the perturbation component in the perturbation vector corresponding to the perturbation target dimension index into the corresponding index position in the injection vector template. After completing the traversal, the injection vector is obtained.

[0110] The injected vector is superimposed on the feature vector of the seed sample to obtain the perturbed feature vector. The perturbed feature vector is then used to replace the original feature vector of the seed sample to generate test data.

[0111] After obtaining the adversarial directions of each state node, they need to be transformed into specific perturbation vectors that can be injected with seed samples, and finally, test data is generated to verify the robustness of the large language model. The adversarial direction itself is a vector with arbitrary magnitude; directly using it as a perturbation vector will lead to uncontrollable perturbation strength. Therefore, it is necessary to normalize the adversarial direction first. The magnitude of the adversarial direction vector is then calculated. That is, taking the square root of the sum of the squares of the components of the vector to obtain its Euclidean norm. Then, the adversarial direction vector is divided by... , obtain the unit adversarial vector Its modulus is strictly equal to 1, retaining only directional information while eliminating amplitude influence. This normalization step ensures that subsequent operations based on the preset perturbation strength... When adjusting the amplitude, the perturbation intensity between different state nodes and different feature dimensions has a unified dimensional benchmark, making the perturbation experiment results comparable.

[0112] Based on preset disturbance intensity against unit adversarial vector Amplitude adjustment is performed to obtain the basic disturbance vector. .at this time The modulus is exactly equal to The direction is consistent with the direction of confrontation. However, directly... Injecting perturbation into the feature space of seed samples may cause the perturbed feature values ​​to exceed reasonable physical or business limits. For example, CPU utilization exceeding 100%, negative memory usage, or extreme outliers in response latency may occur. These out-of-bounds feature values ​​can cause the test data to lose its semantic rationality in real-world operational scenarios, thus affecting the validation conclusions of the large language model. Therefore, it is necessary to extract the original feature values ​​at the perturbation target dimension positions in the seed samples. The upper limit of the disturbance boundary is determined based on the actual value range of this dimension in historical operation and maintenance data. Lower limit of the disturbance boundary The upper and lower limits of the perturbation boundaries can be obtained from the statistical distribution of historical data, such as taking the maximum and minimum values ​​of that dimension in historical data, or determined based on the mean plus or minus a certain number of standard deviations. Alternatively, hard boundaries can be directly specified in conjunction with business rules (such as percentage indicators between 0 and 100). After determining the boundaries, the basic perturbation vector... Each component corresponding to the perturbation target dimension is truncated one by one: if the original eigenvalues The result of adding this component exceeds Then the component is truncated to If the result is lower than Then the component is truncated to Otherwise, the original component values ​​are retained unchanged. The perturbation vector is obtained after truncation. Each component satisfies the constraint that the feature value does not exceed the bound after injection.

[0113] After obtaining the perturbation vector Next, it needs to be precisely injected into the specified dimension position of the seed sample feature vector without affecting the original values ​​of other dimensions. This involves constructing a vector with a number of dimensions equal to the total number of seed sample feature dimensions. Same zero vector As an injection vector template, all components of this zero vector are initialized to 0, representing that no perturbation is applied to any dimension by default. The set of perturbation target dimension indices is denoted as... This includes the dimension numbers that need to be injected with perturbations, as determined during the sensitivity analysis phase. Traversal Each dimension index in ,Will The perturbation component corresponding to this index Fill into the injection vector template The There are several positions. After the traversal is complete, In this case, there is a non-zero value only at the position of the perturbation target dimension, while the values ​​at other positions remain 0. This is the complete injection vector. This design strictly limits the scope of the perturbation to the target dimension, avoiding the introduction of unexpected numerical shifts into non-target dimensions, ensuring the semantic integrity of the test data in the non-perturbation dimensions, and also facilitating the accurate tracing of the independent contribution of each dimension's perturbation to the model output during subsequent analysis.

[0114] Inject vector Compared with the original feature vector of the seed sample By adding elements one by one, we obtain the perturbed feature vector. Since the components of the injected vector in the non-target dimensions are all 0, this addition operation is equivalent to replacing the original feature values ​​only in the target dimension. The remaining dimensions remain unchanged. Replace the original feature vector of the seed sample while retaining other metadata information (such as timestamp, service identifier, alarm type label, etc.) to generate a complete test data record. Repeat the above perturbation injection process for all seed samples, and for each seed sample, different perturbation intensities can be applied. Or different perturbation target dimension index sets Multiple test data points are generated and finally aggregated to form a test dataset covering various anomaly patterns and perturbation levels.

[0115] Each record in the test dataset carries information about the known source of the perturbation, including the seed sample number, the index of the perturbation target dimension, the actual injected perturbation component value, and the corresponding state node position index. This metadata plays a crucial role in the subsequent calculation of the difference matrix between the verification text and the expected result, ensuring that the decomposition result of the difference matrix accurately corresponds to the specific perturbation dimension and state node. This supports the targeted updating of the directional components of the prompt word template and the adjustment of the perturbation intensity. The adaptive adjustment is achieved through a combination of truncation constraints and zero vector injection templates. This process maintains the semantic rationality of the test data while enabling fine-grained control over the location and magnitude of the perturbations, providing high-quality adversarial test samples for evaluating the robustness of large language models in operational anomaly detection scenarios.

[0116] The test dataset is input into the large language model to obtain the validation text. The difference matrix between the validation text and the expected result is calculated. The difference matrix is ​​decomposed into directional and amplitude components. The initial prompt word template is updated based on the directional component to obtain the updated prompt word template. The perturbation intensity is adjusted based on the amplitude component, including:

[0117] The test data in the test dataset is concatenated with the initial prompt word template and then lexical encoded. The lexical encoding is then input into the encoding layer of the large language model to obtain the context representation. The context representation is then input into the decoding layer to generate the output lexical sequence. The output lexical sequence is then decoded to obtain the verification text.

[0118] The verification text and the expected result are lexicalized and semantic vectors are extracted to obtain the verification semantic vector and the expected semantic vector. The element difference between the verification semantic vector and the expected semantic vector is calculated to construct the difference matrix. The difference matrix is ​​decomposed into singular value to obtain the left singular vector matrix and the singular value vector. The first column vector of the left singular vector matrix is ​​extracted as the direction component, and the singular value vector is extracted as the amplitude component.

[0119] Calculate the cosine similarity between the directional component and the semantic vector of each word in the initial prompt word template, and select the word position corresponding to the maximum value as the insertion position. Take the opposite direction of the directional component to obtain the compensation vector. Calculate the cosine similarity between the compensation vector and the semantic vector of each word in the vocabulary, and select the word corresponding to the maximum value as the correction word. Insert the correction word at the insertion position to generate the updated prompt word template.

[0120] The average value of the amplitude components is calculated as the offset strength. The ratio of the offset strength to the preset reference strength is used as the adjustment coefficient to adjust the disturbance strength, thus obtaining the adjusted disturbance strength.

[0121] Each test data point in the test dataset is concatenated with the initial prompt word template to form a complete input sequence. During concatenation, the initial prompt word template serves as a prefix, with the test data appended to it, creating a composite input with contextual guidance information. Lexical encoding is performed on the concatenated input sequence, which involves segmenting the character sequence into discrete lexical units according to the vocabulary mapping rules and converting them into corresponding embedding vectors. The embedding vector sequence is fed into the encoding layer of the large language model. The encoding layer models the dependencies between lexical units in the input sequence using a multi-head self-attention mechanism, outputting a context representation matrix carrying global contextual information. This context representation matrix is ​​then passed to the decoding layer, which generates the output lexical sequence step by step in an autoregressive manner. Each generation step is based on the already generated lexical units and contextual representations, determining the next lexical unit through probability distribution sampling or greedy decoding. Inverse lexicalization is performed on the final output lexical sequence, restoring the discrete lexical units to readable natural language text, which is the verification text.

[0122] After obtaining the verification text, both the verification text and the expected result are lexicalized, converting them into word sequences. A semantic encoder then extracts their corresponding semantic vectors, denoted as the verification semantic vector and the expected semantic vector, respectively. The dimension of the semantic vectors matches the hidden layer dimension of the large language model, enabling the capture of the semantic distribution features of the text in a high-dimensional space. The element-wise difference between the verification and expected semantic vectors is calculated by subtracting each vector along each dimension. The resulting differences are then arranged by row or column to construct a difference matrix. Difference Matrix It comprehensively reflects the distribution of deviations between the verification text and the expected results in the semantic space, where each element corresponds to the amount of deviation in a specific semantic dimension.

[0123] For the difference matrix Perform singular value decomposition to decompose it into a left singular vector matrix. Singular value diagonal matrix and right singular vector matrix ,satisfy Singular value decomposition (SVD) can separate the deviation information contained in the difference matrix into two orthogonal components: direction information and amplitude information. Extracting the left singular vector matrix... The first column vector As a directional component The direction of the principal deviation with the highest energy concentration in the difference matrix represents the direction in which the verification text deviates most significantly from the expected result in the semantic space. Singular value vectors are extracted. (i.e., singular value diagonal matrix) The vector consisting of the main diagonal elements is used as the amplitude component. The magnitude of each element in the matrix reflects the deviation of the difference matrix in the corresponding direction.

[0124] In obtaining directional components Then, calculate The cosine similarity between the semantic vector of each lexical unit in the initial prompt word template and the current deviation direction is calculated. A higher cosine similarity indicates that the semantic direction of that lexical unit is closer to the current deviation direction, meaning that the prompt word content at that position contributes most significantly to the model's output deviation. The lexical unit position corresponding to the maximum cosine similarity is selected as the insertion position. This position is the target position for subsequent corrected lexical insertion. (Regarding the directional component) Reverse the direction to obtain the compensation vector. The compensation vector points in the semantic space in the opposite direction to the main bias direction, representing the semantic compensation direction that needs to be introduced by inserting lexical units. The compensation vector is calculated as follows: The word with the highest cosine similarity to the semantic vectors of all words in the vocabulary is selected as the corrected word. This lexical element is closest to the compensation vector in the semantic space, and can most effectively correct semantic deviations in the prompt word template. At the insertion position... The word will be corrected. Insert the initial prompt word template to generate an updated prompt word template. The updated prompt word template introduces semantic compensation lexical units at key positions, enabling the large language model to more accurately guide the output toward the expected result during subsequent inference, thereby reducing the semantic deviation between the validation text and the expected result.

[0125] In amplitude component Based on this, calculate The offset strength is obtained by taking the arithmetic mean of all elements in the matrix. The offset intensity quantifies the overall deviation magnitude of the difference matrix in each singular direction; a larger value indicates a more significant overall deviation between the verification text and the expected result. The offset intensity... With preset benchmark strength The ratio is defined as the adjustment coefficient. ,Right now Adjustment coefficient This reflects the ratio of the current deviation magnitude to the benchmark level. When this occurs, it indicates that the deviation of the current verification text exceeds the baseline level, and the perturbation strength needs to be appropriately increased to enhance the diversity coverage of the test data; when When the deviation is below the baseline level, the disturbance intensity can be appropriately reduced to avoid introducing excessive noise that could interfere with the model's normal inference. (Adjustment coefficient) For the current preset disturbance intensity Perform a product adjustment to obtain the adjusted disturbance strength. The adjusted perturbation strength will replace the original preset perturbation strength in the calculation of the perturbation vector during the subsequent test dataset generation stage. This allows the perturbation injection process to adaptively respond to changes in model output deviation, thereby ensuring the validity of the test data while avoiding a decrease in test quality caused by excessively large or small perturbation strength.

[0126] The above process uses the directional and amplitude components of the difference matrix for two independent feedback channels: prompt word template optimization and perturbation intensity adaptive adjustment. This achieves synergistic optimization of the model's semantic guidance capability and test data generation capability, effectively improving the robustness and generalization ability of the large language model in the scenario of operation and maintenance anomaly detection.

[0127] like Figure 2 As shown, Figure 2This is a flowchart of the operation and maintenance data anomaly detection and root cause localization based on a large language model according to an embodiment of the present invention.

[0128] The process involves acquiring the operational data to be detected, converting it into semantic text based on an update prompt word template, and inputting it into a large language model to obtain anomaly detection results. Based on these results, the process locates the state node with the highest matching degree in the reverse state sequence as the root cause localization result, including:

[0129] The system acquires the operation and maintenance data to be detected and performs field parsing on the data to obtain a structured field set. The structured field set is then mapped and filled with named placeholders in the update prompt word template to form semantic text for detection.

[0130] The semantically encoded text is input into a large language model to generate an output word sequence. The output word sequence is then decoded to obtain the judgment text.

[0131] Semantic parsing is performed on the judgment text to extract anomaly identification information and anomaly feature description. Based on the anomaly identification information, it is determined whether there is anomaly in the operation and maintenance data to be detected and anomaly detection results are generated.

[0132] An anomaly semantic vector is obtained by semantic vector encoding of the anomaly feature description, and the state nodes in the reverse state sequence are traversed and the state description information of the state nodes is extracted.

[0133] Semantic vector encoding is performed on the state description information to obtain the state semantic vector. The cosine similarity between the anomaly semantic vector and each state semantic vector is calculated to obtain the similarity sequence. The state node with the largest value in the similarity sequence is selected as the root cause localization result.

[0134] After acquiring the operational data to be tested, the first step is to perform field parsing. This data typically includes various types such as log text, time-series metrics, and alarm records, each with a different field structure. The field parsing process splits the raw data according to a predefined field pattern, extracting key fields such as service name, host identifier, error code, metric value, and timestamp to form a structured field set. For log data, fields are extracted using regular expression matching or delimiter segmentation; for metric data, labels and values ​​for each dimension are parsed according to key-value pair format; and for alarm data, fields such as alarm level, triggering conditions, and impact scope are extracted according to alarm protocol specifications. Each field in the structured field set is stored as a key-value pair to ensure accurate matching in subsequent mapping and population steps.

[0135] The update prompt template predefines several named placeholders, each corresponding to a specific field in the operational data to be detected. When mapping and filling the structured field set with the named placeholders, the field value is replaced in the placeholder position according to the correspondence between the placeholder name and the field key name, thus transforming the original operational data into semantically complete detection text. If a named placeholder does not have a corresponding field in the structured field set, it is filled with a preset default descriptor to avoid gaps in the semantic text that could lead to ambiguity in the large language model. The completed detection semantic text combines the specific numerical information of the operational data with the semantic structure provided by the prompt template, providing a formatted and semantically complete input for subsequent large language model inference.

[0136] When performing semantic encoding on the detected text, the text is segmented according to the vocabulary of a large language model. Each word is converted into a corresponding word index sequence, and positional encoding information is appended before being input into the encoding layer of the large language model. The large language model sequentially extracts features from the word sequence through multiple attention mechanisms and feedforward networks. Finally, the decoding layer generates the output word sequence step by step in an autoregressive manner. At each generation step, the decoding layer selects the next word based on the generated word sequence and the context representation output by the encoding layer, using probability distribution sampling or a greedy decoding strategy, until a terminator is generated or the maximum generation length is reached. The output word sequence is inversely mapped and decoded according to the vocabulary to restore the word index sequence to readable text, which is the judgment text. The judgment text contains the comprehensive analysis conclusions of the large language model on the operational data to be detected, covering anomaly judgment opinions, anomaly type descriptions, and possible fault feature descriptions.

[0137] When performing semantic parsing on the judgment text, two key types of content are extracted from the judgment text according to pre-defined structured parsing rules: anomaly identification information and anomaly feature description. Anomaly identification information is a Boolean or enumerated marker used to characterize whether the operational data to be detected is abnormal and the severity level of the anomaly, such as "normal," "minor anomaly," or "serious anomaly." The semantic parsing process extracts anomaly identification information from specific locations in the judgment text through keyword matching or slot extraction based on predefined templates. Based on the anomaly identification information, it is determined whether the operational data to be detected is abnormal. If the anomaly identification information indicates "normal," a detection result of no anomaly is generated, and the subsequent root cause localization process is terminated; if the anomaly identification information indicates the presence of an anomaly, the anomaly feature description is further extracted, and the root cause localization step continues. The anomaly feature description contains natural language descriptions of the abnormal phenomenon, such as "memory usage continues to climb and triggers OOM" or "database connection pool exhaustion causes request timeouts." These descriptions carry the core semantic features of the anomaly and are the key basis for subsequent semantic vector encoding and state node matching.

[0138] When performing semantic vector encoding on anomaly feature descriptions, the anomaly feature description text is input into the semantic encoder, and after multiple transformations, the corresponding anomaly semantic vector is generated in a high-dimensional vector space. The semantic encoder uses the same vector space definition as the encoding layer of the large language model, ensuring that the abnormal semantic vector and the subsequent state semantic vector are in the same metric space, thus guaranteeing the effectiveness of cosine similarity calculation. When traversing each state node in the reverse state sequence, the pre-stored state description information is extracted from each state node. This state description information is the result of textualizing the semantic content of each state node during the reverse reasoning stage to generate the reverse state sequence. The description covers elements such as the fault mode, triggering conditions, and scope of influence corresponding to that state node.

[0139] Semantic vector encoding is performed on the state description information of each state node to obtain the state semantic vector corresponding to each state node. ,in This is the index of the state node in the reverse state sequence. Calculate the anomaly semantic vector. With each state semantic vector Cosine similarity between The calculation method is to divide the inner product of the two vectors by the product of their respective magnitudes, that is... After traversing all state nodes, a similarity sequence is obtained, consisting of the cosine similarity of each state node. The range of cosine similarity values ​​is... The closer the value is to 1, the more consistent the direction of the abnormal semantic vector and the corresponding state semantic vector are in the semantic space, that is, the more similar the fault semantics described by the two.

[0140] The state node corresponding to the highest cosine similarity score in the similarity sequence is selected and output as the root cause localization result. The position of this state node in the reverse state sequence reflects the origin of the fault in the system's state evolution chain, and its state description information is the most explanatory candidate root cause description for the current anomaly. In actual operation and maintenance scenarios, if the difference between the maximum and second-largest similarity values ​​in the similarity sequence is lower than a preset discrimination threshold, several state nodes with the highest similarity ranking can be output simultaneously as a set of candidate root causes for further manual evaluation by operations and maintenance personnel, avoiding misjudgments due to insufficient confidence in a single localization result. Finally, the root cause localization result and the anomaly detection result are packaged together into a structured output, including the anomaly type, root cause node identifier, root cause description text, and similarity confidence score, providing a complete decision-making basis for subsequent automated alarm handling or manual intervention.

[0141] A second aspect of this invention provides a system for detecting and analyzing operational anomalies based on a large model, comprising:

[0142] Build a configuration unit to construct a large language model including an encoding layer and a decoding layer and configure the initial prompt word template;

[0143] The data conversion unit is used to acquire historical operation and maintenance data and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template.

[0144] The seed labeling unit is used to input the training semantic text into the large language model to obtain the decision text, encode the decision text into semantic vectors and map it to the state space to obtain the projection coordinates, and label the seed samples according to the distribution and dispersion of the projection coordinates.

[0145] The perturbation generation unit is used to perform reverse reasoning on the seed sample to generate a reverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset.

[0146] The validation and adjustment unit is used to input the test dataset into the large language model to obtain the validation text, calculate the difference matrix between the validation text and the expected result, decompose the difference matrix into directional and amplitude components, update the initial prompt word template according to the directional component to obtain the updated prompt word template, and adjust the perturbation intensity according to the amplitude component.

[0147] The root cause localization unit is used to acquire the operation and maintenance data to be detected, convert the operation and maintenance data to be detected into semantic text based on the updated prompt word template, and input it into the large language model to obtain the judgment text. The state node with the highest matching degree in the reverse state sequence is located based on the judgment text as the root cause localization result.

[0148] A third aspect of the present invention provides an electronic device, comprising:

[0149] processor;

[0150] Memory used to store processor-executable instructions;

[0151] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0152] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0153] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for operation and maintenance anomaly detection and root cause analysis based on a large language model, characterized in that, include: Construct a large language model including encoding and decoding layers and configure initial prompt word templates; Acquire historical operation and maintenance data, and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template; The training semantic text is input into the large language model to obtain the decision text. The decision text is then encoded into semantic vectors and mapped to the state space to obtain the projection coordinates. Seed samples are labeled according to the distribution and dispersion of the projection coordinates. Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset. The test dataset is input into the large language model to obtain the verification text. The difference matrix between the verification text and the expected result is calculated. The difference matrix is ​​decomposed into directional components and amplitude components. The initial prompt word template is updated according to the directional components to obtain the updated prompt word template. The perturbation intensity is adjusted according to the amplitude components. The system acquires the operation and maintenance data to be detected, converts the data into semantic text based on the update prompt word template, and inputs it into a large language model to obtain the judgment text. The system then locates the state node with the highest matching degree in the reverse state sequence based on the judgment text as the root cause localization result.

2. The method according to claim 1, characterized in that, Obtaining historical operation and maintenance data, and converting the historical operation and maintenance data into training semantic text based on the initial prompt word template includes: Collect log information, indicator information, alarm information, and topology information as historical operation and maintenance data; Extract the trigger timestamp of the alarm information as the time alignment reference point, calculate the time offset of the log information and indicator information relative to the time alignment reference point, determine the time window boundary based on the distribution statistics of the time offset, extract the log information and indicator information within the time window boundary and combine them with the alarm information to form a time synchronization dataset. Based on the topology information, a propagation path is constructed starting from the service node corresponding to the alarm information. Log information and indicator information of each node are extracted from the time-series synchronization dataset along the propagation path. The extracted log information, indicator information, alarm information from the time-series synchronization dataset, and propagation path are organized into a causal event chain according to the propagation path order. Named placeholders are extracted from the initial prompt word template, and log information, indicator information, alarm information and propagation path in the causal event chain are filled into the corresponding named placeholder positions to generate training semantic text.

3. The method according to claim 1, characterized in that, The training semantic text is input into the large language model to obtain the decision text. The decision text is then semantically encoded and mapped to the state space to obtain the projected coordinates. Seed samples are labeled according to the distribution and dispersion of the projected coordinates, including: The training semantic text is input into the large language model. The probability distribution of multiple candidate outputs is extracted from the decoding layer of the large language model. The probability distribution is sampled to generate multiple judgment texts. The multiple judgment texts are semantically encoded to obtain multiple judgment semantic vectors. The similarity distribution between the multiple judgment semantic vectors is calculated and statistically aggregated to obtain the judgment consistency score. The judgment semantic vector corresponding to the judgment text with the highest judgment consistency score is selected as the final judgment semantic vector. The training semantic text is semantically encoded to obtain the input semantic vector. The difference between the final decision semantic vector and the input semantic vector is calculated as the decision offset vector. The decision offset vector is then mapped to the state space through dimensionality reduction transformation to obtain the projection coordinates. Cluster analysis is performed on the projected coordinates to obtain cluster centers and cluster boundaries. The minimum distance from the projected coordinates to the cluster boundaries is calculated as the boundary distance, and the distance from the projected coordinates to the cluster centers is calculated as the center distance. The ratio of the boundary distance to the center distance is calculated as the distribution dispersion. The training semantic text corresponding to the projected coordinates with a distribution dispersion exceeding a preset dispersion threshold is marked as seed samples.

4. The method according to claim 1, characterized in that, Perform reverse reasoning on the seed sample to generate an inverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes with sensitivity exceeding a preset sensitivity threshold to generate a test dataset, including: The seed sample is input into the large language model and forward propagation is performed to obtain the decision output. A computation graph from the decision output to the seed sample input is constructed. The decision output is fixed and backpropagation is performed along the computation graph. The gradient activation state of each layer during the backpropagation process is recorded as a state node. The state nodes are organized according to the propagation order from the output layer to the input layer to form a reverse state sequence. The gradient magnitude of each state node in the reverse state sequence is calculated and normalized to obtain the sensitivity of each state node. State nodes with sensitivity exceeding the preset sensitivity threshold are selected. The corresponding input feature dimension of the state nodes with sensitivity exceeding the preset sensitivity threshold is located by tracing back along the computation graph. The gradient direction of the state nodes with sensitivity exceeding the preset sensitivity threshold is extracted and the gradient direction is reversed to generate the adversarial direction. Construct a perturbation vector along the adversarial direction according to a preset perturbation intensity, inject the perturbation vector at the perturbation target dimension to generate test data, and add the test data to the test dataset.

5. The method according to claim 4, characterized in that, A perturbation vector is constructed along the adversarial direction according to a preset perturbation intensity. The perturbation vector is then injected at the target dimension to generate test data, including: Calculate the vector magnitude of the adversarial direction, and normalize the adversarial direction based on the vector magnitude to obtain a unit adversarial vector; The basic perturbation vector is obtained by adjusting the amplitude of the unit adversarial vector based on the preset perturbation intensity. The original feature values ​​of the perturbation target dimension position in the seed sample are extracted. The upper limit and lower limit of the perturbation boundary are determined according to the value range of the original feature values. The basic perturbation vector is truncated based on the upper limit and lower limit of the perturbation boundary to obtain the perturbation vector. Construct a zero vector with the same number of dimensions as the total number of feature dimensions of the seed sample as the injection vector template. Traverse the perturbation target dimension index and fill the perturbation component in the perturbation vector corresponding to the perturbation target dimension index into the corresponding index position in the injection vector template. After completing the traversal, the injection vector is obtained. The injected vector is superimposed on the feature vector of the seed sample to obtain the perturbed feature vector. The perturbed feature vector is then used to replace the original feature vector of the seed sample to generate test data.

6. The method according to claim 1, characterized in that, The test dataset is input into the large language model to obtain the validation text. The difference matrix between the validation text and the expected result is calculated. The difference matrix is ​​decomposed into directional and amplitude components. The initial prompt word template is updated based on the directional component to obtain the updated prompt word template. The perturbation intensity is adjusted based on the amplitude component, including: The test data in the test dataset is concatenated with the initial prompt word template and then lexical encoded. The lexical encoding is then input into the encoding layer of the large language model to obtain the context representation. The context representation is then input into the decoding layer to generate the output lexical sequence. The output lexical sequence is then decoded to obtain the verification text. The verification text and the expected result are lexicalized and semantic vectors are extracted to obtain the verification semantic vector and the expected semantic vector. The element difference between the verification semantic vector and the expected semantic vector is calculated to construct the difference matrix. The difference matrix is ​​decomposed into singular value to obtain the left singular vector matrix and the singular value vector. The first column vector of the left singular vector matrix is ​​extracted as the direction component, and the singular value vector is extracted as the amplitude component. Calculate the cosine similarity between the directional component and the semantic vector of each word in the initial prompt word template, and select the word position corresponding to the maximum value as the insertion position. Take the opposite direction of the directional component to obtain the compensation vector. Calculate the cosine similarity between the compensation vector and the semantic vector of each word in the vocabulary, and select the word corresponding to the maximum value as the correction word. Insert the correction word at the insertion position to generate the updated prompt word template. The average value of the amplitude components is calculated as the offset strength. The ratio of the offset strength to the preset reference strength is used as the adjustment coefficient to adjust the disturbance strength, thus obtaining the adjusted disturbance strength.

7. The method according to claim 1, characterized in that, The process involves acquiring the operational data to be detected, converting it into semantic text based on an update prompt word template, and inputting it into a large language model to obtain anomaly detection results. Based on these results, the process locates the state node with the highest matching degree in the reverse state sequence as the root cause localization result, including: The system acquires the operation and maintenance data to be detected and performs field parsing on the data to obtain a structured field set. The structured field set is then mapped and filled with named placeholders in the update prompt word template to form semantic text for detection. The semantically encoded text is input into a large language model to generate an output word sequence. The output word sequence is then decoded to obtain the judgment text. Semantic parsing is performed on the judgment text to extract anomaly identification information and anomaly feature description. Based on the anomaly identification information, it is determined whether there is anomaly in the operation and maintenance data to be detected and anomaly detection results are generated. An anomaly semantic vector is obtained by semantic vector encoding of the anomaly feature description, and the state nodes in the reverse state sequence are traversed and the state description information of the state nodes is extracted. Semantic vector encoding is performed on the state description information to obtain the state semantic vector. The cosine similarity between the anomaly semantic vector and each state semantic vector is calculated to obtain the similarity sequence. The state node with the largest value in the similarity sequence is selected as the root cause localization result.

8. A system for detecting operational anomalies and analyzing root causes based on a large language model, used to implement the method as described in any one of claims 1-7, characterized in that, include: Build a configuration unit to construct a large language model including an encoding layer and a decoding layer and configure the initial prompt word template; The data conversion unit is used to acquire historical operation and maintenance data and convert the historical operation and maintenance data into training semantic text based on the initial prompt word template. The seed labeling unit is used to input the training semantic text into the large language model to obtain the decision text, encode the decision text into semantic vectors and map it to the state space to obtain the projection coordinates, and label the seed samples according to the distribution and dispersion of the projection coordinates. The perturbation generation unit is used to perform reverse reasoning on the seed sample to generate a reverse state sequence, calculate the sensitivity of each state node in the state sequence, and inject perturbations into state nodes whose sensitivity exceeds a preset sensitivity threshold to generate a test dataset. The validation and adjustment unit is used to input the test dataset into the large language model to obtain the validation text, calculate the difference matrix between the validation text and the expected result, decompose the difference matrix into directional and amplitude components, update the initial prompt word template according to the directional component to obtain the updated prompt word template, and adjust the perturbation intensity according to the amplitude component. The root cause localization unit is used to acquire the operation and maintenance data to be detected, convert the operation and maintenance data to be detected into semantic text based on the updated prompt word template, and input it into the large language model to obtain the judgment text. The state node with the highest matching degree in the reverse state sequence is located based on the judgment text as the root cause localization result.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Ophthalmic prognosis visual simulation system and training method thereof

    CN121983328A

  • System and method for agentic artificial intelligence based root cause analysis in hybrid distributed systems

    US20260140812A1