Interface abnormality analysis method based on time log fingerprint and multi-modal fusion

CN122845401APending Publication Date: 2026-09-29NANJING BESTLINK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611209010.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-11
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

对于固定阈值法或者关键词/规则匹配法来说,指标监控与日志检测相互独立,难以捕捉故障发生前指标与日志的协同偏移,既易漏检无错误日志的隐性亚健康异常,也易将业务高峰、瞬时抖动误判为故障,难以判断异常是真实故障还是正常业务波动

Benefits of technology

[0052]同时,本发明的基于时序-日志指纹与多模态融合的接口异常分析方法通过轻量化两级架构设计,先通过实时轻量算法快速筛选可疑窗口,再进行可疑窗口内的深度融合分析,兼顾检测实时性与分析精准度,适配高并发、低开销的运行要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845401A_ABST
    Figure CN122845401A_ABST
Patent Text Reader

Abstract

This invention relates to interface anomaly analysis technology in the field of network communication, and discloses an interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion. The method includes: constructing a multimodal fingerprint vector that fuses time-series and log components within a time window; calculating its Mahalanobis distance to a long-term dynamic baseline vector; and identifying windows with a distance greater than a dynamic threshold as suspicious windows. For suspicious windows, a dual-branch sequence feature encoding is used to extract a global log semantic vector and an indicator feature vector, respectively. The log semantic vector and indicator feature vector are input into a cross-modal cross-attention layer for verification to generate a fusion vector. A classification output module outputs the real-time classification status of the interface as normal, sub-healthy, or under warning. This invention enables collaborative offset monitoring of indicators and logs, reduces single-dimensional false alarms, significantly improves the pre-emptive prevention and interpretability of latent interface anomalies, and reduces operation and maintenance troubleshooting costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication technology, and in particular to interface anomaly identification and analysis technology, specifically to an interface anomaly analysis method based on time-log fingerprinting and multimodal fusion. Background Technology

[0002] With the large-scale deployment of microservices, distributed systems, and cloud-native architectures, data interaction between systems via API interfaces has become the mainstream model. Unlike the monolithic architecture era, after microservice decomposition, business logic is broken down into numerous independent services. These services interact via interfaces such as HTTP, RPC, and message queues (MQ), with API calls spanning networks, multiple processes, and multiple servers. Interfaces represent the communication boundaries between different services and modules. Issues such as network jitter, incorrect parameters, dependency degradation, resource exhaustion, and version incompatibility can all trigger interface anomalies, causing business errors, request timeouts, and even chain collapses. The sources of failure become more complex, and the propagation paths and chains become more insidious. An anomaly in a single downstream interface can propagate upstream layer by layer, triggering cascading failures.

[0003] Traditional fault handling relies heavily on system log tracing, resulting in delayed fault detection and time-consuming localization, making it difficult to handle large-scale distributed scenarios. Currently, interface anomaly analysis and fault detection mainly rely on monitoring metrics, trace data, logs, and business message data. Based on observational data such as log data, trace data, original interface messages, and time-series data of metrics (e.g., QPS, concurrency, thread pool status), automatic anomaly identification and fault tracing are achieved through fixed threshold comparison, rule / keyword matching, and detection methods based on machine learning and deep learning models.

[0004] Performance monitoring based on fixed thresholds collects operational metrics such as interface response time, request latency, error rate, and QPS in real time. Anomalies are identified by pre-configured static thresholds or dynamically updated exception thresholds. Alarms are triggered when metrics exceed the thresholds, focusing only on single-dimensional metric changes. Anomaly detection based on log rule / keyword matching uses regular expressions to match explicit error log keywords such as ERROR, Timeout, and Fail, or configures fixed log template rules, relying on explicit anomaly features in the log text for identification. For fixed threshold or keyword / rule matching methods, metric monitoring and log detection are independent, making it difficult to capture the coordinated shift between metrics and logs before a failure occurs. This easily leads to missed hidden sub-health anomalies without error logs and misjudgments of business peaks or momentary fluctuations as failures, making it difficult to determine whether an anomaly is a genuine failure or normal business fluctuation. Furthermore, it relies primarily on post-event anomaly classification and fault location, failing to identify potential risks before failures occur, and thus failing to meet the needs of proactive prevention and maintenance of core interfaces.

[0005] Existing deep learning interface anomaly detection models are mostly single-modal input and post-event analysis, such as One-Class SVM, LSTM, and TCN. They are usually used for identification and classification after a fault occurs, but they cannot capture weak offset signals before the fault occurs. Moreover, they lack cross-modal verification and interpretability mechanisms, and cannot distinguish between false anomalies such as instantaneous jitter, GC fluctuations, and network interruptions. They have black box and unexplainable problems, which greatly increases the cost of operation and maintenance troubleshooting. Summary of the Invention

[0006] In view of the problems and defects of existing technologies, the purpose of this invention is to provide an interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion. This method establishes a strong semantic association between the interface's operational status (metric data) and business logic (log data). Through a multimodal spatiotemporal prediction model, the core metrics of the interface (response time, latency, error rate) are used as time-series features, and system logs generated within the same time period are used as text features. By capturing the collaborative offset of these two types of data before a fault occurs, accurate fault warnings are achieved.

[0007] According to a first aspect of the present invention, an interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion is proposed.

[0008] An interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion is characterized by the following steps:

[0009] S100: Extract time series components and log components respectively within each time window of preset length and without overlap, and fuse the time series components and log components to construct the multimodal fingerprint vector of the current time window;

[0010] S200: Calculate the deviation between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and combine it with the dynamic threshold to determine whether the system is normal or has suspicious deviations from the normal state, and determine the suspicious time window.

[0011] S300: For a received suspicious time window, obtain the original log sequence and original indicator sequence within the window;

[0012] S400: Input the original log sequence and the original indicator sequence into the dual-branch deep multimodal spatiotemporal prediction model. Extract the global log semantic vector and local temporal features through the log coding branch and the indicator coding branch respectively. Then perform cross-modal fusion verification through the cross attention layer and output the fusion vector of multimodal features.

[0013] S500: The fused vector is input into the fully connected layer for linear transformation and nonlinear activation. The output result is normalized using a normalized exponential function. Based on a maximum value indexing strategy, the real-time classification status of the current time window interface is output. The real-time classification status includes three states: normal, sub-healthy, and warning.

[0014] S600: The real-time classification status is converted into a graded early warning signal through mapping rules, and the early warning signal level, early warning information, time window start and end time and indicator value are packaged and encapsulated into a structured object and sent to the alarm system.

[0015] As an optional implementation, in step S100, the timing component within each time window is composed of a three-dimensional sub-vector consisting of the average response time, average latency, and average error rate of all interface requests within the window, which is obtained by statistically calculating the average of the sampled data of the corresponding performance index within the current window.

[0016] The log component is obtained by templated log messages in the current window, statistically analyzing the distribution of various log templates in the current window, and then converting them into a sparse vector of fixed dimensions using TF-IDF weights.

[0017] The time-series component and log component of the current time window are concatenated to construct the multimodal fingerprint vector of the current time window, and the interface's running status indicator data and business logic log data are strongly correlated at the semantic level.

[0018] As an optional implementation, for the statistics of log templates within each time window, each log message is templated, and the variable values, IDs, IP addresses, and timestamps are replaced with fixed placeholders <*>, while the unchanging words and structures are retained. Similar logs are grouped into the same log template and assigned a unique template identifier.

[0019] Collect all unique log templates to form a set; count the frequency of each type of log template within each time window to obtain the distribution data of log templates.

[0020] As an optional implementation, in step S200, based on the system's steady-state reference vector, the long-term dynamic baseline vector is asynchronously updated using EWMA:

[0021] Multiply the current baseline vector by a preset first weighting factor, multiply the steady-state reference vector by a second weighting factor, and add the products of the two to obtain the updated long-term dynamic baseline vector.

[0022] Wherein, the second weighting factor is the learning rate of receiving new normal data, and the sum of the first weighting factor and the second weighting factor is one;

[0023] The stable state reference vector refers to a verified window fingerprint vector that represents the system in a healthy operating state. It is the only legitimate data source for baseline updates, and the multimodal fingerprint vector is only marked as a stable state reference vector and allowed to participate in baseline updates when the multimodal fingerprint vector in the current time window is determined to be non-abnormal.

[0024] As an optional implementation, in step S200, during the cold start window of the system's initial operation, the multimodal fingerprint vectors of all time windows are collected and the arithmetic mean is calculated as the initial baseline vector.

[0025] Calculate the standard fractional variant and coefficient of variation of the mean response time for each time window within the cold start window; wherein, the standard fractional variant is obtained by calculating the absolute difference between the mean response time of the current time window and the historical global mean, and then dividing it by the global standard deviation; the coefficient of variation is obtained by calculating the ratio between the standard deviation of the response time within the current time window and the mean response time;

[0026] When the standard score variant is less than or equal to a preset significant deviation threshold and the coefficient of variation is not significantly higher than the average level of the cold start window, the current time window is determined to be normal, and its corresponding multimodal fingerprint vector is marked to construct a stable state reference vector.

[0027] As an optional implementation, in step S200, the deviation is determined by calculating the Mahalanobis distance or Euclidean distance between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and compared with the dynamic threshold. If the deviation exceeds the dynamic threshold, the time window is determined to be a suspicious window.

[0028] The dynamic threshold is determined based on the historical mean and historical standard deviation of Mahalanobis distance or Euclidean distance within a preset historical time period. The historical standard deviation is multiplied by the dynamic sensitivity coefficient and then added to the historical mean to obtain the dynamic threshold.

[0029] As an optional implementation, in step S400, the dual-branch deep multimodal spatiotemporal prediction model includes a log coding branch, an index coding branch, a cross-modal cross-attention layer, and a fully connected layer.

[0030] The log encoding branch receives the original log message input extracted within the suspicious time window, performs word segmentation and embedding processing through a pre-trained distributed representation model, and obtains a global log semantic vector representing the business logic of the entire time window after mean pooling.

[0031] The index encoding branch has a one-dimensional convolutional layer and a connected gated recurrent unit; the one-dimensional convolutional layer is used for local pattern extraction, and the original time series of the index, including response time, latency and error rate, is input into the one-dimensional convolutional layer to extract local temporal features; then it is input into the gated recurrent unit for long-range dependency capture, and the output is an index feature vector representing the dynamic trend of the index throughout the time window. The gated recurrent unit adopts a bidirectional gated recurrent unit layer.

[0032] The cross-modal attention layer uses the log semantic vector as the query matrix and the indicator feature vector as the key matrix and value matrix, respectively, and inputs them into the cross-modal attention layer for cross-modal fusion and verification. By calculating the matching degree between the query matrix and the indicator feature vector, an attention weight is obtained, and the indicator feature vector is weighted based on the attention weight to obtain a fusion vector that integrates multimodal features.

[0033] As an optional implementation, in step S400, the multimodal feature fusion process based on cross-modal cross-attention includes:

[0034] The log semantic vector L is transformed linearly to generate the query matrix Q, and the indicator feature vector M is transformed into the key matrix K and the value matrix V respectively.

[0035] By calculating the transpose matrix K of the query matrix Q and the key matrix K. T The dot product is calculated and then divided by a preset scaling factor to obtain the attention score matrix.

[0036] The attention scoring matrix is ​​normalized using a normalized exponential function to obtain attention weights, which are used to enhance the semantically relevant parts of the indicator features.

[0037] Then, the attention weights are used to sum the indicator feature vector M to obtain the fusion vector F, so that the model can automatically focus on the relevant semantic paragraphs in the log sequence when the indicator deviates.

[0038] As an optional implementation, the method further includes: performing interpretability processing on the output interface status warning results, specifically including the following steps:

[0039] The attention weights calculated by the cross-modal attention layer are recorded and averaged on the time axis to obtain the average attention score at each time point, and then arranged in descending order.

[0040] Find the physical moments when the indicators corresponding to the top few indices with the highest average attention scores shifted.

[0041] Retrieve original log messages within a preset time deviation range before and after the physical time; if a corresponding log exists, extract its corresponding log template;

[0042] The start and end times of the current time window, the indicator name, the current value of the indicator, and the retrieved log templates or logs without relevant tags are packaged to generate structured early warning information.

[0043] As an optional implementation, the distributed representation model, one-dimensional convolutional layer, bidirectional gated recurrent unit layer, and fully connected layer of the deep multimodal network model are parameterized through an offline training phase, specifically including the following process:

[0044] Collect historical multimodal sample pairs and complete max-min normalization and lexical encoding, and simultaneously label the corresponding normal, sub-healthy, and warning status tags;

[0045] The difference between the predicted probability and the true label in the current iteration is calculated using the weighted cross-entropy loss function.

[0046] After the index encoding branch and fully connected layer converge, the network parameters of the distributed representation model are unfrozen, and the entire network parameters are backpropagated and the parameter gradients are updated with a fine-tuned learning rate lower than the initial learning rate until the loss function of the model on the validation set no longer decreases for several consecutive cycles. The training is then stopped, the optimal weight parameters are saved, and the solidified deep multimodal network model is obtained.

[0047] According to a second aspect of the present invention, a computing system is provided, comprising:

[0048] One or more processors;

[0049] A memory, on which computer programs are stored;

[0050] The process by which the one or more processors implement the aforementioned interface anomaly analysis method based on time-log fingerprinting and multimodal fusion when executing the computer program.

[0051] Combining the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion from the above embodiments, this method utilizes lightweight time-series anomaly detection and deep learning-based multimodal spatiotemporal early warning analysis. First, it quickly filters suspicious time windows. Then, it performs deep multimodal analysis on the logs and time-series features within these windows. Based on cross-modal interactive attention analysis using indicator features and log semantic features, it confirms genuine anomalies or normal fluctuations caused by business peaks, achieving highly real-time and low-overhead accurate analysis. Based on baseline anomaly identification through the rapid filtering process, even without any error logs, the system can still capture sub-health signals if the response time pattern deviates. Simultaneously, through subsequent multimodal causal verification of logs / indicators, the real-time performance and accuracy of interface anomaly identification are improved, filtering out normal fluctuations such as business peaks and instantaneous network jitter. In particular, the time-series dynamic baseline and multimodal verification logic can effectively identify implicit sub-health states of interfaces without error logs, reducing false alarms caused by single-dimensional fluctuations (such as simple instantaneous network jitter).

[0052] Meanwhile, the interface anomaly analysis method based on time-log fingerprinting and multimodal fusion of the present invention adopts a lightweight two-level architecture design. First, it quickly filters suspicious windows through real-time lightweight algorithms, and then performs in-depth fusion analysis within the suspicious windows. It takes into account both the real-time detection and the accuracy of analysis, and is suitable for high concurrency and low overhead operation requirements.

[0053] It should be understood that all combinations of the foregoing concepts and the additional concepts described in more detail below may be considered part of the inventive subject matter of this disclosure, provided that such concepts do not contradict each other. Furthermore, all combinations of the claimed subject matter are considered part of the inventive subject matter of this disclosure.

[0054] The foregoing and other aspects, embodiments, and features of the teachings of the present invention will be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the invention, such as features and / or beneficial effects of exemplary embodiments, will become apparent from the following description or may be learned through practice of specific embodiments according to the teachings of the present invention. Attached Figure Description

[0055] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various figures may be denoted by the same reference numeral. For clarity, not every component is labeled in each figure. Embodiments of various aspects of the invention will now be described by way of example and with reference to the accompanying drawings.

[0056] Figure 1 This is a flowchart illustrating an interface anomaly analysis method based on time-log fingerprinting and multimodal fusion according to an embodiment of the present invention.

[0057] Figure 2This is a schematic diagram of a suspicious window identification process based on baseline deviation according to an embodiment of the present invention.

[0058] Figure 3 This is a schematic diagram illustrating the principle of a dual-branch deep multimodal spatiotemporal prediction model according to an embodiment of the present invention.

[0059] Figure 4 This is a schematic diagram of a multimodal feature fusion process based on cross-modal cross-attention according to an embodiment of the present invention.

[0060] Figure 5 This is a schematic diagram illustrating the process of handling the interpretability of the output interface state according to an embodiment of the present invention. Detailed Implementation

[0061] To better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.

[0062] Various aspects of the invention are described in this disclosure with reference to the accompanying drawings, which illustrate numerous illustrative embodiments. The embodiments of this disclosure are not necessarily intended to encompass all aspects of the invention. It should be understood that the various concepts and embodiments described above, as well as those described in more detail below, can be implemented in any of many ways, because the concepts and embodiments disclosed herein are not limited to any particular implementation. Furthermore, some aspects of the invention disclosed may be used alone or in any suitable combination with other aspects of the invention disclosed.

[0063] {Example 1}

[0064] Combination Figures 1-4 As shown in the embodiment of the present invention, the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion combines the time-series features of interface performance indicators with the textual features of system logs to construct a multimodal fingerprint vector, realizing a strong semantic-level correlation between the interface's operating status and business logic. Based on this, a stable state reference vector is used, and the long-term dynamic baseline vector is dynamically updated using EWMA to achieve baseline drift protection. By asynchronously updating the baseline and combining Z-score variants and coefficient of variation (CV) to identify the interface's sub-health state, Mahalanobis distance and dynamic thresholds are used to quickly filter suspicious windows. Then, combined with a dual-branch deep multimodal cross-modal fusion model design, a dual-branch structure of log semantic encoding and indicator time-series feature extraction is used, combined with a cross-attention layer to achieve causal verification between indicators and logs, completing a rapid real-time classification and determination of the interface as normal / sub-healthy / warning, realizing collaborative offset monitoring of indicators and logs, reducing single-dimensional false alarms, and improving the pre-emptive prevention and interpretability of hidden anomalies.

[0065] Combination Figure 1 , Figure 2 , Figure 3As shown, the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion, according to one example, includes the following steps:

[0066] S100: Extract time series components and log components respectively within each time window of preset length and without overlap, and fuse the time series components and log components to construct the multimodal fingerprint vector of the current time window;

[0067] S200: Calculate the deviation between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and combine it with the dynamic threshold to determine whether the system is normal or has suspicious deviations from the normal state, and determine the suspicious time window.

[0068] S300: For a received suspicious time window, obtain the original log sequence and original indicator sequence within the window;

[0069] S400: Input the original log sequence and the original indicator sequence into the dual-branch deep multimodal spatiotemporal prediction model. Extract the global log semantic vector and local temporal features through the log coding branch and the indicator coding branch respectively. Then perform cross-modal fusion verification through the cross attention layer and output the fusion vector of multimodal features.

[0070] S500: The fused vector is input into the fully connected layer for linear transformation and nonlinear activation. The output result is normalized using a normalized exponential function. Based on a maximum value indexing strategy, the real-time classification status of the current time window interface is output. The real-time classification status includes three states: normal, sub-healthy, and warning.

[0071] S600: The real-time classification status is converted into a graded early warning signal through mapping rules, and the early warning signal level, early warning information, time window start and end time and indicator value are packaged and encapsulated into a structured object and sent to the alarm system.

[0072] As an optional implementation, in step S100, the timing component within each time window is composed of a three-dimensional sub-vector consisting of the average response time, average latency, and average error rate of all interface requests within the window, which is obtained by statistically calculating the average of the sampled data of the corresponding performance index within the current window.

[0073] The log component is obtained by templated log messages in the current window, statistically analyzing the distribution of various log templates in the current window, and then converting them into a sparse vector of fixed dimensions using TF-IDF weights.

[0074] The time-series component and log component of the current time window are concatenated to construct the multimodal fingerprint vector of the current time window, and the interface's running status indicator data and business logic log data are strongly correlated at the semantic level.

[0075] As an optional implementation, for the statistics of log templates within each time window, each log message is templated, and the variable values, IDs, IP addresses, and timestamps are replaced with fixed placeholders <*>, while the unchanging words and structure are retained. Similar logs are grouped into the same log template and assigned a unique template identifier.

[0076] Collect all unique log templates to form a set; count the frequency of each type of log template within each time window to obtain the distribution data of log templates.

[0077] As an optional implementation, in step S200, based on the system's steady-state reference vector, the long-term dynamic baseline vector is asynchronously updated using EWMA:

[0078] Multiply the current baseline vector by a preset first weighting factor, multiply the steady-state reference vector by a second weighting factor, and add the products of the two to obtain the updated long-term dynamic baseline vector.

[0079] Wherein, the second weighting factor is the learning rate of receiving new normal data, and the sum of the first weighting factor and the second weighting factor is one;

[0080] The stable state reference vector refers to a verified window fingerprint vector that represents the system in a healthy operating state. It is the only legitimate data source for baseline updates, and the multimodal fingerprint vector is only marked as a stable state reference vector and allowed to participate in baseline updates when the multimodal fingerprint vector in the current time window is determined to be non-abnormal.

[0081] As an optional implementation, in step S200, during the cold start window of the system's initial operation, the multimodal fingerprint vectors of all time windows are collected and the arithmetic mean is calculated as the initial baseline vector.

[0082] Calculate the standard fractional variant and coefficient of variation of the mean response time for each time window within the cold start window; wherein, the standard fractional variant is obtained by calculating the absolute difference between the mean response time of the current time window and the historical global mean, and then dividing it by the global standard deviation; the coefficient of variation is obtained by calculating the ratio between the standard deviation of the response time within the current time window and the mean response time;

[0083] When the standard score variant is less than or equal to a preset significant deviation threshold and the coefficient of variation is not significantly higher than the average level of the cold start window, the current time window is determined to be normal, and its corresponding multimodal fingerprint vector is marked to construct a stable state reference vector.

[0084] As an optional implementation, in step S200, the deviation is determined by calculating the Mahalanobis distance or Euclidean distance between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and compared with the dynamic threshold. If the deviation exceeds the dynamic threshold, the time window is determined to be a suspicious window.

[0085] The dynamic threshold is determined based on the historical mean and historical standard deviation of Mahalanobis distance or Euclidean distance within a preset historical time period. The historical standard deviation is multiplied by the dynamic sensitivity coefficient and then added to the historical mean to obtain the dynamic threshold.

[0086] As an optional implementation, in step S400, the dual-branch deep multimodal spatiotemporal prediction model includes a log coding branch, an index coding branch, a cross-modal cross-attention layer, and a fully connected layer.

[0087] The log encoding branch receives the original log message input extracted within the suspicious time window, performs word segmentation and embedding processing through a pre-trained distributed representation model, and obtains a global log semantic vector representing the business logic of the entire time window after mean pooling.

[0088] The index encoding branch has a one-dimensional convolutional layer and a connected gated recurrent unit; the one-dimensional convolutional layer is used for local pattern extraction, and the original time series of the index, including response time, latency and error rate, is input into the one-dimensional convolutional layer to extract local temporal features; then it is input into the gated recurrent unit for long-range dependency capture, and the output is an index feature vector representing the dynamic trend of the index throughout the time window. The gated recurrent unit adopts a bidirectional gated recurrent unit layer.

[0089] The cross-modal attention layer uses the log semantic vector as the query matrix and the indicator feature vector as the key matrix and value matrix, respectively, and inputs them into the cross-modal attention layer for cross-modal fusion and verification. By calculating the matching degree between the query matrix and the indicator feature vector, an attention weight is obtained, and the indicator feature vector is weighted based on the attention weight to obtain a fusion vector that integrates multimodal features.

[0090] As an optional implementation, in step S400, the multimodal feature fusion process based on cross-modal cross-attention includes:

[0091] The log semantic vector L is transformed linearly to generate the query matrix Q, and the indicator feature vector M is transformed into the key matrix K and the value matrix V respectively.

[0092] By calculating the transpose matrix K of the query matrix Q and the key matrix K. T The dot product is calculated and then divided by a preset scaling factor to obtain the attention score matrix.

[0093] The attention scoring matrix is ​​normalized using a normalized exponential function to obtain attention weights, which are used to enhance the semantically relevant parts of the indicator features.

[0094] Then, the attention weights are used to sum the indicator feature vector M to obtain the fusion vector F, so that the model can automatically focus on the relevant semantic paragraphs in the log sequence when the indicator deviates.

[0095] As an optional implementation, the interface anomaly analysis method of the present invention further includes interpretability processing of the output interface status warning results, specifically including the following steps:

[0096] The attention weights calculated by the cross-modal attention layer are recorded and averaged on the time axis to obtain the average attention score at each time point, and then arranged in descending order.

[0097] Find the physical moments when the indicators corresponding to the top few indices with the highest average attention scores shifted.

[0098] Retrieve original log messages within a preset time deviation range before and after the physical time; if a corresponding log exists, extract its corresponding log template;

[0099] The start and end times of the current time window, the indicator name, the current value of the indicator, and the retrieved log templates or logs without relevant tags are packaged to generate structured early warning information.

[0100] It should be understood that the distributed representation model, one-dimensional convolutional layer, bidirectional gated recurrent unit layer, and fully connected layer of the aforementioned deep multimodal network model have their parameters fixed during the offline training phase, specifically including the following processes:

[0101] Collect historical multimodal sample pairs and complete max-min normalization and lexical encoding, and simultaneously label the corresponding normal, sub-healthy, and warning status tags;

[0102] The difference between the predicted probability and the true label in the current iteration is calculated using the weighted cross-entropy loss function.

[0103] After the index encoding branch and fully connected layer converge, the network parameters of the distributed representation model are unfrozen, and the entire network parameters are backpropagated and the parameter gradients are updated with a fine-tuned learning rate lower than the initial learning rate until the loss function of the model on the validation set no longer decreases for several consecutive cycles. The training is then stopped, the optimal weight parameters are saved, and the solidified deep multimodal network model is obtained.

[0104] {Example 2}

[0105] This embodiment provides a more detailed example of the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion from the aforementioned embodiments. This example of interface anomaly analysis mainly includes a suspicious window filtering process based on a dynamic baseline and a deep multimodal fusion analysis process for suspicious data.

[0106] Phase 1: Suspicious Window Screening Based on Dynamic Baselines

[0107] Step 1: Construct a multimodal fingerprint vector based on the time-series component and the log component.

[0108] In this embodiment, logs can be collected and sent using the Fluentd tool, an existing log collector.

[0109] For each time window W, taking 10 seconds as an example, extract a fused temporal component v. metric With log component v log The composite vector V w As a multimodal fingerprint vector: V w =[v metric v log ].

[0110] Wherein, the time component v metric A three-dimensional sub-vector is constructed from the average response time, average latency, and average error rate of all interface requests within window W. This sub-vector is obtained by statistically calculating the average of the performance metric sample data within the current window. Log component v log The distribution of various log templates within the current window is statistically analyzed and then transformed into a sparse vector of fixed dimensions using TF-IDF weights.

[0111] In an embodiment of the invention, a fixed-length time window W of 10 seconds is set. The system generates a new time window every 10 seconds, with no overlap between windows. For each time window W, the system retrieves all log lines whose timestamps fall within that window from the log stream. Each log entry includes: a timestamp, a log level (INFO / WARN / ERROR), and a log message body.

[0112] For log template statistics, each log message is templated, and the variable values, IDs, IP addresses, and timestamps are replaced with fixed placeholders <*>, while the unchanging words and structure are retained. Similar logs are grouped into the same log template, and each log template is assigned a unique template identifier.

[0113] Specifically, after templated processing of all log entries within the window, all unique log templates are collected to form a set. Each template has a unique template ID, such as template ID=1,2,3... etc. For example, assuming there are 100 log entries in the window and 3 different templates are generated, the records would be as follows:

[0114] Template A: User login successful, userId= <num>;

[0115] Template B: Connection timeout to database, retry= <num>;

[0116] Template C: High memory usage, current= <num>%.

[0117] Then, the log template set for this window is {A, B, C}, which can actually be stored as an ID set, such as {1, 5, 9}. If there are no logs in the window: the set is empty.

[0118] The frequency of various templates appearing in the statistics window is used to obtain the distribution data of log templates.

[0119] Step 2: Asynchronously update the baseline vector B based on EWMA (Exponentially Weighted Moving Average). new .

[0120] In this example, the baseline vector is updated according to the following formula:

[0121] B new =(1-α)B + αV stable ;

[0122] Here, α learning rate represents the weight of the impact of new data on the long-term baseline, which determines the speed at which the system accepts the new normal, and is usually taken in the range of 0.01-0.05.

[0123] It should be understood that if the learning rate α is large, the baseline will quickly follow business fluctuations, such as flash sales or celebrations; if the learning rate α is small, the baseline is more stable and can better identify slow-onset performance degradation, such as progressive memory leaks.

[0124] In the formula, parameter V stable This parameter represents the stable state reference vector. Only window vectors that are determined to be normal will participate in baseline updates, preventing the baseline from being contaminated by abnormal data and achieving baseline drift protection.

[0125] In this way, if the business volume rises slowly during the midday peak, the baseline will move smoothly accordingly and will not trigger false alarms.

[0126] In an embodiment of the present invention, the aforementioned stable state reference vector V stable This refers to a verified window fingerprint vector that represents a healthy operating state of the system and is the only legitimate data source for baseline updates. stable It is not a newly generated vector, but rather a vector from the current window vector V. w Selected from the list. As mentioned above, only when V... w It is marked as V only when it meets the non-anomaly criteria. stable It is also permitted to be used to modify the baseline.

[0127] For the system, during the initial cold start window, for example, a range of 15-30 minutes can be set, the vectors of all time windows within this period are collected and the arithmetic mean of the vectors is calculated, which is used as the initial baseline B0.

[0128] Simultaneously, the mean response time µ of all window vectors (W=10s) within the cold start window is calculated. w The Z-score variant Z and the coefficient of variation CV. The mean response time µ. w The Z-score variant Z is based on the global mean µ base Compared with the global standard deviation σ base The calculation yielded:

[0129] Z=|µ w -µ base | / σ base ;

[0130] CV=σ w / µ w .

[0131] Here, Z represents the number of times the current average deviates from the historical average (in standard deviations). For example, if the average response time for all requests changes from 100ms to 150ms, the Z score will jump rapidly (e.g., from 0 to 5). Due to µ... w Significantly pulled away from µ base Furthermore, due to the overall consistent pace, the molecules become larger, causing the Z score to exceed the preset threshold. In this example, Z>3 is set as a significant anomaly.

[0132] It should be understood that the aforementioned µ w σ w These represent the mean and standard deviation of the response time within the current window (W=10s), respectively. (Using µ...) w σ w The standard deviation (CV) is calculated as a ratio to the mean to measure the degree of fluctuation in interface performance. The larger the CV, the more uneven the data distribution, thus accurately identifying the sub-healthy state where the mean is not yet fully poisoned but extreme spikes have already appeared.

[0133] For example, most requests are fast, but a very small number of requests become extremely slow due to occasional garbage collection, specific slow SQL queries, or single server failures. If the CV (volume response) is significantly higher than the average level of the cold start window, the sampling point is considered to have experienced severe fluctuations, exhibiting sub-health characteristics.

[0134] Therefore, by using the Z-score variants and CVs of each time window within the cold start window, normal windows are filtered out, and the V calculated based on their window vectors is then used. w Incorporating the baseline accelerates its formation and constructs a stable state reference vector V. stable .

[0135] Step 3: Based on the updated baseline vector B new Calculate the window vector V of the current window. w The deviation is measured and a suspicious case is identified.

[0136] In this invention, the current window vector V is used. w Compared with baseline B new The deviation S is represented by Mahalanobis distance or weighted Euclidean distance.

[0137] Combination Figure 2 The example shown in this embodiment uses Mahalanobis distance as an example to illustrate the correlation between different indicators:

[0138] ;

[0139] Where C represents the covariance matrix.

[0140] Furthermore, the calculated Mahalanobis distance S is compared with the dynamic threshold T. If the Mahalanobis distance S is greater than or equal to the dynamic threshold T, it is determined that there is an unstable suspicious window in the current window, that is, it deviates from the normal state, and it is determined to be a suspicious window, and further enters the subsequent deep multimodal analysis. Otherwise, it is determined that the window is normal and enters the next window calculation and identification.

[0141] As an optional implementation, the dynamic threshold T is determined based on the historical mean and standard deviation of the Mahalanobis distance S. For ease of explanation, this example uses data from the past 1-2 hours:

[0142] T=µ s + kσ s ;

[0143] Where, µ s With σ s These refer to the historical mean and standard deviation of the Mahalanobis distance S, respectively, and k represents the dynamic sensitivity coefficient. A smaller dynamic sensitivity coefficient k indicates a more sensitive system, particularly suitable for core interfaces such as payment interfaces; a larger k indicates a more lenient system, suitable for non-core interfaces such as query interfaces.

[0144] Therefore, if the Mahalanobis distance S is normal (within the allowable range of the dynamic threshold T), it indicates that the current window state is normal, and the calculated window vector V can be used. m Fine-tune the dynamic baseline vector update. If the Mahalanobis distance S is abnormal (exceeding the dynamic threshold T), pause the baseline update.

[0145] Step 4: Based on the dual-branch deep multimodal spatiotemporal prediction model, use its log coding branch and time-series coding branch to perform multimodal spatiotemporal prediction model fusion analysis on the screened suspicious data to determine whether the aforementioned offset is a benign offset caused by business growth or a deterioration caused by interface failure, thereby achieving accurate analysis and judgment of interface anomalies.

[0146] Combination Figure 3 As mentioned above, the dual-branch deep multimodal spatiotemporal prediction model includes a log coding branch, a temporal coding branch, a cross-modal attention layer, and a fully connected layer.

[0147] In a specific embodiment, the dual-branch deep multimodal spatiotemporal prediction model adopts a deep multimodal Transformer model, which has a dual-branch network, one branch for handling log coding and the other branch for handling temporal coding.

[0148] Combination Figure 3 As shown, when the system receives a suspicious window, it retrieves the following raw data of that window:

[0149] Log sequence: Each raw log message (untemplated) in the window is sorted by time; a maximum of the first 256 log messages are retained, and if more are exceeded, they are truncated; if fewer are not retained, they are filled with empty log messages.

[0150] Metric sequence: The original time series of each metric (response time, latency, and error rate) within the window; for each metric, all metric values ​​are grouped into a sequence.

[0151] Combination Figure 3 As shown, log encoding and metric encoding are performed through two branches respectively.

[0152] In the log encoding branch, N original logs (N ≤ 256) within the window are input into a pre-trained distributed representation model; this example uses Distil BERT. Each log is segmented and converted into a vector sequence. Then, the [CLS] bit feature of each log is extracted, or global average pooling is used to obtain the embedding vector E of a single log. i (256 dimensions). Average pooling is performed on all log vectors within the window, calculated using the following formula:

[0153] ;

[0154] Thus, a 256-dimensional vector L representing the entire window log is obtained. This vector integrates global semantic information of the business logic within the current window and can identify weak features such as retries and waiting hidden in non-error logs.

[0155] The metric encoding branch has a one-dimensional convolutional layer (1D-CNN) and a connected gated recurrent unit (GRU). The branch input dimension is (tw, D), where tw is the time step (sampling points within the window) and D is the metric dimension (response time, latency, error rate).

[0156] In this architecture, the one-dimensional convolutional layer (1D-CNN) is used for local pattern extraction. The gated recurrent unit (GRU) employs a bidirectional GRU layer for long-range dependency capture.

[0157] In this example, 1D-CNN uses a convolutional kernel of size 3, i.e., Kernel Size=3, Stride=1. By sliding the calculation on the time axis, it extracts the local temporal features of the index sequence, captures the instantaneous jumps (such as Spikes) and local fluctuation patterns of the index through convolution operations, and outputs feature maps.

[0158] The Gated Recurrent Unit (GRU) is used as a sequence processing network to capture the evolutionary dependencies of the index sequence on the forward and reverse time axes. The convolutional feature map is input into the bidirectional GRU layer. The GRU determines how much historical state information to retain through the Reset Gate and Update Gate. Finally, the hidden layer outputs a 256-dimensional vector M, which represents the dynamics of the index in the entire window and characterizes the dynamic evolution trend of the index.

[0159] Furthermore, cross-modal fusion output is performed in the cross-modal cross-attention layer. The log semantic vector is used as the query matrix, and the indicator feature vector is used as the key matrix and value matrix, respectively. These are input into the cross-modal cross-attention layer for cross-modal fusion and verification. By calculating the matching degree between the query matrix and the indicator feature vector, an attention weight is obtained. Based on the attention weight, the indicator feature vector is weighted to obtain the fusion vector that integrates multimodal features.

[0160] Combination Figure 4 The example shown illustrates a multimodal feature fusion process based on cross-modal cross-attention, which includes:

[0161] The log semantic vector L is transformed linearly to generate the query matrix Q, and the indicator feature vector M is transformed into the key matrix K and the value matrix V respectively.

[0162] By calculating the transpose matrix K of the query matrix Q and the key matrix K. T The dot product is calculated and then divided by a preset scaling factor to obtain the attention score matrix.

[0163] The attention scoring matrix is ​​normalized using a normalized exponential function to obtain attention weights (length 256), which are used to enhance the semantically relevant parts of the indicator features.

[0164] Then, the attention weights are used to sum the indicator feature vector M to obtain the fusion vector F (length 256), which enables the model to automatically focus on the relevant semantic segments in the log sequence when the indicator deviates.

[0165] As an example, attention weights are calculated as follows:

[0166] ;

[0167] Among them, W Q W K W V Let d be a learnable weight matrix. k This is the scaling factor.

[0168] This allows the model to automatically focus on relevant semantic segments in the log sequence when the metric deviates, thus enabling mutual verification of the causal relationship between metric fluctuations and log content.

[0169] Finally, the fusion vector F is input to the fully connected layer defined by the classification weight matrix. Through linear transformation and nonlinear activation, the fusion vector F is mapped to three-dimensional components representing the interface state. Then, the Softmax function is used to normalize these three-dimensional components, outputting the confidence probability values ​​for the interface being in normal, sub-healthy, and warning states (the sum of the three probabilities is 1). Finally, the real-time classification state of the interface is determined according to the Argmax maximum probability strategy.

[0170] Specifically, the calculation of the fully connected layer includes:

[0171] Z out =F*W cl +b;

[0172] Among them, W cl The classification weight matrix has a dimension of 3×256 and is used to project the fused 256-dimensional multimodal feature space onto the three-dimensional category space, corresponding to normal, sub-healthy, and warning. The bias term b has a dimension of 3×1 and is used to correct the offset of the model output.

[0173] Furthermore, nonlinear activations can be added to the fully connected layers. Taking the ReLU activation function as an example, this enhances the nonlinear expressive power of the model, and is calculated as follows:

[0174] A = max(0, Z out );

[0175] Finally, input the Softmax function to calculate the output:

[0176] ;

[0177] Among them, P i This represents the probability that the current window belongs to category i, where i = 1, 2, 3.

[0178] Furthermore, based on the Argmax maximum probability strategy: Output = argmax(P i The category with the highest probability value is selected as the final interface status determination result.

[0179] As an optional implementation, during the training phase, the system compares the output probabilities with pre-labeled fault labels, calculates the difference between the predicted probabilities and the true labels using a weighted cross-entropy loss function, and iteratively trains by minimizing the cross-entropy loss function, collaboratively optimizing the parameters of the front-end dual-branch network and the terminal fully connected layer. Especially for sub-healthy samples, the model enhances its sensitivity to latent deviations in indicators by adjusting the weight bias of the output layer.

[0180] Through iterative training until the training termination conditions are met, the trained and solidified model is deployed for online diagnosis and analysis.

[0181] As an example, the training process includes the following steps:

[0182] First, during the data preparation phase, historical monitoring data is segmented into fixed 10-second time windows. Raw log messages within the same time window are correlated with indicator time series to form multimodal sample pairs. Simultaneously, each window is labeled with three categories: normal, sub-healthy, and warning. Min-Max normalization (scaling to 0-1) is performed on the indicator series; log text is tokenized and uniformly padded to a length of 256.

[0183] Then, feature encoding branch initialization is performed, pre-trained BERT weights are loaded and set to feature extraction mode; the convolutional kernels of 1D-CNN and the gating units of GRU are initialized with random normal distribution;

[0184] Then, a dual-branch parallel encoding process is performed. Log data is fed into BERT to generate vector L, and metric data is processed through 1D-CNN+GRU to generate vector M. Vectors L and M are then input into a cross-modal attention layer to calculate attention weights and generate a fused vector F. Finally, the fused vector F undergoes linear transformation and activation function processing in a fully connected layer, and is then mapped to a classification decision using the Softmax function to output the predicted probabilities of the three states.

[0185] During training, the weighted cross-entropy loss function is used to calculate the difference between the predicted probability P and the true label. Based on the loss value, the gradient of each layer's parameters is calculated using the backpropagation algorithm for gradient backpropagation. During iterative training, the AdamW optimizer is used to update the weights of the convolutional layers, GRU layers, and fully connected layers according to the gradient direction. The updated weights are then fed into the next round of forward propagation. Through iterative training, the model performance is spirally improved until the branches converge.

[0186] After the metric branch converges, the BERT parameters are unfrozen and fine-tuned with a very low learning rate, thus improving semantic understanding and making it more suitable for operational scenarios. Performance is evaluated on an independent validation set; training stops if the loss function no longer decreases for 10 consecutive epochs. The optimal weight parameters from the validation set are saved to generate a fixed prediction model.

[0187] Step 6: Perform interpretability processing on the output interface status warning results.

[0188] After a traditional AI model issues an alarm, the black box nature and lack of interpretability make it difficult for operations and maintenance personnel to troubleshoot the alarm. The method in this invention uses attention scores to reverse locate the physical moment and retrieve logs, achieving a combination of proactive prevention and fault interpretability. This significantly improves the proactive prevention and interpretability of hidden interface anomalies and reduces the cost of operations and maintenance troubleshooting.

[0189] In an embodiment of the present invention, the attention weight vector (256 values) calculated by the cross-modal attention layer is recorded. Each value corresponds to a position in the index vector M, that is, a time point or a feature channel.

[0190] To simplify the process, we focus only on the time dimension of the indicator sequence: we average the attention weights over time points to obtain the average attention score for each time point.

[0191] If the metric vector M contains multiple time steps, the attention mechanism should generate a corresponding number of weight scalars, and identify the top N indices with the largest weight scalars. These indices directly correspond to the physical moments when the metric experiences significant shifts. At the moment t with the largest weights, the mechanism then... max Then, search for time t. max The original logs appearing within ±Δt (e.g., 2 seconds before and after). If no logs are found, mark them as unrelated.

[0192] Finally, using a predefined alarm message template, specific values ​​and keywords are filled in to generate a natural language explanation.

[0193] For example: Template 1 (with logs):

[0194] Within the window start time to window end time, the metric name rises from the normal baseline value to the current value, and at the same time, a message of type [log template] appears in the log. The two are highly correlated, and it is recommended to check the [suggested action].

[0195] Template 2 (No log):

[0196] If abnormal fluctuations occur in the [metric name] (current value [current value], historical average [average value]) between the [window start time] and [window end time], and no obvious related logs are found, it is recommended to check the system load or external dependencies.

[0197] Fill example:

[0198] Window time: 10:00:00-10:00:10;

[0199] Metric Name: Average Latency;

[0200] Normal baseline value: 50ms (obtained from historical average);

[0201] Current value: 85ms;

[0202] Log template: Connection timeout to database, retry= <num>;

[0203] Recommended action: Check the database connection pool status.

[0204] The example output alarm message is as follows:

[0205] Between 10:00:00 and 10:00:10, the average latency increased from 50ms to 85ms, and messages of the type 'Connection timeout to database, retry=3' appeared in the logs. The two are strongly correlated, and it is recommended to check the database connection pool status.

[0206] As an optional implementation, the output category can be mapped to a level:

[0207] Normal: Green;

[0208] Sub-health: Yellow alert (WARNING);

[0209] Warning: Red alert (EMERGENCY).

[0210] Finally, the warning level, its corresponding explanatory text, window time, and relevant indicator values ​​are packaged into a JSON structure and sent to the alarm system.

[0211] {Example 3}

[0212] Based on the implementation process of the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion in the above embodiments, this invention also proposes a computing system, including one or more processors and a memory. The memory stores a computer program. When the aforementioned one or more processors execute the aforementioned computer program, they implement the process of the interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion in any of the aforementioned embodiments.

[0213] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.< / num> < / num> < / num> < / num>

Claims

1. An interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion, characterized in that, Includes the following steps: S100: Extract time series components and log components respectively within each preset length and non-overlapping time window, and fuse the time series components and log components to construct the multimodal fingerprint vector of the current time window; S200: Calculate the deviation between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and combine it with the dynamic threshold to determine whether the system is normal or has suspicious deviations from the normal state, and determine the suspicious time window. S300: For a received suspicious time window, obtain the original log sequence and original indicator sequence within the window; S400: Input the original log sequence and the original indicator sequence into the dual-branch deep multimodal spatiotemporal prediction model. Extract the global log semantic vector and local temporal features through the log coding branch and the indicator coding branch respectively. Then perform cross-modal fusion verification through the cross attention layer and output the fusion vector of multimodal features. S500: The fusion vector is input into the fully connected layer for linear transformation and nonlinear activation. The output result is normalized by the normalized exponential function. The real-time classification status of the current time window interface is output according to the maximum value index strategy. The real-time classification status includes three states: normal, sub-healthy, and warning. as well as S600: The real-time classification status is converted into a graded early warning signal through mapping rules, and the early warning signal level, early warning information, time window start and end time and indicator value are packaged and encapsulated into a structured object and sent to the alarm system.

2. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 1, characterized in that, In step S100, the timing component within each time window is composed of a three-dimensional sub-vector consisting of the average response time, average latency, and average error rate of all interface requests within the window. This is obtained by statistically calculating the average of the sampled data of the corresponding performance indicators within the current window. The log component is obtained by templated log messages in the current window, statistically analyzing the distribution of various log templates in the current window, and then converting them into a sparse vector of fixed dimensions using TF-IDF weights. The time-series component and log component of the current time window are concatenated to construct the multimodal fingerprint vector of the current time window, and the interface's running status indicator data and business logic log data are strongly correlated at the semantic level.

3. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 2, characterized in that, For the statistics of log templates within each time window, each log message is templated, and the variable values, IDs, IP addresses, and timestamps are replaced with fixed placeholders <*>, while the unchanging words and structure are preserved. Similar logs are grouped into the same log template and assigned a unique template identifier. Collect all unique log templates to form a set; The frequency of various log templates appearing within each time window is statistically analyzed to obtain log template distribution data.

4. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 1, characterized in that, In step S200, based on the system's steady-state reference vector, the long-term dynamic baseline vector is asynchronously updated using EWMA: Multiply the current baseline vector by a preset first weighting factor, multiply the steady-state reference vector by a second weighting factor, and add the products of the two to obtain the updated long-term dynamic baseline vector. Wherein, the second weighting factor is the learning rate of receiving new normal data, and the sum of the first weighting factor and the second weighting factor is one; The stable state reference vector refers to a verified window fingerprint vector that represents the system in a healthy operating state. It is the only legitimate data source for baseline updates, and the multimodal fingerprint vector is only marked as a stable state reference vector and allowed to participate in baseline updates when the multimodal fingerprint vector in the current time window is determined to be non-abnormal.

5. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 1, characterized in that, In step S200, during the cold start window of the system's initial operation, the multimodal fingerprint vectors of all time windows are collected and the arithmetic mean is calculated as the initial baseline vector; Calculate the standard fractional variant and coefficient of variation of the mean response time for each time window within the cold start window; wherein, the standard fractional variant is obtained by calculating the absolute difference between the mean response time of the current time window and the historical global mean, and then dividing it by the global standard deviation; the coefficient of variation is obtained by calculating the ratio between the standard deviation of the response time within the current time window and the mean response time; When the standard score variant is less than or equal to a preset significant deviation threshold and the coefficient of variation is not significantly higher than the average level of the cold start window, the current time window is determined to be normal, and its corresponding multimodal fingerprint vector is marked to construct a stable state reference vector.

6. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 1, characterized in that, In step S200, the deviation is determined by calculating the Mahalanobis distance or Euclidean distance between the multimodal fingerprint vector of the current time window and the long-term dynamic baseline vector, and compared with the dynamic threshold. If the deviation exceeds the dynamic threshold, the time window is determined to be a suspicious window. The dynamic threshold is determined based on the historical mean and historical standard deviation of Mahalanobis distance or Euclidean distance within a preset historical time period. The historical standard deviation is multiplied by the dynamic sensitivity coefficient and then added to the historical mean to obtain the dynamic threshold.

7. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 1, characterized in that, In step S400, the dual-branch deep multimodal spatiotemporal prediction model includes a log coding branch, an index coding branch, a cross-modal cross-attention layer, and a fully connected layer; The log encoding branch receives the original log message input extracted within the suspicious time window, performs word segmentation and embedding processing through a pre-trained distributed representation model, and obtains a global log semantic vector representing the business logic of the entire time window after mean pooling. The index encoding branch has a one-dimensional convolutional layer and a connected gated recurrent unit; the one-dimensional convolutional layer is used for local pattern extraction, and the original time series of the index, including response time, latency and error rate, is input into the one-dimensional convolutional layer to extract local temporal features; then it is input into the gated recurrent unit for long-range dependency capture, and the output is an index feature vector representing the dynamic trend of the index throughout the time window. The gated recurrent unit adopts a bidirectional gated recurrent unit layer. The cross-modal attention layer uses the log semantic vector as the query matrix and the indicator feature vector as the key matrix and value matrix, respectively, and inputs them into the cross-modal attention layer for cross-modal fusion and verification. By calculating the matching degree between the query matrix and the indicator feature vector, an attention weight is obtained, and the indicator feature vector is weighted based on the attention weight to obtain a fusion vector that integrates multimodal features.

8. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to claim 7, characterized in that, In step S400, the multimodal feature fusion process based on cross-modal cross-attention includes: The log semantic vector L is transformed linearly to generate the query matrix Q, and the indicator feature vector M is transformed into the key matrix K and the value matrix V respectively. By calculating the transpose matrix K of the query matrix Q and the key matrix K. T The dot product is calculated and then divided by a preset scaling factor to obtain the attention score matrix. The attention scoring matrix is ​​normalized using a normalized exponential function to obtain attention weights, which are used to enhance the semantically relevant parts of the indicator features. Then, the attention weights are used to sum the indicator feature vector M to obtain the fusion vector F, so that the model can automatically focus on the relevant semantic paragraphs in the log sequence when the indicator deviates.

9. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to any one of claims 1-8, characterized in that, The method further includes: performing interpretability processing on the output interface status warning results, specifically including the following steps: The attention weights calculated by the cross-modal attention layer are recorded and averaged on the time axis to obtain the average attention score at each time point, and then arranged in descending order. Find the physical moments when the indicators corresponding to the top few indices with the highest average attention scores shifted. Retrieve original log messages within a preset time deviation range before and after the physical time; if a corresponding log exists, extract its corresponding log template; The start and end times of the current time window, the indicator name, the current value of the indicator, and the retrieved log templates or logs without relevant tags are packaged to generate structured early warning information.

10. The interface anomaly analysis method based on time-series-log fingerprinting and multimodal fusion according to any one of claims 1-8, characterized in that, The distributed representation model, one-dimensional convolutional layer, bidirectional gated recurrent unit layer, and fully connected layer of the deep multimodal network model are parameterized during the offline training phase, specifically including the following process: Collect historical multimodal sample pairs and complete max-min normalization and lexical encoding, and simultaneously label the corresponding normal, sub-healthy, and warning status tags; The difference between the predicted probability and the true label in the current iteration is calculated using the weighted cross-entropy loss function. After the index encoding branch and fully connected layer converge, the network parameters of the distributed representation model are unfrozen, and the entire network parameters are backpropagated and the parameter gradients are updated with a fine-tuned learning rate lower than the initial learning rate until the loss function of the model on the validation set no longer decreases for several consecutive cycles. The training is then stopped, the optimal weight parameters are saved, and the solidified deep multimodal network model is obtained.