Terminal operation and maintenance management method and platform for medical big data platform

Through the deep sequence learning model, the multimodal operation and maintenance data of the medical big data platform is analyzed, and the problems of high false alarm rate and low detection rate in traditional methods are solved, and the comprehensive perception of system status and dynamic abnormality detection are achieved, which improves operation and maintenance efficiency and service quality.

CN120371581AActive Publication Date: 2025-07-25WUHAN SHENGBOHUI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510454752.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The operation and maintenance methods of traditional medical big data platforms have problems with high false alarm rates and low detection rates. They cannot effectively capture the comprehensive analysis ability of multi-dimensional data and the deep understanding of time series data, and are difficult to adapt to the evolution of system behavior, resulting in operation and maintenance personnel being tired of dealing with false alarms and missing out on real system anomalies.

Method used

The deep sequence learning model is adopted, by collecting multimodal historical operation and maintenance data, building a historical timing feature vector sequence, training a deep sequence learning model to capture the context state of the normal workflow mode, and using the baseline model to calculate reconstruction errors, combining potential context representations and abnormal mode criterions to generate early warning signals.

Benefits of technology

It significantly improves the accuracy and intelligence level of abnormal detection, reduces the false alarm rate, improves the detection rate of real abnormalities, can identify complex time dependencies and seasonal patterns, dynamically adjusts abnormal judgment standards, reduces the work burden of operation and maintenance personnel, and improves system reliability and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371581A_ABST
    Figure CN120371581A_ABST
Patent Text Reader

Abstract

The invention provides a terminal operation and maintenance management method and platform for a medical big data platform, and the method comprises the steps: firstly collecting multi-mode historical operation and maintenance data of a platform terminal in a normal working state, and constructing a time sequence feature vector sequence after preprocessing and fusion; a deep sequence learning model training process is then utilized to learn the potential representation to capture the context state of the normal workflow mode and establish a reconstructed baseline model. Acquiring real-time operation and maintenance data of the terminal, constructing a feature vector, extracting potential context representation through the trained model, and calculating a reconstruction error; a current workflow state is identified based on the potential representation, and a context-aware anomaly metric value is calculated in conjunction with state information and reconstruction errors. And analyzing the time evolution characteristic of the abnormal metric value and comparing the time evolution characteristic with a preset abnormal mode criterion to judge whether the terminal has workflow abnormality, and if so, generating an early warning signal containing abnormal evolution characteristic description. The method has the effect of improving the operation and maintenance detection accuracy of the terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of operation and maintenance management, and specifically relates to a terminal operation and maintenance management method and platform for a medical big data platform. Background Art

[0002] Traditional monitoring methods for medical big data platforms usually adopt isolated single-index analysis and lack the comprehensive analysis ability for multi-dimensional data. The health status of the system is often the result of the combined action of multiple indicators. The fluctuation of a single indicator may be normal, but a specific combination pattern of multiple indicators may indicate problems. Secondly, the existing technologies generally lack a deep understanding of time-series data and cannot effectively capture long-term dependencies and seasonal patterns. System behavior usually has complex time correlations, such as daily, weekly, and monthly fluctuations, and these patterns are difficult to describe with simple rules.

[0003] On the other hand, traditional methods are difficult to adapt to the evolution of system behavior. With the development of business, the system load and usage patterns are constantly changing, and fixed monitoring rules quickly become obsolete. Manually adjusting these rules is not only time-consuming and laborious, but also prone to introducing human errors. Moreover, the existing technologies lack an understanding of the system state context, and the same indicator value may have completely different meanings in different scenarios. For example, the normal ranges of CPU usage during peak and off-peak business periods may vary greatly. These limitations have led to the dilemma of high false alarm rates and low detection rates, making operation and maintenance personnel exhausted in dealing with a large number of false alarms, while at the same time they may miss real system anomalies, ultimately affecting service quality and user experience. Summary of the Invention

[0004] The present invention provides a terminal operation and maintenance management method and platform for a medical big data platform to solve the problems of high false alarm rates and low detection rates existing in the traditional operation and maintenance process.

[0005] In a first aspect, the present invention provides a terminal operation and maintenance management method for a medical big data platform, and the method includes the following steps:

[0006] Collect multi-modal historical operation and maintenance data of the platform terminal in a normal workflow mode;

[0007] Perform preprocessing and fusion processing on the multi-modal historical operation and maintenance data to construct a historical time-series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal;

[0008] Use the historical time-series feature vector sequence to train a preset deep sequence learning model. The model training process includes learning the latent representation of the historical time-series feature vector sequence to capture the context state of the normal workflow mode, and establishing a baseline model for reconstructing the historical time-series feature vector sequence based on the latent representation;

[0009] Construct a real-time time series feature vector sequence based on the real-time operation and maintenance data of the currently obtained platform terminal, extract the latent context representation of the real-time time series feature vector sequence using the trained deep sequence learning model, and calculate the reconstruction error of the real-time time series feature vector sequence using the baseline model;

[0010] Identify the inferred workflow state in which the platform terminal is currently located based on the latent context representation, and calculate the context-aware anomaly metric value in combination with the inferred workflow state and the reconstruction error;

[0011] Analyze the time evolution characteristics of the anomaly metric value, and compare it with the preset anomaly pattern criterion according to the time evolution characteristics to determine whether there is a workflow pattern anomaly in the platform terminal. If there is a workflow pattern anomaly in the platform terminal, generate a warning signal including a description of the abnormal evolution characteristics of the platform terminal.

[0012] Optionally, the preprocessing and fusion processing of the multi-modal historical operation and maintenance data to construct a historical time series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal includes the following steps:

[0013] Perform time alignment and data parsing processing on the multi-modal historical operation and maintenance data at a unified sampling time interval, and parse the key information fields of the multi-modal historical operation and maintenance data;

[0014] Calculate and extract the derived correlation features representing the immediate interaction relationship between different modal data;

[0015] Encode and fuse the key information fields and the derived correlation features into a single high-dimensional feature vector, and arrange the high-dimensional feature vectors in chronological order to form a historical time series feature vector sequence.

[0016] Optionally, the training of the preset deep sequence learning model using the historical time series feature vector sequence includes the following steps:

[0017] Use the encoder in the preset deep sequence learning model to map the historical time series feature vector sequence to the latent representation space;

[0018] Train the deep sequence learning model to perform multi-task learning. The learning tasks at least include: reconstructing the original input sequence based on the decoder in the deep sequence learning model and according to the latent context representation in the latent representation space to minimize the model reconstruction error; predicting the time series feature vector of the next time step based on the latent context representation to minimize the model prediction error;

[0019] After the deep sequence learning model completes multi-task learning, retain the encoder for real-time extraction of the latent context representation, and retain the reconstruction task part as the baseline model.

[0020] Optionally, the step of identifying the current inference workflow state of the platform terminal based on the potential context representation includes the following steps:

[0021] Perform clustering analysis on the potential context representation in the potential representation space, and define discrete inference workflow state clusters according to the clustering analysis results;

[0022] Calculate the spatial distance between the potential context representation of the real-time time series feature vector sequence and the centers of each inference workflow state cluster;

[0023] Statistically analyze the local density and neighboring sample distribution characteristics of the potential context representation of the real-time time series feature vector sequence in the potential representation space;

[0024] Assign the current target inference workflow state to the platform terminal by combining the spatial distance, local density, and neighboring sample distribution characteristics.

[0025] Optionally, the step of calculating the context-aware anomaly metric value by combining the inference workflow state and the reconstruction error includes the following steps:

[0026] For each inference workflow state, calculate and store the multi-dimensional error statistical model in the inference workflow state based on the reconstruction error of the samples belonging to the inference workflow state in the historical time series feature vector sequence;

[0027] According to the currently assigned target inference workflow state, look up the corresponding target multi-dimensional error statistical model;

[0028] Compare the reconstruction error of the real-time time series feature vector sequence with the target multi-dimensional error statistical model, calculate the statistical significance and unexpectedness in the target inference workflow state, and calculate the context-aware anomaly metric value by combining the statistical significance and unexpectedness.

[0029] Optionally, the step of analyzing the time evolution characteristics of the anomaly metric value and comparing it with the preset anomaly pattern criterion to determine whether there is an anomaly in the workflow pattern of the platform terminal includes the following steps:

[0030] Maintain the anomaly metric value as a sequence of continuous anomaly metric values within a time window;

[0031] Calculate the dynamic evolution index of the anomaly metric value sequence, and the dynamic evolution index includes trend slope, fluctuation amplitude change rate, and autocorrelation change;

[0032] Select the corresponding target evolution template from the predefined normal evolution pattern template library according to the target inference workflow state. The evolution templates in the normal evolution pattern template library characterize the normal fluctuation range and dynamic characteristics of the anomaly metric value in different states;

[0033] Compare the dynamic evolution indicators with the target evolution template, and calculate the comprehensive deviation degree of the dynamic evolution indicators from the target evolution template;

[0034] When the comprehensive deviation degree exceeds the preset threshold, it is determined that there is an abnormal workflow mode in the platform terminal, and the comprehensive deviation degree and the target evolution template are used as the description of the abnormal evolution characteristics.

[0035] Optionally, after generating the warning signal including the description of the abnormal evolution characteristics of the platform terminal, the following steps are further included:

[0036] Integrate the description of the abnormal evolution characteristics and the current inferred workflow state of the platform terminal into comprehensive abnormal characteristics;

[0037] Based on the reconstruction error analysis, determine the original feature dimension in the real-time time series feature vector sequence that has the highest correlation with the comprehensive abnormal characteristics;

[0038] Combine the comprehensive abnormal characteristics and the original feature dimension, and use the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility.

[0039] Optionally, the step of combining the comprehensive abnormal characteristics and the original feature dimension and using the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility includes the following steps:

[0040] Use the comprehensive abnormal characteristics and the original feature dimension as query conditions;

[0041] Retrieve the target nodes and target paths with the highest correlation with the query conditions in the pre-constructed platform association knowledge graph. The target nodes include terminal components, terminal service ports, or terminal resource ports, and the target paths represent the dependency or influence relationship between the target nodes;

[0042] Take the target nodes as potential fault sources, calculate the correlation scores of each potential fault source according to the target paths, and sort the potential fault sources by the correlation scores.

[0043] In a second aspect, the present invention also provides a terminal operation and maintenance management platform for a medical big data platform, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the terminal operation and maintenance management method for the medical big data platform as described in the first aspect.

[0044] In a third aspect, the present invention also provides a computer-readable storage medium, on which instructions are stored. When the instructions are executed by a processor, the processor is configured to execute the terminal operation and maintenance management method for the medical big data platform as described in the first aspect.

[0045] The beneficial effects of the present invention are as follows:

[0046] Through deep sequence learning and context-aware analysis, the present invention significantly improves the accuracy and intelligence level of anomaly detection in platform terminals, effectively solving many challenges faced by traditional monitoring systems. Compared with traditional methods that rely on fixed thresholds and static rules, the present invention can adaptively learn the normal behavior patterns of the system and establish a dynamic baseline, thereby greatly reducing the false alarm rate and increasing the detection rate of real anomalies. By processing multi-modal data fusion, the present invention overcomes the limitations of single-index monitoring and realizes a comprehensive perception of the system state. Especially in capturing temporal features, the present invention can identify complex time-dependent relationships and seasonal patterns, which are difficult to achieve by traditional methods. In addition, the context-aware ability of the present invention enables it to dynamically adjust the anomaly judgment criteria according to the workflow state of the system, avoiding the "one-size-fits-all" problem in traditional methods. In practical applications, the present invention can not only detect anomalies in a timely manner, but also provide descriptions of anomaly evolution characteristics, providing more valuable decision-making support for operation and maintenance personnel. This intelligent anomaly detection method significantly reduces the workload of operation and maintenance personnel, shortens the problem response time, improves system reliability, provides strong technical support for the digital transformation of enterprises, and ultimately realizes the dual improvement of operation and maintenance efficiency and service quality. Description of the Drawings

[0047] Figure 1 It is a schematic flowchart of a terminal operation and maintenance management method for a medical big data platform in one embodiment of the present application. Detailed Embodiments

[0048] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0049] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0050] Figure 1It is a schematic flowchart of a terminal operation and maintenance management method for a medical big data platform in an embodiment. It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in Figure 1 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps. As Figure 1 shown, a terminal operation and maintenance management method for a medical big data platform disclosed by the present invention specifically includes the following steps:

[0051] S101. Collect multi-modal historical operation and maintenance data of the platform terminal in the normal workflow mode.

[0052] Among them, in the terminal operation and maintenance management of the medical big data platform, it is first necessary to collect multi-modal historical operation and maintenance data of the platform terminal in the normal workflow mode. These data include but are not limited to: system performance indicators such as CPU usage rate, memory occupancy, disk I / O rate, network traffic, application response time, database query latency, and number of user sessions; at the same time, it also includes text data such as error information, warning information, and operation records in the log file; as well as user interaction behavior data, business process execution data, etc. The collection process is implemented through a monitoring agent program deployed on the terminal, and a data point is recorded at a fixed time interval (such as every 5 minutes) to ensure the continuity and timeliness of the data. For example, for a medical imaging processing terminal, the CPU usage rate curve, memory occupancy change, network transmission rate, and application log when processing imaging tasks will be recorded simultaneously. These data together constitute the complete behavior characteristics of the terminal in the normal working state, providing a reference benchmark for subsequent anomaly detection.

[0053] S102. Perform preprocessing and fusion processing on the multi-modal historical operation and maintenance data to construct a historical time series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal.

[0054] Among them, for the preprocessing and fusion processing of the collected multi-modal historical operation and maintenance data, time alignment needs to be carried out first, unifying data from different sources and different sampling frequencies to the same time scale, such as resampling all data to one data point every 5 minutes. Then data cleaning is performed, including removing outliers (such as incorrect data with CPU usage exceeding 100%), filling in missing values (using methods such as linear interpolation or previous value filling), and standardization processing (normalizing data with different dimensions to the same range). Next, key information fields are extracted, such as extracting structured information such as error types and severity levels from system logs. Further, derivative correlation features are calculated, such as the ratio of CPU usage to memory occupancy, the correlation between network traffic and the number of database queries, etc. These features can reflect the interaction relationships between different modal data. Finally, all processed features are encoded and fused into a high-dimensional feature vector (such as a 300-dimensional vector), and arranged in chronological order to form a historical time series feature vector sequence, which comprehensively reflects the dynamic behavior characteristics of the platform terminal during historical normal operation.

[0055] S103. Use the historical time series feature vector sequence to train a preset deep sequence learning model. The model training process includes learning the latent representation of the historical time series feature vector sequence to capture the context state of the normal workflow pattern, and establishing a baseline model for reconstructing the historical time series feature vector sequence based on the latent representation.

[0056] Among them, use the constructed historical time series feature vector sequence to train a preset deep sequence learning model. This model usually adopts LSTM (Long Short-Term Memory Network) or Transformer architecture, which can effectively capture the long-term dependencies in time series data. The training process first maps the input historical time series feature vector sequence to a low-dimensional latent representation space through an encoder (such as compressing the 300-dimensional original features into 50-dimensional latent representations), and these latent representations can capture the context state information of the normal workflow pattern. The model adopts a multi-task learning framework and simultaneously performs two key tasks: one is the reconstruction task, which remaps the latent representation back to the original feature space through a decoder and minimizes the reconstruction error (using the mean square error loss function); the other is the prediction task, which predicts the feature vector of the next time step based on the current latent representation. For example, for a medical image processing terminal, the model can learn the normal change patterns of indicators such as CPU, memory, and disk I / O when processing large image files. After training is completed, the encoder is retained for real-time feature extraction, and the reconstruction part is retained as the baseline model for subsequent anomaly detection.

[0057] S104. Construct a real-time time series feature vector sequence based on the currently obtained real-time operation and maintenance data of the platform terminal, use the trained deep sequence learning model to extract the latent context representation of the real-time time series feature vector sequence, and calculate the reconstruction error of the real-time time series feature vector sequence using the baseline model.

[0058] Among them, according to the real-time operation and maintenance data of the current platform terminal obtained, a real-time time series feature vector sequence is constructed according to the same preprocessing process as the historical data. Specifically, multi-modal operation and maintenance data of the terminal are collected every fixed time interval (such as 5 minutes), and time alignment, data cleaning, feature extraction and fusion are performed to form real-time feature vectors with the same format as the historical data. Then, these real-time feature vectors are input into the encoder of the trained deep sequence learning model to extract their latent context representations. For example, for a medical information system terminal that is performing a database backup task, the encoder will compress its current multi-dimensional features such as CPU usage, memory occupancy, and disk write rate into a compact latent representation. Then, the reserved baseline model (decoder part) is used to try to reconstruct the original real-time feature vectors from these latent representations, and the error (such as Euclidean distance or cosine similarity) between the reconstruction result and the actual observed value is calculated. These reconstruction errors reflect the degree of deviation of the current terminal behavior from the historical normal mode, providing a basic measure for anomaly detection.

[0059] S105. Identify the inferred workflow state in which the platform terminal is currently located based on the latent context representation, and calculate a context-aware anomaly metric value by combining the inferred workflow state and the reconstruction error.

[0060] Among them, based on the extracted latent context representation, first, the latent representation space is divided into multiple discrete workflow state clusters through clustering analysis (such as K-means or DBSCAN algorithms), and each cluster represents a typical working mode (such as data processing state, idle state, backup state, etc.). Then, calculate the spatial distance between the real-time latent representation and the center of each state cluster, and analyze its local density (the density of surrounding sample points) and the distribution characteristics of neighboring samples (such as the class distribution of the nearest neighbor samples). Based on these features, assign the most matching inferred workflow state to the current terminal. For example, if the current latent representation of a medical imaging terminal is closest to the "high-load image processing" state cluster, it is identified as this workflow state. Then, look up the error distribution model corresponding to this state from the pre-established multi-dimensional error statistical model library, compare the current reconstruction error with this model, and calculate the statistical significance (such as z-score) and the degree of unexpectedness (such as Mahalanobis distance). Finally, calculate a context-aware anomaly metric value by combining these metrics. This value not only considers the absolute magnitude of the reconstruction error but also the normal error fluctuation range in the current workflow state, providing a more accurate anomaly assessment.

[0061] S106. Analyze the time evolution characteristics of the abnormal metric values, compare them with the preset abnormal pattern criteria according to the time evolution characteristics, and determine whether there is an abnormal workflow pattern in the platform terminal. If there is an abnormal workflow pattern in the platform terminal, generate a warning signal containing a description of the abnormal evolution characteristics of the platform terminal.

[0062] Among them, the continuously calculated abnormal metric values are maintained as a sequence within a fixed time window (such as the most recent 30 minutes), and its time evolution characteristics are analyzed. Specifically, three types of dynamic evolution indicators are calculated: trend slope (using linear regression to calculate the rising or falling trend of the abnormal metric values), change rate of fluctuation amplitude (calculating the expansion or contraction rate of the fluctuation range of the abnormal metric values), and autocorrelation change (analyzing the periodicity and persistence of the abnormal metric value sequence). According to the currently inferred workflow state, select the corresponding template from the predefined normal evolution pattern template library. These templates define the normal fluctuation range and dynamic characteristics of the abnormal metric values under different working states. For example, for the "database backup" state, its normal template may show a pattern where the abnormal metric value first rises rapidly and then falls slowly. Compare the calculated dynamic evolution indicators with the template and calculate the comprehensive deviation degree. When the deviation degree exceeds the preset threshold (such as the 95% confidence interval), it is determined that there is an abnormal workflow pattern in the terminal, and a warning signal containing a description of the abnormal evolution characteristics (such as "the abnormal metric value continues to rise and the fluctuation amplitude expands abnormally") is generated for the operation and maintenance personnel to further analyze and process.

[0063] In one implementation manner, the preprocessing and fusion processing of the multi-modal historical operation and maintenance data to construct a historical time series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal includes the following steps:

[0064] Perform time alignment and data parsing processing on the multi-modal historical operation and maintenance data at a unified sampling time interval, and parse out the key information fields of the multi-modal historical operation and maintenance data;

[0065] Calculate and extract the derived correlation features characterizing the immediate interaction relationship between different modal data;

[0066] Encode and fuse the key information fields and the derived correlation features into a single high-dimensional feature vector, and arrange the high-dimensional feature vectors in chronological order to form a historical time series feature vector sequence.

[0067] In this embodiment, the terminal operation and maintenance data sources of the medical big data platform are diverse and the collection frequencies are different. Therefore, time alignment processing needs to be carried out first. During specific implementation, a unified sampling time interval (such as 5 minutes) is selected, and all data is resampled to this time scale. For example, the CPU usage rate may be collected every 1 minute, while the log data may be generated irregularly. Through time alignment, they are unified onto the same time axis. For high-frequency data, aggregation methods such as taking the average value or the maximum value are used for downsampling; for low-frequency data, methods such as linear interpolation are used for interpolation and supplementation. After time alignment, data parsing processing is carried out to convert the original data in different formats into structured information. For example, for system performance indicators, the numerical value and its unit are extracted; for log texts, key information fields such as error types, severity levels, and affected components are extracted using regular expressions or natural language processing techniques; for network traffic data, fields such as source address, target address, and transmission protocol are parsed. This step ensures the consistency of data from different sources in the time dimension and converts unstructured or semi-structured data into a structured form that can be quantitatively analyzed, laying a foundation for subsequent feature extraction.

[0068] After obtaining the key information fields, it is necessary to further explore the interaction relationships between different modal data to generate derived correlation features. These features can capture the overall state of the system that cannot be reflected by single-modal data. During specific implementation, first, basic statistical correlation features are calculated, such as the ratio of CPU usage rate to memory occupancy rate, which reflects the balance state of computing resources; the correlation coefficient between network traffic and the number of database queries, which reflects the coordination of data transmission and processing. Secondly, time window features are constructed, such as calculating the lag correlation between disk I / O and CPU usage rate in the past 30 minutes to capture the causal relationship of resource usage. Thirdly, pattern conversion features are extracted, such as detecting the conversion pattern from low network traffic to high CPU usage, which may indicate the processing stage after data download. Composite metrics can also be calculated, such as the "resource pressure index" (weighted average combining CPU, memory, and disk I / O). For the specific business logic of the medical platform, features such as "imaging processing efficiency" (the ratio of the number of processed images to resource consumption) can be defined. These derived correlation features greatly enrich the data expression ability, enabling the system to understand the complex working state of the terminal and the mutual influence between modalities, providing a more comprehensive perspective for anomaly detection.

[0069] Encode and fuse the key information fields and derivative correlation features obtained in the first two steps to construct a unified feature representation. For numerical features (such as CPU usage rate, memory occupancy rate, etc.), perform standardization processing to map them to the interval [0, 1] or a standard normal distribution, eliminating the dimension difference. For categorical features (such as error types, component names, etc.), use one-hot encoding or embedding encoding to convert them into numerical vectors. For text features (such as log messages), use TF-IDF or word embedding techniques to convert them into vectors of a fixed dimension. Concatenate all the encoded features in a predefined order to form a single high-dimensional feature vector (such as 300 dimensions). For example, the feature vector of a medical image processing terminal at a certain moment may include: the first 100 dimensions are system performance indicators, the middle 100 dimensions are log text features, and the last 100 dimensions are derivative correlation features. Arrange the high-dimensional feature vectors at different time points in chronological order to form a historical time-series feature vector sequence, such as {V1, V2,..., V n}, where each Vi represents a complete feature vector at a time point. This unified time-series feature representation retains the time evolution characteristics of the original data and integrates multi-modal information, providing structured input for subsequent deep sequence learning models and enabling the models to learn the complex dynamic behavior patterns in the normal operating state of the terminal.

[0070] In one implementation, the training of a preset deep sequence learning model using the historical time-series feature vector sequence includes the following steps:

[0071] Use the encoder in the preset deep sequence learning model to map the historical time-series feature vector sequence to the latent representation space;

[0072] Train the deep sequence learning model to perform multi-task learning, and the learning tasks at least include: based on the decoder in the deep sequence learning model and according to the latent context representation in the latent representation space, reconstruct the original input sequence to minimize the model reconstruction error; based on the latent context representation, predict the time-series feature vector of the next time step to minimize the model prediction error;

[0073] After the deep sequence learning model completes multi-task learning, retain the encoder for real-time extraction of the latent context representation, and retain the reconstruction task part as the baseline model.

[0074] In this embodiment, the encoder part of the deep sequence learning model is responsible for compressing and mapping the high-dimensional historical time-series feature vector sequence into a low-dimensional latent representation space, achieving data dimensionality reduction and key information extraction. Specifically, when implementing, a bidirectional long short-term memory network (BiLSTM) or a Transformer encoder architecture is adopted, and these architectures can effectively capture the long-term dependencies and context information in the sequence data. Taking BiLSTM as an example, the input layer receives a sequence of historical time-series feature vectors {X1, X2,..., X n}, where each X i is a d-dimensional vector (e.g., d = 300). The BiLSTM layer contains two LSTM units, forward and backward, which process the data from the beginning to the end and from the end to the beginning of the sequence respectively. Each direction's LSTM unit contains h hidden neurons (e.g., h = 128). The outputs of the two directions are concatenated at each time step to form a 2h-dimensional hidden state. Subsequently, through a fully connected layer, the 2h-dimensional hidden state is mapped to a k-dimensional latent representation space (e.g., k = 50), obtaining the sequence {Z1, Z2,..., Z n}, where each Z i is a k-dimensional vector, representing the latent context representation at time point i. This mapping significantly reduces the data dimensionality (from 300 to 50), while retaining the key time-series patterns and state information in the original sequence, providing a compact and information-rich representation for subsequent tasks.

[0075] The deep sequence learning model simultaneously performs two key tasks, reconstruction and prediction, through a multi-task learning framework, enabling the model to comprehensively understand the internal structure and time-series evolution law of the data. In the reconstruction task, the decoder receives the sequence of latent context representations {Z1, Z2,..., Z n} output by the encoder and attempts to reconstruct the original input sequence {X1, X2,..., X n}. The decoder adopts a structure symmetric to the encoder, such as an LSTM layer followed by a fully connected layer, mapping the k-dimensional latent representation back to the d-dimensional original feature space, obtaining the reconstructed sequence {X1, X2,..., X n}. The reconstruction loss function uses the mean squared error (MSE), and the calculation formula is L_recon = (1 / n)∑ i ||X i - X i || 2 . At the same time, in the prediction task, another prediction network (usually an LSTM or a feed-forward network) receives the latent representation Z t at time point t and predicts the feature vector X t+1 at the next time point t + 1. The prediction loss function also uses MSE, and the calculation formula is L_pred = (1 / n - 1)∑ i ||Xi+1 -X i+1 || 2 The final total loss function is the weighted sum of two parts: L_total = αL_recon + βL_pred, where α and β are weight hyperparameters (e.g., α = 0.7, β = 0.3). The model is trained using the Adam optimizer with a learning rate of 0.001, a batch size of 64, and 100 training epochs.

[0076] After the deep sequence learning model completes multi-task learning training, the model structure is decomposed into two key components for the real-time anomaly detection system. First, the trained encoder part is retained, which has learned how to effectively compress high-dimensional raw feature vectors into low-dimensional latent context representations. During the real-time operation phase, the encoder receives the real-time feature vector sequence {X_t-w+1, X_t-w+2,..., X_t} within the current time window (where w is the window size, e.g., w = 10) and outputs the corresponding latent context representation sequence {Z_t-w+1, Z_t-w+2,..., Z_t}. These latent representations capture the essential features of the current system state and contain temporal context information. At the same time, the network components related to the reconstruction task (i.e., the decoder part) are retained as the baseline model to evaluate the deviation of the current system state from the normal mode. Specifically, the baseline model receives the latent representation Z_t output by the encoder, attempts to reconstruct the raw feature vector X_t, and calculates the reconstruction error e_t = ||X_t - X_t||. This decomposition design enables the anomaly detection system to have an efficient modular structure: the encoder is responsible for feature extraction and dimensionality reduction, providing a compact state representation; the baseline model is responsible for anomaly evaluation, quantifying the anomaly degree of the system behavior through the reconstruction error. These two components together constitute the core engine of anomaly detection, which can monitor the system state in real time and identify potential anomalies.

[0077] In one implementation manner, the method for identifying the current inferred workflow state of the platform terminal based on the latent context representation includes the following steps:

[0078] Perform clustering analysis on the latent context representations in the latent representation space, and define discrete inferred workflow state clusters according to the clustering analysis results;

[0079] Calculate the spatial distances between the latent context representations of the real-time temporal feature vector sequence and the centers of each inferred workflow state cluster;

[0080] Statistically analyze the local density and neighboring sample distribution characteristics of the latent context representations of the real-time temporal feature vector sequence in the latent representation space;

[0081] Allocate the current target inferred workflow state for the platform terminal by combining the spatial distance, local density, and neighboring sample distribution characteristics.

[0082] In this embodiment, a large number of latent context representation samples are extracted from historical normal operation data by using a trained encoder, and these samples form a point cloud in a k-dimensional latent representation space (such as k = 50). The Gaussian Mixture Model (GMM) clustering algorithm is applied to these point cloud data. This algorithm assumes that the data is generated by a mixture of multiple Gaussian distributions and can capture the probability distribution characteristics of different working states. In specific implementation, first, the optimal number of clusters m (such as m = 8) is determined by methods such as the Bayesian Information Criterion (BIC) or the silhouette coefficient, which represents the possible number of main workflow states of the medical platform terminal. Subsequently, the GMM algorithm estimates the parameters of each Gaussian distribution through the Expectation-Maximization (EM) iterative optimization process, including the mean vector μ i (representing the i-th cluster center), the covariance matrix Σ i (representing the shape and orientation of the cluster), and the mixing weight π i (representing the prior probability of this cluster). Each cluster represents a specific workflow state, such as "image data processing", "patient information query", "system idle", etc. For easy understanding and interpretation, the k-dimensional clustering results can be visualized as two-dimensional or three-dimensional graphs through dimensionality reduction techniques such as t-SNE or PCA, and semantic labels are assigned to each cluster in combination with domain knowledge. This clustering analysis divides the continuous latent representation space into a finite number of discrete workflow state clusters, providing a structured reference framework for subsequent state inference and anomaly detection, enabling the system to understand the working state of the terminal at different times.

[0083] In the real-time monitoring stage, the system continuously receives the sequence of time-series feature vectors of the terminal and maps it to the latent representation space through a pre-trained encoder to obtain the current latent context representation Z_current. To determine the relationship between Z_current and each predefined workflow state cluster, it is necessary to calculate its spatial distance from each cluster center. Considering the non-uniformity of the latent representation space and the shape differences of each cluster, the Mahalanobis distance is used as the distance metric, which takes into account the covariance structure of the data. For the i-th cluster, the Mahalanobis distance calculation formula is: d_i = √[(Z_current - μ_i)^TΣ_i^(-1)(Z_current - μ_i)], where μ_i is the center vector of the i-th cluster and Σ_i is its covariance matrix.

[0084] For example, assume that a medical terminal is currently processing a large amount of imaging data. The Mahalanobis distance between its potential context representation Z_current and the "imaging processing" cluster center may be 1.2, while the distance from the "patient information query" cluster center is 5.7. The system calculates the distances from Z_current to all m cluster centers, forming a distance vector D = [d_1, d_2,..., d_m]. These distance values not only reflect the similarity between the current state and various typical workflow states but also provide an important basis for subsequent state assignment. The distance calculation results can be converted into similarity scores, such as through the Gaussian kernel function s_i = exp(-d_i^2 / 2σ^2), where σ is the bandwidth parameter that controls the rate of similarity decay with distance. This distance-based similarity quantifies the matching degree between the current state and each workflow pattern, providing the primary reference index for workflow state inference.

[0085] In addition to calculating the distances to the cluster centers, it is also necessary to analyze the local characteristics of the current potential context representation Z_current in the potential representation space, which helps to evaluate whether it belongs to a normal workflow pattern or a potential outlier. First, calculate the local density ρ of Z_current, which represents the sample density in its surrounding area. The local density can be obtained through the kernel density estimation method: ρ = (1 / n) ∑_i K((Z_current - Z_i) / h), where Z_i is a sample point in the historical sample set, K is the kernel function (such as the Gaussian kernel), h is the bandwidth parameter, and n is the total number of samples. A high local density indicates that the current state is located in the high-frequency region of the historical data and may represent a common workflow pattern; a low local density may imply an abnormal or rare state.

[0086] Second, analyze the distribution characteristics of the k-nearest neighbor samples of Z_current (e.g., k = 20), including: (1) the distribution of the cluster labels of the nearest neighbor samples, calculate the occurrence frequencies of each cluster label to form a probability vector P = [p_1, p_2,..., p_m], where p_i represents the proportion of the i-th cluster label among the nearest neighbors; (2) the temporal distribution of the nearest neighbor samples, check whether these samples come from different time periods or are concentrated in a specific period; (3) the spatial dispersion of the nearest neighbor samples, quantified by calculating the average distance or variance of the nearest neighbor samples. This local characteristic analysis provides more fine-grained context information for workflow state inference, enabling the system to understand the position and meaning of the current state in the overall workflow pattern.

[0087] Based on the multi-dimensional information obtained from the first three steps, a comprehensive decision-making mechanism is used to assign the most suitable workflow state to the terminal. First, a state assignment score function is constructed, which comprehensively considers three key factors: the Mahalanobis distance D = [d_1, d_2,..., d_m] from the cluster center, the local density ρ, and the cluster label distribution P = [p_1, p_2,..., p_m] of the neighboring samples. For the i-th cluster, its state assignment score is calculated as: Score_i = w_1·exp(-d_i^2 / σ_1^2) + w_2·ρ·p_i + w_3·(1 - H(P))·p_i, where w_1, w_2, w_3 are weight parameters, σ_1 is the distance normalization parameter, and H(P) is the entropy of the label distribution, which is used to quantify the uncertainty of the distribution.

[0088] The first term of this score function prefers the state close to the cluster center, the second term prefers the state with high local density and large proportion among the neighbors, and the third term gives an extra reward when the neighboring label distribution is concentrated. After calculating the scores of all clusters, two assignment strategies can be adopted: hard assignment - assign the terminal to the workflow state with the highest score; soft assignment - calculate the normalized probability distribution of each state, indicating the possibility that the terminal is in multiple workflow states simultaneously. At the same time, the confidence of the state assignment will also be calculated, such as the difference between the score and the second-highest score or the entropy of the score. Assignments with low confidence will be marked as "transition state" or "mixed state", which is very important for understanding the workflow transition process of the terminal. This comprehensive decision-making mechanism not only considers the similarity between the current state and the typical workflow pattern, but also considers the local data distribution characteristics, can accurately identify the working state of the terminal, and can give a reasonable state inference even in the transition period of workflow conversion or complex scenarios of multi-task parallelism, providing reliable context information for subsequent anomaly detection and performance optimization.

[0089] In one of the implementation manners, the steps for calculating the context-aware anomaly metric value by combining the inferred workflow state and the reconstruction error are as follows:

[0090] For each inferred workflow state, based on the reconstruction errors of the samples belonging to the inferred workflow state in the historical time series feature vector sequence, calculate and store the multi-dimensional error statistical model under the inferred workflow state;

[0091] According to the currently assigned target inferred workflow state, look up the corresponding target multi-dimensional error statistical model;

[0092] Compare the reconstruction error of the real-time time series feature vector sequence with the target multi-dimensional error statistical model, calculate the statistical significance and unexpectedness under the target inferred workflow state, and calculate the context-aware anomaly metric value by combining the statistical significance and unexpectedness.

[0093] In this embodiment, for each identified workflow state cluster (such as "image processing", "patient information query", etc.), an exclusive error statistical model needs to be constructed to characterize the reconstructed error distribution characteristics during normal operation in this state. First, all sample points belonging to a specific workflow state are screened out from the historical data, and these sample points are identified by the labels obtained from the aforementioned clustering analysis. For workflow state i, all its corresponding historical samples {X1 i , X2 i ,..., X n i} are collected. These samples are input into the pre-trained autoencoder to obtain the reconstruction results {X1 i , X2 i ,..., X n i}. The reconstruction error vector e j i = X j i - X j i is calculated, where the dimension of each error vector is the same as that of the original feature vector.

[0094] Considering that the importance and variability of different feature dimensions are different, a multi-dimensional Gaussian distribution is used to model the error distribution instead of simply using scalar errors. For workflow state i, the mean vector μ i = (1 / n)∑ j e j i and the covariance matrix Σ i = (1 / n)∑ j (e j i - μ i )(e j i - μ i ) T . This pair of parameters (μ i , Σ i ) constitutes the multi-dimensional error statistical model of this workflow state, capturing the correlation and variation patterns between different feature dimensions.

[0095] To handle the computational challenges that may be brought by high-dimensional data, the error vectors can be reduced to a lower dimension through principal component analysis (PCA), reducing the computational complexity while retaining the main variation information. In addition, to enhance the robustness of the model to outliers, a pruning strategy can be adopted to remove the 5% samples with the largest reconstruction errors in each workflow state and then calculate the statistical model. Finally, the error statistical model (μ i , Σ i is constructed and stored for each workflow state i.), these models will serve as the benchmarks for subsequent anomaly detection.

[0096] After the real-time monitoring system assigns a target inference workflow state to the terminal, it is necessary to immediately invoke an error statistical model that matches this state for context-related anomaly detection. Specifically, in implementation, first obtain the current workflow state identifier i of the terminal from the previous step (e.g., i = 2, corresponding to the "patient data analysis" state). The system then retrieves the corresponding multi-dimensional error statistical model parameters (μ i , Σ i ) from the pre-constructed model library, where μ i is a d-dimensional mean vector and Σ i is a d×d-dimensional covariance matrix (or a low-dimensional representation after PCA dimensionality reduction).

[0097] Considering that the terminal may be in the transition region of multiple workflow states or performing multiple tasks simultaneously, the system also supports a soft state assignment method. In this case, the terminal is assigned a state probability distribution P = [p1, p2,..., p m , where p i represents the probability that the terminal is in the i-th workflow state. Accordingly, the system will retrieve multiple relevant error statistical models and construct a mixture Gaussian model based on probability weights: μ mix = ∑ i p i μ i , Σ mix = ∑ i p i (Σ i +(μ i -μ mix )(μ i -μ mix ) T ). This mixture model can more accurately represent the normal error distribution of the terminal in the state transition or multi-task scenario. This dynamic model selection mechanism ensures the context relevance of anomaly detection, enabling the system to adjust the anomaly judgment criteria according to the current working state of the terminal and avoiding the high false alarm rate problem that may be brought by static thresholds.

[0098] After obtaining the target error statistical model, the system conducts anomaly assessment on the real-time monitoring data. First, the time series feature vector sequence X_current within the current time window is reconstructed through a pre-trained autoencoder to obtain the reconstructed sequence X_current, and the reconstruction error vector e_current = X_current - X_current is calculated. To comprehensively evaluate the degree of anomaly, it is analyzed from two complementary dimensions: statistical significance and unexpectedness. Statistical significance measures the degree to which the current observation deviates from the normal pattern, which is achieved by calculating the Mahalanobis distance of the reconstruction error vector under the target multi-dimensional Gaussian distribution: MD = √[(e_current - μ_target)^TΣ_target^(-1)(e_current - μ_target)], where μ_target and Σ_target are the parameters of the target error statistical model. The Mahalanobis distance takes into account the variability and correlation of different dimensions and provides a unified measure for multi-dimensional anomalies. For high-dimensional data, the Mahalanobis distance can be converted into a p-value, representing the probability of observing the current or more extreme errors: p = 1 - CDF(χ 2 _d, MD 2 ), where χ 2 _d is the chi-square distribution with degree of freedom d, and CDF is its cumulative distribution function. A smaller p-value (e.g., p < 0.01) indicates a highly statistically significant anomaly.

[0099] Unexpectedness evaluates the novelty of the current error pattern, which is quantified by calculating the local density comparison between the reconstruction error vector and historical error samples: S = -log(ρ_current / ρ_avg), where ρ_current is the local density of the current error vector, and ρ_avg is the average local density of historical errors in the target workflow state. A high degree of unexpectedness indicates that the current error pattern is rare in the historical data and may represent a new type of anomaly.

[0100] Finally, the system combines these two metrics to calculate a context-aware anomaly metric value: A = w1·(-log(p)) + w2·S, where w1 and w2 are weight parameters (e.g., w1 = 0.6, w2 = 0.4). When the anomaly metric value A exceeds a preset threshold (e.g., A > 5), an anomaly alarm is triggered, and at the same time, detailed features of the anomaly are provided, including the feature dimension with the greatest contribution, the temporal pattern of the anomaly, and the similarity analysis with historical anomalies. This multi-dimensional anomaly assessment mechanism can not only detect anomalies that significantly deviate from the normal pattern but also identify potential problems that are not statistically significant but have a novel pattern, greatly improving the accuracy and interpretability of anomaly detection. For example, although the CPU usage rate of a medical terminal in the "image processing" state is within the normal range, the combination with the memory usage pattern is abnormal, and the system can capture this multi-dimensional associated anomaly, while traditional single-dimensional monitoring may ignore it.

[0101] In one implementation, analyzing the time-evolution characteristics of the abnormal metric values and comparing the time-evolution characteristics with a preset abnormal pattern criterion to determine whether there is an abnormal workflow pattern in the platform terminal includes the following steps:

[0102] Maintain the abnormal metric values as a sequence of consecutive abnormal metric values within a time window;

[0103] Calculate the dynamic evolution indicators of the abnormal metric value sequence. The dynamic evolution indicators include trend slope, change rate of fluctuation amplitude, and change of autocorrelation;

[0104] Select the corresponding target evolution template from a predefined normal evolution pattern template library according to the target inferred workflow state. The evolution templates in the normal evolution pattern template library characterize the normal fluctuation range and dynamic characteristics of abnormal metric values in different states;

[0105] Compare the dynamic evolution indicators with the target evolution template and calculate the comprehensive deviation degree of the dynamic evolution indicators deviating from the target evolution template;

[0106] When the comprehensive deviation degree exceeds the preset threshold, it is determined that there is an abnormal workflow pattern in the platform terminal, and the comprehensive deviation degree and the target evolution template are used as descriptions of abnormal evolution characteristics.

[0107] In this implementation, the maintenance of the abnormal metric value sequence adopts a sliding time window mechanism, continuously recording and updating the abnormal metric values A1, A2,..., A at the most recent N time points (e.g., N = 120, corresponding to 2 hours) n . Whenever a new abnormal metric value A is calculated n+1 , add it to the end of the sequence, and at the same time remove the earliest value A1 to keep the window size constant. This sliding window design can not only capture the short-term fluctuations of anomalies but also reflect their long-term evolution trends. To ensure data quality, apply median filtering (window size of 3) to the original abnormal metric values to eliminate instantaneous noise, and perform min-max normalization to map the abnormal metric values to the [0, 1] interval for easy comparison between different time periods and different terminals.

[0108] The sequence storage adopts a circular buffer structure to ensure efficient and fixed memory usage, which is suitable for long-running monitoring systems. For each terminal, additionally maintain the statistical summary of abnormal metric values, including the mean, standard deviation, maximum value, minimum value, and their occurrence time points within the sliding window, to facilitate quick access to key statistical features without repeated calculation. To handle possible data missing (such as monitoring gaps caused by network interruptions), use linear interpolation to fill short-term missing values, and for long-term missing values exceeding a preset threshold (such as 5 consecutive points), mark them as invalid intervals and exclude them in subsequent analysis.

[0109] In addition, for anomaly patterns at different time scales, multiple time windows of different lengths (such as 10 minutes, 2 hours, 24 hours) are maintained simultaneously to form a set of multi-scale anomaly metric value sequences, enabling the system to detect both rapidly emerging anomalies and slowly evolving anomaly patterns. For example, a medical imaging processing terminal may exhibit a resource usage pattern with a morning peak, a lunchtime low, and an afternoon stability on normal working days. By maintaining a sufficiently long time window, this periodic pattern can be distinguished from a true anomaly.

[0110] Based on the maintained anomaly metric value sequences, three key dynamic evolution indicators are calculated to comprehensively characterize the time evolution characteristics of anomaly behaviors. First, the trend slope is obtained by applying linear regression to the anomaly metric value sequence {A1, A2,..., A n} within the time window: where t i is the time point, and t and are the mean values of time and anomaly metrics respectively. To enhance the sensitivity to short-term trends, exponential weighted regression is used, assigning higher weights to recent data: w_i = α(1 - α)^(n - i), where α = 0.1 is the smoothing factor. A positive slope indicates an increasing degree of anomaly, a negative slope indicates a decreasing degree of anomaly, and the absolute value of the slope reflects the rate of change.

[0111] The rate of change of the fluctuation amplitude measures the change in the intensity of anomaly fluctuations and is obtained by calculating the relative change of a sliding variance window. The time window is equally divided into k sub-windows, and the standard deviations σ1, σ2,..., σ k of the anomaly metric values within each sub-window are calculated. Then, the relative change rate of the standard deviations of adjacent sub-windows is calculated: amp_change = (1 / (k - 1)) ∑ i=1 k-1 |(σ i+1 - σ i ) / σ i |. This indicator can effectively capture the process of the system transitioning from a stable state to an unstable state, even when the average anomaly metric value remains unchanged.

[0112] The change in autocorrelation is evaluated by calculating the autocorrelation coefficients at different time lags and their changes to assess the periodicity and regularity of the anomaly pattern. For a time lag τ, the autocorrelation coefficient is calculated: Select multiple key time lags τ1, τ2,..., τ m (such as 5 minutes, 30 minutes, 1 hour), and calculate the change rate of the autocorrelation spectrum: acf_change = (1 / m) ∑ j |ACF_current(τ j ) - ACF_previous(τ j)|, where ACF_current and ACF_previous are the autocorrelation coefficients of the current window and the previous window respectively. The change in autocorrelation can detect the rhythm change of abnormal patterns. For example, the resource usage peak that originally occurred once an hour suddenly becomes irregular. These three types of dynamic evolution indicators together constitute the multi-dimensional feature vector [slope, amp_change, acf_change] of abnormal behavior evolution, providing a comprehensive dynamic feature description for subsequent comparison with the normal evolution template.

[0113] The normal evolution pattern template library is a pre-constructed knowledge base that stores the normal dynamic behavior patterns of abnormal metric values in various workflow states. For each workflow state i, the template library contains a set of evolution templates T_i = {T_i1, T_i2,..., T_ik}, and each template T_ij represents a typical normal evolution pattern in that state. Each evolution template consists of three parts: the normal range interval of dynamic evolution indicators, the typical time series pattern, and the context condition description.

[0114] The normal range interval of dynamic evolution indicators defines the expected change range of the trend slope, the change rate of fluctuation amplitude, and the change in autocorrelation under normal circumstances, usually expressed as [min_value, max_value] or mean ± standard deviation. For example, the normal trend slope range in the "image processing" state may be [-0.05, 0.08], indicating that in this state, the abnormal metric value may increase slightly but not grow sharply. The typical time series pattern captures the characteristic form of the abnormal metric value sequence in the time dimension, such as periodic fluctuations, progressive growth, or stable plateau periods, etc. These patterns can be represented by parametric curves, Fourier coefficients, or wavelet transform coefficients. The context condition describes the specific scenario conditions applicable to the template, such as time period (working hours / non-working hours), system load level (high / medium / low), or external events (such as regular maintenance), etc. This enables the system to select the most matching template according to the current context.

[0115] After determining the target inference workflow state i of the terminal, the system retrieves the set of all templates T_i corresponding to this state from the template library, and then filters out the most matching target evolution template T_target according to the current context conditions (such as time, load, etc.). If there are multiple matching templates, a weighted combined template can be constructed, and the weights are determined based on the context matching degree. This dynamic template selection mechanism ensures the context adaptability of anomaly detection and can adjust the expected range of normal behavior according to different scenarios.

[0116] The calculated dynamic evolution index is compared with the selected target evolution template in multiple dimensions to quantify the degree of deviation. First, for the trend slope index, calculate its standardized deviation from the normal range of the template: d_slope = max(0, (slope-upper_bound) / range_width) or max(0, (lower_bound-slope) / range_width), where upper_bound and lower_bound are the normal upper and lower limits of the slope defined in the template, and range_width is the range width. This standardization process makes the deviation of different indicators comparable. A value of 0 means it is within the normal range, a positive value means it deviates from the normal range, and the larger the value, the more serious the deviation.

[0117] Similarly, the standardized deviations d_amp and d_acf of the fluctuation amplitude change rate and autocorrelation change are calculated. Considering that the importance of different dynamic indicators may vary depending on the workflow state, a state-related weight vector w = [w_slope, w_amp, w_acf] is introduced. For example, for the "patient monitoring" state that requires stability, w = [0.3, 0.5, 0.2] may be set to pay more attention to the change in fluctuation amplitude.

[0118] In addition to the range deviation, the matching degree of the time series pattern needs to be evaluated. The dynamic time warping (DTW) distance or Euclidean distance is calculated between the current abnormal metric value sequence and the typical time series pattern in the template to obtain the pattern deviation d_pattern. The DTW distance is particularly suitable for processing time series comparison because it allows nonlinear alignment of the sequence on the time axis and can tolerate slight expansion and contraction of the pattern in time.

[0119] Finally, the comprehensive deviation degree is calculated by weighted combination of various deviation indicators: deviation = w_slope·d_slope+w_amp·d_amp+w_acf·d_acf+w_pattern·d_pattern, where the weight reflects the relative importance of each indicator in a specific workflow state. To enhance the interpretability, the contribution percentage of each indicator to the total deviation can also be calculated to help quickly locate the main manifestations of the anomaly. For example, the comprehensive deviation of a medical terminal in the "data transmission" state is 2.7, of which 75% comes from the abnormal increase in fluctuation amplitude, indicating that the use of system resources has experienced abnormal and drastic fluctuations.

[0120] Based on the calculated comprehensive deviation degree, it is determined whether there is an abnormal workflow pattern by comparing with a preset threshold. The threshold setting adopts an adaptive mechanism and is dynamically adjusted according to the workflow state, time period, and historical deviation distribution. Specifically, for workflow state i and context condition c, the threshold is calculated as: threshold(i,c) = base_threshold(i) × context_factor(c) × (1 + adaptive_term), where base_threshold(i) is the base threshold for state i (such as 3.0), context_factor(c) is the context adjustment factor (such as 1.2 at night to tolerate greater deviation), and adaptive_term is dynamically adjusted based on the percentile of the historical deviation distribution, such as set as the 95th percentile of the deviation degree in the same state and time period in the past 30 days.

[0121] When the comprehensive deviation degree exceeds the threshold, the system determines that there is an abnormal workflow pattern in the terminal and generates a detailed description of the abnormal evolution characteristics. The description content includes: (1) basic information of the anomaly, such as detection time, duration, the workflow state it belongs to, and the comprehensive deviation degree; (2) the main manifestation forms of the anomaly, such as "abnormal upward trend", "sharp increase in fluctuation amplitude", or "sudden disappearance of periodicity", etc., and mark the specific deviation values and contribution ratios of each indicator; (3) similarity analysis with historical anomalies, find the most similar historical anomaly events and their handling methods; (4) assessment of the possible influence range and severity of the anomaly, comprehensively judged based on the workflow criticality and deviation degree.

[0122] In one implementation manner, after generating the warning signal including the description of the abnormal evolution characteristics of the platform terminal, the following steps are further included:

[0123] Integrate the description of the abnormal evolution characteristics and the currently inferred workflow state of the platform terminal into comprehensive abnormal characteristics;

[0124] Based on the reconstruction error analysis, determine the original feature dimension with the highest correlation degree with the comprehensive abnormal characteristics in the real-time time series feature vector sequence;

[0125] Combine the comprehensive abnormal characteristics and the original feature dimension and use the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility.

[0126] In this embodiment, the integration of anomaly evolution feature description and inference workflow status adopts a multi-level feature fusion mechanism to construct a comprehensive integrated anomaly feature representation. First, the quantitative indicators (such as comprehensive deviation degree 2.7, trend slope deviation 1.5, fluctuation amplitude change rate deviation 0.8, autocorrelation change deviation 0.4) and qualitative descriptions (such as "sharp increase in fluctuation amplitude", "sudden disappearance of periodicity") in the anomaly evolution feature description are separated and extracted. The quantitative indicators are directly used as the numerical components of the feature vector, while the qualitative descriptions are converted into standardized feature labels and corresponding severity scores through a predefined semantic mapping table. For example, "sharp increase in fluctuation amplitude" is mapped to the feature label "amplitude_fluctuation" and a severity score of 0.85 (with a full score of 1).

[0127] Meanwhile, status identifiers (such as "patient data processing", "image analysis", or "drug dispensing") and their associated context information are extracted from the inference workflow status, including the typical resource consumption patterns of the status, criticality levels (such as "high", "medium", "low"), and expected durations. The workflow status information is converted into a fixed-dimensional vector representation through a status embedding matrix, which is obtained in advance through unsupervised learning of a large amount of normal workflow execution data and can capture the semantic similarities between different workflow statuses.

[0128] Next, the anomaly evolution feature vector is concatenated with the workflow status embedding vector, and a comprehensive anomaly feature is generated through an attention fusion network. This network calculates the association weights between the anomaly features and each dimension of the workflow status, highlighting the most critical anomaly features in the current workflow status. For example, in the "real-time patient monitoring" workflow, anomaly features related to system stability will obtain higher weights; while in the "batch data analysis" workflow, anomaly features related to resource efficiency may have higher weights. The finally generated comprehensive anomaly feature is a multi-dimensional vector, including the type, severity, time characteristics, workflow context, and their association scores of the anomaly. This structured representation not only retains the detailed features of the anomaly but also incorporates workflow context information, providing rich and accurate input for subsequent root cause analysis. For example, for a medical imaging processing terminal, the comprehensive anomaly feature may indicate that in the "emergency CT scan processing" workflow status, there is a periodic abnormal fluctuation in system resource utilization, and this anomaly is highly correlated with the image processing stage of the workflow.

[0129] Based on the comprehensive anomaly features, the original feature dimensions most relevant to the anomaly are accurately located through reconstruction error analysis technology. Specifically, a data subset during the anomaly occurrence period needs to be extracted from the real-time time-series feature vector sequence continuously collected by the monitoring system. These original feature vectors usually contain dozens or even hundreds of dimensions, covering various aspects of the system running state such as CPU usage, memory occupancy, network traffic, disk I / O, and application-level metrics. To improve the analysis efficiency, principal component analysis (PCA) is applied to the original feature vectors for dimensionality reduction, retaining the principal components required to explain 90% of the variance, usually reducing the dimensions to 20 - 30% of the original. Next, the pre-trained autoencoder model is used to reconstruct the dimensionality-reduced feature vectors. The autoencoder consists of an encoder and a decoder. The encoder compresses the input features into a latent representation, and the decoder attempts to reconstruct the original input from the latent representation. The autoencoder trained on normal operating data can accurately reconstruct the normal mode but has a poor reconstruction effect on the abnormal mode. Calculate the reconstruction error of the feature vector at each time point: error(t) = ||X(t) - X'(t)||2, where X(t) is the original feature vector and X'(t) is the reconstructed feature vector.

[0130] To determine the original dimensions most relevant to the comprehensive anomaly features, calculate the distribution of the reconstruction error across each original feature dimension. For each original feature dimension i, calculate its reconstruction error contribution: contrib(i) = |Xi(t) - X'i(t)| / sumj|Xj(t) - X'j(t)|. Dimensions with high reconstruction error contributions are usually closely related to the anomaly phenomenon. Further, calculate the cross-correlation coefficient between the reconstruction error time series and the time series of each anomaly metric in the comprehensive anomaly features to identify the feature dimensions that show significant correlation before and after the anomaly occurs. Finally, based on the reconstruction error contribution and cross-correlation analysis, select the top K original feature dimensions with the highest degree of association (e.g., K = 5) as potential root cause indicators.

[0131] Based on the comprehensive anomaly features and the identified key original feature dimensions, multi-path reasoning is performed through the platform-associated knowledge graph to generate a list of potential fault root causes. The platform-associated knowledge graph is a pre-constructed structured knowledge base that contains the complex dependencies and influence paths among system components, services, resources, and configuration items. The nodes in the knowledge graph represent system entities (such as database servers, network switches, application services, etc.), the edges represent the relationships between entities (such as "depends on", "affects", "contains", etc.), and the edge weights reflect the strength of the relationships. The knowledge graph is constructed through three approaches: expert knowledge encoding, historical fault case learning, and system topology automatic discovery. Root cause inference first locates the set of nodes S corresponding to the key original feature dimensions in the knowledge graph. For example, if "database connection pool utilization rate" is identified as a key feature, the "database connection pool" node in the knowledge graph is located. Starting from these nodes, a bidirectional graph traversal algorithm is executed, tracing back the possible fault sources upward and exploring the possible influence scope downward. During the traversal process, the edge weights are dynamically adjusted according to the comprehensive anomaly features, so that the paths that match the current anomaly feature pattern obtain higher weights. For example, if the anomaly manifests as "periodic increase in latency", the path weights related to scheduled tasks, cache expiration, etc. will be increased.

[0132] For each potential root cause node r identified during the traversal, calculate its likelihood score: score(r) = base_prob(r) × path_strength(r, S) × feature_match(r), where base_prob(r) is the basic fault probability of node r (based on historical statistics), path_strength(r, S) is the path strength from node r to the feature node set S, and feature_match(r) is the matching degree between the features of node r and the current comprehensive anomaly features. Finally, output the list of potential fault root causes in descending order of likelihood scores. Each root cause item includes: root cause description, likelihood score, impact path description, recommended verification steps, and repair solutions. For example, for the anomaly of a medical imaging processing terminal, the system may output: "1. Improper database connection pool configuration (likelihood: 0.87) - The periodic connection reconstruction is caused by too short maximum survival time of the connection. It is recommended to check the maxLifetime parameter of the connection pool; 2. Network bandwidth limitation (likelihood: 0.65) - The bandwidth limit is triggered during the transmission of large images. It is recommended to check the network QoS configuration; 3. Storage I / O bottleneck (likelihood: 0.42) - The image data writing competes for I / O resources with other applications. It is recommended to check the storage performance monitoring". Such structured root cause analysis results greatly shorten the fault diagnosis time and improve the repair efficiency, especially suitable for the rapid location and solution of faults in complex medical platform environments.

[0133] In one of the embodiments, the steps of combining the comprehensive anomaly features and the original feature dimensions and using the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility are as follows:

[0134] Use the comprehensive anomaly features and the original feature dimensions as query conditions;

[0135] Retrieve the target nodes and target paths with the highest degree of association with the query conditions in the pre-constructed platform association knowledge graph. The target nodes include terminal components, terminal service ports, or terminal resource ports, and the target paths represent the dependency relationships or influence relationships between the target nodes;

[0136] Take the target nodes as potential fault sources, calculate the correlation scores of each potential fault source according to the target paths, and sort the potential fault sources by the correlation scores.

[0137] In this embodiment, the process of converting the comprehensive anomaly features and the original feature dimensions into structured query conditions uses feature vectorization and semantic enhancement techniques. First, the comprehensive anomaly features usually contain information in multiple dimensions, such as anomaly types (such as "response latency", "resource exhaustion"), severity levels (such as 0.85 points, with a full score of 1 point), time characteristics (such as "periodic", "persistent", "sudden"), and context information (such as "occurring in the data processing stage"). These anomaly features are encoded as a multi-dimensional vector, where each dimension represents a characteristic. For example, the anomaly type is represented by one-hot encoding, the severity level is directly represented by a numerical value, and the time characteristics are encoded by predefined categories. Key information is extracted from the original feature dimensions, including feature names (such as "CPU usage rate", "number of memory allocation failures"), anomaly values (such as "95%", "growth rate of 200%"), anomaly duration (such as "lasting for 30 minutes"), etc. These original feature information are also converted into structured vector representations. To enhance the semantic understanding ability of the query, domain ontology mapping is applied to the feature names to unify features with different expressions but similar semantics into standard terms. For example, both "memory shortage" and "memory exhaustion" are mapped to the standard term "memory_exhaustion".

[0138] Next, the comprehensive anomaly feature vector and the original feature vector are concatenated to form a unified query vector. To improve query efficiency, principal component analysis is applied to the query vector for dimensionality reduction, retaining the principal components required to explain 90% of the variance. At the same time, to handle the correlation between features, the feature correlation matrix is calculated, highly correlated feature groups are identified, and their weights are appropriately adjusted in the query to avoid a certain type of correlated features from overly influencing the query results. The finally formed query condition is a structured multi-dimensional vector, containing the core features of the anomaly and key original indicators, and has undergone dimensionality reduction and weight adjustment processing. For example, for an anomaly occurring in a medical image processing terminal, the final query condition may include key information such as "response latency (0.9)", "periodicity (0.8)", "peak memory usage (95%)", "database connection pool exhaustion (frequency: every 5 minutes)", etc., laying a foundation for accurate retrieval in the knowledge graph later.

[0139] Based on the constructed query condition, multi-modal semantic retrieval is performed in the platform-associated knowledge graph to identify the most relevant target nodes and paths. The platform-associated knowledge graph is a complex multi-level network structure, containing four types of core elements: nodes (representing entities such as system components, services, resources, etc.), edges (representing the dependency or influence relationships between entities), attributes (describing the characteristics of nodes and edges), and semantic labels (providing additional domain knowledge). The knowledge graph is constructed and continuously updated through three methods: system topology automatic discovery, historical fault case learning, and expert knowledge encoding. The retrieval process first uses a vector similarity matching algorithm to calculate the cosine similarity between the query vector and the feature vectors of each node in the knowledge graph. The feature vector of each node contains information such as its historical behavior pattern, associated anomaly features, resource consumption characteristics, etc. The similarity calculation takes into account the importance weights of the features, and higher weights are assigned to key anomaly features. For example, if "peak memory usage" in the query condition is identified as a high-importance feature, a higher weight is obtained for this dimension in the similarity calculation.

[0140] Next, starting from the top K nodes with the highest similarity (e.g., K = 10), execute the bidirectional graph traversal algorithm. Trace upwards to possible source nodes of influence (such as dependent services, shared resources), and explore downwards to nodes that may be affected (such as services that depend on the current component). Apply a heuristic pruning strategy during the traversal process, preferentially exploring paths with a high semantic relevance to the query conditions and restricting the maximum traversal depth (usually 3 - 5 layers) to control the computational complexity. For each path during the traversal, calculate the path relevance score: path_score = Σ(node_similarity × edge_weight × decay_factor^depth), where node_similarity is the similarity between the node and the query, edge_weight is the weight of the edge (reflecting the strength of the relationship), decay_factor is the depth decay factor (usually taking values between 0.7 - 0.9), and depth is the depth of the path.

[0141] Finally, select the top M paths (e.g., M = 20) with the highest relevance scores as the target paths, and the set of nodes connected by these paths constitutes the target node set. For example, for a periodic response delay anomaly in a medical data processing terminal, the retrieval may identify highly relevant paths such as "database connection pool" → "database service" → "storage system", as well as target nodes such as "database connection pool", "database service", "storage system", "network switch", etc. These target nodes and paths provide a clear scope and direction for subsequent root cause analysis, greatly narrowing the search space for fault location.

[0142] Based on the retrieved target nodes and target paths, calculate the relevance score of each target node as a potential fault root cause through a multi-dimensional scoring model and generate a ranking result. The scoring model comprehensively considers four key dimensions: the direct relevance of the node itself to the query conditions, the topological importance of the node in the target path, the historical failure probability of the node, and the temporal relevance between the node state and the current anomaly. Calculate the direct relevance of the node, that is, the semantic matching degree between the node features and the query conditions. Use the weighted cosine similarity method to assign higher weights to the key features in the query conditions (such as specific anomaly patterns or resource metrics). For example, if the query condition emphasizes "periodic memory usage peak", the relevance of similar patterns appearing in the node's historical behavior will be higher. Secondly, evaluate the topological importance of the node in the target path. Calculate the centrality metrics of the node, including degree centrality (the number of connected edges), betweenness centrality (the frequency of the node being on the shortest path), and eigenvector centrality (a recursive definition considering the importance of neighbor nodes). Nodes with high topological importance are usually common dependencies of multiple components and have a wide range of fault impacts.

[0143] It is also necessary to calculate the historical failure probability of the incorporated nodes. Based on historical failure records, the failure frequency and conditional probability of nodes in similar scenarios can be calculated. For example, if the "database connection pool" has failed multiple times in the past "high-concurrency access" scenarios, its failure probability assessment will be higher in the current similar scenarios. Finally, it is also necessary to analyze the temporal correlation between the node status and the current anomaly. Calculate the cross-correlation function between the time series of node monitoring metrics and the time series of anomaly occurrences, and identify the nodes that show significant changes before the anomaly occurs. Nodes with strong temporal correlation are more likely to be the root cause rather than the symptom. Ultimately, calculate the correlation score for each target node by integrating the above four dimensions. Then, sort the potential root causes of failure in descending order according to the correlation score, and generate an explanatory description for each root cause, including: root cause name, correlation score, critical impact path, anomaly feature matching points, and recommended verification and repair steps. For example, the sorting result may show: "1. Improper database connection pool configuration (score: 0.89) - The maximum connection survival time is too short, resulting in periodic connection reconstruction. It is recommended to check the maxLifetime parameter; 2. Storage I / O bottleneck (score: 0.76) - Large file operations conflict with database writes. It is recommended to check the storage performance monitoring; 3. Network bandwidth limitation (score: 0.65) - Data transfer peaks trigger traffic limiting. It is recommended to check the network QoS configuration". This structured root cause analysis result intuitively presents the failure probability ranking, helping technicians quickly locate the core of the problem and improve the efficiency of troubleshooting.

[0144] The present invention also discloses a terminal operation and maintenance management platform for a medical big data platform, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the terminal operation and maintenance management method for a medical big data platform as described in the first aspect.

[0145] Among them, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be used. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. The present application does not make any restrictions on this.

[0146] Among them, the memory can be an internal storage unit of the computer device, for example, the hard disk or memory of the computer device, or an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital card (SD), or a flash card (FC) equipped on the computer device. Moreover, the memory can also be a combination of the internal storage unit and the external storage device of the computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store the data that has been output or will be output. This application does not limit this.

[0147] The present invention also discloses a computer-readable storage medium. Instructions are stored on the computer-readable storage medium. When the instructions are executed by a processor, the processor is configured to execute the terminal operation and maintenance management method for the medical big data platform described in any of the above embodiments.

[0148] Among them, the computer program can be stored in a machine-readable medium. The computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some middleware form, etc. The machine-readable medium includes any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code. It should be noted that the machine-readable medium includes, but is not limited to, the above components.

[0149] Among them, through this computer-readable storage medium, the transmission line comprehensive fault detection method in the above embodiment is stored in the computer-readable storage medium, and is loaded and executed on the processor to facilitate the storage and application of the above method.

[0150] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the protection scope of this application is limited to these examples; under the idea of this application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of this application as described above. For the sake of brevity, they are not provided in detail.

[0151] One or more embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omission, modification, equivalent substitution, improvement, etc. made within the spirit and principle of one or more embodiments of this application shall be included in the protection scope of this application.

Claims

1. A terminal operation and maintenance management method for a medical big data platform, characterized in that Including the following steps: Collect multi-modal historical operation and maintenance data of the platform terminal in the normal workflow mode; Preprocess and fuse the multi-modal historical operation and maintenance data to construct a historical time-series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal; Use the historical time-series feature vector sequence to train a preset deep sequence learning model. The model training process includes learning the latent representation of the historical time-series feature vector sequence to capture the context state of the normal workflow mode, and establishing a baseline model for reconstructing the historical time-series feature vector sequence based on the latent representation; Construct a real-time time-series feature vector sequence according to the currently obtained real-time operation and maintenance data of the platform terminal, use the trained deep sequence learning model to extract the latent context representation of the real-time time-series feature vector sequence, and use the baseline model to calculate the reconstruction error of the real-time time-series feature vector sequence; Based on the latent context representation, identify the inferred workflow state where the platform terminal is currently located, and calculate a context-aware anomaly metric value by combining the inferred workflow state and the reconstruction error; Analyze the time evolution characteristics of the anomaly metric value, and compare it with the preset anomaly pattern criterion according to the time evolution characteristics to determine whether there is an abnormal workflow mode for the platform terminal. If there is an abnormal workflow mode for the platform terminal, generate a warning signal including a description of the abnormal evolution characteristics of the platform terminal.

2. The terminal operation and maintenance management method for a medical big data platform according to claim 1, wherein The preprocessing and fusion processing of the multi-modal historical operation and maintenance data to construct a historical time-series feature vector sequence reflecting the historical normal dynamic behavior of the platform terminal includes the following steps: Perform time alignment and data parsing processing on the multi-modal historical operation and maintenance data at a unified sampling time interval, and parse the key information fields of the multi-modal historical operation and maintenance data; Calculate and extract derivative correlation features characterizing the immediate interaction relationship between different modal data; Encode and fuse the key information fields and derivative correlation features into a single high-dimensional feature vector, and arrange the high-dimensional feature vectors in chronological order to form a historical time-series feature vector sequence.

3. The terminal operation and maintenance management method for a medical big data platform according to claim 1, characterized in that The training of the preset deep sequence learning model using the historical time-series feature vector sequence includes the following steps: Use the encoder in the preset deep sequence learning model to map the historical time-series feature vector sequence to the latent representation space; Train the deep sequence learning model to perform multi-task learning. The learning tasks at least include: reconstructing the original input sequence based on the decoder in the deep sequence learning model and the latent context representation in the latent representation space to minimize the model reconstruction error; predicting the time-series feature vector of the next time step based on the latent context representation to minimize the model prediction error; After the deep sequence learning model completes multi-task learning, retain the encoder for real-time extraction of the latent context representation, and retain the reconstruction task part as the baseline model.

4. The terminal operation and maintenance management method for a medical big data platform according to claim 3, characterized in that The identification of the inferred workflow state where the platform terminal is currently located based on the latent context representation includes the following steps: Perform clustering analysis on the latent context representation in the latent representation space, and define discrete inferred workflow state clusters according to the clustering analysis results; Calculate the spatial distance between the latent context representation of the real-time time-series feature vector sequence and the centers of each inferred workflow state cluster; Statistically analyze the local density and the distribution characteristics of neighboring samples of the latent context representation of the real-time time series feature vector sequence in the latent representation space; Based on the spatial distance, local density, and the distribution characteristics of neighboring samples, assign the current target inference workflow state to the platform terminal.

5. The terminal operation and maintenance management method for a medical big data platform according to claim 4, wherein, The steps for calculating the context-aware anomaly metric value by combining the inference workflow state and the reconstruction error are as follows: For each inference workflow state, calculate and store the multi-dimensional error statistical model under the inference workflow state based on the reconstruction error of the samples belonging to the inference workflow state in the historical time series feature vector sequence; According to the currently assigned target inference workflow state, search for the corresponding target multi-dimensional error statistical model; Compare the reconstruction error of the real-time time series feature vector sequence with the target multi-dimensional error statistical model, calculate the statistical significance and the degree of unexpectedness under the target inference workflow state, and calculate the context-aware anomaly metric value by combining the statistical significance and the degree of unexpectedness.

6. The terminal operation and maintenance management method for a medical big data platform according to claim 4, characterized in that, The steps for analyzing the time evolution characteristics of the anomaly metric value and comparing them with the preset anomaly pattern criteria to determine whether there is an anomaly in the workflow pattern of the platform terminal are as follows: Maintain the anomaly metric value as a sequence of consecutive anomaly metric values within a time window; Calculate the dynamic evolution indicators of the anomaly metric value sequence. The dynamic evolution indicators include the trend slope, the change rate of the fluctuation amplitude, and the change of autocorrelation; Select the corresponding target evolution template from the predefined normal evolution pattern template library according to the target inference workflow state. The evolution templates in the normal evolution pattern template library characterize the normal fluctuation range and dynamic characteristics of the anomaly metric value under different states; Compare the dynamic evolution indicators with the target evolution template, and calculate the comprehensive deviation degree of the dynamic evolution indicators deviating from the target evolution template; When the comprehensive deviation degree exceeds the preset threshold, it is determined that there is an anomaly in the workflow pattern of the platform terminal, and the comprehensive deviation degree and the target evolution template are used as the description of the abnormal evolution characteristics.

7. The terminal operation and maintenance management method for a medical big data platform according to claim 1, wherein, After generating the warning signal containing the description of the abnormal evolution characteristics of the platform terminal, the following steps are further included: Integrate the description of the abnormal evolution characteristics and the current inference workflow state of the platform terminal into the comprehensive abnormal characteristics; Based on the reconstruction error, analyze the original feature dimension with the highest correlation degree with the comprehensive abnormal characteristics in the real-time time series feature vector sequence; Combine the comprehensive abnormal characteristics and the original feature dimension, and use the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility.

8. The terminal operation and maintenance management method for a medical big data platform according to claim 7, characterized in that, The steps for combining the comprehensive abnormal characteristics and the original feature dimension and using the pre-constructed platform association knowledge graph to infer and output a set of potential fault sources sorted by possibility are as follows: Use the comprehensive abnormal characteristics and the original feature dimension as the query conditions; Retrieve the target nodes and target paths with the highest correlation degree with the query conditions in the pre-constructed platform association knowledge graph. The target nodes include terminal components, terminal service ports, or terminal resource ports, and the target paths represent the dependency relationships or influence relationships between the target nodes. Take the target node as the potential root cause of the fault, calculate the correlation scores of each potential root cause of the fault according to the target path, and sort the potential root causes of the fault through the correlation scores.

9. A terminal operation and maintenance management platform for a medical big data platform, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the terminal operation and maintenance management method for the medical big data platform according to any one of claims 1 to 8.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, When executed by the processor, the instruction causes the processor to be configured to execute the terminal operation and maintenance management method for the medical big data platform according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Improved fault diagnosis algorithm based on multi-modal information fusion

    CN118536068A

  • Intelligent operation and maintenance method, system and equipment based on deep learning and medium

    CN119168626A

  • Network traffic abnormity monitoring method and device based on BiLSTM-Att network

    CN119232490A

  • Data anomaly detection method and apparatus

    WO2023123941A1

Cited By

  • Intelligent walking stick early warning method and system based on multi-parameter physiological monitoring

    CN120938373A

  • Equipment operation and maintenance method and system based on multi-modal large model

    CN121052807A

  • Abnormality detection method, electronic equipment and storage medium

    CN121327721A

  • Method and system for monitoring abnormal pressure of fuel oil common rail pipe of marine main engine

    CN121959367A

  • System operation multivariate time series data anomaly detection method and electronic equipment

    CN122310382A