Cloud platform abnormal state early warning method based on deep time sequence feature extraction
Patent Information
- Application Number
- CN202610991295.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-05
- Publication Date
- 2026-09-25
AI Technical Summary
若平台资源已接近安全边界,未经验证的自愈动作可能造成资源过冲、请求失败率上升或异常传播范围扩大
本发明通过建立观测暗窗,并将观测暗窗实例与监控对照实例进行同源业务请求对照,能够显现被主动健康检查、连接保活及缓存预热掩盖的运行异常,降低隐性异常的漏检率。通过对成对运行时序数据进行请求阶段对齐、归一化及深度时序特征提取,能够消除宿主机计时起点、资源量纲和采样缺失造成的干扰,提高跨服务实例状态比较的准确性。
Smart Images

Figure CN122817031A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud platform operation monitoring technology, and in particular to a cloud platform abnormal state early warning method based on deep time series feature extraction. Background Technology
[0002] As online learning, remote examinations, and live course streaming services migrate to cloud platforms, these systems typically adopt a microservice architecture and maintain service availability and collect operational data through proactive health checks, connection keep-alive, cache preheating, and resource monitoring. However, proactive monitoring operations may alter the actual operational status of service instances.
[0003] For example, proactive health checks may establish database or downstream service connections in advance, connection keep-alive may extend connection lifecycles, and cache preheating may load business data and code pages ahead of time. Therefore, some connection rebuilding, cache misses, code page loading, or database handshake anomalies may be temporarily masked under continuous monitoring, only to reappear when real business requests arrive. Existing anomaly detection methods primarily rely on resource metrics, request failure rates, or response times of individual service instances, making it difficult to identify hidden anomalies masked by proactive monitoring operations.
[0004] Furthermore, existing cloud platforms typically execute self-healing actions such as scaling up, connection adjustments, cache rebuilding, traffic switching, or instance rebuilding immediately upon detecting anomalies. If platform resources are already close to security limits, unverified self-healing actions may cause resource overload, increased request failure rates, or an expanded scope of anomaly propagation. Existing technologies lack a mechanism to evaluate the anomaly suppression effect and resource side effects using small-scale, rollback-enabled actions before formal self-healing, making it difficult to safely determine the execution order, scale of action, and phased advancement conditions of self-healing actions. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a cloud platform abnormal state early warning method based on deep temporal feature extraction, thereby solving the technical problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A cloud platform anomaly state early warning method based on deep temporal feature extraction includes the following steps: S1. Obtain the resource configuration, dependency topology, and load data of the target microservice, select a pair of service instances and set them as the observation dark window instance and the monitoring control instance, pause the active monitoring operation of the observation dark window instance, send the same source business request, and form a pair of runtime sequence datasets. S2. Perform request phase alignment and normalization on the paired runtime time series datasets to form a dual instance aligned time series matrix. Input it into the deep time series feature extraction model to obtain the observation dark window time series features and the monitoring control time series features, forming a dark window difference feature sequence. S3. Based on the fault-free samples, establish the normal differential range, identify the hidden abnormal segments in the dark window differential feature sequence, combine the differential start time, resource channel weight and resource call relationship to construct the propagation path, determine the candidate resource channels, and generate a set of candidate self-healing actions and abnormal state warnings. S4. Generate reversible rescue probes for the candidate self-healing action set, collect probe response time series data, calculate dark window differential attenuation and rescue action impedance, construct reinforcement learning state and output self-healing action ranking results. S5. Select the target self-healing action from the self-healing action ranking results. The rescue action impedance has not diverged and the dark window differential attenuation meets the threshold. After the action is executed in stages, the observation dark window is re-established to form a verification dark window differential feature sequence. Based on the verification normal ratio, the confidence of the latent anomaly and the propagation path, it is determined whether the latent anomaly has been eliminated. If it has not been eliminated, the reinforcement learning state is updated.
[0007] S1 specifically includes: acquiring the instance status, resource configuration, dependency topology, and historical load data of the target microservice; selecting paired service instances based on resource configuration differences and historical load interval overlap rates; determining the observation dark window instance and monitoring control instance; generating instance pairing identifiers; pausing the active health check, connection keep-alive, and cache preheating of the observation dark window instance, while retaining host-level liveness detection; determining the duration of the observation dark window based on the operation cycle and security observation duration; synchronously collecting processor, memory, network, and disk operation data of the paired service instances; generating a business request sequence based on the request metadata of the upstream gateway; configuring a request lineage identifier for the same business request; configuring a shadow execution identifier for write requests and routing them to isolated replicas; merging request phase data with runtime data to form a paired runtime sequence dataset.
[0008] S2 specifically includes: associating paired runtime sequence datasets based on instance pairing identifiers, request lineage identifiers, and shadow execution identifiers; converting sampled records into relative stage positions according to request stages; performing resampling, missing value processing, and quantile normalization to form a dual-instance aligned time series matrix; inputting the dual-instance aligned time series matrix into a deep time series feature extraction model with parameter-shared time series coding branches to extract short-term mutation features, long-term cumulative features, and cross-resource propagation features to obtain observation dark window time series features and monitoring control time series features; comparing the observation dark window time series features with the monitoring control time series features according to the same request stage, resource channel, and relative stage position to determine the differential amplitude, differential start time, differential duration, propagation level, and resource channel weight to form a dark window differential feature sequence.
[0009] S3 specifically includes: based on the historical fault-free dark window differential feature sequence, establishing a normal differential range according to request type, resource load range, and resource channel; calculating the standardized exceedance degree and implicit anomaly confidence; identifying implicit anomaly segments with fault precursor attributes and determining the monitoring cover-up operation identifier; constructing a propagation path by combining the differential start time of the implicit anomaly segment, resource channel weight, dependent service topology, and actual resource call relationship; determining the resource channel located before the response resource channel and meeting the threshold as a candidate resource channel, forming a candidate resource channel record; querying the cloud platform control interface registry based on the candidate resource channels to determine the control interface and action type that supports rollback, setting the maximum scale of action, resource security upper limit, and rollback conditions, forming a candidate self-healing action set, and outputting anomaly status warnings.
[0010] S4 specifically includes: generating reversible rescue probes acting on target candidate resource channels based on the candidate self-healing action set; sequentially executing the probe pre-baseline, probe action, probe rollback, and resource recovery to form a probe response time-series dataset; performing deep time-series feature extraction on the probe response time-series dataset to determine the response gain, maximum overshoot amplitude, recovery time, number of affected resources, and dark window differential attenuation; calculating the rescue action impedance and determining the impedance status based on this; constructing a reinforcement learning state based on abnormal state warning, candidate resource channel records, rescue action impedance, resource remaining ratio, and operational performance indicators; constructing the candidate self-healing actions, action scale, and duration as reinforcement learning actions; and outputting a self-healing action ranking result containing execution priority, recommended total action scale, and recommended stage duration through the reinforcement learning strategy model.
[0011] S5 specifically includes: selecting target self-healing actions from the self-healing action sorting results whose impedance state is not divergent, whose dark window differential attenuation exceeds the threshold, and whose resource conditions are met; determining the execution stage and stage differential convergence target based on the recommended total action scale and minimum adjustment step size; setting action control locks for target self-healing actions and executing them in stages, maintaining the observation dark window state, calculating real-time rescue action impedance, and controlling stage advancement or rollback based on stage differential convergence target, resource safety upper limit, and rollback conditions; establishing an observation dark window after the target self-healing action is completed, forming a verification dark window differential feature sequence, and judging whether the hidden anomaly has been eliminated based on the verification normal ratio, the confidence of the hidden anomaly, and the propagation path. If it has not been eliminated, returning to the corresponding step to update the decision; if it has been eliminated, adjusting the monitoring configuration corresponding to the monitoring cover-up operation and removing the abnormal state warning.
[0012] The beneficial effects of this invention are as follows: This invention establishes an observation window and compares the observation window instances with monitoring control instances using same-source business requests. This enables the detection of operational anomalies masked by proactive health checks, connection keep-alive, and cache preheating, reducing the missed detection rate of latent anomalies. By aligning, normalizing, and extracting deep temporal features from paired runtime sequence data for request phases, interference caused by host timing start points, resource dimensions, and sampling deficiencies can be eliminated, improving the accuracy of cross-service instance status comparisons.
[0013] This invention, by analyzing the differential start time, differential duration, resource channel weights, and actual resource call relationships, can construct anomaly propagation paths and locate candidate resource channels with fault precursor attributes before a significant anomaly occurs. By applying small-scale, rollback-capable, reversible rescue probes to candidate self-healing actions, it is possible to obtain the anomaly suppression effect, resource overshoot degree, and recovery time of the action before formal self-healing, avoiding the direct execution of high-risk self-healing actions.
[0014] This invention uses a combination of rescue action impedance and a reinforcement learning strategy model to determine the execution priority, recommended total action scale, and recommended stage duration of candidate self-healing actions. This approach balances anomaly mitigation effectiveness with resource security, improving the adaptability of self-healing decisions. By executing target self-healing actions in stages and controlling the progression or rollback based on real-time rescue action impedance, stage differential convergence targets, and resource security limits, and then verifying the anomaly elimination results using a validation dark window differential feature sequence, the risk of resource overload and secondary failures induced during the self-healing process can be reduced. Attached Figure Description
[0015] Figure 1 This is a flowchart of the cloud platform abnormal state early warning method based on deep temporal feature extraction according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example: Figure 1 As shown, this embodiment provides a cloud platform abnormal state early warning method based on deep temporal feature extraction, including the following steps: S1. Obtain the resource configuration, dependency topology, and load data of the target microservice, select a pair of service instances and set them as the observation dark window instance and the monitoring control instance, pause the active monitoring operation of the observation dark window instance, send the same source business request, and form a pair of runtime sequence datasets. S2. Perform request phase alignment and normalization on the paired runtime time series datasets to form a dual instance aligned time series matrix. Input it into the deep time series feature extraction model to obtain the observation dark window time series features and the monitoring control time series features, forming a dark window difference feature sequence. S3. Based on the fault-free samples, establish the normal differential range, identify the hidden abnormal segments in the dark window differential feature sequence, combine the differential start time, resource channel weight and resource call relationship to construct the propagation path, determine the candidate resource channels, and generate a set of candidate self-healing actions and abnormal state warnings. S4. Generate reversible rescue probes for the candidate self-healing action set, collect probe response time series data, calculate dark window differential attenuation and rescue action impedance, construct reinforcement learning state and output self-healing action ranking results. S5. Select the target self-healing action from the self-healing action ranking results. The rescue action impedance has not diverged and the dark window differential attenuation meets the threshold. After the action is executed in stages, the observation dark window is re-established to form a verification dark window differential feature sequence. Based on the verification normal ratio, the confidence of the latent anomaly and the propagation path, it is determined whether the latent anomaly has been eliminated. If it has not been eliminated, the reinforcement learning state is updated.
[0018] S1 specifically includes the following sub-steps: S110. Determine the target microservice based on the microservice architecture and select paired service instances. The target microservice refers to a microservice specified by the cloud platform management terminal and provided by at least two independently schedulable service instances.
[0019] The cloud platform management terminal receives the target service identifier consisting of the service name, service version, and deployment namespace. It reads the instance identifier and instance running status corresponding to the target microservice from the microservice registry center, and reads the container image summary, vCPU quota, memory limit, network bandwidth limit, disk I / O limit, and deployment node type of each service instance from the container orchestration control plane. It reads the downstream service identifier, call direction, and downstream service version actually called by each service instance in the past 7 days from the call chain repository, and reads the request processing volume, processor utilization, and memory utilization of the same business period in the past 7 days from the monitoring time series database.
[0020] Service instances with identical sets of dependent nodes, dependent edges, downstream service versions, and container image digests are identified as candidate instances, and the resource configuration difference value between any two candidate instances is calculated: ; In the formula, J represents the resource allocation difference value; J represents the number of resource types participating in the comparison. and These are the resource configuration values of the j-th type for the two candidate instances, respectively. Let be the resource allocation weight corresponding to the j-th type of resource, and the sum of all resource allocation weights is 1.
[0021] The resource types included in the comparison are vCPU quota, memory limit, network bandwidth limit, and disk I / O limit. Historical load intervals are constructed using the 5th and 95th percentile values of request throughput, processor utilization, and memory utilization, respectively, and the overlap rate of historical load intervals is calculated as the ratio of the intersection length to the union length of two historical load intervals.
[0022] When the resource configuration difference is no greater than 0.10, and the overlap rate of the above three types of historical load intervals is no less than 0.70, the corresponding two candidate instances are determined as paired service instances. When there are multiple instance combinations that meet the conditions, the instance combination with the smallest resource configuration difference and the largest historical load interval overlap rate is selected in sequence. The two selected service instances are determined as the first service instance and the second service instance, respectively. Instance pairing identifiers are generated based on the target service identifier, the observation dark window batch number, the first service instance identifier, and the second service instance identifier. The instance pairing identifiers are written into all subsequent sampling records.
[0023] For example, both the first and second service instances are configured with 4 vCPUs and 8GB of memory, with network bandwidth limits of 500 Mbit / s and 480 Mbit / s respectively. The remaining resource configurations are the same, and the calculated resource configuration difference value is 0.01, which meets the instance pairing conditions.
[0024] In this embodiment, the cloud platform and the target microservice are deployed in a container cluster (such as a Kubernetes cluster). The hardware configuration of the computing node is, for example, 8-core vCPU and 16GB memory. The host operating system adopts the Linux kernel operating system.
[0025] S120. Establish an observation window and collect runtime data from paired service instances. The observation window refers to a limited time interval during which, while the first service instance continues to receive business requests, active monitoring operations that would maintain dependent connections, refresh business caches, or load code pages in advance are suspended from being sent to its target service process, while passive runtime data that does not call the target service process's business interfaces are continuously collected.
[0026] The first service instance is set as the observation dark window instance, and the second service instance is set as the monitoring control instance. Proactive monitoring operations include proactive health checks, connection keep-alive, and cache preheating. Proactive health checks refer to periodically calling the target service process and accessing the database, cache, or downstream services for checks. Connection keep-alive refers to periodically sending keep-alive messages to maintain database connections, cache connections, or downstream service connections. Cache preheating refers to proactively reading course configurations, user permissions, video indexes, or code pages before real business requests arrive.
[0027] Retain host-level liveness detection, which is only used to determine whether the container process is alive and does not call business interfaces, to prevent the container orchestration control plane from misjudging the observed dark window instance as a faulty instance and performing a restart.
[0028] Read the active health check cycle, connection keep-alive cycle, and cache warm-up cycle from the service mesh configuration, connection pool configuration, and cache task configuration, respectively, and determine the duration of the observation dark window: ; In the formula, To observe the duration of the dark window; For safe observation duration; This is a proactive health check-up cycle; To maintain the connection during the keep-alive cycle; This is the cache warm-up period.
[0029] The duration of security observation is determined based on the minimum number of available instances and the service level target for the target microservice. During the observation window, the number of available service instances other than the observation window instances must not be lower than the minimum number of available instances. If the request failure rate of the observation window instances exceeds the 99th percentile of the request failure rate in the 30 minutes prior to the start of the observation window, or if the first response time exceeds the service level target for three consecutive sampling periods, the observation window will be terminated early.
[0030] Processor data originates from the processor performance monitoring unit, memory data from the container control group, network data from the kernel network interface counter, and disk data from the block device input / output interface. Two service instances are synchronously collected by the same sampling controller at the same sampling period of 1 to 5 seconds. Each sampling record includes the instance pairing identifier, instance role, sampling sequence number, monotonic counting time, resource type, metric name, and metric value.
[0031] For example, when the active health check cycle is 10s, the connection keep-alive cycle is 30s, the cache warm-up cycle is 60s, and the security observation duration is 180s, the observation dark window duration is 120s.
[0032] S130. Generate a business request sequence and form a pairwise runtime sequence dataset. A business request sequence refers to a set of requests formed by arranging the request metadata recorded by the upstream gateway of the target microservice according to the request arrival time and gateway request sequence number. The request set maintains consistency in request interface identifier, parameter size, dependency relationship, sequence position and adjacent arrival interval in the observation dark window instance and the monitoring control instance.
[0033] The request metadata includes the request interface identifier, request method, parameter byte length, dependency chain identifier, arrival time, tenant isolation identifier, and interface write attributes. User tokens, answer content, and identity fields are not recorded. Based on the microservice interface definition file and the presence of database writes, message publications, or object storage writes in the historical call chain, it is determined whether the business request involves state writes.
[0034] Configure the same request lineage identifier for the execution branches of the same original business request in two service instances to associate the runtime data generated by the same business request; configure a shadow execution identifier for business requests involving state writing to route database transactions, cache writes, message publishing and object storage writes to the corresponding isolated state replicas.
[0035] Isolated replicas refer to the first and second write-time replication spaces established based on database snapshots at the same time, as well as their respective independent cache namespaces, shadow message topics, and shadow object directories. Shadow message topics are not connected to formal business consumers, and shadow object directories are not read by formal business users.
[0036] Read-only service requests are sent to the observation dark window instance and the monitoring control instance respectively, and write requests are sent to the corresponding isolated state replicas respectively; the time difference between the sending of two requests corresponding to the same request spectrum identifier shall not exceed 10% of one sampling period, otherwise the request shall be determined as an invalid pairing request.
[0037] Data on connection establishment, cache access, code page loading, thread waiting, database handshake, network transmission, storage access, and initial response are collected according to the request lineage identifier and shadow execution identifier. This data is then merged with the passive runtime data generated in S120 to obtain a paired runtime sequence dataset. Each data record includes an instance pairing identifier, request lineage identifier, shadow execution identifier, instance role, request phase identifier, sampling sequence number, monotonic counting time, indicator name, and indicator value.
[0038] For example, the learning progress save request is written to the first write-time copy space and the second write-time copy space in the observation dark window instance and the monitoring control instance, respectively, and a progress update message is published to the corresponding shadow message topic, thereby comparing the database handshake, transaction wait and message sending sequence of the two service instances without modifying the actual learning progress.
[0039] S2 specifically includes the following sub-steps: S210. Read the paired runtime sequence datasets and form a dual-instance aligned time series matrix. Complete the dual-instance data association based on the instance pairing identifier, request lineage identifier, and shadow execution identifier. The instance pairing identifier is used to limit the same target microservice, the same observation window batch, and the same paired service instances. The request lineage identifier is used to associate the execution process of the same original business request in the observation window instance and the monitoring control instance. The shadow execution identifier is used to distinguish the isolated execution branch entered by the write request.
[0040] The timing anchor for a business request entering a service instance comes from the request entry event recorded by the service mesh proxy; the timing anchor for dependency call initiation and return comes from the distributed call chain record; and the timing anchor for the completion of a business request comes from the first byte of the response sent event.
[0041] Since the monotonic counting time start points are different on different host machines, the request phase is divided by adjacent time-series anchor points, and the sampling records within each request phase are converted into relative phase positions: ; In the formula, Let be the relative stage position of the i-th sampling record in instance p; This is the monotonic counting time of the sampling record; and These are the monotonic count times for the start and end time sequence anchors of the current request phase, respectively.
[0042] The same request phase in two service instances is resampled to 21 equally spaced positions between 0 and 1, while retaining the original request phase duration. If a metric is missing for no more than two consecutive sampling periods, linear interpolation is performed using valid values before and after the missing point; if it is missing for more than two consecutive sampling periods, the corresponding metric value is set to 0, and a missing status field is added, where valid samples are recorded as 1 and missing samples as 0. If the valid sampling ratio of any service instance is less than 80%, or if there is a missing request entry event or request completion event, the corresponding request is determined to be an invalid pairing request.
[0043] Extract various metrics from historical records within the past 30 days that simultaneously met the criteria of no service level target defaults, no container restarts, and no failure rollbacks. Then, use these extracted metrics to perform quantile normalization on the current data. ; In the formula, Let h be the normalized value of the h-th index at the i-th resampling position; These are the original indicator values; , and These are the 50th, 95th, and 5th percentile values of the h-th indicator, respectively.
[0044] The instance roles, request phases, resource channels, normalized metric values, missing states, and original phase durations are arranged according to their resampling positions to obtain a dual-instance aligned time-series matrix. Resource channels include connection resource channels, cache resource channels, code page resource channels, thread resource channels, database resource channels, network resource channels, storage resource channels, and response resource channels.
[0045] For example, when the database handshake phase of the observation dark window instance and the monitoring control instance lasts for 80ms and 40ms respectively, both phases are mapped to the same 21 resampling positions, while retaining the original phase durations of 80ms and 40ms respectively.
[0046] S220. Input the dual-instance aligned temporal matrix into the deep temporal feature extraction model to obtain the temporal features of the observed dark window and the monitoring control, respectively. The deep temporal feature extraction model includes two temporal coding branches with identical network structures and sharing all model parameters to avoid pseudo-differences caused by differences in coding parameters. Each temporal coding branch includes, in sequence, a first causal convolutional layer, a second dilated causal convolutional layer, a one-way gated recurrent layer, a resource association layer, and a feature output layer.
[0047] The first causal convolutional layer extracts local changes using three consecutive resampling positions as the convolution window. The second dilated causal convolutional layer has a dilation rate of 2 to expand the temporal perception range. The unidirectional gated recurrent layer accumulates the preceding states according to the sampling order. The resource association layer calculates the temporal association weights between eight types of resource channels. The feature output layer outputs the 32-dimensional temporal feature vectors corresponding to each request stage.
[0048] The model training data comes from historical pairwise runtime sequence datasets, including normal samples that have been confirmed to be fault-free by operation and maintenance event records, as well as abnormal samples formed in the isolated test environment by stopping connection keep-alive, clearing cache, limiting vCPU quota, delaying database handshake, increasing network round-trip time, or limiting storage input and output.
[0049] The model training employed a joint loss function for optimization. This joint loss function included: the mean squared error (MSE) temporal reconstruction loss for predicting resource indicators at the next resampled location, and the mean absolute error (MAE) first response prediction loss for predicting the first response duration based on the characteristics of each resource channel. The model training used the Adam optimizer with an initial learning rate of 0.001 and a batch size of 64. After model training converged, the model parameters were fixed.
[0050] Short-term mutation features are the changes in the current resampling position relative to the previous position and the average of the previous three positions. Long-term cumulative features are the cumulative state information between the starting position of the current request phase and the current resampling position. Cross-resource propagation features are the directional correlation information formed when the changes in the previous resource channel precede those in the next resource channel in time, and the introduction of the previous resource channel features can reduce the prediction error of the next resource channel.
[0051] Connection establishment duration, cache miss duration, and database handshake duration are all determined by the monotonic count time difference between the corresponding start and end events; code page load density is determined by dividing the number of code page loads within the request phase by the duration of the request phase; thread wait ratio is determined by dividing the thread wait duration by the sum of the thread runtime and the thread wait duration; network resource metrics include packet retransmission count, socket wait queue length, and network round-trip time; storage resource metrics include disk queue depth, input / output wait time, and read / write throughput; the first response duration is determined by the time difference between the request entering the service instance and the first byte of the response leaving the service instance.
[0052] S230. Form a dark window differential feature sequence. Based on the request spectrum identifier, compare the observation dark window time series features and the monitoring control time series features item by item according to the same request stage, the same resource channel, and the same relative stage position. The average of the absolute differences of the corresponding dimensions of the two 32-dimensional time series feature vectors is determined as the differential amplitude. The differential amplitude distribution of each resource channel is statistically analyzed from the historical fault-free paired samples of the past 30 days, and the 95th percentile value is used as the differential noise threshold of the corresponding resource channel.
[0053] When the differential amplitude of a resource channel exceeds the differential noise threshold for three consecutive resampling positions, the actual time corresponding to the first over-limit position is determined as the differential start time. Subsequently, when the differential amplitude does not exceed the differential noise threshold for two consecutive resampling positions, the actual time length between the differential start time and the position before the first fall-back position is determined as the differential duration.
[0054] If only one resampling location exceeds the limit, it is identified as transient noise. The propagation level is determined according to the differential start time of each resource channel; if the preceding resource channel is at least one sampling period earlier than the following resource channel, and the two resource channels are directly related in the dependency service topology or resource call relationship of the current request, it is determined that the propagation will proceed from the preceding resource channel to the following resource channel.
[0055] Calculate resource channel weights using the first response prediction output: ; ; In the formula, Let r be the resource channel weight of the r-th resource channel; This represents the total number of resource channels; q represents the average change in the first response prediction after replacing the r-th resource channel with the historical normal median value; q is the summation index of the traversed resource channels; This represents the average change in the predicted first response value after replacing the q-th resource channel with the historical normal median value; This represents the number of effective resampling locations. and These are the initial response predictions when using all resource channels and when replacing the r-th resource channel, respectively.
[0056] The instance pairing identifier, request lineage identifier, shadow execution identifier, request stage identifier, resource channel, differential amplitude, differential noise threshold, differential start time, differential duration, propagation level, resource channel weight, observation dark window time series characteristics, and monitoring control time series characteristics are combined in chronological order to obtain the dark window differential feature sequence.
[0057] S3 specifically includes the following sub-steps: S310. Establish the normal differential range and identify hidden abnormal segments. Read the dark window differential feature sequence formed in S230, and read the historical dark window differential feature sequence of the target microservice in the past 30 days from the monitoring time series database; combine the operation and maintenance event records, container orchestration event records and service level target records, remove samples that have experienced service level target violations, service instance restarts or migrations, database master-slave switching, cache failures, network interruptions, failure rollbacks or manual failure work orders, and determine the remaining samples as historical fault-free samples.
[0058] Request types are categorized based on the request interface identifier, request method, whether state writing is involved, dependency chain identifier, and parameter size range. The parameter size range is defined based on the 25th, 50th, and 75th percentile values of the byte length of historical request parameters. Resource load ranges are defined based on processor utilization, memory utilization, number of active requests, and database connection occupancy in the sampling period prior to the request entering the target microservice.
[0059] For the same request type, the same resource load range, and the same resource channel, the 5th percentile and 95th percentile of the difference amplitude of historical fault-free samples are used as the lower limit and upper limit of the normal difference range, respectively. When there are fewer than 30 sets of historical fault-free samples, the range is expanded to the adjacent resource load range and all historical fault-free samples of the target microservice, and a baseline sample shortage flag is generated.
[0060] Calculate the degree of standardization exceedance for the r-th resource channel: ; In the formula, is the standardized over-limit degree of the r-th resource channel; is the current differential amplitude; and are the upper limit and lower limit of the corresponding normal differential range respectively.
[0061] When 3 consecutive resampling positions in the same resource channel exceed the upper limit of the normal differential range, the first over-limit position is determined as the starting position of the hidden abnormal section; when the subsequent 2 consecutive resampling positions fall back into the normal differential range, the position preceding the first falling-back position is determined as the ending position of the hidden abnormal section.
[0062] In response to 3 consecutive resampling positions of the resource channel exceeding the limit, the actual time corresponding to the first over-limit position is determined as the occurrence time of the first response abnormality; when the starting time of the hidden abnormal section of the resource channel is earlier than the occurrence time of the first response abnormality by at least one sampling period, the resource channel is determined to have the property of failure precursor.
[0063] The hidden anomaly confidence is determined according to the following formula: ; In the formula, is the hidden anomaly confidence; is the average value of the standardized over-limit degree in the hidden abnormal section; is the ratio of the duration of the hidden abnormal section to the duration of the current request phase; is the proportion of requests with abnormal resource channels of the same type among the past 5 consecutive requests of the same type. When the hidden anomaly confidence is not less than 0.60, the corresponding hidden abnormal section is output.
[0064] To determine the specific monitoring masking operation, three sub-observation dark windows are established respectively, each sub-observation dark window only suspends one type of active monitoring operation among active health check, connection keep-alive and cache preheating, and keeps the other two types running; the active monitoring operation that can make the same hidden abnormal section reappear is determined as the monitoring masking operation, and a monitoring masking operation identifier is generated.
[0065] S320, constructing a propagation path and determining candidate resource channels. A resource channel refers to an operation link that undertakes the function of the same type of computing resources when the target microservice processes service requests, including connection resource channels, cache resource channels, code page resource channels, thread resource channels, database resource channels, network resource channels, storage resource channels and response resource channels. A propagation path refers to a directed resource channel sequence formed by at least 2 resource channels according to the occurrence order of anomalies and the actual resource call relationship.
[0066] Read the dependency service topology formed by S110, and use the call chain events, cache origin events, database access events, thread scheduling events, network connection events and storage access events corresponding to the current observation dark window batch to make corrections, retaining only the dependency edges that the current request actually passed through.
[0067] The propagation edge from resource channel r to resource channel s must simultaneously satisfy the following conditions: the differential start time of resource channel r is at least one sampling period earlier than that of resource channel s; the two resource channels have a direct resource call relationship; the two latent anomaly segments overlap in time, or the interval between the end of the previous latent anomaly segment and the start of the next latent anomaly segment does not exceed two sampling periods; the latent anomaly confidence of resource channel r is not lower than 0.60; and the resource channel weight of resource channel r is not lower than the candidate channel weight threshold. The candidate channel weight threshold is the larger of the 75th percentile value of the resource channel weight corresponding to the historical fault-free sample under the same request type and resource load range, and 0.10.
[0068] When multiple upstream resource channels point to the same downstream resource channel, the propagation edge with the earliest differential start time is selected first; when the differential start times are the same, the propagation edge with the largest resource channel weight is selected; when both the differential start time and the resource channel weight are the same, it is retained as a parallel propagation branch. When at least 3 out of 5 consecutive requests of the same type form the same starting resource channel and propagation direction, the corresponding propagation paths are merged into a duplicate propagation path.
[0069] Resource channels that precede response resource channels in the propagation path, possess pre-fault warning attributes, and meet the candidate channel weight threshold are identified as candidate resource channels, forming a candidate resource channel record. The candidate resource channel record includes instance pairing identifier, request lineage identifier, target microservice identifier, candidate resource channel, upstream resource channel, downstream resource channel, anomaly start and end times, implicit anomaly confidence level, resource channel weight, number of repetitions, and propagation level.
[0070] S330. Generate a set of candidate self-healing actions and output an abnormal status warning. Based on the candidate resource channels, read the applicable resource channel, control interface identifier, action type, parameter upper and lower limits, minimum adjustment step size, duration range, rollback interface, and calling permissions from the cloud platform control interface registry.
[0071] The container resource adjustment interface originates from the container orchestration control plane, the request ratio adjustment interface originates from the service mesh control plane, the database connection adjustment interface originates from the connection pool management interface, the cache rebuild interface originates from the cache management interface, and the storage path switching interface originates from the storage routing control interface. Candidate self-healing actions are generated only for registered control interfaces that have invocation permissions and support rollback, and the candidate resource channel bound to each candidate self-healing action is determined as the target candidate resource channel.
[0072] Connection resource channels correspond to actions such as releasing failed connections, adjusting connection pool lower limits, or limiting connection establishment rates; cache resource channels correspond to actions such as rebuilding specified cache partitions, adjusting cache warm-up periods, or switching cache access paths; code page resource channels correspond to actions such as code page preloading, service instance rebuilding, or image layer switching; thread resource channels correspond to actions such as adjusting thread concurrency, adjusting vCPU quotas, or limiting request concurrency; database resource channels correspond to actions such as adjusting the number of connections, switching read-only replicas, or limiting transaction concurrency; network resource channels correspond to actions such as adjusting request allocation ratios or switching network paths; and storage resource channels correspond to actions such as switching storage replicas, adjusting input / output quotas, or switching data access paths.
[0073] The response resource channel serves only as the termination resource channel in the propagation path and does not directly generate candidate self-healing actions. The maximum scope of candidate self-healing action 'a' is: ; In the formula, The maximum effective scale of candidate self-healing action a; This refers to the maximum allowable adjustment amount for the control interface; This represents the maximum adjustment that the current remaining resources can support. This is the 95th percentile of the size of similar actions that have not resulted in a default on service level targets in the past 30 days.
[0074] Resource remaining quantity refers to the absolute amount of resources available for allocation after deducting the current occupied quantity and the reserved quantity from the upper limit of resource allocation. Rollback conditions are triggered when any resource indicator exceeds the resource safety limit for two consecutive sampling periods, the request failure rate exceeds the service level target, the first response time deteriorates for three consecutive sampling periods, or the dark window differential amplitude increases for three consecutive sampling periods.
[0075] Each candidate self-healing action records the action identifier, target candidate resource channel, control interface identifier, rollback interface identifier, action type, maximum impact scale, minimum adjustment step size, duration range, resource safety upper limit, rollback conditions, implicit anomaly confidence, resource channel weight, propagation path, and action reversibility identifier, forming a set of candidate self-healing actions.
[0076] S4 specifically includes the following sub-steps: S410: Generate and apply a reversible rescue probe. Read the candidate self-healing action set formed in S330, and verify the target candidate resource channel, control interface identifier, rollback interface identifier, maximum action scale, duration range, resource safety limit, and rollback conditions. A reversible rescue probe refers to a control action that calls the same control interface corresponding to a candidate self-healing action, applies a control action to the same target candidate resource channel with a smaller than maximum action scale, limited duration, and the ability to restore the channel to its pre-application state.
[0077] Adjustments to thread concurrency, vCPU quotas, code page preloading, and cache partitions apply to instances observed in the dark window. Adjustments involving shared databases, shared caches, shared network paths, or shared storage paths only apply to isolated state replicas, shadow message topics, shadow network paths, or shadow storage directories established by S130. Candidate self-healing actions that cannot be limited in scope or lack a rollback interface are marked as undetectable actions.
[0078] The probe action scale corresponding to candidate self-healing action a is: ; In the formula, The probe action scale corresponding to candidate self-healing action a; To maximize the scale of action; This is the minimum adjustment step size for the control interface; The detection ratio is set to 0.05 to 0.15.
[0079] Integer parameters are rounded down according to the minimum adjustment step size. Each reversible rescue probe test includes the following phases in sequence: pre-probe baseline phase, probe action phase, probe rollback phase, and resource recovery phase. The pre-probe baseline phase lasts for 5 sampling cycles; the duration of the probe action phase is the maximum value among the shortest duration of the control interface, the 95th percentile of the completion time of the same type of request, and the three sampling cycles, and must not exceed the maximum duration limited by S330; the probe rollback phase starts from the call to the rollback interface and continues until the control interface returns a successful parameter recovery; the resource recovery phase starts from the parameter recovery time and continues until each resource indicator returns to within 10% of the pre-probe baseline value for three consecutive sampling cycles.
[0080] The next repetitive test can only be performed after the previous reversible rescue probe has completed rollback and resource recovery; a total of 3 sets of valid test results were obtained, and the median value of the corresponding response features of the 3 sets was used as the final probe response feature. This forms the probe response time series dataset.
[0081] S420. Extract probe response features and calculate rescue action impedance. Input the probe response time-series dataset into the S220 depth time-series feature extraction model, and use the quantile normalization parameters determined in S210. The pre-probe baseline value is the median value of the corresponding resource index within the pre-probe baseline stage.
[0082] Response gain refers to the ratio of the maximum change of the resource indicator relative to the pre-probe reference value to the proportion of the probe action scale to the maximum action scale, and is normalized by dividing by the 99th percentile of the historical fault-free value of the corresponding resource indicator; maximum overshoot amplitude refers to the maximum deviation of the resource indicator from the median value of the last 3 valid sample values during the probe action phase, and is normalized by dividing by the difference between the resource safety upper limit and the pre-probe reference value; recovery time refers to the time taken for the rollback interface confirmation parameters to recover to the point where the resource indicator returns to within 10% of the pre-probe reference value for 3 consecutive sampling cycles; number of affected resources refers to the number of resource channels whose response amplitude after probe action exceeds the 95th percentile of the historical fault-free probe record.
[0083] The differential attenuation of the dark window is calculated according to the following formula: ; In the formula, The differential attenuation of the dark window corresponding to candidate self-healing action a; The average differential amplitude of the target candidate resource channel in the latent anomaly segment is applied to the probe before it is applied; This represents the average differential amplitude under the same request type and resource load range during the stable phase of probe operation.
[0084] The resistance to rescue actions is calculated using the following formula: ; In the formula, The resistance to the rescue action corresponding to candidate self-healing action a; Normalized response gain; This represents the normalized maximum overshoot amplitude. This is the ratio of the recovery time to the longest recovery time. This is the ratio of the number of affected resource channels to the total number of resource channels.
[0085] When the differential attenuation of the dark window is not greater than 0, the rescue action impedance is marked as invalid impedance. For the same candidate self-healing action, incremental tests are performed at 1x, 2x, and 3x probe action scales. If the impedance growth rate of two consecutive action levels exceeds 20%, the response gain increase exceeds twice the probe scale increase, the recovery time exceeds the longest recovery time, the differential attenuation of the dark window turns non-positive, or the rollback condition is triggered, the impedance state is marked as divergent. If none of the above conditions occur, it is marked as non-divergent. If at least two action levels cannot be formed, it is marked as undeterminable.
[0086] S430. Generate self-healing action ranking results using reinforcement learning. Reinforcement learning refers to a machine learning method that uses the current running state of the cloud platform as the reinforcement learning state, candidate self-healing actions and their parameters as reinforcement learning actions, constructs reinforcement learning rewards based on the changes in dark window difference and resource response after the action is executed, and updates the action selection strategy according to the correspondence between state, action and reward.
[0087] The reinforcement learning state is composed of fixed dimensions, including the abnormal state warning level, the confidence level of latent anomalies, the differential amplitude and weight of each resource channel, the propagation path length, the propagation level of the target candidate resource channel, the rescue action impedance and impedance status of each candidate self-healing action, the remaining resource ratio of processor, memory, network, storage, and database connections, the request failure rate, the first response time, and the executed action identifier. The remaining resource ratio refers to the ratio of the remaining resource amount to the corresponding resource configuration limit.
[0088] If there are fewer than 10 candidate self-healing actions, they are filled with 0. If there are more than 10, the top 10 are retained according to the confidence of latent anomalies, resource channel weights, and dark window differential attenuation. The candidate self-healing action identifier, 25%, 50%, 75%, or 100% of the maximum effect size, and 1, 2, or 3 times the shortest duration of the control interface are combined into reinforcement learning actions. Candidate self-healing actions with impedance states of divergence, invalid impedance, indeterminate, or without a rollback interface are set as unselectable actions.
[0089] The reward for enhanced learning is determined by the following formula: ; In the formula, The reward for the t-th state transition is the reinforcement learning reward. This represents the differential attenuation amount during the dark window. Normalized resource response gain; This represents the normalized maximum overshoot amplitude. This represents the increase in the request failure rate relative to the baseline value before the action was executed. This is the rollback flag; it is 0 if no rollback is triggered, and 1 if a rollback is triggered. to It is a non-negative return weight.
[0090] The reinforcement learning strategy model specifically employs a deep action value network based on a deep Q-network (DQN). The network structure sequentially includes an input layer, at least two fully connected hidden layers with 128 and 64 neurons respectively (using the ReLU activation function), and an output layer that outputs the Q-value corresponding to each reinforcement learning action. During model training, an experience replay mechanism is used. Training data comes from reversible rescue probe records, formal self-healing action records, rollback records, and current probe response records from the past 30 days. For each valid candidate self-healing action, the action scale and duration with the highest action value are selected. The selected action scale is determined as the recommended total action scale, and the selected duration is determined as the recommended stage duration. Self-healing actions are then ranked from highest to lowest action value.
[0091] S5 specifically includes the following sub-steps: S510: Filter target self-healing actions and generate phased execution configurations. Read the self-healing action sorting results generated in S430, and read historical action records from the monitoring time series database that are the same as the target microservice, target candidate resource channel and action type in the past 30 days, and have not triggered rollback, and whose hidden anomalies have been eliminated after formal execution. Use the 25th percentile of the dark window differential attenuation amount as the preset attenuation threshold; if there are fewer than 20 historical valid actions, use 0.20 as the temporary preset attenuation threshold.
[0092] Candidate self-healing actions are checked according to execution priority. Candidate self-healing actions with the following characteristics are identified as target self-healing actions: impedance status is not diverging, dark window differential attenuation is greater than the preset attenuation threshold, both control interface and rollback interface can be called, recommended total action scale is not less than twice the minimum adjustment step size, target candidate resource channel is still in the current propagation path and the remaining resource amount is sufficient to support execution. If there are no candidate self-healing actions that meet the conditions, an abnormal state warning is maintained and a no-safety target self-healing action flag is output.
[0093] Determine the reference action size for a single execution phase based on the recommended total action size and minimum adjustment step size: ; In the formula, The reference action scale for a single execution phase of the self-healing action a; Recommended total action scale; Minimum adjustment step size; This indicates rounding up to the nearest integer.
[0094] The number of execution phases is determined by the following formula: ; In the formula, The number of execution stages for the target self-healing action 'a'. The scale of the stage effect of each execution stage is taken as follows The final execution phase takes the remaining value after deducting the aforementioned cumulative effect from the total recommended effect size.
[0095] The phase difference convergence objective for the k-th execution phase is: ; In the formula, k is the sequence number of the current execution stage. The phase difference convergence objective for the k-th execution phase; The average differential amplitude of the target candidate resource channel before the target self-healing action is executed; This represents the differential attenuation of the dark window obtained from S420.
[0096] For example, when the recommended total action size is 6 database connections and the minimum adjustment step size is 1 connection, the reference action size for a single execution phase is 2 connections, and the number of execution phases is 3.
[0097] S520 executes the target self-healing action in stages and implements real-time security controls. Before formal execution, it reads the control interface identifier, rollback interface identifier, and calling permissions from the cloud platform control interface registry, and sets action control locks on the target candidate resource channels to prevent other auto-scaling containers, connection pool regulators, cache management tasks, or resource scheduling tasks from modifying the same resource parameters simultaneously.
[0098] During each execution phase of the target self-healing action, the observation dark window state of the observation dark window instance and the monitoring state of the monitoring control instance remain unchanged, and paired running data are collected and processed according to S120-S230 to form a phase dark window differential feature sequence.
[0099] Processor and memory data are sourced from the processor performance monitoring unit and container control group; network and storage data are sourced from the kernel network interface counter and block device interface; database connection usage is sourced from the connection pool management interface; request failure rate and first response time are sourced from the service mesh proxy; and control plane response is sourced from the control interface call log. The median value of resource metrics from the five sampling periods prior to the start of the current execution phase is used as the baseline value for this phase, and the real-time rescue action impedance is determined using the same calculation method as S420.
[0100] When the current execution phase reaches the phase duration, the real-time rescue action impedance does not exceed 120% of the probe rescue action impedance corresponding to the same action scale, all resource indicators do not exceed the phase resource safety limit for three consecutive sampling periods, the current average differential amplitude is not higher than the phase differential convergence target, the request failure rate meets the service level target, and the control interface returns that the current phase parameters have taken effect, the next execution phase will begin.
[0101] If the current average difference magnitude has not yet reached the stage difference convergence target, but continues to decrease for three consecutive sampling periods and other safety conditions are met, extend the stage stability observation period by one period; if the stage difference convergence target is still not reached after the extension, execute the rollback.
[0102] When the following events occur: real-time rescue action impedance exceeds limits, resource indicators exceed the stage resource safety limit for two consecutive sampling periods, request failure rate exceeds service level target, average differential amplitude increases for three consecutive sampling periods, control interface execution fails, or a new downstream resource channel appears in the propagation path, the rollback interface is called to restore the snapshot of control parameters before the start of the current execution stage.
[0103] If the rollback is successful and the propagation path remains unchanged, the real-time rescue action impedance, rollback reason, and remaining resource ratio are written into the reinforcement learning state, and the process returns to S430 to regenerate the self-healing action ranking result. If the propagation path changes or a new candidate resource channel appears, the process returns to S310 to re-identify the hidden abnormal segment and candidate resource channel. If the rollback fails, the action control lock is released, automatic rescue is stopped, and a level three abnormal state warning is generated.
[0104] S530. Re-establish the observation dark window and verify whether the latent anomaly has been eliminated. After the target self-healing action is completed, when the control interface returns the execution completion status, the processor, memory, network, storage and database connection indicators have not exceeded the resource safety limit for 5 consecutive sampling periods, the request failure rate meets the service level target, there are no incomplete rollback operations, and the current resource load range is consistent with the resource load range corresponding to the business request sequence in S130, the cloud platform is determined to have entered the verification stable state and the action control lock is released.
[0105] Before generating the verification request sequence, the first and second copy-on-write spaces are re-established based on the same database snapshot after the target self-healing action is completed. The corresponding independent cache namespaces, shadow message topics, and shadow object directories are cleared so that the verification requests start from the same baseline state.
[0106] The observation window is re-established according to S120, prioritizing the use of the business request sequence saved in S130. When the token, timestamp, or data version in the original request becomes invalid, the current request metadata is read from the upstream gateway of the target microservice. Requests with identical request interface identifiers, request methods, write attributes, parameter size ranges, dependency chain identifiers, and resource load ranges, and whose arrival intervals differ by no more than 10% between adjacent requests, are selected to form a verification request sequence. The verification request sequence must include at least 20 valid paired requests, and each request type that triggers a hidden anomaly must have at least 5 valid paired requests.
[0107] The validation dark window differential feature sequence is formed according to S210-S330, and the validation normal ratio is calculated: ; In the formula, To verify the normal ratio; To verify the number of valid pairing requests where the dark window difference feature is within the corresponding normal difference range; To verify the total number of valid pair requests in the request sequence.
[0108] When the verification success rate is not less than 0.90, the confidence level of the latent anomaly of the original target candidate resource channel is less than 0.60, the original propagation path is no longer formed, the response resource channel is within the normal differential range, no new candidate resource channels appear, and the request failure rate and first response time meet the service level target, the latent anomaly is determined to have been eliminated.
[0109] After the hidden anomalies are eliminated, only the active monitoring operations corresponding to the monitoring masking operation identifier are configured and adjusted: the active health check for accessing downstream dependencies is replaced with a bypass liveness check, the continuous connection keep-alive is adjusted to periodic connection reconstruction verification, the fixed periodic cache preheating is adjusted to delayed preheating based on the predicted service arrival time or on-demand cache filling after the service request arrives, and the adjusted monitoring configuration is written to the microservice configuration center to remove the abnormal state warning.
[0110] If verification fails and the original propagation path remains unchanged, and no new candidate resource channel appears, return to S410; if the propagation path changes or a new candidate resource channel appears, return to S310; if only the action sorting needs to be updated based on the newly added execution result, return to S430. Each abnormal status warning corresponds to no more than 3 rounds of automatic rescue. If the hidden abnormality is not eliminated after 3 consecutive rounds, or if rollback fails, control interface fails, or resource security limits are continuously exceeded, automatic rescue is stopped and manual handling is initiated.
[0111] All the above formulas are performed using dimensionless numerical calculations; the relevant formulas are based on empirical models that approximate the real situation, obtained through extensive data collection and software simulation fitting. The preset parameters and thresholds involved in the formulas can be conventionally set and adjusted by those skilled in the art according to the physical constraints of the actual application scenario.
[0112] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0114] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A cloud platform anomaly state early warning method based on deep temporal feature extraction, characterized in that, Includes the following steps: S1. Obtain the resource configuration, dependency topology, and load data of the target microservice, select a pair of service instances and set them as the observation dark window instance and the monitoring control instance, pause the active monitoring operation of the observation dark window instance, send the same source business request, and form a pair of runtime sequence datasets. S2. Perform request phase alignment and normalization on the paired runtime time series datasets to form a dual instance aligned time series matrix. Input it into the deep time series feature extraction model to obtain the observation dark window time series features and the monitoring control time series features, forming a dark window difference feature sequence. S3. Based on the fault-free samples, establish the normal differential range, identify the hidden abnormal segments in the dark window differential feature sequence, combine the differential start time, resource channel weight and resource call relationship to construct the propagation path, determine the candidate resource channels, and generate a set of candidate self-healing actions and abnormal state warnings. S4. Generate reversible rescue probes for the candidate self-healing action set, collect probe response time series data, calculate the dark window differential attenuation and rescue action impedance, construct the reinforcement learning state and output the self-healing action ranking results.
2. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 1, characterized in that, Also includes: S5. Select the target self-healing action from the self-healing action ranking results. The rescue action impedance has not diverged and the dark window differential attenuation meets the threshold. After the action is executed in stages, the observation dark window is re-established to form a verification dark window differential feature sequence. Based on the verification normal ratio, the confidence of the latent anomaly and the propagation path, it is determined whether the latent anomaly has been eliminated. If it has not been eliminated, the reinforcement learning state is updated.
3. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 1, characterized in that, S1 specifically includes: Obtain the instance status, resource configuration, dependency topology, and historical load data of the target microservice; select paired service instances based on resource configuration differences and historical load interval overlap rates; determine the observation dark window instance and the monitoring control instance; and generate instance pairing identifiers. Suspend proactive health checks, connection keep-alive, and cache preheating for instances in the dark window, while retaining host-level liveness detection. Determine the duration of the dark window based on the operation cycle and security observation duration, and synchronously collect processor, memory, network, and disk operation data for paired service instances.
4. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 3, characterized in that, Also includes: The system generates a business request sequence based on the request metadata from the upstream gateway, configures a request lineage identifier for the same business request, configures a shadow execution identifier for write requests and routes them to an isolated replica, and merges request phase data with runtime data to form a pair of runtime sequence datasets.
5. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 1, characterized in that, S2 specifically includes: Based on instance pairing identifiers, request lineage identifiers, and shadow execution identifiers, pairwise runtime sequence datasets are associated. Sampling records are converted into relative stage positions according to request stages, and resampling, missing value processing, and quantile normalization are performed to form a dual-instance aligned time series matrix. By inputting the dual-instance aligned time series matrix into a deep time series feature extraction model with parameter-shared time series coding branches, short-term mutation features, long-term cumulative features, and cross-resource propagation features are extracted to obtain observation dark window time series features and monitoring control time series features. By comparing the observation dark window time series characteristics with the monitoring control time series characteristics according to the same request stage, resource channel, and relative stage position, the differential amplitude, differential start time, differential duration, propagation level, and resource channel weight are determined to form a dark window differential feature sequence.
6. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 1, characterized in that, S3 specifically includes: Based on the historical fault-free dark window differential feature sequence, a normal differential range is established according to request type, resource load range and resource channel. The standardized over-limit degree and the confidence of hidden anomalies are calculated. Hidden anomaly segments with fault precursor attributes are identified and the monitoring cover-up operation identifier is determined. By combining the differential start time of the hidden anomaly segment, the resource channel weight, the dependent service topology, and the actual resource call relationship, a propagation path is constructed. Resource channels that are located before the response resource channel and meet the threshold are identified as candidate resource channels, forming a candidate resource channel record.
7. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 6, characterized in that, Also includes: Based on the candidate resource channel query cloud platform control interface registry, determine the control interface and action type that support rollback, set the maximum scale of action, resource security limit and rollback conditions, form a set of candidate self-healing actions and output abnormal status warnings.
8. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 1, characterized in that, S4 specifically includes: Based on the set of candidate self-healing actions, a reversible rescue probe is generated that acts on the target candidate resource channel. The probe pre-baseline, probe action, probe rollback and resource recovery are executed sequentially to form a probe response time series dataset. Deep time-series feature extraction is performed on the probe response time-series dataset to determine the response gain, maximum overshoot amplitude, recovery time, number of affected resources, and dark window differential attenuation. Based on this, the rescue action impedance is calculated and the impedance status is determined. Based on abnormal state early warning, candidate resource channel records, rescue action impedance, resource remaining ratio and operational performance indicators, a reinforcement learning state is constructed. Candidate self-healing actions, their scale of action and duration are constructed as reinforcement learning actions. The reinforcement learning strategy model outputs a ranking result of self-healing actions that includes execution priority, recommended total scale of action and recommended stage duration.
9. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 2, characterized in that, S5 specifically includes: Select target self-healing actions from the self-healing action ranking results. The target self-healing actions are selected if the impedance state is not divergent, the dark window differential attenuation exceeds the threshold, and the resource conditions are met. The execution stage and the stage differential convergence target are determined according to the recommended total action scale and the minimum adjustment step size. Set action control locks for the target self-healing action and execute them in stages, maintain the observation dark window state, calculate the real-time rescue action impedance, and control the stage advancement or rollback based on the stage difference convergence target, resource safety limit and rollback conditions.
10. The cloud platform abnormal state early warning method based on deep temporal feature extraction according to claim 9, characterized in that, Also includes: After the target self-healing action is completed, an observation dark window is established to form a verification dark window differential feature sequence. Based on the verification normal ratio, the confidence of the latent anomaly and the propagation path, it is determined whether the latent anomaly has been eliminated. If it has not been eliminated, return to the corresponding step to update the decision. If it has been eliminated, adjust the monitoring configuration corresponding to the monitoring cover-up operation and remove the anomaly state warning.