Self-adaptive adjustment Internet operation and maintenance strategy generation method and self-adaptive adjustment Internet operation and maintenance strategy generation system

By collecting data through distributed sensors and edge nodes, conducting spatiotemporal correlation analysis and generating dynamic threshold baselines, the resource allocation problem of traditional Internet operation and maintenance strategies under burst traffic is solved, and real-time adjustment and resource optimization of adaptive operation and maintenance strategies are achieved, improving service stability and efficiency in scenarios such as e-commerce promotions.

CN120811892AInactive Publication Date: 2025-10-17SHENZHEN SHENMA NETWORK TECH CO LTD
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202511307992.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional Internet operation and maintenance strategy generation methods are difficult to adjust in a timely manner under sudden traffic conditions such as e-commerce promotions, resulting in improper resource allocation and affecting service stability and efficiency.

Method used

Through distributed sensors and edge nodes, dynamic operation and maintenance environment data is collected to generate multi-dimensional time series data sets, conduct spatiotemporal correlation analysis, identify steady-state behavior patterns and abnormal boundaries, generate dynamic threshold baselines, combine with real-time streaming data for matching verification, generate adaptive operation and maintenance strategies, and achieve elastic resource scheduling and fault switching.

Benefits of technology

It achieves real-time adaptive adjustment of operation and maintenance strategies, reduces the duration of fault impact, optimizes resource allocation, and improves operation and maintenance efficiency and service stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811892A_ABST
    Figure CN120811892A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive adjustment Internet operation and maintenance strategy generation method and system, and relates to the technical field of Internet, and the method comprises the steps: 1, dynamically collecting the operation state data of a target operation and maintenance environment through a distributed sensor and an edge node, and generating a multi-dimensional time series data set; 2, performing space-time correlation analysis on the multi-dimensional time sequence data set, determining a reference data node, constructing a two-dimensional correlation structure, and generating a dynamic judgment interval; and step 3, respectively selecting monitoring sample sets in the inner domain and the outer domain of the dynamic judgment interval, generating a trajectory feature sequence according to the time evolution relationship of the sample sets, and calculating a dynamic correction coefficient based on the trajectory feature sequence. According to the method, the self-adaptive circulation control is formed by dynamically adjusting the threshold baseline, the judgment interval and the strategy generation rule, the accuracy and effectiveness of the internet operation and maintenance strategy are improved, and the stability of internet operation is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet, in particular to an adaptive adjustment Internet operation strategy generation method and system. BACKGROUND

[0002] In the e-commerce promotion, the traditional operation strategy generation method has some limitations, and the traditional strategy mostly sets fixed rules according to past data, which is difficult to adjust in time when the traffic bursts, for example, before double 11, the beauty category receives 3 times more than expected traffic due to live broadcast, and the preset resource scheme cannot be quickly scheduled, resulting in page loading delay exceeding 8 seconds; for example, during the promotion, the time-limited discount of a certain home appliance brand causes a sharp increase in detail page clicks, and the traditional strategy has no temporary response mechanism, and the operation manual rule change takes 20 minutes, during which the system is frequently stalled.

[0003] In addition, the traditional strategy does not consider the differentiated needs of businesses, which is easy to cause resource waste or deficiency, the platform allocates cache resources to all categories according to average demand, and during the promotion, the cache of the clothing category is in short supply and the query delay is caused, while the utilization rate of the cache of the home furnishing category is less than 40%; for example, during the payment peak, the database performance is required to be high, but the traditional strategy does not have special expansion rules, and the shared resources cause payment delay and partial order failure. SUMMARY

[0004] The technical problem to be solved by the present application is to provide an adaptive adjustment Internet operation strategy generation method and system, which can adaptively adjust the operation strategy and improve the operation efficiency and service stability.

[0005] To solve the above technical problems, the technical scheme of the present application is as follows: In a first aspect, an adaptive adjustment Internet operation strategy generation method is provided, which comprises: Step 1: dynamically collecting the running state data of the target operation environment through distributed sensors and edge nodes to generate a multi-dimensional time series data set; Step 2: performing spatio-temporal correlation analysis on the multi-dimensional time series data set to determine the reference data node and build a two-dimensional correlation structure to generate a dynamic judgment interval; Step 3: selecting a monitoring sample set in the inner domain and the outer domain of the dynamic judgment interval, generating a trajectory feature sequence according to the time evolution relationship of the sample set, and calculating a dynamic correction coefficient based on the trajectory feature sequence; Step 4: using the dynamic correction coefficient to perform feature extraction and correlation analysis on the multi-dimensional time series data set, identifying the steady state behavior pattern and abnormal boundary of the operation environment, and generating a dynamic threshold baseline; Step 5: matching and verifying the dynamic threshold baseline with real-time flow data, combining the periodic phase alignment mechanism to generate an abnormal type and current operation object bottleneck positioning result; Step 6, based on the abnormal classification and the current operation object bottleneck positioning result, integrate resource constraint conditions and service level agreement, generate and execute operation strategy set including resource elastic scheduling rule, fault switching link and service degradation plan; Step 7, according to the state evolution data obtained after the execution of the operation strategy set, adjust the dynamic threshold baseline, the decision interval configuration and the strategy generation rule, and form the adaptive loop control parameter set.

[0006] The second aspect is an adaptive adjustment Internet operation strategy generation system, comprising: A data acquisition module is configured to dynamically acquire running state data of a target operation environment through distributed sensors and edge nodes, and generate a multi-dimensional time series data set. An association analysis module is configured to perform spatio-temporal association analysis on the multi-dimensional time series data set, determine a reference data node and build a two-dimensional association structure, and generate a dynamic decision interval. A correction coefficient module is configured to select a monitoring sample set in the inner domain and the outer domain of the dynamic decision interval, generate a trajectory feature sequence according to the time evolution relationship of the sample set, and calculate a dynamic correction coefficient based on the trajectory feature sequence. A threshold generation module is configured to use the dynamic correction coefficient to perform feature extraction and correlation analysis on the multi-dimensional time series data set, identify the steady-state behavior pattern and the abnormal boundary of the operation environment, and generate a dynamic threshold baseline. An abnormal bottleneck module is configured to match and verify the dynamic threshold baseline with real-time flow data, generate an abnormal type and a current operation object bottleneck positioning result in combination with a periodic phase alignment mechanism. An operation strategy module is configured to integrate resource constraint conditions and service level agreement based on the abnormal classification and the current operation object bottleneck positioning result, generate and execute an operation strategy set including resource elastic scheduling rule, fault switching link and service degradation plan. An adaptive adjustment module is configured to adjust the dynamic threshold baseline, the decision interval configuration and the strategy generation rule according to the state evolution data obtained after the execution of the operation strategy set, and form an adaptive loop control parameter set.

[0007] The third aspect is a computing device, comprising: One or more processors; A storage device is configured to store one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the method.

[0008] The fourth aspect is a computer readable storage medium, the computer readable storage medium stores a program, the program is executed by a processor to implement the method.

[0009] The above solution of the present invention includes at least the following beneficial effects: Through a dynamic collection architecture of distributed sensors and edge nodes, real-time capture and preliminary processing of operation and maintenance data is achieved, with the latency of generating multi-dimensional time series datasets controlled to milliseconds. Combined with a dynamic threshold baseline and an anomaly identification mechanism featuring periodic phase alignment, this system automatically matches real-time streaming data with the threshold baseline, minimizing the duration of fault impact on services. The spatiotemporal relationship between data is captured based on a spatiotemporal coupling matrix, avoiding the bias of single-dimensional analysis. Dynamic threshold baselines are generated through steady-state behavior recognition, and the thresholds can be adaptively adjusted based on the operating state. During policy generation, dynamic resource allocation is achieved through resource elastic scheduling rules. By combining bottleneck location results with resource constraints, resource allocation ratios are automatically calculated. During off-peak periods, idle servers are dispatched to other demand nodes, while during peak periods, core service nodes are pre-loaded and bandwidth allocation is optimized, achieving full automation of the operation and maintenance process. State evolution data after the execution of the operation and maintenance policy is used as feedback to dynamically adjust the response sensitivity of the dynamic threshold baseline, the boundary tolerance range of the judgment interval, and the priority of the policy generation rules. Trends are identified through state evolution data, and server load thresholds are automatically lowered to avoid overload, allowing the operation and maintenance policy to adapt to dynamic changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flowchart of a method for generating an adaptively adjusted Internet operation and maintenance strategy provided by an embodiment of the present invention.

[0011] Figure 2 Schematic diagram of an adaptively adjusted Internet operation and maintenance strategy generation system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0013] like Figure 1 As shown, an embodiment of the present invention proposes a method for generating an adaptively adjusted Internet operation and maintenance strategy, the method comprising the following steps: Step 1: Dynamically collect the operating status data of the target operation and maintenance environment through distributed sensors and edge nodes to generate a multi-dimensional time series data set; Step 2: Perform spatiotemporal correlation analysis on the multi-dimensional time series dataset, determine the benchmark data nodes, construct a two-dimensional correlation structure, and generate a dynamic decision interval; Step 3, respectively selecting a monitoring sample set in the inner domain and the outer domain of the dynamic determination interval, generating a trajectory feature sequence according to the time evolution relationship of the sample set, and calculating a dynamic correction coefficient based on the trajectory feature sequence; Step 4, using the dynamic correction coefficient, performing feature extraction and correlation analysis on the multi-dimensional time series data set, identifying the steady state behavior pattern and abnormal boundary of the operation and maintenance environment, and generating a dynamic threshold baseline; Step 5, matching and verifying the dynamic threshold baseline with real-time stream data, combining the periodic phase alignment mechanism to generate abnormal type and current operation object bottleneck positioning result; Step 6, based on the abnormal classification and current operation object bottleneck positioning result, integrating resource constraint conditions and service level agreement, generating and executing an operation strategy set including resource elasticity scheduling rules, fault switching link and service degradation contingency; Step 7, according to the state evolution data obtained after executing the operation strategy set, adjusting the dynamic threshold baseline, the determination interval configuration and the strategy generation rule, forming a self-adaptive cycle control parameter set.

[0014] In the embodiment of the application, the running state data is dynamically collected by the distributed sensor and the edge node, which can cover the whole scene of the target operation and maintenance environment, and can also capture performance indicators and business indicators synchronously. The generated multi-dimensional time series data set can fully reflect the running situation, avoid the operation and maintenance judgment deviation caused by one-sided data, reduce the misjudgment problem caused by data missing, determine the reference data node through space-time correlation analysis, and then construct a two-dimensional correlation structure integrating space topology relationship and time fluctuation characteristics, which can accurately capture the linkage law of data in different space-time scenes. The dynamic determination interval generated based on this can flexibly adjust the range according to the running state, rather than a fixed interval threshold, effectively avoiding the problem of insufficient adaptation of fixed interval to dynamic data, making the data classification more in line with the actual running situation, selecting sample sets in the inner domain and the outer domain of the dynamic determination interval, generating trajectory feature sequences according to the time evolution, which can intuitively present the state change law of normal and abnormal edges. The dynamic correction coefficient calculated based on the sequence can real-time calibrate the fluctuation amplitude of the multi-dimensional time series data, eliminate the interference factors in the data collection process, make the feature extraction and correlation analysis more in line with the real running state, and reduce the operation and maintenance strategy failure caused by data deviation.

[0015] After processing the data with dynamic correction coefficients, the steady-state behavior pattern is identified through feature extraction and correlation analysis, and the abnormal boundary is determined to generate a dynamic threshold baseline. The threshold can be adjusted adaptively according to the state. The dynamic threshold baseline is matched and verified with real-time flow data, combined with the periodic phase alignment mechanism, which can quickly identify the type of abnormality and accurately locate the bottleneck of the operation and maintenance object, reduce the impact of the fault on the business duration, and based on the results of abnormal classification and bottleneck positioning, integrate resource constraints and service level agreements to generate a strategy set including resource elasticity scheduling, fault switching and service degradation, which can solve the problem specifically and avoid invalid operation and maintenance operations to improve the efficiency of operation and maintenance.

[0016] In a preferred embodiment of the present application, step 1, the running state data of the target operation and maintenance environment is dynamically collected by distributed sensors and edge nodes to generate a multi-dimensional time series data set, In an embodiment of the present application, first, the coverage boundary of the target operation and maintenance environment is determined, including the hardware devices, network links and business systems that need to be monitored; then, combined with the operation and maintenance requirements, the running state indicators to be collected are sorted out, including basic performance indicators such as CPU usage, memory occupancy, disk read / write rate and network bandwidth utilization; business-related indicators such as the number of business requests per second, request success / failure rate, business response delay and environmental indicators such as device operating temperature and computer room humidity; a complete list of collection indicators is formed, and distributed sensors are deployed to each monitoring object according to the collection range, such as installing temperature sensors and performance sensors on the mainboard of each server and deploying flow sensors on network link nodes; at the same time, data receiving mechanisms are configured on the edge nodes of the operation and maintenance environment, such as edge servers and gateway devices close to the monitoring objects, and the collection trigger mechanism is set, including time trigger, such as collecting CPU usage every 100 milliseconds and collecting network bandwidth every 1 second, and event trigger, such as automatically increasing the collection frequency of the indicator when the business request volume exceeds the preset base value, to ensure dynamic response to running state changes.

[0017] After the edge node receives the raw data transmitted by each distributed sensor, it first performs data validity check to eliminate obviously abnormal invalid data, such as values beyond the reasonable range caused by sensor failure, and blank values caused by data transmission interruption. Then, it performs data format uniform processing to convert heterogeneous data output by different sensors, such as Celsius data of temperature sensors and Mbps data of flow sensors, into a unified structured data format, ensuring data compatibility and correlation. Finally, it preliminarily associates multiple types of index data of the same monitoring object according to the collection timestamp, such as binding the CPU usage, memory occupancy, and business response delay data of a certain server at the same time point to form a preliminary data unit of a single object and multiple indexes, and then aggregates all single object and multiple index data units preprocessed by the edge nodes to the data integration link. First, it classifies the data according to the monitoring object dimension, such as grouping all server-related data into one category and network link data into another category, to clearly define the spatial dimension attribute of the data. Then, it sorts the multiple index data of each type of monitoring object according to the collection timestamp in chronological order to form sequence data with time as the axis, reflecting the time dimension attribute of the data. Finally, it aligns the timestamps to associate and match the data of different monitoring objects and different indexes at the same time point, such as associating the CPU usage of server A, the bandwidth occupancy of network link B, and the request volume data of business system C at a certain time, to finally generate a multi-dimensional time series data set containing monitoring object-index type-time-value multi-dimensional information and arranged in chronological order.

[0018] In a preferred embodiment of the present application, step 2, the spatio-temporal association analysis of the multi-dimensional time series data set, the determination of the reference data node and the construction of the two-dimensional association structure, and the generation of the dynamic judgment interval can include: In the embodiment of the present application, in step 220, based on the spatial topology connection relationship and the time dimension fluctuation characteristics of the multi-dimensional time series data set, the key indicator node having a dominant representation effect on the target operation and maintenance object is identified, and the key indicator node is determined as the reference data node. Specifically, first, a multi-dimensional time series data set is collected, which covers all monitoring node data of the target operation and maintenance object, such as large equipment, industrial systems, etc., including real-time monitoring values, past change records of each node, and installation positions of each node in the physical space, line connection relationships with other nodes, such as cable connection, signal transmission path, etc. When analyzing the spatial topology connection relationship, the connection type of each node with other nodes is sorted one by one, and it is distinguished whether it is a direct physical connection, such as that device A is directly connected with device B through a pipeline, or an indirect logical connection, such as that the data of device C needs to be forwarded to the control center through device D; the connection number of each node, that is, how many nodes are directly connected with it, is counted, and the connection frequency in a unit time, such as the number of data interactions in each hour, is calculated; meanwhile, the connection interruption situation is recorded, the interruption frequency, the number of interruptions per day and the duration of single interruption are calculated, so as to evaluate the stability of the connection. Based on these data, the spatial importance of each node is preliminarily sorted, and the more the connection number, the higher the frequency, the stronger the stability of the node, the higher the spatial importance.

[0019] When analyzing the time dimension fluctuation characteristics, for the time series data of each node, data segments are extracted according to fixed time intervals, such as every minute or every hour, the fluctuation amplitude in each segment is calculated, that is, the difference between the maximum value and the minimum value in the time period; the fluctuation frequency is counted, that is, the number of times that the data changes from rising to falling or from falling to rising; through comparison of a plurality of consecutive periods, the fluctuation period is determined, such as that the data of a node appears a peak value every 8 hours, and the period is 8 hours; meanwhile, the node data fluctuation is matched with the known state of the target operation and maintenance object, such as normal operation, slight abnormality and serious fault, and the probability that the node appears a specific fluctuation, such as a value exceeding the normal range, when the operation and maintenance object is in a certain state is calculated, for example, the proportion of the number of times that the data of node E appears a large fluctuation to the total number of faults when the operation and maintenance object fails.

[0020] The key indicator nodes are identified by combining the spatial and time dimensions, and the spatial importance and time fluctuation correlation degree are divided into 1-5 levels, with level 1 being the lowest and level 5 being the highest. For example, a certain node is connected to 10 nodes and the connection is stable, and the spatial importance is rated as level 5. A certain node appears specific fluctuations in 90% of the fault cases, and the time correlation degree is rated as level 5. According to the characteristics of the operation and maintenance object, the weights of the two dimensions are set. For example, for a system with dense spatial layout, the spatial weight is set to 60%, and the time weight is set to 40%. The comprehensive score of each node (spatial level x spatial weight + time level x time weight) is calculated, and the top 3-5 nodes with the highest comprehensive score are selected as the key indicator nodes of the target operation and maintenance object, which have a dominant representation effect on the state of the target operation and maintenance object, i.e. the reference data nodes.

[0021] In step 221, the distance influence coefficient of each node in the multi-dimensional time series data set and the reference data node in the spatial topology structure is calculated based on the reference data node as the core. Specifically, the spatial topology structure type formed by all nodes is determined first with the reference data node determined in step 220 as the core. If it is an industrial production line, it may be a chain structure, and the nodes are connected in turn according to the production process. If it is a power system, it may be a grid structure, and there are multiple connection paths between nodes. For each non-reference node, the spatial topology distance between it and the reference node is calculated. If it is a physical space distance, the straight-line distance between the two points is obtained using an actual measurement tool such as a laser range finder. If it is a logical space distance, such as a data transmission network, the number of transit nodes that need to be passed through from the node to the reference node is counted. For example, if node F needs to pass through nodes G and H to reach the reference node, the logical distance is 2.

[0022] When calculating the distance influence coefficient, a negative correlation mechanism between distance and influence coefficient is first established. The closer the distance, the greater the influence coefficient. The spatial distances of all nodes are normalized to find the maximum distance and the minimum distance between the reference node and all nodes. The result of (maximum distance - actual distance of a node) ÷ (maximum distance - minimum distance) is used as the initial coefficient, which is in the range of 0-1. Then, the connection strength between nodes is modified. The connection strength is evaluated by data transmission rate, data volume transmitted per unit time, and connection stability, i.e. the proportion of uninterrupted operation time. The higher the transmission rate and the stronger the stability, the greater the correction weight, such as setting the weight range to 0-0.2. The final distance influence coefficient is the initial coefficient plus the correction weight, ensuring that the result does not exceed 1. For example, if the initial coefficient of a certain node is 0.7 and the connection strength correction weight is 0.15, the final distance influence coefficient is 0.85.

[0023] At step 222, based on the distance influence coefficient, the synchronization degree of the fluctuation trend of each node and the reference data node in the time dimension is evaluated, specifically including: based on the distance influence coefficient obtained at step 221, assigning a corresponding weight to each node, the distance influence coefficient being the weight value, ranging from 0 to 1, selecting the same time interval, such as 24 consecutive hours, recording data every 10 minutes, extracting the time series data fluctuation segment of each node and the reference data node, when comparing the fluctuation trend, first, the consistency of the fluctuation direction is counted, at each data sampling point, the data of the node and the reference node is recorded as rising, falling or stable, the proportion of the number of times of the same direction to the total sampling times is calculated, such as 80 times of the same direction in 100 sampling points, the proportion is 80%, then the proportion relationship of the fluctuation amplitude is calculated, the fluctuation amplitude of the node is divided by the fluctuation amplitude of the reference node to obtain the proportion value of each fluctuation, the proportion of the number of times of the proportion value within a predetermined reasonable range, such as 0.8-1.2, is counted, the more stable the proportion, the higher the synchronization, finally, the time difference of the fluctuation peak / valley is calculated, the time points of the peak values of the node and the reference node are recorded, and the time difference between the two is calculated, such as the peak value of the node appearing at 10:05 and the peak value of the reference node appearing at 10:03, the time difference is 2 minutes, and the proportion of the number of times of the time difference within a threshold, such as 5 minutes, is counted.

[0024] The three indicators are weighted and calculated, the proportion of the same direction, the proportion of the stable amplitude, and the proportion of the qualified time difference are multiplied by the node weight, the distance influence coefficient, and the sum of the three results is the synchronization degree of the fluctuation trend in the time dimension, the higher the value, the more synchronized the fluctuation in time, for example, the weight of a certain node is 0.8, the proportion of the same direction is 80% (0.8), the proportion of the stable amplitude is 70% (0.7), and the proportion of the qualified time difference is 90% (0.9), then the synchronization degree is 0.8x0.8+0.8x0.7+0.8x0.9=1.92, the numerical value may exceed 1 according to the calculation method, and can be normalized.

[0025] Step 223, integrate the distance influence coefficient and the time dimension synchronization degree to establish a space-time coupling matrix, generate a two-dimensional correlation structure including spatial dimension correlation and time dimension correlation, specifically including: integrating the distance influence coefficient (spatial dimension) and the time synchronization degree (time dimension), first normalizing the two parameters, if the time synchronization degree exceeds 1, scaling to the range of 0-1 according to the maximum value, ensuring that the two are in the same dimension, selecting the integration rule according to the space-time correlation characteristics of the operation and maintenance object, if the spatial correlation has a greater impact on the operation and maintenance state, such as the physical vibration transmission of large machinery, then the weighted sum of distance influence coefficient × 0.6 + time synchronization degree × 0.4 is adopted; if the space-time correlation is equally important, then the two are multiplied, the greater the product, the higher the coupling strength, for example, the distance influence coefficient of a certain node is 0.8, and the time synchronization degree is 0.7, if the weighted sum is adopted, the coupling strength is 0.8×0.6+0.7×0.4=0.76.

[0026] Construct a space-time coupling matrix, the rows and columns of the matrix correspond to all nodes, including the reference node, the value in the ith row and jth column of the matrix represents the space-time coupling strength between node i and node j, if one of them is a reference node, such as node A as a reference, then directly fill in the coupling strength between node j and A; if both are non-reference nodes, such as nodes B and C, then take the average of the coupling strength between the two and the reference node as the coupling strength between them, such as the strength between B and A is 0.7, and the strength between C and A is 0.6, then the strength between B and C is 0.65; through the matrix, both the spatial distance influence of a certain node to other nodes and the time fluctuation synchronization can be viewed, thereby forming a two-dimensional correlation structure including spatial dimension and time dimension correlation.

[0027] Step 224, according to the parameter distribution of the two-dimensional correlation structure, calculate the stability boundary threshold of the correlation strength under different operating conditions of the target operation and maintenance object, and generate a dynamic judgment interval in the parameter space of the two-dimensional correlation structure based on the stability boundary threshold, specifically including: collecting the two-dimensional correlation structure parameters of the target operation and maintenance object under different operating conditions, the operating conditions include full load operation, half load operation, start-up stage, shutdown stage, high temperature environment, low temperature environment, etc., for each condition, continuously record the distance influence coefficient, time synchronization degree, space-time coupling strength and other parameters of each node, statistical parameter distribution characteristics under each operating condition, calculate the average value, the total sum of all sampling values divided by the sampling number, the maximum value, the maximum parameter value in the sampling period, the minimum parameter value in the sampling period, the standard deviation, reflecting the deviation of the parameter value from the average value, the smaller the deviation, the more stable the parameter, for example, under full load condition, the average value of a certain parameter is 0.6, and the standard deviation is 0.05.

[0028] The stability boundary threshold is calculated according to the operation and maintenance precision requirement, for example, 3 times of the standard deviation for high-precision operation and maintenance, and 2 times of the standard deviation for conventional operation and maintenance, and the average value of the parameters is taken as the center, the average value + 2x standard deviation is taken as the upper boundary by extending upwards, and the average value - 2x standard deviation is taken as the lower boundary by extending downwards, for example, the average value is 0.6, and the standard deviation is 0.05, so that the upper boundary is 0.7, and the lower boundary is 0.5, that is, the stability boundary threshold under the operation condition is 0.5-0.7; a dynamic judgment interval is generated, in a coordinate system with a spatial dimension parameter as the horizontal axis and a time dimension parameter as the vertical axis, according to the stability boundary threshold of different operation conditions, a rectangular interval corresponding to each condition is divided, the horizontal axis range is the upper and lower boundaries of the spatial parameter, and the vertical axis range is the upper and lower boundaries of the time parameter, when the operation condition of the operation and maintenance object changes, such as switching from full load to half load, the threshold interval corresponding to the condition is automatically called to realize dynamic updating of the judgment interval, and ensure that the judgment of the node data association state always matches the current operation condition.

[0029] The reference data node is identified by comprehensively considering the spatial topological connection relationship and the time dimension fluctuation characteristic, the problem of insufficient representativeness of the reference node caused by only relying on a single dimension is avoided, the reference node can truly reflect the core state of the target operation and maintenance object, the accuracy of the overall analysis is improved, the distance influence coefficient is calculated based on the reference node, the spatial correlation degree of each node and the reference node is quantified, the influence range and intensity difference in the spatial dimension are clear, the correlation analysis deviation caused by the fuzzy spatial relationship is avoided, the spatial correlation analysis is more targeted and operable, the time dimension fluctuation synchronization degree is evaluated in combination with the distance influence coefficient, the spatial correlation is considered, and the dynamic cooperativity in time is also embodied, the synchronization misjudgment caused by isolated analysis of the time fluctuation is avoided, the rationality of the time dimension correlation analysis is improved, the spatio-temporal coupling matrix is established by integrating the distance influence coefficient and the time synchronization degree, the correlation relationship in the spatial and time dimensions is systematized and structured, a double-dimension correlation structure is formed, the limitation of single-dimension analysis is broken through, the complex correlation between nodes is comprehensively reflected, the stability boundary threshold is calculated based on different operation conditions and the dynamic judgment interval is generated, so that the correlation analysis result can adapt to the dynamic change of the operation and maintenance object, the misjudgment of the fixed threshold when the operation condition changes is avoided, the flexibility and adaptability of the operation and maintenance object state judgment are improved, and a reliable judgment basis is provided for accurate operation and maintenance.

[0030] In a preferred embodiment of the present application, the step 3, the monitoring sample set is selected in the inner domain and the outer domain of the dynamic judgment interval, the trajectory feature sequence is generated according to the time evolution relationship of the sample set, and the dynamic correction coefficient is calculated based on the trajectory feature sequence, which can include: In the embodiment of the present application, step 330, the first continuous timestamp monitoring sample set is extracted in the inner domain of the dynamic determination interval, and the second continuous timestamp monitoring sample set is extracted in the outer domain, specifically including: first, determining the specific definition standard of the inner domain and the outer domain of the dynamic determination interval, the inner domain refers to the parameter space region of all nodes, including distance influence coefficient, time synchronization degree, etc., which are completely within the stability boundary threshold range determined in step 224, that is, the numerical value of each parameter is neither lower than the lower boundary threshold nor higher than the upper boundary threshold; the outer domain refers to the region where at least one node parameter value exceeds the stability boundary threshold range, for example, the distance influence coefficient of a certain node is lower than the lower boundary, or the time synchronization degree of a certain node is higher than the upper boundary, when extracting the first continuous timestamp monitoring sample set (inner domain sample set), first filter the data from the past running database of the target operation and maintenance object, set the filtering condition as that all node parameters within the continuous timestamp are located in the inner domain, traverse the time sequence in the database, find the continuous time segment that meets the condition, such as 72 hours of continuous data from 00:00 on October 1, 2023 to 00:00 on October 4, 2023, determine the timestamp interval, set according to the collection frequency of operation and maintenance data, such as recording data every 10 minutes, the timestamps are 00:00, 00:10, 00:20, etc., extract the complete data corresponding to all timestamps in this time period, including the real-time parameter value of each node, the operation and maintenance object running state label at the corresponding time point, such as normal, low load, etc., if there is missing data of individual timestamps in the continuous time segment, the average value of the adjacent timestamps before and after should be used for filling, finally forming the first monitoring sample set, the sample number is the total duration divided by the timestamp interval, such as 72 hours x 6 = 432 samples.

[0031] When extracting the second continuous timestamp monitoring sample set (outer domain sample set), the same time length and timestamp interval as the inner domain sample set are used to filter the time segment with at least one node parameter in the outer domain from the database, if there are multiple segments that meet the condition, prefer to select the segment similar to the running condition of the inner domain sample set, such as full load running state; if there is not enough long continuous segment in the outer domain, multiple non-continuous segments can be selected for splicing, such as 3 24-hour segments spliced into 72 hours, but it is necessary to ensure that the timestamps at the splicing position are continuous, and the splicing position is recorded, the data of the 3 timestamps before and after the splicing point should be excluded in subsequent analysis to avoid abnormal fluctuations caused by splicing, extract the complete parameter data of each timestamp, which is completely consistent with the node parameters contained in the inner domain sample set, form the second monitoring sample set, the sample number is the same as the first sample set, such as 432 samples.

[0032] Step 331, for the first monitoring sample set, extract the operation and maintenance state transition path according to the preset time window, generate the inner domain trajectory feature sequence; synchronously extract the operation and maintenance state transition path for the second monitoring sample set, generate the outer domain trajectory feature sequence, specifically including: when processing the first monitoring sample set (inner domain sample set), first determine the specific parameters of the preset time window, the window length is determined according to the state change period of the operation and maintenance object, for example, for a system whose state may change obviously every 2-3 hours, the window length is set to 2 hours, and the sliding step is half of the window length, that is, 1 hour, Ensure that there is overlap between windows to capture continuous changes, traverse the inner domain sample set, and divide the continuous timestamp data into multiple time windows in chronological order, such as the first window containing 12 timestamp data from 00:00-02:00, the second window containing 12 timestamp data from 01:00-03:00, and so on. Calculate the average value of all samples in each window, for each parameter of each node in the window, such as distance influence coefficient, add the parameter values of all timestamps in the window, and then divide by the number of timestamps to get the window average value of the parameter; integrate the parameter average values of all nodes into a vector as the operation and maintenance state value corresponding to the window.

[0033] Compare the operation and maintenance state values of the adjacent two windows, calculate the absolute value of the difference of the corresponding parameters in the two vectors, and take the average value of all differences as the state difference degree; if the state difference degree exceeds the preset small threshold, such as 3% of the range of the inner domain state value, used to filter meaningless small fluctuations, then it is determined that a state transition occurs, the state value before the transition, the state value after the transition and the time point of the transition occurrence, the overlapping boundary time of the two windows, arrange all the operation and maintenance state values of the windows in chronological order, mark the transition mark at the position where the state transition occurs, and form a complete inner domain operation and maintenance state transition path.

[0034] When generating the inner domain trajectory feature sequence, the key features of each window are extracted from the state transition path. If no state transition occurs in the current window, the duration is the window length. If a transition occurs, the time interval from the end of the last transition to the occurrence of the current transition is taken. For the window where the transition occurs, the difference between the state value after the transition and the state value before the transition is calculated. For the window where no transition occurs, the difference is recorded as 0. The statistical unit of time is, for example, the number of state transitions in 24 hours. The state values of the three consecutive windows are compared. If the value of the next window is greater than the value of the previous window, it is recorded as rising. If the value of the next window is less than the value of the previous window, it is recorded as falling. If the values of the next window and the previous window alternate, it is recorded as fluctuation. These features are arranged in the order of time windows. Each window corresponds to an element containing the above four features, forming the inner domain trajectory feature sequence. The sequence length is consistent with the number of time windows. For example, 72 hours of data are divided into 2-hour windows and 1-hour steps. 71 windows can be obtained. The sequence length is 71. When processing the second monitoring sample set (the outer domain sample set), the same time window length (2 hours), sliding step (1 hour), and state difference threshold (3%) as the inner domain are used. The same steps are used to extract the operation and maintenance state transition path of the outer domain, calculate the window state value, determine the transition, arrange the path, and extract the same four features. The outer domain trajectory feature sequence is formed in the order of time windows. The sequence length is consistent with the inner domain, ensuring that the two sequences can be compared element by element.

[0035] Step 332, based on the inner domain trajectory feature sequence and the outer domain trajectory feature sequence, the deviation index is calculated by comparing the statistical distribution characteristics of the two sequences, and the convergence index is calculated by analyzing the state change trend of the two sequences. The deviation index and the convergence index are integrated to generate a dynamic correction coefficient. Specifically, when calculating the deviation index, the statistical distribution characteristics of the inner domain and the outer domain trajectory feature sequence are compared. The state duration feature in the two sequences is calculated. The average duration of all windows in the inner domain and the average duration of all windows in the outer domain are calculated. The absolute difference between the two averages (the absolute value of the difference between the average of the outer domain and the average of the inner domain) is calculated. The state change amplitude, state transition frequency, and continuous state change direction (direction conversion to numerical value: rising = 1, falling = -1, fluctuation = 0) features are calculated in the same way. The absolute differences of the four features are added and then divided by the total number of features, which is 4, to obtain the average difference rate. The larger the value, the more obvious the deviation of the feature averages of the two sequences.

[0036] For each feature, the variances of the inner domain sequence and the outer domain sequence are calculated respectively, reflecting the dispersion degree of the feature value, the square sum of the difference between the feature value of each window and the mean value of the sequence, divided by the number of windows; the ratio of the outer domain variance to the inner domain variance is calculated, if the outer domain variance is 0, the ratio is recorded as 0; if the inner domain variance is 0 and the outer domain variance is not 0, the ratio is recorded as the maximum value 5; the variance ratio of the four features is added and then divided by 4 to obtain the variance difference rate, the value closer to 1, the more similar the dispersion degree of the two sequences; the larger or smaller the value, the more significant the difference. For each feature, the feature values of the inner domain sequence are sorted from small to large, divided into 4 equal intervals, such as 0-25%, 25%-50%, 50%-75%, and 75%-100%, to determine the value range of each interval; the proportion of the number of each feature value in the outer domain sequence falling within the above 4 intervals is counted, such as 30% of the outer domain state duration feature falling within the 0-25% interval of the inner domain; the proportion of each interval in the outer domain and the proportion of the corresponding interval in the inner domain are calculated, theoretically the absolute difference is 25%, the difference of the 4 intervals is added to obtain the shape difference of the feature; the shape differences of the four features are added and then divided by 4 to obtain the shape difference degree, the larger the value, the more obvious the distribution shape difference of the two sequences; the mean difference rate, the variance difference rate and the shape difference degree are added with a weight of 1:1:1 to obtain the total deviation degree index, the larger the value, the more significant the statistical distribution difference of the two sequences.

[0037] When calculating the convergence index, the trend of state change of the two sequences is analyzed, and for each feature, a trend line is drawn for the inner domain and outer domain sequences respectively with the time window number as the horizontal axis (1, 2, 3, …) and the feature value as the vertical axis, which is formed by connecting the midpoints of each window feature value; the overall direction of the trend line is calculated, if the average value of the feature values of the last 5 windows is greater than that of the first 5 windows, it is recorded as an upward trend (+1); if it is smaller, it is recorded as a downward trend (-1); if it is close, it is recorded as a stable trend (0); the number of trends with the same direction of the four features in the two sequences is counted, and the consistency ratio of the trend direction is obtained by dividing by 4, the higher the value, the more consistent the direction; for the features with the same trend direction, the trend change rate is calculated, the inner domain rate = (the last window feature value - the first window feature value) ÷ the total number of windows; the outer domain rate is calculated in the same way; the ratio of the outer domain rate to the inner domain rate is calculated, if the ratio is within the range of 0.9-1.1, it is considered that the rate is close, recorded as 1; otherwise, the score is deducted according to the deviation, such as 0.8 for the ratio of 0.8 and 0.8 for the ratio of 1.2; for all features with the same direction, the average value of the above scores is taken to obtain the trend rate proximity ratio, the higher the value, the closer the rate; the total change of the inner domain sequence from the first window to the last window is calculated, the sum of the absolute values of (the last value - the initial value) of each feature; the total change of the outer domain sequence is calculated in the same way, the ratio of the total change of the outer domain to the total change of the inner domain is calculated, if the ratio is within the range of 0.8-1.2, it is recorded as 1, otherwise it is adjusted according to the deviation, such as 0.6 for the ratio of 0.6, to obtain the final trend convergence ratio, the trend direction consistency ratio, the final trend convergence ratio are added with weights of 3:3:2 to obtain the convergence index, the closer the value to 1, the more similar the change trend of the two sequences, the lower the value, the greater the trend difference.

[0038] The deviation index is normalized by dividing the current deviation index by the maximum deviation, and the maximum deviation value obtained by comparing all inner domain and outer domain sequences from the operation and maintenance database, to obtain a deviation normalization value in the range of 0-1, 1 indicates the maximum deviation; the convergence index is processed in reverse, 1-convergence index is calculated, the larger the value, the more divergent the trend; the dynamic correction coefficient is integrated by weighted multiplication, dynamic correction coefficient = deviation normalization value × (1-convergence index) × 1.0 (weighting coefficient), the final result is a value in the range of 0-1, the larger the value, the more significant the difference between the outer domain sample and the inner domain sample, the greater the degree of correction required.

[0039] By defining the boundary of the inner domain and the outer domain, selecting continuous sample sets with the same length and consistent timestamp intervals, the comparability of the inner domain and the outer domain samples is ensured; the requirement for sample integrity ensures that the samples can truly reflect the typical state of the inner domain and the outer domain, avoids feature deviation caused by improper sample selection, extracts the state transition path by matching the time window with the state change period of the operation and maintenance object and combining the sliding step, which can capture the dynamic change of the state and filter out small noise, so that the state transition path is more in line with the actual operation and maintenance state; the trajectory features of the inner domain and the outer domain are extracted synchronously to form sequences, which ensures that the feature dimensions of the two sequences are consistent, improves the representativeness of the feature sequence and the reliability of the analysis, and comprehensively quantifies the distribution difference of the trajectory features of the inner domain and the outer domain by calculating the deviation index in multiple dimensions, avoiding deviation misjudgment caused by single-dimensional analysis.

[0040] In a preferred embodiment of the present application, step 4, using a dynamic correction coefficient, performing feature extraction and correlation analysis on the multi-dimensional time series data set, identifying the steady-state behavior pattern and abnormal boundary of the operation and maintenance environment, and generating a dynamic threshold baseline, can include: In the embodiment of the present application, step 440, the fluctuation amplitudes of each dimension in the multi-dimensional time series data set are weighted and adjusted using a dynamic correction coefficient to generate a calibrated multi-dimensional time series data set, which specifically includes: first, determine each dimension in the multi-dimensional time series data set, such as the distance influence coefficient of each node, the time synchronization degree, the monitoring value, etc., and extract the original fluctuation amplitude of each dimension. For each dimension, divide the data into fixed time windows, such as 1 hour, calculate the difference between the maximum and minimum values of the data in each window as the original fluctuation amplitude of the window, and traverse all time windows to obtain the original fluctuation amplitude sequence of each dimension.

[0041] The dynamic correction coefficient generated in step 332 has a value range of 0-1, and each dimension is assigned a correction weight according to its importance, with important dimensions having a high weight, such as a benchmark node-related dimension weight of 1.2 and a non-benchmark node weight of 0.8. The fluctuation amplitude of each dimension is weighted and adjusted by multiplying the original fluctuation amplitude of each time window of the dimension by (1+dynamic correction coefficient x correction weight) to obtain the calibrated fluctuation amplitude. If the dynamic correction coefficient is large, it indicates that the difference between the outer domain and the inner domain is significant, and the fluctuation amplitude needs to be adjusted more significantly. The original data of each dimension is adjusted according to the calibrated fluctuation amplitude to form a multi-dimensional time series data set that has been calibrated for each dimension. If the original data exceeds the range corresponding to the calibrated fluctuation amplitude, it is proportionally shrunk to within the range, such as adjusting the original value to the upper limit value if it exceeds the upper limit by 10%. Data that does not exceed the range remains unchanged, and a multi-dimensional time series data set that has been calibrated for each dimension is finally formed.

[0042] Step 441, principal component analysis is performed on the calibrated multi-dimensional time series data set to extract key feature dimensions reflecting the core state of the operation and maintenance environment, and a key feature dimension set is formed, specifically including: for the calibrated multi-dimensional time series data set, first calculate the change contribution of each dimension, count the sum of the fluctuation amplitude of each dimension after calibration in all time windows, divide by the sum of the fluctuation amplitudes of all dimensions, and obtain the contribution proportion of the dimension to the overall data change. The higher the contribution, the greater the impact of the fluctuation of the dimension on the operation and maintenance state; sort the dimensions by contribution from high to low, and accumulate the contribution in turn. When the cumulative contribution reaches a preset threshold, such as 85%, indicating that it covers most of the data changes, stop accumulating. At this time, the selected dimension is the core feature dimension; correlation screening is performed on the core feature dimension, and the numerical consistency of any two core dimensions, such as the proportion of time windows that rise or fall at the same time, is calculated. If the consistency of the two dimensions exceeds 70%, the dimension with higher contribution is retained and the other is removed. The finally retained dimensions are combined to form a key feature dimension set, such as containing 3-5 dimensions including the reference node time synchronization degree, the core node distance influence coefficient, etc.

[0043] Step 442, according to the time series data in the key feature dimension set, the time lag correlation measure value between the feature dimensions is calculated, and a time lag correlation relationship structure is constructed, specifically including: selecting any two dimensions from the key feature dimension set, such as dimension A and dimension B, extracting their time series data, and each timestamp corresponds to a numerical value; set the time lag range, determine the maximum number of time stamps of the time lag according to the state response speed of the operation and maintenance object, such as setting 5 time stamps for a fast responding device and 10 time stamps for a slow responding device, and each time stamp interval is 10 minutes. For the time series of dimension A, shift 1, 2, …, to the maximum time lag timestamp in turn to obtain multiple shifted sequences, such as shifting 2 time stamps, i.e. the t time data of dimension A corresponds to the t+2 time data of dimension B; for each offset, calculate the consistency of the change trend of the shifted dimension A and the original dimension B, count the number of times that both of them rise, fall, or rise and fall in the same time window, and use (the number of times of rising + the number of times of falling) ÷ the total number of windows as the similarity ratio. The similarity ratio is the correlation measure value at this time lag, ranging from 0 to 1, and the higher the value, the more related the two dimensions at this time lag. Repeat the above operation for all dimension combinations, record the highest correlation measure value of each combination at different time lags, i.e. the measure value corresponding to the most relevant time lag, and construct a time lag correlation relationship structure. Take the key feature dimension as the row and column to form a matrix, and the element in the ith row and jth column of the matrix is the highest correlation measure value of dimension i and dimension j and the corresponding time lag.

[0044] At step 443, based on the time lag correlation relationship structure and the calibrated multi-dimensional time series data set, the stable state behavior rule of the operation and maintenance environment is identified through data aggregation characteristics, and a probability distribution cluster representation is obtained. Specifically, based on the time lag correlation relationship structure, dimension combinations with correlation measurement values higher than a preset threshold are screened out, and the calibrated data of these dimensions are taken as analysis objects. The data aggregation characteristics are identified. For the screened dimension data, consecutive time segments are divided in chronological order, such as one segment per 24 hours. The average value of each dimension in each segment is calculated, and these average values are taken as feature points of the segment. The feature points of different segments are compared. If the numerical difference of multiple feature points is less than 5%, these segments are classified into the same aggregation group. The stable state behavior rule is judged. For each aggregation group, the number of time segments it contains, the proportion and duration of the total segment number, and the total duration of continuous occurrence are counted. If the proportion is more than 30% and the duration is more than 72 hours, the feature point range corresponding to the aggregation group is a stable state behavior rule, and a probability distribution cluster representation is generated. For each stable state behavior rule, the numerical distribution of the feature points in each dimension is counted, the numerical range and occurrence frequency of each dimension are calculated, and these distribution characteristics are combined according to the dimension combination to form a probability distribution cluster. All the distribution clusters corresponding to the stable state behavior rules jointly constitute the probability distribution cluster representation.

[0045] At step 444, according to the statistical characteristics of the probability distribution cluster representation, the dynamic confidence interval range of the abnormal boundary under the preset confidence level is determined, and a dynamic threshold baseline is generated. Specifically, the preset confidence level is determined. For each distribution cluster in the probability distribution cluster representation, the statistical characteristics of each dimension are calculated, including the average value, the standard deviation, the maximum value, and the minimum value. The dynamic confidence interval of a single distribution cluster is calculated. Taking the average value of each dimension as the center, the range is expanded by 95% confidence level. Taking the average value ± 2 × standard deviation as the interval range of the dimension. Under 99% confidence level, the average value ± 3 × standard deviation is taken. The interval ranges of all dimensions are combined to form the confidence interval corresponding to the distribution cluster. The confidence intervals of all distribution clusters are integrated. The intersection of the confidence intervals of each dimension in all distribution clusters is taken as the minimum range covering all cluster intervals, which is taken as the overall dynamic confidence interval of the dimension. The upper and lower limits of the overall dynamic confidence interval are determined as the abnormal boundary. Data exceeding the upper limit or below the lower limit is considered as potential abnormality, and a dynamic threshold baseline is generated. The abnormal boundaries of each dimension are arranged in chronological order to form a threshold baseline that dynamically adjusts with time and operating conditions.

[0046] The data calibration can reflect the difference degree between the inner domain and the outer domain by dynamically correcting the coefficient to weight and adjust the fluctuation amplitude of each dimension, and the interference of abnormal fluctuation in the original data on subsequent analysis is avoided; the weight is allocated according to the importance of the dimension, so that the calibration of the core dimension is more accurate, the reliability of the data set is improved, the redundant dimensions with little influence on the operation and maintenance state are removed by calculating the change contribution degree and screening the dimensions with high cumulative contribution degree, and the analysis complexity is reduced; the highly correlated dimensions are removed through correlation screening, the information is avoided from being repeated, the key feature dimension set can more accurately reflect the core state of the operation and maintenance environment, the dynamic correlation between the feature dimensions is captured by calculating the correlation measurement value under different time lags, and the limitation of only analyzing synchronous correlation is broken through; the time-lag correlation relationship clearly presents the correlation strength and time difference between the dimensions, provides a key correlation basis for identifying the steady-state behavior rule, and makes the analysis more in line with the dynamic characteristics of the operation and maintenance environment; the steady-state behavior rule is identified through data aggregation characteristics, and the stable state that repeatedly appears in the operation and maintenance environment can be accurately captured; the probability distribution cluster quantifies the numerical distribution and occurrence frequency of the steady state in each dimension, so that the steady state rule is transformed from qualitative description to quantitative description, the dynamic confidence interval is calculated based on the statistical characteristics of the probability distribution cluster and the preset confidence level, and the probability that the normal fluctuation is misjudged as an abnormality is reduced; the dynamic threshold baseline is adjusted with the running condition and time, and can adapt to the change of the operation and maintenance environment.

[0047] In a preferred embodiment of the present application, step 5, matching and verifying the dynamic threshold baseline with the real-time stream data, combining the periodic phase alignment mechanism, generating the abnormal type and the current operation and maintenance object bottleneck positioning result, can include: In the embodiment of the present application, step 550, according to the business cycle characteristics of the target operation and maintenance object, the real-time stream data is divided into continuous time segment units to form a phase-aligned time window sequence, specifically including: first, analyze the business cycle characteristics of the target operation and maintenance object, collect at least 3 complete cycle operation data, such as the daily cycle of an e-commerce platform and the weekly cycle of a factory production line, and count the rules that repeatedly appear in the data, such as 9:00-12:00 every day as the peak period, 12:00-14:00 as the flat period, 14:00-20:00 as the sub-peak period, and 20:00-9:00 the next day as the trough period, determine the cycle length, such as 24 hours and the time division of each stage, 4 stages, each stage 3-12 hours; according to the cycle characteristics, the real-time stream data is divided into continuous time segment units in time sequence, the length of each unit is consistent with the length of the stage in the cycle, such as 3 hours for the peak period, corresponding to 9:00-12:00 every day, and the start time of each unit must be strictly aligned with the start time of the corresponding stage in the cycle when dividing, such as the 9:00-12:00 unit of today aligns with the 9:00-12:00 unit of yesterday, to avoid time offset.

[0048] Form a phase-aligned time window sequence, arrange all segmented time segment units in chronological order, each unit as a window in the sequence, the phase of the window determined by its position in the cycle, such as all 9:00-12:00 windows belong to the peak phase, and finally form a phase-aligned time window sequence.

[0049] Step 551, based on the time window sequence, perform time dimension alignment operation on real-time stream data in each time window unit and dynamic threshold baseline, and calculate the matching deviation value in each time window unit, specifically including: based on the time window sequence generated in step 550, determine the time axis parameters of each time window unit, including the start time, end time and duration of the window, such as 9:00-12:00 window, starting at 9:00, ending at 12:00, lasting 3 hours, and the phase label of the window in the business cycle, perform time dimension alignment operation on real-time stream data in each window unit, extract its timestamp, and match it with the timestamp of the corresponding phase window in the dynamic threshold baseline, to ensure that each time point of real-time data is aligned with the same relative time point in the baseline, if there is a timestamp missing in real-time data, then use the average value of the adjacent time points to fill in, to ensure the integrity of the alignment.

[0050] Calculate the matching deviation value in each time window unit, for each aligned time point, calculate the deviation of real-time stream data value from the threshold range of the corresponding time point in the dynamic threshold baseline, if the real-time value is within the baseline threshold range, the deviation is 0; if the real-time value exceeds the range, the deviation value = real-time value-upper limit value (0.05); if it is lower than the lower limit, the deviation value = lower limit value-real-time value (0.05), take the average value of all time point deviation values in the window (total deviation value ÷ time point number) as the matching deviation value of the window unit.

[0051] Step 552, according to the matching deviation value, the distribution characteristics of the deviation value on the time axis are analyzed, and the abnormal type is identified, specifically including: collecting all the matching deviation values of the time window units, arranging them in time axis order, analyzing the distribution characteristics of the deviation values, counting the number of windows whose deviation values are greater than the preset slight threshold, and calculating the proportion of the number of windows in the total number of windows; At the same time, record the maximum deviation value and the average deviation value, the deviation duration distribution, observe the number of windows with continuous deviation >0.05, judge whether it is a short-term burst or a long-term duration, calculate the difference value of adjacent window deviation values, if the difference value of three continuous values is positive, it is an increasing trend; If the difference value of three continuous values is negative, it is a decreasing trend; If the positive and negative values are alternated, it is a fluctuation trend; According to the distribution characteristics, the abnormal type is identified, if the deviation amplitude is large (>0.1), the duration is short (≤2 windows), and the change trend is sudden increase, it is determined as a sudden abnormality, if the deviation amplitude gradually increases, the duration is long (>5 windows), and the change trend is increasing, it is determined as a gradual abnormality, if the deviation only appears in a specific phase window and repeats periodically, it is determined as a periodic abnormality; If the deviation amplitude fluctuates greatly, there is no obvious trend, and the duration is indefinite, it is determined as a fluctuation type abnormality.

[0052] Step 553, according to the time dimension alignment result, the correlation imbalance node of resource consumption index and performance index is located, and the abnormal type and the correlation imbalance node are integrated to generate the operation and maintenance object bottleneck positioning result, specifically including: according to the time dimension alignment result of step 551, the real-time data and baseline data of resource consumption index and performance index in each time window unit are extracted, the correlation imbalance node is located, and the correlation degree of resource consumption index and performance index is calculated. Under normal circumstances, the two should be positively correlated. For each node, the number of times that the resource consumption increases but the performance index does not improve or the resource consumption is normal but the performance index is abnormal is counted, if the proportion of the number of times in the total number of windows is >30%, the node is a correlation imbalance node; Integrate the abnormal type and the correlation imbalance node, match the abnormal type identified in step 552 with the correlation imbalance node located, analyze whether the abnormality is caused by node imbalance, combine the physical location and function role of the node, determine the specific location of the bottleneck, and finally generate the operation and maintenance object bottleneck positioning result.

[0053] By analyzing the business cycle characteristics and segmenting the phase-aligned time window, the analysis of real-time stream data can be accurately matched with the periodicity of the operation and maintenance object, avoiding misjudgment caused by time phase misalignment, ensuring that the matching check is carried out in the same cycle stage, and ensuring one-to-one correspondence between real-time stream data and dynamic threshold baseline at each time point through timestamp accurate alignment and missing value filling. The calculation of matching deviation value comprehensively considers the deviation degree of each time point and takes the average value, which can fully reflect the overall matching situation in the window, avoid the interference of single time point fluctuation on the result, make the deviation quantization more objective, provide reliable numerical basis for abnormal type identification, and accurately describe the dynamic characteristics of the abnormality by analyzing the deviation distribution characteristics from three dimensions of amplitude, duration and change trend, breaking through the limitation of judging abnormality only by deviation size. The feature-based abnormal type classification makes the abnormality nature clearer, improves the efficiency and accuracy of abnormal response, and locates the imbalance node by associating resource consumption and performance indicators, which can directly lock the source of the abnormality and avoid blind troubleshooting, so that the bottleneck positioning result has both qualitative and quantitative information.

[0054] In a preferred embodiment of the application, based on the abnormal classification and the current operation and maintenance object bottleneck positioning result, the resource constraint condition and the service level agreement are integrated to generate and execute an operation and maintenance strategy set including resource elasticity scheduling rules, fault switching links and service degradation preplans, which can include: In the embodiment of the application, step 660, according to the abnormal type identification result, selects the corresponding basic strategy framework in the pre-stored strategy template library, specifically including: first, the specific content of the abnormal type identification result is determined, including the category of the abnormality, the severity of the abnormality, the influence range of the abnormality and other key information, the pre-stored strategy template library is stored according to the dimensions of abnormal type, severity, influence range, etc., and each classification corresponds to at least one basic strategy framework, the framework includes the core structure of the strategy, such as resource scheduling module, fault handling module, degradation trigger condition and other general components, but does not contain specific parameters; compare the abnormal type identification result with the classification label of the template library: if the abnormal type is network congestion and the influence range is local cluster, search the classification with the label of network congestion-local cluster in the template library, and select the basic strategy framework with the highest matching degree in the classification as the basis of the current strategy.

[0055] Step 661, inject the key node information in the operation and maintenance object bottleneck coordinate positioning report into the basic policy framework, convert it into device location parameters, generate a topology-aware policy draft, which specifically includes: the key node information in the operation and maintenance object bottleneck coordinate positioning report includes the physical location of the bottleneck node, the network topology location, the device identifier, the role of the node in the service link, etc.; the basic policy framework has preset parameters placeholders related to device location, and the key node information is corresponded to the placeholders one by one, after all the placeholder replacements are completed, the basic policy framework is converted from a general structure to a policy draft containing specific device location information, i.e. a topology-aware policy draft.

[0056] Step 662, based on the resource total amount limit in the resource constraint condition, calculate the resource allocation proportion of each scheduling unit in the topology-aware policy draft to generate an elastic scheduling proportion table, which specifically includes: the resource total amount limit in the resource constraint condition includes specific values such as the total number of available CPU cores, the total memory capacity, the total storage bandwidth, the total network throughput, and the number of schedulable servers, and the scheduling unit in the topology-aware policy draft refers to an independent service or device group that needs resource support, each scheduling unit has its current resource demand, for each resource type, the proportion of the demand of a single scheduling unit to the total amount is calculated, for example, in CPU resources, scheduling unit A needs 100 cores and the total amount is 500 cores, so the CPU proportion is 100 / 500=20%; scheduling unit B needs 200 cores, the proportion is 40%, and so on, if the total demand of all scheduling units exceeds the total amount limit, such as total demand 600 cores>total amount 500 cores, then the demand proportion of each unit is used for equal proportion reduction, the final CPU allocation of scheduling unit A is 100x(500 / 600)≈83 cores, the proportion adjustment is 83 / 500≈16.6%; scheduling unit B is 200x(500 / 600)≈167 cores, the proportion is approximately 33.4%, and so on, the resource allocation proportion of all scheduling units is arranged into a table, i.e. an elastic scheduling proportion table, which clearly shows the allocation proportion and specific value of each unit to CPU, memory, network, etc.

[0057] Step 663, according to the service level agreement in the business processing order, the service unit in the front of the order in the elastic scheduling proportion table is configured with the main execution path and the standby switching path, and the main and standby path configuration results are obtained, which specifically includes: the service processing order in the service level agreement (SLA) is sorted according to the business priority, such as core transaction business > ordinary query business > log backup business, the sorting basis includes business influence range, user importance, interruption loss, etc., the smaller the priority value, the earlier the sorting, such as core transaction business priority 1, ordinary query 2; according to the business processing order of SLA, the business units in the elastic scheduling proportion table are sorted, such as core transaction unit ranked first, ordinary query unit ranked second; select the service unit in the front of the order, usually the first 30% or priority ≤3, configure the main execution path for it, select the path with the shortest link, the lowest load and the lowest failure rate by analyzing the load rate, response delay, historical stability and other indicators of the nodes in the network topology, determine each node, port and forwarding rule in the path, configure the standby switching path for the same business unit, select the path which is physically isolated from the main path and has suboptimal performance, ensure that the main and standby paths have no common single point of failure node, record the specific node sequence of the main path, the condition of triggering switching, the activation mode of the standby path, etc., form the main and standby path configuration results.

[0058] Step 664, based on the main and standby path configuration results, match the pre-stored service degradation rule library, bind the matched degradation rules with the elastic scheduling proportion table, and obtain the proportion table with bound degradation rules, which specifically includes: the pre-stored service degradation rule library contains multiple rules, each rule contains trigger condition, degradation measure, applicable business type and other information, extract the key features in the main and standby path configuration results, such as the maximum load threshold of the main and standby path, the type of business unit, the influence range of path failure, etc., match the above features with the rule trigger condition in the service degradation rule library, if the trigger condition of the rule is completely consistent with the path feature, select the rule as the matching result, bind the matched degradation rule with the corresponding business unit in the elastic scheduling proportion table, add the trigger condition, specific measures and execution priority of the degradation rule under the entry of the business unit in the proportion table, such as when the main and standby path load > 90%, the CPU allocation proportion is temporarily reduced to 80% of the original proportion, the non-core query interface is closed, after all matching and binding are completed, the proportion table with bound degradation rules is formed, which contains resource allocation proportion and corresponding degradation rules.

[0059] Step 665, the master and backup path configuration results and the proportion table of the binding degradation rule are integrated to generate a set of operation and maintenance strategies that can be distributed and executed, specifically including: checking the consistency of the master and backup path configuration results and the proportion table of the binding degradation rule, for example, confirming whether the resource requirements of the business unit in the path configuration match the allocation proportion in the proportion table, such as the core transaction unit in the primary path requiring 100 core CPUs, and the allocation proportion corresponding to 100 cores in the proportion table, no conflict; confirming whether the resource adjustment measures in the degradation rule are compatible with the path switching conditions, such as the CPU allocation being reduced after degradation, which will not cause abnormal path load, integrating the two types of information, classifying the node sequence, switching condition of the master and backup path, and the resource allocation proportion in the proportion table, and the degradation rule according to the business unit, such as the information of the core transaction unit including CPU allocation proportion 20%, primary path A, backup path B, and degradation rule X; converting into an executable format, according to the interface requirements of the execution subject, converting the integrated information into specific instructions, summarizing all the instructions, and forming an operation and maintenance strategy set sorted by execution order, such as first executing resource allocation, then configuring the path, and finally loading the degradation rule, to ensure that each strategy can be distributed and executed by different devices.

[0060] By matching the pre-stored template based on the exception type, the redundant work of building the strategy from zero is avoided, and the strategy generation time is shortened; at the same time, the pre-stored template is verified to ensure the rationality and applicability of the basic strategy framework, reduce the risk of operation and maintenance failure caused by defects in the strategy structure, inject the specific location information of the bottleneck node into the strategy framework, and make the strategy from generalization to topology awareness, ensure that the strategy can accurately point to the device and associated nodes where the bottleneck actually exists, avoid invalid operations caused by ambiguous location during strategy execution, improve the operation and maintenance pertinence, calculate the allocation proportion based on the total resource limit, ensure that resource scheduling is performed within the actual available range, and avoid resource conflicts or waste caused by excessive allocation; the elastic scheduling proportion table clearly defines the resource proportion of each unit, and improves the rationality of resource utilization.

[0061] In a preferred embodiment of the present application, the step 7, the state evolution data obtained after the execution of the operation and maintenance strategy set is used to adjust the dynamic threshold baseline, the determination interval configuration and the strategy generation rule to form a set of adaptive loop control parameters, which can include: In the embodiment of the present application, in step 770, the system state evolution data generated after the execution of the operation and maintenance policy instruction set is collected, and the variation rate of the key performance indicator in the policy action period is extracted, specifically including: determining the execution time point of each instruction in the operation and maintenance policy instruction set, taking the moment when the first instruction starts to execute as the starting time of data collection, and taking the moment when the last instruction is completely executed as the ending time, starting the high-frequency data collection mechanism during this period, and the collection frequency is set according to the characteristics, for example, for the server response time and other fast-changing indicators, data is collected every 0.5 seconds; for the memory occupancy and other relatively stable indicators, data is collected every 5 seconds, and the collection content includes not only the specific values of the key performance indicators, but also the accurate time of each collection, the execution state of the policy instruction and the system environment parameters; the pre-set key performance indicator data column is selected from the collected raw data, the null values caused by collection failure, the abnormal values exceeding the reasonable physical range and the repeated records are removed, and for a small amount of missing intermediate data, the average value of the adjacent two effective data is used for filling to ensure the continuity of the data sequence.

[0062] The cleaned key performance indicator data is sorted according to the collection time stamp from small to large to form a complete time sequence, the entire policy action period is divided into multiple continuous sub-periods according to the execution node of the policy instruction, such as the instruction 1 execution period, the interval period from the completion of instruction 1 to the start of instruction 2, the instruction 2 execution period, etc., and the indicator variation in each sub-period is calculated respectively, in each sub-period, the starting time indicator value and the ending time indicator value of the period are selected, the difference between the two is calculated, the ending time value is subtracted from the starting time value, if the difference is positive, it means that the indicator shows an upward trend; if it is negative, it means that the indicator shows a downward trend, the difference is divided by the duration of the sub-period to obtain the average variation rate in the sub-period, and the average variation rates of all sub-periods are weighted to finally obtain the comprehensive variation rate of the key performance indicator in the entire policy action period.

[0063] Step 771, adjust the response sensitivity of the dynamic threshold baseline according to the change rate of the key performance indicators, generate a sensitivity-updated dynamic threshold baseline, specifically including: the response sensitivity of the dynamic threshold baseline is initially set to a reference value, such as 1.0, and three-level rate threshold intervals are pre-set, a high-rate interval, such as greater than 5 units / second, a medium-rate interval, such as 1-5 units / second, and a low-rate interval, such as less than 1 unit / second, different intervals correspond to different sensitivity adjustment amplitudes, and at the same time, the upper and lower limits of the sensitivity are specified, such as a minimum of 0.3 and a maximum of 2.0, to avoid the adjusted sensitivity from exceeding a reasonable range, when there are multiple key performance indicators, the change rate of each indicator is calculated respectively, and then a weighted sum is performed according to the importance of each indicator, such as CPU usage weight 0.3, memory occupancy rate weight 0.2, and response time weight 0.5, to obtain the overall comprehensive change rate of the system; if the comprehensive change rate is in the high-rate interval, it means that the system state changes dramatically, and the response sensitivity needs to be increased by 30% based on the current value, such as if the current sensitivity is 1.0, then the adjusted sensitivity is 1.3; if it is in the low-rate interval, it means that the system state is stable, and the response sensitivity is reduced by 20%, such as if the current sensitivity is 1.0, then the adjusted sensitivity is 0.8; if it is in the medium-rate interval, the sensitivity remains unchanged, if the adjusted sensitivity exceeds the pre-set upper and lower limits, the closest boundary value is taken, such as if the calculated sensitivity is 2.2, the actual value is 2.0, the adjusted response sensitivity is substituted into the generation logic of the dynamic threshold baseline, for example, in the original baseline calculation logic, the higher the sensitivity, the greater the change amplitude of the baseline with the latest 3 data points; the lower the sensitivity, the more the baseline refers to the average value of the past 10 data points, the threshold baseline value of each time point is recalculated according to the new sensitivity, and a complete sensitivity-updated dynamic threshold baseline sequence is formed.

[0064] Step 772, analyze the abnormal processing process using the sensitivity-updated dynamic threshold baseline, and correct the boundary tolerance range of the decision interval configuration, specifically including: collecting all recorded abnormal events during policy execution, including the trigger time of each abnormality, the specific value of the key performance indicator at the trigger time, the abnormal level based on the original decision interval configuration, the handling measures taken, the start and end time of the handling, the recovery of the indicator after handling, and other detailed information, for each abnormal event, compare the indicator value at the trigger time with the sensitivity-updated dynamic threshold baseline to determine whether the event is still abnormal under the new baseline standard, for example, the normal interval of the original baseline is [80, 120], and the indicator value of a certain event is 125, which is judged as abnormal; the normal interval of the new baseline is [75, 125], so the event is normal under the new baseline standard, which is a misjudgment, and the number of misjudgment events and missed events and the corresponding indicator value deviation are also counted.

[0065] The maximum deviation value of the index value of the misjudgment event beyond the original boundary, such as the original upper boundary 120, the index value 125 of a certain misjudgment event, the deviation value is 5, and the minimum deviation value of the index value of the missed judgment event and the original boundary, such as the original upper boundary 120, the index value 122 of a certain missed judgment event, the deviation value is 2, if the number of misjudgment events accounts for more than 10%, it means that the original boundary tolerance range is too small; if the number of missed judgment events accounts for more than 10%, it means that the original boundary tolerance range is too large, when the tolerance range needs to be expanded, the above maximum deviation value is taken as the basis, the original boundary is expanded outward by 80% of the deviation value, such as the original upper boundary 120, the maximum deviation 5, then the new upper boundary is 120+5x80%=124; when the tolerance range needs to be reduced, the minimum deviation value is taken as the basis, the original boundary is contracted inward by 60% of the deviation value, such as the original upper boundary 120, the minimum deviation 2, then the new upper boundary is 120-2x60%=118.8, for the case of misjudgment and missed judgment, the comprehensive influence of the two deviations is calculated, and the problem type with higher proportion is preferentially modified.

[0066] Step 773, based on the modified boundary tolerance range, the strategy execution performance value is calculated, the action level parameter of the strategy generation rule is adjusted, specifically including: selecting the time of system returning to normal state after strategy execution, abnormal event recurrence rate, system resource consumption change amount and the like as the index for evaluating the strategy execution performance, according to the modified boundary tolerance range, the judgment standard of normal and abnormal state is redefined, for each evaluation index, the actual data after strategy execution is compared with the ideal data based on the new judgment standard, for example, the shorter the system returns to normal state, the higher the performance, the reciprocal of the ratio of the actual recovery time to the ideal recovery time is taken as the performance score of the index; the lower the abnormal event recurrence rate, the higher the performance, (1-actual recurrence rate) is taken as the performance score of the index, the performance scores of all evaluation indexes are weighted and averaged to obtain the strategy execution performance value, the current action level parameter of each rule in the strategy generation rule is obtained, which is used to represent the priority of the rule in generating the operation and maintenance strategy, the calculated strategy execution performance value is compared with the preset performance threshold value, if the performance value is higher than the high performance threshold value, it means that the corresponding strategy generation rule is better, the action level parameter is increased (such as increasing 1 level based on the original parameter value); if the performance value is lower than the low performance threshold value, it means that the rule effect is poor, the action level parameter is reduced (such as reducing 1 level based on the original parameter value); if the performance value is within the normal performance threshold value, the action level parameter remains unchanged.

[0067] Step 774, integrating the dynamic threshold baseline, the boundary tolerance range and the action level parameter into the adaptive loop control parameter set, specifically including: respectively performing validity check on the sensitivity updated dynamic threshold baseline obtained in step 771, the corrected boundary tolerance range obtained in step 772 and the adjusted action level parameter obtained in step 773, checking whether they are within the preset reasonable value range, whether they meet the basic logic of system operation, analyzing the relevance among the three parameters, ensuring that the dynamic threshold baseline and the boundary tolerance range remain consistent in the judgment logic, and the adjustment of the action level parameter can be adapted to the changes of the previous two parameters, avoiding contradictions among the parameters, and the dynamic threshold baseline, the boundary tolerance range and the action level parameter that pass the validity check and are associated and matched are summarized and arranged, combined together according to the preset format and structure to form a complete parameter set, that is, the adaptive loop control parameter set.

[0068] By collecting system state evolution data and extracting the change rate of key performance indicators, the change of the system under the action of the operation and maintenance strategy can be more accurately captured, the perception of the system to its own state is more accurate, the response sensitivity of the dynamic threshold baseline is adjusted according to the change rate of the key performance indicators, the baseline can be flexibly adjusted according to the actual change speed of the system, timely response is made when the index changes rapidly, and stability is maintained when the index changes slowly, the adaptability of the baseline to different system states is improved, the boundary tolerance range of the determination interval configuration is corrected by using the updated dynamic threshold baseline, the misjudgment and omission in abnormal determination can be reduced, the abnormal determination is more accurate, the strategy execution efficiency value is calculated based on the corrected boundary tolerance range, and the action level parameter of the strategy generation rule is adjusted accordingly, which can make the rules with good action effect have higher priority in strategy generation, so that more reasonable and effective operation and maintenance strategies are generated.

[0069] As shown in Figure 2 The embodiment of the application also provides an adaptive adjustment internet operation and maintenance strategy generation system, which comprises: A data collection module is configured to collect running state data of a target operation and maintenance environment dynamically through distributed sensors and edge nodes, and generate a multi-dimensional time series data set; An association analysis module is configured to perform spatio-temporal association analysis on the multi-dimensional time series data set, determine a reference data node, construct a two-dimensional association structure, and generate a dynamic determination interval; A correction coefficient module is configured to select a monitoring sample set in the inner domain and the outer domain of the dynamic determination interval, generate a trajectory feature sequence according to the time evolution relationship of the sample set, and calculate a dynamic correction coefficient based on the trajectory feature sequence; A threshold generation module is configured to perform feature extraction and correlation analysis on the multi-dimensional time series data set by using the dynamic correction coefficient, identify a steady-state behavior pattern and an abnormal boundary of the operation and maintenance environment, and generate a dynamic threshold baseline. An abnormal bottleneck module is configured to match and verify a dynamic threshold baseline with real-time flow data, generate an abnormal type and a current operation and maintenance object bottleneck positioning result in combination with a period phase alignment mechanism; An operation and maintenance strategy module is configured to generate and execute an operation and maintenance strategy set including resource elastic scheduling rules, fault switching links and service degradation preplans based on the abnormal classification and the current operation and maintenance object bottleneck positioning result, integrate resource constraint conditions and service level agreements; An adaptive adjustment module is configured to adjust a dynamic threshold baseline, a decision interval configuration and a strategy generation rule according to state evolution data obtained after execution of the operation and maintenance strategy set, and form an adaptive loop control parameter set.

[0070] It should be noted that the system corresponds to the above method, and all implementation manners in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.

[0071] Embodiments of the application also provide a computing device, comprising a processor and a memory storing a computer program, wherein the computer program is executed by the processor to perform the method described above. All implementation manners in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.

[0072] Embodiments of the application also provide a computer readable storage medium storing instructions, wherein the instructions are executed on a computer to make the computer perform the method described above. All implementation manners in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.

[0073] The above is the preferred embodiment of the application, and it should be noted that for those skilled in the art, without departing from the principles of the application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the application.

Claims

1. A method for generating an adaptively adjusted Internet operation and maintenance strategy, characterized in that: The method comprises: Step 1: Dynamically collect the operating status data of the target operation and maintenance environment through distributed sensors and edge nodes to generate a multi-dimensional time series data set; Step 2: Perform spatiotemporal correlation analysis on the multi-dimensional time series dataset, determine the benchmark data nodes, construct a two-dimensional correlation structure, and generate a dynamic decision interval; Step 3: Select monitoring sample sets in the inner and outer domains of the dynamic judgment interval, generate trajectory feature sequences according to the time evolution relationship of the sample sets, and calculate the dynamic correction coefficient based on the trajectory feature sequences; Step 4: Using the dynamic correction coefficient, perform feature extraction and correlation analysis on the multi-dimensional time series data set to identify the steady-state behavior patterns and abnormal boundaries of the operation and maintenance environment and generate a dynamic threshold baseline; Step 5: Match and verify the dynamic threshold baseline with the real-time streaming data, and combine the periodic phase alignment mechanism to generate the anomaly type and the bottleneck location result of the current operation and maintenance object; Step 6: Based on the anomaly classification and the bottleneck location results of the current operation and maintenance object, resource constraints and service level agreements are integrated to generate and execute an operation and maintenance policy set including resource elastic scheduling rules, failover links, and service degradation plans; Step 7: Based on the state evolution data obtained after the operation and maintenance strategy set is executed, the dynamic threshold baseline, the decision interval configuration and the strategy generation rule are adjusted to form an adaptive loop control parameter set.

2. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 1, characterized in that: Perform spatiotemporal correlation analysis on multi-dimensional time series data sets, determine benchmark data nodes, build a two-dimensional correlation structure, and generate dynamic decision intervals, including: Based on the spatial topological connection relationship and time dimension fluctuation characteristics of the multi-dimensional time series data set, the key indicator nodes that have a dominant role in representing the target operation and maintenance object are identified and determined as benchmark data nodes; Taking the benchmark data node as the correlation core, the distance influence coefficient between each node in the multi-dimensional time series dataset and the benchmark data node in the spatial topology structure is calculated; Based on the distance influence coefficient, the degree of synchronization between each node and the benchmark data node in the time dimension is evaluated; The distance influence coefficient and the degree of synchronization of the time dimension are integrated to establish a spatiotemporal coupling matrix, generating a two-dimensional correlation structure including spatial dimension correlation and temporal dimension correlation. According to the parameter distribution of the two-dimensional correlation structure, the stability boundary threshold of the correlation strength under different operating conditions of the target operation and maintenance object is calculated, and based on the stability boundary threshold, a dynamic judgment interval is generated in the parameter space of the two-dimensional correlation structure.

3. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 2, wherein: A monitoring sample set is selected from the inner and outer domains of the dynamic determination interval, and a trajectory feature sequence is generated based on the temporal evolution of the sample set. The dynamic correction coefficient is calculated based on the trajectory feature sequence, including: Extracting a first continuous time stamp monitoring sample set in an inner domain of the dynamic decision interval, and extracting a second continuous time stamp monitoring sample set in an outer domain; For the first monitoring sample set, the operation and maintenance state transition path is extracted according to the preset time window to generate the inner domain trajectory feature sequence; simultaneously, the operation and maintenance state transition path is extracted for the second monitoring sample set to generate the outer domain trajectory feature sequence; Based on the inner domain trajectory characteristic sequence and the outer domain trajectory characteristic sequence, the deviation index is calculated by comparing the statistical distribution characteristics of the two sequences, and the convergence index is calculated by analyzing the state change trends of the two sequences. The deviation index and the convergence index are integrated to generate a dynamic correction coefficient.

4. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 3, wherein: Using dynamic correction coefficients, we perform feature extraction and correlation analysis on multi-dimensional time series data sets to identify steady-state behavior patterns and abnormal boundaries in the operation and maintenance environment and generate dynamic threshold baselines, including: Use dynamic correction coefficients to perform weighted adjustment on the fluctuation amplitude of each dimension in the multi-dimensional time series data set to generate a calibrated multi-dimensional time series data set; Perform principal component analysis on the calibrated multi-dimensional time series data set to extract key feature dimensions that reflect the core status of the operation and maintenance environment, forming a key feature dimension set; Based on the time series data in the key feature dimension set, the time lag correlation measurement value between the feature dimensions is calculated and the time lag correlation relationship structure is constructed; Based on the time-lag correlation structure and the calibrated multi-dimensional time series data set, the steady-state behavior of the operation and maintenance environment is identified through the data aggregation characteristics, and the probability distribution cluster representation is obtained; According to the statistical characteristics of the probability distribution cluster representation, the dynamic confidence interval range of the abnormal boundary at the preset confidence level is determined to generate a dynamic threshold baseline.

5. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 4, characterized in that: The dynamic threshold baseline is matched and verified with real-time streaming data. Combined with the periodic phase alignment mechanism, the anomaly type and the bottleneck location results of the current operation and maintenance object are generated, including: Based on the business cycle characteristics of the target operation and maintenance object, the real-time streaming data is divided into continuous time segments to form a phase-aligned time window sequence; Based on the time window sequence, the real-time streaming data in each time window unit is aligned with the dynamic threshold baseline in the time dimension, and the matching deviation value in each time window unit is calculated; According to the matching deviation value, the distribution characteristics of the deviation value on the time axis are analyzed to identify the abnormal type; Based on the time dimension alignment results, locate the associated imbalance nodes between resource consumption indicators and performance indicators, and integrate the abnormality types and associated imbalance nodes to generate the bottleneck location results of the operation and maintenance object.

6. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 5, characterized in that: Based on anomaly classification and current O&M bottleneck location results, resource constraints and service level agreements are integrated to generate and execute an O&M strategy set that includes resource elastic scheduling rules, failover links, and service degradation plans, including: According to the abnormal type identification result, select the corresponding basic policy framework in the pre-stored policy template library; Inject key node information from the bottleneck coordinate location report of the operation and maintenance object into the basic policy framework, convert it into device location parameters, and generate a topology-aware policy draft; Based on the total resource limit in the resource constraint conditions, the resource allocation ratio of each scheduling unit in the topology-aware policy draft is calculated to generate a flexible scheduling ratio table; According to the service processing order in the service level agreement, for the service units ranked higher in the elastic scheduling ratio table, the primary execution path and the backup switching path are configured to obtain the primary and backup path configuration results; Based on the primary and backup path configuration results, the pre-stored service degradation rule library is matched, and the matched degradation rule is bound to the elastic scheduling ratio table to obtain the ratio table of bound degradation rules; The primary and backup path configuration results and the proportion table of binding degradation rules are integrated to generate a set of operation and maintenance policies that can be distributed and executed.

7. The method for generating an adaptively adjusted Internet operation and maintenance strategy according to claim 6, wherein: Based on the state evolution data obtained after the execution of the operation and maintenance strategy set, the dynamic threshold baseline, decision interval configuration and strategy generation rules are adjusted to form an adaptive loop control parameter set, including: Collect system state evolution data generated after the execution of the operation and maintenance strategy instruction set, and extract the change rate of key performance indicators during the strategy action period; Adjusting the response sensitivity of the dynamic threshold baseline according to the rate of change of the key performance indicator to generate a dynamic threshold baseline with updated sensitivity; Use the dynamic threshold baseline after sensitivity update to analyze the abnormality handling process and correct the boundary tolerance range of the judgment interval configuration; Calculate the strategy execution effectiveness value based on the revised boundary tolerance range and adjust the action level parameters of the strategy generation rule; The dynamic threshold baseline, boundary tolerance range and action level parameters are integrated into an adaptive loop control parameter set.

8. An adaptively adjusted Internet operation and maintenance strategy generation system, the system implementing the method according to any one of claims 1 to 7, characterized in that: include: The data acquisition module is used to dynamically collect the operating status data of the target operation and maintenance environment through distributed sensors and edge nodes to generate a multi-dimensional time series data set; The correlation analysis module is used to perform spatiotemporal correlation analysis on multi-dimensional time series data sets, determine the benchmark data nodes, build a two-dimensional correlation structure, and generate dynamic decision intervals; The correction coefficient module is used to select monitoring sample sets in the inner and outer domains of the dynamic judgment interval, generate trajectory feature sequences according to the time evolution relationship of the sample sets, and calculate the dynamic correction coefficient based on the trajectory feature sequences; The threshold generation module is used to perform feature extraction and correlation analysis on multi-dimensional time series data sets using dynamic correction coefficients, identify steady-state behavior patterns and abnormal boundaries of the operation and maintenance environment, and generate dynamic threshold baselines; The abnormal bottleneck module is used to match and verify the dynamic threshold baseline with real-time streaming data, and combine the periodic phase alignment mechanism to generate the abnormal type and the bottleneck location results of the current operation and maintenance object; The operation and maintenance strategy module is used to integrate resource constraints and service level agreements based on anomaly classification and bottleneck location results of the current operation and maintenance object, and generate and execute an operation and maintenance strategy set including resource elastic scheduling rules, fault switching links and service degradation plans; The adaptive adjustment module is used to adjust the dynamic threshold baseline, judgment interval configuration and strategy generation rules according to the state evolution data obtained after the execution of the operation and maintenance strategy set to form an adaptive cycle control parameter set.

9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Automobile gateway CAN upgrade security rollback decision management method and system

    CN121070408A

  • Automotive Gateway CAN Upgrade Security Rollback Decision Management Method and System

    CN121070408B

  • Intelligent diagnosis-based operation and maintenance software adaptive optimization method and system

    CN121277545A

  • Operation and maintenance data management system and method based on large model

    CN121279743A

  • Operation and maintenance management method and system for process industry equipment

    CN121544244A