Information system intelligent operation and maintenance scheduling method and system
Patent Information
- Application Number
- CN202610989792.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
问题一,现有运维调度方法缺乏对观测数据在采集、识别、执行各阶段中时间偏移的分阶段系统化建模能力,无法形成与运行状态关联的分阶段延迟修正机制,导致修正后的观测数据难以准确反映信息化系统资源的真实时序状态;
1、本发明通过将信息化系统的数据处理过程划分为采集阶段、识别阶段和执行阶段,分别对各阶段的观测延迟进行独立建模与统计分析,有效避免了将不同来源延迟混合处理所导致的不可分解性问题,实现了对各阶段延迟偏移特性范围的精细化刻画,显著提升了系统观测数据时间对齐修正的准确度。
Smart Images

Figure CN122816804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance scheduling technology, specifically to an intelligent operation and maintenance scheduling method and system for information systems. Background Technology
[0002] With the widespread application of cloud computing and big data technologies, the scale and complexity of information systems continue to grow. Intelligent operation and maintenance scheduling methods based on resource load data are widely used in automated resource scheduling, elastic scaling, and other operation and maintenance tasks. Existing technologies typically employ continuous collection of system observation data such as CPU utilization, memory usage, and network throughput to achieve real-time perception of system operating status and automatic generation of scheduling strategies. However, from the generation of system observation data to its use in scheduling decisions, it needs to go through multiple continuous processing stages, such as data acquisition, state identification, and policy execution. Each stage introduces a different degree of time delay, resulting in a discrepancy between the timestamp of the observation data and the actual time when the system state occurs. Existing technologies usually treat system observation delay as a holistic indicator for uniform correction, or completely ignore the impact of delay on scheduling decisions. They lack the ability to systematically model the stage-by-stage causes of delay, making it difficult to accurately characterize the dynamic changes in delay under different operating states. As a result, the corrected observation data still fails to reflect the true temporal state of the system. At the same time, the characteristics of delay at each stage under different operating states are significantly different. Existing technologies lack a mechanism for adaptively selecting delay correction parameters based on the operating state category. They usually use a uniform correction strategy to process observation data under all states, making it difficult to maintain high time alignment accuracy under different operating states.
[0003] In summary, the existing technology has the following technical problems when used: Problem 1: Existing operation and maintenance scheduling methods lack the ability to systematically model the time offset of observation data in each stage of collection, identification and execution, and cannot form a stage delay correction mechanism associated with the operating status, resulting in the corrected observation data failing to accurately reflect the true temporal status of information system resources. Problem 2: Existing methods fail to adaptively determine the corresponding delay correction parameters based on the differences in delay characteristics at each stage under different operating conditions. They typically adopt a uniform delay correction strategy, resulting in insufficient delay correction accuracy and difficulty in ensuring consistency between system observation data and the actual state in complex dynamic environments. Summary of the Invention
[0004] To achieve the above objectives, the present invention provides the following technical solution: An intelligent operation and maintenance scheduling method for an information system, the method comprising: S1: Obtain resource load data of the information system, and annotate the resource load data with corresponding timestamp information to form a system observation data sequence; S2: Based on the system observation data sequence, the operation process of the information system is divided into states to determine the current operation state category. Based on the historical system observation data sequence, the observation delay data under different operation state categories is statistically analyzed to obtain the offset characteristic range of the observation delay data under different operation state categories. S3: Based on the offset characteristic range corresponding to the current operating state category, select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value, and perform independent time compensation on the timestamps of the corresponding data records in the acquisition stage, identification stage, and execution stage of the system observation data sequence to obtain the corrected system observation data; S4: Generate a phased delay correction strategy based on the corrected system observation data and add it to the intelligent operation and maintenance scheduling strategy library.
[0005] Furthermore, a system observation data sequence is formed, including: Obtain resource load data from the information system and annotate the resource load data with timestamp information to form resource load data with timestamps; The resource load data is continuously segmented according to a preset sliding time window to obtain multiple time segments with time overlap. Based on the aforementioned time overlap relationship, the resource load data within adjacent time segments are subjected to time continuity correlation processing, and the time sequence of the data in each time segment after correlation processing is reconstructed to form a system observation data sequence.
[0006] Furthermore, determine the current operating status category, including: Based on the resource load data corresponding to different time segments in the system observation data sequence, the resource load characteristics of each time segment are extracted; The resource load characteristics corresponding to different time segments in the historical system observation data sequence are used as a historical feature sample set, and the historical feature sample set is labeled with a status to form a set of operating status categories. The resource load characteristics of the current time segment are matched with the reference characteristics corresponding to the set of running status categories to determine the running status category corresponding to the current time segment; Time consistency processing is performed on the running status categories corresponding to multiple time segments obtained based on the sliding time window to determine the current running status category.
[0007] Furthermore, a set of operational status categories is formed, including: Unsupervised clustering is performed on the historical feature sample set to divide the historical feature samples into multiple clusters. Each cluster corresponds to a historical operating state category, and the cluster center of each cluster is used as the reference feature of the corresponding historical operating state category to form an operating state category set.
[0008] Furthermore, statistical analysis was conducted on the observation delay data under different operational status categories, including: The observation delay data includes system acquisition delay data, system identification delay data, and system execution delay data; The system acquisition delay data is the time difference between the occurrence of resource load data in the information system and the time of acquisition and recording. The system identification delay data is the time difference between the completion of resource load data acquisition and the completion of operation status category identification. The system execution delay data is the time difference between the completion of information system operation status category identification and the time when the resource scheduling strategy takes effect.
[0009] Furthermore, the offset characteristic range of the observed delay data under different operating state categories is obtained, including: Based on historical system observation data sequences, the observation delay data of multiple historical time segments under the same operating state category are statistically analyzed, the statistical distribution parameters of the observation delay data are calculated, and the offset characteristic range corresponding to the operating state category is determined.
[0010] Furthermore, the corrected system observation data are obtained, including: Obtain the offset characteristic range corresponding to the current running status category, and select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value from the offset characteristic range; The system observation data is obtained by independently compensating the timestamps of the corresponding data records in the acquisition, identification and execution phases of the system observation data sequence using the system acquisition delay correction benchmark value, the system identification delay correction benchmark value and the system execution delay correction benchmark value respectively, so as to obtain the corrected system observation data.
[0011] Furthermore, selecting the system acquisition delay correction reference value, the system identification delay correction reference value, and the system execution delay correction reference value from the offset characteristic range includes: Obtain multiple historical sample time segments under the current running status category and the base timestamp corresponding to each historical sample time segment; Multiple candidate correction baseline values are preset within the offset characteristic ranges corresponding to the system acquisition delay data, system identification delay data, and system execution delay data, respectively. Time compensation is performed on the original observation delay data of the multiple historical sample time segments using each candidate correction benchmark value. The compensated timestamp is compared with the corresponding benchmark timestamp, and the sample pass rate corresponding to each candidate correction benchmark value is calculated. The sample pass rate is the percentage of samples whose error value between the compensated timestamp and the benchmark timestamp is lower than a preset error threshold. The candidate correction benchmark value with the highest sample pass rate within the offset characteristic range of each stage is determined as the delay correction benchmark value for the corresponding stage.
[0012] Furthermore, a phased delay correction strategy is generated based on the corrected system observation data and added to the intelligent operation and maintenance scheduling strategy library, including: If the sample pass rate reaches or exceeds the preset ratio threshold, a mapping relationship is established between the current operating status category and the collection delay correction benchmark value, the identification delay correction benchmark value and the execution delay correction benchmark value, a phased delay correction strategy is generated and added to the intelligent operation and maintenance scheduling strategy library; If the sample pass rate is lower than the preset ratio threshold, the offset characteristic range is expanded, and multiple candidate correction benchmark values are re-preset within the expanded offset characteristic range. The candidate value with the highest pass rate is re-determined as the delay correction benchmark value for the corresponding stage. This process is repeated until the sample pass rate reaches or exceeds the preset ratio threshold.
[0013] Furthermore, the information system intelligent operation and maintenance scheduling system includes: The data acquisition module is used to acquire resource load data of the information system and mark the resource load data with corresponding timestamp information to form a system observation data sequence. The state recognition module is used to classify the state of the information system operation process according to the system observation data sequence, determine the current operation state category, and statistically analyze the observation delay data under different operation state categories based on the historical system observation data sequence to obtain the offset characteristic range of the observation delay data under different operation state categories. The delay correction module is used to select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value based on the offset characteristic range corresponding to the current operating state category. It then independently compensates the timestamps of the data records corresponding to the acquisition, identification, and execution phases in the system observation data sequence to obtain the corrected system observation data. The strategy generation module is used to generate phased delay correction strategies based on the corrected system observation data and add them to the intelligent operation and maintenance scheduling strategy library.
[0014] This invention provides an intelligent operation and maintenance scheduling method and system for information systems. It has the following beneficial effects: 1. This invention divides the data processing of an information system into an acquisition stage, an identification stage, and an execution stage, and independently models and statistically analyzes the observation delay of each stage. This effectively avoids the indivisibility problem caused by mixing delays from different sources, achieves a fine characterization of the delay offset range of each stage, and significantly improves the accuracy of the system's observation data time alignment correction.
[0015] 2. This invention establishes a mapping relationship between operating status categories and delay correction benchmark values for each stage, forming an intelligent operation and maintenance scheduling strategy library. This enables the information system to adaptively match the corresponding staged delay correction strategy based on the real-time identified operating status categories, maintaining high time alignment accuracy under different operating states. At the same time, the reliability of the correction benchmark values is ensured through a sample verification mechanism, effectively solving the problem that the unified correction strategy in the prior art is difficult to adapt to dynamic environmental changes, and providing an accurate and reliable observation data foundation for operation and maintenance scheduling decisions. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the intelligent operation and maintenance scheduling method for an information system according to the present invention. Figure 2 This is a schematic diagram of the intelligent operation and maintenance scheduling system of the information system of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] like Figure 1 As shown, the intelligent operation and maintenance scheduling method for information systems includes: S1: Obtain resource load data of the information system, and annotate the resource load data with corresponding timestamp information to form a system observation data sequence; S2: Based on the system observation data sequence, the operation process of the information system is divided into states to determine the current operation state category. Based on the historical system observation data sequence, the observation delay data under different operation state categories is statistically analyzed to obtain the offset characteristic range of the observation delay data under different operation state categories. S3: Based on the offset characteristic range corresponding to the current operating state category, select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value, and perform independent time compensation on the timestamps of the corresponding data records in the acquisition stage, identification stage, and execution stage of the system observation data sequence to obtain the corrected system observation data; S4: Generate a phased delay correction strategy based on the corrected system observation data and add it to the intelligent operation and maintenance scheduling strategy library.
[0019] Forming a system observation data sequence, including: Obtain resource load data from the information system and annotate the resource load data with timestamp information to form resource load data with timestamps; The resource load data is continuously segmented according to a preset sliding time window to obtain multiple time segments with time overlap. Based on the aforementioned time overlap relationship, the resource load data within adjacent time segments are subjected to time continuity correlation processing, and the time sequence of the data in each time segment after correlation processing is reconstructed to form a system observation data sequence.
[0020] In a specific embodiment of the present invention, the process of acquiring resource load data of an information system and annotating the resource load data with timestamp information to form time-stamped resource load data is used to establish a system observation data sequence. The resource load data may include, but is not limited to, multi-dimensional performance indicators that reflect the operating status of the information system, such as CPU utilization, memory usage, network throughput, disk I / O indicators, or task queue length.
[0021] It should be noted that resource load data reflects the real-time occupancy of various computing resources during the operation of an information system. By continuously collecting resource indicators such as CPU, memory, storage, and network, the operating status changes of the information system under different business pressures can be dynamically depicted, thus providing objective data for operation and maintenance personnel or automated operation and maintenance systems. Resource load data is not only used to describe the current instantaneous operating status of the information system, but also to reflect the trend of information system load changes over time in the form of time series, thereby revealing the evolution process of the information system from low load, balanced load to high load or bottleneck state.
[0022] Specifically, during operation, the information system continuously collects the aforementioned resource load data at a fixed or dynamic sampling frequency, and adds a timestamp to each data record corresponding to the collection time, so that each resource load data corresponds to its occurrence time, forming a set of system observation data with time stamps.
[0023] Based on this, the time-stamped resource load data is continuously segmented according to a preset sliding time window to divide the observation data on the continuous time axis into multiple local time segments. The preset sliding time window is used to limit the time length covered by each time segment, such as a fixed-length time interval (e.g., 5 minutes, 10 minutes, or an adaptive length interval), and slides on the time axis with a fixed step size or an overlapping step size to obtain multiple adjacent time segments with time overlap. The time overlap relationship means that there is partial intersection between adjacent time segments on the time axis, so that the later time segment not only contains new observation data, but also retains some historical data of the previous time segment, thereby ensuring the continuity and smooth transition of the system observation data in the time dimension and avoiding state abrupt changes or information loss due to hard segmentation.
[0024] It should be noted that since resource load data in information systems usually have obvious periodic and sudden characteristics, such as load surges or instantaneous request peaks during peak business periods, sliding time windows can enhance the sensitivity to local abnormal changes through time overlap mechanisms, while reducing the bias of a single time slice in the overall status judgment, thus providing a stable time structure basis for operational status identification and delay analysis.
[0025] In a specific embodiment of the present invention, based on the aforementioned temporal overlap relationship, resource load data within adjacent time segments are subjected to temporal continuity correlation processing, and the time sequence of the correlated time segment data is reconstructed to form a continuously changing and temporally consistent system observation data sequence. Specifically, overlapping data in adjacent time segments are aligned and correlated. For overlapping data with differences, consistency correction is performed on resource load data in different time segments through time interpolation, weighted fusion, or smoothing, so that the resource load value corresponding to the same time point remains continuous and consistent across different segments, thereby eliminating the boundary breakage effect caused by window segmentation. Further, after completing the temporal continuity correlation processing, the resource load data within each time segment are reordered and spliced according to timestamps, so that the time boundaries between segments are smoothly transitioned, and the overall data satisfies the monotonically increasing and continuous evolution characteristics in the time dimension. Through the above correlation processing and sequence reconstruction operations, the discrete segment data originally represented by multiple overlapping time windows is transformed into a system observation data sequence that is continuous and consistent on the time axis, without obvious breaks, and can truly reflect the evolution process of the information system's operating state.
[0026] Determine the current running status category, including: Based on the resource load data corresponding to different time segments in the system observation data sequence, the resource load characteristics of each time segment are extracted; The resource load characteristics corresponding to different time segments in the historical system observation data sequence are used as a historical feature sample set, and the historical feature sample set is labeled with a status to form a set of operating status categories. The resource load characteristics of the current time segment are matched with the reference characteristics corresponding to the set of running status categories to determine the running status category corresponding to the current time segment; Time consistency processing is performed on the running status categories corresponding to multiple time segments obtained based on the sliding time window to determine the current running status category.
[0027] This forms a set of running status categories, including: Unsupervised clustering is performed on the historical sample set to divide the historical feature samples into multiple clusters. Each cluster corresponds to a historical operating state category, and the cluster center of each cluster is used as the reference feature of the corresponding historical operating state category to form a set of operating state categories.
[0028] It should be noted that after the system observation data sequence is divided into multiple time segments, each time segment is not directly used for classification. Instead, it is first subjected to feature processing to turn the time segment into a comparable state description vector. The historical system observation data sequence refers to the set of resource load data continuously collected from the information system over a long historical period. This set is organized in chronological order and contains corresponding timestamp information to characterize the changes in the operating status of the information system at different historical moments.
[0029] It should be noted that the resource load characteristics and observation latency characteristics of an information system are not fixed during actual operation, but change dynamically with the load level, business pressure, and resource competition of the information system. Therefore, the observation latency data under different operating states have significant state-dependent differences. By introducing operating state categories, the operation process of the information system can be divided into several statistically consistent state intervals, such as high load state, low load state, or resource bottleneck state. The corresponding observation latency data can be statistically analyzed under different state categories, thereby establishing a mapping relationship between "operating state and observation latency".
[0030] In a specific embodiment of the present invention, extracting resource load features for each time segment includes performing feature extraction processing on each time segment based on resource load data corresponding to different time segments in the system observation data sequence to form resource load features.
[0031] In a specific embodiment of the present invention, the process of constructing the set of operating state categories includes taking the resource load features corresponding to different time segments in the historical system observation data sequence as a set of historical feature samples. Since historical operating states usually do not have manually predefined labels, an unsupervised clustering method is used to automatically divide the set of historical feature samples, and samples with similar resource load features are grouped into the same cluster. Specifically, K-means, density-based clustering algorithms, or other unsupervised learning methods can be used to divide the distribution of historical feature samples in the feature space, thereby forming multiple clusters. Each cluster corresponds to a historical operating state category, such as high load state, low load state, IO intensive state, or resource balanced state.
[0032] It should be noted that the purpose of using unsupervised clustering to construct the set of operational state categories is to address the problems of lack of manual annotation of historical feature samples and the difficulty in clearly defining the boundaries of operational states in advance. For example, in a distributed information system, three-dimensional feature vectors such as CPU utilization, memory usage, and network throughput are extracted from historical resource load data: Sample A: (CPU 85%, Memory 90%, Network 80%), Sample B: (CPU 88%, Memory 92%, Network 75%), Sample C: (CPU 30%, Memory 40%, Network 20%), Sample D: (CPU 88%, Memory 92%, Network 75 ... (35% memory, 45% network), after inputting the above feature samples into the K-means clustering algorithm, the algorithm automatically divides them into multiple clusters according to the distance relationship between the samples in the feature space. For example: Cluster 1: {A, B} corresponds to the "high load operation status category", Cluster 2: {C, D} corresponds to the "low load operation status category"; and the cluster center of each cluster is, for example: high load center: (87%, 91%, 78%), low load center: (32%, 42%, 22%).
[0033] In one specific embodiment of the present invention, the resource load characteristics of the current time segment are matched with the reference characteristics in the set of operating state categories to determine the operating state category corresponding to the current time segment; wherein, the similarity matching can be implemented based on Euclidean distance, cosine similarity or other distance measurement methods to measure the degree of closeness between the current features and the reference features of each historical state category.
[0034] In one specific embodiment of the present invention, time consistency processing is performed on the running state categories corresponding to multiple time segments obtained based on sliding time windows to determine the current running state category. The purpose is to eliminate the instability in state classification caused by overlapping calculations or short-term fluctuations in the sliding window, thereby obtaining a stable current running state category. Running state categories with the same identification results in adjacent or overlapping time segments are continuously merged to form the same running state category. Short-term state changes with a duration less than a preset time threshold are filtered or smoothed to eliminate state jitter caused by local fluctuations, thereby obtaining a stable current running state category. The preset time threshold... The value is determined by professionals based on the frequency of state transitions, business response requirements, and the scheduling cycle of the information system in the historical system operation logs. It is usually set to the minimum time scale that can cover the duration of a typical state to avoid frequent state transitions caused by short-term fluctuations. For example, in the resource scheduling scenario of a distributed information system, analysis of historical logs shows that the average duration of the information system's operating state is about 3 minutes and short-term fluctuations mostly occur within 30 seconds. Therefore, the preset time threshold is set to 60 seconds or 90 seconds to filter state changes with a duration of less than the threshold, thereby ensuring the stability and consistency of the operating state categories in the time dimension.
[0035] The observed delay data under different operating status categories were statistically analyzed, including: The observation delay data includes system acquisition delay data, system identification delay data, and system execution delay data; The system acquisition delay data is the time difference between the occurrence of resource load data in the information system and the time of acquisition and recording. The system identification delay data is the time difference between the completion of resource load data acquisition and the completion of operation status category identification. The system execution delay data is the time difference between the completion of information system operation status category identification and the time when the resource scheduling strategy takes effect.
[0036] In a specific embodiment of the present invention, the process of statistically analyzing observation delay data under different operating state categories is used to characterize the time lag characteristics of data flow in the information system under different operating load conditions, thereby providing a quantitative basis for subsequent delay correction and resource scheduling strategies. Specifically, system acquisition delay data characterizes the time difference between the actual occurrence of the resource load state and its recording by the system acquisition module and its writing into the system observation sequence; this delay reflects the responsiveness of the data acquisition link and monitoring mechanism. System identification delay data characterizes the time difference between the completion of system observation data acquisition and the completion of operating state category identification; this delay reflects the responsiveness of operating state category identification. System execution delay data characterizes the time difference between the completion of operating state category identification and the actual execution and effectiveness of the resource scheduling strategy; this delay reflects the propagation and landing delay of scheduling instructions in the execution link.
[0037] It should be noted that by dividing the operation of the information system into three consecutive processing stages—"collection, identification, and execution"—and defining corresponding observation delay data for each stage, the overall observation delay of the information system is decomposed into three sub-delay components that can be independently modeled and statistically analyzed. For example, under high load conditions, the system collection delay data may increase significantly due to data reporting congestion, while the system execution delay data may increase due to increased queuing time in the resource scheduling queue. Under low load conditions, various delays are usually at a low and stable level. That is, different observation delay data exist under different operating state categories. Through the above method, this invention achieves a phased decomposition of the system's observation delay, making the delay no longer a single overall indicator, but a structured feature that can dynamically change with the operating state, thereby improving the refinement and adaptability of subsequent delay correction and resource scheduling strategies.
[0038] The offset characteristic range of the observed delay data under different operating state categories is obtained, including: Based on historical system observation data sequences, the observation delay data of multiple historical time segments under the same operating state category are statistically analyzed, the statistical distribution parameters of the observation delay data are calculated, and the offset characteristic range corresponding to the operating state category is determined.
[0039] In a specific embodiment of the present invention, obtaining the offset characteristic range of observation delay data under different operating state categories includes: extracting observation delay data within each time segment based on multiple historical time segments corresponding to the same operating state category in a historical system observation data sequence, wherein the observation delay data includes system acquisition delay data, system identification delay data, and system execution delay data; further, performing statistical analysis on the observation delay data under the same operating state category to obtain statistical distribution parameters of the observation delay data under that operating state category, wherein the statistical distribution parameters include, but are not limited to, mean, variance, standard deviation, quantiles, or confidence intervals, used to characterize the number of observation delays under that operating state category. The central tendency and fluctuation range of the data are determined; based on the statistical distribution parameters, the offset characteristic range corresponding to the operating state category is determined. Specifically, the offset characteristic range is used to characterize the typical fluctuation range of the observation delay data under the operating state category. For example, the delay offset characteristic range is determined by the mean ± standard deviation or by confidence interval, thereby characterizing the variable interval characteristics of the observation delay under the operating state category, rather than a single fixed value. For example, under the high load operating state category, the mean of the system acquisition delay is 2.8 seconds and the standard deviation is 0.6 seconds, obtained by statistical analysis of historical data. Then, its offset characteristic range can be determined as [2.2 seconds, 3.4 seconds], which is used to represent the typical fluctuation range of the acquisition delay under this state.
[0040] It should be noted that even within the same operational status category, there are still differences in the observed latency data. This is because the same operational status category contains a variety of micro-level operational fluctuation factors. Specifically, although the resource load of an information system is generally at a relatively consistent level within the same operational status category, it is still affected by factors such as instantaneous request fluctuations, local resource competition, and changes in the system scheduling queue within specific time segments. This results in a certain degree of fluctuation in the observed latency on a short time scale. Therefore, the observed latency data within the same operational status category exhibits a fluctuating distribution around a certain statistical center value.
[0041] The corrected system observation data obtained include: Obtain the offset characteristic range corresponding to the current running status category, and select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value from the offset characteristic range; The system observation data is obtained by independently compensating the timestamps of the corresponding data records in the acquisition, identification and execution phases of the system observation data sequence using the system acquisition delay correction benchmark value, the system identification delay correction benchmark value and the system execution delay correction benchmark value respectively, so as to obtain the corrected system observation data.
[0042] Selecting the system acquisition delay correction reference value, the system identification delay correction reference value, and the system execution delay correction reference value from the offset characteristic range includes: Obtain multiple historical sample time segments under the current running status category and the base timestamp corresponding to each historical sample time segment; Multiple candidate correction baseline values are preset within the offset characteristic ranges corresponding to the system acquisition delay data, system identification delay data, and system execution delay data, respectively. Time compensation is performed on the original observation delay data of the multiple historical sample time segments using each candidate correction benchmark value. The compensated timestamp is compared with the corresponding benchmark timestamp, and the sample pass rate corresponding to each candidate correction benchmark value is calculated. The sample pass rate is the percentage of samples whose error value between the compensated timestamp and the benchmark timestamp is lower than a preset error threshold. The candidate correction benchmark value with the highest sample pass rate within the offset characteristic range of each stage is determined as the delay correction benchmark value for the corresponding stage.
[0043] In a specific embodiment of the present invention, selecting the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value from the offset characteristic range specifically includes: obtaining multiple historical sample time segments under the current operating state category and the benchmark timestamp corresponding to each historical sample time segment; the historical sample time segment refers to multiple consecutive time periods obtained by dividing the historical system observation data sequence according to a preset time window, and each time segment contains the resource load data and the corresponding timestamp information within that time period.
[0044] It should be noted that historical sample time segments are historical data that have already occurred. Therefore, the reference timestamps corresponding to each data record can be obtained retrospectively through various means independent of the software acquisition chain. The reference timestamp refers to the reference value of the actual moment when the resource load state occurred in the physical world, and its source is independent of the conventional data acquisition and recording chain of the information system. Specifically, the reference timestamp can be obtained in the following ways: For verification of acquisition delay, the hardware interrupt record timestamp of the data packet or I / O request arriving on the physical device can be obtained from the hardware layer of the information system and used as the reference timestamp of the actual moment when the resource load state occurred; For verification of identification delay, the time reference aligned across nodes in the distributed tracing system or the moment confirmed by manual verification can be used as the reference time. For verifying execution delays, the resource allocation completion time recorded at the underlying resource scheduler can be used as the baseline timestamp. Since the above-mentioned baseline timestamp comes from independent reference sources such as the hardware layer, distributed tracing, or underlying scheduling records, it is not affected by software layer acquisition delays, recognition delays, and execution delays. Therefore, it can be used as a true benchmark for judging the correction effect. That is, in the offline verification stage, since the time has already occurred, the information system has the conditions and sufficient time to check the hardware logs, distributed tracing records, or scheduler underlying logs, thereby obtaining more accurate time reference information than during online operation. However, in the online operation stage, the information system cannot obtain the corresponding baseline timestamp for every piece of real-time data. Therefore, it is necessary to use the correction baseline value verified in the offline stage to quickly compensate for the real-time data.
[0045] In a specific embodiment of the present invention, multiple candidate correction benchmark values are preset within the offset characteristic ranges corresponding to the acquisition delay, recognition delay, and execution delay, respectively. Specifically, multiple candidate correction benchmark values are selected using a preset method based on the offset characteristic ranges of the system acquisition delay data, the system recognition delay data, and the system execution delay data. The preset method can be uniform sampling (i.e., selecting multiple candidate values uniformly within the range according to a preset step size), or random sampling or quantile sampling based on historical distribution. For example, for the offset characteristic range of the acquisition delay [1.2s, 3.8s], uniform sampling with a step size of 0.2s can preset the set of candidate correction benchmark values as {1.2, 1.4, 1.6, 1.8, 2.0, 2.2, 2.4, 2.6, 2.8, 3.0, 3.2, 3.4, 3.6, 3.8} (unit: seconds).
[0046] In one specific embodiment of the present invention, time compensation is performed on the original observation delay data of the plurality of historical sample time segments using each candidate correction benchmark value. The compensated timestamps are compared with the corresponding benchmark timestamps, and the sample pass rate corresponding to each candidate correction benchmark value is calculated. Specifically, taking the system acquisition delay data as an example, let a certain candidate correction benchmark value of the system acquisition delay data be... The original timestamp of a data record in any historical sample time segment (i.e., the timestamp recorded by the information system software) is The reference timestamp corresponding to this data record (i.e., the actual time of occurrence obtained from independent reference sources such as hardware logs) is: The compensated timestamp is , ; the compensated timestamp Compared with the base timestamp Perform a comparison and calculate the error value e. When e is less than or equal to a preset error threshold, the sample correction is deemed successful; otherwise, the sample correction is deemed unsuccessful. The preset error threshold is set by a person skilled in the art based on the actual time accuracy requirements of the information system. For any candidate correction benchmark value, all data records in all historical sample time segments are traversed, the number of samples that have passed correction is counted, and the proportion of the number of samples that have passed correction to the total number of samples is calculated, which is the sample pass rate corresponding to the candidate correction benchmark value. For example, if there are 100 historical samples, and a certain candidate correction benchmark value makes the error value between the compensated timestamp and the benchmark timestamp of 91 of the samples lower than the preset error threshold, then the sample pass rate of the candidate correction benchmark value is 91%. For recognition delay and execution delay, within their respective offset characteristic ranges, for their respective sets of candidate correction benchmark values, the same method as above is used to perform time compensation on the data records of the recognition stage or execution stage in the historical sample time segment with each candidate correction benchmark value, and the compensated timestamp is compared with the corresponding benchmark timestamp to count the sample pass rate corresponding to each candidate correction benchmark value.
[0047] It should be noted that, in the above process, "raw observation delay data" refers to the original timestamp information of each data record contained in the historical sample time segment. This timestamp is recorded by the software layer of the information system during data acquisition, identification, or execution. By using the candidate correction benchmark value as a time compensation amount to correct the timestamp, and comparing the corrected timestamp with the independently acquired benchmark timestamp, the correction effect of the candidate correction benchmark value can be quantitatively evaluated. The "error value" reflects the residual time difference between the compensated timestamp and the actual occurrence time. The lower the error value, the closer the candidate correction benchmark value is to the actual delay value, and the better the correction effect. The higher the error value, the worse the correction effect. Since the benchmark timestamp comes from independent reference sources such as hardware records, it does not contain the acquisition, identification, or execution delay of the software layer. Therefore, the above error value can accurately measure the compensation accuracy of the candidate correction benchmark value for the delay of the corresponding stage.
[0048] In a specific embodiment of the present invention, the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value are used respectively to independently compensate the timestamps of the corresponding data records in the acquisition stage, identification stage, and execution stage of the system observation data sequence, so as to obtain the corrected system observation data. Specifically, each data record in the system observation data sequence contains resource load data and its corresponding timestamp. According to the stage in the information system processing flow of the data record, where the stage includes the acquisition stage, identification stage, or execution stage, the timestamp is compensated by using the correction benchmark value of the corresponding stage.
[0049] It should be noted that, due to the different delay mechanisms in the acquisition, identification, and execution phases, the delay correction benchmark values for each phase are determined through independent sample verification processes. Therefore, the correction benchmark values for the three phases are independent of each other, achieving phased independent time compensation for the system observation data sequence. After the above phased time alignment correction, the timestamps of each data record in the system observation data sequence have completed the delay compensation corresponding to their respective phases. The corrected system observation data obtained thereby eliminates the time deviation caused by the superposition of delays in data acquisition, status identification, and policy execution, and can more realistically reflect the actual evolution of the information system's resource load status on the time axis.
[0050] A phased delay correction strategy is generated based on the corrected system observation data and added to the intelligent operation and maintenance scheduling strategy library, including: If the sample pass rate reaches or exceeds the preset ratio threshold, a mapping relationship is established between the current operating status category and the collection delay correction benchmark value, the identification delay correction benchmark value and the execution delay correction benchmark value, a phased delay correction strategy is generated and added to the intelligent operation and maintenance scheduling strategy library; If the sample pass rate is lower than the preset ratio threshold, the offset characteristic range is expanded, and multiple candidate correction benchmark values are re-preset within the expanded offset characteristic range. The candidate value with the highest pass rate is then re-determined as the delay correction benchmark value for the corresponding stage. This process is repeated until the sample pass rate reaches or exceeds the preset ratio threshold.
[0051] In a specific embodiment of the present invention, a phased delay correction strategy is generated based on the corrected system observation data and added to the intelligent operation and maintenance scheduling strategy library. Specifically, this includes: determining whether the sample pass rate obtained through the above sample verification process reaches or exceeds a preset percentage threshold, wherein the preset percentage threshold is set by those skilled in the art based on the actual requirements of the information system for correction reliability, for example, it can be set to 90%; if the sample pass rate reaches or exceeds the preset percentage threshold, it indicates that the currently selected phased delay correction benchmark value has been effectively verified by historical samples and has high credibility. In this case, a mapping relationship is established between the current operating state category and the collected delay correction benchmark value, the identified delay correction benchmark value, and the executed delay correction benchmark value; for example, {Operating state category: "High load state", collected delay correction benchmark value: 2.0s, identified delay correction benchmark value: 1.5s, executed delay correction benchmark value: 0.8s, sample pass rate: 94%}, and the above mapping relationship is used as a phased delay correction strategy and added to the intelligent operation and maintenance scheduling strategy library.
[0052] If the sample pass rate is lower than the preset proportion threshold, it indicates that the current offset characteristic range setting may not be accurate enough, resulting in the candidate correction benchmark value selected within this range not being able to pass the verification of a sufficient proportion of samples. By expanding the offset characteristic range, for example, increasing the confidence level from 95% to 99%, or increasing the standard deviation multiple from 1.96 times to 2.58 times, a wider range of delay values can be covered. Within the expanded offset characteristic range, multiple candidate correction benchmark values are re-preset, and time compensation, comparison, and sample pass rate statistics are performed on each candidate correction benchmark value again using multiple historical sample time segments under the current operating status category and the benchmark timestamps corresponding to each historical sample time segment. The candidate value with the highest pass rate is re-determined as the delay correction benchmark value for the corresponding stage. Through iterative iteration until the sample pass rate reaches or exceeds the preset proportion threshold, the finally verified staged delay correction benchmark value is associated with the current operating status category to generate a staged delay correction strategy and add it to the intelligent operation and maintenance scheduling strategy library.
[0053] It should be noted that the intelligent operation and maintenance scheduling strategy library contains a mapping relationship between multiple operating status categories and their corresponding phased delay correction benchmark values. During the actual operation of the information system, after the current operating status category is identified in real time, the information system searches for a phased delay correction strategy that matches the operating status category from the intelligent operation and maintenance scheduling strategy library. It then uses the acquisition delay correction benchmark value, identification delay correction benchmark value, and execution delay correction benchmark value recorded in the strategy to independently compensate the timestamps of the corresponding data records in the acquisition, identification, and execution phases of the system's observed data sequence. Through the construction and use of the above strategy library, the information system can adopt phased delay correction parameters adapted to different operating states to achieve differentiated time alignment correction.
[0054] like Figure 2 As shown, the intelligent operation and maintenance scheduling system for information systems is applied to the intelligent operation and maintenance scheduling method for information systems described above. The system includes: The data acquisition module is used to acquire resource load data of the information system and mark the resource load data with corresponding timestamp information to form a system observation data sequence. The state recognition module is used to classify the state of the information system operation process according to the system observation data sequence, determine the current operation state category, and statistically analyze the observation delay data under different operation state categories based on the historical system observation data sequence to obtain the offset characteristic range of the observation delay data under different operation state categories. The delay correction module is used to select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value based on the offset characteristic range corresponding to the current operating state category. It then independently compensates the timestamps of the data records corresponding to the acquisition, identification, and execution phases in the system observation data sequence to obtain the corrected system observation data. The strategy generation module is used to generate phased delay correction strategies based on the corrected system observation data and add them to the intelligent operation and maintenance scheduling strategy library.
[0055] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code, which, when executed by the one or more processors, can perform the intelligent operation and maintenance scheduling method and system for the information system as described above.
[0056] The systems and methods according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the intelligent operation and maintenance scheduling method and system for information systems provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs.
[0057] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0058] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An intelligent operation and maintenance scheduling method for an information system, characterized in that, The method includes: S1: Obtain resource load data of the information system, and annotate the resource load data with corresponding timestamp information to form a system observation data sequence; S2: Based on the system observation data sequence, the operation process of the information system is divided into states to determine the current operation state category. Based on the historical system observation data sequence, the observation delay data under different operation state categories is statistically analyzed to obtain the offset characteristic range of the observation delay data under different operation state categories. S3: Based on the offset characteristic range corresponding to the current operating state category, select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value, and perform independent time compensation on the timestamps of the data records corresponding to the acquisition stage, identification stage, and execution stage in the system observation data sequence to obtain the corrected system observation data; S4: Generate a phased delay correction strategy based on the corrected system observation data and add it to the intelligent operation and maintenance scheduling strategy library.
2. The intelligent operation and maintenance scheduling method for an information system according to claim 1, characterized in that, Forming a system observation data sequence, including: Obtain resource load data from the information system and annotate the resource load data with timestamp information to form resource load data with timestamps; The resource load data is continuously segmented according to a preset sliding time window to obtain multiple time segments with time overlap. Based on the aforementioned time overlap relationship, the resource load data within adjacent time segments are subjected to time continuity correlation processing, and the time sequence of the data in each time segment after correlation processing is reconstructed to form a system observation data sequence.
3. The intelligent operation and maintenance scheduling method for an information system according to claim 2, characterized in that, Determine the current running status category, including: Based on the resource load data corresponding to different time segments in the observation data sequence of the information system, the resource load characteristics of each time segment are extracted. The resource load characteristics corresponding to different time segments in the historical system observation data sequence are used as a historical feature sample set, and the historical feature sample set is labeled with a status to form a set of operating status categories. The resource load characteristics of the current time segment are matched with the reference characteristics corresponding to the set of running status categories to determine the running status category corresponding to the current time segment; Time consistency processing is performed on the running status categories corresponding to multiple time segments obtained based on the sliding time window to determine the current running status category.
4. The intelligent operation and maintenance scheduling method for an information system according to claim 3, characterized in that, This forms a set of running status categories, including: Unsupervised clustering is performed on the historical feature sample set to divide the historical feature samples into multiple clusters. Each cluster corresponds to a historical operating state category, and the cluster center of each cluster is used as the reference feature of the corresponding historical operating state category to form an operating state category set.
5. The intelligent operation and maintenance scheduling method for an information system according to claim 4, characterized in that, Statistics were compiled on the observed delay data under different operating status categories, including: The observation delay data includes system acquisition delay data, system identification delay data, and system execution delay data; The system acquisition delay data is the time difference between the occurrence of resource load data in the information system and the time of acquisition and recording. The system identification delay data is the time difference between the completion of resource load data acquisition and the completion of operation status category identification. The system execution delay data is the time difference between the completion of information system operation status category identification and the time when the resource scheduling strategy takes effect.
6. The intelligent operation and maintenance scheduling method for an information system according to claim 5, characterized in that, The offset characteristic range of the observed delay data under different operating state categories is obtained, including: Based on historical system observation data sequences, the observation delay data of multiple historical time segments under the same operating state category are statistically analyzed, the statistical distribution parameters of the observation delay data are calculated, and the offset characteristic range corresponding to the operating state category is determined.
7. The intelligent operation and maintenance scheduling method for an information system according to claim 6, characterized in that, The corrected system observation data obtained include: Obtain the offset characteristic range corresponding to the current running status category, and select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value from the offset characteristic range; The system observation data is obtained by independently compensating the timestamps of the corresponding data records in the acquisition, identification and execution phases of the system observation data sequence using the system acquisition delay correction benchmark value, the system identification delay correction benchmark value and the system execution delay correction benchmark value respectively, so as to obtain the corrected system observation data.
8. The intelligent operation and maintenance scheduling method for an information system according to claim 7, characterized in that, Selecting the system acquisition delay correction reference value, the system identification delay correction reference value, and the system execution delay correction reference value from the offset characteristic range includes: Obtain multiple historical sample time segments under the current running status category and the base timestamp corresponding to each historical sample time segment; Multiple candidate correction baseline values are preset within the offset characteristic ranges corresponding to the system acquisition delay data, system identification delay data, and system execution delay data, respectively. Time compensation is performed on the original observation delay data of the multiple historical sample time segments using each candidate correction benchmark value. The compensated timestamp is compared with the corresponding benchmark timestamp. The sample pass rate corresponding to each candidate correction benchmark value is calculated. The sample pass rate is the percentage of samples whose error value between the compensated timestamp and the benchmark timestamp is lower than a preset error threshold. The candidate correction benchmark value with the highest sample pass rate within the offset range of each stage is determined as the delay correction benchmark value for the corresponding stage.
9. The intelligent operation and maintenance scheduling method for an information system according to claim 8, characterized in that, A phased delay correction strategy is generated based on the corrected system observation data and added to the intelligent operation and maintenance scheduling strategy library, including: If the sample pass rate reaches or exceeds the preset ratio threshold, a mapping relationship is established between the current operating status category and the collection delay correction benchmark value, the identification delay correction benchmark value and the execution delay correction benchmark value, a phased delay correction strategy is generated and added to the intelligent operation and maintenance scheduling strategy library; If the sample pass rate is lower than the preset ratio threshold, the offset characteristic range is expanded, and multiple candidate correction benchmark values are re-preset within the expanded offset characteristic range. The candidate value with the highest pass rate is re-determined as the delay correction benchmark value for the corresponding stage. This process is repeated until the sample pass rate reaches or exceeds the preset ratio threshold.
10. An intelligent operation and maintenance scheduling system for information systems, characterized in that: The system, which is applied to the intelligent operation and maintenance scheduling method for an information system as described in any one of claims 1-9, comprises: The data acquisition module is used to acquire resource load data of the information system and mark the resource load data with corresponding timestamp information to form a system observation data sequence. The state recognition module is used to classify the state of the information system operation process according to the system observation data sequence, determine the current operation state category, and statistically analyze the observation delay data under different operation state categories based on the historical system observation data sequence to obtain the offset characteristic range of the observation delay data under different operation state categories. The delay correction module is used to select the system acquisition delay correction benchmark value, the system identification delay correction benchmark value, and the system execution delay correction benchmark value based on the offset characteristic range corresponding to the current operating state category. It then independently compensates the timestamps of the data records corresponding to the acquisition, identification, and execution phases in the system observation data sequence to obtain the corrected system observation data. The strategy generation module is used to generate phased delay correction strategies based on the corrected system observation data and add them to the intelligent operation and maintenance scheduling strategy library.