Intelligent regulation and control method for monitoring period of DevOps platform

By obtaining and analyzing multi-dimensional indicator data flow on the DevOps platform, combining event level model and fuzzy logic judgment technology, intelligent evaluation and scheduling of system status and event urgency are realized, solving the problems of inflexible monitoring cycles and insane scheduling strategies, and improving the system's operating efficiency and response capabilities.

CN119938313AActive Publication Date: 2025-05-06THREE GORGES ENVIRONMENTAL TECH CO LTD +1

Patent Information

Application Number
CN202411861773.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The monitoring cycle lacks flexibility in the monitoring cycle, insufficient multi-dimensional indicator analysis, and insufficient scheduling and disposal strategies in the monitoring of existing DevOps platforms.

Method used

By obtaining the multi-dimensional index data flow in the system operation, a multi-dimensional vector representing the system state is formed, combining event level models, rule bases and fuzzy logic judgment technology, quantitative processing and logical operations of system state and event urgency are realized, and fuzzy scheduling strategies and flexible management plans are triggered.

Benefits of technology

It improves the efficient operation of the system under different load conditions, discovers potential abnormal situations in advance, reduces the risk of business interruption caused by system failure, and improves the system's response speed and processing capabilities to respond to emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938313A_ABST
    Figure CN119938313A_ABST
Patent Text Reader

Abstract

An intelligent regulation and control method for a DevOps platform monitoring period comprises the following steps: firstly, acquiring a multi-dimensional index data stream of a system, and forming an instant multi-dimensional vector through stream-oriented calculation, window processing, feature combination and the like so as to judge the state of the system; and quantifying event information by using an event level model to determine an emergency degree. A confidence vector is obtained through operation of a preset rule base, a corresponding scheduling strategy is triggered, and resources are allocated according to the fuzzy logic and the event emergency degree. Meanwhile, the event priority is determined in combination with expert experience and fuzzy logic, and a flexible management plan is executed and a result is evaluated accordingly. And finally, according to the rule base evaluation, dynamically adjusting the rule weight ratio, and optimizing the judgment on the multi-dimensional indexes of the system. According to the invention, dynamic regulation and control of a monitoring period, accurate evaluation of systems and events, intelligent scheduling and handling are realized, and the stability, high efficiency and self-adaptability of a DevOps platform are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of system monitoring, and in particular relates to a method for intelligently controlling the monitoring cycle of a DevOps platform. Background Art

[0002] Existing monitoring methods are insufficient for comprehensive analysis and in-depth exploration of multi-dimensional system metrics. DevOps platforms generate massive amounts of multi-dimensional data during operation, including system load, service response time, and various types of event information. Simply collecting and analyzing this data in isolation makes it difficult to fully and accurately grasp the overall operational status of the system. For example, focusing solely on system load or service response time may overlook the potential impact of event types and numbers on system status. Furthermore, the lack of effective data integration and feature extraction methods makes it impossible to extract key information that truly represents system status from this complex data, making it difficult to accurately assess system health and provide early warning of potential issues.

[0003] Furthermore, traditional monitoring and handling mechanisms lack intelligent scheduling strategies and flexible response plans when faced with system anomalies or incidents. When the system experiences an unhealthy state or an emergency, they are often limited to simple alerts and limited resource allocation according to pre-set fixed rules. They are unable to comprehensively assess and implement more reasonable and efficient scheduling measures based on multiple factors, such as the urgency of the incident, the actual system status, and historical experience. Furthermore, there is a lack of effective evaluation and feedback mechanisms for the results of handling operations, preventing timely optimization and adjustment of management strategies based on the effectiveness of handling operations. This results in a lack of capacity to cope with complex and changing system operations.

[0004] In summary, the technical problems to be solved by the present invention are: the lack of flexibility in monitoring cycles, insufficient multi-dimensional indicator analysis, and insufficient intelligence in scheduling and handling strategies in existing DevOps platform monitoring. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a DevOps platform monitoring cycle intelligent control method. Compared with the traditional fixed scheduling strategy, the present invention can better adapt to various changes in the system operation process, improve resource utilization, avoid over-allocation or under-allocation of resources, and ensure that the system can operate efficiently under different load conditions. In order to solve the above technical problems, the technical solution adopted by the present invention is: A DevOps platform monitoring cycle intelligent control method, the steps are as follows: S1. During the operation of the DevOps platform, a multi-dimensional indicator data stream is obtained during the system operation. The multi-dimensional indicator data stream includes system load, service response time, and the specific type and number of events that occurred during the collection period. A multi-dimensional vector representing the system status is formed based on the multi-dimensional indicator data stream to obtain a real-time multi-dimensional vector. S2. Quantify the event information contained in the DevOps platform's real-time multidimensional vector using the established event level model. Match each event with a corresponding level parameter, combine the level parameters of each event into a level vector, and determine the urgency of the corresponding event based on the specific values ​​contained in the level vector. S3. Using a preset rule base, perform logical operations on the system status information values ​​in the DevOps platform's real-time multidimensional vector and the event urgency values ​​in the level vector. Derived from the rule base, the confidence vector of the corresponding system status information and the confidence vector of the corresponding event urgency information under the given conditions are obtained. S4. If more than a set number of system status information in the DevOps platform system status information confidence vector indicates an unhealthy state, a system status alarm instruction is issued to the system, and the fuzzy scheduling strategy mechanism is triggered based on the numerical value of the event urgency; S5. Obtain the expert experience processing rule sets corresponding to different types of events. Use fuzzy logic judgment technology combined with the expert experience processing rule sets to comprehensively calculate the real-time indicators and trigger events of the DevOps platform system, and determine the priority parameters based on the comprehensive calculation of the system resource status and urgency of the event. S6. Determine several corresponding flexible management plans based on the priority parameters, execute specific handling operation instruction sets according to the processes in the flexible management plans, obtain instruction handling results and feed them back to the information processing link for result evaluation; S7. Determine the dynamically adjusted rule weight ratio based on the rule base evaluation result information, re-judge the multi-dimensional indicators of the DevOps platform system based on the updated rule base, and determine the adjustment direction and parameter values ​​when the information processing mechanism is executed next time.

[0006] Preferably, the sub-steps of S1 are: S1.1, according to the system load data stream, using a stream computing engine to continuously receive system indicator data, and processing the collected load data by setting a window function to obtain a load average within a set time window; Assume that the system load data flow is: ,in Represents a time series, and the window function is , The load average ,in is the number of data points in the window, For the window time points; S1.2. Fusion system load average and service delay information flow, calculate the delay within a fixed time range, and obtain statistical data; suppose the service delay information flow is , the sliding window size is , service delay statistics: ; in The first time points; S1.3. Integrate the statistical data of the load mean and service delay to form a two-dimensional information list to represent the comprehensive information of the system load and service delay. Let the integrated two-dimensional list be: ; S1.4. Obtain a two-dimensional list of all event types and statistical times that occurred during the window period; suppose the event type set is: ,in is the total number of event types, and the corresponding number of events is , then the two-dimensional list of event types and times ; S1.5. Correlate the load and service delay information list with the event type and frequency list, count the number of events that occurred within the set time window, and count the number of events of different categories within the set time window to obtain a third list of the number of occurrences and types of each category, and merge them into a third list. Let the merged third list be: ; S1.6. Use feature combination technology to integrate the three lists output in the previous step into a multi-dimensional data stream and transform it into a data form that can comprehensively reflect the multi-dimensional characteristics of the system; let the feature combination function be , then the combined data: ; S1.7. The data after feature combination processing is regarded as a set of high-dimensional feature vector data; let the high-dimensional feature vector be ,in is the feature vector dimension, For combined data The eigenvalues ​​obtained after mapping or transformation; S1.8. Use time series related clustering technology to calculate time series correlation to explore the potential laws of data in time series; suppose the time series clustering algorithm is , the clustering result is ; S1.9. Calculate the correlation with the instantaneous time by constructing a time decay coefficient. If the time decay coefficient of data at a certain moment is greater than a preset value at a certain time, it is discarded to ensure the timeliness of the data. Suppose the time decay coefficient is: ,in is the attenuation factor, is the reference time point, like , If the time attenuation coefficient threshold is set to a preset value, the corresponding data will be discarded; S1.10. Obtain a list set with time information; let the set be: ; in is the number of data points, For the The time of the data point, is the corresponding eigenvector; S1.11. Use the PCA principal component analysis method to reduce the dimensionality of the principal components of the time list set obtained above, extract key information, and reduce the data dimension; let the PCA transformation matrix be , the principal component vector after dimensionality reduction is: ; in is the dimension after dimensionality reduction, , the PCA algorithm specifically calculates the covariance matrix: ; in is the mean of the eigenvector, and then solve the eigenvalue and eigenvector of the covariance matrix, and select the eigenvector corresponding to the larger eigenvalue to form the transformation matrix ; S1.12. Extract the vector group of the main components in the set, and obtain the main component with the largest proportion at each moment as the feature data to obtain the immediacy multidimensional vector; let the immediacy multidimensional vector be: ; in is the number of principal components finally selected, For The principal component values ​​selected according to the proportion; S1.13. Calculate the instantaneous multidimensional vector in an unsupervised manner, calculate the Euclidean distance from the vector to the center of the known system state category, and determine the minimum distance to the center of all known vectors; suppose the set of known system state category centers is: ; in is the number of system state categories, vector Go to Category Center The Euclidean distance of: ; in Category Center No. eigenvalues; S1.14, determine which system state is most similar; if the distance between the immediacy vector and the center of each known system state category is less than the state distance threshold, then the immediacy vector is classified into the most similar system state, and the calculated immediacy vector is classified into the most similar system state; let the state distance threshold be ,like ,but Belong to category ; S1.15. The state similarity results output from multiple time periods are processed through an adaptive deep neural network. Each system state is output to the deep network. If the network has a new output, the parameters of a certain system state are updated. The trained deep neural network is used as the decision basis for the immediate state calculation.

[0007] Preferably, the S2 method is: suppose the event level model is , the event information is , then the event level parameter , let the level vector be ,in is the number of event types and the urgency of the event according to The elements in are obtained through mapping relationships.

[0008] Preferably, the S3 sub-step is: S3.1. Collect the multi-dimensional vectors and level vectors generated during the operation of the DevOps platform system. The information obtained includes the system status information value and the event urgency value, and build an information database. Suppose the system status information value is , the event urgency value is , the information database is ; S3.2. Normalize the system status information values ​​in the database. According to the distribution of the event urgency levels, divide the event urgency values ​​into multiple intervals to obtain the normalized database and the level value information of the divided intervals. Let the normalization function be , the normalized system status information value is: , Assume that the event urgency is divided into intervals , the corresponding level value is ; S3.3. Divide the information into two parts, one is the state information group and the other is the level information group. Perform clustering on the state information group, calculate the distance between the other state information values ​​and each cluster center point generated by clustering, and obtain cluster data. Suppose the state information group clustering algorithm is , the clustering results are: , Let the cluster center point set be , Status information value To cluster center Distance: ; in for No. eigenvalues, for No. eigenvalues, is the feature dimension; S3.4. Preset a threshold and compare the system status information value in the database with the threshold. If it is higher than the threshold, the current value is determined to be a high-risk state, the corresponding status information group is a high-risk state, and an alarm instruction for the high-risk state is output to obtain an alarm instruction; set the threshold to ,like , it is determined to be a high-risk state and an alarm instruction is output ; S3.5. Use the alarm command to query the numerical information of the level information group. According to the rules defined in the rule base, if the queried level information group information matches the high urgency threshold of the level numerical information, an alarm signal for the event is generated, the event alarm signal data is output, and the information data of the emergency event is obtained; S3.6. Use the training set data to train a convolutional neural network classification model, set the network loss function and tune the network to obtain a mature network. The network inputs emergency information data, determines the matching relationship between the output emergency information and the confidence level of the system status information, and generates a confidence level association list. Let the training set be: ; in is the training sample index, For the true confidence level label, the network loss function uses the mean square error loss function: ; in is the number of training samples, For the The confidence level of the network output of each sample is calculated. The network parameters are adjusted by the gradient descent algorithm to minimize the loss function. After obtaining a mature network, the input emergency information data is processed to obtain a confidence level association list. .

[0009] Preferably, the sub-steps of S4 are: S4.1. Collect various status parameters of the DevOps platform system during operation, construct a multi-dimensional system status information set, and obtain the system state vector after data cleaning. S=(s1, s2; si), where si represents the value of the i-th state information; S4.2. Compare the value si of each state information in the system state vector S with the pre-set healthy range, and calculate the number of unhealthy state information to obtain the number N of unhealthy states. Determine whether the number N of unhealthy states is greater than the set threshold T. If N>T, output an alarm information set and generate an alarm instruction A. If N≤T, return to step S4.1 and continue collecting system state information. Assume the healthy range is: ; in For the The health range of status information, if , then the status information is unhealthy, count the number of unhealthy states: ; in is the indicator function, and the conditions are met , otherwise ,like , output alarm information set and alarm instructions ; S4.3. Obtain the urgency values ​​of multiple events corresponding to the DevOps platform system alarm information set, normalize the urgency values, and obtain the normalized urgency value E=(e1, e2; ei), where ei represents the urgency value of the i-th event; let the normalization function be ,but ; S4.4. Based on the normalized urgency value E, a fuzzy logic-based algorithm is used to determine the triggered dispatch level L. The core rule of the fuzzy logic algorithm is to determine the final dispatch level through fuzzy reasoning based on the correspondence between different urgency value ranges and pre-set dispatch levels, while considering the combined impact of the urgency of multiple events. The corresponding dispatch level is obtained. Assume that the fuzzy logic system has the following input fuzzy set: , corresponding to fuzzy sets with different ranges of urgency, Output fuzzy set , corresponding to the fuzzy sets of different scheduling levels; Fuzzy inference rule table , which represents the correspondence between different input fuzzy sets and output fuzzy sets. ; S4.5. After determining the scheduling level L, multiple scheduling strategies that meet this level are selected from a pre-established scheduling strategy library. These strategies are then integrated to obtain a fused scheduling strategy G. The DevOps platform scheduling engine executes the fused scheduling strategy G, and a resource allocation information data stream is sent through the communication module to determine a resource reallocation plan. Assume that the scheduling policy library is ,in is the number of strategies, selection and scheduling level Matching policy subset: , Fusion scheduling strategy: ,in is the strategy fusion function, and the resource allocation information data flow is set to , according to the fusion scheduling strategy Sure The content of the resource allocation is realized through the resource allocation amount and allocation object to achieve resource reallocation.

[0010] Preferably, the sub-steps of S5 are: S5.1. Collect historical DevOps platform operation and maintenance data and expert processing experience, classify and analyze them, and build a set of event processing rules and corresponding historical resource status data sets. S5.2. When an event is triggered, the real-time status information of the DevOps platform system is obtained synchronously based on multiple monitoring devices. The operating status information matches the type and urgency of the triggering event and the status information is assigned according to the preset fuzzy set interval. The fuzzy set and membership function parameters are called according to the event type query to calculate the membership of the assigned status information to obtain a number of operating status membership data. According to the expert experience rules set in the expert knowledge base, the rule sets corresponding to different states in the same type of emergency event records are selected to obtain a targeted solution set. The membership is jointly calculated with the expert rule set to obtain a normalized membership value. The normalized value is weighted to obtain the event priority. After the fuzzy set is delineated according to the status information, the fuzzy set is matched with the indicator membership in the generated operating data to obtain the calculated priority data. The priority data is then used to update the event record. According to the event urgency and event type, the calculated priority data is compared to obtain the processing rule priority ranking. The ranking result is input into the existing equipment expert rule base to guide the generation of operation and maintenance tasks.

[0011] Preferably, in S5.2, let the fuzzy set be: , corresponding to fuzzy sets of different state information or event types; The membership function is: ,in Status information or event-related data; The running status membership data is set as: ,in It is status information or number of event types; The expert experience rule set is: ,in is the number of rules; Event priority is set to: ,in is the weight coefficient; The processing rule priority order is set to: ,in is the sorting function.

[0012] Preferably, the sub-steps of S6 are: S6.1. Based on the content and priority information of the emergency plan group, the support vector machine algorithm is used to generate the initial instruction set of each disposal group; the support vector machine algorithm finds the optimal classification hyperplane based on the mapping relationship between the priority parameters and the characteristics of the emergency plan group through learning and training of a large amount of historical disposal data, thereby generating an initial instruction set corresponding to the current priority; S6.2. Obtain the initial instruction set of the disposal group generated by S6.1, use the initial instruction set of the disposal group as the input of the process execution module, obtain an output result set, and send the output result set to the information flow module; S6.3. Through the data information flow of the information flow module, a corresponding data relationship topology graph G=(V, E) is constructed for the result set; where V is a set of nodes, representing each data result in the result set, and E is a set of edges, representing the logical relationship or influence relationship between the results; ei,j is used to represent the edge between nodes vi and vj; S6.4. Based on the feedback value analysis, the correlation between the two results Ri,j=F(vi,vj) is obtained, and the function F is a function of the correlation between the two data, Ri,j The calculation of will be determined in combination with the feedback value; set the correlation threshold δ, if Ri,j is greater than δ, then add the edge ei,j to the graph, otherwise do not add it; S6.5, obtain the data relationship topology graph, build a prediction model through the graph neural network GNN and calculate the representation vector Z=GNN(G) of the node in the graph; the graph neural network will train and update the model parameters θ based on the local information and feedback values ​​of each node in the topology graph; the node representation vector Zv Used to describe the node status, Z=(zv1, zv2; zv|v|); the graph neural network can mine the deep structural features within the result set data by learning the relationship and feedback information between nodes in the topological graph, thereby providing a more comprehensive and accurate node representation vector for subsequent evaluation; S6.6, based on the representation vector Z of the graph node calculated in the previous step, transmit the information to the evaluation group, use the collected data for training, and generate a conditional decision model; S6.7, use the random forest algorithm to process the categorical variables in the data; the random forest algorithm analyzes and processes the categorical variables in the node representation vector by constructing multiple decision trees, and integrates the results of multiple decision trees to improve the accuracy and stability of the evaluation; S6.8, use the path analysis in the decision tree to generate the result evaluation probability P=DT (Zv) where DT Represents a decision tree trained on training samples. The function is used to predict the probability of an optimal outcome. S6.9: When the probability of a negative outcome is greater than 60%, the flexibility value in the original management plan is updated and synchronized with the execution entity in the first step. After adjusting the flexibility of the management plan, a new initial instruction set for the treatment group is generated. S6.10: Based on the outcome evaluation probabilities generated by the decision tree and the treatment group adjustment information, a pre-built dynamic flow prediction model is used to determine the new treatment outcome set prediction data for the treatment group. S6.11. This new set of treatment group results is sent to the treatment result update module via a data interface. S6.12. The results are verified and updated according to the module's defined logic. S6.13. After verification, the treatment result update module data is input into the designed classification model to generate a priority update model. S6.14. The classification model is used to annotate the data type and update the priorities of each plan in the management plan.

[0013] Preferably, the sub-steps of S7 are: S7.1. Obtaining evaluation result information based on the initial rule base; S7.2. Obtain the rule base evaluation result information and calculate the new weight value for each rule based on the original weight; S7.3. Update the weights of the original rule base according to the new weight values ​​of each rule to form an updated rule base; S7.4. Use the update library to judge the multi-dimensional indicators of the DevOps platform system and obtain the new indicators corresponding to the original indicators; S7.5. Calculate the adjusted value of each dimension indicator for the new and old indicators, and clarify the adjustment direction within the execution interval; S7.6. If the indicator adjustment value exceeds the preset range, the difference comparison algorithm is used to obtain the corresponding change gradient; S7.7. Determine the operating parameter values ​​of each dimension in the information processing mechanism through the change gradient of each dimension and the adjustment value of each dimension indicator. A DevOps platform monitoring cycle intelligent control system adopts the DevOps platform monitoring cycle intelligent control method.

[0014] The present invention can achieve the following beneficial effects: 1. The present invention obtains multi-dimensional indicator data streams during system operation and uses complex data processing processes, such as streaming computing, window functions, clustering technology, principal component analysis, etc., to accurately form a multi-dimensional vector that characterizes the system state, thereby accurately judging the system state. Compared with traditional monitoring methods, the present invention can more keenly capture the comprehensive impact of system load, service response time and various events on the system operation state, discover potential abnormal situations in advance, and significantly advance the warning time of system failures, effectively reducing the risk of business interruption due to system failures. Based on unsupervised learning to calculate the distance between the instant multi-dimensional vector and the center of the known system state category, it can automatically identify the similarity between the current state of the system and various typical states, realize intelligent classification and precise positioning of the system state, and provide a reliable basis for subsequent targeted treatment measures.

[0015] 2. The present invention uses a specially constructed event level model, combined with event information in multi-dimensional indicator data, to quantify the urgency of the event. Compared with traditional empirical judgment or simple threshold setting, this data-driven and model-based analysis method can more objectively and accurately evaluate the impact of events on system operations, and avoid misjudgment of the urgency of events due to human error or unreasonable rule setting. According to the numerical value of the urgency of the event, the corresponding alarm and processing mechanism can be triggered in a timely and effective manner to ensure that when faced with an emergency, operation and maintenance personnel can obtain key information and take action in the first time, greatly improving the system's response speed and processing capabilities to cope with emergencies.

[0016] 3. Compared with the traditional fixed scheduling strategy, the present invention can better adapt to various changes in the system operation process, improve resource utilization, avoid over-allocation or under-allocation of resources, and ensure that the system can operate efficiently under different load conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings and examples: Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0018] The preferred solution is Figure 1 As shown in the figure, a DevOps platform monitoring cycle intelligent control method includes the following steps: S1. During the operation of the DevOps platform, a multi-dimensional indicator data stream is obtained during the system operation. The data stream includes system load, service response time, and the specific type and number of events that occurred during the collection period. A multi-dimensional vector representing the system status is formed based on the multi-dimensional indicator data stream to obtain an instant multi-dimensional vector; specifically: S1.1. Based on the system load data stream, a stream computing engine is used to continuously receive system indicator data. The collected load data is processed by setting a window function to obtain the load average within the set time window. Assume that the system load data stream is (in represents a time series), the window function is , then the load average (in is the number of data points in the window, For the window time point).

[0019] S1.2, the fusion system load average and service delay information flow, the service delay information flow uses the sliding window method to calculate the delay within a fixed time range to obtain statistical data. Suppose the service delay information flow is , the sliding window size is , service delay statistics (in The first time point).

[0020] S1.3. Integrate the statistical data of the above system load average and service delay to form a two-dimensional information list to represent the comprehensive information of system load and service delay. Let the integrated two-dimensional list be .

[0021] S1.4. Obtain a two-dimensional list of all event types and statistical times that occurred within the window period. Suppose the event type set is (in is the total number of event types), the corresponding number of events is , then the two-dimensional list of event types and times .

[0022] S1.5. Associate the load and service delay information list with the event type and frequency list, count the number of events that occurred within the set time window, and count the number of different types of events within the set time to obtain a third list of the number of occurrences and types of each category, and merge them into a third list. Let the merged third list be .

[0023] S1.6. Use feature combination technology to integrate the three lists output in the previous step into a multi-dimensional data stream and transform it into a data form that can comprehensively reflect the multi-dimensional characteristics of the system. Let the feature combination function be , then the combined data .

[0024] S1.7. The data after feature combination processing is used as a set of high-dimensional feature vector data. Let the high-dimensional feature vector be (in is the feature vector dimension), For combined data The eigenvalue obtained after some mapping or transformation.

[0025] Such as Euclidean distance: , Where $x,y$ are two data points, Respectively eigenvalues) S1.8. Use time series related clustering technology to calculate time series correlation to explore the potential laws of data in time series. Suppose the time series clustering algorithm is , the clustering result is The specific time series clustering algorithm can adopt, for example, a density-based time series clustering algorithm. Its calculation process involves complex operations such as distance measurement between data points and density calculation, which will not be expanded in detail here.

[0026] S1.9. By constructing a time decay coefficient to calculate the correlation with the instantaneous time, if the time decay coefficient of the data at a certain moment is greater than a preset value at a certain time, it will be discarded to ensure the timeliness of the data. Let the time decay coefficient be (in is the attenuation factor, is the reference time point), if ( is the preset time attenuation coefficient threshold), the corresponding data is discarded.

[0027] S1.10. Get a list set with time information. Let the set be: (in is the number of data points, For the The time of the data point, is the corresponding eigenvector).

[0028] S1.11. Use the PCA principal component analysis method to reduce the dimensionality of the principal components of the time list set obtained above, extract key information, and reduce the data dimension. Let the PCA transformation matrix be , the principal component vector after dimensionality reduction is (in is the dimension after dimensionality reduction), , the PCA algorithm specifically calculates the covariance matrix (in is the mean of the eigenvectors), then solve the eigenvalues ​​and eigenvectors of the covariance matrix, and select the eigenvector corresponding to the larger eigenvalue to form the transformation matrix .

[0029] S1.12. Extract the vector group of the main components in the set, and obtain the main component with the largest proportion at each moment as the feature data to obtain the instant multidimensional vector. Let the instant multidimensional vector be: (in is the number of principal components finally selected), For The principal component values ​​selected according to the proportion.

[0030] S1.13. Calculate the instantaneous multidimensional vector in an unsupervised manner, calculate the Euclidean distance from the vector to the center of the known system state category, and determine the minimum distance from the center of all known vectors. Suppose the set of centers of the known system state categories is (in is the number of system state categories), vector Go to Category Center The Euclidean distance of: (in Category Center No. eigenvalues).

[0031] S1.14, determine which system state is most similar; if the distance is less than the state distance threshold, the calculated immediacy vector is classified as the most similar system state. Let the state distance threshold be ,like ,but Belong to category .

[0032] S1.15. Process the state similarity results output from multiple time periods through an adaptive deep neural network. Each system state is output to the deep network. If the network has a new output, the parameters of a certain system state are updated. The trained deep neural network is used as the basis for the decision of the instant state calculation. Let the deep neural network be NN, and the input be the state similarity results of multiple time periods. (in is the number of time periods), the output is the system status category prediction , the network training process adopts the back propagation algorithm to minimize the loss function (such as the cross entropy loss function: ,in is the true label, is the network prediction value) to update the network parameters.

[0033] S2. Through the event level model that has been built, the event information contained in the instant multi-dimensional vector of the DevOps platform is quantified, and the corresponding level parameters are matched to each event that occurs. The level parameters of each event are combined into a level vector, and the corresponding event urgency value is determined according to the specific value contained in the level vector; Supplementary explanation: The event level model is constructed based on historical event data, event impact range, business criticality and other factors. Each level parameter is determined by combining the analysis of a large amount of sample data with expert experience to ensure the accuracy and rationality of the event urgency assessment. Suppose the event level model is , the event information is , then the event level parameter , let the level vector be (in is the number of event types), the event urgency value According to The elements in are obtained through some mapping relationship (such as weighted summation, etc.), for example: (in For the weights for each event type).

[0034] S3. Using a preset rule base, perform logical operations on the system status information values ​​in the DevOps platform's real-time multidimensional vector and the event urgency values ​​in the level vector. The confidence vector of the corresponding system status information and the confidence vector of the corresponding event urgency information under the given conditions are obtained based on the rule base inference. Specifically: S3.1. Collect the multi-dimensional vectors and level vectors generated during the operation of the DevOps platform system. The information obtained includes the system status information value and the event urgency value, and build an information database. Assume that the system status information value is , the event urgency value is , the information database is .

[0035] S3.2. Perform normalization on the system status information values ​​in the database. According to the distribution of the event urgency levels, divide the event urgency values ​​into multiple intervals, and obtain the normalized database and the level value information of the divided intervals. Let the normalization function be , the normalized system status information value is , let the event urgency be divided into intervals , the corresponding level value is .

[0036] S3.3. Divide the information into two parts, one is the state information group and the other is the level information group. Perform clustering operation on the state information group. For each cluster center point generated by clustering, calculate the distance between other state information values ​​and it to obtain cluster data.

[0037] Suppose the state information group clustering algorithm is , the clustering result is , Let the cluster center point set be , status information value To cluster center Distance: (in for No. eigenvalues, for No. eigenvalues, is the feature dimension).

[0038] S3.4, preset a threshold and compare the system status information value distance in the database with the threshold. If it is higher than the threshold, the current value is determined to be a high-risk state, the corresponding status information group is a high-risk state, and the alarm instruction of the high-risk state is output to obtain the alarm instruction. Set the threshold to ,like , it is determined to be a high-risk state and an alarm instruction is output .

[0039] S3.5. Use the alarm command to query the numerical information of the level information group. Define the rules according to the rule base. If the queried level information group information matches the high urgency threshold of the level numerical information, then generate an alarm signal for the event, output the event alarm signal data, and obtain the information data of the emergency event. Transition description: After determining the high-risk state, the convolutional neural network classification model is used because it can learn the complex nonlinear relationship between the system status information value and the event urgency value. Through training with a large amount of sample data, it can more accurately determine the confidence matching relationship between the two, thereby providing a more reliable basis for subsequent decision-making. Assume that the convolutional neural network classification model is CNN, and the input is the combination of the system status information value and the event urgency value in the high-risk state. , the output is the confidence matching relationship , the convolutional neural network structure contains convolutional layers (such as convolution kernel Perform convolution operation on the input data ,in represents convolution operation), pooling layer (such as maximum pooling operation ), fully connected layers, etc., and adjust the parameters of each layer through training to optimize the model performance.

[0040] S3.6. Use the training set data to train a convolutional neural network classification model, set the network loss function and tune the network to obtain a mature network. The network inputs the emergency information data, determines the matching relationship between the output emergency information and the confidence level of the system status information, and generates a confidence level association list. Let the training set be (in is the training sample index, is the true confidence level label), the network loss function can use the mean square error loss function: (in is the number of training samples, For the The confidence level of the network output of each sample is obtained by adjusting the network parameters through optimization methods such as gradient descent algorithm to minimize the loss function. After obtaining a mature network, the input emergency information data is processed to obtain a confidence level association list. .

[0041] S4. If the DevOps platform system status information confidence vector contains more than a set number of system status information that is in an unhealthy state, a system status alarm instruction is issued to the system, and the fuzzy scheduling strategy mechanism is triggered based on the numerical judgment of the event urgency; specifically: S4.1. Collect various status parameters of the DevOps platform system during operation, construct a multi-dimensional system status information set, and obtain the system state vector after data cleaning. S=(s1, s2; si), where si represents the value of the i-th state information; S4.2. Compare the value si of each state information in the system state vector S with the pre-set healthy range, and calculate the number of unhealthy state information to obtain the number N of unhealthy states; determine whether the number N of unhealthy states is greater than the set threshold T. If N>T, output an alarm information set and generate an alarm instruction A; if N≤T, return to step S4.1 and continue to collect system state information. Let the healthy range be (in For the The health range of the status information), if , then the status information is unhealthy, count the number of unhealthy states: (in is the indicator function, and the conditions are met , otherwise ),like , output alarm information set and alarm instructions .

[0042] S4.3. Obtain the urgency values ​​of multiple events corresponding to the DevOps platform system alarm information set, normalize the urgency values, and obtain the normalized urgency value E=(e1, e2; ei), where ei represents the urgency value of the i-th event. Let the normalization function be ,but .

[0043] S4.4. According to the normalized urgency value E, a fuzzy logic-based algorithm is used to determine the triggered dispatch level L. The core rule of the fuzzy logic algorithm is to determine the final dispatch level through fuzzy reasoning based on the correspondence between different urgency value ranges and pre-set dispatch levels, while considering the comprehensive impact of the urgency of multiple events; the corresponding dispatch level is obtained. Assume that the fuzzy logic system has an input fuzzy set (corresponding to fuzzy sets with different urgency ranges), output fuzzy sets (corresponding to different scheduling levels of fuzzy sets), fuzzy inference rule table (indicates the correspondence between different input fuzzy sets and output fuzzy sets), by fuzzifying the input ; If the triangular membership function is used: , in Fuzzy set The parameters are then calculated based on the inference rule table and finally defuzzified; Such as the center of gravity method: , in is the output fuzzy set The representative value of .

[0044] S4.5. After determining the scheduling level L, multiple scheduling strategies that meet this level are selected from the pre-established scheduling strategy library. After the scheduling strategies are integrated, the integrated scheduling strategy G is obtained. The integrated scheduling strategy G is executed by the DevOps platform scheduling engine, and the resource allocation information data stream is sent through the communication module to determine the resource reallocation plan. Let the scheduling strategy library be (in is the number of strategies), select and schedule the level Matching policy subset , fusion scheduling strategy (in is a strategy fusion function, such as weighted summation or rule-based combination, etc.), the resource allocation information data flow is set to , according to the fusion scheduling strategy Sure The specific content of the resource allocation, such as the resource allocation amount, allocation object and other information, can be used to achieve resource reallocation.

[0045] S5. Obtain the expert experience processing rule sets corresponding to different types of events that have been set up, use fuzzy logic judgment technology combined with the expert experience processing rule sets to comprehensively calculate the real-time indicators and triggering events of the DevOps platform system, and determine the priority parameters after comprehensive calculation of the system resource status and urgency events; specifically: Construct an expert knowledge base logic block diagram: First, collect the DevOps platform historical operation and maintenance data, expert processing experience and other information, classify and analyze this information, and construct an event processing rule set and a corresponding historical resource status data set. Then, when an event is triggered, the real-time status information of the DevOps platform system is obtained synchronously based on multiple monitoring devices. The running status information matches the triggering event type and urgency, and the status information is assigned according to the preset fuzzy set interval. The fuzzy set and membership function parameters are called for event type query to calculate the membership of the assigned status information to obtain a number of running status membership data. According to the expert experience rules set in the expert knowledge base, the rule sets corresponding to different states in the same type of emergency event records are selected to obtain a targeted solution set. The membership is calculated jointly with the expert rule set to obtain a normalized membership value. The normalized value is weighted to obtain the event priority. After the fuzzy set is delineated according to the status information, the fuzzy set is matched with the indicator membership in the generated running data to obtain the calculated priority data. The priority data is then used to update the event record. According to the event urgency and event type, the range is compared from the calculated priority data to obtain the processing rule priority ranking. The ranking result is input into the existing equipment expert rule base to guide the generation of operation and maintenance tasks. Let the fuzzy set be (corresponding to fuzzy sets of different state information or event types), the membership function is (in is status information or event-related data), the running status membership data is set to (in is the state information or the number of event types), the expert experience rule set is (in is the number of rules), the event priority is set to (in is the weight coefficient), the priority order of processing rules is set to (in is the sorting function).

[0046] S6. Determine several corresponding flexible management plans based on the priority parameters, execute specific handling operation instruction sets according to the process in the flexible management plan, obtain the instruction handling results and feed them back to the information processing link for result evaluation; specifically: S6.1. Based on the plan group content and priority information, the support vector machine algorithm is used to generate the initial instruction set for each treatment group. The support vector machine algorithm uses the mapping relationship between priority parameters and plan group characteristics, and through training on a large amount of historical treatment data, it finds the optimal classification hyperplane, thereby generating the initial instruction set corresponding to the current priority, providing the starting operation basis for the subsequent treatment process. The specific method is as follows: Assume that the support vector machine model is SVM and the input is the priority parameter and characteristics of the emergency response team ,in is the number of features, and the output is the initial instruction set for the treatment group: ; The support vector machine constructs a decision function: ,in is the number of training samples, is the Lagrange multiplier, is the training sample label, is a kernel function, such as the radial basis kernel function , is the kernel parameter, is the bias term, which is determined by solving the optimization problem and to build the model.

[0047] S6.2. Obtain the initial instruction set of the disposal group generated in the previous step, use it as the input of the process execution module, obtain an output result set, and send the output result set to the information flow module; let the process execution module be , the output result set is .

[0048] S6.3. Build a corresponding data relationship topology diagram for the result set through the data information flow of the information flow module ;in, Is a collection of nodes, representing each data result in the result set, It is a set of edges, which represents the logical relationship or influence relationship between the results; Representation node and The edge between S6.4. Analyze the correlation between the two results based on the feedback value ,function is the function of the correlation between two data, The calculation will be determined in combination with the feedback value; set the correlation threshold ,like Greater than , then add an edge to the graph , otherwise not added; let the correlation function is some similarity metric function, such as cosine similarity: ; in represents the vector dot product, Represents the vector norm.

[0049] S6.5. Obtain the data relationship topology graph, build a prediction model through the graph neural network (GNN), and calculate the representation vectors of the nodes in the graph: ; The graph neural network will train and update the model parameters based on the local information and feedback values ​​of each node in the topology graph ; Node representation vector Used to describe the node status, ;Graph neural network uses message passing mechanism, such as nodes in each layer The update formula is: ; in For nodes The set of neighbor nodes of and For the The layer's weight matrix and bias vector, Is the activation function, such as the ReLU function ), after multiple layers of updating, the node representation vector is obtained .

[0050] S6.6. The representation vector of the graph node calculated in the previous step , transmit the information to the evaluation group, use the collected data to train and generate a conditional decision model; let the evaluation group model be , the training data is ,in is the decision result label of the corresponding node, which is obtained through training of a machine learning algorithm .

[0051] S6.7. Use the random forest algorithm to process the categorical variables in the data. The random forest algorithm constructs multiple decision trees, analyzes and processes the categorical variables in the node representation vector, and integrates the results of multiple decision trees to improve the accuracy and stability of the evaluation. The specific method is: Assume that the random forest is composed of A decision tree, , for the input node representation vector , each decision tree outputs a decision result ,The final result is obtained through voting or averaging to obtain a comprehensive decision result.

[0052] S6.8. Use the path analysis in the decision tree to generate the result evaluation probability: Here, DT represents a decision tree trained on training samples, and the function is used to predict the probability of the result being evaluated as excellent. The decision tree is constructed by recursively partitioning the training data. For example, the optimal partitioning attribute is selected based on indicators such as information gain or the Gini index, and the tree structure is constructed. Then, the evaluation probability is obtained based on the path of the input data in the tree.

[0053] S6.9. When the probability of judging that the result is not good is greater than 60%, the flexibility value in the original management plan is updated, and the updated information is synchronized to the execution entity of the first step. After adjusting the flexibility of the management plan, a new initial instruction set of the disposal group is generated; let the flexibility value be , the update rule is ,in The new disposal group initial instruction set is the adjustment amount determined based on the assessment results. Regenerate based on updated flexibility values ​​and other relevant information.

[0054] S6.10. Based on the result evaluation probability generated by the decision tree and the treatment group adjustment information, use the pre-built dynamic flow prediction model to determine the new treatment result set prediction data for the treatment group; let the dynamic flow prediction model be DFM and the input be the result evaluation probability and disposal group adjustment information , the output is the new treatment result set prediction data ,The dynamic flow prediction model can adopt a time series prediction model or ,sequence model based on deep learning, such as the LSTM model, which ,predicts future results by learning the sequence of historical ,processing results.

[0055] S6.11. Send the new treatment group result set to the treatment result update module through a data interface; S6.12. Verify and update the results according to the logic defined in the module; S6.13. After verification, update the module data of the treatment result and input it into the designed classification model to generate a priority update model; let the classification model be CM and the input be the verification result data. , the output is the priority update model ,The classification model can adopt support vector machine classifier, neural network classifier, ,etc., and build the model through classification learning of historical result data.

[0056] S6.14. The data is labeled by calling the classification model, and the priority of each plan in the management plan is updated.

[0057] S7. Determine the dynamically adjusted rule weights based on the rule base evaluation results. Re-evaluate the multi-dimensional indicators of the DevOps platform system based on the updated rule base to determine the adjustment direction and parameter values ​​for the next execution of the information processing mechanism. Specifically: S7.1. Obtain evaluation result information based on the initial rule base; let the evaluation result information be .

[0058] S7.2. Obtain the evaluation result information of the rule base and calculate the new weight value of each rule; suppose the rule base is ,in is the number of rules, the original weight is , the new weight calculation function is , then the new weight ,For example It can be a weighting function that adjusts the importance of different rules according to the evaluation results, such as ,in is the adjustment coefficient, For evaluation result information and rules Relevant parts.

[0059] S7.3, update the weight of the original rule base according to the new weight value of each rule to form an updated rule base; the updated rule base .

[0060] S7.4. Use the update library to judge the multi-dimensional indicators of the DevOps platform system and obtain the new indicators corresponding to the original indicators; let the system multi-dimensional indicators be , the judgment function is , then the new indicator ,For example It can be an inference or calculation function based on the rule base, which processes the indicators according to different rules and weights to obtain new indicator values.

[0061] S7.5. Calculate the adjusted value of each dimension indicator for the new and old indicators of each dimension, and clarify the adjustment direction under the execution interval; set the indicator adjustment value to ,in ,according to The positive or negative value determines the adjustment direction.

[0062] S7.6. If the indicator adjustment value exceeds the preset range, the difference comparison algorithm is used to obtain the corresponding change gradient; Difference comparison algorithm description: The difference comparison algorithm calculates the difference between the indicator adjustment value and the preset range boundary value, and then combines the historical trend of the indicator change and the data distribution characteristics to determine the corresponding change gradient, thereby quantifying the speed and direction of the indicator change. Assume the preset range is ,in For the The preset range of an indicator, if Beyond , change gradient ,in The time interval can be corrected by taking into account historical trends and data distribution characteristics, such as moving average or statistical distribution fitting.

[0063] S7.7. Determine the operating parameter values ​​of each dimension in the information processing mechanism by using the change gradients of each dimension and the adjustment values ​​of each dimension indicator; let the operating parameters of the information processing mechanism be ,according to and Sure The adjustment amount, such as (in and is the adjustment coefficient), and the adjusted operating parameter value is obtained . The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. Equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A DevOps platform monitoring cycle intelligent control method, characterized in that The following steps are involved: S1. During the operation of the DevOps platform, obtain the multi-dimensional indicator data stream in the system operation, which includes the system load, service response time, and the specific type and quantity of each event that occurred during the collection period. According to the multi-dimensional indicator data stream, a multi-dimensional vector representing the system status is formed to obtain an instant multi-dimensional vector; S2. Use the event level model that has been built to quantify the event information contained in the real-time multi-dimensional vector of the DevOps platform, match the corresponding level parameters to each event that occurs, and form a level vector with the level parameters of each event. Determine the urgency value of the corresponding event based on the specific values ​​contained in the level vector; S3, using a preset rule base to perform logical operations on the system status information value in the DevOps platform real-time multidimensional vector and the event urgency value in the level vector, and inferring the confidence vector of the corresponding system status information and the confidence vector of the corresponding event urgency information under the condition according to the rule base; S4. If more than a set number of system status information in the DevOps platform system status information confidence vector is in an unhealthy state, a system status alarm instruction is issued to the system, and the fuzzy scheduling strategy mechanism is triggered according to the numerical judgment of the event urgency; S5. Obtain the expert experience processing rule sets corresponding to different types of events that have been set up, use fuzzy logic judgment technology combined with expert experience processing rule sets, comprehensively calculate the real-time indicators and triggering events of the DevOps platform system, and determine the priority parameters after comprehensive calculation of the system resource status and urgency events; S6. Determine a number of corresponding flexible management plans through priority parameters, execute specific handling operation instruction sets according to the process in the flexible management plan, obtain instruction handling results and feed them back to the information processing link for result evaluation; S7. Determine the dynamically adjusted rule weight ratio based on the rule base evaluation result information, re-judge the multi-dimensional indicators of the DevOps platform system based on the updated rule base, and determine the adjustment direction and parameter values ​​when the information processing mechanism is executed next time.

2. According to claim 1, a DevOps platform monitoring cycle intelligent control method is characterized by: The sub-steps of S1 are: S1.

1. According to the system load data stream, a stream computing engine is used to continuously receive system indicator data, and the collected load data is processed by a set window function to obtain the load average within the set time window; Assume the system load data flow is: ,in represents a time series, and the window function is , The load average ,in is the number of data points in the window, For the window time point; S1.2, integrate the system load average and service delay information flow, calculate the delay within a fixed time range, and obtain statistical data; suppose the service delay information flow is , the sliding window size is , service delay statistics: ; in The first time point; S1.

3. Integrate the statistical data of the load mean and the service delay to form a two-dimensional information list to represent the comprehensive information of the system load and the service delay; let the integrated two-dimensional list be: ; S1.

4. Obtain a two-dimensional information list of all event types and statistical times occurring within the window period; suppose the event type set is: ,in is the total number of event types, and the corresponding number of events is , then the two-dimensional list of event types and times ; S1.

5. Associate the list of load and service delay information with the list of event types and times, count the number of events that occur within the set time window, and count the number of events of different categories within the set time to obtain a third list of the number of occurrences and types of each category, and merge them into a third list; Suppose the third list after fusion is: ; S1.

6. Use feature combination technology to integrate the three lists output in the previous step into a multi-dimensional data stream and convert them into a data form that can comprehensively reflect the multi-dimensional characteristics of the system; Assume that the feature combination function is , then the combined data: ; S1.

7. The data after feature combination processing is taken as a set of high-dimensional feature vector data; let the high-dimensional feature vector be ,in is the feature vector dimension, For combined data The feature value obtained after mapping or transformation; S1.

8. Use time series related clustering technology to calculate time series correlation to explore the potential laws of data in time series; suppose the time series clustering algorithm is The clustering result is ; S1.

9. By constructing a time decay coefficient to calculate the correlation with the instantaneous time, if the time decay coefficient of the data at a certain moment is greater than a preset value of a certain time, it is discarded to ensure the timeliness of the data; suppose the time decay coefficient is: ,in is the attenuation factor, is the reference time point, like , If it is the preset time attenuation coefficient threshold, the corresponding data is discarded; S1.

10. Get a list set with time information; let the set be: ; in is the number of data points, For the The time of the data point, is the corresponding eigenvector; S1.

11. Use the PCA principal component analysis method to reduce the dimension of the principal components of the time list set obtained above, extract key information, and reduce the data dimension; let the PCA transformation matrix be , the principal component vector after dimensionality reduction is: ; in is the dimension after dimensionality reduction, , the PCA algorithm specifically calculates the covariance matrix: ; in is the mean of the eigenvectors, and then solves the eigenvalues ​​and eigenvectors of the covariance matrix, and selects the eigenvectors corresponding to the larger eigenvalues ​​to form the transformation matrix ; S1.

12. Extract the vector group of the main components in the set, and obtain the main component with the largest proportion of components at each moment as the feature data to obtain the instant multidimensional vector; let the instant multidimensional vector be: ; in is the number of principal components finally selected, For The principal component values ​​selected according to the proportion; S1.

13. Calculate the instantaneous multidimensional vector in an unsupervised manner, calculate the Euclidean distance from the vector to the center of the known system state category, and determine the minimum distance from the center point of all known vectors; suppose the set of centers of the known system state categories is: ; in is the number of system state categories, vector Go to Category Center The Euclidean distance of: ; in Category Center No. eigenvalues; S1.14, determine which system state is most similar; if the distance between the immediacy vector and the center of each known system state category is less than the threshold value of the state distance, then the immediacy vector is classified into the most similar system state, and the calculated immediacy vector is classified into the most similar system state; let the state distance threshold be ,like ,but Belongs to category ; S1.

15. The state similarity results output in multiple time periods are processed by an adaptive deep neural network. Each system state is output to the deep network. If the network has a new output, the parameters of a certain system state are updated. The trained deep neural network is used as the decision basis for the immediate state calculation.

3. According to claim 1, a DevOps platform monitoring cycle intelligent control method is characterized by: The S2 method is: Assume the event level model is , the event information is , then the event level parameter , let the level vector be ,in is the number of event types and the urgency level of the event according to The elements in are obtained through mapping relationships.

4. According to claim 1, a DevOps platform monitoring cycle intelligent control method is characterized in that: The S3 sub-steps are: S3.

1. Collect multidimensional vectors and level vectors generated during the operation of the DevOps platform system. The information obtained includes system status information values ​​and event urgency values, and build an information database. Assume the system status information value is , the event urgency value is , the information database is ; S3.

2. Perform normalization on the system status information values ​​in the database, divide the event urgency values ​​into multiple intervals according to the level distribution of the event urgency, obtain the normalized database, and the level value information of the divided intervals; let the normalization function be , the normalized system status information value is: , Assume that the event urgency interval is , the corresponding level value is ; S3.3, divide the information into two parts, one is the state information group, and the other is the level information group. Perform clustering operation on the state information group, calculate the distance between other state information values ​​and each cluster center point generated by clustering, and obtain clustering data; suppose the state information group clustering algorithm is , the clustering results are: , Let the cluster center point set be , Status information value To cluster center Distance: ; in for No. eigenvalues, for No. eigenvalues, is the feature dimension; S3.4, preset a threshold and compare the numerical distance of the system status information in the database with the threshold. If it is higher than the threshold, the current value is determined to be a high-risk state, the corresponding status information group is a high-risk state, and an alarm instruction for the high-risk state is output to obtain an alarm instruction; set the threshold to ,like , it is judged as a high-risk state and an alarm instruction is output ; S3.

5. Use the alarm command to query the numerical information of the level information group. According to the rules defined in the rule base, if the queried level information group information matches the high urgency threshold of the level numerical information, an alarm signal of the event is generated, the event alarm signal data is output, and the information data of the emergency event is obtained; S3.

6. Use the training set data to train a convolutional neural network classification model, set the network loss function and tune the network, obtain a mature network, input the network information data of the emergency event, determine the matching relationship between the output emergency event information and the confidence level of the system status information, and generate a confidence level association list; let the training set be: ; in is the training sample index, For the true confidence level label, the network loss function uses the mean square error loss function: ; in is the number of training samples, For the The confidence level of the network output of each sample is calculated. The network parameters are adjusted through the gradient descent algorithm to minimize the loss function. After obtaining the mature network, the input emergency information data is processed to obtain the confidence level association list. .

5. According to claim 1, a DevOps platform monitoring cycle intelligent control method is characterized by: The sub-steps of S4 are: S4.

1. Collect various status parameters of the DevOps platform system during operation, build a multi-dimensional system status information set, and obtain the system status vector after data cleaning; S=(s1, s2; si), where si represents the value of the i-th state information; S4.

2. Compare the value si of each state information in the system state vector S with the preset healthy range, and obtain the number of unhealthy states N by counting the number of unhealthy state information; determine whether the number of unhealthy states N is greater than the set threshold T. If N>T, output an alarm information set and generate an alarm instruction A; if N≤T, return to step S4.1 to continue collecting system state information; suppose the healthy range is: ; in For the The healthy range of status information, if , then the status information is unhealthy, count the number of unhealthy states: ; in is the indicator function, satisfying the condition , otherwise ,like , output alarm information set And alarm instructions ; S4.

3. Obtain the urgency values ​​of multiple events corresponding to the DevOps platform system alarm information set, normalize the urgency values, and obtain the normalized urgency value E=(e1, e2; ei), where ei represents the urgency value of the ith event; let the normalization function be ,but ; S4.

4. According to the normalized urgency value E, a fuzzy logic-based algorithm is used to determine the triggered dispatch level L. The core rule of the fuzzy logic algorithm is to determine the final dispatch level through fuzzy reasoning based on the correspondence between different urgency value ranges and pre-set dispatch levels, while considering the comprehensive impact of the urgency of multiple events; Get the corresponding dispatch level; set up A fuzzy logic system has input fuzzy sets: , corresponding to fuzzy sets with different ranges of urgency, Output fuzzy set , corresponding to the fuzzy sets of different scheduling levels; Fuzzy inference rules table , represents the correspondence between different input fuzzy sets and output fuzzy sets, by fuzzifying the input ; S4.

5. After determining the scheduling level L, multiple scheduling strategies that meet the level are selected from the pre-established scheduling strategy library, and the scheduling strategies are integrated to obtain the integrated scheduling strategy G; the integrated scheduling strategy G is executed by the DevOps platform scheduling engine, and the resource allocation information data stream is sent through the communication module to determine the resource reallocation plan; Assume that the scheduling strategy library is ,in is the number of strategies, selection and scheduling level Matching policy subset: , Fusion scheduling strategy: ,in is the strategy fusion function, and the resource allocation information data flow is set to , according to the fusion scheduling strategy Sure The resource reallocation is achieved through the resource allocation amount and allocation object.

6. According to claim 1, a DevOps platform monitoring cycle intelligent control method is characterized by: The sub-steps of S5 are: S5.

1. Collect historical operation and maintenance data and expert processing experience of the DevOps platform, classify and analyze the historical operation and maintenance data and expert processing experience, and build a set of event processing rules and a corresponding set of historical resource status data; S5.

2. When an event is triggered, the real-time status information of the DevOps platform system is obtained synchronously based on multiple monitoring devices. The running status information matches the type and urgency of the triggering event, and the status information is assigned according to the preset fuzzy set interval. The fuzzy set and membership function parameters are called according to the event type query to calculate the membership of the assigned status information to obtain a number of running status membership data. According to the expert experience rules set in the expert knowledge base, the rule sets corresponding to different states in the same type of emergency event records are selected to obtain a targeted solution set. The membership is jointly calculated with the expert rule set to obtain a normalized membership value, and the normalized value is weighted to obtain the event priority. After the fuzzy set is delineated according to the status information, the fuzzy set is matched with the indicator membership in the generated running data to obtain the calculated priority data, and then the priority data is used to update the event record. According to the event urgency and event type, the range is compared from the calculated priority data to obtain the processing rule priority ranking, and the ranking result is input into the existing equipment expert rule base to guide the generation of operation and maintenance tasks.

7. The method for intelligently controlling the monitoring cycle of a DevOps platform according to claim 6 is characterized in that: In S5.2, let the fuzzy set be: , corresponding to fuzzy sets of different state information or event types; The membership function is: ,in is status information or event-related data; The running status membership data is set as: ,in It is the status information or the number of event types; The expert experience rule set is: ,in is the number of rules; Event priorities are set to: ,in is the weight coefficient; The processing rule priority order is set to: ,in is the sorting function.

8. The method for intelligently controlling the monitoring cycle of a DevOps platform according to claim 1, characterized in that: The sub-steps of S6 are: S6.

1. Based on the content and priority information of the emergency plan group, the support vector machine algorithm is used to generate the initial instruction set of each treatment group; the support vector machine algorithm finds the optimal classification hyperplane based on the mapping relationship between the priority parameters and the characteristics of the emergency plan group through learning and training of a large amount of historical treatment data, thereby generating an initial instruction set corresponding to the current priority; S6.

2. Obtain the initial instruction set of the treatment group generated by S6.1, use the initial instruction set of the treatment group as the input of the process execution module, obtain an output result set, and send the output result set to the information flow module; S6.

3. Through the data information flow of the information flow module, a corresponding data relationship topology graph G=(V, E) is constructed for the result set; wherein V is a set of nodes, representing each data result in the result set, and E is a set of edges, representing the logical relationship or influence relationship between the results; ei,j is used to represent the edge between nodes vi and vj; S6.

4. According to the feedback value analysis, the correlation between the two results Ri,j=F (vi, vj) is obtained, and the function F is a function of the correlation between the two data, Ri,j The calculation of will be determined in combination with the feedback value; set the correlation threshold δ, if Ri,j is greater than δ, then add the edge ei,j to the graph, otherwise do not add it; S6.5, obtain the data relationship topology graph, build a prediction model through the graph neural network GNN and calculate the representation vector Z=GNN(G) of the nodes in the graph; the graph neural network will train and update the model parameters θ according to the local information and feedback values ​​of each node in the topology graph; the node representation vector Zv Used to describe the node status, Z=(zv1, zv2; zv|v|); The graph neural network can mine the deep structural features of the result set data by learning the relationship and feedback information between nodes in the topological graph, so as to provide a more comprehensive and accurate node representation vector for subsequent evaluation; S6.

6. According to the representation vector Z of the graph node calculated in the previous step, the information is transmitted to the evaluation group, and the conditional decision model is generated by training with the collected data; S6.

7. The random forest algorithm is used to process the categorical variables in the data; The random forest algorithm analyzes and processes the categorical variables in the node representation vector by constructing multiple decision trees, and integrates the results of multiple decision trees to improve the accuracy and stability of the evaluation; S6.

8. Using the path analysis in the decision tree, the result evaluation probability P=DT (Zv) is generated, where DT Represents a decision tree trained on the training sample, and the function is used to predict the probability of the result evaluation being excellent; S6.9, when the probability of judging that the result is not excellent is greater than 60%, the flexibility value in the original management case is updated, and the updated information is synchronized to the execution subject of the first step. After adjusting the flexibility of the management case, a new initial instruction set of the treatment group is generated; S6.10, according to the result evaluation probability generated by the decision tree and the treatment group adjustment information, the pre-built dynamic flow prediction model is used to determine the new treatment result set prediction data of the treatment group; S6.

11. Send this new disposal group result set to the disposal result update module through a data interface; S6.

12. Verify and update the results according to the logic defined in the module; S6.

13. After verification, input the disposal result update module data into the designed classification model to generate a priority update model; S6.

14. Type-label the data by calling the classification model and update the priority of each plan in the management case.

9. The method for intelligently controlling the monitoring cycle of a DevOps platform according to claim 1, characterized in that: Sub-steps of S7: S7.

1. Obtaining evaluation result information according to the initial rule base; S7.2, obtain the rule base evaluation result information and calculate the original weight of each rule to obtain a new weight value; S7.3, updating the weight of the original rule base according to the new weight value of each rule to form an updated base; S7.

4. Use the update library to judge the multi-dimensional indicators of the DevOps platform system and obtain the new indicators corresponding to the original indicators; S7.

5. Calculate the adjustment value of each dimension indicator for the new and old indicators of each dimension and clarify the adjustment direction under the execution interval; S7.

6. If the index adjustment value exceeds the preset range, the difference comparison algorithm is used to obtain the corresponding change gradient; S7.

7. Determine the operating parameter values ​​of each dimension in the information processing mechanism through the change gradient of each dimension and the adjustment value of each dimension indicator.

10. A DevOps platform monitoring cycle intelligent control system, characterized by: A DevOps platform monitoring cycle intelligent control method according to any one of claims 1-9 is adopted.

Citation Information

Patent Citations

  • Molten iron balance expert scheduling method based on state evaluation

    CN115685922A

  • Workshop production equipment operation monitoring system based on distributed control

    CN118707914A

  • Data-driven limo safety monitoring system and method

    CN118711279A

  • State control method, device and equipment of monitoring terminal and storage medium

    CN118859731A

  • Cardiac bypass intraoperative decision support system based on fuzzy logic control

    CN118983055A

Cited By

  • Power distribution network control method considering communication quality

    CN120566427A

  • High-stability proprietary IP communication control method and system

    CN120602615A

  • Personalized reconstruction construction method and system for power grid Internet of Things platform

    CN120874290A

  • Personalized Transformation and Construction Methods and Systems for Power Grid Internet of Things Platforms

    CN120874290B

  • MOM software quality optimization method and system based on dynamic fuzzy scenario

    CN121233448A