A management method and system for a highly available policy engine

Through encrypted storage, version control and genetic algorithm scheduling, combined with abnormal detection and trend prediction, the problems of frequent policy updates, unbalanced resource utilization and untimely fault handling in quantitative policy management are solved, and the stable operation and security of the high-availability policy engine are achieved.

CN119961771BActive Publication Date: 2025-07-08GOING INT INNOVATIVE TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510449112.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the prior art, the management and resource scheduling of quantitative strategies are difficult to cope with the problems of frequent policy updates, unbalanced resource utilization, difficulty in detecting abnormalities, and untimely troubleshooting, resulting in unstable policy operation and insufficient security.

Method used

The high-availability policy engine management method is adopted to encrypt the storage and version control of policy files, dynamic scheduling is carried out in combination with genetic algorithms, and anomaly detection and trend prediction models are used to ensure the security and version consistency of policy files, and efficient resource utilization and rapid recovery of failures are achieved.

Benefits of technology

It improves the flexibility and robustness of policy file management, realizes balanced utilization of resources, ensures the stable operation of the policy engine under abnormal conditions, and provides continuous and reliable decision-making services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961771B_ABST
    Figure CN119961771B_ABST
Patent Text Reader

Abstract

The present application discloses a management method and system for a highly available policy engine, relating to the field of policy engines, including: storing the encrypted policy file on multiple data nodes through distributed storage, and storing the historical versions of the policy file; scheduling policy instances to run on target nodes according to the operation mode and scheduling constraints defined in the policy file; judging whether the policy instances are operating abnormally by collecting the operation status data of the scheduled and running policy instances, and if so, rescheduling the corresponding policy instances; collecting the performance parameters of the policy instances running on the target nodes; judging whether there are abnormalities in the policy instances according to the collected performance parameters through an anomaly detection algorithm, and predicting the usage trend of the policy instances for node resources in a future period of time. Aiming at the problem in the prior art that the policy engine cannot effectively cope with frequent policy updates, the present application improves the overall resource utilization rate of the system under the conditions of meeting node resources and policy constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of policy engines, and particularly to a management method and system for a highly available policy engine. Background Art

[0002] With the rapid development of the financial industry, quantitative trading has been widely applied globally due to its advantages of guiding investments with scientific methodologies and executing trades with strict discipline. The scale of funds managed by institutions such as quantitative private equity and hedge funds continues to grow, and the research teams behind fund managers are also expanding. The number of quantitative strategies they develop has increased geometrically. These quantitative strategies form a complete systematic method from data collection, factor construction, signal generation, weight optimization, order execution, etc., and quickly capture market opportunities through programmed trading means, which is more efficient and objective than manual trading.

[0003] However, in the actual operation of quantitative strategies, especially in the live trading environment, it is a huge challenge to ensure that each strategy can run long-term, stably, and securely. Quantitative strategies go through a complete life cycle from development, testing to live trading, and many management issues will be involved during this period. First of all, the core of quantitative strategies is the strategy files developed by researchers, which contain important information such as the logic, parameters, and configurations of the strategies. How to effectively manage a large number of strategy files, ensure their security and version consistency, and avoid the risk of strategy confusion caused by human errors is a common problem faced by quantitative trading institutions. Secondly, the demand for computing resources varies greatly among different strategies. When hundreds of strategies run concurrently, how to reasonably schedule and allocate hardware resources such as CPU, memory, and network, not only to maximize resource utilization but also to avoid interference between strategies caused by resource contention is extremely technically challenging. Thirdly, in a complex market environment, the operation of quantitative strategies is often unpredictable and requires real-time monitoring and evaluation. The traditional manual inspection mode is difficult to handle hundreds of strategies, and a highly automated intelligent operation and maintenance means is urgently needed. Finally, quantitative trading runs 24 / 7 without interruption, and any abnormality of a strategy may bring inestimable losses to the fund. Therefore, a perfect abnormal emergency handling mechanism must be established to ensure that faults are detected and isolated in the first time when a strategy fails, and can be quickly restored and restarted.

[0004] In summary, the large-scale operation of quantitative strategies places high demands on the infrastructure. The traditional manual management mode is no longer sustainable, and a new generation of highly available policy engine systems is urgently needed. Summary of the Invention

[0005] In view of the problem that the policy engine in the prior art cannot effectively handle frequent policy updates, this application provides a management method and system for a highly available policy engine. By performing version control on policy files and using a genetic algorithm for dynamic scheduling of policy instances, the overall resource utilization rate of the system is improved under the conditions of meeting node resources and policy constraints.

[0006] One aspect of this application provides a management method for a highly available policy engine, including: S1, encrypting the policy file, storing the encrypted policy file on multiple data nodes through distributed storage, and storing the historical versions of the policy file; S2, scheduling the policy instance to run on the target node according to the running mode and scheduling constraints defined in the policy file; S3, by collecting the running status data of the policy instance scheduled to run in step S2, determining whether the policy instance is running abnormally. If it is abnormal, reschedule the corresponding policy instance; S4, collecting the performance parameters of the policy instance running on the target node in step S2, where the performance parameters include the node resource usage and the running status of the policy instance; S5, according to the collected performance parameters, determining whether the policy instance has an abnormality through an anomaly detection algorithm, and predicting the usage trend of the policy instance for node resources in the next period of time. In this application, the highly available policy engine is a software system that can continuously, stably, and efficiently execute policies. By adopting a series of technical means such as fault tolerance, redundancy, monitoring, and scheduling, it ensures that the policy can still operate normally under various abnormal conditions, thereby providing continuous and reliable decision-making services for the upper-layer services.

[0007] Among them, the scheduling constraint refers to a series of restrictive conditions that need to be met when allocating a policy instance to a target running node to ensure that the policy instance can run normally and meet the requirements of the business and the system. Common scheduling constraints include: Resource constraint: The CPU, memory, disk, etc. resources required by the policy instance during operation shall not exceed the available capacity of the node; Affinity constraint: Some policy instances need to be deployed on the same node or rack as a specific service or instance to reduce communication latency; Mutual exclusion constraint: Some policy instances cannot be deployed on the same node as a specific service or instance to avoid mutual interference; Data locality constraint: Some policy instances need to be deployed as close as possible to their data sources on the same node to improve data access efficiency; The scheduling system needs to embed these constraints into the optimization objective and evaluation index during the scheduling decision to find an optimal deployment plan that meets the constraint conditions.

[0008] Anomaly detection algorithms refer to a class of algorithms that automatically identify abnormal behaviors that deviate from normal patterns by analyzing the behavioral data of systems or components. In the high-availability policy engine scenario, anomaly detection algorithms are mainly used for anomaly detection of the running status and performance parameters of policy instances, to detect faults in a timely manner and take recovery measures. The basic assumption of anomaly detection algorithms is that the data samples of normal behaviors account for the majority, the data samples of abnormal behaviors are sparse, and are significantly different from normal samples. Through data mining and machine learning techniques, anomaly detection algorithms can automatically learn the pattern features of normal behaviors and identify anomalies based on this benchmark. In this application, anomaly detection algorithms can adopt statistical methods, machine learning methods such as K-Means clustering: assuming that normal data clusters, and abnormal data is far from all cluster centers; LOF local outlier factor: measuring the density difference between a data point and its surrounding neighbors, and points with significantly lower density than neighbors are regarded as anomalies; OC-SVM one-class support vector machine: mapping the data to a high-dimensional space, enclosing most of the normal data within a hypersphere, and points outside the sphere are regarded as anomalies, etc. Combinatorial algorithms can also be adopted.

[0009] Further, S1, encrypt the policy file, and store the encrypted policy file on multiple data nodes through distributed storage, and store the historical versions of the policy file, including: S11, encrypt the policy file using an asymmetric encryption algorithm to generate the encrypted policy file; S12, use the Reed-Solomon data redundancy algorithm to divide the encrypted policy file into blocks and add check blocks; S13, store the data blocks and check blocks on multiple data nodes respectively. When some data nodes fail, recover the original policy file through the data blocks and check blocks on other nodes; S14, each time the policy file is modified, encrypt and store the new version of the policy file, and establish a parent-child link with the previous version to construct a version tree of the policy file; S15, when a policy instance requests to load the policy file, first obtain the latest version of the policy file from the data node. If the access to the latest version fails, traverse the version tree to obtain the nearest historical version.

[0010] Among them, Reed-Solomon (RS) is an error-correcting code algorithm based on the data block level. By attaching several redundant check blocks outside the original data blocks, even if some data blocks are lost or damaged, the original data can be restored through the remaining intact data blocks and check blocks. The basic principle of the RS algorithm is as follows: The original data is divided into n data blocks, and then m check blocks are generated based on the data blocks to form an (n, m) RS erasure code. When the total number of lost or damaged data blocks and check blocks does not exceed m, the lost data blocks can be reconstructed by solving a system of linear equations. The RS algorithm has the advantages of simple encoding and decoding calculations and strong error-correcting ability, and is widely used in data redundancy and fault tolerance in fields such as disk arrays, distributed storage, and satellite communications. In the management method of the highly available policy engine, the RS algorithm is used to perform data redundancy on the encrypted policy file, which can tolerate the failure of some data nodes and improve the reliability of policy file storage. A policy file refers to a digital file that records policy definitions, configuration parameters, execution logic, etc. in a specific format. The policy file organizes each part of the policy in a structured and modular manner, facilitating the management, update, sharing, and reuse of the policy. Common policy file formats include XML, YAML, JSON, etc. In the management method of the highly available policy engine, the policy file is the core resource for policy management and execution, and its reliable storage and version management directly affect the high availability of the policy engine.

[0011] A parent-child link is a way to organize the change relationships between multiple versions of a file in a version control system. Each version of the file is regarded as a node, and each change generates a new sub-version. The sub-version node points to the parent-version node through a link, and the parent-version node also records the links of all sub-version nodes, ultimately forming a tree-structured version library.

[0012] Furthermore, in S2, according to the running mode and scheduling constraints defined in the policy file, the policy instance is scheduled to run on the target node, including: using an algorithm such as the genetic algorithm to solve the optimal scheduling plan, and realizing the allocation and deployment of the policy instance on the target node under the conditions of meeting the node resources and policy constraints; S21, reading the running mode and scheduling constraint parameters defined in the policy file as the input of the scheduling algorithm; S22, collecting the resource usage of each current target node as the input of the scheduling algorithm; where the resource usage includes CPU and memory usage; S23, using the genetic algorithm to solve the optimal scheduling plan, and allocating and deploying the policy instance to the target node under the conditions of meeting the node resources and scheduling constraints; S24, according to the optimal scheduling plan, distributing the policy instance to the target node through container orchestration Kubernetes. Each target node starts or updates the corresponding policy instance according to the received instruction and allocates computing resources.

[0013] Further, in S23, a genetic algorithm is used to solve the optimal scheduling scheme, including: setting the parameters of the genetic algorithm according to the number of policy instances and the number of candidate target nodes, where the parameters include the initial population size, crossover probability, mutation probability, and termination generation; adopting a priority-based encoding strategy to encode the deployment mapping relationship from policy instances to target nodes as an N-tuple, which serves as the chromosome of a scheduling scheme; where N is the number of policy instances, and each component of the tuple represents the number of the target node to which a policy instance is deployed; randomly generating chromosomes to form an initial population; calculating the fitness of the scheduling scheme corresponding to each chromosome; adopting a tournament selection algorithm to randomly select a group of chromosomes from the current population, and taking the pair of chromosomes with the highest fitness as the parents; performing multi-point crossover recombination on the parent chromosomes according to the set crossover probability to generate new offspring chromosomes; performing a mutation operation on the generated offspring chromosomes according to the mutation probability, randomly selecting one or more gene positions, and randomly replacing the values of the selected gene positions with the numbers of other candidate nodes to obtain the mutated offspring chromosomes; adding the offspring chromosomes obtained by crossover recombination and the mutated offspring chromosomes to the next generation population, repeating the crossover and mutation operations until the size of the new population reaches the initial population size; repeating the above iterations until the number of iterations reaches the set termination generation, and decoding the chromosome with the highest fitness in the final population to obtain the optimal scheduling scheme.

[0014] Further, adopting a priority-based encoding strategy to encode the deployment mapping relationship from policy instances to target nodes as an N-tuple, which serves as the chromosome of a scheduling scheme, includes: setting the priority of policy instances according to the execution frequency of policy instances; sorting the tuple components corresponding to policy instances according to the set priority; determining the value range of each gene position in the chromosome according to the number of candidate target nodes; adopting gray coding to map the number of each candidate target node to a binary string, which serves as the gene encoding of the corresponding node; sequentially splicing the gene encodings of all candidate target nodes in the order of tuple components to form a complete chromosome; where each gene position of the chromosome corresponds to the deployment node of a policy instance.

[0015] Among them, in the scheduling optimization method of the high-availability policy engine, an N-tuple is used to represent a scheduling scheme, that is, the deployment mapping relationship from policy instances to target nodes. This N-tuple can be expressed as (P1, P2,......, PN), where N is the total number of policy instances, and Pi represents the target node number where the i-th policy instance is deployed. Each element Pi in the N-tuple is called a tuple component, representing the scheduling decision of a policy instance. For example, there are 5 policy instances {P1, P2, P3, P4, P5}, and the candidate target node numbers are {1, 2, 3}. A possible scheduling scheme is (2, 1, 3, 2, 1), which means: policy instance P1 is deployed to node 2; policy instance P2 is deployed to node 1; policy instance P3 is deployed to node 3; policy instance P4 is deployed to node 2; policy instance P5 is deployed to node 1; here, 2, 1, 3, 2, 1 are the 5 tuple components in the N-tuple of this scheduling scheme, representing the target node allocation decisions of 5 policy instances. By rearranging and combining the tuple components, different scheduling schemes of policy instances can be traversed and searched to find the optimal solution for deployment constraints and optimization objectives. This is the basic idea of scheduling optimization. Tuple components are an abstract representation method for scheduling decision variables in the scheduling optimization model.

[0016] Gray Code is a binary coding method, and its characteristic is that there is only one different binary bit between two adjacent coding values. Through Gray Code, a discrete value range can be mapped to a set of binary codes. In the scheduling optimization method of the high-availability policy engine, gray coding the numbers of candidate target nodes can better meet the gene mutation requirements of the genetic algorithm and improve the search efficiency. There is only one different binary bit between adjacent values of Gray Code. For example, there is only the highest bit changing from 0 to 1 between 1 (01) and 2 (11). This coding method has better continuity and smoothness and is more suitable for the crossover and mutation operations of the genetic algorithm. If standard binary coding is directly adopted, there may be multiple different bits between adjacent codes. For example, between 1 (01) and 2 (10), which will lead to a large jump after crossover and mutation and is not conducive to retaining excellent genes. In the scheduling optimization method of the high-availability policy engine, the Gray Code of candidate node numbers is concatenated in the order of N-tuple components to form a complete chromosome coding. The genetic algorithm searches for an optimized scheduling scheme by performing operations such as selection, crossover, and mutation on these gene bits. As a coding strategy of the genetic algorithm, Gray Code can better guide the search direction and accelerate the optimization solution process.

[0017] Furthermore, calculate the fitness of the scheduling scheme corresponding to each chromosome: , where F is the fitness function value, and its value range is [0, 1]. The larger the value, the better the comprehensive performance of the scheduling scheme; They are the weight coefficients of three optimization objectives: load balancing, criticality, and failure rate, respectively, and ; The specific values can be dynamically adjusted according to the importance of each objective during system operation; G is the Gini coefficient, which is used to measure the imbalance of load distribution among nodes. The calculation formula is: , where N is the number of candidate target nodes, and are the load amounts of the i-th node and the j-th node, respectively; SLA is the service level agreement satisfaction rate of critical policy instances. The calculation formula is: , where S is the number of critical policy instances, is the weight coefficient of the s-th critical policy instance, is the SLA satisfaction indicator variable of the s-th critical policy instance; P is the failure rate of the scheduling scheme. The calculation formula is: , where N is the number of candidate target nodes, is the failure rate of the n-th node, is the proportion of the number of policy instances deployed on the n-th node to the total number of instances.

[0018] Furthermore, in S5, based on the collected performance parameters, use the anomaly detection algorithm to determine whether there are anomalies in the policy instances and predict the future usage trend of the policy instances on node resources, including: S51, based on the performance parameters collected in step S4, construct the resource usage time series of the policy instances. The time series includes indicators such as CPU usage rate, memory usage rate, and response time; S52, use the One Class SVM model to perform anomaly detection on the time series of each indicator; among them, the OneClass SVM model determines whether the newly collected data points are abnormal through feature vector mapping and boundary learning; S53, based on the anomaly detection results of multiple indicators, use the anomaly state machine to determine whether there are anomalies in the policy instances; S54, use the LSTM neural network to predict the future usage trend of resources. The input of the LSTM neural network is the time series of multiple indicators, and the output is the predicted value for a future period; S55, based on the predicted node resource usage trend, determine the load situation of the policy instances in the future period.

[0019] Among them, One Class SVM (Support Vector Machine) is a commonly used unsupervised anomaly detection model, especially suitable for scenarios where there is only normal data and lack of anomaly data labels. Different from the traditional binary classification SVM, One Class SVM only uses data of one class to train the model, learns an optimal hypersphere, tightly wraps the normal data points within the hypersphere, and isolates the anomaly data points outside the hypersphere. The basic assumption of One Class SVM is that normal data accounts for the majority and is relatively compactly distributed, while anomaly data accounts for the minority and has a significant difference from the distribution of normal data. In the anomaly detection of the highly available policy engine, for the time series of indicators such as CPU, memory, and response time of each policy instance, a One Class SVM model is trained to determine whether the newly collected indicator data points are abnormal. One Class SVM can adaptively learn the distribution characteristics of indicator data in the normal state and has good detection generalization ability for unknown anomaly patterns.

[0020] The anomaly state machine is a finite state automaton model used to describe the transition process of the system state driven by events. In the field of anomaly detection, it is mainly used to comprehensively judge and characterize the overall anomaly state of the system. The system starts from the normal state. When certain anomaly events or conditions are met, the transition to the anomaly state will be triggered. After entering the anomaly state, if the recovery conditions are met, it will transfer back to the normal state, thus forming a state transition diagram. In the anomaly detection of policy instances, taking the One Class SVM anomaly detection results of multiple indicators (such as CPU, memory, etc.) as input events to drive the state transition of the anomaly state machine can judge the overall anomaly situation of the policy instance. The anomaly state machine synthesizes the local anomaly signals in multiple dimensions into a global anomaly judgment through the state transition logic, and distinguishes the severity and duration of the anomaly, which can more comprehensively and accurately reflect the health level of the complex system and provide a reference basis for anomaly location and decision-making.

[0021] Further, in S52, an One Class SVM model is used to perform anomaly detection on the time series of each metric, including: segmenting the time series data by a sliding window, extracting the statistical features of each window to form a statistical feature vector; obtaining the historical sequence data during the normal operation of the policy instance to construct a training data set; using the training data set to train the One Class SVM model; wherein, the One Class SVM model maps the feature vectors in the training data set from the original space to a high-dimensional space through a kernel function; in the high-dimensional space, the One Class SVM model obtains the smallest hypersphere containing a number of training samples greater than a threshold, and uses the smallest hypersphere as the decision boundary for distinguishing normal samples and abnormal samples; inputting the newly collected time series data into the trained One Class SVM model to obtain the mapped feature vector; calculating the distance from the mapped feature vector to the smallest hypersphere; comparing the calculated distance with a preset threshold, and if it exceeds the preset threshold, determining the time series data of the corresponding time window as abnormal.

[0022] Further, in S53, according to the anomaly detection results of multiple metrics, an anomaly state machine is used to determine whether there is an anomaly in the policy instance, including: obtaining the results of the anomaly detection of the time series data of multiple metrics, and the results include the anomaly status of each metric in each time window; wherein, the anomaly status is represented by a boolean value, 1 represents anomaly, and 0 represents normal; according to the anomaly status of each metric in each time window, calculating the comprehensive anomaly degree of the policy instance in the corresponding time window; arranging the comprehensive anomaly degrees of the policy instance in multiple consecutive time windows in chronological order to form an anomaly degree time series; using the pre-constructed anomaly state machine model to determine whether there is an anomaly in the obtained anomaly degree time series, wherein, the anomaly state machine model is constructed by a directed graph, the nodes in the graph represent the anomaly status of the policy instance, and the edges represent the transition conditions and probabilities between the anomaly statuses.

[0023] Another aspect of the present application further provides a management system for a highly available policy engine, which is used to execute a management method for a highly available policy engine of the present application.

[0024] Compared with the prior art, the advantages of the present application are as follows:

[0025] By encrypting the storage, version controlling, and data redundancy of the policy file, it effectively solves the problems such as inconsistent storage and version chaos of the policy file caused by frequent policy updates in the prior art. Managing the policy file by version not only ensures the atomicity of policy release but also enables rapid rollback to the previous stable version in case of policy update failure, greatly improving the flexibility and robustness of policy file management. At the same time, technologies such as asymmetric encryption and Reed-Solomon erasure code are used to encrypt and provide data redundancy for the stored policy file, ensuring that confidential policies cannot be illegally stolen and the original policy can still be restored in case of failure of some storage nodes, thus fundamentally guaranteeing the security and reliability of policy storage.

[0026] Using a genetic algorithm for dynamic scheduling of policy instances can achieve better load balancing while meeting node resource and policy constraint conditions, improving the overall resource utilization efficiency of the system. Most traditional policy instance scheduling methods use static rules or simple dynamic algorithms such as polling. Such methods are difficult to adapt to the complex and changeable policy operation environment and are prone to uneven phenomena where some nodes are overloaded while others are idle. The genetic algorithm constructs a fitness function with load balancing, key policy SLA satisfaction rate, and solution failure rate as optimization objectives by simulating the process of biological evolution, and abstracts the deployment mapping relationship from policy instances to target nodes as a chromosome using a priority-based coding method. Through iterative evolution, it searches for the optimal solution in the population. Since the genetic algorithm has the ability to globally optimize, it can quickly search for a scheduling solution that meets multiple constraints and continuously evolves and optimizes with the dynamic changes in the policy operation environment, thus maximizing the load balancing of each node while ensuring the performance of key policies and making full use of limited computing resources.

[0027] Combining anomaly detection and trend prediction models, it monitors the running state of policy instances in real time and anticipates the resource usage in the future for a period of time, providing a reliable basis for the dynamic scheduling decision of policy instances. The One Class SVM model determines the abnormal state of time series data through technical means such as feature vector mapping and boundary learning. Compared with simple threshold comparison methods, One Class SVM can adaptively learn the behavior patterns during the normal operation of policy instances and has better detection and tolerance capabilities for abnormal behaviors caused by various unknown faults. The LSTM model establishes a non-linear correlation between system performance indicators through deep learning algorithms and mines the evolution trends contained in time series data, accurately predicting the demand changes of policy instances for system resources within a period of time. The combination of anomaly detection and trend prediction can not only detect current running faults in a timely manner but also plan ahead, providing optimization space and time dimensions for dynamic scheduling, which is beneficial for the policy engine to continuously and stably provide services externally.

[0028] The abnormal state machine model comprehensively analyzes the anomaly detection results of multiple indicators, evaluates the health status of the policy instance from a global perspective, and overcomes the semantic gap between the anomaly of a single indicator and the anomaly of the policy instance as a whole. Traditional operation and maintenance monitoring often sets an abnormal threshold for a single indicator such as CPU and memory, and issues an alarm once an indicator exceeds the threshold, which is prone to produce many false positives and missed positives. Whether a policy instance is abnormal cannot be simply judged by whether one or two indicators are abnormal, but must be considered from the perspective of its behavioral evolution process. The abnormal state machine model abstracts the law of the evolution of the abnormal state of the policy instance over time through a directed graph, and combines the anomaly detection snapshots of each indicator to accurately depict the complete process of the policy instance from local anomaly to global failure, and can quickly locate and isolate faults based on this, thereby minimizing the impact of abnormal policy instances on the system and improving the system's fault tolerance. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0030] Figure 1 It is an exemplary flowchart of a management method of a high-availability policy engine according to some embodiments of the present application. DETAILED DESCRIPTION

[0031] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0032] like Figure 1 As shown, the policy file is encrypted, and the encrypted policy file is stored in a distributed manner on multiple data nodes, and the historical version of the policy file is stored; the policy instance is scheduled to run on the target node according to the operation mode and scheduling constraints defined in the policy file; by collecting the operation status data of the scheduled policy instance, it is determined whether the policy instance is running abnormally, and if it is abnormal, the corresponding policy instance is rescheduled; the performance parameters of the policy instance running on the target node are collected, and the performance parameters include the node resource usage and the operation status of the policy instance; based on the collected performance parameters, it is determined whether the policy instance is abnormal through the anomaly detection algorithm, and the usage trend of the node resources by the policy instance in the future period is predicted.

[0033] Specifically, in S1, the policy file is encrypted, and the encrypted policy file is distributed and stored on multiple data nodes through a distributed system, and the historical versions of the policy file are stored, including: S11, the policy file is encrypted using an asymmetric encryption algorithm to generate an encrypted policy file; First, a pair of RSA keys, including a public key and a private key, is generated for each owner of the policy file. The public key is used to encrypt the policy file, and the private key is used to decrypt the encrypted policy file. The key length of the RSA key can be flexibly set according to the requirements of the security level, such as 1024 bits, 2048 bits, etc. After the policy file is created or modified, the public key of the policy owner is used to encrypt the policy file. Different padding schemes, such as PKCS#1 v1.5, OAEP, etc., can be used in the encryption process to enhance the encryption strength. The encrypted policy file will be stored and distributed in ciphertext form. When a policy instance needs to read and execute the policy file, it is first necessary to verify the identity of the user to which the policy instance belongs. Only users who have passed identity authentication and authorization can use their corresponding private keys to decrypt the encrypted policy file to obtain the policy logic and parameter definitions in plaintext form. Unauthorized users cannot decrypt the sensitive content in it without knowing the private key even if they obtain the encrypted policy file. Through the asymmetric encryption mechanism, access control of the plaintext content of the policy file is achieved. The policy file is always in ciphertext state during storage and transmission, and it will only be decrypted into plaintext in the policy instance authorized by the policy owner when the policy is actually executed. This can minimize the exposure range and time of the plaintext policy file and effectively prevent the leakage of confidential information such as policy logic and parameters.

[0034] In terms of data reliability in S12, the Reed-Solomon data redundancy algorithm is introduced. This algorithm divides the encrypted policy file into blocks and calculates and generates check blocks based on the block data. The data blocks and check blocks are distributed and stored on multiple data nodes. When some data nodes fail, the complete original policy file can be restored using the data blocks and check blocks on other nodes. For example, if RS(12, 4) coding is used, the policy file is divided into 12 data blocks and 4 check blocks, for a total of 16 blocks. These 16 blocks can be stored on 16 data nodes. If any 4 nodes fail, the original policy file can still be reconstructed based on the remaining 12 blocks, greatly improving the fault tolerance against node failures.

[0035] S13. Store the data blocks and parity blocks on multiple data nodes respectively. When some data nodes fail, the original policy file can be restored through the data blocks and parity blocks on other nodes; S14. In terms of version management, every time the policy file is modified, the new version of the policy file will be encrypted and stored, and a parent-child link will be established with the previous version to build a version tree of the policy file. The version tree records the change history of the policy file and forms a series of file snapshots representing the states at different time points. When the policy file needs to be loaded, the policy instance first obtains the latest version. If the access to the latest version fails, it will backtrack and search for the nearest available historical version on the version tree. Through the version tree and the backtracking mechanism, both the latest modification of the policy file can be obtained at any time, and it can be quickly restored to the previous stable version when the update fails, effectively coping with various abnormal situations that may be caused by frequent policy updates.

[0036] S15. In terms of policy operation management, according to the different operation modes of the policy file in actual applications, it is divided into batch mode and overlay mode. In batch mode, a policy file will generate multiple policy instances, and each instance is responsible for processing different data nodes to achieve parallel computing; in overlay mode, a policy file will generate only one policy instance, and this instance is responsible for processing all data nodes to achieve serial computing. Different operation modes have different scheduling constraint conditions such as resource requirements, concurrency limits, and node affinity requirements. At the same time, it is also necessary to collect the operation status data of the policy instance in real time, including key indicators such as the CPU and memory utilization rates, and the processing latency of events, to accurately evaluate the execution of the policy instance.

[0037] S2. Schedule the policy instance to run on the target node according to the operation mode and scheduling constraints defined in the policy file, including: using algorithms such as genetic algorithms to solve the optimal scheduling plan, and realizing the allocation and deployment of the policy instance on the target node under the condition of meeting the node resources and policy constraints. The node resource usage includes CPU utilization rate, memory usage, and disk I / O, which reflect the load of the target node; the operation status of the policy instance includes the event processing delay of the policy and the order execution situation, which reflect the execution effect of the policy instance.

[0038] S21. Read the running mode and scheduling constraint parameters defined in the policy file as the input of the scheduling algorithm. The policy file is usually written in structured markup languages such as JSON and YAML, and declares the attributes of the policy in the form of key-value pairs. The scheduling algorithm reads the policy file and parses the parameters related to the running mode and scheduling constraints, mainly including: Running mode: Specify the parallel deployment method of the policy instance. Common values include: "Batch": Batch mode, where multiple policy instances are generated from one policy file to process data on different nodes respectively; "Overlay": Overlay mode, where one policy instance is generated from one policy file to process data on all nodes. For example, if "runMode": "Batch" is defined in a certain policy file, it means that the policy runs in batch mode. The scheduling algorithm extracts this parameter and decides whether to deploy policy instances in parallel on multiple nodes accordingly. Resource requirements: The computing resources required for the operation of the policy instance, including: "cpu": The number of CPU cores required, such as "cpu": 2; "memory": The required memory space, such as "memory": "4GB"; "gpu": The number of GPU cards required, such as "gpu": 1. The scheduling algorithm filters the target nodes that meet the resource requirements of the policy instance based on these parameter values to avoid the failure of policy instance startup due to insufficient resources. Concurrency limit: To avoid overloading the policy instance, the maximum number of events or transactions that can be processed concurrently can be set. For example, "max Concurrency": 1000 means that the policy instance can process at most 1000 events simultaneously; "maxOrders": 5000 means that the policy instance can process at most 5000 transactions simultaneously. The scheduling algorithm refers to these limits to control the number of policy instances on a single node and prevent the node from being overloaded. Target node affinity: Limit the range of nodes where the policy instance is deployed, supporting two modes: white list and black list: "node Affinity": "in: node1, node2" means that the policy instance can only be deployed on node1 or node2; "node Affinity": "not in: node3" means that the policy instance cannot be deployed on node3. The scheduling algorithm reads the affinity expression and performs matching and filtering among the candidate nodes.

[0039] S22. Collect the resource usage of each current target node as the input of the scheduling algorithm. The scheduling algorithm needs to keep track of the load levels of each target node in real time to reasonably allocate policy instances. Regularly collect the resource usage of nodes through the monitoring component, focusing on two metrics: CPU usage rate and memory usage. CPU usage rate = (1 - idle CPU time / total CPU time) * 100%. In this embodiment, through the / proc / stat file in Linux, collect the total CPU time and idle time within a time window (such as 1 minute): cpu 841001 10116 889278 6010608 22267 0 17761 0 0 0; where the first column is the total CPU time and the fourth column is the idle CPU time, both in 1 / 100 seconds. Suppose the total times collected twice are T1 = 7763264 and T2 = 7824728 respectively, and the idle times are I1 = 6003511 and I2 = 6028102 respectively. Then the CPU usage rate within this time window is: CPU usage rate = (1 - (I2 - I1) / (T2 - T1)) * 100% = 60.33%. Memory usage = used memory / total memory * 100%. In this embodiment, through the free command in Linux, collect the memory usage snapshot of the node. The total column is the total memory and the used column is the used memory, both in KB. Suppose the total memory is 8GB = 8388608KB and the used memory is 1.6GB = 1677722KB. Then the memory usage is: Memory usage = 1677722 / 8388608 * 100% = 20.00%. See Table 1, Node Memory Status Table for details.

[0040] Table 1 Node Memory Status Table

[0041] Total Memory Used Free Shared Buffers / Cache Available 8169348 1680404 4066836 495620 2422108 5812632

[0042] The monitoring component periodically collects the CPU usage rate and memory usage data of all target nodes and constructs time-series data in the following format. Each row records the resource usage of a certain node at a certain moment, all represented as floating-point decimals. See Table 2, Node Resource Time-Series Table for details.

[0043] Table 2 Node Resource Time-Series Table

[0044] Timestamp Node CPU Usage Memory Usage 1620620705 node1 0.6033 0.2000 1620620705 node2 0.4288 0.3155 1620620705 node3 0.7500 0.8360

[0045] The scheduling algorithm queries the resource usage of nodes within a certain time range (such as the last 5 minutes) from the time-series database, takes the average value of each metric, and forms the following data snapshot as the input for scheduling decisions. See Table 3, Scheduling Decision Input Table for details.

[0046] Table 3 Scheduling Decision Input Table

[0047] Node Average CPU Usage Average Memory Usage node1 0.5933 0.1988 node1 0.4012 0.3008 node1 0.7707 0.8298

[0048] The running parameters defined in the policy file and the node resource usage are pre - processed through steps such as collection, parsing, and aggregation, and are transformed into structured data that can be directly used by the scheduling algorithm. Based on this, the scheduling algorithm evaluates the available resources of each node, weighs the concurrency limit and affinity constraints of policy instances, and dynamically generates an optimal node deployment plan.

[0049] S23. Use the genetic algorithm to solve the optimal scheduling plan. Under the conditions of meeting node resources and scheduling constraints, allocate and deploy policy instances to target nodes, including: Setting genetic algorithm parameters: According to the number N of policy instances to be scheduled and the number M of candidate target nodes, set the following parameters: Initial population size P: Generally, the value ranges from 20 to 100. While ensuring population diversity, control the computational overhead, and P can be set to 50; Crossover probability Pc: Generally, the value ranges from 0.4 to 0.9. The larger the value, the faster the population updates, and Pc can be set to 0.8; Mutation probability Pm: Generally, the value ranges from 0.001 to 0.1. The larger the value, the higher the population diversity, and Pm can be set to 0.05; Termination generation T: Generally, the value ranges from 100 to 500. The larger the value, the higher the quality of the solution, but the slower the convergence speed, and T can be set to 200. For example, if there are currently 100 policy instances and 10 candidate nodes, then N = 100, M = 10, and the genetic algorithm parameters can be set as: P = 50, Pc = 0.8, Pm = 0.05, T = 200.

[0050] Chromosome encoding: Adopt a priority - based encoding strategy. Encode a scheduling plan as an integer tuple of length N. Each component of the tuple takes values from 1 to M, indicating which node the corresponding numbered policy instance is deployed on. The encoding process is as follows: Calculate the priority of policy instances according to their historical execution frequencies. The higher the frequency, the greater the priority. For example: policy1: 15 times / minute, priority = 5; policy2: 10 times / minute, priority = 4; policy3: 8 times / minute, priority = 3;......; Sort the policy instances from high to low according to priority, with the higher - priority ones in front of the tuple, forming the encoding order of policy instances, such as: policy1 > policy2 > policy3 >......; According to the number M of candidate nodes, determine that the value range of each gene position is from 1 to M, and use gray coding to map the node numbers to binary strings of equal length as the gene coding of the nodes. For example, when M = 10, see Table 4, the gene coding table of policy instances.

[0051] Table 4 Gene coding table of policy instances

[0052] Policy Instance Priority Deployment Node Gene Encoding Policy1 5 Node1 000 Policy2 4 Node2 001 Policy3 3 Node3 011

[0053] According to the encoding order of policy instances, the gene encodings of each node are spliced in sequence to form a complete chromosome. For example: policy1-node1, policy2-node2, policy3-node3. Mapping to a chromosome: 000001011 means that policy1 is deployed on node1, policy2 is deployed on node2, and policy3 is deployed on node3, and so on.

[0054] Initialize the population: Randomly generate P chromosomes as the initial population of the genetic algorithm, as shown in Table 5, the example table of the chromosomes in the initial population.

[0055] Table 5 Example table of the chromosomes in the initial population

[0056] Chromosome Number Chromosome Example 1 010,011,000,111,001 2 001,010,011,000,110 3 011,001,010,100,000 4 111,000,001,011,010

[0057] Fitness evaluation: Calculate the fitness value of the scheduling scheme corresponding to each chromosome. The fitness function is: , and the value range is [0, 1]. The larger the value, the better the scheme. Among them: are the weight coefficients of the three optimization goals of load balancing, critical policy SLA, and failure rate respectively: They can be dynamically adjusted according to the importance of the three goals during operation. For example, increase during high load, and increase during the business peak, and increase when nodes are frequently abnormal, and always ensure . G is the Gini coefficient, which measures the imbalance of the load distribution among nodes: , where N is the number of nodes, and are the loads of the i-th and j-th nodes, which can be represented by CPU, memory occupancy rate, or the number of policy instances. The value range of G is [0, 1], and the smaller the value, the more balanced the load. For example, for the scheduling scheme represented by a certain chromosome, the number of policy instances of each node is: 18, 9, 11, 20, 2, then: G = 1 / 10 * (|18 - 9| + |18 - 11| + |18 - 20| + |18 - 2| + |9 - 11| + |9 - 20| + |9 - 2| + |11 - 20| + |11 - 2| + |20 - 2|) = 0.64; it shows that the load distribution of this scheme among nodes is not balanced enough. SLA is the SLA satisfaction rate of critical policy instances: ; where S is the total number of critical policy instances, is the weight of the s-th critical instance, which is generally proportional to its priority, It is the SLA satisfaction indication variable for the s-th critical instance, which is 1 when deployed on a node that meets the performance requirements and 0 otherwise. The SLA value ranges from [0, 1], and the larger it is, the better the performance requirements of the critical policy instance are met. For a certain chromosome, the weights of 3 critical policy instances are 0.5, 0.3, and 0.2 respectively, and the SLAs of 2 instances are satisfied. Then: SLA = (0.5×1 + 0.3×1 + 0.2×0) / 3 = 0.6, indicating that the performance guarantee of the critical policy instance by this solution is relatively good. P is the failure rate of the scheduling solution: ; where N is the total number of nodes, is the failure rate of the n-th node, is the proportion of the number of policy instances deployed on the n-th node. The value of P ranges from [0, 1], and the smaller it is, the higher the reliability of the solution. For a certain chromosome, assuming the failure rates of 3 nodes are 0.01, 0.05, and 0.001 respectively, and the proportion of the number of instances they deploy is 0.5, 0.3, and 0.2, then: , indicating that the overall failure rate of this solution is relatively low.

[0058] Genetic operators: Use operators such as tournament selection, multi-point crossover, and random mutation to evolve the current population to generate the next generation population. Use tournament selection, randomly select K individuals each time, and take the pair with the highest fitness as the parent. Generally, K = 3 - 5. Use multi-point crossover, randomly select several nodes as crossover points, align the parent chromosomes at the crossover points, and then exchange the corresponding gene segments to recombine new offspring chromosomes. For example: parent1: 000|011|010; parent2: 001|100|111; child1: 000|100|010; child2: 001|011|111. Use random mutation, randomly select some gene positions of the chromosome with probability Pm and replace them with other values to introduce new search directions. For example: original: 011000101010; mutated: 010000001010. Use the elitist retention strategy, directly copy the top E individuals with the highest fitness in the current population to the next generation. Generally, E = 1 - 5 to avoid losing the optimal solution.

[0059] Iteration termination: Repeatedly execute genetic operators such as selection, crossover, and mutation to continuously generate new populations until the number of iterations reaches T, or when there is no obvious improvement in the optimal fitness value of the population for several consecutive generations, the algorithm terminates. Result decoding: Decode the chromosome with the highest fitness in the final population to obtain the optimal scheduling plan. For example, best chromosome: 010 000011 001 ......; decoded as: policy1 -> node2; policy2 -> node1; policy3 -> node4; policy4 -> node2; thus completing the optimal scheduling solution process based on the genetic algorithm.

[0060] Execution and distribution of the S24 scheduling plan: Convert the optimal scheduling plan into Kubernetes resource configuration files such as Deployment and StatefulSet, and define attributes such as the image, resource requests, and environment variables of each policy instance. Then, use the kubectl apply command to distribute the configuration file to the Kubernetes cluster. The scheduler on the Master node allocates Pods to the target worker nodes, and the container engine (such as Docker) on the nodes pulls the image and starts the policy container instance. When the policy is updated or the node load changes, the optimal scheduling plan can be recalculated through the above steps, compare the similarities and differences between the new and old plans, and only perform migrations or scaling on the policy instances in the different parts to achieve incremental rolling updates and reduce scheduling overhead. The genetic algorithm can efficiently solve complex policy instance scheduling problems. Through means such as coding mapping, fitness evaluation, and genetic evolution, it searches for the optimal solution globally and combines cloud-native technology stacks such as Kubernetes to achieve flexible execution and dynamic update of the scheduling plan, thereby building an intelligent policy instance management system to strongly support the parallelization and large-scale operation of policies.

[0061] S3. By collecting the running status data of the policy instances scheduled and run in step S2, determine whether the policy instances are running abnormally. If so, reschedule the corresponding policy instances. Specifically, define the abnormal judgment metrics: According to the characteristics of the policy instances, set a group of key metrics reflecting their running status, such as: Liveness Probe: Such as whether the process exists and whether the port is available. Once the detection fails, the instance is regarded as abnormal; Readiness Probe: Such as whether the response latency times out and whether the dependent services are reachable. When the readiness probe fails, the instance cannot process new requests; Event Latency: The average time for the instance to process an event. If it is significantly higher than the historical level, it indicates that the instance may have a performance bottleneck; Resource Usage: The proportion of resources such as CPU, memory, and disk occupied by the instance on the node. When it exceeds the set threshold, the instance is regarded as abnormal. State data collection: Through monitoring components such as Prometheus, regularly collect the running status metric data of each policy instance, such as: The detection results (success or failure) of the liveness probe and the readiness probe; The distribution statistics of event latency (TP50 / TP90 / TP99, etc.); Resource metrics such as CPU usage, memory usage, and disk usage. The state data is stored in the monitoring database in the form of time series. Each data point contains the metric name, timestamp, label (policy name + instance ID), and value.

[0062] Abnormal instance judgment: Compare the collected instance state data with the thresholds of the abnormal judgment metrics to identify abnormal policy instances, such as: Instances where the liveness probe or the readiness probe fails continuously for multiple times (such as 3 times); Instances where the TP90 value of event latency continuously exceeds 100 ms; Instances where the CPU or memory usage continuously exceeds 80%. Abnormal judgment can be achieved through preset Prometheus alarm rules. When the conditional expression of the alarm rule is continuously satisfied for a specified duration, the corresponding alarm is triggered. The alarm information contains the label of the abnormal instance. Rescheduling of abnormal instances: After receiving the abnormal alarm, the scheduling system queries and deletes the Kubernetes resource object (such as Pod) of the abnormal instance according to the policy name and instance ID in the alarm, and releases the node resources occupied by it. Then the scheduling system re-executes the scheduling algorithm in S2, calculates a new scheduling plan excluding the faulty node, and creates a new policy instance to fill the deleted abnormal instance to ensure that the replica count of the policy meets the requirements. The newly created policy instance will fill the vacancy of the abnormal instance and be scheduled to a healthy node to restore the high availability state of the entire policy.

[0063] S4. Collect the performance parameters of the policy instances running on the target nodes in step S2. The performance parameters include the node resource usage and the running status of the policy instances. Specifically, for the node resource usage, key metrics such as CPU, memory, disk, and network are concerned: node_cpu_usage: the CPU usage rate of the node; node_cpu_load1 / 5 / 15: the average CPU load of the node in the last 1 minute, 5 minutes, and 15 minutes; node_memory_usage: the memory usage rate of the node; node_memory_available: the available memory amount of the node; node_file system_usage: the disk partition usage rate of the node; node_network_receive / transmit_bytes: the number of received / transmitted bytes of the node network.

[0064] For the running status of the policy instances, key metrics related to service quality and faults are concerned: instance_event_latency: the latency distribution of instance event processing; instance_order_count: the total number of orders successfully executed by the instance; instance_order_amount: the total amount of orders successfully executed by the instance; instance_failure_count: the number of orders with failed or abnormal execution of the instance; instance_failure_rate: the execution failure rate of the instance; instance_message_backlog: the number of messages to be processed by the instance.

[0065] By continuously collecting and monitoring these two types of performance parameters, the scheduling system can real-time understand the running conditions such as the load, capacity, and SLA of all nodes and policy instances in the cluster, and accordingly evaluate the actual effect of the scheduling scheme: evaluate whether the node resource allocation is reasonable, whether the load is balanced, and whether there is overload or waste; evaluate whether the performance guarantee of key policy instances is in place and whether there is a risk of SLA violation; evaluate the fault level and fault mode of nodes and instances and whether there are high-availability vulnerabilities.

[0066] The scheduling system can associate these performance parameters with the scheduling model and continuously optimize the scheduling algorithm in a data-driven manner: using load data to guide the scheduling strategy to dynamically balance the load, performance, and stability goals; using SLA data to assist scheduling decisions and enhance the identification and isolation of long-tail transactions; using fault data to evaluate the health of nodes and instances and improve the fault tolerance of the scheduling scheme; collecting the effectiveness data of different scheduling decisions and continuously iteratively optimizing the scheduling model with AI algorithms such as reinforcement learning. The running state of the policy instance and the node resource usage are important inputs and feedback for the scheduling system. By collecting and analyzing these performance parameters, the scheduling system can perceive the business semantics, adapt to business changes, and ultimately make the scheduling scheme continuously evolve towards the optimal direction, providing a solid data foundation for the stable and efficient operation of the quantitative policy.

[0067] S5. According to the collected performance parameters, use the anomaly detection algorithm to determine whether there is an anomaly in the policy instance and predict the future trend of the policy instance's use of node resources for a period of time, including: S51. According to the performance parameters collected in S4, construct time series of multiple resource usage metrics for each policy instance. Common metrics include: CPU usage rate: the percentage of the node's CPU resources occupied by the instance, indicating the computing load intensity of the instance. Memory usage rate: the percentage of the node's memory resources occupied by the instance, indicating the storage load intensity of the instance. Response time: the average time taken for the instance to process requests, indicating the service quality level of the instance. Each data point in the time series contains the metric name, timestamp, policy instance label, and value.

[0068] S52. Use the One Class SVM model to perform anomaly detection on the time series of each metric; for the time series of each resource usage metric, perform anomaly detection independently using the One Class SVM model: data segmentation and feature extraction: set a sliding window of a fixed size (such as 1 hour), and segment the time series into several segments according to the window; extract features from each time window to obtain a set of statistical feature values, such as mean, standard deviation, quantiles, etc., to form a feature vector; the window sliding step size is 10 minutes, so a one-hour time series is extracted into 6 feature vectors.

[0069] Construction of normal sample training set: Select historical data of several days when the policy instance runs normally as the training set; Extract the time series of the training set into multiple feature vectors. Training of One Class SVM model: Train the One Class SVM model with the feature vectors of the training set; One Class SVM maps the feature vectors from the original space to a high-dimensional space through a kernel function (such as the RBF kernel); In the high-dimensional space, One Class SVM obtains the smallest hypersphere containing most of the training samples as the boundary of normal samples. Anomaly detection: Extract the new time series to be detected into feature vectors and input them into the model; One Class SVM maps it to a high-dimensional space and calculates the distance d to the smallest hypersphere; If d is greater than a given threshold (generally the 90th percentile of the sample distance), it is determined as an anomaly.

[0070] S53, According to the anomaly detection results of multiple indicators, judge whether there is an anomaly in the policy instance through the anomaly state machine; including: Conduct anomaly detection on multiple key indicators such as the CPU, memory, and response time of the instance respectively (such as using One Class SVM), and the detection result is the anomaly state of each indicator in each time window, represented by the boolean value 0 (normal) or 1 (anomaly). The detection results of a certain instance in 6 consecutive windows are: cpu_anomaly: [0, 0, 1, 1, 1, 0]; memory_anomaly: [0, 0, 0, 1, 0, 0]; latency_anomaly: [0, 1, 1, 1, 0, 0].

[0071] Calculate the comprehensive anomaly degree of the instance in each time window, and count the anomaly ratio of each indicator as the comprehensive anomaly degree of this time window. Anomaly degree = number of anomaly indicators / total number of indicators. The anomaly degrees of the 6 time windows in the above example are: [0 / 3, 1 / 3, 2 / 3, 3 / 3, 1 / 3, 0 / 3] = [0, 0.33, 0.67, 1, 0.33, 0]. Obtain the anomaly degree time series, arrange the anomaly degrees of the instance in consecutive multiple time windows in chronological order to form an anomaly degree time series reflecting the overall anomaly state change trend of the instance. The anomaly degree time series in the above example is: anomaly_score: [0, 0.33, 0.67, 1, 0.33, 0].

[0072] Use an anomaly state machine to determine whether the anomaly degree sequence is abnormal; pre-construct an anomaly state machine model to depict the evolution process of an instance from normal to abnormal. The state machine is represented by a directed graph, where nodes represent the abnormal states of the instance and edges represent state transition conditions and probabilities. Common abnormal states include: normal, suspicious, anomalous, faulty, etc. The transition conditions on the edges are usually related to the anomaly degree threshold. Match the anomaly degree sequence with the state machine, calculate the most likely state transition path, and output the final abnormal state of the instance. In the above example, the matching process of the anomaly degree sequence [0, 0.33, 0.67, 1, 0.33, 0] in the state machine is as follows:

[0073] Normal--anomaly_score≥0.33-->Suspicious;Suspicious--anomaly_score≥0.67-->Anomalous;

[0074] Anomalous--anomaly_score≤0.33-->Suspicious;

[0075] Suspicious--anomaly_score≤0.33-->Normal;

[0076] Normal--0.33-->Suspicious--0.67-->Anomalous--1.0-->Anomalous--0.33-->Suspicious--0.0-->Normal. The final state is Normal, indicating that the instance only had a temporary anomaly once and has returned to normal. The anomaly state machine flexibly represents different stages and paths in the anomaly evolution process in the form of a graph model. By matching the anomaly degree sequence in the state machine, different anomaly patterns can be adaptively identified, and non-abrupt anomalies such as progressive, intermittent, and periodic anomalies can be discovered.

[0077] S54. Use an LSTM (Long Short-Term Memory) network to predict the future resource usage trend of the policy instance: construct a multivariate time series: construct the indicators such as the CPU, memory, and response time of the instance into a multi-column time series data; each row is the value of each indicator at a time point, as shown in Table VI, the data table of the policy instance resource usage time series.

[0078] Table VI Data Table of the Policy Instance Resource Usage Time Series

[0079] Timestamp CPU Usage Memory Usage Response Time (ms) 1620620600 0.2 0.6 20 1620620660 0.25 0.62 22 1620620720 0.18 0.65 25

[0080] Data preprocessing: Interpolate and fill missing values, and smooth outlier values; normalize the data of each indicator to eliminate the influence of dimensions; perform differencing on the data at a fixed step size to extract the relative change trend of the data. Construct training set and test set: Select a part of the historical data (such as 70%) as the training set, and the remaining part as the test set; use the sliding window method to divide the time series data into multiple input-output pairs: (X, y); each X contains historical observations of a fixed length (such as 12), and y contains the target prediction values for several (such as 3) future time points.

[0081] Build an LSTM prediction model: The input layer of the model receives the feature matrix X of the multivariate time series; several LSTM layers and Dense fully connected layers are used in the middle layer to extract time series features and fit the non-linear trend; the output layer predicts the multi-indicator values y for several future time points; use loss functions such as mean squared error, and adopt the Adam optimizer to train the model parameters. Model evaluation and optimization: Evaluate the prediction accuracy of the model on the test set, and calculate evaluation indicators such as MAE, MAPE, and RMSE; optimize by adjusting hyperparameters and trying different model structures. Model inference and prediction: Use the trained LSTM model to input the latest time series data and predict the future resource usage trend.

[0082] S55, after using LSTM prediction to obtain the resource usage trend of the policy instance in the future for a period of time (such as 1 hour), the future load situation can be judged: compare the predicted values of CPU and memory usage rates with the current values to judge whether there is a significant growth trend and whether the growth rate exceeds the threshold (such as 30%); analyze the change trend of the response time to judge whether there is a deterioration trend and whether the severity affects the service quality; infer whether the instance will continue to be in a high-load state in the future and whether the high-load duration exceeds the threshold (such as 30 minutes); accordingly, divide the future load state of the instance into: idle, normal, busy, and overloaded levels.

Claims

1. A management method for a highly available policy engine, characterized in that, Including: S1, encrypt the policy file, store the encrypted policy file on multiple data nodes through distributed storage, and store the historical versions of the policy file; S2, schedule the policy instance to run on the target node according to the running mode and scheduling constraints defined in the policy file; S3, by collecting the running status data of the policy instance scheduled and running in step S2, determine whether the policy instance runs abnormally. If abnormal, reschedule the corresponding policy instance; S4, collect the performance parameters of the policy instance running on the target node in step S2, where the performance parameters include the node resource usage and the running status of the policy instance; S5, according to the collected performance parameters, determine whether the policy instance has an abnormality through an anomaly detection algorithm, and predict the usage trend of the policy instance for node resources in the future for a period of time; S2, schedule the policy instance to run on the target node according to the running mode and scheduling constraints defined in the policy file, including: S21, read the running mode and scheduling constraint parameters defined in the policy file as the input of the scheduling algorithm; S22, collect the resource usage of each current target node as the input of the scheduling algorithm; where the resource usage includes CPU and memory usage; S23, use the genetic algorithm to solve the optimal scheduling plan, and allocate and deploy the policy instance to the target node under the condition of meeting the node resources and scheduling constraints; S24, according to the optimal scheduling plan, distribute the policy instance to the target node through container orchestration Kubernetes. Each target node starts or updates the corresponding policy instance through the container engine according to the received instruction and allocates computing resources.

2. The management method of the highly available policy engine according to claim 1, characterized in that: S1, encrypt the policy file, store the encrypted policy file on multiple data nodes through distributed storage, including: S11, use an asymmetric encryption algorithm to encrypt the policy file to generate the encrypted policy file; S12, use the data redundancy algorithm Reed-Solomon to divide the encrypted policy file into blocks and add check blocks; S13, store the data blocks and check blocks on multiple data nodes respectively. When some data nodes fail, restore the original policy file through the data blocks and check blocks on other nodes; S14, each time the policy file is modified, encrypt and store the new version of the policy file, and establish a parent-child link with the previous version to construct the version tree of the policy file; S15, when the policy instance requests to load the policy file, first obtain the latest version of the policy file from the data node. If the access to the latest version fails, traverse the version tree to obtain the nearest historical version.

3. The management method of the highly available policy engine according to claim 1, characterized in that: S23, use the genetic algorithm to solve the optimal scheduling plan, including: Set the parameters of the genetic algorithm according to the number of policy instances and the number of candidate target nodes. The parameters include the initial population size, crossover probability, mutation probability, and termination generation; Adopt a priority-based encoding strategy, and encode the deployment mapping relationship from policy instances to target nodes as an N-tuple, which serves as the chromosome of a scheduling scheme; where N is the number of policy instances, and each component of the tuple represents the number of the target node where a policy instance is deployed. Randomly generate chromosomes to form an initial population. Calculate the fitness of the scheduling scheme corresponding to each chromosome. Adopt a tournament selection algorithm, randomly select a group of chromosomes from the current population, and use the pair of chromosomes with the highest fitness as the parents; perform multi-point crossover recombination on the parent chromosomes according to the set crossover probability to generate new offspring chromosomes. Perform mutation operations on the generated offspring chromosomes according to the mutation probability, randomly select one or more gene positions, and randomly replace the values of the selected gene positions with the numbers of other candidate nodes to obtain the mutated offspring chromosomes. Add the offspring chromosomes obtained by crossover recombination and the mutated offspring chromosomes to the next-generation population, and repeat the crossover and mutation operations until the scale of the new population reaches the initial population size. Repeat the above iterations until the number of iterations reaches the set termination generation, and decode the chromosome with the highest fitness in the final population to obtain the optimal scheduling scheme.

4. The management method of the highly available policy engine according to claim 3, characterized in that: Adopt a priority-based encoding strategy, and encode the deployment mapping relationship from policy instances to target nodes as an N-tuple, which serves as the chromosome of a scheduling scheme, including: Set the priorities of policy instances according to the execution frequencies of the policy instances. Sort the tuple components corresponding to the policy instances according to the set priorities. Determine the value range of each gene position in the chromosome according to the number of candidate target nodes. Adopt gray coding to map the number of each candidate target node to a binary string, which serves as the gene coding of the corresponding node. Sequentially splice the gene codings of all candidate target nodes in the order of tuple components to form a complete chromosome; where each gene position of the chromosome corresponds to the deployment node of a policy instance.

5. The management method of the highly available policy engine according to claim 4, characterized in that: Calculate the fitness of the scheduling scheme corresponding to each chromosome: Where F is the fitness function value, and the value range is [0, 1]; They are the weight coefficients for the three optimization objectives of load balancing, criticality, and failure rate respectively; G is the Gini coefficient, and the calculation formula is: where N is the number of candidate target nodes, and are the loads of the i-th node and the j-th node, respectively; SLA is the satisfaction rate of the service level agreement of critical policy instances, and the calculation formula is: where S is the number of key policy instances, is the weight coefficient of the s-th key policy instance, is the SLA satisfaction indication variable of the s-th key policy instance; P is the failure rate of the scheduling scheme, and the calculation formula is: where N is the number of candidate target nodes, is the failure rate of the nth node, is the ratio of the number of policy instances deployed on the nth node to the total number of instances.

6. The management method of the highly available policy engine according to claim 1, characterized in that: S5. Determine whether there are abnormalities in policy instances through an anomaly detection algorithm, and predict the usage trend of policy instances for node resources in a future period, including: S51. According to the performance parameters collected in step S4, construct a resource usage time series of policy instances, and the time series includes CPU usage rate, memory usage rate, and response time indicators. S52. Use the OneClass SVM model to perform anomaly detection on the time series of each metric. Among them, the OneClass SVM model determines whether a newly collected data point is abnormal through feature vector mapping and boundary learning; S53. According to the anomaly detection results of multiple metrics, use an anomaly state machine to determine whether there is an anomaly in the policy instance; S54. Use an LSTM neural network to predict the resource usage trend in the future for a period of time. The input of the LSTM neural network is the time series of multiple metrics, and the output is the predicted value for a period of time in the future; S55. According to the predicted node resource usage trend, determine the load situation of the policy instance in the future for a period of time.

7. The management method of the highly available policy engine according to claim 6, characterized in that: S52. Use the OneClass SVM model to perform anomaly detection on the time series of each metric, including: Segment the time series data by a sliding window, extract the statistical features of each window, and form a statistical feature vector; Obtain the historical sequence data when the policy instance is running normally, and construct a training data set; Use the training data set to train the OneClass SVM model. Among them, the OneClass SVM model maps the feature vectors in the training data set from the original space to a high-dimensional space through a kernel function; in the high-dimensional space, the OneClass SVM model obtains the smallest hypersphere containing a number of training samples greater than the threshold, and uses the smallest hypersphere as the decision boundary for distinguishing normal samples and abnormal samples; Input the newly collected time series data into the trained OneClass SVM model to obtain the mapped feature vector; Calculate the distance from the mapped feature vector to the smallest hypersphere; Compare the calculated distance with a preset threshold. If it exceeds the preset threshold, determine the time series data of the corresponding time window as abnormal.

8. The management method of the highly available policy engine according to claim 6, characterized in that: S53. According to the anomaly detection results of multiple metrics, use an anomaly state machine to determine whether there is an anomaly in the policy instance, including: Obtain the results of the anomaly detection of the time series data of multiple metrics. The results include the anomaly status of each metric in each time window. Among them, the anomaly status is represented by a boolean value, 1 means abnormal, and 0 means normal; According to the anomaly status of each metric in each time window, calculate the comprehensive anomaly degree of the policy instance in the corresponding time window; Arrange the comprehensive anomaly degrees of the policy instance in multiple consecutive time windows in chronological order to form an anomaly degree time series; Use a pre-constructed anomaly state machine model to determine whether there is an anomaly in the obtained anomaly degree time series. Among them, the anomaly state machine model is constructed by a directed graph. The nodes in the graph represent the anomaly status of the policy instance, and the edges represent the transition conditions and probabilities between the anomaly statuses.

9. A management system of a highly available policy engine, characterized in that including: At least one processing unit; used to execute instructions to implement the management method of the highly available policy engine according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Distributed cloud container resource scheduling method and system

    CN117909083A

  • Container-based strategy arrangement response method and system and computer storage medium

    CN119440733A