Management method and system for high-availability strategy engine

By versioning and encrypting the storage of quantitative strategy files, and using genetic algorithms for policy instance scheduling, combining anomaly detection and trend prediction models, the management and operation challenges of quantitative strategy in a real-time trading environment are solved, achieving high availability and stability.

CN119961771AActive Publication Date: 2025-05-09GOING INT INNOVATIVE TECH (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510449112.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Quantitative strategies are difficult to operate for a long time, stable and safely in real-time trading environments, especially in the areas of massive policy file management, computing resource scheduling, real-time monitoring and abnormal emergency response.

Method used

Adopt the management method of the high-availability policy engine, and use genetic algorithms to dynamically schedule policy instances by versioning and encrypting storage of policy files, and combine exception detection and trend prediction models to monitor and optimize policy operation in real time.

Benefits of technology

It realizes the secure and reliable storage and version management of policy files, improves system resource utilization, ensures high availability and stability of policies, promptly detects and handles abnormal situations, and avoids potential losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961771A_ABST
    Figure CN119961771A_ABST
Patent Text Reader

Abstract

The invention discloses a management method and system for a high-availability policy engine, and relates to the field of policy engines, and the method comprises the steps: storing an encrypted policy file on a plurality of data nodes in a distributed manner, and storing a historical version of the policy file; scheduling the strategy instance to a target node for operation according to an operation mode and scheduling constraints defined by the strategy file; the method comprises the following steps: judging whether the strategy instance runs abnormally or not by collecting running state data of the strategy instance which is scheduled to run, and rescheduling the corresponding strategy instance if the strategy instance runs abnormally; collecting performance parameters of a strategy instance running on the target node; according to the collected performance parameters, whether the strategy instance is abnormal or not is judged through an anomaly detection algorithm, and the use trend of the strategy instance to the node resources in a period of time in the future is predicted. Aiming at the problem that a strategy engine cannot effectively cope with frequent strategy updating in the prior art, the overall resource utilization rate of a system is improved under the condition that node resources and strategy constraint conditions are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of policy engines, and in particular to a management method and system for a high-availability policy engine. Background Art

[0002] With the rapid development of the financial industry, quantitative trading has been widely used around the world due to its advantages of guiding investment with scientific methodology and executing transactions with strict discipline. The scale of funds managed by quantitative private equity, hedge funds and other institutions continues to grow, and the research team behind the fund managers is also growing. The number of quantitative strategies they develop has increased exponentially. These quantitative strategies have formed a complete set of systematic methods from data collection, factor construction, signal generation, weight optimization, order execution and other links. They quickly capture market opportunities through programmatic trading, which is more efficient and objective than manual trading.

[0003] However, in the actual operation of quantitative strategies, especially in the real-time trading environment, how to ensure that each strategy can run long-term, stably and safely is a huge challenge. From development and testing to the launch of real-time trading, quantitative strategies have to go through a complete life cycle, which involves many management issues. First of all, the core of quantitative strategies is the strategy file developed by researchers, which contains important information such as the logic, parameters, and configuration of the strategy. How to effectively manage massive strategy files, ensure their security and version consistency, and avoid the risk of strategy confusion caused by human error is a common problem faced by quantitative trading institutions. Secondly, different strategies have very different requirements for computing resources. When hundreds of strategies are running concurrently, how to reasonably schedule and allocate hardware resources such as CPU, memory, and network, maximize resource utilization, and avoid interference between strategies caused by resource contention, is extremely technically challenging. Thirdly, in a complex market environment, the operation of quantitative strategies is often treacherous and requires real-time monitoring and evaluation. The traditional manual inspection mode is difficult to cope with hundreds of strategies, and a highly automated intelligent operation and maintenance method is urgently needed. Finally, quantitative trading runs non-stop 24 hours a day, 7 days a week. Any abnormality in a strategy may cause immeasurable losses to the fund. Therefore, a complete abnormal emergency handling mechanism must be established to ensure that when a strategy fails, it is discovered and isolated in the first time, and can be quickly restored and restarted.

[0004] In summary, the large-scale operation of quantitative strategies places high demands on the infrastructure. The traditional manual management model is no longer sustainable, and a highly available new generation strategy engine system is urgently needed. Summary of the invention

[0005] In response to the problem that the policy engine in the prior art cannot effectively cope with frequent policy updates, the present application provides a management method and system for a high-availability policy engine. By performing version control on policy files and using genetic algorithms to dynamically schedule policy instances, the overall resource utilization of the system is improved while satisfying node resource and policy constraints.

[0006] One aspect of the present application provides a management method for a high-availability policy engine, including: S1, encrypting the policy file, storing the encrypted policy file on multiple data nodes through distributed storage, and storing the historical version of the policy file; S2, scheduling the policy instance to run on the target node according to the operation mode and scheduling constraints defined in the policy file; S3, judging whether the policy instance is running abnormally by collecting the operation status data of the policy instance scheduled to run in step S2, and if abnormal, rescheduling the corresponding policy instance; S4, collecting the performance parameters of the policy instance running on the target node in step S2, the performance parameters include the node resource usage and the operation status of the policy instance; S5, judging whether the policy instance is abnormal through the anomaly detection algorithm based on the collected performance parameters, and predicting the use trend of the node resources by the policy instance in the future. In the present application, the high-availability policy engine is a software system that can execute policies continuously, stably and efficiently, and ensures that the policy can maintain normal operation under various abnormal conditions by adopting a series of fault-tolerant, redundancy, monitoring, scheduling and other technical means, thereby providing continuous and reliable decision-making services for upper-level businesses.

[0007] Among them, scheduling constraints refer to a series of restrictions that need to be met when assigning policy instances to target running nodes to ensure that the policy instances can run normally and meet business and system requirements. Common scheduling constraints include: Resource constraints: The CPU, memory, disk and other resources required for the policy instance to run must not exceed the available capacity of the node; Affinity constraints: Some policy instances need to be deployed on the same node or rack as specific services or instances to reduce communication latency; Mutual exclusion constraints: Some policy instances cannot be deployed on the same node as specific services or instances to avoid mutual interference; Data locality constraints: Some policy instances need to be deployed on the same node as their data source as much as possible to improve data access efficiency; The scheduling system needs to embed these constraints into the optimization objectives and evaluation indicators when making scheduling decisions, and find an optimal deployment plan that meets the constraints.

[0008] Anomaly detection algorithm refers to a type of algorithm that automatically identifies abnormal behaviors that deviate from normal patterns by analyzing the behavioral data of a system or component. In the high-availability policy engine scenario, anomaly detection algorithms are mainly used for anomaly detection of policy instance operation status and performance parameters, timely detection of faults and taking recovery measures. The basic assumption of the anomaly detection algorithm is that the data samples of normal behavior account for the majority, and the data samples of abnormal behavior are sparse and significantly different from normal samples. Through data mining and machine learning techniques, the anomaly detection algorithm can automatically learn the pattern characteristics of normal behavior and use this as a benchmark to identify anomalies. In this application, the anomaly detection algorithm can use statistical methods, machine learning methods such as K-Means clustering: assuming that normal data is clustered into clusters and abnormal data is far away from all cluster centers; LOF local anomaly factor: measure the density difference between data points and surrounding neighbors, and points with significantly lower density than neighbors are considered abnormal; OC-SVM single-class support vector machine: map data to high-dimensional space, enclose most normal data in a hypersphere, and consider points outside the sphere as abnormalities, etc. A combined algorithm can also be used.

[0009] Furthermore, S1 encrypts the policy file, stores the encrypted policy file in a distributed manner on multiple data nodes, and stores historical versions of the policy file, including: S11, encrypts the policy file using an asymmetric encryption algorithm to generate an encrypted policy file; S12, uses the data redundancy algorithm Reed-Solomon to divide the encrypted policy file into blocks, and adds check blocks; S13, stores the data blocks and check blocks on multiple data nodes respectively. When some data nodes fail, the original policy file is restored through the data blocks and check blocks on other nodes; S14, each time the policy file is modified, the new version of the policy file is encrypted and stored, and a parent-child link is established with the previous version to build a version tree of the policy file; S15, when a policy instance requests to load a policy file, the latest version of the policy file is first obtained from the data node. If the access to the latest version fails, the version tree is traversed to obtain the most recent historical version.

[0010] Among them, Reed-Solomon (RS) is an error correction code algorithm based on the data block level. By adding several redundant check blocks to the original data blocks, even if some data blocks are lost or damaged, the original data can be restored through the remaining intact data blocks and check blocks. The basic principle of the RS algorithm is: divide the original data into n data blocks, and then generate m check blocks based on the data blocks to form an (n, m) RS erasure code. When the total number of lost or damaged data blocks and check blocks does not exceed m, the lost data blocks can be reconstructed by solving a linear equation system. The RS algorithm has the advantages of simple encoding and decoding calculations and strong error correction capabilities. It is widely used in data redundancy and fault tolerance in the fields of disk arrays, distributed storage, satellite communications, etc. In the management method of the high-availability policy engine, the RS algorithm is used to perform data redundancy on the encrypted policy file, which can tolerate the failure of some data nodes and improve the reliability of policy file storage. A policy file refers to a digital file that records policy definitions, configuration parameters, execution logic, etc. in a specific format. The policy file organizes the various parts of the policy in a structured and modular manner to facilitate the management, update, sharing and reuse of the policy. Common policy file formats include XML, YAML, JSON, etc. In the management method of a high-availability policy engine, the policy file is the core resource for policy management and execution, and its reliable storage and version management directly affect the high availability of the policy engine.

[0011] Parent-child links are a way to organize the change relationship between multiple versions of a file in a version management system. Each version of a file is considered a node, and each change generates a new child version. The child version node points to the parent version node through a link, and the parent version node also records the links of all child version nodes, eventually forming a tree-structured version library.

[0012] Furthermore, S2, according to the operation mode and scheduling constraints defined in the policy file, schedules the policy instance to run on the target node, including: using a genetic algorithm to solve the optimal scheduling scheme, and realizing the allocation and deployment of the policy instance on the target node while satisfying the node resources and policy constraints; S21, reading the operation mode and scheduling constraint parameters defined in the policy file as the input of the scheduling algorithm; S22, collecting the current resource usage of each target node as the input of the scheduling algorithm; wherein the resource usage includes CPU and memory usage; S23, using a genetic algorithm to solve the optimal scheduling scheme, and allocating and deploying the policy instance to the target node while satisfying the node resources and scheduling constraints; S24, according to the optimal scheduling scheme, distributes the policy instance to the target node through the container orchestration Kubernetes, and each target node starts or updates the corresponding policy instance through the container engine according to the received instructions, and allocates computing resources.

[0013] Further, S23, a genetic algorithm is used to solve the optimal scheduling solution, including: according to the number of strategy instances and the number of candidate target nodes, the parameters of the genetic algorithm are set, and the parameters include the initial population size, crossover probability, mutation probability and termination generation; a priority-based encoding strategy is used to encode the deployment mapping relationship from the strategy instance to the target node into an N-tuple as the chromosome of a scheduling solution; wherein N is the number of strategy instances, and each component of the tuple represents the number of a strategy instance deployed to the target node; chromosomes are randomly generated to form an initial population; the fitness of each chromosome corresponding to the scheduling solution is calculated; a tournament selection algorithm is used to randomly select a group of chromosomes from the current population and The pair of chromosomes with the highest fitness is used as the father; the father chromosomes are subjected to multi-point crossover recombination according to the set crossover probability to generate new daughter chromosomes; the generated daughter chromosomes are mutated according to the mutation probability, one or more gene positions are randomly selected, and the values ​​of the selected gene positions are randomly replaced with the numbers of other candidate nodes to obtain the mutated daughter chromosomes; the daughter chromosomes obtained by the crossover recombination and the mutated daughter chromosomes are added to the next generation population, and the crossover and mutation operations are repeated until the size of the new population reaches the initial population size; the above iterations are repeated until the number of iterations reaches the set termination generation, and the chromosome with the highest fitness in the final population is decoded to obtain the optimal scheduling plan.

[0014] Furthermore, a priority-based encoding strategy is adopted to encode the deployment mapping relationship from the policy instance to the target node into an N-tuple as the chromosome of a scheduling scheme, including: setting the priority of the policy instance according to the execution frequency of the policy instance; sorting the tuple components corresponding to the policy instance according to the set priority; determining the value range of each gene bit in the chromosome according to the number of candidate target nodes; using grayscale encoding to map the number of each candidate target node into a binary string as the gene code of the corresponding node; splicing the gene codes of all candidate target nodes in sequence according to the tuple component order to form a complete chromosome; wherein each gene bit of the chromosome corresponds to a deployment node of the policy instance.

[0015] Among them, in the scheduling optimization method of the high-availability policy engine, an N-tuple is used to represent a scheduling scheme, that is, the deployment mapping relationship from a policy instance to a target node. This N-tuple can be expressed as (P1, P2, ..., PN), where N is the total number of policy instances, and Pi represents the target node number of the i-th policy instance deployment. Each element Pi in the N-tuple is called a tuple component, which represents the scheduling decision of a policy instance. For example, there are 5 policy instances {P1, P2, P3, P4, P5}, and the candidate target nodes are numbered {1, 2, 3}. A possible scheduling scheme is (2, 1, 3, 2, 1), which means: policy instance P1 is deployed to node 2; policy instance P2 is deployed to node 1; policy instance P3 is deployed to node 3; policy instance P4 is deployed to node 2; policy instance P5 is deployed to node 1; here 2, 1, 3, 2, 1 are the 5 tuple components in the N-tuple of this scheduling scheme, representing the target node allocation decision of the 5 policy instances. By rearranging and combining tuple components, we can traverse different scheduling schemes of the search strategy instance and find the optimal solution for the deployment constraints and optimization objectives. This is the basic idea of ​​scheduling optimization. Tuple components are an abstract representation method for scheduling decision variables in the scheduling optimization model.

[0016] Gray code is a binary coding method, which is characterized by only one binary bit difference between two adjacent code values. Through gray code, a discrete value range can be mapped into a set of binary codes. In the scheduling optimization method of the high-availability strategy engine, gray code is performed on the numbers of candidate target nodes to better meet the genetic algorithm's gene mutation requirements and improve search efficiency. There is only one binary bit difference between adjacent values ​​of gray code, such as only the highest bit changes from 0 to 1 between 1 (01) and 2 (11). This coding method has better continuity and smoothness and is more suitable for the crossover and mutation operations of genetic algorithms. If standard binary coding is directly used, there may be multiple bits of difference between adjacent codes, such as 1 (01) and 2 (10), which will cause a large jump after crossover and mutation, which is not conducive to retaining excellent genes. In the scheduling optimization method of the high-availability strategy engine, the gray codes of the candidate node numbers are spliced ​​in the order of N-tuple components to form a complete chromosome code. The genetic algorithm searches for an optimized scheduling solution by performing operations such as selection, crossover, and mutation on these gene bits. Grayscale coding, as a coding strategy of genetic algorithms, can better guide the search direction and accelerate the optimization solution process.

[0017] Furthermore, the fitness of each chromosome corresponding to the scheduling scheme is calculated: , where F is the fitness function value, and its value range is [0, 1]. The larger the value, the better the overall performance of the scheduling scheme; are the weight coefficients of the three optimization objectives of load balancing, criticality and failure rate, and ; The specific value can be dynamically adjusted according to the importance of each target during system operation; G is the Gini coefficient, which is used to measure the imbalance of load distribution between nodes. The calculation formula is: , where N is the number of candidate target nodes, and are the loads of the i-th node and the j-th node respectively; SLA is the service level agreement satisfaction rate of the key policy instance, and the calculation formula is: , where S is the number of key strategy instances, is the weight coefficient of the sth key strategy instance, is the SLA satisfaction indicator variable of the sth key policy instance; P is the failure rate of the scheduling scheme, and the calculation formula is: , where N is the number of candidate target nodes, is the failure rate of the nth node, The ratio of the number of policy instances deployed on the nth node to the total number of instances.

[0018] Further, S5, based on the collected performance parameters, determines whether the policy instance has an anomaly through an anomaly detection algorithm, and predicts the usage trend of the policy instance for node resources in the future, including: S51, constructing the resource usage time series of the policy instance according to the performance parameters collected in step S4, the time series includes CPU usage, memory usage and response time indicators; S52, using the One Class SVM model to perform anomaly detection on the time series of each indicator; wherein the OneClass SVM model determines whether the newly collected data point is abnormal through feature vector mapping and boundary learning; S53, based on the anomaly detection results of multiple indicators, determines whether the policy instance has an anomaly through the anomaly state machine; S54, using the LSTM neural network to predict the resource usage trend in the future, the input of the LSTM neural network is the time series of multiple indicators, and the output is the predicted value in the future; S55, based on the predicted node resource usage trend, determines the load of the policy instance in the future.

[0019] Among them, One Class SVM (Support Vector Machine) is a commonly used unsupervised anomaly detection model, which is particularly suitable for scenarios with only normal data and lack of abnormal data labels. Unlike the traditional two-class SVM, One Class SVM only uses one category of data to train the model and learns an optimal hypersphere, which tightly wraps the normal data points in the hypersphere and isolates the abnormal data points outside the hypersphere. The basic assumption of One Class SVM is that normal data is the majority and has a relatively compact distribution, while abnormal data is the minority and has a significant difference in distribution from normal data. In the anomaly detection of the high-availability policy engine, a One Class SVM model is trained for the time series of indicators such as CPU, memory, and response time of each policy instance to determine whether the newly collected indicator data points are abnormal. One Class SVM can adaptively learn the distribution characteristics of indicator data under normal conditions and has good detection and generalization capabilities for unknown abnormal patterns.

[0020] The abnormal state machine is a finite state automaton model that describes the transition process of the system state under event-driven. In the field of anomaly detection, it is mainly used to comprehensively judge and characterize the overall abnormal state of the system. The system starts from the normal state. When certain abnormal events or conditions are met, it will trigger the transition to the abnormal state. After entering the abnormal state, if the recovery conditions are met, it will transfer back to the normal state, thus forming a state transition diagram. In the anomaly detection of the policy instance, the One Class SVM anomaly detection results of multiple indicators (such as CPU, memory, etc.) are used as input events to drive the abnormal state machine to transfer the state, so as to judge the overall abnormal status of the policy instance. The abnormal state machine integrates the local abnormal signals of multiple dimensions into a global abnormal judgment through the state transition logic, and distinguishes the severity and duration of the abnormality. It can more comprehensively and accurately reflect the health level of the complex system and provide a reference for abnormal location and decision-making.

[0021] Further, S52, a One Class SVM model is used to perform anomaly detection on the time series of each indicator, including: segmenting the time series data according to sliding windows, extracting statistical features of each window, and forming a statistical feature vector; obtaining historical sequence data when the policy instance is operating normally, and constructing a training data set; using the training data set to train the One Class SVM model; wherein the One Class SVM model maps the feature vector in the training data set from the original space to the high-dimensional space through a kernel function; in the high-dimensional space, the One Class SVM model obtains the minimum hypersphere containing the number of training samples greater than a threshold, and uses the minimum hypersphere as the decision boundary for distinguishing normal samples from abnormal samples; the newly collected time series data is input into the trained One Class SVM model to obtain the mapped feature vector; the distance from the mapped feature vector to the minimum hypersphere is calculated; the calculated distance is compared with a preset threshold, and if it exceeds the preset threshold, the time series data of the corresponding time window is judged to be abnormal.

[0022] Further, S53, based on the anomaly detection results of multiple indicators, determine whether the strategy instance has an anomaly through an anomaly state machine, including: obtaining the results of anomaly detection of time series data of multiple indicators, the results including the abnormal state of each indicator in each time window; wherein the abnormal state is represented by a Boolean value, 1 represents abnormality, and 0 represents normal; according to the abnormal state of each indicator in each time window, calculate the comprehensive abnormality degree of the strategy instance in the corresponding time window; arrange the comprehensive abnormality degrees of the strategy instance in multiple consecutive time windows in chronological order to form an abnormality degree time series; use a pre-built abnormal state machine model to determine whether the obtained abnormality degree time series has an abnormality, wherein the abnormal state machine model is constructed using a directed graph, the nodes in the graph represent the abnormal state of the strategy instance, and the edges represent the transition conditions and probabilities between abnormal states.

[0023] Another aspect of the present application also provides a management system of a high-availability policy engine, which is used to execute a management method of a high-availability policy engine of the present application.

[0024] Compared with the prior art, the advantages of this application are:

[0025] By encrypting, storing, versioning, and performing data redundancy on policy files, we effectively solve the problems of inconsistent policy file storage and version confusion caused by frequent policy updates in the prior art. Managing policy files by version not only ensures the atomicity of policy release, but also enables rapid rollback to the previous stable version when a policy update fails, greatly improving the flexibility and robustness of policy file management. At the same time, we use asymmetric encryption and Reed-Solomon erasure coding technologies to encrypt and redundancy stored policy files, ensuring that confidential policies will not be illegally stolen, and that the original policies can still be restored in the event of failure of some storage nodes, thereby fundamentally ensuring the security and reliability of policy storage.

[0026] Using genetic algorithms to dynamically schedule policy instances can achieve better load balancing while meeting node resource and policy constraints, thereby improving the overall resource utilization efficiency of the system. Traditional policy instance scheduling mostly uses static rules or simple dynamic algorithms such as polling. Such methods are difficult to adapt to the complex and changeable policy operation environment, and are prone to cause imbalances where individual nodes are overloaded while other nodes are idle. The genetic algorithm simulates the process of biological evolution, constructs a fitness function with load balancing, key policy SLA satisfaction rate, and solution failure rate as optimization goals, and abstracts the deployment mapping relationship from policy instances to target nodes into a chromosome using a priority-based encoding method, and searches for the optimal solution in the population through iterative evolution. Since the genetic algorithm has the ability to search for global optimization, it can quickly search for scheduling solutions that meet multiple constraints, and continuously evolve and optimize with the dynamic changes of the policy runtime environment, thereby maximizing the load balance of each node while ensuring the performance of key policies, so that limited computing resources can be fully utilized.

[0027] Combining anomaly detection and trend prediction models, real-time monitoring of the running status of policy instances and early prediction of resource usage in the future can provide a reliable basis for dynamic scheduling decisions of policy instances. The One Class SVM model uses technical means such as feature vector mapping and boundary learning to determine the abnormal state of time series data. Compared with simple threshold comparison methods, the One Class SVM can adaptively learn the behavior patterns of policy instances during normal operation, and has better detection and tolerance capabilities for abnormal behaviors caused by various unknown faults. The LSTM model uses deep learning algorithms to establish nonlinear correlations between system performance indicators, mine the evolutionary trends contained in time series data, and accurately predict changes in policy instance demand for system resources over a period of time. The combination of anomaly detection and trend prediction can not only detect current operating failures in a timely manner, but also prepare for a rainy day, provide optimized spatial and temporal dimensions for dynamic scheduling, and facilitate the policy engine to continuously and stably provide external services.

[0028] The abnormal state machine model comprehensively analyzes the anomaly detection results of multiple indicators, evaluates the health status of the policy instance from a global perspective, and overcomes the semantic gap between the anomaly of a single indicator and the anomaly of the policy instance as a whole. Traditional operation and maintenance monitoring often sets an abnormal threshold for a single indicator such as CPU and memory, and issues an alarm once an indicator exceeds the threshold, which is prone to produce many false positives and missed positives. Whether a policy instance is abnormal cannot be simply judged by whether one or two indicators are abnormal, but must be considered from the perspective of its behavioral evolution process. The abnormal state machine model abstracts the law of the evolution of the abnormal state of the policy instance over time through a directed graph, and combines the anomaly detection snapshots of each indicator to accurately depict the complete process of the policy instance from local anomaly to global failure, and can quickly locate and isolate faults based on this, thereby minimizing the impact of abnormal policy instances on the system and improving the system's fault tolerance. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0030] Figure 1 It is an exemplary flowchart of a management method of a high-availability policy engine according to some embodiments of the present application. DETAILED DESCRIPTION

[0031] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0032] like Figure 1 As shown, the policy file is encrypted, and the encrypted policy file is stored in a distributed manner on multiple data nodes, and the historical version of the policy file is stored; the policy instance is scheduled to run on the target node according to the operation mode and scheduling constraints defined in the policy file; by collecting the operation status data of the scheduled policy instance, it is determined whether the policy instance is running abnormally, and if it is abnormal, the corresponding policy instance is rescheduled; the performance parameters of the policy instance running on the target node are collected, and the performance parameters include the node resource usage and the operation status of the policy instance; based on the collected performance parameters, it is determined whether the policy instance is abnormal through the anomaly detection algorithm, and the usage trend of the node resources by the policy instance in the future period is predicted.

[0033] Specifically, S1, encrypts the policy file, stores the encrypted policy file on multiple data nodes in a distributed manner, and stores the historical version of the policy file, including: S11, encrypts the policy file using an asymmetric encryption algorithm to generate an encrypted policy file; first, generate a pair of RSA keys for each policy file owner, including a public key and a private key. The public key is used to encrypt the policy file, and the private key is used to decrypt the encrypted policy file. The key length of the RSA key can be flexibly set according to the requirements of the security level, such as 1024 bits, 2048 bits, etc. After the policy file is created or modified, the policy file is encrypted using the public key of the policy owner. The encryption process can use different padding schemes such as PKCS#1 v1.5 and OAEP to enhance the encryption strength. The encrypted policy file will be stored and distributed in the form of ciphertext. When a policy instance needs to read and execute a policy file, it is first necessary to verify the identity of the user to which the policy instance belongs. Only users who have passed identity authentication and authorization can use their corresponding private keys to decrypt the encrypted policy file and obtain the policy logic and parameter definition in plain text. Even if an unauthorized user obtains an encrypted policy file, they cannot decrypt the sensitive content without knowing the private key. Through the asymmetric encryption mechanism, access control of the plaintext content of the policy file is achieved. The policy file is always in ciphertext during storage and transmission. Only when the policy is actually executed will it be decrypted into plaintext in the policy instance authorized by the policy owner. This can minimize the exposure scope and time of the plaintext policy file, and effectively prevent the leakage of confidential information such as policy logic and parameters.

[0034] S12, in terms of data reliability, introduces the Reed-Solomon data redundancy algorithm. This algorithm divides the encrypted policy file into blocks and generates check blocks based on the block data. The data blocks and check blocks are stored in multiple data nodes. When some data nodes fail, the complete original policy file can be restored using the data blocks and check blocks on other nodes. For example, if RS (12, 4) encoding is used, the policy file is divided into 12 data blocks, and 4 check blocks are generated, for a total of 16 blocks. These 16 blocks can be stored on 16 data nodes. If any 4 nodes fail, the original policy file can still be reconstructed based on the remaining 12 blocks, greatly improving the fault tolerance capability against node failures.

[0035] S13, the data blocks and check blocks are stored on multiple data nodes respectively. When some data nodes fail, the original policy file is restored through the data blocks and check blocks on other nodes; S14, in terms of version management, each time the policy file is modified, the new version of the policy file is encrypted and stored, and a parent-child link is established with the previous version to build a version tree of the policy file. The version tree records the change history of the policy file, forming a series of file snapshots representing the status at different time points. When the policy file needs to be loaded, the policy instance first obtains the latest version. If the latest version fails to be accessed, the version tree is backtracked to search for the most recent available historical version. Through the version tree and backtracking mechanism, the latest changes to the policy file can be obtained at any time, and the previous stable version can be quickly restored when the update fails, which effectively deals with various abnormal situations that may be caused by frequent policy updates.

[0036] S15, in terms of policy operation management, according to the different operation modes of policy files in actual applications, it is divided into batch mode and overlay mode. In batch mode, a policy file will generate multiple policy instances, each instance is responsible for processing different data nodes to achieve parallel computing; in overlay mode, a policy file only generates one policy instance, which is responsible for processing all data nodes to achieve serial computing. Different operation modes have different scheduling constraints such as resource requirements, concurrency restrictions, and node affinity requirements. At the same time, it is also necessary to collect the operation status data of the policy instance in real time, including key indicators such as CPU and memory usage, and event processing delay, in order to accurately evaluate the execution of the policy instance.

[0037] S2, according to the operation mode and scheduling constraints defined in the policy file, schedules the policy instance to run on the target node, including: using genetic algorithms to solve the optimal scheduling scheme, and realizing the allocation and deployment of the policy instance on the target node under the conditions of node resources and policy constraints. Node resource usage includes CPU usage, memory usage, and disk I / O, reflecting the load of the target node; the operation status of the policy instance includes the event processing delay and order execution of the policy, reflecting the execution effect of the policy instance.

[0038] S21, read the running mode and scheduling constraint parameters defined in the policy file as the input of the scheduling algorithm; the policy file is usually written in a structured markup language such as JSON and YAML, and declares various attributes of the policy in the form of key-value pairs. The scheduling algorithm reads the policy file and parses the parameters related to the running mode and scheduling constraints, mainly including: Running mode: specifies the parallel deployment method of the policy instance. Common values ​​include: "Batch": batch mode, one policy file generates multiple policy instances to process data from different nodes respectively; "Overlay": overlay mode, one policy file generates one policy instance to process data from all nodes; for example, a policy file defines "runMode": "Batch", which means that the policy is run in batch mode. The scheduling algorithm extracts this parameter and decides whether it is necessary to deploy policy instances in parallel on multiple nodes. Resource requirements: computing resources required for the strategy instance to run, including: "cpu": the number of CPU cores required, such as "cpu": 2; "memory": the required memory space, such as "memory": "4GB"; "gpu": the number of GPU cards required, such as "gpu": 1; the scheduling algorithm selects the target nodes that meet the resource requirements of the strategy instance based on these parameter values ​​to avoid the failure of the strategy instance to start due to insufficient resources. Concurrency limit: to avoid overloading the strategy instance, you can set the maximum number of events or transactions that can be processed concurrently, such as: "max Concurrency": 1000, indicating that the strategy instance can process up to 1000 events at the same time; "maxOrders": 5000, indicating that the strategy instance can process up to 5000 transactions at the same time; the scheduling algorithm refers to these limits to control the number of strategy instances on a single node to prevent the node from being overloaded. Target node affinity: limits the deployment node range of the policy instance, supporting both whitelist and blacklist modes: "node Affinity": "in: node1, node2", indicating that the policy instance can only be deployed on node1 or node2; "node Affinity": "not in: node3", indicating that the policy instance cannot be deployed on node3; the scheduling algorithm reads the affinity expression and matches and filters among the candidate nodes.

[0039] S22, collect the current resource usage of each target node as the input of the scheduling algorithm; the scheduling algorithm needs to grasp the load level of each target node in real time in order to reasonably allocate policy instances. The resource usage of the node is collected regularly through the monitoring component, focusing on the two indicators of CPU usage and memory usage: CPU usage = (1-idle CPU time / total CPU time) * 100%. In this embodiment, the total CPU time and idle time in a time window (such as 1 minute) are collected through the Linux / proc / stat file: cpu 841001 10116 889278 6010608 22267 0 17761 0 0 0; the first column is the total CPU time, and the fourth column is the idle CPU time, both in units of 1 / 100 seconds. Assume that the total time of the two collections is T1=7763264 and T2=7824728, and the idle time is I1=6003511 and I2=6028102 respectively. Then the CPU usage in this time window is: CPU usage = (1-(I2-I1) / (T2-T1))*100%=60.33%. Memory usage = used memory / total memory*100%. In this embodiment, the free command of Linux is used to collect the memory usage snapshot of the node. The total column is the total memory and the used column is the used memory. The unit is KB. Assuming that the total memory is 8GB=8388608KB and the used memory is 1.6GB=1677722KB, the memory usage is: memory usage = 1677722 / 8388608*100%=20.00%. See Table 1, Node Memory Status Table for details.

[0040] Table 1 Node memory status table Total Memory Used idle shared Buffering / Cache Available 8169348 1680404 4066836 495620 2422108 5812632

[0041] The monitoring component periodically collects the CPU usage and memory usage data of all target nodes and constructs them into time series data in the following format. Each line records the resource usage of a node at a certain moment, all expressed as floating point decimals. For details, see Table 2, Node Resource Time Series Table.

[0042] Table 2 Node resource timing table Timestamp node CPU usage Memory usage 1620620705 node1 0.6033 0.2000 1620620705 node2 0.4288 0.3155 1620620705 node3 0.7500 0.8360

[0043] The scheduling algorithm queries the node resource usage within a certain time range (e.g., the last 5 minutes) from the time series database, takes the average value of each indicator, and forms the following data snapshot as the input of the scheduling decision. See Table III, Scheduling Decision Input Table for details.

[0044] Table 3 Scheduling decision input table node Average CPU usage Average memory usage node1 0.5933 0.1988 node1 0.4012 0.3008 node1 0.7707 0.8298

[0045] The operating parameters and node resource usage defined in the policy file are converted into structured data that can be directly used by the scheduling algorithm after preprocessing steps such as collection, parsing, and aggregation. Based on this, the scheduling algorithm evaluates the available resources of each node, weighs the concurrency restrictions and affinity constraints of the policy instance, and dynamically generates the optimal node deployment plan.

[0046] S23, using genetic algorithm to solve the optimal scheduling solution, and allocating and deploying the strategy instances to the target nodes while satisfying the node resources and scheduling constraints; including: setting genetic algorithm parameters: according to the number of strategy instances N to be scheduled and the number of candidate target nodes M, setting the following parameters: initial population size P: generally 20~100, ensuring population diversity while controlling computing overhead, P=50 can be set; crossover probability Pc: generally 0.4~0.9, the larger the value, the faster the population update, Pc=0.8 can be set; mutation probability Pm: generally 0.001~0.1, the larger the value, the higher the population diversity, Pm=0.05 can be set; termination algebra T: generally 100~500, the larger the value, the higher the quality of the solution, but the slower the convergence speed, T=200 can be set. For example, there are currently 100 policy instances and 10 candidate nodes, then N=100, M=10, and the genetic algorithm parameters can be set as: P=50, Pc=0.8, Pm=0.05, T=200.

[0047] Chromosome encoding: A priority-based encoding strategy is used to encode a scheduling scheme into an integer tuple of length N. Each component of the tuple takes a value of 1~M, indicating on which node the policy instance of the corresponding sequence number is deployed. The encoding process is as follows: According to the historical execution frequency of the policy instance, its priority is calculated. The higher the frequency, the greater the priority, such as: policy1: 15 times / minute, priority=5; policy2: 10 times / minute, priority=4; policy3: 8 times / minute, priority=3; ...; Sort the policy instances from high to low according to priority, and the ones with high priority are placed in front of the tuple to form the encoding order of the policy instance, such as: policy1>policy2>policy3>......; According to the number of candidate nodes M, the value range of each gene bit is determined to be 1~M, and the node number is mapped to an equal-length binary string using grayscale encoding as the gene encoding of the node. For example, when M=10, see Table 4, the gene encoding table of the policy instance.

[0048] Table 4. Gene encoding table of strategy examples Strategy Examples Priority Deploy Node Genetic coding Policy1 5 Node1 000 Policy2 4 Node2 001 Policy3 3 Node3 011

[0049] According to the coding order of the policy instance, the gene codes of each node are spliced ​​in sequence to form a complete chromosome, such as: policy1-node1, policy2-node2, policy3-node3. Mapping to chromosome: 000001011 means that policy1 is deployed on node1, policy2 is deployed on node2, policy3 is deployed on node3, and so on.

[0050] Initialize the population: randomly generate P chromosomes as the initial population of the genetic algorithm, as shown in Table 5, an example table of chromosomes of the initial population.

[0051] Table 5. Initial population chromosome example table Chromosome number Chromosome Example 1 010,011,000,111,001 2 001,010,011,000,110 3 011,001,010,100,000 4 111,000,001,011,010

[0052] Fitness evaluation: Calculate the fitness value of each chromosome corresponding to the scheduling scheme. The fitness function is: , the value range is [0, 1], the larger the value, the better the solution. The weight coefficients of the three optimization goals of load balancing, key strategy SLA, and failure rate can be adjusted dynamically according to the importance of the three goals at runtime, such as increasing the , increase during peak business hours , increase when nodes frequently fail , always guarantee G is the Gini coefficient, which measures the imbalance of load distribution among nodes: , where N is the number of nodes, and is the load of the i-th and j-th nodes, which can be expressed by CPU, memory usage or number of policy instances. G takes values ​​of [0, 1], and the smaller it is, the more balanced the load is. For example, for a scheduling scheme represented by a chromosome, the number of policy instances of each node is: 18, 9, 11, 20, 2, then: G=1 / 10*(|18-9|+|18-11|+|18-20|+|18-2|+|9-11|+|9-20|+|9-2|+|11-20|+|11-2|+|20-2|)=0.64; indicating that the node load distribution of this scheme is not balanced enough. SLA is the SLA satisfaction rate of the key policy instance: ; where S is the total number of key strategy instances, is the weight of the sth key instance, which is generally proportional to its priority. It is the SLA satisfaction indicator variable of the sth key instance. It is 1 when deployed on a node that meets the performance requirements, otherwise it is 0. SLA takes values ​​[0, 1]. The larger the value, the better the performance requirements of the key policy instance are met. For example, for a chromosome, the weights of the three key policy instances are 0.5, 0.3, and 0.2, respectively, and the SLA of two instances is met, then: SLA = (0.51 + 0.31 + 0.2 * 0) / 3 = 0.6; indicating that the scheme has a good performance guarantee for the key policy instance. P is the failure rate of the scheduling scheme: ; where N is the total number of nodes, is the failure rate of the nth node, The percentage of strategy instances deployed for the nth node. P takes values ​​of [0, 1], and the smaller the value, the higher the reliability of the solution. For example, for a chromosome, assuming that the failure rates of the three nodes are 0.01, 0.05, and 0.001 respectively, and the percentage of their deployed instances is 0.5, 0.3, and 0.2, then: , indicating that the overall failure rate of this scheme is low.

[0053] Genetic operators: Use operators such as tournament selection, multi-point crossover, and random mutation to evolve the current population and generate the next generation population: Use tournament selection to randomly select K individuals each time, and take the pair with the highest fitness as the father, generally K=3~5; use multi-point crossover to randomly select several nodes as crossover points, align the paternal chromosomes at the crossover points, and then exchange the corresponding gene fragments to recombinant new offspring chromosomes, such as: parent1:000|011|010; parent2:001|100| 111; child1: 000|100|010; child2: 001|011|111; Random mutation is used to randomly select certain gene positions of the chromosome with probability Pm and replace them with other values ​​to introduce new search directions, such as: original: 011000101010; mutated: 010000001010; The elite retention strategy is used to directly copy the first E individuals with the highest fitness in the current population to the next generation, generally E=1~5, to avoid the loss of the optimal solution.

[0054] Iteration termination: Repeatedly execute genetic operators such as selection, crossover, and mutation to continuously generate new populations until the number of iterations reaches T, or the optimal fitness value of the population has not been significantly improved for several consecutive generations, the algorithm terminates. Result decoding: Decode the chromosome with the highest fitness in the final population to obtain the optimal scheduling solution, such as: best chromosome: 010 000011 001 ......; decoded to: policy1 -> node2; policy2 -> node1; policy3 -> node4; policy4 -> node2; thus completing the optimal scheduling solution process based on the genetic algorithm.

[0055] S24 Execution and distribution of scheduling scheme: Convert the optimal scheduling scheme into Kubernetes resource configuration files, such as Deployment, StatefulSet, etc., and define the image, resource request, environment variables and other attributes of each policy instance. Then, the configuration file is distributed to the Kubernetes cluster through the kubectl apply command, and the scheduler of the Master node allocates the Pod to the target working node, and then the container engine (such as Docker) on the node pulls the image and starts the policy container instance. When the policy is updated or the node load changes, the optimal scheduling scheme can be recalculated through the above steps, and the differences between the new and old schemes can be compared. Only the policy instances of the different parts are migrated or scaled to achieve incremental rolling updates and reduce scheduling overhead. Genetic algorithms can efficiently solve complex policy instance scheduling problems. Through encoding mapping, fitness evaluation, genetic evolution and other means, the optimal solution is searched globally, and combined with cloud native technology stacks such as Kubernetes, the flexible execution and dynamic update of scheduling schemes are realized, thereby building an intelligent policy instance management system, which strongly supports the parallelization and scale operation of policies.

[0056] S3, by collecting the running status data of the policy instance scheduled in step S2, determine whether the policy instance is running abnormally. If abnormal, reschedule the corresponding policy instance; specifically, define abnormal judgment indicators: according to the characteristics of the policy instance, set a set of key indicators reflecting its running status, such as: liveness probe: such as whether the process exists, whether the port is available, etc. Once the detection fails, it is considered that the instance is abnormal; readiness probe: such as whether the response delay has timed out, whether the dependent service is reachable, etc., when the readiness probe fails, the instance cannot process new requests; event processing delay (event Latency): the average time for the instance to process an event. If it is significantly higher than the historical level, it means that the instance may have a performance bottleneck; resource utilization (resource Usage): the proportion of the node's CPU, memory, disk and other resources occupied by the instance. When it exceeds the set threshold, it is considered that the instance is abnormal. Status data collection: Through monitoring components such as Prometheus, the running status indicator data of each policy instance is collected regularly, such as: the detection results of the survival probe and the readiness probe (success or failure); the distribution statistics of event processing delay (TP50 / TP90 / TP99, etc.); resource indicators such as CPU usage, memory usage, and disk usage. The status data is stored in the monitoring database in the form of time series, and each data point contains the indicator name, timestamp, label (policy name + instance ID) and value.

[0057] Abnormal instance judgment: Compare the collected instance status data with the threshold of the abnormal judgment indicator to identify abnormal policy instances, such as: instances where the survival probe or readiness probe fails multiple times (such as 3 times); instances where the TP90 value of event processing delay continues to exceed 100ms; instances where the CPU or memory usage rate continues to be higher than 80%. Abnormal judgment can be achieved through the preset Prometheus alarm rules. When the condition expression of the alarm rule continues to meet the specified time, the corresponding alarm is triggered, and the alarm information contains the label of the abnormal instance. Abnormal instance rescheduling: After receiving the abnormal alarm, the scheduling system queries and deletes the Kubernetes resource object (such as Pod) of the abnormal instance according to the policy name and instance ID in the alarm, and releases the node resources occupied by it. Then the scheduling system re-executes the S2 scheduling algorithm, calculates a new scheduling plan for eliminating the faulty node, and creates a new policy instance to fill the deleted abnormal instance to ensure that the number of copies of the policy meets the requirements. The newly created policy instance will fill the vacancy of the abnormal instance and be scheduled to the healthy node to restore the high availability of the entire policy.

[0058] S4, collects the performance parameters of the policy instance running on the target node in step S2. The performance parameters include the node resource usage and the running status of the policy instance. Specifically, the node resource usage focuses on indicators such as CPU, memory, disk, and network: node_cpu_usage: the CPU usage of the node; node_cpu_load1 / 5 / 15: the average CPU load of the node in the last 1 minute, 5 minutes, and 15 minutes; node_memory_usage: the memory usage of the node; node_memory_available: the available memory of the node; node_file system_usage: the disk partition usage of the node; node_network_receive / transmit_bytes: the number of bytes received / sent by the node network.

[0059] The running status of the policy instance, focusing on service quality and fault-related indicators: instance_event_latency: latency distribution of instance processing events; instance_order_count: total number of orders successfully executed by the instance; instance_order_amount: total amount of orders successfully executed by the instance; instance_failure_count: number of orders that failed to execute or were abnormal; instance_failure_rate: instance execution failure rate; instance_message_backlog: number of messages to be processed by the instance.

[0060] By continuously collecting and monitoring these two types of performance parameters, the scheduling system can grasp the load, capacity, SLA and other operating conditions of all nodes and policy instances in the cluster in real time, and evaluate the actual effect of the scheduling plan accordingly: evaluate whether the node resource allocation is reasonable, whether the load is balanced, and whether there is overload or waste; evaluate whether the performance guarantee of key policy instances is in place and whether there is any risk of SLA violation; evaluate the failure level and failure mode of nodes and instances, and whether there are high availability vulnerabilities.

[0061] The scheduling system can associate these performance parameters with the scheduling model and continuously optimize the scheduling algorithm in a data-driven way: use load data to guide the scheduling strategy to dynamically balance load balancing, performance and stability goals; use SLA data to assist scheduling decisions and enhance the identification and isolation of long-tail transactions; use fault data to evaluate the health of nodes and instances to improve the fault tolerance of scheduling solutions; collect effect data of different scheduling decisions, and continuously iterate and optimize scheduling models with the help of AI algorithms such as reinforcement learning. The running status of strategy instances and the usage of node resources are important inputs and feedbacks of the scheduling system. By collecting and analyzing these performance parameters, the scheduling system can perceive business semantics, adapt to business changes, and ultimately enable the scheduling solution to continuously evolve in the optimal direction, providing a solid data foundation for the stable and efficient operation of quantitative strategies.

[0062] S5, based on the collected performance parameters, uses an anomaly detection algorithm to determine whether the policy instance has an anomaly, and predicts the usage trend of the policy instance for node resources in the future, including: S51, based on the performance parameters collected in S4, constructs a time series of multiple resource usage indicators for each policy instance. Common indicators include: CPU usage: the percentage of the node's CPU resources occupied by the instance, indicating the instance's computing load intensity. Memory usage: the percentage of the node's memory resources occupied by the instance, indicating the instance's storage load intensity. Response time: the average time it takes for the instance to process a request, indicating the instance's service quality level. Each data point in the time series contains an indicator name, timestamp, policy instance label, and value.

[0063] S52, use the One Class SVM model to perform anomaly detection on the time series of each indicator; use the One Class SVM model to independently detect anomalies on the time series of each resource usage indicator: Data segmentation and feature extraction: set a sliding window of fixed size (such as 1 hour), and divide the time series into several segments according to the window; perform feature extraction on each time window to obtain a set of statistical feature values, such as mean, standard deviation, quantile, etc., to form a feature vector; the window sliding step is 10 minutes, then the one-hour time series is extracted into 6 feature vectors.

[0064] Construction of normal sample training set: select several days of historical data when the strategy instance is operating normally as the training set; extract the training set time series into multiple feature vectors. One Class SVM model training: use the feature vectors of the training set to train the One Class SVM model; One Class SVM maps the feature vectors from the original space to the high-dimensional space through the kernel function (such as the RBF kernel); in the high-dimensional space, One Class SVM obtains the smallest hypersphere containing most of the training samples as the boundary of the normal sample. Anomaly detection: extract the new time series to be detected as a feature vector input into the model; One Class SVM maps it to the high-dimensional space and calculates the distance d to the smallest hypersphere; if d is greater than a given threshold (generally the 90% quantile of the sample distance), it is judged as an anomaly.

[0065] S53, based on the anomaly detection results of multiple indicators, determine whether the policy instance has an anomaly through the anomaly state machine; including: performing anomaly detection on multiple key indicators of the instance, such as CPU, memory, response time, etc. (such as using OneClass SVM), and the detection result is the abnormal state of each indicator in each time window, represented by Boolean value 0 (normal) or 1 (abnormal). The detection results of an instance in 6 consecutive windows are: cpu_anomaly: [0, 0, 1, 1, 1, 0]; memory_anomaly: [0, 0, 0, 1, 0, 0]; latency_anomaly: [0, 1, 1, 1, 0, 0].

[0066] Calculate the comprehensive anomaly of the instance in each time window, and count the abnormal proportion of each indicator as the comprehensive anomaly of the window period. Anomaly = number of abnormal indicators / total number of indicators. In the above example, the anomaly of the 6 windows is: [0 / 3, 1 / 3, 2 / 3, 3 / 3, 1 / 3, 0 / 3] = [0, 0.33, 0.67, 1, 0.33, 0]. Obtain the anomaly time series, and arrange the anomaly of the instance in multiple consecutive time windows in chronological order to form an anomaly time series that reflects the change trend of the overall abnormal state of the instance. The anomaly time series of the above example is: anomaly_score: [0, 0.33, 0.67, 1, 0.33, 0].

[0067] Use the abnormal state machine to determine whether the abnormal degree sequence is abnormal; pre-build an abnormal state machine model to describe the evolution of the instance from normal to abnormal. The state machine is represented by a directed graph. The nodes represent the abnormal state of the instance, and the edges represent the state transition conditions and probabilities. Common abnormal states include: normal, suspicious, abnormal, faulty, etc. The transition conditions on the edges are usually related to the abnormal degree threshold. Match the abnormal degree sequence with the state machine, calculate the most likely state transition path, and output the final abnormal state of the instance. In the above example, the matching process of the abnormal degree sequence [0, 0.33, 0.67, 1, 0.33, 0] in the state machine is:

[0068] Normal--anomaly_score≥0.33-->Suspicious;Suspicious--anomaly_score≥0.67-->Anomalous;

[0069] Anomalous--anomaly_score≤0.33-->Suspicious;

[0070] Suspicious--anomaly_score≤0.33-->Normal;

[0071] Normal--0.33-->Suspicious--0.67-->Anomalous--1.0-->Anomalous--0.33-->Suspicious--0.0-->Normal. The final state is Normal, indicating that the instance only had a temporary anomaly and has returned to normal. The abnormal state machine flexibly represents the different stages and paths in the abnormal evolution process in the form of a graphical model. By matching the abnormal degree sequence in the state machine, different abnormal modes can be adaptively identified to discover non-sudden anomalies such as gradual, intermittent, and periodic.

[0072] S54, using LSTM (Long Short-Term Memory) network to predict the resource usage trend of the strategy instance in the future: Construct multivariate time series: construct the instance's CPU, memory, response time and other indicators into multiple columns of time series data; each line is the value of each indicator at a time point, as shown in Table 6, strategy instance resource usage time series data table.

[0073] Table 6 Policy instance resource usage time series data table Timestamp CPU usage Memory usage Response time (ms) 1620620600 0.2 0.6 20 1620620660 0.25 0.62 22 1620620720 0.18 0.65 25

[0074] Data preprocessing: interpolate missing values ​​and smooth outliers; normalize each indicator data to eliminate the impact of dimension; differentiate the data at a fixed step size to extract the relative change trend of the data. Construct training and test sets: select a part of the historical data (such as 70%) as the training set and the rest as the test set; use the sliding window method to divide the time series data into multiple input and output pairs: (X, y); each X contains historical observations of a fixed length (such as 12), and y contains the target forecast values ​​for several future time points (such as 3).

[0075] Build an LSTM prediction model: The model input layer receives the feature matrix X of the multivariate time series; the middle layer uses several LSTM layers and Dense fully connected layers to extract time series features and fit nonlinear trends; the output layer predicts the multi-index value y at several future time points; use loss functions such as mean square error and use the Adam optimizer to train model parameters. Model evaluation and optimization: Evaluate the model prediction accuracy on the test set, calculate evaluation indicators such as MAE, MAPE, RMSE; optimize by adjusting hyperparameters and trying different model structures. Model reasoning and prediction: Use the trained LSTM model, input the latest time series data, and predict future resource usage trends.

[0076] S55, after using LSTM to predict the resource usage trend of the strategy instance in the future (such as 1 hour), its future load situation can be judged: compare the predicted value of CPU and memory usage with the current value to determine whether there is a significant growth trend and whether the growth rate exceeds the threshold (such as 30%); analyze the response time change trend to determine whether there is a deterioration trend and whether the severity affects the service quality; infer whether the instance will continue to be in a high-load state in the future and whether the high-load duration exceeds the threshold (such as 30 minutes); based on this, the future load state of the instance is divided into: idle, normal, busy, and overload levels.

Claims

1. A management method for a high-availability policy engine, characterized in that: include: S1, encrypt the policy file, store the encrypted policy file on multiple data nodes in a distributed manner, and store the historical version of the policy file; S2, schedules the policy instance to run on the target node according to the operation mode and scheduling constraints defined in the policy file; S3, by collecting the running status data of the policy instance scheduled in step S2, determining whether the policy instance is running abnormally, and if abnormal, rescheduling the corresponding policy instance; S4, collecting performance parameters of the policy instance running on the target node in step S2, where the performance parameters include node resource usage and the running status of the policy instance; S5, based on the collected performance parameters, determines whether the policy instance has an anomaly through an anomaly detection algorithm, and predicts the usage trend of node resources by the policy instance in the future.

2. The method for managing a high availability policy engine according to claim 1, characterized in that: S1, encrypt the policy file and store the encrypted policy file on multiple data nodes in a distributed manner, including: S11, encrypting the policy file using an asymmetric encryption algorithm to generate an encrypted policy file; S12, using the data redundancy algorithm Reed-Solomon, divides the encrypted policy file into blocks and adds a check block; S13, storing the data blocks and the check blocks on multiple data nodes respectively. When some data nodes fail, the original policy file is restored through the data blocks and check blocks on other nodes. S14, each time the policy file is modified, the new version of the policy file is encrypted and stored, and a parent-child link is established with the previous version to build a version tree of the policy file; S15, when the policy instance requests to load a policy file, the latest version of the policy file is first obtained from the data node. If the latest version fails to be accessed, the version tree is traversed to obtain the most recent historical version.

3. The method for managing a high availability policy engine according to claim 1, characterized in that: S2, according to the operation mode and scheduling constraints defined in the policy file, schedules the policy instance to run on the target node, including: S21, reading the operation mode and scheduling constraint parameters defined in the strategy file as input of the scheduling algorithm; S22, collecting the current resource usage of each target node as the input of the scheduling algorithm; wherein the resource usage includes CPU and memory usage; S23, using a genetic algorithm to solve the optimal scheduling solution, allocating and deploying the policy instance to the target node while satisfying the node resource and scheduling constraints; S24, according to the optimal scheduling plan, the policy instance is distributed to the target node through the container orchestration Kubernetes. Each target node starts or updates the corresponding policy instance through the container engine according to the received instructions and allocates computing resources.

4. The method for managing a high availability policy engine according to claim 3, characterized in that: S23, using genetic algorithm to solve the optimal scheduling solution, including: According to the number of strategy instances and the number of candidate target nodes, the parameters of the genetic algorithm are set, including the initial population size, crossover probability, mutation probability and termination generation number; A priority-based encoding strategy is adopted to encode the deployment mapping relationship from policy instances to target nodes into an N-tuple as the chromosome of a scheduling scheme; where N is the number of policy instances, and each component of the tuple represents the number of a policy instance deployed to the target node; Randomly generate chromosomes to form an initial population; Calculate the fitness of each chromosome corresponding to the scheduling scheme; Using the tournament selection algorithm, a group of chromosomes are randomly selected from the current population, and the pair of chromosomes with the highest fitness is used as the father. The father chromosomes are recombined by multi-point crossover according to the set crossover probability to generate new offspring chromosomes. Perform mutation operation on the generated offspring chromosome according to the mutation probability, randomly select one or more gene positions, and randomly replace the values ​​of the selected gene positions with the numbers of other candidate nodes to obtain the mutated offspring chromosome; Add the offspring chromosomes obtained by crossover recombination and the mutated offspring chromosomes to the next generation population, and repeat the crossover and mutation operations until the size of the new population reaches the initial population size; Repeat the above iterations until the number of iterations reaches the set termination generation, decode the chromosome with the highest fitness in the final population, and obtain the optimal scheduling plan.

5. The method for managing a high availability policy engine according to claim 4, characterized in that: A priority-based encoding strategy is used to encode the deployment mapping relationship from the policy instance to the target node into an N-tuple as the chromosome of a scheduling scheme, including: Set the priority of the policy instance according to its execution frequency; Sort the tuple components corresponding to the policy instances according to the set priority; According to the number of candidate target nodes, determine the value range of each gene position in the chromosome; Grayscale coding is used to map the number of each candidate target node into a binary string as the genetic code of the corresponding node; The gene codes of all candidate target nodes are concatenated in sequence according to the order of tuple components to form a complete chromosome; each gene position of the chromosome corresponds to a deployment node of a strategy instance.

6. The method for managing a high availability policy engine according to claim 5, characterized in that: Calculate the fitness of each chromosome corresponding to the scheduling scheme: ; Among them, F is the fitness function value, and its value range is [0, 1]; are the weight coefficients of the three optimization objectives of load balancing, criticality, and failure rate; G is the Gini coefficient, and the calculation formula is: ; Where N is the number of candidate target nodes, and are the loads of the i-th node and the j-th node respectively; SLA is the service level agreement satisfaction rate of the key policy instance, and the calculation formula is: ; Where S is the number of key strategy instances, is the weight coefficient of the sth key strategy instance, is the SLA satisfaction indicator variable of the sth key policy instance; P is the failure rate of the scheduling scheme, and the calculation formula is: ; Where N is the number of candidate target nodes, is the failure rate of the nth node, The ratio of the number of policy instances deployed on the nth node to the total number of instances.

7. The method for managing a high availability policy engine according to claim 1, characterized in that: S5, using anomaly detection algorithms to determine whether the policy instance has anomalies, and predict the usage trend of node resources by the policy instance in the future, including: S51, constructing a resource usage time series of the policy instance according to the performance parameters collected in step S4, the time series including CPU usage, memory usage and response time indicators; S52, using the One Class SVM model to perform anomaly detection on the time series of each indicator; wherein the One Class SVM model determines whether the newly collected data point is abnormal through feature vector mapping and boundary learning; S53, judging whether the policy instance has an anomaly through an anomaly state machine according to the anomaly detection results of the multiple indicators; S54, using LSTM neural network to predict resource usage trends in the future, the input of LSTM neural network is the time series of multiple indicators, and the output is the predicted value in the future; S55, judging the load of the policy instance in the future based on the predicted node resource usage trend.

8. The method for managing a high availability policy engine according to claim 7, characterized in that: S52, uses the One Class SVM model to perform anomaly detection on the time series of each indicator, including: The time series data is segmented by sliding windows, and the statistical features of each window are extracted to form a statistical feature vector; Obtain historical sequence data when the strategy instance is operating normally and build a training data set; The One Class SVM model is trained using the training data set. The One Class SVM model maps the feature vectors in the training data set from the original space to the high-dimensional space through the kernel function. In the high-dimensional space, the One Class SVM model obtains the minimum hypersphere containing the number of training samples greater than the threshold, and uses the minimum hypersphere as the decision boundary for distinguishing normal samples from abnormal samples. Input the newly collected time series data into the trained One Class SVM model to obtain the mapped feature vector; Calculate the distance from the mapped eigenvector to the minimum hypersphere; The calculated distance is compared with a preset threshold. If it exceeds the preset threshold, the time series data of the corresponding time window is judged as abnormal.

9. The method for managing a high availability policy engine according to claim 7, characterized in that: S53, judging whether the policy instance has an abnormality through an abnormal state machine according to the abnormality detection results of the multiple indicators, including: Obtain the results of anomaly detection of time series data of multiple indicators. The results include the abnormal status of each indicator in each time window. The abnormal status is represented by a Boolean value, 1 for abnormality and 0 for normality. According to the abnormal status of each indicator in each time window, calculate the comprehensive abnormality of the strategy instance in the corresponding time window; Arrange the comprehensive abnormality of the strategy instance in multiple consecutive time windows in chronological order to form an abnormality time series; The pre-built abnormal state machine model is used to determine whether the obtained abnormality degree time series has an abnormality. The abnormal state machine model is constructed using a directed graph. The nodes in the graph represent the abnormal states of the policy instances, and the edges represent the transition conditions and probabilities between the abnormal states.

10. A management system for a high-availability policy engine, characterized in that: include: At least one processing unit; used to execute instructions to implement the management method of the high-availability policy engine as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Distributed cloud container resource scheduling method and system

    CN117909083A

  • Container scheduling method applied to container arrangement system

    CN118964015A

  • Container-based strategy arrangement response method and system and computer storage medium

    CN119440733A

  • Method and system for batch scheduling uniform parallel machines with different capacities based on improved genetic algorithm

    US20180356803A1