Server cluster security control method combined with multi-level encryption strategy joint scheduling
By combining multi-level encryption strategies and joint scheduling methods with reinforcement learning, combined with zero-trust architecture and deep learning algorithms, the problems of low resource scheduling efficiency, insufficient data security protection, and insufficient ability to deal with dynamic security threats in the server cluster are solved, and efficient resource utilization and strong security protection are achieved.
Patent Information
- Application Number
- CN202510423341.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art has low resource scheduling efficiency, insufficient data security protection, and insufficient ability to deal with dynamic security threats in server clusters.
Using a method that combines multi-level encryption strategies and reinforcement learning joint scheduling, the security response mechanism is automatically implemented by evaluating the sensitivity of data in the cluster, dynamically adjusting the encryption strategy and resource allocation, combining the zero-trust architecture and deep learning algorithm for real-time monitoring and abnormal detection.
Customized encryption for data with different sensitivity is realized, the flexibility and security of data encryption is improved, resource utilization efficiency is optimized, the security protection capabilities of the cluster are enhanced, and dynamic security threats can be effectively dealt with.
Smart Images

Figure CN119939637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a server cluster security control method combined with multi-level encryption strategy joint scheduling. Background Art
[0002] With the widespread application of server clusters in fields such as big data and cloud computing, efficient management of cluster resources and data security protection have become key issues. In traditional cluster management solutions, resource scheduling often relies on static configurations and rules, and cannot be dynamically adjusted according to cluster load changes, task requirements, and data sensitivity. This static resource scheduling method not only fails to effectively utilize cluster resources, but also easily leads to computing node overload or resource idleness, affecting the overall performance of the cluster.
[0003] In terms of data security, existing encryption technologies usually adopt a unified encryption strategy, and there is no flexible encryption scheme for data of different sensitivities, resulting in high performance overhead for computing tasks. In addition, traditional encryption methods often have performance bottlenecks in large-scale clusters, especially when processing large amounts of data, where performance loss is more obvious.
[0004] Traditional protection methods for cluster security issues usually rely on trust-based models, ignoring fine-grained access control and dynamic permission management, and cannot effectively prevent unauthorized access and potential internal threats. In addition, the various security threats and complex attack modes in cluster environments pose great challenges to traditional security protection measures, and more intelligent and efficient security control strategies are urgently needed. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a server cluster security control method combined with multi-level encryption strategy joint scheduling, which solves the problems of low resource scheduling efficiency, insufficient data security protection and insufficient ability to cope with dynamic security threats in server clusters.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a server cluster security control method combined with multi-level encryption strategy joint scheduling, comprising the following steps: S1. Perform sensitivity assessment on computing tasks and data within the cluster, and classify data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data according to predetermined rules; S2. The first encryption algorithm is used to encrypt low-sensitivity data, the second encryption algorithm is used to encrypt medium-sensitivity data, and the homomorphic encryption algorithm is used to encrypt high-sensitivity data; S3, a joint scheduling mechanism based on reinforcement learning, dynamically adjusts the allocation of computing resources and task scheduling by obtaining cluster load, computing task requirements, and data sensitivity information in real time; S4. Deploy a zero-trust architecture in the cluster so that each request must be authenticated and authorized, and fine-grained control of task permissions is performed through role-based access control policies; S5, monitor computing tasks and data transmission in the cluster in real time, use deep learning algorithms to detect abnormal cluster behavior, and identify potential security threats; S6. When a potential threat is detected, the security response mechanism is automatically executed to provide protection by isolating abnormal nodes and limiting resource access of malicious tasks.
[0007] Preferably, the sensitivity assessment of the data in the cluster in step S1 specifically includes the following steps: S1.1. Collect data of computing tasks in the cluster and analyze data types, contents and access frequency; S1.2. Use a trained machine learning model or specified statistical method to assess the sensitivity of the data, and classify the data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data based on the assessment results; S1.3. Classify the data according to the evaluation results and specify the corresponding encryption algorithm and security strategy according to the category; S1.4. Update the sensitivity level of data regularly or in real time according to changes in cluster resources, and dynamically adjust the encryption strategy.
[0008] Preferably, the step S2 specifically includes the following steps: S2.1. Encrypt low-sensitivity data using the AES-128 encryption algorithm to generate corresponding encrypted ciphertext; S2.2, encrypt the sensitive data using the AES-256 encryption algorithm to generate the corresponding encrypted ciphertext; S2.3. Encrypt highly sensitive data using a homomorphic encryption algorithm, where the homomorphic encryption algorithm is a Paillier encryption algorithm; S2.4. All encryption processes generate and manage keys through a centralized key management system, including key lifecycle management, key storage and distribution.
[0009] Preferably, the joint scheduling mechanism based on reinforcement learning in step S3 includes: S3.1. Collect the resource usage of the current cluster, including the load, memory and CPU usage of the computing nodes, as well as the computing requirements and priorities of the tasks; S3.2. Based on deep Q learning, dynamically adjust the task scheduling strategy and resource allocation strategy while satisfying the resource allocation constraints; S3.3. According to the encryption level, data sensitivity and computing resource requirements of the task, select a computing node with idle resources that can meet the encryption computing requirements; S3.4. When scheduling computing tasks, consider resource utilization, load balancing, and security.
[0010] Preferably, the implementation of the zero trust architecture in step S4 specifically includes: S4.1. Implement OAuth2.0 or OpenIDConnect-based authentication for each request within the cluster. S4.2. Implement attribute-based access control to control access rights based on the sensitivity of the task, the identity and role information of the requester; S4.3. Each access request is subject to permission verification; S4.4. Dynamically adjust permissions so that permissions are updated in real time based on task requirements and access conditions during the task life cycle.
[0011] Preferably, the step S5 specifically includes: S5.1. Use the Prometheus monitoring system to monitor the computing node resource usage of the cluster in real time, and visualize it through Grafana; S5.2. Use deep learning algorithms to model cluster behavior and identify abnormal patterns that deviate from normal behavior; S5.3. Use model-based anomaly detection methods to automatically trigger alarms and notify administrators when cluster resources or task behaviors are abnormal; S5.4. Combine the monitoring system with the anomaly detection results to generate real-time security reports to help administrators take timely response measures.
[0012] Preferably, the security response mechanism in step S6 includes: S6.1. When a security threat is detected, the malicious node is automatically isolated and the computing task on the node is stopped to prevent the security threat from spreading. S6.2. Limit resource access rights of affected nodes or tasks to reduce their interference with normal computing tasks; S6.3. Automatically adjust cluster access permissions based on the severity of abnormal behavior to prohibit unauthorized users or nodes from accessing sensitive data. S6.4. Keep detailed records of all security incidents, generate audit logs, and provide them to administrators for subsequent analysis and investigation.
[0013] Preferably, the joint scheduling mechanism of reinforcement learning in step S3 optimizes the scheduling strategy through the following mathematical model: ; Where R(t) represents the reward value of the scheduling strategy, ri represents the available capacity of resource i, t j represents the demand of task j, x ij is a Boolean variable indicating whether task j is assigned to resource i, α and β are regularization coefficients, n is the number of resource nodes, m is the number of tasks, and d j is the resource requirement of task j, p j is the priority of task j.
[0014] The present invention provides a server cluster security control method combined with multi-level encryption strategy joint scheduling. It has the following beneficial effects: 1. The present invention implements encryption protection for data of different sensitivity levels by performing sensitivity assessment on computing tasks and data within the cluster and selecting a suitable encryption algorithm based on the assessment results. Compared with the traditional unified encryption scheme, the multi-level encryption strategy of the present invention can perform customized encryption according to the specific sensitivity of the data, greatly improving the flexibility and security of data encryption. Through this flexible encryption management, the encryption burden of low-sensitivity data can be effectively reduced, ensuring that highly sensitive data is more strongly protected, thereby improving resource utilization efficiency while ensuring the overall security of the cluster.
[0015] 2. The present invention can obtain cluster load, task requirements and data sensitivity information in real time and dynamically adjust the allocation of computing resources and task scheduling by introducing reinforcement learning algorithms, especially the joint scheduling mechanism based on deep Q learning. Compared with traditional static resource allocation or simple scheduling algorithms, the dynamic scheduling mechanism based on reinforcement learning can optimize according to the real-time changing workload and task requirements, automatically select the most appropriate computing resources for task scheduling, and improve the utilization efficiency of cluster resources and task processing capabilities.
[0016] 3. The present invention combines zero-trust architecture with deep learning algorithms to achieve strict access control and real-time abnormal behavior detection within the cluster. All requests must be authenticated and authorized. At the same time, fine-grained access control is performed based on the sensitivity of the task and the user's role information. Using deep learning algorithms, the system can monitor cluster behavior in real time and identify abnormal patterns. Once a potential security threat is discovered, it automatically triggers a security response mechanism, such as isolating abnormal nodes, restricting malicious task resource access, etc., thereby effectively preventing the spread of security threats. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] Please refer to the attached Figure 1 The embodiment of the present invention provides a server cluster security control method combined with multi-level encryption strategy joint scheduling, comprising the following steps: S1. Perform sensitivity assessment on computing tasks and data within the cluster, and classify data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data according to predetermined rules; The sensitivity assessment of the data in the cluster in step S1 specifically includes the following steps: S1.1. Collect data of computing tasks in the cluster and analyze data types, contents and access frequency; S1.2. Use a trained machine learning model or specified statistical method to assess the sensitivity of the data, and classify the data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data based on the assessment results; S1.3. Classify the data according to the evaluation results and specify the corresponding encryption algorithm and security strategy according to the category; S1.4. Update the sensitivity level of data regularly or in real time according to changes in cluster resources, and dynamically adjust the encryption strategy.
[0020] Specifically, in step S1.1, the system first needs to collect all relevant computing task data from each computing node in the cluster. Data collection not only includes the type and content of the data, but also considers the access frequency, access pattern, and historical usage of the data. For example, some data may be frequently accessed by different computing nodes, while other data may only be used in specific tasks. Therefore, access frequency and pattern are one of the important bases for sensitivity assessment. This information can be obtained in real time through a cluster management system or a dedicated log analysis system to ensure the comprehensiveness and accuracy of data collection.
[0021] In step S1.2, based on the collected task data, the system uses trained machine learning (such as support vector machines, decision trees, random forests, etc.) or statistical methods to evaluate the sensitivity of the data. Common sensitivity assessment models can comprehensively derive the sensitivity level of each data based on characteristics such as data type (such as financial data, user information), content (such as sensitive fields, privacy data), and access frequency. The model in the data assessment process can be trained with historical data and continuously adjusted and optimized to improve the accuracy of the assessment.
[0022] The output of the sensitivity assessment divides the data into three categories: Low-sensitivity data: This data may be public data or data with low security requirements, and can be protected by conventional encryption or simple access control measures.
[0023] Medium sensitive data: This data requires moderate protection measures and is processed using strong encryption algorithms and appropriate access controls.
[0024] Highly sensitive data: These data are extremely sensitive core data and require the highest level of protection. They are usually encrypted using advanced encryption technologies such as homomorphic encryption, and their access rights are strictly controlled. Formula: Assume that the data set is ; where d i For the i-th data, the data sensitivity assessment is performed using the following model: ; Among them, S (d i ) represents data d i The sensitivity assessment result of is, f is the trained sensitivity assessment model, and the output is the sensitivity level of the data.
[0025] Based on the evaluation results, the system classifies the data according to its sensitivity level. Specifically, for different categories of data, the system specifies different encryption algorithms and access control policies: For low-sensitivity data, simpler encryption algorithms such as AES-128 are used and protected through basic authentication and access control; For medium-sensitive data, the AES-256 encryption algorithm is used to protect it, and stricter access rights are configured for the data; For highly sensitive data, homomorphic encryption algorithms (such as the Paillier encryption algorithm) are used for encryption to ensure that the data can still be processed and calculated even in an encrypted state.
[0026] In this step, the system matches the encryption algorithm with the sensitivity level of the data and ensures that each data category has corresponding protection measures. At the same time, the system will allocate appropriate computing resources and scheduling strategies for different levels of data according to task requirements to ensure that encryption calculations will not have a significant impact on cluster performance.
[0027] As computing tasks in the cluster change and data access patterns adjust, the sensitivity of the data may change. Therefore, in this step, the system needs to update the sensitivity level of the data regularly or based on real-time conditions. This process is achieved through continuous monitoring and data evaluation to ensure that the data is always at the appropriate security protection level throughout its life cycle.
[0028] During the update process, the system will dynamically adjust the data encryption policy based on the new sensitivity assessment results. For example, when the access frequency of certain medium-sensitive data increases significantly, the system may upgrade it to highly sensitive data and switch to a stronger encryption algorithm. At the same time, the system will optimize the resource allocation of encryption computing based on the priority of the task and the changes in computing resources to ensure a balance between data protection and cluster resources.
[0029] This embodiment combines machine learning methods and statistical analysis to evaluate the sensitivity of data, and implements a multi-level encryption strategy based on the evaluation results, so that the data in the cluster can be properly protected according to different security requirements. During the sensitivity evaluation process, the system combines multiple factors such as task type and data access frequency, and uses an intelligent evaluation model to ensure the accuracy of sensitivity classification. This method has high flexibility and scalability, and can effectively respond to various security threats in a cluster environment.
[0030] S2. The first encryption algorithm is used to encrypt low-sensitivity data, the second encryption algorithm is used to encrypt medium-sensitivity data, and the homomorphic encryption algorithm is used to encrypt high-sensitivity data; The S2 step specifically includes the following steps: S2.1. Encrypt low-sensitivity data using the AES-128 encryption algorithm to generate corresponding encrypted ciphertext; S2.2, encrypt the sensitive data using the AES-256 encryption algorithm to generate the corresponding encrypted ciphertext; S2.3. Use homomorphic encryption algorithm to encrypt highly sensitive data. The homomorphic encryption algorithm is Paillier encryption algorithm. S2.4. All encryption processes generate and manage keys through a centralized key management system, including key lifecycle management, key storage and distribution.
[0031] Specifically, in this embodiment, step S2 encrypts the data, which specifically includes the following contents: For low-sensitivity data, the AES-128 (Advanced Encryption Standard, 128-bit key length) encryption algorithm is used. AES-128 is a symmetric encryption algorithm with a relatively simple encryption process and fast calculation speed. It is suitable for scenarios with large data volumes and relatively low security requirements. Although data encrypted by the AES-128 algorithm has certain protection capabilities, due to its low encryption strength, it is suitable for protecting non-sensitive or non-abused data, such as some public system logs or other data with low security requirements.
[0032] In this step, the system uses a 128-bit key to encrypt low-sensitivity data and generate the corresponding encrypted ciphertext. The encrypted data will be stored or transmitted to the computing nodes of the cluster. During the transmission process, the data will be protected to prevent unauthorized access.
[0033] Sensitive data usually includes some important information in the business (such as user information, some commercial data, etc.), so stronger encryption measures are required. Compared with AES-128, AES-256 uses a 256-bit key for encryption, providing stronger security protection and can effectively prevent attacks such as brute force cracking.
[0034] In this step, the system generates a 256-bit key for the medium-sensitive data and encrypts the data using the AES-256 algorithm. The data ciphertext encrypted by AES-256 will be more secure and can withstand stronger attacks, which is suitable for protecting data that has a certain degree of confidentiality but is not extremely sensitive.
[0035] Highly sensitive data is the most confidential and most protected type of data in the cluster, such as financial data, personal privacy data, core business data, etc. In order to ensure that these data can still be calculated and analyzed after being encrypted, this embodiment uses a homomorphic encryption algorithm, especially the Paillier encryption algorithm.
[0036] Homomorphic encryption is a special type of encryption that allows certain calculations to be performed directly on encrypted data without decryption, thus avoiding the risk of exposing data during processing. The Paillier encryption algorithm is an additive homomorphic encryption algorithm that is particularly suitable for scenarios where addition operations need to be performed on encrypted data, such as summing encrypted financial data without decryption. The Paillier algorithm encrypts and decrypts by generating public and private keys, the public key is used to encrypt data, and the private key is used to decrypt data.
[0037] Through this encryption algorithm, highly sensitive data can be calculated in an encrypted state, preventing data leakage while ensuring the security and accuracy of the calculation results.
[0038] Key management is the core part of the encryption system. In this step, all encryption processes rely on a centralized key management system to generate, distribute, and store keys. The key management system not only generates the keys required by the encryption algorithm (such as AES-128, AES-256, Paillier keys), but is also responsible for managing the key life cycle, including key creation, update, revocation, storage, and distribution.
[0039] Specifically, the key management system needs to ensure the secure storage of keys to prevent them from being accessed or leaked by unauthorized users. To this end, keys are usually stored in hardware security modules (HSMs) or key management hardware to provide a high level of protection. The key management system also needs to ensure that the key distribution process is secure and only authenticated and authorized nodes can obtain the keys.
[0040] In addition, the key management system will regularly update keys to reduce the risk of key cracking. For keys that are expired or no longer needed, the system will revoke them to ensure the validity and security of the keys.
[0041] This embodiment uses a multi-level encryption strategy to apply AES-128, AES-256 and Paillier homomorphic encryption algorithms to low-sensitivity data, medium-sensitivity data and high-sensitivity data in the cluster respectively. Different encryption algorithms ensure the security of data of different sensitivity levels while minimizing the consumption of computing resources. All encryption processes use a centralized key management system to generate, store, distribute and manage the life cycle of keys to ensure the security and effectiveness of the keys. This solution provides a flexible and efficient solution for data security in a cluster environment, which helps to optimize the efficiency of cluster resource utilization while ensuring security.
[0042] S3, a joint scheduling mechanism based on reinforcement learning, dynamically adjusts the allocation of computing resources and task scheduling by obtaining cluster load, computing task requirements, and data sensitivity information in real time; The joint scheduling mechanism based on reinforcement learning in step S3 includes: S3.1. Collect the resource usage of the current cluster, including the load, memory and CPU usage of the computing nodes, as well as the computing requirements and priorities of the tasks; S3.2. Based on deep Q learning, dynamically adjust the task scheduling strategy and resource allocation strategy while satisfying the resource allocation constraints; S3.3. According to the encryption level, data sensitivity and computing resource requirements of the task, select a computing node with idle resources that can meet the encryption computing requirements; S3.4. When scheduling computing tasks, consider resource utilization, load balancing, and security.
[0043] The joint scheduling mechanism of reinforcement learning in step S3 optimizes the scheduling strategy through the following mathematical model: ; Where R(t) represents the reward value of the scheduling strategy, r i represents the available capacity of resource i, t j represents the demand of task j, x ijis a Boolean variable indicating whether task j is assigned to resource i, α and β are regularization coefficients, n is the number of resource nodes, m is the number of tasks, and d j is the resource requirement of task j, p j is the priority of task j.
[0044] Specifically, in this embodiment, the resource usage of the cluster includes the load of each computing node, the usage of memory and CPU, and the computing requirements and priority information of the current computing task. Specifically: Compute node load collection: The cluster management system monitors the resource usage of each node in real time. The load status of computing node i is calculated by cpu i and mem i Two indicators are used to measure. The load function of the node can be expressed as: ; Where f is the load function, which means that the load of the node is comprehensively judged by calculating the CPU and memory resource consumption of the node.
[0045] Task computing requirements and priorities: Each computing task j has its computing resource requirements ;in: and are the CPU and memory resource requirements of task j, respectively, j The priority of the task, the task with higher priority will be scheduled first.
[0046] In this way, the system can accurately collect the resource usage and task demand status of all nodes in the cluster, providing data support for subsequent scheduling decisions.
[0047] After collecting cluster resources and task information, the system then uses the deep Q learning method in reinforcement learning to implement dynamic task scheduling and resource allocation. The specific steps are as follows: Definition of state space: The system state space contains the resource status of all computing nodes and the demand information of the tasks to be scheduled. t It can be represented as the resource status of all nodes in the cluster and the collection of tasks currently to be scheduled: ;in, ; indicates the resource status of the i-th node, ; represents the resource requirements and priority of the jth task.
[0048] Definition of action space: Action space a t represents the choice of the current scheduling decision. For each task j, the system can choose whether to assign it to node i action x ij is a Boolean variable indicating whether task j is assigned to resource node i: ; Design of reward function: The goal of the reward function is to maximize the efficiency of cluster resource utilization, optimize load balancing, and consider the priority of tasks. According to the task allocation results, the reward function R(t) can be defined by the following formula: ; Where R(t) represents the reward value of the scheduling strategy, r i represents the available capacity of resource i, t j represents the demand of task j, x ij is a Boolean variable indicating whether task j is assigned to resource i, α and β are regularization coefficients, n is the number of resource nodes, m is the number of tasks, and d j is the resource requirement of task j, p j is the priority of task j.
[0049] This reward function aims to optimize resource scheduling by considering resource utilization and task priority, so that tasks can be effectively scheduled to appropriate computing nodes.
[0050] Q-value update rule: In reinforcement learning, Q-value represents the expected reward of taking a certain action in a certain state. After each scheduling decision, the system updates the Q-value based on the reward obtained. The update of Q-value follows the classic Q-learning formula: ; Among them: α is the learning rate, which controls the speed of learning, γ is the discount factor, which determines the importance of future rewards, and r t is the reward obtained at the current time step. ; is the next state s t+1 The maximum Q value of all possible actions.
[0051] In this embodiment, the scheduling of tasks not only considers their computing resource requirements and priority, but also the encryption level and data sensitivity of the tasks. According to the different encryption requirements of the tasks, the system schedules the tasks to the appropriate computing nodes to meet the security and computing requirements: Encryption level: Low-sensitivity tasks use standard encryption (such as AES-128), medium-sensitivity tasks use stronger encryption (such as AES-256), and high-sensitivity tasks use homomorphic encryption algorithms (such as Paillier encryption).
[0052] Encryption computing resource requirements: Since encryption algorithms consume different computing resources, the system dynamically calculates its resource requirements based on the encryption level of the task. Encryption computing requirements It can be calculated by the following model: ; Among them, f enc is the encryption requirement function, which represents the resource consumption of tasks under different encryption levels.
[0053] Resource allocation: Based on the encryption requirements of the task, the system selects computing nodes that can provide sufficient computing power to ensure that the encryption task can be completed smoothly.
[0054] During the scheduling process, the system takes the following factors into consideration: Resource utilization and load balancing: To avoid overloading certain nodes in the cluster, the system will make a trade-off between load balancing and resource utilization. For example, when the load of a node exceeds 80%, the system will avoid assigning new tasks to it.
[0055] Security: For highly sensitive tasks, the system will give priority to scheduling nodes with reinforced security protection to ensure that the transmission and calculation process of task data meets security requirements.
[0056] S4. Deploy a zero-trust architecture in the cluster so that each request must be authenticated and authorized, and fine-grained control of task permissions is performed through role-based access control policies; The implementation of zero trust architecture in step S4 specifically includes: S4.1. Implement OAuth2.0 or OpenIDConnect-based authentication for each request within the cluster. S4.2. Implement attribute-based access control to control access rights based on the sensitivity of the task, the identity and role information of the requester; S4.3. Each access request is subject to permission verification; S4.4. Dynamically adjust permissions so that permissions are updated in real time based on task requirements and access conditions during the task life cycle.
[0057] Specifically, in this embodiment, in order to ensure that each request undergoes strict identity authentication, the cluster uses an authentication protocol based on OAuth2.0 or OpenIDConnect for identity authentication. These protocols can effectively confirm the identity of the user and ensure the legitimacy of the identity of the requester.
[0058] OAuth2.0 authentication: The OAuth2.0 protocol allows client applications to access resource servers on behalf of users. The client must obtain an authorization token from the authorization server as a credential for the request. Each time a cluster resource is requested, the client needs to carry an OAuth2.0 access token, and the server verifies its identity with the authorization server through the token.
[0059] OpenIDConnect authentication: OpenIDConnect is an extension of the OAuth2.0 protocol that provides additional authentication capabilities. Through OpenIDConnect, the cluster can verify the identity of the requester and obtain the user's basic information (such as name, email, role, etc.). Each request will be authenticated by the identity provider and an ID token will be generated for subsequent permission verification.
[0060] Authentication process: The requester first requests authentication from the authentication server through the OAuth2.0 or OpenIDConnect protocol and obtains an identity token.
[0061] The token is sent along with the request to the cluster server, and the server confirms the identity of the requester by verifying the validity of the token.
[0062] If the identity authentication is successful, the subsequent permission authorization stage will be entered; if the verification fails, the request will be rejected and the corresponding error message will be returned.
[0063] Attribute-based access control is a more flexible and dynamic permission control mechanism that determines access rights by evaluating multiple attributes of the requester (such as identity, role, task sensitivity, etc.). Compared with traditional role-based access control, access control can provide more refined permission management.
[0064] Access control policy: In attribute-based access control, the access control policy relies on a set of defined attributes, which can be divided into the following categories: Requester attributes: including user identity information, role, group, etc.
[0065] Task attributes: including the sensitivity level of the task, the life cycle status of the task, etc.
[0066] Environmental attributes: including request time, geographic location, request source, etc.
[0067] Policy example: For example, a highly sensitive task (task sensitivity is high) can only be accessed by a specific user group (such as the system administrator group) and can only be accessed during working hours (such as 9:00-18:00). This policy can be defined as: ; Through such policy definitions, attribute-based access control can ensure that permission decisions are made dynamically and flexibly in the cluster based on the specific circumstances of the tasks and requesters.
[0068] Implementation method: The access control system in the cluster will match the request with the relevant attributes and determine whether to allow access based on the preset access control policy.
[0069] If the requester meets all policy conditions, the system will allow the request to be executed, otherwise it will return a response denying access.
[0070] In a zero-trust architecture, every request must be verified for permissions, regardless of whether the request is legitimate or from a known user. All requested access permissions are temporary and dynamic, so each access will trigger permission verification.
[0071] Permission verification process: Whenever the cluster receives a request, the system will perform permission verification based on the context of the current request (including the identity of the requester, the content of the request, the sensitivity of the task, etc.).
[0072] Permission verification usually includes the following aspects: Authentication: Checks whether the requester has successfully authenticated via the OAuth2.0 or OpenIDConnect protocol.
[0073] Role and attribute matching: Check whether the requester has sufficient permissions to access the target resource through attribute-based access control or role-based access control model.
[0074] Task sensitivity and access control: Based on the sensitivity of the task and the attributes of the requester, further determine whether the access conditions are met.
[0075] Verification failure: If the permission verification fails, the system will immediately terminate the execution of the request and return the corresponding error code (such as "403Forbidden") to the requester.
[0076] In order to better adapt to changes in the task life cycle and different access modes, permissions need to be adjusted in real time according to task requirements and access conditions. The system can dynamically update permissions to ensure the security of tasks at different stages.
[0077] Trigger conditions for permission adjustment: When the sensitivity of a task changes, the system will automatically adjust the access rights. For example, when a task enters a highly sensitive stage, the system will increase the strictness of the permission and restrict access to users with specific roles.
[0078] When the access frequency of a task increases, the system may temporarily adjust resources and permissions based on the load situation to prevent security risks caused by over-authorization or access to too many resources.
[0079] Implementation method: Based on the life cycle of the task and the access log, the cluster management system will dynamically update the access rights according to the preset rules. For example, if the task enters a sensitive stage, the permission check level will be automatically increased.
[0080] In addition, the cluster will record a log of each permission change for auditing and tracking. Each permission adjustment operation will generate an audit log and synchronize it to the management platform in real time for monitoring and management.
[0081] S5, monitor computing tasks and data transmission in the cluster in real time, use deep learning algorithms to detect abnormal cluster behavior, and identify potential security threats; Step S5 specifically includes: S5.1. Use the Prometheus monitoring system to monitor the computing node resource usage of the cluster in real time, and visualize it through Grafana; S5.2. Use deep learning algorithms to model cluster behavior and identify abnormal patterns that deviate from normal behavior; S5.3. Use model-based anomaly detection methods to automatically trigger alarms and notify administrators when cluster resources or task behaviors are abnormal; S5.4. Combine the monitoring system with the anomaly detection results to generate real-time security reports to help administrators take timely response measures.
[0082] Specifically, in order to achieve real-time monitoring of cluster computing tasks and resources, this embodiment uses Prometheus as the cluster resource monitoring system. Prometheus can regularly pull resource usage indicators of each node in the cluster and save them as time series data. Grafana is used to visualize Prometheus monitoring data so that administrators can intuitively understand the health status of the cluster.
[0083] Prometheus monitoring system: Prometheus is an open source monitoring system that can regularly collect resource usage of each node and support efficient storage, query and alarm. It mainly monitors the following aspects: Compute node resource usage: such as CPU usage, memory usage, disk IO, network bandwidth, etc.
[0084] Task execution status: such as task running time, status, resource requirements, etc.
[0085] Prometheus collects and saves the resource information of each computing node through "exporters" to generate time series data.
[0086] Grafana visualization: Grafana is a visualization tool for Prometheus that presents the real-time collected resource usage data to administrators in the form of charts, dashboards, etc. Through Grafana, administrators can see the resource usage of each node in real time and discover potential performance bottlenecks or anomalies.
[0087] Implementation process: Deploy Prometheus, install a monitoring agent or exporter on each computing node, and configure Prometheus to pull data regularly.
[0088] Configure Grafana to connect with Prometheus, and use Grafana to configure a visual dashboard to display the system's operating status, resource usage, task execution, etc.
[0089] Generate monitoring reports regularly to facilitate administrators to do performance analysis.
[0090] In order to more efficiently identify abnormal behavior in the cluster, the system uses deep learning algorithms to model cluster behavior, identify patterns that deviate from normal behavior, and then detect potential security threats.
[0091] Construction of deep learning models: Resource monitoring data, task execution data, network traffic, etc. collected in the cluster are used as inputs to train models using deep learning algorithms. These models can learn normal operation modes and behavior patterns.
[0092] Autoencoder: With the autoencoder model, the normal mode of cluster behavior is measured by the reconstruction loss of the input and output. When abnormal behavior occurs, the reconstruction error will increase significantly, so it can be used to detect behavioral anomalies.
[0093] Training process: First, collect a large amount of normal operating cluster data, including computing resource usage, task status, data transmission, etc.
[0094] A deep learning algorithm (such as an LSTM model) is trained on this data to obtain a "normal behavior model" of the cluster.
[0095] In actual operation, real-time monitoring data is input into the deep learning model, and the loss value is calculated to determine whether there is an abnormality.
[0096] Abnormal pattern recognition: When the behavior of real-time monitoring data deviates from the normal behavior learned by the model, the model will output a large anomaly score, indicating that the current cluster is abnormal.
[0097] These anomalies include but are not limited to: abnormal computing node load, sudden increase in network traffic, long task execution time, abnormal data transmission speed, etc.
[0098] The model-based anomaly detection method determines whether an anomaly occurs by comparing the cluster behavior with the trained normal behavior model. When an anomaly is found, the system can automatically trigger an alarm mechanism to notify the administrator to take timely countermeasures.
[0099] Anomaly detection method: The system compares the real-time monitoring data with the normal pattern in the deep learning model to determine whether there is an anomaly. If there is a large difference between the output predicted by the model and the actual monitoring data, it means that the cluster behavior deviates from the normal pattern and is judged as abnormal.
[0100] Threshold-based alarm: When the anomaly score of a resource exceeds the set threshold, an alarm is triggered.
[0101] Automatic alarm mechanism: The system can automatically trigger alarms based on pre-defined policies and thresholds. These alarms include: System performance alarm: such as high CPU usage, abnormal memory usage, etc.
[0102] Task execution abnormality alarm: such as task execution timeout, excessive resource request, etc.
[0103] Network anomaly alarm: such as sudden increase in network traffic, abnormal data transmission behavior, etc.
[0104] Notification mechanism: Alarm information will be pushed to the administrator in real time, and the notification content includes the type of abnormality, scope of impact, specific indicators and possible solutions.
[0105] Administrators can be notified via multiple channels, such as email, SMS, instant messaging, etc.
[0106] The generation of real-time security reports is the final output of cluster monitoring and anomaly detection, helping administrators understand the current health status of the cluster and respond quickly to possible security incidents.
[0107] Report generation: During cluster operation, the monitoring system will collect the resource usage of each node in real time, and combine the deep learning anomaly detection results to generate a detailed security report. The report content usually includes: Current resource usage: CPU, memory, and network bandwidth usage of each node.
[0108] Anomaly detection results: whether there is abnormal behavior, the type and scope of the anomaly.
[0109] Alarm log: Detailed record of abnormal events occurring within the system.
[0110] Report format: Real-time security reports can be displayed in various forms such as charts, logs, warnings, etc., and the reports are sent to administrators or security operations teams immediately after they are generated.
[0111] The report includes analysis of abnormal behavior, potential risk assessment, and recommended countermeasures, such as restricting access to abnormal nodes and optimizing resource allocation.
[0112] Countermeasures: Reports help administrators quickly understand the health status of the cluster and promptly detect and respond to potential security threats.
[0113] For example, if abnormal network traffic causes security risks, administrators can immediately isolate related nodes or restrict access to certain services.
[0114] S6. When a potential threat is detected, the security response mechanism is automatically executed to provide protection by isolating abnormal nodes and limiting resource access of malicious tasks.
[0115] The S6 step security response mechanism includes: S6.1. When a security threat is detected, the malicious node is automatically isolated and the computing task on the node is stopped to prevent the security threat from spreading. S6.2. Limit the resource access rights of the affected nodes or tasks to reduce their interference with normal computing tasks; S6.3. Automatically adjust cluster access permissions based on the severity of abnormal behavior to prohibit unauthorized users or nodes from accessing sensitive data. S6.4. Keep detailed records of all security incidents, generate audit logs, and provide them to administrators for subsequent analysis and investigation.
[0116] Specifically, when the system detects a malicious node or a potential security threat, isolation measures must be taken quickly to prevent the threat from spreading to other computing nodes or affecting other parts of the system.
[0117] Malicious node detection: The system uses deep learning models, behavior analysis algorithms, and real-time monitoring data to determine whether a node has abnormal behavior, such as task timeout, abnormal resource usage, abnormal network traffic, etc.
[0118] If it is detected that the node behavior deviates significantly from the normal pattern and meets the characteristics of malicious activities (such as malicious scripts, etc.), the system will determine that the node is a "malicious node".
[0119] Automatically isolate malicious nodes: Once a node is determined to be malicious, the system will immediately isolate the node. Isolation measures include: Network isolation: Disconnect malicious nodes from the cluster's internal network to prevent them from communicating with other normal nodes.
[0120] Task stop: Forcefully stop all computing tasks on the node to prevent malicious tasks from being further executed.
[0121] Resource recycling: Immediately recycle the computing resources occupied by the node and release them to other nodes for normal task processing.
[0122] Prevent the spread of threats: By quickly isolating malicious nodes, malicious behavior can be prevented from spreading to other nodes in the cluster or affecting critical data processing tasks, ensuring the overall security of the cluster.
[0123] After the threat is detected, in addition to isolating the malicious nodes, it is also necessary to limit the resource access rights of the affected nodes or tasks to prevent them from interfering with normal computing tasks.
[0124] Resource access restrictions: For nodes or tasks that are determined to be affected, the system will take measures to restrict their resource access rights, including: CPU / Memory Limit: Set quota limits on the CPU and memory resources of the affected tasks to prevent them from occupying too many resources.
[0125] Network access restriction: Limit the network bandwidth or network access capability of the affected nodes to reduce their interference with other nodes.
[0126] Storage access restriction: Limit the access of affected nodes to the storage system to prevent malicious data tampering or leakage.
[0127] Reduce interference: By precisely controlling the resource usage of the affected nodes or tasks, it is possible to minimize the impact on other normally running computing tasks, while also effectively isolating potential security risks.
[0128] Based on the severity of the detected abnormal behavior, the system needs to dynamically adjust the access permissions in the cluster to ensure that sensitive data in the cluster is not accessed by unauthorized nodes or users.
[0129] Severity Assessment of Abnormal Behavior: The system will assess the severity of the detected abnormal behavior based on its nature and impact. For example, if a node has malicious code or network attack behavior, the system will assess the potential threat of the behavior to the data.
[0130] Severity assessment criteria: The severity of abnormal behavior can be assessed based on factors such as the duration of the abnormal behavior, the scope of impact, and the sensitivity of the task.
[0131] Access rights adjustment: When abnormal behavior is assessed to be serious, the system automatically adjusts access permissions based on the cluster's security policy: Prohibit unauthorized users from accessing sensitive data: If a node is identified as a malicious node, all requests from that node to access sensitive data will be prohibited.
[0132] Strengthened Authentication: Strengthen authentication for all users and nodes to ensure that only authorized nodes can access critical resources and sensitive data.
[0133] Restrict data transmission permissions: restrict data transmission between tasks to prevent sensitive data from being illegally accessed or leaked.
[0134] Dynamic adjustment: Access rights can be adjusted in real time and restored to normal access rights policy after abnormal behavior is fixed or threats are eliminated.
[0135] In the process of executing the security response mechanism, the system will record all security incidents in detail and generate audit logs for administrators to conduct subsequent analysis and investigation to ensure the transparency and traceability of the incidents.
[0136] Security event log: All operations related to security threat detection, isolation, permission adjustment, etc. will be recorded in the log. The record content includes: Event time: The specific time when the security incident occurred.
[0137] Event type: for example, "malicious node isolation", "task stop", "authority adjustment", etc.
[0138] Impact scope: affected nodes, tasks, or data.
[0139] Response measures: security response measures taken, such as network isolation, resource restrictions, etc.
[0140] Audit log generation: Audit logs will be stored in a secure log management system to ensure that the logs are not tampered with or deleted.
[0141] These log records can help administrators track and trace security events and further analyze the source, process, and impact of malicious attacks.
[0142] Subsequent analysis and investigation: Audit logs are not only an important reference for security response, but also provide data support for post-event analysis. Based on the log content, administrators can: Conduct in-depth analysis of malicious behavior and identify attack patterns.
[0143] Evaluate the effectiveness of cluster security policies and adjust security policies based on discovered vulnerabilities.
[0144] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A server cluster security control method combined with multi-level encryption strategy joint scheduling, characterized in that: The following steps are involved: S1. Perform sensitivity assessment on computing tasks and data within the cluster, and classify data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data according to predetermined rules; S2. The first encryption algorithm is used to encrypt low-sensitivity data, the second encryption algorithm is used to encrypt medium-sensitivity data, and the homomorphic encryption algorithm is used to encrypt high-sensitivity data; S3, a joint scheduling mechanism based on reinforcement learning, dynamically adjusts the allocation of computing resources and task scheduling by obtaining cluster load, computing task requirements, and data sensitivity information in real time; S4. Deploy a zero-trust architecture in the cluster so that each request must be authenticated and authorized, and fine-grained control of task permissions is performed through role-based access control policies; S5, monitor computing tasks and data transmission in the cluster in real time, use deep learning algorithms to detect abnormal cluster behavior, and identify potential security threats; S6. When a potential threat is detected, the security response mechanism is automatically executed to provide protection by isolating abnormal nodes and limiting resource access of malicious tasks.
2. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The sensitivity assessment of the data in the cluster in step S1 specifically includes the following steps: S1.
1. Collect data of computing tasks in the cluster and analyze data types, contents and access frequency; S1.
2. Use a trained machine learning model or specified statistical method to assess the sensitivity of the data, and classify the data into low-sensitivity data, medium-sensitivity data, and high-sensitivity data based on the assessment results; S1.
3. Classify the data according to the evaluation results and specify the corresponding encryption algorithm and security strategy according to the category; S1.
4. Update the sensitivity level of data regularly or in real time according to changes in cluster resources, and dynamically adjust the encryption strategy.
3. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The S2 step specifically includes the following steps: S2.
1. Encrypt low-sensitivity data using the AES-128 encryption algorithm to generate corresponding encrypted ciphertext; S2.2, encrypt the sensitive data using the AES-256 encryption algorithm to generate the corresponding encrypted ciphertext; S2.
3. Encrypt highly sensitive data using a homomorphic encryption algorithm, where the homomorphic encryption algorithm is a Paillier encryption algorithm; S2.
4. All encryption processes generate and manage keys through a centralized key management system, including key lifecycle management, key storage and distribution.
4. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The joint scheduling mechanism based on reinforcement learning in step S3 includes: S3.
1. Collect the resource usage of the current cluster, including the load, memory and CPU usage of the computing nodes, as well as the computing requirements and priorities of the tasks; S3.
2. Based on deep Q learning, dynamically adjust the task scheduling strategy and resource allocation strategy while satisfying the resource allocation constraints; S3.
3. According to the encryption level, data sensitivity and computing resource requirements of the task, select a computing node with idle resources that can meet the encryption computing requirements; S3.
4. When scheduling computing tasks, consider resource utilization, load balancing, and security.
5. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The implementation of the zero trust architecture in step S4 specifically includes: S4.
1. Implement OAuth2.0 or OpenIDConnect-based authentication for each request within the cluster. S4.
2. Implement attribute-based access control to control access rights based on the sensitivity of the task, the identity and role information of the requester; S4.
3. Each access request is subject to permission verification; S4.
4. Dynamically adjust permissions so that permissions are updated in real time based on task requirements and access conditions during the task life cycle.
6. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The S5 step specifically includes: S5.
1. Use the Prometheus monitoring system to monitor the computing node resource usage of the cluster in real time, and visualize it through Grafana; S5.
2. Use deep learning algorithms to model cluster behavior and identify abnormal patterns that deviate from normal behavior; S5.
3. Use model-based anomaly detection methods to automatically trigger alarms and notify administrators when cluster resources or task behaviors are abnormal; S5.
4. Combine the monitoring system with the anomaly detection results to generate real-time security reports to help administrators take timely response measures.
7. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The safety response mechanism in step S6 includes: S6.
1. When a security threat is detected, the malicious node is automatically isolated and the computing task on the node is stopped to prevent the security threat from spreading. S6.
2. Limit the resource access rights of the affected nodes or tasks to reduce their interference with normal computing tasks; S6.
3. Automatically adjust cluster access permissions based on the severity of abnormal behavior to prohibit unauthorized users or nodes from accessing sensitive data. S6.
4. Keep detailed records of all security incidents, generate audit logs, and provide them to administrators for subsequent analysis and investigation.
8. The server cluster security control method combined with multi-level encryption strategy joint scheduling according to claim 1 is characterized in that: The joint scheduling mechanism of reinforcement learning in step S3 optimizes the scheduling strategy through the following mathematical model: ;in, R(t) represents the reward value of the scheduling strategy, r i represents the available capacity of resource i, t j represents the demand of task j, x ij is a Boolean variable indicating whether task j is assigned to resource i, α and β are regularization coefficients, n is the number of resource nodes, m is the number of tasks, and d j is the resource requirement of task j, p j is the priority of task j.
Citation Information
Patent Citations
Online service computing power optimization method and system based on cloud computing
CN118484267A
Client information dynamic management and protection method
CN118643480A
Network security early warning isolation system based on cloud computing
CN119172118A
Data security acquisition scheduling system based on big data
CN119293810A
Station area intelligent fusion terminal data processing system based on edge calculation
CN119440800A
Cited By
Server cluster management method, electronic device, storage medium and program product
CN120185952A
Server cluster management method, electronic device, storage medium and program product
CN120185952B
Data encryption strategy making and data full-life-cycle safety guarantee system and method based on artificial intelligence
CN120200862A
Data encryption strategy formulation and data full life cycle security assurance system and method based on artificial intelligence
CN120200862B
Improved Paillier dynamic operator method and system
CN120281461A