A cloud platform-based resource management method and system

By allocating resources and loading initial security policies for big data clusters through the veStack virtualization layer, monitoring network traffic in real time, and dynamically updating boundary firewall policies, the problem of dynamic adaptability and zero-disruption migration in cloud platform resource management is solved. It also enables accurate identification and automatic protection of exposed nodes and abnormal traffic on the public network, reduces hardware costs, and improves management efficiency.

CN120956452BActive Publication Date: 2026-02-03BEIJING XINYUNXINWANG DATA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511025350.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2026-02-03
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

Existing cloud platform resource management methods lack dynamic adaptability in the face of network traffic changes and new security threats, making it difficult to accurately identify risks, resulting in a high risk of threat spread, unreasonable resource allocation, lack of zero-disruption migration guarantees, error-prone manual configuration, time-consuming fault location, and a lack of unified visual management and control interfaces and standardized delivery processes.

Method used

The veStack virtualization layer allocates computing and storage resources to the big data cluster, loads the initial security policy package, monitors network traffic in real time, dynamically updates the boundary firewall policy, performs dynamic resource scheduling based on the cluster load status, generates visual delivery documents, and achieves zero-disruption migration and unified management.

Benefits of technology

It achieves accurate identification of exposed nodes and abnormal traffic on the public network, automatically activates protection rules to block ransomware and DDoS attacks, dynamically allocates resources to reduce hardware costs, reduce human error, improve management efficiency, and provides a visual interface for convenient deployment and version tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956452B_ABST
    Figure CN120956452B_ABST
Patent Text Reader

Abstract

The application provides a cloud platform-based resource management method and system, and relates to the technical field of cloud computing management.The method comprises the following steps: allocating computing resources and storage resources for a big data cluster through a veStack virtualization layer; loading an initial security policy package for the cluster based on the computing resources and the storage resources, wherein the initial security policy package comprises protection analysis rules of a cloud firewall and access control policies of a boundary firewall; monitoring network traffic and access behavior of the cluster resources in real time based on the initial security policy package, identifying active external connection hosts and risk domain names and IP addresses through a protection analysis unit, detecting abnormal communication traffic between public network IP assets and the Internet, and generating a security event evaluation report.The application can accurately identify and respond to threats through initial security policies, real-time monitoring and dynamic policy updating, dynamically schedule resources, realize zero-interruption migration, and improve management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing management technology, and in particular to a resource management method and system based on a cloud platform. Background Technology

[0002] With the rapid development of cloud computing and big data technologies, enterprises' resource management needs for big data clusters on cloud platforms are becoming increasingly complex. However, current cloud-based resource management methods may face many technical challenges in practical applications: some security protection systems lack dynamic adaptability, statically deployed security strategies may be unable to respond in real time to changes in network traffic and new security threats, may not be able to accurately identify the risks of exposed nodes on the public network, and some event responses rely on manual intervention, which may lead to a high risk of threat spread; static resource allocation models may cause resource waste or overload, and some migration processes lack zero-disruption guarantee mechanisms; manual configuration may be prone to errors and fault location is time-consuming, there may be a lack of a unified visual management interface, and some delivery processes also lack standardization and auditable mechanisms. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a resource management method and system based on a cloud platform, which can accurately identify and respond to threats through initial security policies, real-time monitoring and dynamic policy updates; at the same time, dynamically schedule resources and achieve zero-interruption migration, visualized delivery, and improve management efficiency.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, a resource management method based on a cloud platform, the method comprising:

[0006] Step S1: Allocate computing and storage resources to the big data cluster through the veStack virtualization layer;

[0007] Step S2: Based on computing and storage resources, load an initial security policy package for the cluster. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall.

[0008] Step S3: Based on the initial security policy package, monitor the network traffic and access behavior of cluster resources in real time, identify active external hosts, risky domains and IPs through the protection analysis unit, detect abnormal communication traffic between public IP assets and the Internet, and generate a security incident assessment report;

[0009] Step S4: Based on the security incident assessment report, dynamically update the policy management component of the perimeter firewall and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies;

[0010] Step S5: Based on the updated security policy package, combined with the cluster load status and cost optimization objectives, dynamically scale up and down virtualization resources to generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; automatically isolating and migrating the load of cluster nodes with unprocessed security events.

[0011] Step S6: Based on the scheduled resource allocation results and the updated security policy package, visualize the cluster operation status through a unified management interface, and generate a deliverable software installation package and deployment documentation.

[0012] Further, step S1: Allocate computing and storage resources to the big data cluster through the veStack virtualization layer, including:

[0013] A big data cluster framework is built using the veStack virtualization layer. The big data cluster framework includes the logical topology of Hadoop cluster, Doris cluster, Flink cluster, StarRocks cluster and OpenSearch cluster.

[0014] Based on the logical topology of the OpenSearch cluster, initial computing and storage resources are allocated to each cluster, and a resource allocation plan is generated.

[0015] According to the resource allocation plan, cluster components are deployed in the veStack virtualization layer, and the running status of the cluster components is output, including: HDFS 3.3.4, YARN 3.3.4 and MapReduce 23.3.4 components are deployed for the Hadoop cluster, Doris 2.0.7 components are deployed for the Doris cluster, and Flink 1.16.1 and Paimon 0.6.1 components are deployed for the Flink cluster;

[0016] Based on the running status of cluster components, resource quotas are dynamically adjusted, and the allocated computing and storage resources are output.

[0017] Further, step S2: Based on computing and storage resources, load an initial security policy package for the cluster. This initial security policy package includes protection analysis rules for the cloud firewall and access control policies for the boundary firewall, including:

[0018] Based on computing and storage resources, the network exposure surface of cluster nodes is calculated to obtain node network exposure status data. For the nodes in the calculated network exposure status data, nodes bound to public IP addresses are marked as publicly exposed nodes. For the storage utilization data of all cluster nodes, nodes with storage utilization exceeding a preset threshold are marked as data storage nodes. Based on the publicly exposed nodes, data storage nodes, and the logical connection relationships between nodes, a security domain topology containing security labels and logical connection relationships is output.

[0019] Based on the security domain topology, for nodes exposed to the public network, their open ports are scanned. If risky ports are found, inbound traffic port filtering rules are generated, and unnecessary ports are closed according to the business whitelist. For data storage nodes, their outbound traffic application protocols are analyzed. If non-database protocols are found, outbound traffic application control rules are generated, high-risk protocol traffic is blocked, and an effective set of access control policies is output.

[0020] Based on the effective status of the access control policy set, the traffic statistics nodes deployed inside the border firewall identify active outbound behavior through the traffic matrix between the statistical nodes. When the frequency of node outbound connections exceeds the baseline value, the active outbound connection detection engine is triggered to capture risky domain names and IPs. At the same time, the risk awareness nodes deployed in the core switching layer extract communication characteristics from the traffic of the mirrored public network exposed nodes. If traffic matching the risk IP database is detected, the abnormal communication traffic analysis engine is triggered to generate an alarm and output a protection analysis rule deployment report.

[0021] By integrating access control policy sets and protection analysis rule deployment reports, the boundary policy rules are bound to the cloud firewall detection engine to generate a unified policy execution priority list and output the initial security policy package.

[0022] Further, step S3: Based on the initial security policy package, monitor network traffic and access behavior of cluster resources in real time, identify proactively connecting hosts and risky domains and IPs through the protection analysis unit, detect abnormal communication traffic between public IP assets and the Internet, and generate a security incident assessment report, including:

[0023] Based on the initial security policy package, the traffic monitoring probe is activated in the protection analysis unit. According to the engine configuration parameters of the initial security policy package, a lightweight log collection agent is deployed on the cluster nodes. According to the policy priority list, full traffic capture is used for data storage nodes and sampling capture is used for computing nodes. Real-time network traffic logs with node labels are output.

[0024] Based on real-time network traffic logs, the external analysis engine filters out request packets with missing outgoing SYN or ACK responses in the logs, parses the HTTPHost header and DNSQuery field of the request packets, and outputs a list of target domain names and IPs. At the same time, the public network traffic analysis engine mirrors the TCP or UDP session streams of public network IP nodes, calculates the average packet size, traffic burst variance, and port access frequency of the session, and outputs the session feature vector.

[0025] Input the session feature vector into the baseline behavior model, generate abnormal communication alarms for packets whose mean size deviates from the baseline by more than 50% or whose traffic burst variance is greater than the threshold, and output an abnormal communication traffic alarm set; at the same time, query the threat intelligence database for the target domain name and IP list, mark those in the intelligence database as confirmed risk items, and mark those not in the intelligence database but whose access frequency is more than 3 times the baseline value as suspected risk items, and output an active outbound risk list;

[0026] Based on the proactive external connection risk list and abnormal communication traffic alarm set, the risk domain name and IP are associated with the abnormal traffic session to calculate the security incident risk level and generate a security incident assessment report.

[0027] Further, step S4: Based on the security incident assessment report, dynamically update the policy management component of the perimeter firewall and output an updated security policy package, including: application control, URL filtering, and virus protection rules enabled for cluster nodes at risk, and adjusted inbound and outbound traffic access control policies, including:

[0028] Based on the security incident assessment report, the incident classification and risk node list are analyzed, and a risk node protection requirement set including high-risk nodes, medium-risk nodes, and low-risk nodes is output.

[0029] Based on the risk node protection requirements set, application control rules, URL filtering rules, and virus protection rules are created for high-risk nodes, application control rules and URL filtering rules are created for medium-risk nodes, and virus protection rules are created for low-risk nodes, generating a dynamic protection rule set.

[0030] Based on the risk node protection requirement set, for inbound traffic, if the risk node is exposed to the public network, a port access restriction rule is generated; if there is abnormal traffic with DDoS characteristics, a traffic rate restriction rule is generated. For outbound traffic, if the risk node is connected to a risky IP, a target IP blocking rule is generated; if the outbound traffic exceeds the business baseline, a bandwidth limiting rule is generated. The traffic control rule set is recalculated and generated.

[0031] Based on dynamic protection rule sets and traffic control rule sets, a policy execution dependency tree is constructed, a unified policy version identifier is generated, and an updated security policy package is output.

[0032] Further, step S5: Based on the updated security policy package, combined with cluster load status and cost optimization objectives, dynamically scale virtualization resources to generate scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet security risk level standards; automatically isolating and migrating the load of cluster nodes with unprocessed security events, including:

[0033] Based on the updated security policy package, the policy execution dependency tree is parsed to identify clusters that meet the security risk level, a list of risk nodes is extracted and nodes with unprocessed security events are marked, and a resource scheduling instruction set is output, including HDFS and YARN resource expansion instructions for compliant clusters and node isolation and load migration instructions for risk nodes.

[0034] Based on the resource scheduling instruction set, for expansion instructions, the current CPU, memory, and storage utilization of the cluster are collected, the difference between the target performance index and the current state is calculated, and the resource gap is output; for migration instructions, the container group dependencies of the nodes to be migrated are scanned, the container group communication matrix and storage volume mounting topology are drawn, and the load distribution topology is output.

[0035] Based on the resource gap and load distribution topology, resource expansion is sorted by cost priority and allocated from the lowest cost pool until the gap is met, generating a minimum resource allocation scheme; for load migration, the load tolerance of cluster nodes is calculated, the target node with the highest tolerance is selected, and a phased migration order of stateless services first and stateful services is designed to generate a zero-interruption migration path.

[0036] Based on the minimum resource allocation scheme and zero-interruption migration path, the system calls the API through the veStack virtualization layer to add DataNode nodes to the HDFS cluster and allocate YARN container quotas, triggers real-time migration of isolated nodes, and outputs the resource allocation results after scheduling.

[0037] Further, step S6: Based on the scheduled resource allocation results and the updated security policy package, the cluster operating status is visualized through a unified management interface, and a deliverable software installation package and deployment documentation are generated, including:

[0038] Based on the scheduled resource allocation results and the updated security policy package, extract the visualization metadata, parse the cluster resource topology from the resource allocation results, and parse the policy configuration snapshot from the security policy package;

[0039] Based on visual metadata, a dynamic resource heatmap component is constructed by using a unified management interface rendering engine to generate a third-order color gradient based on CPU or memory usage at the node level and overlaying a 2Hz flashing alarm icon on isolated nodes. At the same time, a rule tree topology of application control → URL filtering → virus protection is constructed with the policy version identifier as the root node, and a red pulse highlight border is added to high-risk rules to generate a policy dependency graph component. Finally, an embeddable visual interface component is output.

[0040] Based on cluster resource topology and policy configuration snapshots, a software installation package is built by encapsulating the veStack virtualization layer driver and cluster component deployment scripts, and injecting updated security policy packages. At the same time, the resource topology is converted into AnsiblePlaybook and the policy configuration snapshot is converted into a Markdown format configuration manual to generate deployment documents and deliverables.

[0041] Integrate deliverables and visual interface components, create a delivery zone on a unified management interface, provide a software installation package download channel and embed deployment documents that support online preview, and finally output an auditable delivery report that includes interface access links, installation package hash values, and document version numbers.

[0042] Secondly, a cloud-based resource management system includes:

[0043] The allocation module is used to allocate computing and storage resources to the big data cluster through the veStack virtualization layer;

[0044] The loading module is used to load an initial security policy package for the cluster based on computing and storage resources. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall.

[0045] The detection module is used to monitor network traffic and access behavior of cluster resources in real time based on the initial security policy package. It identifies active external hosts, risky domains and IPs through the protection analysis unit, detects abnormal communication traffic between public IP assets and the Internet, and generates a security incident assessment report.

[0046] The dynamic update module is used to dynamically update the policy management component of the perimeter firewall based on the security incident assessment report, and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies.

[0047] The scheduling module is used to dynamically scale virtualization resources based on updated security policy packages, combined with cluster load status and cost optimization objectives, and generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; and automatically isolating and migrating the load of cluster nodes with unprocessed security events.

[0048] The visualization module is used to visualize the cluster's operating status through a unified management interface based on the scheduled resource allocation results and updated security policy packages, and to generate deliverable software installation packages and deployment documents.

[0049] Thirdly, a computing device includes:

[0050] One or more processors;

[0051] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0052] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0053] The above-described solution of the present invention has at least the following beneficial effects:

[0054] A closed-loop risk management system is formed throughout the entire lifecycle by initial policy loading, real-time monitoring, and dynamic updates. It accurately identifies exposed public network nodes and abnormal traffic, automatically activates protection rules to block threats such as ransomware and DDoS attacks, and utilizes risk isolation and baseline analysis mechanisms to intercept known risks and provide early warnings of unknown threats. Based on security status and load requirements, it dynamically scales up and down and migrates loads, allocates resources based on cost priorities to reduce hardware costs, and employs a zero-disruption migration path to ensure business continuity. Full-process API automation reduces human error, real-time monitoring accelerates fault location, and a visual interface and standardized deliverables enable convenient deployment and version tracking. Baseline behavior models promptly block abnormal traffic, the risk IP database is updated in real-time to defend against the latest threats, and auditable reports and operation logs are available. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating a cloud platform-based resource management method provided by an embodiment of the present invention.

[0056] Figure 2 This is a schematic diagram of a cloud-based resource management system provided by an embodiment of the present invention. Detailed Implementation

[0057] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0058] like Figure 1 As shown, an embodiment of the present invention proposes a resource management method based on a cloud platform, the method comprising the following steps:

[0059] Step S1: Allocate computing and storage resources to the big data cluster through the veStack virtualization layer;

[0060] Step S2: Based on computing and storage resources, load an initial security policy package for the cluster. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall.

[0061] Step S3: Based on the initial security policy package, monitor the network traffic and access behavior of cluster resources in real time, identify active external hosts, risky domains and IPs through the protection analysis unit, detect abnormal communication traffic between public IP assets and the Internet, and generate a security incident assessment report;

[0062] Step S4: Based on the security incident assessment report, dynamically update the policy management component of the perimeter firewall and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies;

[0063] Step S5: Based on the updated security policy package, combined with the cluster load status and cost optimization objectives, dynamically scale up and down virtualization resources to generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; automatically isolating and migrating the load of cluster nodes with unprocessed security events.

[0064] Step S6: Based on the scheduled resource allocation results and the updated security policy package, visualize the cluster operation status through a unified management interface, and generate a deliverable software installation package and deployment documentation.

[0065] In this embodiment of the invention, a closed loop of "protection-monitoring-response" is formed through initial security policy loading, real-time risk monitoring, and dynamic policy updates. This loop can accurately identify security events such as proactive external connection risks and abnormal communication traffic, automatically activate application control and URL filtering rules to reduce risks such as virus intrusion and DDoS attacks, and automatically isolate and migrate the load of nodes with security events to prevent risk spread. The risk status is displayed in real time through a visual interface, improving the efficiency of security event handling. Based on the security status and load requirements, HDFS and YARN resources are added to qualified clusters as needed, and the load is migrated for risky nodes to maximize resource utilization. Furthermore, when expanding resources, local storage is allocated based on cost priority, which is greater than network storage, which is greater than cloud storage, generating a minimum resource allocation plan to reduce costs. A unified management interface is used to visually display resource heatmaps and policy dependencies, helping operations and maintenance personnel to intuitively grasp the cluster status, generate software installation packages, deployment documents, and auditable reports, and standardize the delivery process. Migration operations are designed with a path in the order of "stateless service → stateful service", and mechanisms such as IP migration and storage volume mounting are used to ensure zero service interruption. At the same time, by comparing traffic baselines and session characteristics, traffic that deviates from the normal business mode can be detected in time, providing early warnings and blocking risks to reduce the risk of business interruption.

[0066] In a preferred embodiment of the present invention, step S1 above: allocating computing and storage resources to the big data cluster through the veStack virtualization layer, may include:

[0067] Step S11: Construct a big data cluster framework through the veStack virtualization layer. The big data cluster framework includes the logical topology of Hadoop cluster, Doris cluster, Flink cluster, StarRocks cluster and OpenSearch cluster.

[0068] Step S12: Based on the logical topology of the OpenSearch cluster, allocate initial computing and storage resources to each cluster and generate a resource allocation plan.

[0069] Step S13: According to the resource allocation plan, deploy cluster components in the veStack virtualization layer and output the running status of the cluster components, including: deploying HDFS 3.3.4, YARN 3.3.4 and MapReduce 23.3.4 components for the Hadoop cluster, deploying Doris 2.0.7 components for the Doris cluster, and deploying Flink 1.16.1 and Paimon 0.6.1 components for the Flink cluster;

[0070] Step S14: Based on the running status of cluster components, dynamically adjust resource quotas and output the allocated computing and storage resources.

[0071] In this embodiment of the invention, a standardized cluster architecture is formed, which facilitates unified management of subsequent resource allocation and component deployment. At the same time, it clarifies component dependencies and reduces cluster stability risks caused by topology design defects. Resource allocation is quantified based on business needs to avoid excessive or insufficient initial configuration, and redundant space is reserved to ensure cluster stability when data volume increases. Standardized component configuration parameters ensure that cluster performance meets expectations, and the running status is output in real time to facilitate rapid location of deployment faults. Dynamic adaptation to business load fluctuations avoids resource waste, ensures cluster performance stability, and prevents task timeouts or component crashes due to insufficient resources.

[0072] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0073] In step S11 above, the roles and connections of nodes are designed according to the functional positioning of each cluster. For example, in a Hadoop cluster, NameNode (one primary and one backup) and DataNode (the number of nodes is estimated based on the amount of data, e.g., 50 nodes are needed for 10PB of data, and each node can store 200TB) are planned. In a Flink cluster, JobManager (two high-availability nodes) and TaskManager (set according to the number of concurrent tasks, e.g., 20 nodes are needed for 100 parallel tasks, with 5 task slots per node) are planned.

[0074] Determine the version compatibility of each cluster component. For example, HDFS 3.3.4 needs to be compatible with YARN 3.3.4, and Flink 1.16.1 needs to be compatible with the storage interface of Paimon 0.6.1 to avoid component crashes caused by version conflicts.

[0075] In step S12 above, the Hadoop cluster is calculated based on data storage volume. For example, 10PB of data (3 replicas) requires 30PB of storage. With 20% redundancy, the total storage requirement is 36PB, which is allocated to 50 nodes, with each node storing 720TB. YARN memory is configured at 128GB per node, for a total memory of 6.4TB (50 nodes × 128GB).

[0076] Flink cluster: Calculated by task concurrency, 100 parallel tasks require 16GB of memory per TaskManager node (2GB per task slot), with a total of 20 nodes (100 tasks / 5 slots / node), for a total memory of 320GB.

[0077] Redundancy settings: Set critical components (such as NameNode) to 1:1 hot standby, doubling resource usage; reserve 20%~30% of storage space to cope with data growth.

[0078] In step S13 above, when configuring component parameters, HDFS sets the block size to 256MB and the number of replicas to 3. Calculating the number of storage blocks per node using 720TB / 256MB, it ensures that the node storage utilization does not exceed 80%, i.e., 720TB × 80% = 576TB is used for storage blocks. YARN sets the ResourceManager memory to 16GB and the NodeManager memory to 120GB (with 8GB reserved for system overhead). Each container has a maximum memory of 10GB, thus calculating the number of allocable containers as 120GB / 10GB = 12 per node. Simultaneously, status indicators are calculated, monitoring component startup time (e.g., HDFSNameNode startup should be ≤5 minutes), resource usage baseline (e.g., DorisBE node CPU utilization should initially be ≤20%), and generating a component health score that includes startup success rate, memory leak detection, etc.

[0079] In step S14 above: When calculating resource gaps, if the HDFS storage utilization exceeds 80%, such as the current usage of 30PB / 36PB reaching 83%, then the required storage expansion is 36PB × (83% - 80%) = 1.08PB. Based on a single node capacity of 200TB, approximately 6 new nodes are needed (1.08PB / 200TB ≈ 5.4). If the YARN CPU utilization consistently exceeds 70%, then the required number of vCPU cores is the current total number of cores × (actual utilization - ...). 70%) / 70%, for example, with 50 nodes and 16 cores per node, totaling 800 cores, when the utilization rate is 85%, 800×(85%-70%) / 70%≈171 cores need to be added, which can be allocated to 11 nodes (11×16 cores=176 cores); the quota redistribution strategy is to reclaim 20% of the resources from low-load components (such as StarRocks cluster memory utilization <30%) and allocate them to high-load components (such as Flink cluster) to ensure that the overall resource utilization rate is not less than 60%.

[0080] In a preferred embodiment of the present invention, step S2 above: based on computing resources and storage resources, loading an initial security policy package for the cluster, the initial security policy package including cloud firewall protection analysis rules and perimeter firewall access control policies, may include:

[0081] Step S21: Based on computing and storage resources, calculate the network exposure surface of cluster nodes to obtain node network exposure status data; for the nodes in the calculated network exposure status data, mark the nodes bound to public IP addresses as public network exposed nodes; for the storage utilization data of all cluster nodes, mark the nodes whose storage utilization exceeds a preset threshold as data storage nodes; based on the public network exposed nodes, data storage nodes, and the logical connection relationships between nodes, output the security domain topology containing security labels and logical connection relationships.

[0082] Step S22: Based on the security domain topology, for nodes exposed to the public network, scan their open ports. If there are risky ports, generate inbound traffic port filtering rules and close unnecessary ports according to the business whitelist. For data storage nodes, analyze their outbound traffic application protocols. If there are non-database protocols, generate outbound traffic application control rules, block high-risk protocol traffic, and output the effective access control policy set.

[0083] Step S23: Based on the effective status of the access control policy set, the traffic statistics node deployed inside the border firewall identifies active outbound behavior in the traffic matrix between the statistics nodes. When the outbound frequency of a node exceeds the baseline value, the active outbound detection engine is triggered to capture risky domain names and IPs. At the same time, the risk awareness node deployed in the core switching layer extracts communication characteristics from the traffic of the mirrored public network exposed nodes. If traffic matching the risk IP database is detected, the abnormal communication traffic analysis engine is triggered to generate an alarm and output a protection analysis rule deployment report.

[0084] Step S24: By integrating the access control policy set and protection analysis rule deployment report, the boundary policy rules are bound to the cloud firewall detection engine to generate a unified policy execution priority list and output the initial security policy package.

[0085] In this embodiment of the invention, by clearly defining the attack surface boundary of the cluster, focusing on key protection of publicly exposed nodes, and prioritizing the coverage of security policies for data-intensive nodes based on storage risk-level management, the attack surface of publicly exposed nodes is effectively reduced, the risk of brute-force attacks is lowered, and data storage nodes are prevented from leaking data through unauthorized protocols. Unknown threats are identified based on behavioral baselines, compensating for the lag in the rule base and blocking communication with risky IPs in real time. Rule conflicts are avoided through unified policy management, improving security response efficiency. By leveraging the joint protection of the boundary and cloud firewalls, a multi-level security barrier is formed, achieving a collaborative defense effect of initial screening through boundary blocking and in-depth detection by the cloud firewall.

[0086] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0087] In step S21 above, the IP configuration information of all nodes in the cluster is traversed to determine whether it belongs to a public address range (such as the classic 1.0.0.0 / 8, 2.0.0.0 / 7, etc. IPv4 public network ranges) or whether it is bound to an Elastic IP (EIP). If one of the above conditions is met, the node will be directly marked as a "public network exposed node" and classified as a high-risk object. Taking a certain cluster as an example, it contains 100 nodes, of which 5 nodes will be quickly identified and marked because they are bound to public IPs, becoming the key monitoring targets for security protection.

[0088] A storage utilization threshold (e.g., 80%) is preset, and the storage utilization of each node is collected in real time. When the storage utilization of a node exceeds the threshold, it will be marked as a "data storage node". For example, if a node has a total storage capacity of 1TB and has used 850GB, its utilization rate is calculated to be 85%, which exceeds the 80% threshold. Therefore, it will be marked as a data-intensive node, and data security strategies will be deployed accordingly.

[0089] Based on the IP range of the nodes and the VPC subnet they belong to, the system will logically divide the cluster into regions and then draw the communication links between the nodes. For example, it will clarify the traffic flow between publicly exposed nodes and the boundary firewall and internal data nodes (publicly exposed node → boundary firewall → internal data node) and label each node with a security tag (such as "publicly exposed", "data storage" and "ordinary node"). In this way, a visual security domain topology containing node security tags and logical connection relationships will be formed.

[0090] In step S22 above, a port scanning tool (such as Nmap) is called to detect open ports and compare them with a list of risky ports (such as 21 / 22 / 3389). For example, if a node is found to be opening port 22 (SSH), but the business whitelist only allows ports 80 / 443, a filtering rule to close port 22 is generated.

[0091] Based on the cluster service list (e.g., Hadoop requires ports 9870 / 9000 to be open), retain the necessary ports and close non-whitelisted ports. For example, the Doris cluster only needs ports 9030 / 8030, and all other ports should be blocked.

[0092] Parse the traffic 5-tuple (source IP / port, destination IP / port, protocol) to identify the application layer protocol. For example, if the outbound traffic from a data node contains TCP port 6379 (Redis), it belongs to the database protocol and is allowed to pass. If UDP port 137 (NetBIOS) is found, it belongs to the non-database protocol and a blocking rule is generated. A preset blocking list (such as RDP, Telnet unencrypted protocol) is also provided. For example, if a node is detected to be connecting to the outside via Telnet (port 23), the traffic is directly blocked.

[0093] In step S23 above, collect the traffic logs between nodes within 24 hours, count the outbound frequency of each node (e.g., a node makes an average of 5 outbound connections per hour), and set this value as the baseline value; multiply the baseline value by 1.5 (e.g., 5 times / hour → threshold 7.5 times / hour), and trigger an alarm if the threshold is exceeded.

[0094] Real-time statistics on the frequency of outbound connections of nodes (e.g., a node makes 10 outbound connections in 1 hour, exceeding the threshold), triggering the engine to capture the domain name (e.g., malicious.com) and IP (1.1.1.1) of the outbound target.

[0095] Access authoritative threat intelligence databases such as VirusTotal and AlienVaultOTX, import known malicious IPs (such as ransomware C2 servers and nodes) in batches, and store them according to risk level (e.g., initially import 100,000 high-risk IPs); collect attack IPs from historical security incidents (such as the source IPs of intrusions), manually add them to the local risk database, and form an enterprise-specific threat list.

[0096] Every day at midnight, we connect to external intelligence sources via API to obtain newly added malicious IPs (e.g., 100 new node IPs per day), automatically add them to the risk database and mark the risk type; we capture IPs with abnormal behavior (such as high-frequency port scanning) in real time through firewall / IDS, and add them to the risk database after verification by the security team, thus realizing a closed loop of "discovery-verification-database entry".

[0097] Perform mirror analysis on the traffic bound to public IP nodes to extract multi-dimensional features:

[0098] Basic characteristics: source / destination IP, port, protocol type, session duration (e.g., normal business session < 5 minutes, a session lasting 30 minutes);

[0099] Traffic characteristics: Packet size distribution (e.g., the average size of a normal business packet is 500B, while the average size of a certain traffic packet is 10KB), and traffic burst variance (standard deviation of a certain time period > 500KB / s).

[0100] If the destination IP in the traffic exists in the risk database (e.g., 2.2.2.2 is marked as a C2 server), an "Confirm Risk" alarm is immediately triggered, and the risk type is associated (e.g., "Remote Control"). For IPs not in the database, if the traffic characteristics match the malicious pattern (e.g., large traffic outflow, abnormally long session), a "Suspected Risk" alarm is generated, submitted for manual review, and the risk database is updated.

[0101] The border firewall immediately blocks inbound / outbound traffic from the risky IP (such as 2.2.2.2); it traces the IP's historical communication records to locate the affected node (such as a data node that once transmitted 10GB of data to 2.2.2.2); if the same IP is accessed by multiple nodes simultaneously (such as 3 public network nodes), it automatically upgrades the risk level and triggers an emergency alarm.

[0102] In step S24 above, priority rules are set so that high-risk port blocking (such as 3389) is greater than protocol control (such as non-database protocols) and traffic rate limiting. For example, when both port risk and protocol risk exist, port filtering rules are executed first. If there are contradictions in the scanning policy set (such as a port being allowed and blocked at the same time), the rule with higher priority is selected. For example, port 22 is allowed by the business whitelist, but it is a high-risk port, so the blocking rule (with higher priority) is executed in the end.

[0103] Synchronize border firewall rules (such as port filtering) to the cloud firewall detection engine and configure linkage policies. For example, after the border firewall blocks an IP, the cloud firewall automatically initiates traffic cleaning for that IP. Sort by "risk level + scope of impact". For example, block high-risk ports of publicly exposed nodes (high risk level) first, followed by protocol control of data nodes.

[0104] In a preferred embodiment of the present invention, step S3 above: based on the initial security policy package, real-time monitoring of network traffic and access behavior of cluster resources, identification of proactively connected hosts and risky domain names and IPs through the protection analysis unit, detection of abnormal communication traffic between public IP assets and the Internet, and generation of a security incident assessment report may include:

[0105] Step S31: Based on the initial security policy package, activate the traffic monitoring probe in the protection analysis unit, deploy a lightweight log collection agent on the cluster nodes according to the engine configuration parameters of the initial security policy package, capture full traffic for data storage nodes and sampled traffic for computing nodes according to the policy priority list, and output real-time network traffic logs with node labels.

[0106] Step S32: Based on real-time network traffic logs, the external analysis engine filters out request packets with missing outgoing SYN or ACK responses in the logs, parses the HTTPHost header and DNSQuery field of the request packets, and outputs a list of target domain names and IPs; at the same time, the public network traffic analysis engine mirrors the TCP or UDP session streams bound to public network IP nodes, calculates the average packet size, traffic burst variance, and port access frequency of the session, and outputs the session feature vector.

[0107] Step S33: Input the session feature vector into the baseline behavior model, generate abnormal communication alarms for packets whose average size deviates from the baseline by more than 50% or whose traffic burst variance is greater than the threshold, and output the abnormal communication traffic alarm set; at the same time, query the threat intelligence database for the target domain name and IP list, mark those existing in the intelligence database as confirmed risk items, and mark those not in the intelligence database but whose access frequency is more than 3 times the baseline value as suspected risk items, and output the proactive outbound risk list;

[0108] Step S34: Based on the proactive external connection risk list and abnormal communication traffic alarm set, associate risky domain names and IPs with abnormal traffic sessions, calculate the security event risk level, and generate a security event assessment report.

[0109] In this embodiment of the invention, a differentiated collection strategy is used to balance monitoring accuracy and system overhead. Specifically, full traffic capture ensures the security of critical nodes, while sampling reduces the load on computing nodes. Tagging logs facilitates subsequent correlation analysis and can quickly locate abnormal traffic of publicly exposed nodes. A parallel analysis mechanism improves processing efficiency, with dual engines processing data from different dimensions simultaneously. This accurately extracts domain names, IP lists, and session vectors of abnormal traffic characteristics related to outbound connections, laying a data foundation for risk identification. Combining a known threat database with behavioral analysis achieves a dual protection effect of "rapid interception of known risks + early warning of unknown threats." A global risk view is provided to help security teams quickly locate core threats, and the quantified risk levels also support priority decision-making, allowing for the priority handling of confirmed high-risk related events.

[0110] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0111] In step S31 above, based on the priority list of the initial security policy package, a full traffic capture probe is deployed on the data storage node (such as HDFSDataNode) to ensure 100% traffic log collection; a sampling probe is deployed on the computing node (such as FlinkTaskManager) to randomly sample at a ratio of 10% to 30% (such as collecting 300 out of 1000 packets per second).

[0112] Set the collection frequency to 100ms / time, and the maximum size of each log file to 100MB. If the size exceeds this limit, a new file will be generated continuously. Automatically add node security tags (such as "public network exposed node" or "data storage node") to the collected traffic logs. For example, if a log record originates from node ID=101, automatically associate it with the "data storage node" tag.

[0113] In step S32 above, all outbound traffic is scanned in the logs, and half-open connection packets with missing SYN or ACK responses (such as sending only a SYN packet but not receiving an ACK confirmation) are discarded to reduce false alarms.

[0114] Extract the HTTPHost header (e.g., "example.com") and DNSQuery field (e.g., "www.target.com") from the remaining valid requests, deduplicate them, and generate a list of target domain names; at the same time, extract the destination IP address from the request packet and generate an IP list.

[0115] For each TCP / UDP session bound to a public IP node, calculate the session duration, average packet size (e.g., an average packet size of 1.2KB for a certain session), traffic burst variance (e.g., a standard deviation of traffic fluctuation of 500KB / s for a certain period), and port access frequency (e.g., TCP port 443 is accessed 200 times per minute). Combine the above indicators into a multi-dimensional vector (e.g., [duration = 300s, average packet size = 1.2KB, variance = 500KB / s, port frequency = 200 times / min]).

[0116] Step S33 above involves collecting data from multiple broad and authoritative data sources, including security vendors, internet intelligence sharing organizations, and publicly available security vulnerability reports. For example, many well-known security vendors collect information on malicious activities worldwide through their extensive monitoring networks; internet intelligence sharing organizations gather threat intelligence shared by numerous security experts, researchers, and enterprises; and publicly available security vulnerability reporting platforms, such as CVEs (Common Vulnerability Disclosures), record detailed information related to security vulnerabilities found in various software and systems.

[0117] The massive amounts of collected data are integrated, and in the process, duplicate, redundant and invalid information is removed to ensure the accuracy and consistency of the data. Since the data formats and quality of different data sources vary, the same threat information may be recorded repeatedly in multiple data sources, or the data records may contain errors or be incomplete. Through data integration and cleaning, this messy data can be transformed into usable high-quality intelligence.

[0118] Data is categorized and labeled based on factors such as attack behavior, type, and target. For example, attacks can be categorized by type, such as phishing, malware, and denial-of-service attacks; and by target, they can be individuals, specific organizations, or industries. This categorization and labeling facilitates subsequent querying and analysis of threat intelligence.

[0119] Threats are assessed and categorized based on the severity and impact of attacks, typically into low, medium, and high levels. The assessment process considers multiple factors, such as the potential damage, the scope of impact, and the ease with which the attack can succeed. Threat intelligence databases need to be timely, requiring regular updates to data sources, continuous collection of new threat intelligence, and updates and maintenance of existing data. As cyberattack methods evolve and new threats continue to emerge, only regular updates can ensure that the threat intelligence database reflects the current true cyber threat landscape.

[0120] The list of target domains and IPs is compared with external threat intelligence databases (such as AlienVaultOTX). AlienVaultOTX is a globally authoritative open threat information sharing and analysis network with more than 50,000 participants from 140 countries, contributing more than 4 million threat indicators daily. When a domain (such as "malware.example.com") or IP (such as 2.2.2.2) exists in this threat intelligence database, it can be directly marked as a confirmed risk item. This is because these domains or IPs have been discovered and recorded in the threat intelligence database by other security agencies or researchers, proving that they have carried out or are highly likely to carry out malicious activities.

[0121] For domains or IPs not in the threat intelligence database, the method of calculating the ratio of their access frequency to the historical baseline value is used to determine their risk. For example, if the historical access frequency baseline value of an IP is 10 times / hour, when the calculated ratio of the current access frequency of the IP to the historical baseline value exceeds 3 times (such as the current access frequency being 31 times / hour), it is marked as a suspected risk item. This method is based on statistical principles. When the access behavior of a domain or IP deviates significantly from its normal historical access pattern, there is a high probability of abnormal activity, which requires further attention and investigation.

[0122] In session feature vectors, the average packet size is an important analytical metric. It is compared with the historical baseline. For example, the average packet size under normal business conditions is 500B. When the average packet size in the current session deviates from the historical baseline by more than 50% (such as reaching 1.2KB), the system will trigger an alarm. This is because abnormal changes in packet size may indicate abnormal data transmission in network traffic. For example, there may be malware transmitting a large amount of data, or the network protocol may have been tampered with, causing the packet size to not conform to the normal specification.

[0123] The system determines whether network communication is abnormal by comparing the traffic burst variance with a preset threshold. Assuming the preset threshold is 300KB / s, when the traffic burst variance of a session reaches 500KB / s, the system will generate an abnormal communication alarm. The traffic burst variance reflects the degree of drastic change in network traffic in a short period of time. An excessively high variance indicates that there are sudden and abnormal fluctuations in network traffic, which is often related to network attacks (such as the initial stage of a distributed denial-of-service attack) or other abnormal network activities.

[0124] In step S34 above, the domains or IPs in the proactive outbound risk list are cross-matched with the session characteristics in the abnormal communication traffic alarm set. Taking the confirmed risk IP (2.2.2.2) as an example, when the IP not only exists in the threat intelligence database and is marked as a confirmed risk item, but its related sessions also trigger abnormal packet size alarms (such as the average packet size deviating from the baseline by more than 50%), these two independent risk information are associated and marked as a high-risk event. This process is implemented through automated algorithms to quickly identify multiple abnormal behaviors caused by the same risk source, avoid viewing a single alarm in isolation, and improve the accuracy of risk judgment.

[0125] Based on preset weighting rules, the associated risk events are quantitatively scored to determine the risk level. For example, the weight of confirmed risk items is set to 0.7, and the weight of abnormal traffic is set to 0.3. For the above associated events, assuming that the confirmed risk item is judged as 100% risk and the severity of abnormal traffic alarm is assessed as 80%, the comprehensive score is calculated by the formula: 0.7×100%+0.3×80%=94%. According to the preset score range, the risk level is divided into three levels: high, medium, and low (e.g., above 90 points is high risk, 70-89 points is medium risk, and below 70 points is low risk).

[0126] A comprehensive statistical analysis of various risk events is conducted. The statistics include the number of different types of risk events, such as 5 confirmed risk events and 12 suspected risk events; the distribution of affected nodes, such as 80% of the affected nodes being exposed on the public network, thus identifying key areas for security protection; and the time trend of risk events, revealing a surge in risk events between 2 and 4 a.m., providing a time-based reference for the security team to develop targeted monitoring and protection strategies. These statistics present the cybersecurity situation from multiple perspectives, helping security personnel to grasp the overall risk status.

[0127] Using cluster nodes as the horizontal axis and time as the vertical axis, the risk density is displayed by different shades of color; the darker the color, the higher the risk level of the node at the corresponding time. For example, red represents high risk, yellow represents medium risk, and green represents low risk. Security personnel can use the heatmap to quickly locate high-risk nodes and time periods and intuitively understand the spatiotemporal distribution pattern of risks.

[0128] Using risky IPs, domains, and affected services as nodes and communication connections as edges, a visual graph is constructed. For example, the graph shows that a risky IP (2.2.2.2) is connected to multiple publicly exposed nodes and is associated with database services, web services, etc., clearly presenting the risk propagation path and scope of impact, which facilitates the security team to conduct root cause analysis and formulate blocking strategies.

[0129] In a preferred embodiment of the present invention, step S4 above: dynamically updating the policy management component of the perimeter firewall based on the security incident assessment report, and outputting an updated security policy package, including: application control, URL filtering, and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies, which may include:

[0130] Step S41: Based on the security incident assessment report, analyze the incident classification and risk node list, and output a risk node protection requirement set including high-risk nodes, medium-risk nodes, and low-risk nodes;

[0131] Step S42: Based on the risk node protection requirement set, create application control rules, URL filtering rules, and virus protection rules for high-risk nodes, create application control rules and URL filtering rules for medium-risk nodes, and create virus protection rules for low-risk nodes, generating a dynamic protection rule set.

[0132] Step S43: Based on the risk node protection requirement set, for inbound traffic, if the risk node is exposed to the public network, generate port access restriction rules; if there is abnormal traffic with DDoS characteristics, generate traffic rate restriction rules. For outbound traffic, if the risk node is connected to a risky IP, generate target IP blocking rules; if the outbound traffic exceeds the business baseline, generate bandwidth limiting rules. Recalculate and generate the traffic control rule set.

[0133] Step S44: Based on the dynamic protection rule set and the traffic control rule set, construct the policy execution dependency tree, generate a unified policy version identifier, and output the updated security policy package.

[0134] In this embodiment of the invention, by clearly defining risk levels, differentiated protection is implemented for nodes with different levels of danger, avoiding a "one-size-fits-all" approach and ensuring that security resources are precisely allocated to high-risk nodes, thereby improving protection efficiency. By creating protection rules in a hierarchical manner, while resisting high-risk threats, excessive protection is avoided to prevent interference with business operations, achieving a balance between security and efficiency. Relying on precise traffic control rules, malicious traffic is effectively intercepted, the spread of attacks is curbed, normal business traffic is ensured to flow smoothly, and network stability is enhanced. The use of unified policy management and version control eliminates the risk of rule conflicts, and a complete security policy package enables rapid deployment of protection rules, improving security response speed.

[0135] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0136] Step S41 above involves in-depth analysis of the security incident assessment report, extracting incident classification information from the report, such as confirmed risk incidents (malicious behaviors verified by the threat intelligence database), suspected risk incidents (behaviors that may pose a threat based on behavioral analysis), and identifying the IP addresses of the risk nodes involved in these incidents.

[0137] Based on pre-defined risk level classification criteria, the extracted risk nodes are categorized. The classification criteria comprehensively consider various factors, such as the frequency of communication between the node and known malicious IPs, the number and severity of triggered abnormal alarms, etc. For example, if a node maintains continuous communication with a known malicious IP and simultaneously triggers multiple alarms such as abnormal packet size and traffic bursts, this type of node will be judged as a high-risk node. Nodes that experience a single abnormal traffic fluctuation with a small fluctuation level are classified as low-risk nodes. Nodes in between are classified as medium-risk nodes.

[0138] By integrating nodes of different risk levels and their detailed information, such as the node's function in the cluster (whether it is a data storage node or a computing node) and the type of risk it currently faces (such as ransomware threats or DDoS attack threats), a set of protection requirements is formed. This set of requirements clearly presents the status of each risk node, providing a clear basis for the formulation of subsequent protection rules.

[0139] In step S42 above, after obtaining the set of protection requirements for risk nodes, the system formulates corresponding protection rules for nodes with different risk levels:

[0140] High-risk nodes: Due to the severe threats faced by high-risk nodes, comprehensive protection rules are created for them. Application control rules will strictly control the applications running on the nodes, allowing necessary applications related to business to access the network, blocking unnecessary applications from communicating with the outside world, and preventing malicious programs from transmitting data or launching attacks through the application layer; URL filtering rules will filter the URLs accessed by the nodes in real time based on a list of known malicious websites and phishing pages, and immediately block access requests once a malicious URL is detected; virus protection rules will perform deep scanning of data entering and leaving the nodes, using virus signature database comparison, behavioral analysis and other technologies to identify and remove potential virus threats, ensuring data security.

[0141] Medium-risk nodes: Considering that medium-risk nodes pose certain potential risks, but business operations have high requirements for their network access, application control rules and URL filtering rules are created for them. Application control rules appropriately restrict the network access of applications to reduce potential risks while ensuring the normal operation of business. URL filtering rules can also block malicious websites and reduce the possibility of nodes being attacked by networks.

[0142] Low-risk nodes: Low-risk nodes face relatively low risks. Create virus protection rules for them and prevent virus intrusion by scanning the data for viruses. This ensures basic security while minimizing the impact on business operations.

[0143] When formulating each rule, the system associates it with specific risk nodes and clarifies the effective conditions (such as specific time periods, specific traffic types, etc.) and execution actions (such as blocking connections, issuing alarms, etc.) to ensure that the protection rules can play an accurate and effective role.

[0144] In step S43 above, for each node with concentrated risk node protection needs, the inbound and outbound traffic are analyzed separately:

[0145] Inbound traffic: If a node is exposed to the public network, close unnecessary ports and open necessary ports and restrict access sources according to its business needs; if abnormal traffic with DDoS attack characteristics is detected (such as a large number of requests in a short period of time), set a traffic rate limit to control the traffic limit per unit time.

[0146] Outbound traffic: If a node actively connects to a risky IP, communication with that target IP will be directly blocked; if outbound traffic exceeds the historical service baseline, bandwidth limiting will be set according to the degree of excess to ensure that the traffic is within a reasonable range.

[0147] Step S44 above integrates the dynamic protection rule set (such as application control, URL filtering, and virus protection rules) with the traffic control rule set (such as port restriction, rate restriction, and IP blocking rules). During this process, the system analyzes the logical dependencies between rules. For example, URL filtering rules must be executed first to block access requests from known malicious domains and prevent subsequent rules from processing invalid traffic; application control rules must be executed after URL filtering to ensure that applications that pass security checks are allowed to communicate over the network; traffic control rules (such as port restriction and rate restriction), as the underlying protection, must take effect before application layer rules to quickly block abnormal traffic. Through this hierarchical dependency analysis, the system can determine the optimal execution order of rules and avoid rule conflicts or redundant execution.

[0148] Based on the logical relationships between rules, the system constructs a policy execution dependency tree. This tree starts from the root node and expands downwards to its child nodes according to the execution order:

[0149] Root node: Typically a traffic access point (such as a border firewall interface), responsible for receiving all inbound and outbound traffic;

[0150] Intermediate nodes: Various rule groups arranged by priority, such as: Inbound traffic: Port access restriction rules → DDoS rate restriction rules → URL filtering rules → Application control rules → Virus protection rules; Outbound traffic: Bandwidth rate limiting rules → Target IP blocking rules → URL filtering rules → Application control rules;

[0151] Leaf nodes: Each specific rule instance is associated with a specific risk node and execution action (such as "blocking access requests from IP 2.2.2.2").

[0152] The construction of the dependency tree ensures that rules are executed in sequence. For example, when traffic from the public network is detected, the system will first check the port access restriction rules, then perform URL filtering, and finally apply virus protection rules to form a multi-layered security protection system.

[0153] To facilitate policy management and traceability, the system assigns a unified version identifier (e.g., `v20250624-01`) to the entire rule set, which includes the following information:

[0154] Timestamp: The specific time the rule was generated (e.g., `2025-06-24T14:30:00`).

[0155] Version number: indicates the number of iterations of the strategy (e.g., `01` represents the first version of the day);

[0156] Source identifier: The system module or operator ID that generated the policy (e.g., `security-engine-01`).

[0157] At the same time, the system records metadata such as the rule's creation time, last modification time, summary of the modified content, and the operator, forming a complete change history; this information is crucial in subsequent policy audits, troubleshooting, and compliance checks.

[0158] All rules (including dependency tree structure, version identifiers, and metadata) are packaged to generate a complete security policy package. This policy package typically uses a standard format (such as JSON, XML, or a proprietary binary format) and includes:

[0159] Policy configuration file: Stores detailed configurations for all rules, including rule ID, type, activation conditions, and execution actions;

[0160] Dependency description: Defines the execution order and dependencies between rules;

[0161] Version control information includes version identifier, timestamp, change log, etc.

[0162] Verification information, such as digital signatures or hash values, ensures that the policy packet has not been tampered with during transmission.

[0163] The generated security policy package will be pushed to the policy management component of the perimeter firewall. The firewall will load and execute the new policy to achieve real-time protection of network traffic. The deployment process supports zero-disruption updates, meaning that the new policy will not affect ongoing legitimate connections when it takes effect, ensuring business continuity.

[0164] In a preferred embodiment of the present invention, step S5 above: based on the updated security policy package, combined with the cluster load status and cost optimization objectives, dynamically scaling up and down virtualization resources to generate the scheduled resource allocation result, including: for clusters that meet the security risk level, increasing HDFS and YARN resource quotas as needed; for cluster nodes with unprocessed security events, automatically isolating and migrating the load, which may include:

[0165] Step S51: Based on the updated security policy package, parse the policy execution dependency tree to identify clusters that meet the security risk level, extract the list of risk nodes and mark nodes with unprocessed security events, and output the resource scheduling instruction set, including generating HDFS and YARN resource expansion instructions for compliant clusters and node isolation and load migration instructions for risk nodes.

[0166] Step S52: Based on the resource scheduling instruction set, for expansion instructions, collect the current CPU, memory, and storage utilization of the cluster, calculate the difference between the target performance index and the current state, and output the resource gap; for migration instructions, scan the container group dependencies of the nodes to be migrated, draw the container group communication matrix and storage volume mounting topology, and output the load distribution topology.

[0167] Step S53: Based on the resource gap and load distribution topology, sort the resource pools according to cost priority for resource expansion and allocate resources from the lowest cost pool until the gap is met, generating a minimum resource allocation scheme; calculate the load tolerance of cluster nodes for load migration, select the target node with the highest tolerance, design a phased migration order of stateless services first and then stateful services, and generate a zero-interruption migration path.

[0168] Step S54: Based on the minimum resource allocation scheme and zero-interruption migration path, add DataNode nodes to the HDFS cluster and allocate YARN container quotas through the veStack virtualization layer API call, trigger real-time migration for isolated nodes, and output the resource allocation results after scheduling.

[0169] In this embodiment of the invention, the security-compliant clusters and risky nodes are identified by parsing the policy dependency tree. Priority is given to expanding HDFS and YARN resources for low-risk clusters, avoiding the allocation of resources to nodes with security vulnerabilities. Simultaneously, risky nodes are accurately located and isolation and migration commands are output to prevent the spread of security incidents. Cluster resource utilization is collected and the shortfall is calculated to prevent resource waste or insufficiency caused by blind expansion. A load topology map is drawn by scanning container dependencies, providing data support for migration path design and reducing the risk of service interruption. Resources are allocated according to cost priority from local SSD to HDD to cloud storage to reduce procurement costs. Combined with node load tolerance calculation and a phased migration strategy of "stateless service → stateful service," zero-interruption migration of services is achieved. Resource expansion and node migration are automatically executed through the veStack virtualization layer API, improving scheduling efficiency, monitoring the migration process in real time, and supporting abnormal rollback.

[0170] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0171] In step S51 above, the policy execution dependency tree in the updated security policy package is traversed to extract the risk level label (such as high / medium / low risk) of each cluster node. If the risk level of all nodes in the cluster is ≤ medium risk and there are no unprocessed security events, it is determined to be a "safety compliant cluster".

[0172] Filter out nodes with unprocessed security events from the dependency tree (such as isolated high-risk nodes, nodes that triggered abnormal alarms but were not blocked), generate a list of risk nodes, and label their security event types (such as actively connecting to risky IPs, abnormal traffic alarms).

[0173] For clusters that meet security standards, generate instructions to expand HDFS data block storage quotas (e.g., add 50TB of storage) and YARN container resource expansion instructions (e.g., add 200 vCPU cores); for risky nodes, generate isolation instructions (e.g., disconnect public IP connections) and load migration instructions (e.g., migrate Flink tasks on the node).

[0174] In step S52 above, the CPU utilization rate (e.g., current average 80%), memory utilization rate (e.g., 20GB remaining), and storage utilization rate (e.g., 800TB of HDFS used / 1PB total) of the qualified cluster are collected in real time; compared with the preset target performance indicators (e.g., CPU utilization rate ≤70%, storage remaining ≥300TB), the resource gap is calculated (e.g., 100TB of storage and 20GB of memory need to be added).

[0175] Scan the container group dependencies of the nodes to be migrated (e.g., container A depends on the database service of container B), draw the inter-container communication matrix (record the traffic size and frequency between each container); analyze the storage volume mounting topology (e.g., container C mounts the shared storage volume / data1), and generate a load distribution topology diagram containing dependencies and storage associations.

[0176] In step S53 above, the resource pools are sorted by cost priority: local SSD storage (high cost / high performance) < local HDD storage (medium cost / large capacity) < cloud storage (low cost / elastic expansion); resources are allocated starting from the lowest cost pool, for example, local HDDs are used first to supplement the storage gap, and cloud storage is called when local resources are insufficient, until the gap is met (e.g., allocate 80TB of local HDD first, and then call 20TB of cloud storage).

[0177] Based on the utilization of CPU, memory, and network bandwidth, for example, if node X has 30% CPU remaining and 40% memory remaining, and a tolerance score of 70 out of 100, select the three nodes with the highest tolerance as the target migration nodes; design the migration phases in the order of "stateless services (such as Flink tasks) → stateful services (such as Kafka clusters)", first migrate stateless services (taking 10 minutes), then migrate stateful services (which require data synchronization and take 30 minutes) to ensure uninterrupted service.

[0178] In step S54 above, DataNode nodes are added to the HDFS cluster via the veStack virtualization layer API (e.g., adding 5 physical machines, each with 10TB of storage), and container quotas are allocated to YARN (e.g., adding 200 vCPUs and 400GB of memory). Network isolation is performed on isolated nodes (e.g., blocking their public IP addresses in the firewall), and load migration is triggered simultaneously: first, stateless services are migrated to target node X, and then stateful services are migrated to target node Y. The migration progress is monitored in real time. When the storage volume data synchronization is completed (e.g., 10GB of data migration takes 15 minutes), it is confirmed that the load on the original node has been completely transferred, and the new resource allocation results are output (e.g., the total HDFS storage reaches 1.05PB, and the available YARN resources increase by 20%).

[0179] In a preferred embodiment of the present invention, step S6: based on the scheduled resource allocation results and the updated security policy package, visually displaying the cluster operating status through a unified management interface, and generating a deliverable software installation package and deployment documentation, may include:

[0180] Step S61: Based on the scheduled resource allocation results and the updated security policy package, extract the visualization metadata, parse the cluster resource topology from the resource allocation results, and parse the policy configuration snapshot from the security policy package;

[0181] Step S62: Based on the visual metadata, through the rendering engine of the unified management interface, generate a third-order color gradient by CPU or memory usage for each node and overlay a 2Hz flashing alarm icon on isolated nodes to build a dynamic resource heatmap component; at the same time, build a rule tree topology of application control → URL filtering → virus protection with the policy version identifier as the root node, add a red pulse highlight border to high-risk rules, generate a policy dependency graph component, and finally output an embeddable visual interface component;

[0182] Step S63: Based on the cluster resource topology and policy configuration snapshot, build a software installation package by encapsulating the veStack virtualization layer driver and cluster component deployment scripts, and injecting the updated security policy package. At the same time, convert the resource topology into AnsiblePlaybook and the policy configuration snapshot into a Markdown format configuration manual to generate deployment documents and deliverables.

[0183] Step S64: Integrate deliverables and visual interface components, create a delivery zone in the unified management interface, provide a software installation package download channel and embed deployment documents that support online preview, and finally output an auditable delivery report containing interface access links, installation package hash values, and document version numbers.

[0184] In this embodiment of the invention, the solution lays a solid data foundation for visualization and delivery. Accurate and comprehensive data extraction ensures that the visualization interface and delivery documents truthfully reflect the status of cluster resources and security policies, eliminating information bias. Intuitive visualization helps operations personnel quickly understand cluster resource usage and security policy execution, accurately pinpointing resource bottlenecks and high-risk rules, thereby efficiently carrying out resource allocation and policy optimization, and improving operational efficiency. The standardized delivery model significantly reduces labor costs and operational errors during cluster deployment, facilitating rapid deployment, migration, and maintenance of the cluster, and improving delivery efficiency and quality. Complete and traceable delivery results, coupled with auditable reports, not only facilitate project acceptance and auditing but also clarify responsibility, ensure the accuracy of delivered content, and provide clear records for subsequent version management and updates, comprehensively improving the standardization and reliability of delivery work.

[0185] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:

[0186] Step S61 above parses the cluster resource topology information from the resource allocation results after scheduling, including the node composition of major data clusters (such as Hadoop clusters and Doris clusters), the connection relationships between nodes, and resource allocation (such as the number of CPU cores, memory capacity, and storage quota allocated to each node). At the same time, it extracts policy configuration snapshots from the updated security policy package, covering the policy execution dependency tree, the protection rules corresponding to each risk level node (the specific content of application control, URL filtering, virus protection, etc. rules), policy version identifiers, and other information. Through the accurate extraction of these data, visual metadata is formed, providing basic data for subsequent visualization and delivery document generation.

[0187] In step S62 above, the rendering engine of the unified management interface is used to process the visual metadata on a node-by-node basis. For each node, a three-level color gradient is assigned based on its CPU or memory usage: nodes with usage below a threshold (e.g., 60%) are displayed in green, indicating normal resource usage; nodes with usage between 60% and 80% are displayed in yellow, indicating that resource usage is approaching saturation; nodes with usage exceeding 80% are displayed in red, indicating that resource usage is strained; for nodes in an isolated state, a flashing alarm icon at a 2Hz frequency is superimposed to highlight their abnormal state; based on this, a dynamic resource heatmap component is constructed to intuitively display the resource load distribution of the entire cluster.

[0188] Meanwhile, using the policy version identifier as the root node, a tree-like topology graph is constructed by organizing application control, URL filtering, and virus protection rules according to their execution order and dependencies. For high-risk rules (such as protection rules involving high-risk nodes), a red pulse highlight border is added to make them more eye-catching in the graph. Finally, the dynamic resource heatmap component and the policy dependency graph component are integrated to output a visual interface component that can be embedded into the unified management interface.

[0189] In step S63 above, based on the extracted cluster resource topology and policy configuration snapshots, the software installation package and deployment documentation are constructed. When constructing the software installation package, the veStack virtualization layer driver is encapsulated to ensure its compatibility with the cluster environment. Deployment scripts for various data cluster components are integrated, such as the deployment scripts for HDFS 3.3.4, YARN 3.3.4, and MapReduce 23.3.4 components of the Hadoop cluster, and the deployment script for the Doris 2.0.7 component of the Doris cluster, etc., and an updated security policy package is injected to make the installation package contain all the content required for complete cluster deployment and security protection.

[0190] In terms of generating deployment documentation, the cluster resource topology is converted into Ansible Playbook format for easy use during automated deployment; policy configuration snapshots are converted into Markdown format configuration manuals that record detailed configuration information, effective conditions, and execution actions of security policies, forming standardized deployment documentation; finally, the software installation package and deployment documentation are integrated into a deliverable.

[0191] Step S64 above integrates the deliverables (software installation packages and deployment documents) with the visual interface components, creating a delivery zone in the unified management interface. Within this delivery zone, a download channel for the software installation package is provided, and the deployment documents are displayed in an embedded format, supporting online preview. Simultaneously, unique identification information is generated for each deliverable, including the interface access link, installation package hash value, and document version number. This information is then aggregated to generate an auditable delivery report. The report can also record information such as delivery time and personnel, ensuring the traceability and auditability of the delivery process.

[0192] like Figure 2 As shown, embodiments of the present invention also provide a cloud platform-based resource management system, including:

[0193] The allocation module is used to allocate computing and storage resources to the big data cluster through the veStack virtualization layer;

[0194] The loading module is used to load an initial security policy package for the cluster based on computing and storage resources. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall.

[0195] The detection module is used to monitor network traffic and access behavior of cluster resources in real time based on the initial security policy package. It identifies active external hosts, risky domains and IPs through the protection analysis unit, detects abnormal communication traffic between public IP assets and the Internet, and generates a security incident assessment report.

[0196] The dynamic update module is used to dynamically update the policy management component of the perimeter firewall based on the security incident assessment report, and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies.

[0197] The scheduling module is used to dynamically scale virtualization resources based on updated security policy packages, combined with cluster load status and cost optimization objectives, and generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; and automatically isolating and migrating the load of cluster nodes with unprocessed security events.

[0198] The visualization module is used to visualize the cluster's operating status through a unified management interface based on the scheduled resource allocation results and updated security policy packages, and to generate deliverable software installation packages and deployment documents.

[0199] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0200] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0201] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A resource management method based on a cloud platform, characterized in that, The method includes: Step S1: Allocate computing and storage resources to the big data cluster through the veStack virtualization layer; Step S2: Based on computing and storage resources, load an initial security policy package for the cluster. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall. Step S3: Based on the initial security policy package, monitor the network traffic and access behavior of cluster resources in real time, identify active external hosts, risky domains and IPs through the protection analysis unit, detect abnormal communication traffic between public IP assets and the Internet, and generate a security incident assessment report; Step S4: Based on the security incident assessment report, dynamically update the policy management component of the perimeter firewall and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies; Step S5: Based on the updated security policy package, combined with the cluster load status and cost optimization objectives, dynamically scale up and down virtualization resources to generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; automatically isolating and migrating the load of cluster nodes with unprocessed security events. Step S6: Based on the scheduled resource allocation results and the updated security policy package, visualize the cluster operation status through a unified management interface, and generate a deliverable software installation package and deployment documentation.

2. The resource management method based on a cloud platform according to claim 1, characterized in that, Step S1: Allocate computing and storage resources to the big data cluster through the veStack virtualization layer, including: A big data cluster framework is built using the veStack virtualization layer. The big data cluster framework includes the logical topology of Hadoop cluster, Doris cluster, Flink cluster, StarRocks cluster and OpenSearch cluster. Based on the logical topology of the OpenSearch cluster, initial computing and storage resources are allocated to each cluster, and a resource allocation plan is generated. According to the resource allocation plan, cluster components are deployed in the veStack virtualization layer, and the running status of the cluster components is output, including: HDFS 3.3.4, YARN 3.3.4 and MapReduce 23.3.4 components are deployed for the Hadoop cluster, Doris 2.0.7 components are deployed for the Doris cluster, and Flink 1.16.1 and Paimon 0.6.1 components are deployed for the Flink cluster; Based on the running status of cluster components, resource quotas are dynamically adjusted, and the allocated computing and storage resources are output.

3. The resource management method based on a cloud platform according to claim 2, characterized in that, Step S2: Based on computing and storage resources, load the initial security policy package for the cluster. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the perimeter firewall, including: Based on computing and storage resources, the network exposure surface of cluster nodes is calculated to obtain node network exposure status data. For the nodes in the calculated network exposure status data, nodes bound to public IP addresses are marked as publicly exposed nodes. For the storage utilization data of all cluster nodes, nodes with storage utilization exceeding a preset threshold are marked as data storage nodes. Based on the publicly exposed nodes, data storage nodes, and the logical connection relationships between nodes, a security domain topology containing security labels and logical connection relationships is output. Based on the security domain topology, for nodes exposed to the public network, their open ports are scanned. If risky ports are found, inbound traffic port filtering rules are generated, and unnecessary ports are closed according to the business whitelist. For data storage nodes, their outbound traffic application protocols are analyzed. If non-database protocols are found, outbound traffic application control rules are generated, high-risk protocol traffic is blocked, and an effective set of access control policies is output. Based on the effective status of the access control policy set, the traffic statistics nodes deployed inside the border firewall identify active outbound behavior through the traffic matrix between the statistical nodes. When the frequency of node outbound connections exceeds the baseline value, the active outbound connection detection engine is triggered to capture risky domain names and IPs. At the same time, the risk awareness nodes deployed in the core switching layer extract communication characteristics from the traffic of the mirrored public network exposed nodes. If traffic matching the risk IP database is detected, the abnormal communication traffic analysis engine is triggered to generate an alarm and output a protection analysis rule deployment report. By integrating access control policy sets and protection analysis rule deployment reports, the boundary policy rules are bound to the cloud firewall detection engine to generate a unified policy execution priority list and output the initial security policy package.

4. The resource management method based on a cloud platform according to claim 3, characterized in that, Step S3: Based on the initial security policy package, monitor network traffic and access behavior of cluster resources in real time. Identify proactively connecting hosts, risky domains and IPs through the protection analysis unit, detect abnormal communication traffic between public IP assets and the Internet, and generate a security incident assessment report, including: Based on the initial security policy package, the traffic monitoring probe is activated in the protection analysis unit. According to the engine configuration parameters of the initial security policy package, a lightweight log collection agent is deployed on the cluster nodes. According to the policy priority list, full traffic capture is used for data storage nodes and sampling capture is used for computing nodes. Real-time network traffic logs with node labels are output. Based on real-time network traffic logs, the external analysis engine filters out request packets with missing outgoing SYN or ACK responses in the logs, parses the HTTPHost header and DNSQuery field of the request packets, and outputs a list of target domain names and IPs. At the same time, the public network traffic analysis engine mirrors the TCP or UDP session streams of public network IP nodes, calculates the average packet size, traffic burst variance, and port access frequency of the session, and outputs the session feature vector. Input the session feature vector into the baseline behavior model, generate abnormal communication alarms for packets whose mean size deviates from the baseline by more than 50% or whose traffic burst variance is greater than the threshold, and output an abnormal communication traffic alarm set; at the same time, query the threat intelligence database for the target domain name and IP list, mark those in the intelligence database as confirmed risk items, and mark those not in the intelligence database but whose access frequency is more than 3 times the baseline value as suspected risk items, and output an active outbound risk list; Based on the proactive external connection risk list and abnormal communication traffic alarm set, the risk domain name and IP are associated with the abnormal traffic session to calculate the security incident risk level and generate a security incident assessment report.

5. The resource management method based on a cloud platform according to claim 4, characterized in that, Step S4: Based on the security incident assessment report, dynamically update the policy management component of the perimeter firewall and output an updated security policy package, including: application control, URL filtering, and virus protection rules enabled for cluster nodes at risk, and adjusted inbound and outbound traffic access control policies, including: Based on the security incident assessment report, the incident classification and risk node list are analyzed, and a risk node protection requirement set including high-risk nodes, medium-risk nodes, and low-risk nodes is output. Based on the risk node protection requirement set, application control rules, URL filtering rules, and virus protection rules are created for high-risk nodes, application control rules and URL filtering rules are created for medium-risk nodes, and virus protection rules are created for low-risk nodes, generating a dynamic protection rule set; Based on the risk node protection requirement set, for inbound traffic, if the risk node is exposed to the public network, a port access restriction rule is generated; if there is abnormal traffic with DDoS characteristics, a traffic rate restriction rule is generated. For outbound traffic, if the risk node is connected to a risky IP, a target IP blocking rule is generated; if the outbound traffic exceeds the business baseline, a bandwidth limiting rule is generated. The traffic control rule set is recalculated and generated. Based on dynamic protection rule sets and traffic control rule sets, a policy execution dependency tree is constructed, a unified policy version identifier is generated, and an updated security policy package is output.

6. The resource management method based on a cloud platform according to claim 5, characterized in that, Step S5: Based on the updated security policy package, combined with cluster load status and cost optimization objectives, dynamically scale virtualization resources up and down, generating scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet security risk levels; automatically isolating and migrating the load of cluster nodes with unhandled security events, including: Based on the updated security policy package, the policy execution dependency tree is parsed to identify clusters that meet the security risk level, a list of risk nodes is extracted and nodes with unprocessed security events are marked, and a resource scheduling instruction set is output, including HDFS and YARN resource expansion instructions for compliant clusters and node isolation and load migration instructions for risk nodes. Based on the resource scheduling instruction set, for expansion instructions, the current CPU, memory, and storage utilization of the cluster are collected, the difference between the target performance index and the current state is calculated, and the resource gap is output; for migration instructions, the container group dependencies of the nodes to be migrated are scanned, the container group communication matrix and storage volume mounting topology are drawn, and the load distribution topology is output. Based on the resource gap and load distribution topology, resource expansion is sorted by cost priority and allocated from the lowest cost pool until the gap is met, generating a minimum resource allocation scheme; for load migration, the load tolerance of cluster nodes is calculated, the target node with the highest tolerance is selected, and a phased migration order of stateless services first and stateful services is designed to generate a zero-interruption migration path. Based on the minimum resource allocation scheme and zero-interruption migration path, the system calls the API through the veStack virtualization layer to add DataNode nodes to the HDFS cluster and allocate YARN container quotas, triggers real-time migration of isolated nodes, and outputs the resource allocation results after scheduling.

7. The resource management method based on a cloud platform according to claim 6, characterized in that, Step S6: Based on the scheduled resource allocation results and the updated security policy package, visualize the cluster's operating status through a unified management interface, and generate a deliverable software installation package and deployment documentation, including: Based on the scheduled resource allocation results and the updated security policy package, extract the visualization metadata, parse the cluster resource topology from the resource allocation results, and parse the policy configuration snapshot from the security policy package; Based on visual metadata, a dynamic resource heatmap component is constructed by using a unified management interface rendering engine to generate a third-order color gradient based on CPU or memory usage at the node level and overlaying a 2Hz flashing alarm icon on isolated nodes. At the same time, a rule tree topology of application control → URL filtering → virus protection is constructed with the policy version identifier as the root node, and a red pulse highlight border is added to high-risk rules to generate a policy dependency graph component. Finally, an embeddable visual interface component is output. Based on cluster resource topology and policy configuration snapshots, a software installation package is built by encapsulating the veStack virtualization layer driver and cluster component deployment scripts, and injecting updated security policy packages. At the same time, the resource topology is converted into AnsiblePlaybook and the policy configuration snapshot is converted into a Markdown format configuration manual to generate deployment documents and deliverables. Integrate deliverables and visual interface components, create a delivery zone on a unified management interface, provide a software installation package download channel and embed deployment documents that support online preview, and finally output an auditable delivery report that includes interface access links, installation package hash values, and document version numbers.

8. A cloud-based resource management system, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The allocation module is used to allocate computing and storage resources to the big data cluster through the veStack virtualization layer; The loading module is used to load an initial security policy package for the cluster based on computing and storage resources. The initial security policy package includes the protection analysis rules of the cloud firewall and the access control policies of the boundary firewall. The detection module is used to monitor network traffic and access behavior of cluster resources in real time based on the initial security policy package. It identifies active external hosts, risky domains and IPs through the protection analysis unit, detects abnormal communication traffic between public IP assets and the Internet, and generates a security incident assessment report. The dynamic update module is used to dynamically update the policy management component of the perimeter firewall based on the security incident assessment report, and output the updated security policy package, including: application control, URL filtering and virus protection rules enabled for cluster nodes with risks, and adjusted inbound and outbound traffic access control policies. The scheduling module is used to dynamically scale virtualization resources based on updated security policy packages, combined with cluster load status and cost optimization objectives, and generate the scheduled resource allocation results, including: increasing HDFS and YARN resource quotas as needed for clusters that meet the security risk level; and automatically isolating and migrating the load of cluster nodes with unprocessed security events. The visualization module is used to visualize the cluster's operating status through a unified management interface based on the scheduled resource allocation results and updated security policy packages, and to generate deliverable software installation packages and deployment documents.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Firewall chip data packet buffer management method

    CN101212451A

  • Network protection evaluation information processing method and system

    CN119135448A