Cloud platform design method based on Docker container technology
By building a cloud platform design method based on Docker container technology, adopting distributed management clusters, multi-level network isolation and dynamic container scheduling strategies, the shortcomings of traditional cloud platforms in network isolation and security protection are solved, and the reliability and security of the cloud platform are improved.
Patent Information
- Application Number
- CN202511011370.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional overlay networks cannot provide multi-level network isolation and security protection in cloud platform design, resulting in low cloud platform reliability.
A cloud platform design method based on Docker container technology is constructed, including distributed management clusters, multi-level network isolation strategies, dynamic container instance scheduling strategies, and cluster management state synchronization strategies. Container instance isolation and scheduling are achieved through Docker Overlay network and label configuration, and secure isolation is achieved by combining encrypted communication and Ingress routing mode.
The reliability of the cloud platform has been improved. Through multi-level network isolation strategies and dynamic scheduling strategies, it can effectively block lateral penetration, control the scope of fault impact, and achieve efficient resource utilization and security isolation.
Smart Images

Figure CN120803614A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cloud computers, and particularly relates to a cloud platform design method based on Docker container technology. BACKGROUND
[0002] In recent years, cloud computing technology has developed rapidly, and Docker containers, as a lightweight virtualization technology, play an important role in the construction of cloud platforms. With the help of Docker containers and their orchestration tools, such as Docker Swarm, users can easily build cloud computing platforms and deploy and manage containerized application services. This container-based cloud platform architecture, through unified resource management and efficient scheduling capabilities, makes business deployment more flexible, resource utilization higher, and enables basic resource isolation and dynamic expansion between cluster nodes, meeting the needs of modern cloud computing environments for high performance and high availability.
[0003] However, although the traditional Overlay network supports cross-host container interconnection, it usually only provides simple network isolation, and cannot comprehensively isolate and secure from multiple levels, resulting in low reliability of cloud platform design. SUMMARY
[0004] Therefore, it is necessary to provide a cloud platform design method based on Docker container technology to solve the above technical problems. This method can improve the reliability of cloud platform design.
[0005] The present application adopts the following technical solutions: The present application provides a cloud platform design method based on Docker container technology, comprising: Based on the cloud platform, a distributed management cluster is constructed, a plurality of target application services in container are deployed in the distributed management cluster, and the number of container instances required by each target application service is determined; the distributed management cluster includes at least one management node and a plurality of working nodes; A container instance allocation strategy is constructed; the container instance allocation strategy is used to allocate a plurality of container instances to different working nodes for each target application service according to a load balancing strategy and the number of container instances; A multi-level network isolation strategy is configured for the plurality of target application services; the multi-level network isolation strategy corresponds to different Docker Overlay networks for the container instances of different target application services, and the internal network and the external network of the distributed management cluster are isolated from each other; A dynamic container instance scheduling strategy based on labels is constructed; the dynamic container instance scheduling strategy is used to schedule the container instances of the plurality of target application services in real time; A cluster management state synchronization strategy is constructed; the cluster management state synchronization strategy is used to replicate the cluster management state among multiple management nodes in real time, and synchronize and periodically backup the application state of running container instances, so as to make the running states of all nodes in the distributed management cluster consistent; the cluster management state includes service configuration, container instance replica state, network topology, node label, container scheduling information and security policy.
[0006] Preferably, the construction process of the distributed management cluster specifically includes: A plurality of hosts are obtained as nodes, and any one host is determined as a first management node, and at least two hosts are selected as worker nodes from the nodes other than the first management node; Docker container environments are installed on the first management node and the worker nodes, and network configuration is performed to interconnect all nodes; the management nodes use fixed IP and open control ports used for communication between the management nodes; The Docker Swarm cluster is initialized on the first management node to generate a join token; A join cluster command is executed on each worker node to join the Docker Swarm cluster, so as to construct the distributed management cluster.
[0007] Preferably, the method further includes: In response to improving the fault tolerance capability of the distributed management cluster, Docker container environments are installed on a preset number of hosts and management node join commands are executed, and the preset number of hosts are added as multiple management nodes of the cluster.
[0008] Preferably, the configuration process of the multi-level network isolation strategy specifically includes: At least one custom Docker Overlay network is created inside the distributed management cluster; Container instances of different target application services are respectively allocated to different Docker Overlay networks, so that different target application services are isolated from each other; The service ports provided by the container instances to the outside and the routing mechanism allow each container instance to provide services to the external network, so that the internal network and the external network of the distributed management cluster are isolated from each other.
[0009] Preferably, the method further includes: After creating at least one custom Docker Overlay network, encrypted communication is enabled for all Docker Overlay networks, and all external network accesses are forwarded to target container instances through the entrance of the distributed management cluster.
[0010] Preferably, the execution process of the dynamic container instance scheduling strategy specifically comprises: Tags representing attributes or requirements of all target application services and nodes are configured; the tags include service types, hardware resource attributes, and running environment identifiers; Container instances corresponding to the target application services are scheduled to nodes with corresponding tags according to the tags of the target application services; The running states of the multiple container instances of the target application services are monitored in real time, and the deployment of the container instances is adjusted in real time according to the running states.
[0011] Preferably, the running states of the multiple container instances of the target application services are monitored in real time, and the deployment of the container instances is adjusted in real time according to the running states, and the method specifically comprises: When a node with a specific tag is detected to be faulty or overloaded, the running state of a container instance running on the faulty or overloaded node is determined to be abnormal; If the running state of the container instance is abnormal, the container instance with the abnormal running state is scheduled to a node with the same tag.
[0012] Preferably, the method further comprises: When it is detected that the load of the target application service exceeds a preset threshold, container instances with the same tag as the target application service whose load exceeds the preset threshold are automatically added; The added container instances are allocated to nodes with corresponding tags.
[0013] Preferably, the cluster management state is real-time replicated among multiple management nodes, and the method specifically comprises: A consistent key-value database is established among the multiple management nodes; the key-value database includes all metadata related to container instance scheduling; the metadata includes service configurations, replica states, network topologies, and tags; The cluster management state is real-time replicated among all the management nodes to update the key-value database.
[0014] Preferably, the method further comprises: After the application states of the running container instances are synchronized and periodically backed up, when a container instance is abnormally terminated, a newly started container instance acquires the backed-up application state and recovers according to the acquired application state.
[0015] The application provides a cloud platform design device based on a Docker container technology, which comprises: A construction module is configured to construct a distributed management cluster based on a cloud platform, deploy containerized multiple target application services in the distributed management cluster, and determine the number of container instances required by each target application service; the distributed management cluster comprises at least one management node and multiple working nodes; an allocation module configured to build a container instance allocation strategy; the container instance allocation strategy is used to allocate a plurality of container instances to different worker nodes according to a load balancing strategy and a number of container instances for each target application service; an isolation module configured to configure a multi-level network isolation strategy for the plurality of target application services; the multi-level network isolation strategy corresponds to different Docker Overlay networks for container instances of different target application services, and is used to distribute and manage the internal network and the external network of the distributed management cluster to be isolated from each other; a scheduling module configured to build a label-based dynamic container instance scheduling strategy; the dynamic container instance scheduling strategy is used to schedule the container instances of the plurality of target application services in real time; a synchronization module configured to build a cluster management state synchronization strategy; the cluster management state synchronization strategy is used to synchronize the cluster management state between a plurality of management nodes in real time, and synchronize and periodically backup the application state of the running container instances, so that the running states of all nodes in the distributed management cluster are consistent; the cluster management state includes service configuration, container instance replica state, network topology, node label, container scheduling information and security policy.
[0016] The application provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the cloud platform design method based on the Docker container technology.
[0017] The application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the cloud platform design method based on the Docker container technology when executing the program.
[0018] The above at least one technical solution adopted by the application can achieve the following beneficial effects: a label-based dynamic container instance scheduling strategy is built; the multi-level network isolation strategy corresponds to different Docker Overlay networks for container instances of different target application services, and is used to distribute and manage the internal network and the external network of the distributed management cluster to be isolated from each other, logical isolation domains are divided through virtual networks, and network conflicts between different application services are avoided; different application services are deployed to independent Overlay networks, and when a single service is attacked, horizontal penetration can be effectively blocked, and the fault influence range is controlled within a single network domain. The method can improve the reliability of the cloud platform design. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings, which are included to provide a further understanding of the application and constitute a part of this application, illustrate certain illustrative embodiments of the application and together with the description serve to explain the application. In the drawings:
[0020] Figure 1 A cloud platform design method flowchart based on the Docker container technology is provided in the present application. Figure 2 A cloud platform design method flowchart is provided in the present application. Figure 3 A cloud platform design device schematic diagram based on the Docker container technology is provided in the present application. Figure 4 A computer device for implementing the cloud platform design method based on the Docker container technology is provided in the present application. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in connection with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative work fall within the scope of protection of the present application.
[0022] Devices such as desktop computers, servers, notebook computers, etc. that can execute the solutions of the present application. For the convenience of description, only servers are taken as the execution subjects for description below.
[0023] The technical solutions provided by the embodiments of the present application will be described in detail below in connection with the drawings.
[0024] Figure 1 A cloud platform design method flowchart based on the Docker container technology is provided in the present application, which specifically includes the following steps: S101: Based on the cloud platform, a distributed management cluster is constructed, a plurality of target application services in containers are deployed in the distributed management cluster, and the number of container instances required by each target application service is determined respectively; the distributed management cluster includes at least one management node and a plurality of worker nodes.
[0025] In an exemplary embodiment, the construction process of the distributed management cluster specifically includes: obtaining multiple hosts as nodes, and determining any host as a first management node, and selecting at least two hosts from the nodes other than the first management node as worker nodes; installing a Docker container environment on the first management node and the worker nodes, and performing network configuration to enable interconnection between all nodes; the management nodes use fixed IP and open control ports used for communication between the management nodes; initializing the Docker Swarm cluster on the first management node to generate a join token; executing a join cluster command on each worker node to join the Docker Swarm cluster to construct the distributed management cluster.
[0026] In an exemplary embodiment, the method further includes: in response to improving the fault tolerance capability of the distributed management cluster, installing a Docker container environment on a preset number of hosts and executing a management node join command to add the preset number of hosts as multiple management nodes of the cluster.
[0027] Specifically, when preparing the nodes, hosts with different configurations are selected according to actual business requirements, for example, high-performance CPU, GPU or large-capacity storage nodes, to adapt to different types of containerized workloads; the management nodes are configured with fixed IP addresses to ensure that the work nodes and external systems can stably access their management service interfaces; to improve the scalability and fault tolerance of the cluster, multiple management nodes can be configured to form a highly available distributed management plane, and the number and specifications of the worker nodes can be flexibly expanded according to the load requirements to realize elastic allocation and efficient utilization of cluster resources.
[0028] The Docker container environment is installed on the management nodes and the worker nodes to ensure consistent Docker engine versions to ensure container compatibility and consistency of the cluster, and necessary operating system security reinforcement and system parameter optimization are performed to improve the stability and performance of container operation; in terms of network configuration, the management nodes should use fixed IP addresses, and the worker nodes and the management nodes should be in the same network domain or a network environment that can access each other; to ensure the security and efficiency of cluster communication, the Swarm cluster management port (TCP 2377), the node interconnection communication port (TCP / UDP 7946) and the Overlay network port (UDP 4789) need to be opened on each node, and firewall rules and IP white lists can be configured according to the security policy to ensure that only authorized nodes can access these ports, and port encryption and authentication mechanisms can be enabled to further improve the security of cluster communication.
[0029] It should be noted that the embodiment described in the management node is initialized by executing the docker swarm init command to initialize the Docker Swarm cluster, generate the join token and the management token, and distinguish the joining authority of the worker node and the management node, so as to ensure that the management node in the cluster has the corresponding leadership and scheduling responsibility; In the initialization process, the management node will automatically configure the distributed consistency storage (such as Raft protocol) to synchronize the management state, so as to ensure the real-time consistency of the cluster configuration and scheduling data; When the join token provided by the management node is used to execute the docker swarm join command on each worker node, the fixed IP address of the management node and the cluster management port need to be specified, so as to ensure that the worker node can successfully communicate with the management node and register as part of the cluster; At the same time, the certificate authentication can be performed during the joining process of the worker node, so as to enhance the security of the communication between the nodes in the cluster and avoid unauthorized nodes to access; Finally, a multi-node Swarm cluster architecture is formed, the management node uniformly schedules the container service, the worker node provides the computing resource and the container running environment, and supports the efficient operation and maintenance and resource management of the cloud platform.
[0030] Install Docker on the additional host and execute the management node join command, and use the management token to safely add the host as the second management node of the cluster. The management nodes share the configuration state and service orchestration metadata of the cluster through the distributed consistency protocol (such as Raft), so as to ensure that when any management node fails, other management nodes can automatically take over and continuously perform container scheduling and management, avoiding single point failure leading to service unavailability; Optionally, further increase the number of management nodes in the cluster to build a stronger high-availability distributed management plane, and dynamically adjust the number of management nodes according to the business size and reliability requirements to balance the load distribution and redundancy protection of the management layer. The management nodes can periodically rotate the leadership role or perform health state detection to continuously maintain the stability and consistency of the cluster.
[0031] Specifically, the preset number is set according to specific engineering practice.
[0032] S102: Construct a container instance allocation strategy; the container instance allocation strategy is used to allocate a plurality of container instances to different worker nodes for running according to the load balancing strategy and the number of container instances for each target application service.
[0033] Specifically, the containerized target application service is deployed by using the container orchestration capability of the management node, the number of container instance replicas required by the service and the running strategy are specified, and the cluster management node automatically allocates the container instances to different worker nodes for running according to the scheduling strategy of resource availability, label matching, and load status of the worker nodes, etc., to ensure efficient distribution and load balancing of the service. During deployment, the update strategy and rollback strategy of the service can be set according to specific business requirements to ensure smooth switching of containers during application version upgrade. The cluster supports dynamically increasing or decreasing the number of container instance replicas according to real-time business load conditions or operation and maintenance instructions during application running, and the automatic scaling mechanism can be automatically triggered combined with monitoring data or manually intervened by administrators, realizing elastic scaling of each application service in the cluster and efficient utilization of resources, ensuring rapid expansion to cope with business pressure under high load and automatically recycling resources to reduce costs under low load.
[0034] S103: configuring a multi-level network isolation strategy for the plurality of target application services; the multi-level network isolation strategy corresponds different Docker Overlay networks to the container instances of different target application services, and realizes mutual isolation between the internal network and the external network of the distributed management cluster.
[0035] In an exemplary embodiment, the configuration process of the multi-level network isolation strategy specifically includes: creating at least one custom Docker Overlay network inside the distributed management cluster; respectively allocating the container instances of different target application services to different Docker Overlay networks to realize mutual isolation between different target application services; allowing each container instance to provide services to the external network through the service port and routing mechanism provided by the container instance to realize mutual isolation between the internal network and the external network of the distributed management cluster.
[0036] In an exemplary embodiment, the method further includes: after creating at least one custom Docker Overlay network, enabling encrypted communication for all Docker Overlay networks, and all external network access is forwarded to the target container instance through the ingress of the distributed management cluster.
[0037] Specifically, at least one custom Overlay network is created by using a Docker network driver to carry the communication between containers on different hosts, each Overlay network can configure network name, driver type and subnet segment and other parameters according to the business security and logical isolation requirements, to support more fine-grained network management and virtualization; the containers of different services are respectively allocated to different Overlay networks, to ensure that the services are isolated from each other at the network level, the default intercommunication between Overlay networks is not allowed, only the explicitly configured inter-service link can realize cross-network communication, to avoid potential security threats or data leakage; at the same time, the containers inside the cluster are not exposed to the outside by default, only by specifying the pre-published port and Ingress routing configuration in the container deployment stage, the container is allowed to provide services to the outside, the Ingress routing mode forwards the external request to the target container instance through the cluster node unified entrance, to avoid directly exposing the internal network details of the container, combined with the encrypted Overlay network traffic and the necessary firewall rules, the internal container network and the external network are comprehensively isolated and protected at the network level, to significantly improve the security and reliability of the cluster.
[0038] Specifically, as shown in Figure 2 The multi-level network isolation strategy provided by the present application includes deploying different application services on different Docker Overlay networks, connecting the containers across hosts into isolated virtual networks by using Overlay networks, and realizing the network isolation between services by default intercommunication between Overlay networks.
[0039] The multi-level network isolation strategy described in the embodiments configures multiple Overlay networks on the management node by using the docker network create command, and combines different network names and configuration parameters (such as subnet segment, encryption options, etc.) to ensure the logical separation at the network layer; each service explicitly specifies the connected Overlay network during deployment, to avoid direct communication between different services in the physical network, the Overlay networks are securely isolated by the distributed routing of Swarm, unless the routing or gateway strategy is explicitly configured, otherwise the different Overlay networks will not communicate with each other by default, to ensure that the service networks in different business scenarios are independent and do not interfere with each other, and to improve the security of the overall cloud platform and the isolation capability in the multi-tenant environment.
[0040] The present application provides a cloud platform design method flow chart as shown in Figure 2 The present application provides a cloud platform design method flow chart as shown in Figure 2As shown in the figure, the Overlay network enables encrypted communication mode to encrypt and protect data traffic between containers; and the service adopts Ingress routing mode to publish when providing external network access. All external requests are uniformly forwarded to the target container through the cluster entrance, thereby avoiding direct exposure of the internal network where the container is located.
[0041] The Overlay network encrypted communication mode in this embodiment ensures that the Overlay network data traffic between containers on different hosts is encrypted and encapsulated at the network layer by enabling the encryption option (such as --opt encrypted=true) when creating the network or globally enabling encryption in the Swarm cluster configuration, thereby preventing the data from being sniffed or tampered with during physical network transmission; the service is exposed to the outside world using Swarm's Ingress routing, and all external requests first arrive at the ingress listening port of the cluster node, and are automatically forwarded by the internal routing mesh network to the container instance that actually carries the service, preventing external users from directly accessing the container IP and internal network, and cooperating with firewall and security group rules to significantly improve the security of the cloud platform in public network access scenarios.
[0042] S104: Construct a label-based dynamic container instance scheduling strategy; the dynamic container instance scheduling strategy is used to schedule container instances of multiple target application services in real time.
[0043] In an exemplary embodiment, the execution process of the dynamic container instance scheduling strategy specifically includes: configuring labels representing attributes or requirements for all target application services and nodes; the labels include service type, hardware resource attributes, and operating environment identifiers; scheduling the container instance corresponding to the target application service to the node with the corresponding label according to the label of the target application service; monitoring the operating status of multiple container instances of the target application service in real time, and adjusting the deployment of the container instance in real time according to the operating status.
[0044] Specifically, by adding label information such as "Hardware=GPU", "Environment=Production", and "Region=East China" to each node in the Docker node configuration file or API, and adding corresponding label scheduling policies such as --constraint 'node.labels.Hardware==GPU' in the Docker service create or update configuration of the container service, the scheduler will automatically match these labels during deployment and prioritize deploying container instances to nodes with the required characteristics, achieving a one-to-one correspondence between container and node characteristics, and optimizing resource utilization and the adaptability of application operation.
[0045] Specifically, labels representing hardware resource capabilities (such as GPU, SSD storage), node running environments (such as test, production environment), and service characteristics (such as front-end, back-end module) are pre-configured on the container service and cluster nodes to form a multi-dimensional resource and business attribute description system; when the container service is deployed, the management node dynamically filters the target nodes that meet the label requirements (such as "hardware = GPU" or "environment = production") attached to the container service, to ensure that the container instance can run on the matched node, thereby optimizing the service performance and resource usage.
[0046] In an exemplary embodiment, the running state of the plurality of container instances of the target application service is monitored in real time, and the deployment of the container instances is adjusted in real time according to the running state, specifically including: when a node with a specific label is monitored to be faulty or overloaded, the running state of the container instance running on the faulty or overloaded node is determined to be abnormal; if the running state of the container instance is abnormal, the container instance with the abnormal running state is scheduled to a node with the same label.
[0047] In an exemplary embodiment, the method further includes: when it is detected that the load of the target application service exceeds a preset threshold, automatically increasing a container instance with the same label as the target application service whose load exceeds the preset threshold; and assigning the newly added container instance to a node with the corresponding label.
[0048] Specifically, during the cluster running, the management node continuously monitors the running state and load of each node, and when it is found that some label-associated nodes are faulty, overloaded, or have performance bottlenecks, the system automatically triggers a re-arrangement to migrate the affected container instances to other nodes that meet the corresponding label requirements; in addition, when it is detected that the load of some services exceeds a preset threshold or the business demand changes, the system can automatically adjust the number of container instance replicas and their distribution locations based on the label constraints, to realize dynamic increase and decrease of container deployment and load balancing, further improving the overall scheduling flexibility of the cluster and the availability of the container service.
[0049] Specifically, the label-based orchestration mechanism further includes monitoring the resource utilization of each node and the service load, and when it is detected that a node associated with a certain label is faulty or overloaded, the affected container instances are automatically migrated to other nodes with the same label; when it is detected that the load of a certain service exceeds a preset threshold, the container instance replicas with the service label are automatically increased and assigned to the matched nodes, thereby realizing dynamic increase and decrease of the container instances and automatic adjustment of the deployment location.
[0050] Specifically, the monitoring and dynamic arrangement in the embodiment integrates Docker native commands with external monitoring tools (such as Prometheus or the built-in monitoring module of Swarm) to collect the CPU, memory and network usage of each node and the load of each container service in real time. When the load of a certain node continuously exceeds the standard or the state of the node is abnormal (such as being down or network being disconnected), the management node automatically reschedules the affected container task to other healthy nodes that meet the label requirements based on the label matching strategy, and dynamically increases the container replicas of the corresponding label when detecting a surge in business traffic, so as to ensure the satisfaction of the service level agreement (SLA) and the continuous availability of the business.
[0051] S105: Construct a cluster management state synchronization strategy; the cluster management state synchronization strategy is used to synchronize the cluster management state among the multiple management nodes in real time, and to synchronize and periodically backup the application state of the running container instance, so that the running states of all nodes in the distributed management cluster are consistent; the cluster management state includes service configuration, container instance replica state, network topology, node label, container scheduling information and security policy.
[0052] In an exemplary embodiment, the cluster management state is synchronized among the multiple management nodes in real time, specifically including: establishing a consistent key-value database among the multiple management nodes; the key-value database includes all metadata related to container instance scheduling; the metadata includes service configuration, replica state, network topology and label; the cluster management state is synchronized among all management nodes in real time to update the key-value database.
[0053] Specifically, the cluster management state is synchronized among the multiple management nodes in real time and consistency is checked by using the built-in distributed key-value database of Docker Swarm or external distributed data storage (such as etcd and Consul), so as to ensure that the key metadata including service configuration, container deployment location, network topology and label constraint are kept synchronized and consistent among all management nodes, and when any management node fails, other management nodes can seamlessly take over to avoid service interruption; at the same time, for the running container application state, the business state is periodically backed up and synchronized among the container replicas of the same service by mounting a distributed storage volume, integrating shared storage (such as NFS and Ceph) or using a message queue (such as Kafka), so as to ensure that when a container instance is restarted or switched due to maintenance, upgrade or node failure, the latest business state data can be quickly obtained and restored to the previous working state, avoiding service interruption or data inconsistency caused by state loss, so as to realize automatic and transparent state synchronization at the cluster level and the application level, and further improve the continuity and high availability of the containerized cloud platform.
[0054] Specifically, the interconnection and security protection among the cluster nodes are realized by cooperating the network configuration and secure communication between the management nodes and the worker nodes. The cluster is initialized by the docker swarm init command on the management node and a join token is generated. The worker nodes join the cluster by the docker swarm join command at the fixed IP address specified by the management node. The distributed consistency protocol is used to synchronize the management state among the management nodes, ensuring the overall configuration consistency of the cluster and automatic takeover in case of failure. Through the cooperation of the scheduling instructions between the management nodes and the worker nodes in the cluster and the container orchestration engine, the automatic deployment and load balancing distribution of the containerized target application service are realized. The management node dynamically schedules the container instances to the most suitable worker node according to the available resources, label attributes and real-time load status of the worker node, and automatically adjusts the number of container replicas when the load changes to improve resource utilization efficiency. Through the cooperation of the Docker network driver and the Overlay network encryption channel, multi-level network isolation is created inside the cluster to form an Overlay network isolation layer between different services. By default, the networks are not interconnected, and only when the service is needed, controlled access is performed through the published port and Ingress routing. Further cooperating with the encrypted communication and firewall rules, strict security isolation between the internal network of the cluster and the external network is realized. Through the cooperation of the distributed storage of cluster management metadata and business data and the multi-instance synchronization mechanism, the cluster management state is synchronized in real time among multiple management nodes through the distributed key-value database, and the business data is synchronized and updated among the running container replicas through shared storage or message queue, ensuring that when a container instance or node fails, other nodes or newly started container instances can quickly obtain the latest state information and seamlessly restore the business, avoiding service interruption and data inconsistency. Finally, the entire technical solution realizes the elastic expansion, security isolation, intelligent scheduling and high-availability operation of the cloud platform through the cooperation of node configuration and high-availability mechanism, scheduling engine and dynamic label matching, multi-layer network isolation and security protection, and state synchronization mechanism and container switching, effectively solving the problems of insufficient security isolation, poor scheduling flexibility and lack of business state synchronization in existing cloud platforms.
[0055] The distributed data storage synchronization management state is realized by establishing a consistent key-value database among multiple management nodes. All metadata related to container orchestration, including service configuration, replica state, network topology and label information, are stored in the database and updated and replicated in real time among the management nodes, so that when any management node fails, other management nodes hold the complete and latest cluster state.
[0056] It should be noted that the consistent key-value database in the embodiment can use the Raft distributed storage module built in the Docker Swarm, or in a more complex scenario, externally connect etcd / Consul as a cluster state center, and all management nodes of the cluster are synchronized in real time through heartbeat and data change events to manage the state, such as service scale, distributed location of container instances, Overlay network information, and label strategy metadata, to form a consistent cluster view. When any management node fails, other nodes have the same complete and latest state data, can immediately take over the leading role and continue to schedule and manage tasks, avoiding management single point failure.
[0057] In one embodiment of the present application, the present application provides a cloud platform design method flow chart as shown in Figure 2 As shown in Figure 2 The automatic state synchronization mechanism includes saving key application state data in the container running process to shared storage or synchronizing between container replicas through message communication. When a container instance is abnormally terminated, the newly started container instance can obtain the saved state data and resume execution, thereby ensuring the continuity of the application layer state.
[0058] It should be noted that the state synchronization in the embodiment can use mounting external shared storage volumes (such as NFS, GlusterFS, Ceph) to persist key business data at the application layer, or integrate message queues (such as Kafka, RabbitMQ) in the application architecture for event-driven state updates between multiple instances, to ensure that when a container instance is destroyed during node restart, maintenance, or migration, the newly started container replica can quickly mount the same persistent storage or obtain the latest state from the message queue, implement fast recovery of business session, connection information, and intermediate state, and guarantee seamless continuity of user experience.
[0059] In one embodiment of the present application, as shown in Figure 2 The number of configured management nodes can be multiple to constitute a redundant management plane, improve the fault tolerance of cluster scheduling, and the number of worker nodes can be elastically expanded or shrunk according to business needs. The newly added worker node is immediately included in the container scheduling range after performing the join cluster operation, and the deployment flexibility of the cloud platform is improved.
[0060] The management node in this embodiment can be configured as three or more nodes according to business continuity and availability requirements during actual deployment, and the high availability of the management layer can be maintained through the Raft election mechanism; the number of working nodes can be flexibly adjusted according to business peak and trough periods. After the new working node is installed with Docker and joined through the docker swarm join command, it can be detected by the management node and immediately assigned container tasks, realizing horizontal elastic expansion or contraction of the cluster to meet dynamically changing business needs while reducing overall resource usage costs.
[0061] Docker container technology is a cloud platform design method based on the embodiment of the present invention. Through the network configuration and secure communication of the management node and the working node, the unified management of the cluster and the elastic expansion of the node high availability are realized with the help of the docker swarm init and docker swarm join commands. The management node uses the distributed consistency protocol to keep the management metadata updated synchronously among the management nodes. When any management node fails, the other management nodes immediately take over the leading role to realize the seamless switching of cluster management. During the deployment of the container service, the management node efficiently allocates the container instance to the appropriate node according to the available resources, label configuration and real-time load of each working node in combination with the scheduling strategy. The cluster supports dynamic scaling to ensure the maximum resource utilization. Through the cooperation of the Docker network driver and the encryption configuration of the Overlay network, a multi-level network isolation is constructed. Services are isolated from each other by default and data traffic is encrypted. The Ingress routing only opens the necessary ports to ensure Ensure the security isolation of internal and external networks; combined with label-based dynamic container orchestration, management nodes perform intelligent scheduling according to multi-dimensional labels such as business attributes and hardware capabilities, realize dynamic migration and scaling of container instances, and cooperate with the state synchronization mechanism at the management and application layers through distributed data storage or shared storage, message queues, etc. to ensure that the cluster container instances can quickly recover to the latest state when the load fluctuates or the node fails, ensuring business continuity and high availability. Ultimately, through the combination of multi-dimensional technical features of multi-layer security network isolation, label intelligent orchestration, state synchronization and dynamic container scheduling, it effectively solves the problems of insufficient multi-tenant security isolation, poor scheduling flexibility and inefficient synchronization of business status in existing technologies, and significantly improves the security, elasticity and reliability of the cloud platform.
[0062] When applying the cloud platform design method based on Docker container technology provided by the present invention, it is not necessary to Figure 1 The steps are executed in the order shown. The specific execution order of the steps can be determined according to needs, and the present invention does not limit this.
[0063] The cloud platform design method based on the Docker container technology provided in the above one or more embodiments of the present application, based on the same idea, the present application also provides a corresponding cloud platform design device based on the Docker container technology, such as Figure 3 as shown.
[0064] Figure 3 The cloud platform design device based on the Docker container technology provided by the present application is shown in the schematic diagram, which comprises: The construction module 301 is configured to construct a distributed management cluster based on the cloud platform, deploy a plurality of target application services in a container in the distributed management cluster, and determine the number of container instances required by each target application service; the distributed management cluster comprises at least one management node and a plurality of working nodes.
[0065] The allocation module 302 is configured to construct a container instance allocation strategy; the container instance allocation strategy is used to allocate a plurality of container instances to different working nodes for each target application service according to the load balancing strategy and the number of container instances.
[0066] The isolation module 303 is configured to configure a multi-level network isolation strategy for a plurality of target application services; the multi-level network isolation strategy corresponds to different Docker Overlay networks for the container instances of different target application services, and the internal network and the external network of the distributed management cluster are isolated from each other.
[0067] The scheduling module 304 is configured to construct a dynamic container instance scheduling strategy based on a label; the dynamic container instance scheduling strategy is used to perform real-time scheduling on the container instances of a plurality of target application services.
[0068] The synchronization module 305 is configured to construct a cluster management state synchronization strategy; the cluster management state synchronization strategy is used to real-time replicate the cluster management state among a plurality of management nodes, and synchronize and periodically backup the application state of the running container instances, so that the running states of all nodes in the distributed management cluster are consistent; the cluster management state comprises service configuration, container instance copy state, network topology, node label, container scheduling information and security strategy.
[0069] The specific limitations of the cloud platform design device based on the Docker container technology can be referred to the limitations of the cloud platform design method based on the Docker container technology in the above, which will not be repeated here. The various modules in the above cloud platform design device based on the Docker container technology can be realized by software, hardware and their combinations in whole or in part. The above various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above various modules by the processor.
[0070] The application also provides a computer readable storage medium, which stores a computer program. Figure 1 The application provides a cloud platform design method based on a Docker container technology.
[0071] The application also provides a computer readable storage medium, which stores a computer program. Figure 3 The application also provides a computer readable storage medium, which stores a computer program. Figure 3 As shown in the structural schematic diagram of the computer device, the computer device comprises a processor, an internal bus, a network interface, a memory and a non-volatile memory, and can also comprise other hardware required by a business. Figure 1 The application provides a cloud platform design method based on a Docker container technology.
[0072] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the computer program can include the processes of the above-mentioned embodiment methods. In the embodiments of the application, any reference to a memory, a storage, a database or other media can include at least one of a non-volatile memory and a volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory or an optical memory. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration but not limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM).
[0073] The technical features of the above-mentioned embodiments can be combined in any manner, and for the sake of brevity, not all possible combinations of the technical features in the above-mentioned embodiments are described, however, as long as the combinations of the technical features do not exist, they should be considered as the range disclosed by the application.
Claims
1. A cloud platform design method based on Docker container technology, characterized in that: include: Building a distributed management cluster based on a cloud platform, deploying multiple containerized target application services in the distributed management cluster, and determining the number of container instances required for each target application service; the distributed management cluster includes at least one management node and multiple worker nodes; Constructing a container instance allocation strategy; the container instance allocation strategy is used to allocate multiple container instances to different working nodes for each target application service based on the load balancing strategy and the number of container instances; Configure a multi-level network isolation strategy for multiple target application services; the multi-level network isolation strategy corresponds to different Docker Overlay networks for container instances of different target application services, and isolates the internal network and external network of the distributed management cluster from each other; Build a dynamic container instance scheduling strategy based on labels; The dynamic container instance scheduling strategy is used to schedule container instances of multiple target application services in real time; Construct a cluster management state synchronization strategy; the cluster management state synchronization strategy is used to replicate the cluster management state in real time between multiple management nodes, and synchronize and regularly back up the application state of running container instances to ensure that the operating state of all nodes in the distributed management cluster is consistent; the cluster management state includes service configuration, container instance replica state, network topology, node labels, container scheduling information and security policies.
2. The method according to claim 1, wherein The construction process of the distributed management cluster specifically includes: Obtain multiple hosts as nodes, determine any one host as the first management node, and select at least two hosts from other nodes except the first management node as working nodes; Installing a Docker container environment on the first management node and worker nodes, and performing network configuration to interconnect all nodes; the management node uses a fixed IP address and opens a control port used for communication between management nodes; Initialize the Docker Swarm cluster on the first management node to generate a join token; Execute the join cluster command on each working node to join the Docker Swarm cluster to build the distributed management cluster.
3. The method according to claim 2, wherein The method further comprises: In response to improving the fault tolerance of the distributed management cluster, a Docker container environment is installed on a preset number of hosts and a management node joining command is executed to add the preset number of hosts as multiple management nodes of the cluster.
4. The method according to claim 1, wherein The configuration process of the multi-level network isolation strategy specifically includes: Creating at least one custom Docker Overlay network within the distributed management cluster; Assign container instances of different target application services to different Docker Overlay networks to isolate different target application services from each other. The service ports and routing mechanisms provided by container instances allow each container instance to provide services to the external network, thereby isolating the internal network and external network of the distributed management cluster.
5. The method according to claim 4, wherein The method further comprises: After creating at least one custom Docker Overlay network, enable encrypted communication for all Docker Overlay networks, and forward all external network access to the target container instance through the ingress of the distributed management cluster.
6. The method according to claim 1, wherein The execution process of the dynamic container instance scheduling policy specifically includes: Configure labels representing attributes or requirements for all target application services and nodes; the labels include service type, hardware resource attributes, and operating environment identifiers; Schedule the container instance corresponding to the target application service to the node with the corresponding label based on the label of the target application service; Monitor the running status of multiple container instances of the target application service in real time, and adjust the deployment of container instances in real time according to the running status.
7. The method according to claim 6, wherein The real-time monitoring of the running status of multiple container instances of the target application service and the real-time adjustment of the deployment of the container instances according to the running status specifically include: When a node with a specific label is detected to be faulty or overloaded, the running status of the container instance running on the faulty or overloaded node is determined to be abnormal; If the running status of the container instance is abnormal, the container instance with the abnormal running status will be scheduled to the node with the same label.
8. The method according to claim 7, wherein The method further comprises: When it is detected that the load of the target application service exceeds the preset threshold, the container instance with the same label as the target application service whose load exceeds the preset threshold is automatically added; Assign the newly added container instances to the nodes with the corresponding labels.
9. The method according to claim 1, wherein The real-time replication of cluster management status between multiple management nodes specifically includes: Establish a consistent key-value database across multiple management nodes; the key-value database includes all metadata related to container instance scheduling; the metadata includes service configuration, replica status, network topology, and labels; All management nodes replicate the cluster management status in real time to update the key-value database.
10. The method according to claim 1, wherein The method further comprises: After synchronizing and regularly backing up the application states of running container instances, when a container instance terminates abnormally, the newly started container instance will obtain the backed-up application states and restore based on the obtained application states.
Citation Information
Patent Citations
Application program automatic deployment method and system based on cloud native
CN116755794A
Cloud platform rapid deployment system and method
CN119396422A
Modularization-based Kubernetes cluster automatic deployment system
CN120104145A
Apparatus, systems and methods for container based service deployment
US20170257432A1
Cited By
Metadata management system and method for distributed storage system
CN121187514A