Containerized multi-copy deployment management method and device, terminal equipment and storage medium
Patent Information
- Application Number
- CN202511050009.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-25
AI Technical Summary
In containerized multi-replica deployment scenarios, how can we improve the efficiency and reliability of replica management to ensure business continuity and data consistency?
The container replica periodically updates its own status information to the real-time status list created and maintained by the cluster platform, and actively reads the status information of other replicas to build a decentralized collaborative management mechanism. Based on the IP and node health level, the master container replica is dynamically elected according to the preset master-slave election rules, and an automatic switching process is triggered when the master container is detected to be abnormal.
It improves the operating efficiency and stability of the cluster, enables self-healing from faults, enhances the effectiveness and reliability of replica management, and ensures business continuity and data consistency.
Smart Images

Figure CN121008818A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a containerized multi-copy deployment management method, apparatus, terminal device, and storage medium. Background Technology
[0002] In today's rapidly evolving digital landscape, containerization technology has become a crucial tool for building, deploying, and managing applications. It creates images of multiple applications from traditional computer software systems, then deploys them in multiple containers. Each container environment runs an application independently, collaboratively fulfilling the functions of the original computer software system. Traditional computer software systems typically use a master-slave redundant dual-machine setup to ensure system reliability. The application on the master machine handles business operations, while the application on the slave machine is in hot standby mode. When the master machine fails, the slave machine becomes the master, and the application on the slave machine takes over business processing, thus maintaining stable system operation.
[0003] In existing technologies, containerization technology is combined with clustering technology to ensure system reliability by running multiple replicas. Applications can be further divided into stateless applications and stateful applications. In stateless applications, there is no master-slave distinction among multiple replicas. During business processing, only one application needs to respond and process to achieve the goal. However, in stateful applications, there is a clear master-slave distinction. The current master replica application is responsible for processing business, while other replicas are in a hot standby state. This leads to complex master-slave management operations between multiple replica containers.
[0004] In containerized multi-replica deployment scenarios, how to improve the efficiency and reliability of replica management to ensure business continuity and data consistency is a problem that needs to be considered. Summary of the Invention
[0005] This application provides a containerized multi-replica deployment management method, apparatus, terminal device, and storage medium, which can effectively improve the efficiency and reliability of replica management in containerized multi-replica deployment scenarios to ensure business continuity and data consistency.
[0006] In a first aspect, embodiments of this application provide a containerized multi-replica deployment management method, including:
[0007] After a container replica starts, it periodically updates its own status information to a real-time status list. The real-time status list is created and maintained by the cluster platform in shared storage and is used to record the status information of each container replica, including IP address and node health level.
[0008] Read the status information of other container replicas in the real-time status list;
[0009] Based on the IP address and node health level in the real-time status list, a master container replica is dynamically elected according to the preset master-slave election rules.
[0010] When an anomaly is detected in the current primary container replica, an automatic failover process is triggered to re-elect a primary container replica.
[0011] In one possible implementation of the first aspect, the method further includes:
[0012] Monitor the resource utilization of the host node to which the container replica belongs;
[0013] Based on the resource utilization rate, determine the node health level of the host node to which the container replica belongs.
[0014] In one possible implementation of the first aspect, the step of dynamically electing a master container replica based on the IP address and node health level in the real-time status list according to a preset master-slave election rule includes:
[0015] The active replica with the highest node health level is selected as the primary container replica. The active replica refers to the container replica in the real-time status list whose time interval between the updated timestamp and the current time does not exceed a preset interval threshold.
[0016] If the nodes have the same health level, the active replica with the smallest IP address will be selected as the primary container replica.
[0017] In one possible implementation of the first aspect, the step of detecting an anomaly in the current primary container replica includes:
[0018] Detect whether the time interval between the update timestamp of the main container replica in the real-time status list and the current time exceeds a preset interval threshold;
[0019] If the time interval exceeds the preset interval threshold, the current main container replica is determined to be abnormal.
[0020] In one possible implementation of the first aspect, the step of detecting an anomaly in the current primary container replica includes:
[0021] If the main container replica status is marked as paused in the real-time status list, then the current main container replica is determined to be abnormal.
[0022] Alternatively, when its own replica is the primary container replica, if an active replica with a higher health level than its own replica node is detected in the real-time status list, then the current primary container replica is determined to be abnormal.
[0023] In one possible implementation of the first aspect, the step of triggering the automatic switchover process and re-electing a primary container replica includes:
[0024] Reset the replica status of its own replica in the real-time status list to the pending arbitration status;
[0025] The master container replica election is re-executed according to the preset master-slave election rules.
[0026] In one possible implementation of the first aspect, the status information further includes a replica status and a status control bit, the status control bit being used to identify control commands sent by the cluster platform; the method further includes:
[0027] When its own replica is the primary container replica and the status control bit is set to the switch to slave control command, the self-replica is downgraded, the replica status is updated to the slave container replica status, and an automatic switchover process is triggered.
[0028] When its own replica is a slave container replica and the status control bit is marked as a master switch control command, the own replica is upgraded, and the replica status is updated to the master container replica status.
[0029] Secondly, embodiments of this application provide a containerized multi-replica deployment management device, including:
[0030] The state synchronization unit is used to periodically update its own state information to the real-time state list after the container replica starts. The real-time state list is created and maintained by the cluster platform in shared storage and is used to record the state information of each container replica, including IP address and node health level.
[0031] A status information reading unit is used to read the status information of other container copies in the real-time status list;
[0032] The election unit is used to dynamically elect a master container replica based on the IP address and node health level in the real-time status list and according to the preset master-slave election rules.
[0033] The automatic failover unit is used to trigger the automatic failover process and re-elect a master container replica when an abnormality is detected in the current master container replica.
[0034] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the containerized multi-copy deployment management method as described in the first aspect above.
[0035] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the containerized multi-copy deployment management method as described in the first aspect above.
[0036] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the containerized multi-copy deployment management method described in the first aspect above.
[0037] In this embodiment, container replicas periodically update their own status information to the real-time status list created and maintained by the cluster platform, and actively read the status information of other replicas, thus constructing a decentralized collaborative management mechanism. This frees the master replica management process from dependence on a central control node. Based on IP and node health level, a master container replica is dynamically elected according to preset master-slave election rules. By combining dynamic performance indicators with static identifiers, the quality of the master replica is improved while ensuring the uniqueness of the election results. This optimizes management efficiency from the source and helps improve the overall operating efficiency and stability of the cluster. When a master replica anomaly is detected, an autonomous switching process is immediately triggered. Through the cohesive execution of anomaly detection and re-selection actions, fault self-healing is efficiently achieved, effectively upgrading replica management from a passive response to a proactive protection system. This significantly improves the effectiveness and reliability of replica management while reducing system complexity, ensuring business continuity and data consistency. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the implementation of the containerized multi-replica deployment management method provided in this application embodiment;
[0040] Figure 2 This is a flowchart illustrating a specific implementation of the containerized multi-replica deployment management method provided in this application for determining the health level of a node;
[0041] Figure 3 This is a flowchart illustrating a specific implementation of step S103 in the containerized multi-replica deployment management method provided in this application embodiment;
[0042] Figure 4 This is a flowchart illustrating a specific implementation of the containerized multi-replica deployment management method provided in this application for detecting anomalies in the primary container replica;
[0043] Figure 5 This is a flowchart illustrating a specific implementation of the containerized multi-replica deployment management method provided in this application, specifically the re-election of the primary container replica.
[0044] Figure 6 This is a flowchart illustrating a specific implementation of updating the replica status in the containerized multi-replica deployment management method provided in this application embodiment;
[0045] Figure 7 This is a structural block diagram of the containerized multi-replica deployment management device provided in the embodiments of this application;
[0046] Figure 8 This is a schematic diagram of the terminal device provided in the embodiments of this application. Detailed Implementation
[0047] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0048] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0049] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0050] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0051] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0052] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0053] The containerized multi-replica deployment management method provided in this application is applicable to various types of terminal devices that require containerized multi-replica deployment management. Specific terminal devices may include mobile phones, tablets, wearable devices, laptops, ultra-mobile personal computers (UMPCs), desktop computers, and servers, etc. This application does not impose any restrictions on the specific type of terminal device.
[0054] Figure 1 This paper illustrates the implementation flow of a containerized multi-replica deployment management method provided in an embodiment of this application. The method flow includes steps S101 to S104. The execution end of this method flow can be a container replica or its host node. The specific implementation principle of each step is as follows:
[0055] Step S101: After the container replica starts, it periodically updates its own status information to the real-time status list.
[0056] A container copy is an application copy created based on containerization technology, and each copy can run the application independently.
[0057] The real-time status list is created and maintained by the cluster platform in shared storage (such as the Redis real-time library), and is used to record the status information of each container replica. It is the carrier for storing the status information of all container replicas.
[0058] In one possible implementation, the status information includes a unique identifier for the container replica, an update timestamp, replica status, IP address, and node health level. The unique identifier is a unique identifier for the container replica within the network. The replica status is an important identifier in the container replica status information, used to intuitively reflect the replica's role and operational stage in the system. The IP address is a network identifier assigned to the container replica after startup, used for data transmission and communication, and is also unique. The update timestamp indicates the time when the container replica status information was updated. The node health level identifies the performance of the container replica and is used to measure the operational status of the host node to which the container belongs.
[0059] After the container replica starts, it periodically updates its own status information to the real-time status list at certain time intervals.
[0060] One possible implementation involves setting a timer in the container replica's own program code or using the operating system's scheduled task function. Once the container replica starts, the timer begins working, performing an update operation every set interval (e.g., 1 second). During the update, the container replica calls the relevant application programming interface (API) to encapsulate its unique identifier, update timestamp, replica status, IP address, and node health level into a specified data format, and then sends this information to a real-time status list on shared storage.
[0061] In this embodiment, the cluster platform maintains a real-time status list for applications in shared storage. Each container replica periodically updates its own status to the real-time status list. By periodically updating its own status information, the system ensures that the information of the container replica stored in the real-time status list is up-to-date. This allows the system to monitor the actual operating status of each replica in real time. When other container replicas perform operations such as master-slave election, they can make decisions based on accurate information, avoiding problems such as incorrect elections or inability to elect due to outdated information. This ensures the timeliness and accuracy of data, which is the cornerstone of the efficient operation of the entire system.
[0062] As one possible implementation of this application Figure 2 This paper illustrates a specific implementation process for determining node health levels in the containerized multi-replica deployment management method provided in this application embodiment, detailed below:
[0063] A1: Monitor the resource utilization of the host node to which the container replica belongs.
[0064] Resource utilization reflects the resource load status. By monitoring resource utilization, we can gain a comprehensive understanding of the resource load status of the host node to which the container replica belongs, providing original and crucial data for accurately assessing the node's health level. Lower resource utilization indicates better performance.
[0065] A2: Determine the node health level of the host node to which the container replica belongs based on the resource utilization rate.
[0066] The resource usage of host nodes is quantified into specific health levels, and the master-slave election rules provide an objective and quantifiable assessment basis. In this embodiment, a mapping relationship between resource utilization and node health level is pre-built. Based on this mapping relationship and the monitored resource utilization, the node health level of the host node to which the container replica belongs is determined.
[0067] In one possible implementation, resource utilization includes key indicators such as CPU utilization, memory utilization, disk utilization, and network bandwidth utilization. These key indicators directly reflect the resource consumption of host nodes in terms of computing, data storage, and transmission. For example, CPU utilization represents the proportion of time the CPU spends processing tasks within a certain period, while memory utilization reflects the proportion of currently used memory space in the total memory space. Based on key indicators such as CPU utilization, memory utilization, disk utilization, and network bandwidth utilization, and their corresponding preset weights, a comprehensive status score for the container replica is calculated. Then, combined with a preset mapping relationship between the comprehensive status score and node health level, the node health level corresponding to the container replica is determined. Comprehensive Status Score (S) = (CPU Utilization * 0.3 + Memory Utilization * 0.3 + Disk Utilization * 0.2 + Network Bandwidth Utilization * 0.2) * 100. For example, node health levels can be divided into three levels: healthy, good, and alarm. A comprehensive status score in the range [0, 30] indicates a healthy level, (30, 70] indicates a good level, and a score above 70 indicates an alarm level.
[0068] As one possible implementation of this application, the preset weights corresponding to key indicators such as CPU utilization, memory utilization, hard disk utilization, and network bandwidth utilization can be set according to the application scenario of the container replica and / or the business type associated with the container replica. Under different business scenarios and / or different business types, the preset weights corresponding to the same key indicator may differ.
[0069] For example, in the first application scenario, CPU utilization corresponds to the first preset weight, while in the second application scenario, CPU utilization corresponds to the second preset weight. The first and second preset weights are not equal, and the preset weights corresponding to memory utilization, hard disk utilization, and network bandwidth utilization are adjusted accordingly.
[0070] For example, when a container replica is associated with a first service type that has high bandwidth requirements, the network bandwidth utilization rate corresponds to the third preset weight. When a container replica is associated with a second service type that has low bandwidth requirements, the network bandwidth utilization rate corresponds to the fourth preset weight. The third preset weight is greater than the fourth preset weight, and the preset weights corresponding to memory utilization rate, disk utilization rate and network bandwidth utilization rate are adjusted accordingly.
[0071] Determining the weights of key resource utilization indicators based on application scenarios and business types helps ensure the alignment between node health levels and application scenarios or related business types, thereby ensuring the accuracy of container replica node health level determination and improving the reliability of replica master-slave election.
[0072] Step S102: Read the status information of other container replicas in the real-time status list.
[0073] Shared storage, as the core location for data storage, provides a unified data access point for all container replicas, ensuring data consistency and sharing. Any container replica can read the status information of other container replicas from the real-time status list of shared storage. The active replica in the current environment is identified by reading the status information. The active replica is the online container instance currently eligible to participate in master-slave arbitration, representing the dynamic election and switching decisions.
[0074] Step S103: Based on the IP address and node health level in the real-time status list, dynamically elect a master container replica according to the preset master-slave election rules.
[0075] The preset master-slave election rules are pre-defined election logic. In this embodiment, the election logic comprehensively considers two key factors: replica performance (node health level) and network identifier (IP address), ensuring the fairness and rationality of the election process. Dynamic master election means that the master container replica is not fixed but changes according to the actual situation.
[0076] Following preset master-slave election rules, the election is conducted based on real-time IP addresses and node health levels, enabling the rapid and accurate selection of the most suitable master container replica from numerous container replicas. This dynamic election method allows for timely adjustments to the master container replica selection based on the real-time status of each container replica within the cluster. This ensures the master container replica is always a relatively suitable one within the cluster, thereby enhancing the overall system reliability and stability and preventing the entire cluster service from being affected by master container replica failures or poor performance.
[0077] As one possible implementation of this application Figure 3 A specific implementation flow of step S103 in the containerized multi-replica deployment management method provided in this application embodiment is shown below:
[0078] B1: Prioritize the activity instance with the highest node health level as the main container instance.
[0079] In one possible implementation, the active copy also includes a pre-specified container copy. This pre-specification can be specified by the user or based on actual operational needs.
[0080] In one possible implementation, an active replica refers to a container replica in the real-time status list whose update timestamp and the current time interval do not exceed a preset interval threshold (e.g., 3 seconds). The timestamp mechanism ensures that replicas are active, preventing invalid replicas that have not been updated for a long time from being included in the election, which would reduce election reliability.
[0081] In this embodiment, only active replicas are eligible to participate in the election of the primary container replica. The node health levels of the active replicas are compared, and the one with the highest health level is selected as the primary container replica. Prioritizing the active replica with the highest node health level as the primary container replica ensures that the primary container replica runs on a host node with good performance and sufficient resources, thus providing a strong guarantee for its stable operation. Simultaneously, the active replica selection mechanism can exclude container replicas whose status information is not updated in a timely manner due to faults or network problems, preventing them from participating in the competition for the primary container replica, reducing the risk of primary container replica failure, and improving the reliability and stability of the entire cluster.
[0082] B2: If the node health levels are the same, the active replica with the smallest IP address will be selected as the primary container replica.
[0083] In this embodiment, when nodes have the same health level, IP address is introduced as a secondary election criterion. When multiple active replicas have the same node health level, their IP addresses are compared, and the active replica with the smallest IP address is selected as the primary container replica. IP address comparison can be implemented through simple lexicographical comparison or other predefined comparison rules.
[0084] In this embodiment, the election of container replicas uses a two-layer judgment logic that prioritizes health level and assists IP address. This takes into account the actual performance of replica operation (node health level) and resolves election conflicts in the same state through standardized rules (IP address), ensuring that the election process of the primary replica is fast, accurate and dynamically adaptable in containerized multi-replica scenarios.
[0085] In one possible implementation, the primary container replica election is triggered after a first preset time delay following the container replica startup delay, or when a primary container replica is detected.
[0086] In one possible implementation, the first election is performed after a first preset time delay following the startup of the container copy.
[0087] In one possible implementation, when it is detected that there is no master container copy in the active container copy, the election is performed after a second preset time.
[0088] The first preset time and the second preset time are time intervals preset according to the system's operating characteristics and election process requirements. They can be the same or different.
[0089] As one possible implementation of this application Figure 4 This paper illustrates a specific implementation process for detecting anomalies in the primary container replica within the containerized multi-replica deployment management method provided in this application embodiment, detailed below:
[0090] C1: Detect whether the time interval between the update timestamp of the primary container replica in the real-time status list and the current time exceeds a preset interval threshold. The update timestamp accurately reflects the update time of the replica status information and is an important time basis for determining whether the replica is operating normally.
[0091] C2: If the time interval exceeds the preset interval threshold, the current main container replica is determined to be abnormal.
[0092] In a containerized multi-replica deployment environment, replicas need to update their status information in a timely manner to ensure that the system understands their operational status. By detecting the time interval between the update timestamp and the current time, the system can quickly determine whether the primary container replica is active, avoiding business processing anomalies or data loss due to undetected anomalies in the primary container replica, thus improving the system's fault early warning capability.
[0093] As one possible implementation of this application, if the primary container replica status is marked as paused in the real-time status list, then the current primary container replica is determined to be abnormal. Paused update is a special replica status marker. When the primary container replica is unable to update its information to the real-time status list normally due to program crashes, resource exhaustion, human intervention, or other reasons, the system or cluster platform can actively mark its status as paused. This marker directly indicates that the replica can no longer maintain status updates as expected, and is a direct manifestation of an abnormal state.
[0094] As one possible implementation of this application, when the primary replica is the main container replica, if an active replica with a higher health level than the primary replica node is detected in the real-time status list, the current primary container replica is determined to be abnormal. When the primary replica is the main container replica, the health levels of other active replica nodes need to be monitored in real time. Once a replica with a higher health level is found, the primary replica is determined to be in an abnormal state. This abnormality mainly stems from the replica's performance no longer having an advantage, which may affect business processing efficiency and system stability.
[0095] An anomaly detection mechanism based on node health level comparison can dynamically assess the performance advantage of the primary container replica within the system. In a containerized multi-replica environment, the operating status of each replica changes dynamically with resource usage. By continuously comparing health levels, replicas with superior performance can be identified in a timely manner, avoiding problems such as slow business processing and resource waste caused by performance degradation of the primary container replica.
[0096] Step S104: When an abnormality is detected in the current primary container replica, an automatic switchover process is triggered to re-elect a primary container replica.
[0097] When the primary container replica fails, the automatic failover process can be initiated quickly to prevent business interruption due to primary replica failure. By re-electing the primary container replica, the system can quickly resume normal operation, ensuring business continuity. Meanwhile, based on the previously determined election rules and accurate status information, the newly elected primary container replica also possesses good performance and reliability, effectively preventing data loss, ensuring data consistency, and guaranteeing the stable and efficient operation of the containerized multi-replica deployment system.
[0098] In one possible implementation, container replicas automatically detect and confirm anomalies in the current primary container replica. In another possible implementation, the cluster platform monitors for anomalies in all container replicas, including the current primary container replica. When an anomaly is detected in the current primary container replica, an anomaly notification is sent to each container replica. Upon receiving the anomaly notification, the container replica triggers an automatic failover process to re-elect a primary container replica.
[0099] As one possible implementation of this application Figure 5 This paper illustrates a specific implementation process of re-electing the primary container replica in the containerized multi-replica deployment management method provided in this application embodiment, detailed below:
[0100] D1: Reset the replica status of its own replica in the real-time status list to the pending arbitration status.
[0101] The "Pending Arbitration" state is a special intermediate state among replica states. It's a temporary, isolated buffer state that container replicas enter during the primary replica switchover process. When an anomaly is detected in the current primary container replica and the automatic switchover process is triggered, each replica sets its own replica status in the real-time status list to "Pending Arbitration." This operation clearly indicates to the system that the replica is currently waiting to re-participate in the primary replica election and is temporarily not assuming the fixed responsibilities of a primary or secondary replica. This avoids election errors or business processing anomalies due to replica status confusion during the switchover process. Resetting its own replica status to "Pending Arbitration" provides a clear status identifier and starting point for the automatic switchover process.
[0102] D2: Re-execute the master container replica election according to the preset master-slave election rules.
[0103] In this embodiment, the election is carried out strictly in accordance with the preset master-slave election rules, which can comprehensively consider performance and network aspects to select the most suitable replica to undertake the main business processing tasks as the new master container replica.
[0104] In one possible implementation, after a third preset time delay, the master container replica election is re-executed according to the preset master-slave election rules. The third preset time is also a time interval pre-set based on system operating characteristics and election process requirements, for example, 2 seconds. The purpose of the delay is to provide a buffer time for the system, ensuring that each replica has sufficient time to complete state resets, information synchronization, and other preparatory work before the re-election, avoiding situations where some replicas' information is not updated or their states are inconsistent due to the election starting too early, thus affecting the accuracy of the election results.
[0105] As one possible implementation of this application, the status information also includes replica status and a status control bit, wherein the status control bit is used to identify control commands sent by the cluster platform. Replica status includes Main, Slave, toMain (ready to Main), and toSlave (ready to Slave), etc. The status control bit is a specific identifier, usually existing in numerical code form (e.g., 0 - no command, 1 - switch to Slave command, 2 - switch to Master command), and is set and sent by the cluster platform. As the management center of the entire containerized system, the cluster platform uses the status control bit to issue manual switching control commands to the replicas, thereby intervening in the replica roles.
[0106] Figure 6 This paper illustrates a specific implementation process for updating the replica status in the containerized multi-replica deployment management method provided in this application embodiment, detailed below:
[0107] E1: When its own replica is the primary container replica and the status control bit is set to the switch to slave control command, the own replica is downgraded, the replica status is updated to the slave container replica status, and an automatic switchover process is triggered.
[0108] When a container replica is currently in the master container replica state (i.e., its replica state is Main), and its status control bit receives a switch-to-slave control command sent by the cluster platform (e.g., the status control bit is set to 1), that replica will perform a degradation operation. The degradation process includes updating its own replica state to slave container replica state and actively triggering an automatic failover process to re-elect a new master container replica. During this process, the replica will suspend the business processing tasks it is currently undertaking as the master replica and release business processing permissions to the newly elected master replica.
[0109] By manually downgrading, the system can proactively relinquish the primary replica role when needed (such as when the primary replica node has insufficient resources or when maintenance is required), thus avoiding performance degradation or business interruption caused by forced operation.
[0110] E2: When its own replica is a slave container replica and the status control bit is marked as a master switch control command, the own replica is upgraded, and the replica status is updated to the master container replica status.
[0111] If a container replica is currently in a slave state and its status control bit receives a master switch command from the cluster platform (e.g., the status control bit is set to 2), that replica will perform an upgrade operation. The upgrade operation involves directly updating its replica status to the master container replica status (i.e., Main) and assuming the business processing responsibilities of the master replica. During the upgrade process, the replica will perform necessary initialization and data synchronization operations according to system requirements to ensure normal business processing.
[0112] Manual upgrades are suitable for scenarios where a specific replica needs to be designated as the primary replica. Through manual upgrades, the system can allocate resources more flexibly and optimize business processing efficiency.
[0113] In this embodiment, the combination of replica status and status control bits enables manual control of master-slave replica switching, complementing the automatic anomaly detection and switching mechanism. The manual control function allows the system to more flexibly adjust master-slave replica configurations when facing complex management needs (such as load balancing adjustments, node maintenance, and performance optimization), while the automatic switching mechanism ensures high availability under abnormal conditions. The combination of these two features guarantees system stability and reliability while enhancing manageability and adaptability, comprehensively improving the overall performance and operational efficiency of the containerized multi-replica deployment system.
[0114] As can be seen from the above, in this embodiment, the container replica periodically updates its own status information to the real-time status list created and maintained by the cluster platform, and actively reads the status information of other replicas, thus constructing a decentralized collaborative management mechanism. This frees the master replica management process from dependence on the central control node. Based on the IP and node health level, the master container replica is dynamically elected according to the preset master-slave election rules. By combining dynamic performance indicators with static identifiers, the quality of the master replica is improved while ensuring the uniqueness of the election results. This optimizes management efficiency from the source and is conducive to improving the operating efficiency and stability of the entire cluster. When an abnormality of the master replica is detected, the autonomous switching process is immediately triggered. Through the cohesive execution of anomaly detection and re-selection actions, fault self-healing is efficiently achieved, effectively upgrading replica management from a passive response to a proactive protection system. This significantly improves the effectiveness and reliability of replica management while reducing system complexity, ensuring business continuity and data consistency.
[0115] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0116] Corresponding to the containerized multi-replica deployment management method described in the above embodiments, Figure 7 A structural block diagram of a containerized multi-replica deployment management device provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0117] Reference Figure 7 The containerized multi-replica deployment management device includes: a state synchronization unit 71, a state information reading unit 72, an election unit 73, and an automatic switching unit 74, wherein:
[0118] The state synchronization unit 71 is used to periodically update its own state information to the real-time state list after the container replica starts. The real-time state list is created and maintained by the cluster platform in the shared storage and is used to record the state information of each container replica. The state information includes IP address and node health level.
[0119] The status information reading unit 72 is used to read the status information of other container copies in the real-time status list;
[0120] Election unit 73 is used to dynamically elect a master container replica based on the IP address and node health level in the real-time status list and according to a preset master-slave election rule;
[0121] Automatic switching unit 74 is used to trigger the automatic switching process and re-elect a primary container replica when an abnormality is detected in the current primary container replica.
[0122] As one possible implementation of this application, the containerized multi-replica deployment management device further includes:
[0123] The utilization monitoring unit is used to monitor the resource utilization of the host node to which the container replica belongs;
[0124] The health level determination unit is used to determine the node health level of the host node to which the container replica belongs based on the resource utilization rate.
[0125] As one possible implementation of this application, the election unit 73 is specifically used for:
[0126] The active replica with the highest node health level is selected as the primary container replica. The active replica refers to the container replica in the real-time status list whose time interval between the updated timestamp and the current time does not exceed a preset interval threshold.
[0127] If the nodes have the same health level, the active replica with the smallest IP address will be selected as the primary container replica.
[0128] As one possible implementation of this application, the containerized multi-replica deployment management device further includes an anomaly detection unit, used for:
[0129] Detect whether the time interval between the update timestamp of the main container replica in the real-time status list and the current time exceeds a preset interval threshold;
[0130] If the time interval exceeds the preset interval threshold, the current main container replica is determined to be abnormal.
[0131] As one possible implementation of this application, the anomaly detection unit is further configured to:
[0132] If the main container replica status is marked as paused in the real-time status list, then the current main container replica is determined to be abnormal.
[0133] Alternatively, when its own replica is the primary container replica, if an active replica with a higher health level than its own replica node is detected in the real-time status list, then the current primary container replica is determined to be abnormal.
[0134] As one possible implementation of this application, the automatic switching unit 74 is specifically used for:
[0135] Reset the replica status of its own replica in the real-time status list to the pending arbitration status;
[0136] The master container replica election is re-executed according to the preset master-slave election rules.
[0137] As one possible implementation of this application, the status information further includes a replica status and a status control bit, wherein the status control bit is used to identify control commands sent by the cluster platform; the containerized multi-replica deployment management device further includes:
[0138] The first update processing unit is used to downgrade its own replica to a slave container replica state and trigger an automatic switching process when its own replica is a primary container replica and the status control bit is a switch control command.
[0139] The second update processing unit is used to upgrade its own replica and update the replica status to the master container replica status when its own replica is a slave container replica and the status control bit is identified as a master switch control instruction.
[0140] As can be seen from the above, in this embodiment, the container replica periodically updates its own status information to the real-time status list created and maintained by the cluster platform, and actively reads the status information of other replicas, thus constructing a decentralized collaborative management mechanism. This frees the master replica management process from dependence on the central control node. Based on the IP and node health level, the master container replica is dynamically elected according to the preset master-slave election rules. By combining dynamic performance indicators with static identifiers, the quality of the master replica is improved while ensuring the uniqueness of the election results. This optimizes management efficiency from the source and is conducive to improving the operating efficiency and stability of the entire cluster. When an abnormality of the master replica is detected, the autonomous switching process is immediately triggered. Through the cohesive execution of anomaly detection and re-selection actions, fault self-healing is efficiently achieved, effectively upgrading replica management from a passive response to a proactive protection system. This significantly improves the effectiveness and reliability of replica management while reducing system complexity, ensuring business continuity and data consistency.
[0141] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0142] This application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements... Figures 1 to 6 The steps of any containerized multi-replica deployment management method are represented.
[0143] This application embodiment also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements... Figures 1 to 6 The steps of any containerized multi-replica deployment management method are represented.
[0144] This application also provides a computer program product that, when run on a terminal device, causes the terminal device to execute the implementation of... Figures 1 to 6 The steps of any containerized multi-replica deployment management method are represented.
[0145] Figure 8 This is a schematic diagram of a terminal device provided in an embodiment of this application. For example... Figure 8 As shown, the terminal device 8 in this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, it implements the steps in the various containerized multi-copy deployment management method embodiments described above, for example... Figure 1Steps S101 to S104 are shown. Alternatively, when the processor 80 executes the computer program 82, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 7 The functions of units 71 to 74 are shown.
[0146] For example, the computer program 82 may be divided into one or more modules / units, which are stored in the memory 81 and executed by the processor 80 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program 82 in the terminal device 8.
[0147] The terminal device 8 may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will understand that... Figure 8 This is merely an example of terminal device 8 and does not constitute a limitation on terminal device 8. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device 8 may also include input / output devices, network access devices, buses, etc.
[0148] The processor 80 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0149] The memory 81 can be an internal storage unit of the terminal device 8, such as a hard disk or memory of the terminal device 8. The memory 81 can also be an external storage device of the terminal device 8, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 8. Furthermore, the memory 81 can include both internal and external storage units of the terminal device 8. The memory 81 is used to store the computer program and other programs and data required by the terminal device. The memory 81 can also be used to temporarily store data that has been output or will be output.
[0150] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0151] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0152] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0153] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0154] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A containerized multi-replica deployment management method, characterized in that, include: After a container replica starts, it periodically updates its own status information to a real-time status list. The real-time status list is created and maintained by the cluster platform in shared storage and is used to record the status information of each container replica, including IP address and node health level. Read the status information of other container replicas in the real-time status list; Based on the IP address and node health level in the real-time status list, a master container replica is dynamically elected according to the preset master-slave election rules. When an anomaly is detected in the current primary container replica, an automatic failover process is triggered to re-elect a primary container replica.
2. The method according to claim 1, characterized in that, The method further includes: Monitor the resource utilization of the host node to which the container replica belongs; Based on the resource utilization rate, determine the node health level of the host node to which the container replica belongs.
3. The method according to claim 1, characterized in that, The step of dynamically electing a master container replica based on the IP address and node health level in the real-time status list according to a preset master-slave election rule includes: The active replica with the highest node health level is selected as the primary container replica. The active replica refers to the container replica in the real-time status list whose time interval between the updated timestamp and the current time does not exceed a preset interval threshold. If the nodes have the same health level, the active replica with the smallest IP address will be selected as the primary container replica.
4. The method according to claim 1, characterized in that, The steps for detecting anomalies in the current primary container replica include: Detect whether the time interval between the update timestamp of the main container replica in the real-time status list and the current time exceeds a preset interval threshold; If the time interval exceeds the preset interval threshold, the current main container replica is determined to be abnormal.
5. The method according to claim 1, characterized in that, The steps for detecting anomalies in the current primary container replica include: If the main container replica status is marked as paused in the real-time status list, then the current main container replica is determined to be abnormal. Alternatively, when its own replica is the primary container replica, if an active replica with a higher health level than its own replica node is detected in the real-time status list, then the current primary container replica is determined to be abnormal.
6. The method according to claim 1, characterized in that, The step of triggering the automatic failover process and re-electing the primary container replica includes: Reset the replica status of its own replica in the real-time status list to the pending arbitration status; The master container replica election is re-executed according to the preset master-slave election rules.
7. The method according to any one of claims 1 to 6, characterized in that, The status information also includes replica status and a status control bit, wherein the status control bit is used to identify control commands sent by the cluster platform; the method further includes: When its own replica is the primary container replica and the status control bit is set to the switch to slave control command, the self replica is downgraded, the replica status is updated to the slave container replica status, and an automatic switchover process is triggered. When its own replica is a slave container replica and the status control bit is marked as a master switch control command, the own replica is upgraded, and the replica status is updated to the master container replica status.
8. A containerized multi-replica deployment management device, characterized in that, include: The state synchronization unit is used to periodically update its own state information to the real-time state list after the container replica starts. The real-time state list is created and maintained by the cluster platform in shared storage and is used to record the state information of each container replica, including IP address and node health level. A status information reading unit is used to read the status information of other container copies in the real-time status list; The election unit is used to dynamically elect a master container replica based on the IP address and node health level in the real-time status list and according to the preset master-slave election rules. The automatic failover unit is used to trigger the automatic failover process and re-elect a master container replica when an abnormality is detected in the current master container replica.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the containerized multi-replica deployment management method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the containerized multi-replica deployment management method as described in any one of claims 1 to 7.
Citation Information
Cited By
Data reading method and system for mirror image copy of distributed parallel file system
CN122132359A
Data reading method and system for distributed parallel file system mirror copy
CN122132359B