System for non-blocking scheduling and coordinating database cluster of cloud native Postgres Operator

By introducing weighted connection pool management and system resource weighted timeout retry modules in Postgres Operator, the coordination interruption caused by Executor downtime or long-term non-response is solved, and resource allocation is optimized under high load conditions, improving system stability and resource utilization.

CN120011142APending Publication Date: 2025-05-16HIGHGO SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology causes the Postgres Operator to be unable to continue to coordinate when Executor is down or unresponsive for a long time, and under high load conditions, resource allocation is unreasonable, resulting in system stability and resource waste problems.

Method used

The non-blocking coordination system using cloud-native Postgres Operator includes a weighted connection pool management module and a system resource weighted timeout retry module. High-weight commands are preferred, and when the Executor coordination fails, the timeout time and number of retry times are adjusted according to the system resource idle situation and coordination weight.

Benefits of technology

It realizes automatic fault retry and timeout retry when Executor is down or unresponsive for a long time, avoids coordination interruption, improves system stability and resource utilization, and reduces the risk of coordination abnormalities under system resource waste and high load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011142A_ABST
    Figure CN120011142A_ABST
Patent Text Reader

Abstract

The invention provides a system for a non-blocking scheduling database cluster of a cloud native Postgres Operator, and the system comprises a weighted connection pool management module which is used for the socket connection of an Executor actuator; the system resource weighting overtime retry module is used for carrying out retry when Executor coordination fails; the system resource monitoring module is used for monitoring system resources and caching the system resources; and the alarm module is used for log recording and error alarm when Executor coordination fails. According to the method, when the coordination command is executed, fault retry and overtime retry can be automatically carried out, and when the set maximum retry number is reached, the Executor automatically quits the coordination process, so that the problem of coordination interruption caused by Executor downtime or long-time no response is avoided, and it is ensured that a fault node can be recovered in time; and in a high-load environment, system resource allocation is more reasonable by reasonably occupying and releasing system allocation resources during operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database technology, and in particular to a system for non-blocking coordination of a database cluster using a cloud-native Postgres Operator. Background Art

[0002] In the database field, the current way for Postgres Operator to achieve cluster coordination is to directly execute commands in the container through Executor. Executor is one of the core components to ensure the smooth operation of the database cluster, and is responsible for handling various operational tasks in the Kubernetes environment. Postgres Operator is designed to simplify the deployment, management, and maintenance of PostgreSQL databases on Kubernetes, and support high availability and disaster recovery in enterprise-level environments. In a high-availability environment, Executor manages database replication and failover operations. When a problem occurs in the primary node, PostgresOperator automatically switches to the backup node to ensure the continuous availability of the cluster. It monitors the status of database nodes to ensure that the system can respond to failures quickly.

[0003] However, when the executed command fails or becomes unresponsive due to an unprepared environment or network anomalies, the Executor will enter a continuous waiting state until the Postgres Operator is restarted, which directly causes the Postgres Operator to be unable to continue to coordinate the cluster. The existing solution for handling the failure of the Executor to cause the Postgres Operator to fail is to manually restart the Operator and check the logs to determine whether the coordination after restarting the Postgres Operator is successful. For the idea of ​​​​executor crash recovery, you can use the method of failure / timeout retry and limit the number of retries to restore the Executor from the crash and unresponsive state to the working state and record the log of the execution process.

[0004] However, the existing Executor unresponsiveness recovery method requires manual intervention throughout the entire process, and there is a high probability that repeated operations will be required. First, the operation and maintenance personnel need to confirm whether the cluster, network and other environments are normal, and then restart the Postgres Operator. When the Postgres Operator enters the coordination, it is necessary to combine the logs to repeatedly confirm whether the Executor will become unresponsive again. If at a certain moment, the Executor becomes unresponsive again, it will be difficult for the operation and maintenance personnel to discover it in the first time, and the Postgres Operator will not be able to restart itself. At this time, if other abnormalities occur in the cluster, the Postgres Operator will not be able to coordinate the cluster, which will cause the cluster status to be abnormal at the least, and the cluster will be unable to provide services to the outside world at the worst, thereby affecting upstream and downstream businesses.

[0005] That is to say, the timeout retry mechanism can usually solve the problem of Executor failure to respond to a certain extent, but this mechanism is designed to perform idempotent retries on functions or instructions, while ignoring the impact of system resource factors such as CPU, memory, and IO on the retry results. The Kubernetes environment often faces high-load scenarios. In this scenario, not only will the efficiency of the retry mechanism be greatly reduced, but the already limited resources will also be occupied by invalid retries, so that the timeout retry mechanism does not improve system stability. Summary of the invention

[0006] The technical problem to be solved by the present invention is how to avoid the problem of long-term blocking or inability to continue in the coordination process. And how to allocate resources under high load conditions to achieve stability in the coordination work. In view of this, the present invention provides a non-blocking coordination database cluster system of cloud native Postgres Operator.

[0007] The technical solution adopted by the present invention is that the system for non-blocking coordination of database clusters of the cloud-native Postgres Operator comprises: The weighted connection pool management module is used for the socket connection of the Executor executor. The connection weight of the socket connection is different according to the weight of the command executed by the Executor executor. When the Executor coordinates the coordination, the priority is determined according to the weight of the socket connection. The system resource weighted timeout retry module is used to retry when Executor coordination fails.

[0008] In one embodiment, the system further comprises: The system resource monitoring module is used to monitor and cache system resources.

[0009] In one embodiment, the system further comprises: The alarm module is used for logging and error alarms when Executor coordination fails.

[0010] In one embodiment, the system resource weighted timeout retry module is further used to: When Executor coordination fails, the coordination timeout is positively correlated with the coordination weight; and negatively correlated with the idleness of system resources. In response to the coordination failure, the coordination timeout increases until the configured timeout upper limit is triggered or the configured maximum number of retries is reached, at which time the Executor coordination retries are stopped.

[0011] In one embodiment, in the weighted connection pool management module, multiple upper limits of the number of socket connections are configured for each level of weight, so as to respond to coordination requests and execute coordination commands in parallel.

[0012] In one implementation, a corresponding alarm mode is determined according to the command weight.

[0013] In one embodiment, when the alarm mode is to record a log, the alarm system increases the recording of the log to a corresponding set weight level.

[0014] Another aspect of the present invention further provides an electronic device, comprising a system for non-blocking coordination of a database cluster using the cloud-native Postgres Operator as described above.

[0015] By adopting the above technical solution, the present invention has at least the following advantages: The present invention provides a non-blocking coordination database cluster system for a cloud-native Postgres Operator, which automatically performs fault retries and timeout retries when the Executor executes a coordination command. When the set maximum number of retries is reached, the Executor automatically exits the coordination process. This function significantly improves the stability of the system, avoids coordination interruptions caused by Executor downtime or long-term unresponsiveness, and ensures that faulty nodes can be restored in a timely manner;

[0016] Moreover, in a high-load environment, by reasonably occupying and releasing the resources allocated by the runtime system, the system resource allocation can be made more reasonable, and the working efficiency of the cluster system can be greatly improved, avoiding the frequent coordination anomalies of the Operator caused by the high-pressure environment, and then causing the cluster coordination avalanche, which greatly optimizes the cluster's ability to cope with high-load pressure and reduces the waste of system resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1A system execution logic diagram of a non-blocking coordinated database cluster of a cloud-native Postgres Operator according to an embodiment of the present invention; Figure 2 FIG. 4 is a schematic diagram of a tuning process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention is described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0019] It should be understood that the terms "comprises", "including", "having", "includes" and / or "comprising", when used in this specification, indicate the presence of the stated features, wholes, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, parts and / or combinations thereof. In addition, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire listed features rather than modifying the individual elements in the list. In addition, when describing embodiments of the present application, "may" is used to mean "one or more embodiments of the present application". And, the term "exemplary" is intended to refer to an example or illustration.

[0020] As used herein, the terms "substantially," "approximately," and similar terms are used as terms of approximation, not degree, and are intended to account for the inherent variations in measurements or calculations that would be recognized by those of ordinary skill in the art.

[0021] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined in this article.

[0022] The first embodiment of the present invention is a system for non-blocking coordination of a database cluster using a cloud-native Postgres Operator. Figure 1 As shown, including: The weighted connection pool management module is used for the socket connection of the Executor executor. The connection weight of the socket connection is different according to the weight of the command executed by the Executor executor. When the Executor coordinates the coordination, the priority is determined according to the weight of the socket connection. The system resource weighted timeout retry module is used to retry when Executor coordination fails.

[0023] In this embodiment, the system further includes: a system resource monitoring module, which is used to monitor and cache system resources.

[0024] In this embodiment, the system further includes: an alarm module for logging and alarming errors when Executor coordination fails.

[0025] In this embodiment, the system resource weighted timeout retry module is further used to: When Executor coordination fails, the coordination timeout is positively correlated with the coordination weight; and negatively correlated with the idleness of system resources. In response to the coordination failure, the coordination timeout increases until the configured timeout upper limit is triggered or the configured maximum number of retries is reached, at which time the Executor coordination retries are stopped.

[0026] In this embodiment, in the weighted connection pool management module, multiple upper limits of the number of socket connections are configured for each level of weight, so as to respond to coordination requests and execute coordination commands in parallel.

[0027] In this embodiment, the corresponding alarm mode is determined according to the command weight.

[0028] In this embodiment, when the alarm mode is to record a log, the alarm system increases the recording of the log to a corresponding set weight level.

[0029] The system provided in this embodiment will be described in detail below in combination with practical applications.

[0030] The system resource monitoring system is used to monitor system resources, such as CPU, memory, disk IO, network status, etc. The resource status of system resources is cached, and the cache time can be set through configuration and used in a complete coordination process through context.

[0031] The weighted connection pool management system is used for the socket connection of the Executor executor. The commands executed by different Executor executors have different weights, so the corresponding connection weights are different. The connection pool can retain different numbers of idle socket connections according to the usage of environmental resources, reduce the consumption of system resources when creating socket connections, and flexibly maintain the socket connection in a reasonable range. When the Executor is coordinating, the coordination with high weight can be completed faster, and the coordination process with low weight has a lower execution priority.

[0032] The system resource weighted timeout retry system is used for the retry process when the Executor coordination fails. In the system, after the Executor fails, the context Context is used to control the timeout exit of the coordination process. In this module, the Context timeout is positively correlated with the Executor coordination weight. The higher the weight, the longer the timeout. It is negatively correlated with the idleness of system resources. The fewer the idle resources in the system, the longer the timeout. As the number of failures increases, the timeout increases until the upper limit of the timeout or the number of retries is triggered. The upper limit value is obtained from the configuration file to prevent the Executor from blocking subsequent programs.

[0033] The alarm system is used for logging when Executor coordination fails and for alarming critical errors. First, unlike general alarm systems, when system resources are limited, the alarm system's log recording will be automatically upgraded to the corresponding setting level. This restriction level can be configured through the configuration file, and only the level below the test environment can be configured as the debug level. Secondly, when system resources are tight, the error information of Executor coordination failure will be automatically determined based on the weight of its execution command to determine whether it needs to be pushed in real time or only recorded without pushing, thereby reducing request pressure, reducing system resources occupied by the alarm system, further improving the utilization rate of system resources in other work, and reducing system load.

[0034] Exemplarily, this embodiment adds the following parameters related to system resource monitoring, connection pool system, timeout retry system, and alarm system to the Postgres Operator configuration file: 1. env: the environment of the system, dev is the development environment, staging is the test environment, and prod is the production environment 2. cache_interval: system resource cache time, default is 10 seconds 3. max_resource_threshold: The maximum system resource threshold. When CPU, Memory, and IO exceed the limit threshold, resource contraction intervention is performed on the connection pool, timeout retry, and alarm system. The default value is 0.8 4. socket_num_1: The upper limit of the number of sockets with the highest weight, the default value is 10 5. socket_num_2: The upper limit of the number of general weighted sockets, the default value is 7 6. socket_num_3: The upper limit of the number of sockets with the lowest weight, the default value is 5 7. socket_limit: The minimum number of sockets at any level in the thread pool, the default is 2, the range is 1-5 8. timeout_duration: The basic timeout period for Executor to execute the coordination command. The default value is 2 seconds.

[0035] 9. timeout_interval: The timeout increment for Exector to execute the coordination command. The default value is 2 seconds.

[0036] 10. retry_limit: The number of retries after timeout or failure when executing the coordination command. The default value is 10.

[0037] 11. retry_wight_multiplier_1: The highest weight, retry time multiplier, range is 1.5-2, default is 1.5 12. retry_wight_multiplier_2: general weight, retry time multiplier, range value is 1-1.5, default is 1 13. retry_wight_multiplier_3: minimum weight, retry time multiplier, range value is 0.5-1, default is 1 14. timeout_final: final timeout limit, minimum 10 seconds, maximum 180 seconds, default 60 seconds 15. log_level: log record level, the default is info, optional are debug, info, error, critical 16. insufficient_resource_log_level: The default value is error, and the optional values ​​are info, error, critical During the execution of the coordination command by the Operator's Executor, the socket connection of the corresponding weight is first obtained through the connection pool, and the command is executed through the connection. In case of failure, the timeout retry module will set the timeout period of this coordination process. The timeout period is (timeout_interval+timeout_duration*retry_count)*retry_wight_multiplier_N. During the execution of the Executor, if the command fails to respond after the timeout period, the coordination command will be automatically terminated without waiting. If the execution of a command with a higher weight fails, an alarm will be issued, and if the execution of a command with a lower weight fails, a log will be recorded.

[0038] refer to Figure 2 A detailed process of using the system provided in this embodiment to perform coordination is as follows: S1, Operator sends a coordination request; S2, the connection pool controls the number of connections based on system resources; S3, Executor receives resource usage information from resource monitoring; S4, after the Executor accepts the request, it applies to the connection pool for the number of connections with the corresponding weight according to the resource usage; S5, the coordination process connects concurrently, executes the coordination command in parallel, and returns the connection to the connection pool after success; S6: If the execution fails or times out, a recursive callback is performed and the process returns to step 3 until the retry limit is exceeded.

[0039] The second embodiment of the present invention is an electronic device, which includes a system for non-blocking coordination of database clusters of the cloud-native Postgres Operator as described in the first embodiment, which can be understood as a specific physical device including the above system and can also be used to implement the coordination process of the system.

[0040] Compared with the prior art, the advantages of the present invention include at least: (1) This invention proposes an Executor crash recovery method specifically for Postgres Operator. The core of the method is to introduce a timeout retry mechanism when executing the coordination command. When the Executor becomes unresponsive, the system can automatically retry through the set timeout strategy, avoiding the problem of long-term blocking or inability to continue the coordination process. This method ensures that the cluster coordination task can continue to execute under abnormal circumstances, significantly improving the reliability and recovery capabilities of the cluster;

[0041] (2) The present invention adjusts the waste of system resources under high load conditions by reasonably scaling the connection pool, reducing resource usage of the alarm system under high load conditions, and giving priority to high-weight commands by the Executor to avoid low-weight commands from occupying too many resources. This gives the Postgres Operator the stability to complete cluster coordination when coping with system resource pressure. (3) The present invention provides a variety of customizable functions by flexibly setting the configuration parameters of the Postgres Operator. For example, users can set the upper limit of the number of retries, adjust the basic interval time of the command execution timeout, and control the interval between each retry through the incremental mechanism (i.e., the timeout step). This parameterized configuration method makes cluster management more flexible and automated, effectively reducing the intervention and pressure of operation and maintenance personnel. The Operator can also further optimize the situation in which high-weight coordination commands are prone to failure due to insufficient resources in high-load systems by giving priority to the execution and retries of commands with different weights, greatly improving the Operator's coordination capabilities;

[0042] (4) During the coordination process, if a command times out or fails to exit, the system will automatically record a detailed log of the entire process to ensure that all operations are recorded. These log records not only help operation and maintenance personnel to trace the entire coordination process during subsequent analysis and troubleshooting, but can also serve as a reference when entering the next retry to help optimize the system's coordination mechanism. In the alarm system, the use of resources by the alarm system is reasonably allocated based on the resource occupancy situation, reducing the problem of insufficient system resources caused by the high frequency of resource occupancy by the alarm system under high load. Through the above mechanism, the transparency and controllability of the coordination are greatly improved. Under high load conditions, the stability of the cluster can also be effectively guaranteed, further reducing the risk of cluster unavailability due to unexpected failures.

[0043] In summary, through this invention, the operation and maintenance personnel only need to perform simple configuration before starting the Postgres Operator, so that the Executor can automatically perform fault retries and timeout retries when executing the coordination command. When the set maximum number of retries is reached, the Executor will automatically exit the coordination process. This function significantly improves the stability of the system, avoids the coordination interruption caused by the Executor downtime or long-term unresponsiveness, and ensures that the faulty node can be restored in time; In addition, under high-load environments, by properly occupying and releasing system resources during runtime, the system resource allocation becomes more reasonable, which greatly improves the efficiency of the cluster system and avoids frequent tuning anomalies in Operators caused by high-pressure environments, which in turn causes a cluster tuning avalanche. This greatly optimizes the cluster's ability to cope with high-load pressure and reduces system resource waste. The present invention allows operation and maintenance personnel to monitor the execution of commands in the cluster coordination process in real time by regularly viewing log files. This log monitoring function enables operation and maintenance personnel to quickly discover potential problems and take timely measures, thereby further reducing the risk of the cluster entering an unhealthy state or being unable to provide external services;

[0044] Furthermore, since the present invention is applied under high-load conditions, the information recorded by the log system is concise and critical enough, and operation and maintenance personnel can quickly locate problems caused by the execution of higher-weight commands without having to look for errors in large logs. This greatly improves the efficiency of problem troubleshooting or human intervention. The high-weight alarm strategy can also enable operation and maintenance personnel to ignore the troubles of low-authority alarms, focus on solving problems caused by failures to execute high-weight commands, and promptly eliminate problems caused by coordination, effectively improving operation and maintenance efficiency and system health.

[0045] Through the description of the specific implementation methods, a deeper and more specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose should be obtained. However, the accompanying drawings are only for reference and illustration purposes and are not intended to limit the present invention.

Claims

1. A non-blocking coordination database cluster system for cloud-native Postgres Operator, characterized in that: include: The weighted connection pool management module is used for the socket connection of the Executor executor. The connection weight of the socket connection is different according to the weight of the command executed by the Executor executor. When the Executor coordinates the coordination, the priority is determined according to the weight of the socket connection. The system resource weighted timeout retry module is used to retry when Executor coordination fails.

2. The system for non-blocking coordination of database clusters using the cloud-native Postgres Operator according to claim 1, characterized in that: The system further comprises: The system resource monitoring module is used to monitor and cache system resources.

3. The system for non-blocking coordination of database clusters of cloud-native Postgres Operator according to claim 1, characterized in that: The system further comprises: The alarm module is used for logging and error alarms when Executor coordination fails.

4. The system for non-blocking coordination of database clusters of cloud-native Postgres Operator according to claim 1, characterized in that: The system resource weighted timeout retry module is further used to: When Executor coordination fails, the coordination timeout is positively correlated with the coordination weight; and negatively correlated with the idleness of system resources. In response to the coordination failure, the coordination timeout increases until the configured timeout upper limit is triggered or the configured maximum number of retries is reached, at which time the Executor coordination retries are stopped.

5. The system for non-blocking coordination of database clusters by the cloud-native Postgres Operator according to claim 1, characterized in that: In the weighted connection pool management module, multiple upper limits of the number of socket connections are configured for each level of weight, which are used to respond to coordination requests and execute coordination commands in parallel.

6. The system for non-blocking coordination of database clusters of the cloud-native Postgres Operator according to claim 1, characterized in that: According to the command weight, a corresponding alarm mode is determined.

7. The system for non-blocking coordination of database clusters of cloud-native Postgres Operator according to claim 6, characterized in that: When the alarm mode is to record a log, the alarm system increases the recording of the log to a corresponding set weight level.

8. An electronic device, characterized in that: The electronic device includes a system for non-blocking coordination of a database cluster of the cloud-native Postgres Operator according to any one of claims 1 to 7.