Database cluster inspection control method and system and server
Through the database cluster inspection and control method of multimodal data fusion analysis and dynamic strategy selection mechanism, the problem of long failure recovery time in traditional manual inspection mode is solved, and rapid fault diagnosis and automatic repair are achieved.
Patent Information
- Application Number
- CN202510828784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional manual inspection model is difficult to cope with the multi-dimensional health assessment, performance bottleneck positioning and root cause analysis of database clusters, resulting in the failure recovery time being high for a long time.
The multimodal data fusion analysis method is used to check the health status of the database cluster in real time, and automatically repair it through the dynamic policy selection mechanism, and the handling progress is synchronized through the multi-channel alarm platform.
It significantly improves the inspection and control effect of the database cluster, realizes rapid fault diagnosis and automatic repair, and reduces the fault recovery time.
Smart Images

Figure CN120336342A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database operation and maintenance, and in particular to a method, system and server for database cluster inspection and control. Background Art
[0002] As a core infrastructure, database clusters have been deeply integrated into key business scenarios such as finance, telecommunications, energy, and transportation. Database clusters store and process data through the cooperation of multiple nodes. While improving the system throughput, they also pose higher requirements for dynamic load balancing, data consistency guarantee, and disaster tolerance and fault tolerance capabilities. However, with the expansion of the cluster scale, the maintenance of the cluster becomes more complex, and the traditional manual inspection mode is difficult to meet complex operation and maintenance requirements such as multi-dimensional health assessment, performance bottleneck positioning, and root cause analysis of faults.
[0003] The manual inspection mechanism commonly adopted in the current industry has significant limitations. Its essence is a passive response mode, relying on operation and maintenance personnel to regularly execute preset scripts or manual inspections, and it is difficult to cope with non-linear fault causes such as sudden traffic shocks and node hardware aging, resulting in a long average fault recovery time. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, system and server for database cluster inspection and control. This method can perform real-time health checks, performance analysis, and fault diagnosis on the database cluster through multi-modal data fusion analysis, and can automatically repair based on a dynamic policy selection mechanism when problems are found, and can synchronously push the handling progress through a multi-channel alarm platform, significantly improving the inspection and control effect of the database cluster.
[0005] In a first aspect, an embodiment of the present invention provides a method for database cluster inspection and control, the method comprising: Obtain a first database and a second database included in the database cluster, and obtain the startup status of the first database and the second database; If the first database corresponding to the startup status is not started, obtain the first node status corresponding to the first database and the second node status corresponding to the second database; Determine the first node type corresponding to the first database according to the first node status, and determine the second node type corresponding to the second database according to the second node status; If the first node type and the second node type satisfy a preset first node relationship, obtain the first node update time corresponding to the first node status and the second node update time corresponding to the second node status; Determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
[0006] Optionally, before the steps of obtaining the first database and the second database included in the database cluster and obtaining the startup status of the first database and the second database, the method further includes: Obtaining network status parameters corresponding to the first database, and determining the network connection status of the first database based on the network status parameters; If the network connection status is an abnormal state, then determine whether the first database is started based on the startup status; If the first database is not started, then obtain a first status value corresponding to the first node status, and determine whether the first status value is the same as a preset first status threshold; If they are the same, then control the first database to complete shutdown.
[0007] Optionally, after the steps of obtaining the first database and the second database included in the database cluster and obtaining the startup status of the first database and the second database, the method further includes: If the first database corresponding to the startup status is in the startup state, then determine whether the first database is in the fast recovery state based on the startup status; If the first database is in the fast recovery state, then obtain the duration corresponding to the fast recovery state; If the duration is less than a preset timeout threshold, then control the first database to create a log file, and control the log file to record the current timestamp of the first database.
[0008] Optionally, after determining whether the first database is in the fast recovery state based on the startup status, it further includes: If the first database is not in the fast recovery state, then determine whether the first database contains a log file; If it contains, then delete the log file.
[0009] Optionally, after obtaining the duration corresponding to the fast recovery state, it further includes: If the duration is not less than the timeout threshold, then obtain a second status value corresponding to the second database based on the second node status, and determine whether the second status value is the same as a preset second status threshold; If they are the same, then set the first database to the standby mode, and determine whether the first database contains a log file; If it contains, then delete the log file.
[0010] Optionally, before obtaining the first node status corresponding to the first database and the second node status corresponding to the second database, the method further includes: Obtaining deployment parameters corresponding to the first database, and determining whether the first database is in the single-machine mode based on the deployment parameters; If so, control the first database to complete startup.
[0011] Optionally, after determining whether the first database is in the single-machine mode based on the deployment parameters, it further includes: If not, determine whether the second database has been started based on the startup status; If so, control the first database to complete startup.
[0012] Optionally, after determining the second node type corresponding to the second database according to the second node status, the method further includes: If the first node type and the second node type satisfy the preset second node relationship, control the first database to complete startup.
[0013] In a second aspect, the present invention provides a database cluster inspection and control system, which includes: A startup status probe module, configured to obtain the first database and the second database included in the database cluster, and obtain the startup status of the first database and the second database; A node status probe module, configured to obtain the first node status corresponding to the first database and the second node status corresponding to the second database if the first database corresponding to the startup status has not been started; A node type analysis module, configured to determine the first node type corresponding to the first database according to the first node status, and determine the second node type corresponding to the second database according to the second node status; An update time analysis module, configured to obtain the first node update time corresponding to the first node status and the second node update time corresponding to the second node status if the first node type and the second node type satisfy the preset first node relationship; A policy execution module, configured to determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
[0014] In a third aspect, an embodiment of the present invention further provides a server, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the database cluster inspection and control method provided in the first aspect.
[0015] In a fourth aspect, an embodiment of the present invention further provides a storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the steps of the database cluster inspection and control method provided in the first aspect.
[0016] A database cluster inspection and control method, system and server provided by an embodiment of the present invention, in the process of inspecting and controlling a database cluster, the method first obtains a first database and a second database included in the database cluster, and obtains the startup status of the first database and the second database; if the first database corresponding to the startup status is not started, then obtain the first node status corresponding to the first database and the second node status corresponding to the second database; then determine the first node type corresponding to the first database according to the first node status, and determine the second node type corresponding to the second database according to the second node status; if the first node type and the second node type satisfy a preset first node relationship, then obtain the first node update time corresponding to the first node status and the second node update time corresponding to the second node status; finally, determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status. This method can perform health checks, performance analysis, and fault diagnosis on the database cluster in real time through multi-modal data fusion analysis, and can automatically repair based on a dynamic policy selection mechanism when problems are found, and can synchronously push the handling progress through a multi-channel alarm platform, significantly improving the inspection and control effect of the database cluster.
[0017] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0018] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a database cluster inspection and control method provided by an embodiment of the present invention; Figure 2 It is a flowchart before step S101 in a database cluster inspection and control method provided by an embodiment of the present invention; Figure 3 It is a flowchart after step S101 of a database cluster inspection and control method provided by an embodiment of the present invention; Figure 4 In step S301 of a database cluster inspection and control method provided by an embodiment of the present invention, it is a flowchart after determining whether the first database is in a fast recovery state based on the startup state; Figure 5 In step S302 of a database cluster inspection and control method provided by an embodiment of the present invention, it is a flowchart after obtaining the duration corresponding to the fast recovery state; Figure 6 In step S102 of a database cluster inspection and control method provided by an embodiment of the present invention, it is a flowchart before obtaining the first node state corresponding to the first database and the second node state corresponding to the second database; Figure 7 It is a flowchart of another database cluster inspection and control method provided by an embodiment of the present invention; Figure 8 It is a simplified flowchart of a database cluster inspection and control method provided by an embodiment of the present invention; Figure 9 It is a schematic structural diagram of a database cluster inspection and control system provided by an embodiment of the present invention; Figure 10 It is a schematic structural diagram of another database cluster inspection and control system provided by an embodiment of the present invention; Figure 11 It is a schematic structural diagram of a server provided by an embodiment of the present invention; Figure 12 It is a system architecture diagram of a server provided by an embodiment of the present invention.
[0021] Icon: 910 - Startup state probe module; 920 - Node state probe module; 930 - Node type analysis module; 940 - Update time analysis module; 950 - Policy execution module; 1010 - Distributed probe module; 1020 - Intelligent analysis engine; 1030 - Policy executor; 101 - Processor; 102 - Memory; 103 - Bus; 104 - Communication interface. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] As a core infrastructure, database clusters have been deeply integrated into key business scenarios such as finance, telecommunications, energy, and transportation. Database clusters store and process data through the collaboration of multiple nodes. While improving the system throughput, they also pose higher requirements for dynamic load balancing, data consistency guarantee, and disaster tolerance and fault tolerance capabilities. However, with the expansion of the cluster scale, the maintenance of the cluster becomes more complex, and the traditional manual inspection mode is difficult to meet complex operation and maintenance requirements such as multi-dimensional health assessment, performance bottleneck location, and root cause analysis of faults.
[0024] The manual inspection mechanism commonly adopted in the current industry has significant limitations. Its essence is a passive response mode, relying on operation and maintenance personnel to regularly execute preset scripts or conduct manual inspections, which is difficult to cope with non-linear fault causes such as sudden traffic shocks and node hardware aging. According to statistics from relevant institutions, more than 65% of database downtime events are due to potential fault hazards not being discovered in a timely manner, and there are obvious cognitive limitations in the manual inspection in aspects such as abnormal index threshold judgment and correlation analysis of associated events, resulting in a long-term high average fault recovery time. Based on this, the present invention provides a database cluster inspection control method, system, and server. This method can perform real-time health checks, performance analysis, and fault diagnosis on the database cluster through multi-modal data fusion analysis, and can automatically repair based on a dynamic policy selection mechanism when problems are found, and can synchronously push the handling progress through a multi-channel alarm platform, significantly improving the inspection control effect of the database cluster.
[0025] To facilitate the understanding of this embodiment, first, a database cluster inspection control method disclosed in the embodiments of the present invention will be introduced in detail. As Figure 1 shown, this method includes: Step S101, obtain the first database and the second database included in the database cluster, and obtain the startup status of the first database and the second database.
[0026] The database cluster usually includes multiple databases, at least the first database and the second database. The startup status of the first database and the second database is used to represent the startup situation of the first database and the second database, and may include: started, not started, starting, startup failed, etc. The acquisition of the startup status can be obtained by reading the startup flag of the relevant process log, or the execution situation of the above databases can be obtained through thread processing.
[0027] Step S102, if the first database corresponding to the startup status is not started, then obtain the first node status corresponding to the first database and the second node status corresponding to the second database.
[0028] The node status is obtained by reading the data on a specific page in the relevant database. For example, when using the dual - page verification method to obtain the node status of the first database and the second database, the node status of page 10 and page 11 can be obtained according to the specific requirements of these two databases. It should be noted that the above - mentioned page 10 and page 11 are only for reference, and the page numbers are not limited to the above in the actual scenario.
[0029] Step S103: Determine the first node type corresponding to the first database according to the first node status, and determine the second node type corresponding to the second database according to the second node status.
[0030] In the process of obtaining the node type, the data storage format in the database can be combined. By obtaining the bytes corresponding to the node type in the database and parsing them to obtain the parsed value, the corresponding node type can be determined. For example, for the first database, by obtaining the specific position in a specific row on page 10, the original field parsed is: 02 00 00 00, and the corresponding node type is obtained as 2 after decimal conversion. The rule for obtaining the node type on page 11 is similar to that on page 10. In the specific implementation process, the time points corresponding to these two node types can be judged, and the node type corresponding to the newer time node is used as the node type of the first database. The acquisition of the second node type is the same as that of the first node type and will not be elaborated here.
[0031] Step S104: If the first node type and the second node type satisfy the preset first - node relationship, obtain the first node update time corresponding to the first node status and the second node update time corresponding to the second node status.
[0032] If the first node type is 2, the second node type is 2, and the first - node relationship is that they are the same. At this time, the first node type and the second node type satisfy the first - node relationship. At this time, the node update times corresponding to the first node status and the second node status are respectively obtained through the underlying data acquisition instructions. The process of obtaining the node update time is also based on the specific position in a specific row on a specific page of the database. For example, for the first database, the first node update time can be obtained by obtaining the specific position in a specific row on page 2, and the original field parsed is: 00 3c 2b 66 01, and its corresponding field name is ll_lasttime. After converting it from hexadecimal to decimal, the parsed value is 23472956, which is used as the first node update time. The acquisition of the second node update time is the same as that of the first node update time and will not be elaborated here.
[0033] Step S105: Determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
[0034] By comparing the update time of the first node and the update time of the second node, it is determined which node has a newer logical log, and then when the logical log in the first node corresponding to the first database is the newest, the first database is controlled to complete startup.
[0035] Optionally, before step S101 of obtaining the first database and the second database included in the database cluster and obtaining the startup status of the first database and the second database, as Figure 2 shown, the method further includes: Step S201, obtaining the network status parameters corresponding to the first database, and determining the network connection status of the first database based on the network status parameters; Step S202, if the network connection status is an abnormal state, then determining whether the first database has been started based on the startup status; Step S203, if the first database has not been started, then obtaining the first status value corresponding to the first node status, and determining whether the first status value is the same as a preset first status threshold; Step S204, if they are the same, then controlling the first database to complete shutdown.
[0036] During the actual execution process, it is necessary to first verify whether the network of the first database is normal, specifically by determining the network connection status through the corresponding network status parameters. If the network connection status is abnormal, then it is determined whether the first database has been started based on the startup status. If it has been started, then the first status value corresponding to the first node status is obtained, and it is determined whether the first status value is the same as the first status threshold. Specifically, the first status threshold is 2, and when the first node status value is also 2, then the first database is controlled to complete shutdown.
[0037] Optionally, after step S101 of obtaining the first database and the second database included in the database cluster and obtaining the startup status of the first database and the second database, as Figure 3 shown, the method further includes Step S301, if the first database corresponding to the startup status is in the startup state, then determining whether the first database is in the fast recovery state based on the startup status; Step S302, if the first database is in the fast recovery state, then obtaining the duration corresponding to the fast recovery state; Step S303, if the duration is less than a preset timeout threshold, then controlling the first database to create a log file, and controlling the log file to record the current timestamp of the first database.
[0038] In an actual scenario, it is also necessary to verify the startup status of the first database. When the first database is in the startup status, it is determined whether the first database is in the fast recovery state, that is, Fast Recovery, based on the startup status. If the first database is in the Fast Recovery state, the duration of the Fast Recovery state is obtained to determine whether the Fast Recovery of the first database times out. If the duration is less than the preset timeout threshold, indicating that there is no timeout, a log file is created and the current timestamp is recorded.
[0039] Optionally, after determining whether the first database is in the fast recovery state based on the startup status, as Figure 4 shown, it further includes: Step S401, if the first database is not in the fast recovery state, determine whether the first database contains a log file; Step S402, if it contains, delete the log file.
[0040] If the first database is not in the Fast Recovery state, it is directly determined whether the first database contains a log file, and if it exists, it is directly deleted.
[0041] Optionally, after obtaining the duration corresponding to the fast recovery state, as Figure 5 shown, it further includes: Step S501, if the duration is not less than the timeout threshold, obtain the second status value corresponding to the second database based on the second node status, and determine whether the second status value is the same as the preset second status threshold; Step S502, if they are the same, set the first database to the standby mode, and determine whether the first database contains a log file; Step S503, if it contains, delete the log file.
[0042] If the duration is not less than the preset timeout threshold, indicating that the Fast Recovery state has timed out, at this time, the second status value of the second database is verified. By comparing the second status value with the preset second status threshold, if they are the same, that is, the second status value is 5 and the second status threshold is also 5, the server corresponding to the first database is set to the standby server mode, and the contained log file is deleted.
[0043] Optionally, before obtaining the first node status corresponding to the first database and the second node status corresponding to the second database, as Figure 6 shown, the method further includes: Step S601, obtain the deployment parameters corresponding to the first database, and determine whether the first database is in the single - machine mode based on the deployment parameters; Step S602, if so, control the first database to complete startup.
[0044] Specifically, the deployment parameters can be judged by the node type. If the node type is 0, the first database is in the single-machine mode. If double-page verification is involved, obtain which time node is newer in these two pages through the deployment parameters, and then obtain the node type corresponding to the data page of the newer time node. If the node type is 0, the first database is in the single-machine mode. The first database in the single-machine mode is in an offline state, and at this time, the first database can be directly started.
[0045] Optionally, after judging whether the first database is in the single-machine mode based on the deployment parameters, it includes: if it is not in the single-machine mode, judge whether the second database has been started based on the startup state; if the second database is in the startup state, directly control the first database to complete startup.
[0046] Optionally, after determining the second node type corresponding to the second database according to the second node state, the method further includes: if the first node type and the second node type satisfy the preset second node relationship, control the first database to complete startup. If the first node type is 2 and the second node type corresponds to 3, and the two satisfy the second node relationship, the first database is directly started at this time.
[0047] Such as Figure 7 shown in the flowchart of another method for database cluster inspection and control. The first database corresponding to the local database in the figure, and the other database corresponds to the second database. From Figure 7 the flowchart, it can be seen that this method includes three cases, specifically as follows: Case 1, the local network is abnormal. At this time, it is necessary to verify whether the local network is normal. If the local network is abnormal, verify whether the local database has been started. If the local database has been started, verify whether the node type of the local database is 2. If it is 2, execute onmode -ky to close the local database, otherwise it is automatically processed by the database configuration failover.
[0048] The acquisition of the node and timestamp of the local database can be implemented based on the binary direct reading technology. The underlying data acquisition instructions involved are: # Obtain the 10th page of the ROOT page dd if=rootdbs1 skip=10 bs=2k count=1 | hexdump -C # Obtain the 11th page of the ROOT page dd if=rootdbs1 skip=11 bs=2k count=1 | hexdump -C The skip parameter therein precisely jumps to the page position, and hexdump -C displays the normalized hexadecimal + ASCII view; bs=2k supports customizable page size (default 2KB).
[0049] After obtaining the data information of page 10, the original data is as follows:
[0050] When obtaining the node type in the above data, the target behavior is the data at the beginning of line 00000050. The field position starts from the 47th position in line 00000050 and ends at the 58th position in line 00000050, and the result obtained is 02 00 0000. After mapping it with a specific data structure, the following result is obtained:
[0051] When obtaining the timestamp, the target behavior is the data at the beginning of line 000007f0. The field position starts from the 47th position in line 000007f0 and ends at the 58th position in line 000007f0, and the result obtained is 00 ad 29 66 01. After converting the hexadecimal to decimal by reading from right to left, the parsed value is 23472557, and the specific result is as follows:
[0052] Subsequently, continue to obtain the node status and update time, and dynamically assign the timestamp. When a new checkpoint is written to the database, the timestamp of the latest active page will be updated to a larger value, and then determine which time point of page 10 and page 11 is updated, and obtain the node type of the updated time node. If the local node type is 2, then execute onmode -ky, otherwise it is automatically processed by the database configuration failover.
[0053] Case 2, the local network is normal and the current database is started. The prerequisite is that the local network is connected and the database service has been started. Then execute the status check process. First, check the current status type of the database; if it is in the Fast Recovery state, read the record time in the log file and calculate the difference between the record time and the current time (unit: seconds). Then verify the local database status type. If the local database status is Fast Recovery, then determine whether the duration of the Fast Recovery state is greater than or equal to FastRecoveryTimeout. Specifically query the log file to find the difference between the time recorded in the file content and the current time as the duration.
[0054] If the duration is greater than or equal to FastRecoveryTimeout, verify whether the status of another database (Database2) is 5. If the status of the other database is 5, set the database server (Database1) to standby server mode and delete the log file.
[0055] If the duration is less than FastRecoveryTimeout, create a log file and record the current timestamp. If the status of the local database is not Fast Recovery, check whether the local log file exists. If the log file exists, delete the log file.
[0056] Case 3: The local network is normal and the current database is not started. The prerequisite at this time is that the local network is connected and the database service is not started. First, verify whether the current database is in single - machine mode; if it is in single - machine mode, start the database.
[0057] Verify whether the current database is in single - machine mode in the following way. First, determine which time point on page 10 and page 11 is updated, and obtain the node type of the updated time node. If the node type is 0, the current database is in single - machine mode; if the node type is 0, the current database is in single - machine mode.
[0058] If the current database is not in single - machine mode, verify whether database Database2 is in the started state; if database Database2 is in the started state, directly start the local database. If database Database2 is in the unstarted state, continue to use the double - page verification algorithm to obtain the node status of the local database.
[0059] Obtain the node status of page 10 and page 11 of the local database Database1, and in addition, obtain the node status of page 10 and page 11 of another database Database2. If the local node type is 2 and the other node type is 3, start the local database. If the local node type is 2 and the other node type is 2, use the underlying data acquisition instruction to obtain the data of page 2 and page 3 of the ROOT page.
[0060]
[0061] Taking the data on page 2 above as an example, the target behavior corresponding to the timestamp is the data starting from line 000007f0. The field position starts from the 47th position of line 000007f0 and ends at the 58th position of line 000007f0. Specifically as follows:
[0062] For the case of whether there is a continuation page, the corresponding target behavior is the data at the beginning of line 00000010, which is as follows:
[0063] If HeadNext>0 and HeadPrev>0, it means there is a continuation page, denoted by isContinuesPage.
[0064] The information of the first slot, with the starting position at the 24th byte and the length of 48 bytes, is as follows:
[0065] Starting from the second slot, the length of each slot is 32 bytes. The starting position of the second slot is the 72nd byte, which is as follows:
[0066] The core algorithms involved in the above process include the following: Data preprocessing: Split the string by line: Split HexData into a string array lines by the newline character \n; Filter valid data lines, collect data starting from a specific line (00000040); If IsContinuesPage is false, skip all lines before this line; When encountering a line starting with *, stop data collection.
[0067] The mathematical expression of the above process is as follows: Input data set (The set of original lines split by newline characters, that is, each line of information in the data of page 2); Predicate function: Starting with "*"; Starting with " "; .
[0068] Transformation function: (Take the substring of the 10th - 58th characters); The algorithm process can be described as:
[0069] In the above formula, represents the end line starting with * in the input data set ; is the task start condition; is the task interruption condition.
[0070] The key mathematical features are as follows: Set filtering: ; All elements are after the first element that satisfies ; Does not contain any element that satisfies ; Interval truncation: Apply a linear transformation to each :
[0071] Boundary implicit constraint: Assume , otherwise an array out-of-bounds exception is triggered Group processing: The total length of each group of data is 144 characters. Extract the interval [24, 124) from each group to obtain a hexadecimal string S of length 100 characters. Starting from , the first group of data is , and the second group of data is , that is, the end of each group is the start of the next group.
[0072] Check the key bits as follows: Extract the first 4 characters (corresponding to 2 bytes) of S and convert them to a decimal number V1; Convert V1 to a binary string F; If F ends with 10 or 11, perform subsequent parsing operations.
[0073] Numeric parsing is as follows: Parse a 4-byte integer (offset 0x0C): Extract from the extracted S, calculate the character position: starting from 0x0C * 2 = 24 characters, extract 8 characters (4 bytes) and convert them to a decimal number V2, the unique identifier (may be null); Parse a 4-byte integer (offset 0x1C): Extract from the extracted S, calculate the character position: starting from 0x1C * 2 = 36 characters, extract 8 characters (4 bytes) and convert them to a decimal number V3, the usage status flag.
[0074] The time complexity of this algorithm is O(n), and the space complexity is O(1) (only linear scanning of the input data is required during processing). The above algorithm is the result of the inventor's creative work and is named HEXTILE (Hex-Triggered IntervalLayer Extractor).
[0075] For data acquisition instructions, it is as follows:
[0076] Verify the sizes of ll_lasttime for Page 2 and Page 3 using the double-page verification algorithm, and obtain the time update node data V2 and V3. If HeadNext>0 && HeadPrev>0, there are consecutive pages.
[0077] The traversal query logic for consecutive pages is as follows: 1. Loop initialization.
[0078] Use HeadNext as the starting index, and the upper limit of the loop count is HeadPrev times, that is, the traversal index range: .
[0079] 2. Obtain consecutive page data.
[0080] Perform the following operations for each iteration:
[0081] 3. Extract key fields.
[0082] Extract two key fields from the consecutive page information object: V2: Usage status flag (may be null); V3: Usage status flag.
[0083] 4. Termination condition detection. When it is detected that the V2 field is not null: Obtain V2 and V3 as the values in the consecutive page.
[0084] 5. Mathematical representation ; where: .
[0085] Continue to use the double-page verification algorithm to verify, obtain the update time ll_lasttime, obtain local V2, V3; use the above method to continue to obtain the update time ll_lasttime of another node, and obtain other V2, V3.
[0086] Compare the logical log status of the local node and the peer node to determine which node's log is updated. If localV2>other V2, the local node's log is the latest; if local V2 = other V2 && local V3 >= other V3, the local node's log is the latest; otherwise, the other node's log is the latest.
[0087] If the local node log is up-to-date, execute the command to start the local database. After the execution on its Database1 database is completed, start a thread to change the current execution flag isRun to false and record the current time endTime. Verify that currentTime - startTime > T && isRun == false, then continue with the automatic database maintenance detection; otherwise, wait briefly and continue the next check.
[0088] Figure 7 The process in Figure 8 is relatively complex. Refer to the simplified flowchart of a database cluster inspection and control method shown in
[0089] As can be seen from the database cluster inspection and control method mentioned in the above embodiments, this method can perform real-time health checks, performance analysis, and fault diagnosis on the database cluster through multi-modal data fusion analysis, and can automatically repair based on the dynamic policy selection mechanism when problems are found, and can synchronously push the handling progress through the multi-channel alarm platform, significantly improving the inspection and control effect of the database cluster.
[0090] Corresponding to the database cluster inspection and control method provided in the foregoing embodiments, the embodiments of the present invention provide a database cluster inspection and control system, as Figure 9 shown. The system includes: A startup status probe module 910, configured to obtain the first database and the second database included in the database cluster, and obtain the startup status of the first database and the second database; A node status probe module 920, configured to obtain the first node status corresponding to the first database and the second node status corresponding to the second database if the first database corresponding to the startup status has not been started; A node type analysis module 930, configured to determine the first node type corresponding to the first database according to the first node status, and determine the second node type corresponding to the second database according to the second node status; An update time analysis module 940, configured to obtain the first node update time corresponding to the first node status and the second node update time corresponding to the second node status if the first node type and the second node type satisfy a preset first node relationship; A policy execution module 950, configured to determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
[0091] As Figure 10Schematic diagram of the structure of another database cluster inspection and control system shown, including: a distributed probe module 1010, an intelligent analysis engine 1020, and a policy executor 1030. Among them, the distributed probe module 1010 corresponds to a startup status probe module 910 and a node status probe module 920, the intelligent analysis engine 1020 corresponds to a node type analysis module 930 and an update time analysis module 940, and the policy executor 1030 corresponds to a policy execution module 950. The intelligent analysis engine 1020 uses binary direct reading, a double-page hot switch detection algorithm, and the HEXTILE (Hex-Triggered Interval Layer Extractor) algorithm and a hybrid parsing engine to achieve multi-dimensional fault location. The policy executor integrates advanced primary-backup switching, database cluster startup and shutdown, ensuring that the repair instructions automatically switch the execution entity in case of node failure, and the availability reaches 99.99%.
[0092] As can be seen from the database cluster inspection and control system mentioned in the above embodiments, this system can perform real-time health checks, performance analysis, and fault diagnosis on the database cluster through multi-modal data fusion analysis, and can automatically perform repairs based on a dynamic policy selection mechanism when problems are found, and can synchronously push the handling progress through a multi-channel alarm platform, significantly improving the inspection and control effect of the database cluster.
[0093] The database cluster inspection and control system provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing embodiments of the database cluster inspection and control method. For a brief description, for the parts not mentioned in the system embodiments, reference can be made to the corresponding content in the foregoing embodiments of the database cluster inspection and control method.
[0094] This embodiment also provides a server, and the schematic diagram of the structure of this server is as Figure 11 shown. This device includes a processor 101 and a memory 102; among them, the memory 102 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the steps of the above-mentioned database cluster inspection and control method.
[0095] Figure 11 The server shown also includes a bus 103 and a communication interface 104, and the processor 101, the communication interface 104, and the memory 102 are connected through the bus 103.
[0096] Among them, the memory 102 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The bus 103 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0097] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and send the encapsulated IPv4 packet or IPv4 packet to the user terminal through the network interface.
[0098] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method may be completed by the integrated logic circuit in the hardware of the processor 101 or instructions in software form. The above-mentioned processor 101 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure may be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0099] Such as Figure 12The system architecture diagram of the shown server, taking the GBase8s HDR cluster as an example, interacts through a pre-built GBase8s HDR database cluster and a switch / gateway. This method is applied to different databases to achieve second-level exception perception (30-second delay, dynamically configurable); the fault location accuracy rate > 99% (measured data); 99% of common faults can be self-healed without manual intervention.
[0100] An embodiment of the present invention also provides a storage medium on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the database cluster inspection and control method in the foregoing embodiment.
[0101] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, equipment, and methods can be implemented in other ways. The system embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0102] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0104] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0105] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A database cluster inspection and control method, characterized in that The method includes: Obtain a first database and a second database included in a database cluster, and obtain the startup status of the first database and the second database; If the first database corresponding to the startup status is not started, obtain a first node status corresponding to the first database and a second node status corresponding to the second database; Determine a first node type corresponding to the first database according to the first node status, and determine a second node type corresponding to the second database according to the second node status; If the first node type and the second node type satisfy a preset first node relationship, obtain a first node update time corresponding to the first node status and a second node update time corresponding to the second node status; Determine the logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
2. The database cluster inspection and control method according to claim 1, wherein Before the step of obtaining the first database and the second database included in the database cluster, and obtaining the startup status of the first database and the second database, the method further includes: Obtain network status parameters corresponding to the first database, and determine the network connection status of the first database based on the network status parameters; If the network connection status is an abnormal status, determine whether the first database is started based on the startup status; If the first database is not started, obtain a first status value corresponding to the first node status, and determine whether the first status value is the same as a preset first status threshold; If they are the same, control the first database to complete shutdown.
3. The database cluster inspection and control method according to claim 1, wherein After the step of obtaining the first database and the second database included in the database cluster, and obtaining the startup status of the first database and the second database, the method further includes: If the first database corresponding to the startup status is in the startup state, determine whether the first database is in a fast recovery state based on the startup status; If the first database is in the fast recovery state, obtain the duration corresponding to the fast recovery state; If the duration is less than a preset timeout threshold, control the first database to create a log file, and control the log file to record the current timestamp of the first database.
4. The database cluster inspection and control method according to claim 3, characterized in that After determining whether the first database is in the fast recovery state based on the startup status, it further includes: If the first database is not in the fast recovery state, determine whether the log file is included in the first database; If it is included, delete the log file.
5. The database cluster inspection and control method according to claim 3, wherein After obtaining the duration corresponding to the fast recovery state, it further includes: If the duration is not less than the timeout threshold, obtain a second status value corresponding to the second database based on the second node status, and determine whether the second status value is the same as a preset second status threshold; If they are the same, set the first database to standby mode, and determine whether the log file is included in the first database; If it is included, delete the log file.
6. The database cluster inspection and control method according to claim 1, characterized in that Before obtaining the first node status corresponding to the first database and the second node status corresponding to the second database, the method further includes: Obtaining deployment parameters corresponding to the first database, and determining whether the first database is in a single-machine mode based on the deployment parameters; If so, controlling the first database to complete startup.
7. The database cluster inspection and control method according to claim 6, characterized in that, After determining whether the first database is in a single-machine mode based on the deployment parameters, it further includes: If not, determining whether the second database has been started based on the startup status; If so, controlling the first database to complete startup.
8. The database cluster inspection and control method according to claim 1, wherein After determining the second node type corresponding to the second database according to the second node status, the method further includes: If the first node type and the second node type satisfy a preset second node relationship, controlling the first database to complete startup.
9. A database cluster inspection and control system, characterized in that, The system includes: A startup status probe module, configured to obtain a first database and a second database included in a database cluster, and obtain startup statuses of the first database and the second database; A node status probe module, configured to, if the first database corresponding to the startup status has not been started, obtain a first node status corresponding to the first database and a second node status corresponding to the second database; A node type analysis module, configured to determine a first node type corresponding to the first database according to the first node status, and determine a second node type corresponding to the second database according to the second node status; An update time analysis module, configured to, if the first node type and the second node type satisfy a preset first node relationship, obtain a first node update time corresponding to the first node status and a second node update time corresponding to the second node status; A policy execution module, configured to determine a logical log status in the first database based on the first node update time and the second node update time, and control the first database to complete startup according to the logical log status.
10. A server, characterized in that, It includes a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the database cluster inspection and control method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Main-standby switching method and device of cloud database and electronic equipment
CN115827300A
Cluster repairing method and device
CN115904822A
Self-starting method of database cluster, storage medium and equipment
CN115982283A
Self-starting method of database cluster, storage medium and equipment
CN116186162A
Fault processing method, device and equipment and readable storage medium
CN118018463A
Cited By
Database system event response method and device, equipment and storage medium
CN121579176A