Big data platform operation and maintenance method and device, computer equipment and readable storage medium
By automatically judging and improving the current node as the main node, the problem of high operation and maintenance complexity after the main node fails in traditional big data platforms is solved, and high availability and rapid failover are achieved.
Patent Information
- Application Number
- CN202410164466.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-02-05
AI Technical Summary
In traditional big data platforms, after the main node fails, manual intervention is required to switch backup nodes, resulting in increased operation and maintenance complexity and long system unavailability, affecting business operations.
By determining whether the current node is the primary node, if otherwise, a verification request will be sent regularly. If the primary node fails, the current node will be promoted to the primary node, and preset maintenance processes will be performed, such as server registration, plug-in management, task balance and health monitoring.
It realizes rapid failover when the master node fails, ensures high system availability, simplifies operation and maintenance processes, and reduces manual intervention time.
Smart Images

Figure CN120353650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database operation and maintenance, and more particularly, to a method, device, computer device, and readable storage medium for operating and maintaining a big data platform. Background Art
[0002] In a traditional big data platform, there is usually a master node responsible for managing and coordinating the operation of the entire system, such as task allocation, resource management, etc. However, if the master node has problems or fails, the entire system may not work properly, thus affecting the normal operation of the business. Although some systems adopt a master-slave structure to improve availability, after the master node fails, manual intervention is required to switch to the standby node, which not only takes time but also increases the complexity of operation and maintenance. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, device, computer device, and readable storage medium for operating and maintaining a big data platform.
[0004] In a first aspect, an embodiment of the present invention provides a method for operating and maintaining a big data platform, including:
[0005] Obtain the current node and determine whether the current node is the master node;
[0006] If so, execute a preset maintenance process based on the master node;
[0007] If not, send a master node verification request to the master node at a preset interval;
[0008] In the case where the feedback result represented by the master node verification request indicates that the master node has failed, use the current node as the master node.
[0009] In a possible implementation manner, the executing a preset maintenance process based on the master node includes:
[0010] Perform server registration for proxy nodes based on the master node, and add the proxy nodes in the case of successful registration;
[0011] Store the node data of the proxy nodes in a database;
[0012] Send the configured plugins to the proxy nodes using a query plugin list;
[0013] Send the configured solutions to the proxy nodes;
[0014] Request the server resource utilization rate from the proxy nodes and execute a task balancing process.
[0015] In a possible implementation, the preset maintenance process executed based on the master node includes:
[0016] Obtain a new plug-in;
[0017] On the basis of successfully uploading the new plug-in, record the basic information of the new plug-in in the database;
[0018] Broadcast the new plug-in to the proxy nodes;
[0019] In response to a request instruction for the new plug-in function, check whether the new plug-in exists in the plug-in directory of the proxy node;
[0020] If so, feedback that it already exists;
[0021] If not, store the new plug-in in the plug-in directory.
[0022] In a possible implementation, the preset maintenance process executed based on the master node includes:
[0023] Judge whether there is a target task with the same name, and if not, select a task solution and the task information required for the task corresponding to the task solution;
[0024] Request the server resource usage rate from the proxy node;
[0025] Determine the idle proxy nodes according to the server resource usage rate, and send a new task request to the idle proxy nodes;
[0026] In response to a request for a new task interface, check whether the target task exists based on the idle proxy nodes;
[0027] If so, feedback that there is an existing task;
[0028] If not, complete the task information required for the task and start the target task.
[0029] In a possible implementation, the preset maintenance process executed based on the master node includes:
[0030] Request the list of executing tasks that the proxy node is running;
[0031] Judge whether the task list stored in the database is consistent with the executing task list;
[0032] In the case where the task list stored in the database is inconsistent with the executing task list, determine the missing tasks of the proxy node;
[0033] According to the missing tasks, determine the task information required for the missing tasks, and send a new task request to the proxy node;
[0034] In response to receiving the new task request, find whether there is a solution to the new task corresponding to the new task request based on the idle proxy nodes;
[0035] If so, feedback that the task already exists;
[0036] If not, complete the information required for the new task and start the new task.
[0037] In a possible implementation manner, the execution of the preset maintenance process based on the master node includes:
[0038] Send a liveness verification request to the proxy nodes at a preset interval;
[0039] If the liveness verification request feedbacks that the proxy node fails, send an alarm message;
[0040] Query the task list of the proxy node and request a health status detection interface for the task list;
[0041] If the tasks in the task list are unhealthy, request to determine the server resource usage rate of the surviving proxy nodes;
[0042] Send a new task request to the idle surviving proxy nodes so that the idle surviving proxy nodes re-execute the tasks in the task list.
[0043] In a possible implementation manner, the execution of the preset maintenance process based on the master node includes:
[0044] When there is a newly added proxy node, trigger a judgment on whether the newly added proxy node is configured with a liveness task;
[0045] In the case where there is no configured liveness task, request the server resource usage rate from each proxy node and filter out the idle proxy nodes;
[0046] Send a new task request to the idle proxy nodes;
[0047] In response to receiving the new task request, find whether there is a solution to the new task corresponding to the new task request based on the idle proxy nodes;
[0048] If so, feedback that the task already exists;
[0049] If not, complete the information required for the new task and start the new task by the idle proxy node.
[0050] In a second aspect, an embodiment of the present invention provides a big data platform operation and maintenance device, including:
[0051] An acquisition module, configured to acquire a current node and determine whether the current node is a master node;
[0052] An operation and maintenance module, configured to, if so, execute a preset maintenance process based on the master node; if not, send a master node verification request to the master node at a preset interval; and in the case that the feedback result represented by the master node verification request indicates that the master node fails, use the current node as the master node.
[0053] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the big data platform operation and maintenance method in at least one possible implementation manner of the first aspect.
[0054] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the big data platform operation and maintenance method in at least one possible implementation manner of the first aspect.
[0055] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a big data platform operation and maintenance method, device, computer device and readable storage medium disclosed by the present invention, by determining whether the current node is a master node. If so, then execute a preset maintenance process based on the master node; if not, then send a verification request to the master node at a preset interval. When the feedback result obtained from the verification request indicates that the master node has failed, the system will promote the current node to be the new master node. This method can ensure that when a problem occurs with the master node in the big data platform, a quick failover can be performed, thus ensuring the high availability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a schematic flowchart of the steps of the big data platform operation and maintenance method provided by the embodiment of the present invention;
[0058] Figure 2 It is a schematic block diagram of the structure of the big data platform operation and maintenance device provided by the embodiment of the present invention;
[0059] Figure 3A structural schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0061] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.
[0062] To solve the technical problems in the foregoing background art, Figure 1 A flowchart of a big data platform operation and maintenance method provided by an embodiment of the present disclosure is shown below, and the big data platform operation and maintenance method will be introduced in detail.
[0063] Step S201: Obtain the current node and determine whether the current node is the master node;
[0064] Step S202: If it is, execute a preset maintenance process based on the master node;
[0065] Step S203: If not, send a master node verification request to the master node at a preset interval;
[0066] Step S204: In the case where the feedback result represented by the master node verification request indicates that the master node has failed, use the current node as the master node.
[0067] In an embodiment of the present invention, by way of example, assume that you are managing a big data cluster composed of multiple nodes (servers). When you start the operation and maintenance process, you need to first determine whether the node you are operating on is the master node. For example, you are checking the status of node A. If node A is the master node, then you can execute a preset maintenance process on this node. For example, it may be necessary to update software, check the system's running status, or perform other maintenance tasks. If node A is not the master node, then it needs to send a verification request to the master node regularly to confirm whether the master node is running normally. For example, node A sends a "Are you okay?" message to the master node every 5 minutes. If the master node does not respond to node A's verification request, or it is found from the information returned by the master node that the master node has failed, then node A will take over and become the new master node. For example, if the master node does not respond to node A's verification request within the specified time, then the system will start a failover procedure to set node A as the new master node to ensure the normal operation of the system.
[0068] In an embodiment of the present invention, the foregoing step S204 may be implemented through the following examples.
[0069] (1) Perform server registration for the proxy node based on the master node. In the case of successful registration, add the proxy node.
[0070] (2) Store the node data of the proxy node in the database.
[0071] (3) Send the configured plugins to the proxy node using the query plugin list.
[0072] (4) Send the configured solution to the proxy node.
[0073] (5) Request the server resource utilization rate from the proxy node and execute the task balancing process.
[0074] In an embodiment of the present invention, by way of example, in this example, the master node may be managing a series of proxy nodes (Agents). When a new Agent needs to be added, the master node is responsible for registering this new Agent. For example, if you want to add node B as a new Agent, then you first need to complete the registration of node B on the master node. After successful registration, the master node will store the relevant information of the Agent (such as IP address, running status, etc.) in the database. For example, the master node will save information such as the IP address of node B, the operating system version, and the CPU utilization rate to the database. The master node will view the current available plugin list and send the configured plugins to the newly registered Agent. For example, the master node sends plugins such as the json_to_sql data conversion plugin and the mongo_to_greenplum full extraction plugin to node B. In addition to plugins, the master node also sends the configured solution to the Agent. For example, the master node may send a solution containing the detailed steps on how to synchronize data to node B. To ensure the smooth operation of the entire system, the master node also queries the Agent for its server resource usage. For example, the master node may ask node B about its current CPU and memory utilization rates. Based on this information, the master node can perform task scheduling to maintain load balancing. For example, if the resource utilization rate of node B is low, the master node may assign more tasks to node B for execution.
[0075] In an embodiment of the present invention, the foregoing step S204 may be implemented through the following examples.
[0076] (1) Obtain the newly added plugin.
[0077] (2) On the basis of successfully uploading the newly added plugin, record the basic information of the newly added plugin in the database.
[0078] (3) Broadcast the newly added plugin to the proxy nodes;
[0079] (4) In response to a request instruction for the newly added plugin function, check whether the newly added plugin exists in the plugin directory of the proxy node;
[0080] (5) If so, feedback that it already exists;
[0081] (6) If not, store the newly added plugin in the plugin directory.
[0082] In an embodiment of the present invention, by way of example, assume that a new plugin needs to be added to the system, such as a new data cleaning plugin. First, you need to obtain this newly added plugin on the master node. After obtaining the plugin, upload it to the master node, and after the upload is successful, record the basic information of this newly added plugin (such as plugin name, version, function description, etc.) in the database. Then, the master node broadcasts this newly added plugin to all Agents. For example, the master node will notify all Agents: "We now have a new data cleaning plugin". When a task requires the use of this newly added plugin, the master node will check the plugin directory of each Agent to see if this newly added plugin already exists. For example, if a task requires the use of the new data cleaning plugin, the master node will check whether node B already has this plugin. If the check result shows that node B already has this plugin, the master node will feedback "already exists". If node B does not have this plugin yet, then the master node will send this newly added plugin to node B and save it in its plugin directory. It should be noted that the process of obtaining a new solution is basically the same as the process of obtaining a new plugin described above.
[0083] In an embodiment of the present invention, the foregoing step S204 may be executed by the following example.
[0084] (1) Determine whether there is a target task with the same name, and if not, select a task solution and the information required for the task corresponding to the task solution;
[0085] (2) Request the server resource usage rate from the proxy node;
[0086] (3) Determine the idle proxy nodes based on the server resource usage rate, and send a new task request to the idle proxy nodes;
[0087] (4) In response to a request for the newly added task interface, check whether the target task exists based on the idle proxy nodes;
[0088] (5) If so, feedback that the task already exists;
[0089] (6) If not, complete the information required for the task and start the target task.
[0090] In an embodiment of the present invention, by way of example, assume that you need to add a new data synchronization task to the system. First, you need to check on the master node whether there is already a task with the same name. If not, then you need to select a suitable task solution and prepare the relevant task information. For example, you may select a synchronization solution using the json_to_sql plugin and prepare the connection information for the source database and the target database. After determining the task solution, the master node needs to query all Agents about their server resource usage. For example, the master node will ask node B about its current CPU and memory usage rates. Based on the resource usage information returned by the Agents, the master node can determine which Agents are currently idle. Then, the master node will send a new task request to these idle Agents. For example, if the resource usage rate of node B is the lowest, then the master node will send a request to node B to create a new data synchronization task. When an Agent receives a new task request, it will look locally to see if the target task already exists. For example, node B will check if there is already a data synchronization task with the same name. If node B already has this task, then it will feedback to the master node "task already exists". If not, then node B will complete the task configuration based on the task information provided by the master node and start this new data synchronization task.
[0091] In an embodiment of the present invention, the foregoing step S204 can be implemented through the following example.
[0092] (1) Request the list of execution tasks being run by the proxy node;
[0093] (2) Determine whether the task list stored in the database is consistent with the execution task list;
[0094] (3) In the case where the task list stored in the database is inconsistent with the execution task list, determine the missing tasks of the proxy node;
[0095] (4) Based on the missing tasks, determine the task information required for the missing tasks and send a new task request to the proxy node;
[0096] (5) In response to receiving the new task request, based on the idle proxy nodes, find out whether there is a solution for the new task corresponding to the new task request;
[0097] (6) If so, feedback that the task already exists;
[0098] (7) If not, complete the information required for the new task and start the new task.
[0099] In an embodiment of the present invention, by way of example, the master node needs to know which tasks each Agent is currently running. For example, the master node may request from node B the list of tasks it is currently executing. The master node will compare the task list obtained from the Agent with the task list stored in the database. For example, if the task list recorded in the database is "A, B, C", and the task list returned by node B is "A, B", then the master node knows that these two lists are inconsistent. If it is found that the two task lists are inconsistent, the master node can determine which tasks are missing from the Agent. In the above example, the master node will find that node B is missing task "C". Once the missing tasks are determined, the master node will prepare the relevant task information and send a request to the Agent to create a new task. For example, the master node will prepare the information for task "C" and send a request to node B to create a new task "C". When the Agent receives the request to create a new task, it will look locally to see if there is already a solution for this task. For example, node B will check if there is already a solution for task "C". If node B already has a solution for this task, then it will feedback to the master node "task already exists". If not, then node B will complete the task configuration based on the task information provided by the master node and start this new task "C".
[0100] In an embodiment of the present invention, the foregoing step S204 may be implemented through the following example.
[0101] (1) Send a liveness verification request to the agent node at a preset interval;
[0102] (2) If the liveness verification request feedbacks that the agent node fails, send an alarm message;
[0103] (3) Query the task list of the agent node and request a health status detection interface for the task list;
[0104] (4) If the tasks in the task list are unhealthy, request to determine the server resource utilization rate of the surviving agent nodes;
[0105] (5) Send a request to create a new task to the idle surviving agent nodes so that the idle surviving agent nodes re-execute the tasks in the task list.
[0106] In an embodiment of the present invention, by way of example, the master node needs to periodically check whether each Agent is still running normally. For example, the master node may send a message "Are you okay?" to Node B every 5 minutes. If Node B does not respond to the liveness verification request, or it is found from the information returned that Node B has failed, then the master node will send an alarm message. For example, the master node may notify the administrator by email or text message: "Node B has failed". In addition to checking whether the Agent itself is running normally, the master node also needs to check whether the tasks being executed by the Agent are healthy. For example, the master node will request the list of tasks currently being executed from Node B and perform a health check on these tasks. If it is found that a certain task of Node B is unhealthy (for example, there is no progress all the time, or it fails repeatedly), then the master node needs to find other Agents to take over this task. First, the master node will request the server resource usage of all surviving Agents. The master node will find the currently most idle Agent based on the returned resource usage and send a request to create a new task to it. For example, if it is found that Node C is currently idle, then the master node will send a request to Node C: "Please take over the task that Node B has failed".
[0107] In an embodiment of the present invention, the foregoing step S204 may be executed by the following example.
[0108] (1) When there is a newly added agent node, trigger a judgment on whether the newly added agent node is configured with a liveness task;
[0109] (2) In the case of no configured liveness task, request the server resource utilization rate from each agent node and filter out the idle agent nodes;
[0110] (3) Send a new task request to the idle agent node;
[0111] (4) In response to the receipt of the new task request, check whether there is a solution to the new task corresponding to the new task request based on the idle agent node;
[0112] (5) If so, feedback that the task already exists;
[0113] (6) If not, complete the information required for the new task, and start the new task by the idle agent node.
[0114] In an embodiment of the present invention, exemplarily, when a new Agent (e.g., Node D) is added to the system, the master node needs to determine whether this new Agent has been configured with the tasks to be executed. If Node D has not been configured with any tasks, then the master node will first request all Agents for their server resource usage, and then find out the currently most idle Agent. For example, if the returned result shows that Node D is currently very idle, then the master node will consider Node D as an idle Agent. After determining the idle Agent, the master node can send a request to it to create a new task. For example, the master node may send a request to Node D: "Please start executing the data synchronization task." When the Agent receives the request to create a new task, it will look up locally whether it already has a solution for this task. For example, Node D will check whether it already has a solution for the data synchronization task. If Node D already has a solution for this task, then it will feedback to the master node "task already exists". If not, then Node D will complete the task configuration according to the task information provided by the master node and start this new data synchronization task.
[0115] In summary, the embodiment of the present invention provides a highly available server platform, mainly composed of a server side and an agent side, which can implement functions such as task management, plugin management, solution management, task monitoring, and server health monitoring. Among them, the server side is mainly responsible for the management of overall functions, while the agent side is responsible for the execution of specific tasks and the management of server resources.
[0116] The following is a summary of this system:
[0117] High availability: The system adopts a master-standby scheme, that is, a master node and a standby node. When the master node has problems, the standby node can take over the work of the master node, thus ensuring the high availability of the system.
[0118] Task management: The server side is responsible for allocating tasks to each agent for execution. Each task consists of one or more shell scripts, which are called synchronization solutions. The execution status of the task will be monitored in real time. If a task execution fails or an Agent node fails, the system will immediately issue an alarm.
[0119] Plugin management: The system completes the data synchronization function by using various plugins. For example, there are plugins dedicated to handling dynamic changes in database data, plugins for converting JSON data into SQL, plugins for full extraction of table data, and full extraction plugins from MongoDB to Greenplum, etc.
[0120] Solution Management: Each synchronization task consists of one or several synchronization solutions. A synchronization solution is a set of plugins that execute according to a specific order and logic. The server side is responsible for managing these synchronization solutions, including the creation, modification, deletion of solutions, and corresponding task allocation.
[0121] Server Monitoring: The system periodically checks the health status and resource usage of each Agent node, and performs task scheduling based on this information to achieve balanced task allocation.
[0122] Generally speaking, through the cooperation between the server side and the agent side, this system realizes the automated management and efficient execution of the data synchronization function.
[0123] Please refer to Figure 2 , Figure 2 For a big data platform operation and maintenance device 110 provided by an embodiment of the present invention, it includes:
[0124] An acquisition module 1101, configured to acquire the current node and determine whether the current node is the master node;
[0125] An operation and maintenance module 1102, configured to, if so, execute a preset maintenance process based on the master node; if not, send a master node verification request to the master node at a preset interval; in the case where the feedback result represented by the master node verification request indicates that the master node fails, use the current node as the master node.
[0126] It should be noted that the implementation principle of the foregoing big data platform operation and maintenance device 110 can refer to the implementation principle of the foregoing big data platform operation and maintenance device 110, which will not be elaborated here. It should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the big data platform operation and maintenance device 110 can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above big data platform operation and maintenance device 110. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0127] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0128] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned big data platform operation and maintenance device 110. Figure 3 As shown, Figure 3 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a big data platform operation and maintenance device 110, a memory 111, a processor 112 and a communication unit 113.
[0129] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be realized through one or more communication buses or signal lines. The big data platform operation and maintenance device 110 includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the big data platform operation and maintenance device 110 stored in the memory 111, such as the software function modules and computer programs included in the big data platform operation and maintenance device 110.
[0130] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned big data platform operation and maintenance method.
[0131] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of the disclosure and its practical application, to thereby enable others skilled in the art to best utilize the disclosure and various embodiments with different modifications as are suited to the particular application contemplated.
Claims
1. A method for operation and maintenance of a big data platform, characterized in that, including: Obtain the current node and determine whether the current node is the master node; If so, execute a preset maintenance process based on the master node; If not, send a master node verification request to the master node at a preset interval; In the case where the feedback result represented by the master node verification request indicates that the master node fails, use the current node as the master node.
2. The method according to claim 1, characterized in that The executing the preset maintenance process based on the master node includes: Perform server registration for proxy nodes based on the master node, and add the proxy nodes in the case of successful registration; Store the node data of the proxy nodes into the database; Send the configured plugins to the proxy nodes using the query plugin list; Send the configured solutions to the proxy nodes; Request the server resource usage rate from the proxy nodes and execute a task balancing process.
3. The method according to claim 1, characterized in that The executing the preset maintenance process based on the master node includes: Obtain new plugins; On the basis of successfully uploading the new plugins, record the basic information of the new plugins into the database; Broadcast the new plugins to the proxy nodes; In response to a request instruction for the new plugin function, check whether the new plugins exist in the plugin directory of the proxy nodes; If so, feedback that they already exist; If not, store the new plugins into the plugin directory.
4. The method according to claim 1, wherein The executing the preset maintenance process based on the master node includes: Judge whether there are target tasks with the same name, and in the case of non-existence, select a task solution and the task information required for the task solution; Request the server resource usage rate from the proxy nodes; Determine idle proxy nodes according to the server resource usage rate, and send a new task request to the idle proxy nodes; In response to a request for a new task interface, check whether the target task exists based on the idle proxy nodes; If so, feedback that the task already exists; If not, complete the task information required and start the target task.
5. The method according to claim 1, wherein The executing the preset maintenance process based on the master node includes: Request the list of executing tasks of the proxy nodes; Judge whether the task list stored in the database is consistent with the executing task list; In the case where the task list stored in the database is inconsistent with the executing task list, determine the missing tasks of the proxy nodes; According to the missing tasks, determine the task information required for the missing tasks, and send a new task request to the proxy nodes; In response to receiving the new task request, check whether there is a solution for the new task corresponding to the new task request based on the idle proxy nodes; If so, feedback that the task already exists; If not, complete the information required for the new task and start the new task.
6. The method according to claim 1, wherein The executing the preset maintenance process based on the master node includes: Send a survival verification request to the proxy nodes at a preset interval; If the survival verification request feedbacks that the proxy node fails, send an alarm message; Query the task list of the proxy nodes and request a health status detection interface for the task list; If the tasks in the task list are unhealthy, request to determine the server resource usage rate of the surviving proxy nodes; Send a new task request to the idle surviving agent nodes so that the idle surviving agent nodes can re-execute the tasks in the task list.
7. The method according to claim 1, characterized in that The execution of the preset maintenance process based on the master node includes: When there is a newly added agent node, trigger a judgment on whether the newly added agent node is configured with a survival task; In the case where there is no configured survival task, request the server resource utilization rate from each agent node and filter out the idle agent nodes; Send a new task request to the idle agent nodes; In response to the receipt of the new task request, search for a solution to the new task corresponding to the new task request based on the idle agent nodes; If so, feedback that the task already exists; If not, complete the information required for the new task, and start the new task by the idle agent node.
8. An operation and maintenance device for a big data platform, characterized in that, It includes: An acquisition module for acquiring the current node and judging whether the current node is the master node; An operation and maintenance module for, if so, executing a preset maintenance process based on the master node; If not, send a master node verification request to the master node at a preset interval; In the case where the feedback result represented by the master node verification request indicates that the master node fails, use the current node as the master node.
9. A computer device, characterized in that, The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the big data platform operation and maintenance method according to any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the big data platform operation and maintenance method according to any one of claims 1-7.
Citation Information
Patent Citations
Master-slave service system and master node fault recovery method and device
CN108964948A
Database node cluster fault migration method and device
CN110635941A
Node autonomous method, system and device of node cluster and electronic equipment
CN112035215A
Equipment access method, node equipment, server and storage medium
CN116266801A
A Building Method of Healthcare Information System Access Control Software Architecture Design Process under Clustering Medical Information Environment
KR101583318B1