Data processing method and device, medium and program product
By analyzing the processor utilization of the working nodes in the database cluster, the data sets of high-load tasks are stored in the backup nodes for processing, which solves the problem of idle and wasted backup nodes, improves the processing efficiency and resource utilization of the cluster, and ensures data consistency and timeliness.
Patent Information
- Application Number
- CN202511178784.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-21
AI Technical Summary
In a database cluster, backup nodes provide services when working nodes fail, but remain idle when working nodes are operating normally, resulting in resource waste and an inability to share the load of working nodes, reducing the overall processing efficiency and resource utilization of the cluster.
By analyzing the processor utilization of the working nodes, we filter out the historical tasks of the nodes with high processor utilization, and store their data sets in the idle backup nodes. We use the backup nodes to process the new tasks and reduce the load on the working nodes.
It improves the overall processing efficiency and resource utilization of the database cluster, reduces the resource consumption of working nodes, and ensures data consistency and timeliness.
Smart Images

Figure CN120743546A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database technology, and more specifically to a data processing method, device, medium and program product. Background Art
[0002] In a database cluster, backup nodes are usually set up to ensure service stability. When a working node fails, the backup node is usually used to replace the failed working node to process data.
[0003] However, when there are fewer failed working nodes in the cluster, the backup nodes will be idle, resulting in resource waste. Summary of the Invention
[0004] In view of the above problems, the present application provides a data processing method, device, medium and program product.
[0005] According to the first aspect of the present application, a data processing method is provided, including: analyzing the respective working states of multiple working nodes in a cluster within a predetermined historical period, and determining the respective processor utilization rates of the multiple working nodes within the predetermined historical period; in response to the presence of a first working node among the multiple working nodes whose processor utilization rate is greater than a preset threshold, screening the multiple historical tasks according to the respective historical processor utilization rates of the multiple historical tasks executed by the first working node within the predetermined historical period to obtain a target historical task; storing a target data set required for processing the target historical task in a first backup node, the first backup node being determined from the multiple idle backup nodes of the cluster; in response to receiving a newly added task, when it is determined that the newly added task is used to perform a target operation on the target data set, sending the newly added task to the first backup node, so as to use the first backup node to process the newly added task.
[0006] The second aspect of the present application provides a data processing device, including: an analysis module, used to analyze the respective working states of multiple working nodes in a cluster within a predetermined historical period, and determine the respective processor utilization rates of the multiple working nodes within the predetermined historical period; a screening module, used to, in response to the presence of a first working node among the multiple working nodes whose processor utilization rate is greater than a preset threshold, screen the multiple historical tasks according to the respective historical processor utilization rates of the multiple historical tasks executed by the first working node within the predetermined historical period, and obtain a target historical task; a storage module, used to store the target data set required for processing the target historical task to a first backup node, and the first backup node is determined from the multiple idle backup nodes of the cluster; a sending module, used to, in response to receiving a newly added task, send the newly added task to the first backup node when it is determined that the newly added task is used to perform a target operation on the target data set, so as to use the first backup node to process the newly added task.
[0007] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0008] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0009] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0010] According to an embodiment of the present application, when the processor utilization of the first working node is high, the historical tasks executed by the first working node within a predetermined historical period are screened, and the target historical tasks that cause the processor utilization of the first working node to be high are determined. By storing the target data set required for executing the target historical tasks in the first backup node, the first backup node can be used to execute new tasks that perform target operations on the target data set, thereby fully utilizing the idle backup node resources while reducing the resource consumption of the first working node and improving the overall processing efficiency and resource utilization of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings.
[0012] Figure 1An application scenario diagram of the data processing method, device, medium, and program product according to an embodiment of the present application is shown.
[0013] Figure 2 A flow chart of a data processing method according to an embodiment of the present application is shown.
[0014] Figure 3 A schematic diagram of storing a distributed data table in a first backup node according to an embodiment of the present application is shown.
[0015] Figure 4 A flowchart of processing the second data processing task according to an embodiment of the present application is shown.
[0016] Figure 5 A schematic diagram of processing a newly added task according to an embodiment of the present application is shown.
[0017] Figure 6 A schematic diagram of a second backup node loading data according to an embodiment of the present application is shown.
[0018] Figure 7 The figure shows a structural block diagram of a data processing device according to an embodiment of the present application.
[0019] Figure 8 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0021] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0023] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0024] A small number of backup nodes are configured in the database cluster. When a working node in the cluster fails, the backup node automatically mounts the failed working node's external storage volume, transforming it into a working node and providing normal data read and write services. If all working nodes in the cluster are operating normally, the backup node will not mount any other working node's external storage volume. This means that the working node remains idle and has no tasks to execute. This highly resource-intensive device remaining idle is a significant waste of resources. Furthermore, even if other working nodes are very busy, the backup node will not participate in any cluster tasks, thus not contributing to cluster efficiency.
[0025] An embodiment of the present application provides a data processing method, including: analyzing the respective working states of multiple working nodes in a cluster within a predetermined historical period, and determining the respective processor utilization rates of the multiple working nodes within the predetermined historical period; in response to the presence of a first working node among the multiple working nodes whose processor utilization rate is greater than a preset threshold, screening the multiple historical tasks according to the respective historical processor utilization rates of the multiple historical tasks executed by the first working node within the predetermined historical period to obtain a target historical task; storing a target data set required for processing the target historical task in a first backup node, the first backup node being determined from multiple idle backup nodes in the cluster; in response to receiving a newly added task, when it is determined that the newly added task is used to perform a target operation on the target data set, sending the newly added task to the first backup node, so as to utilize the first backup node to process the newly added task, thereby sharing the load pressure of the first working node, improving the task processing efficiency of the cluster, and improving the resource utilization rate of the cluster, thereby improving the overall performance of the cluster.
[0026] Figure 1 An application scenario diagram of the data processing method, device, medium, and program product according to an embodiment of the present application is shown.
[0027] like Figure 1As shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a database cluster 105. The database cluster 105 may include multiple nodes, such as a first node 105_1, a second node 105_2, and a third node 105_3. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the database cluster 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0028] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a database cluster 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0029] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0030] The database cluster 105 can provide data processing services, such as providing support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103, analyzing and processing received user request data, and feeding back processing results (such as data obtained or generated according to user requests, etc.) to the terminal devices.
[0031] It should be noted that the data processing method provided in the embodiment of the present application can generally be executed by the database cluster 105. Accordingly, the data processing device provided in the embodiment of the present application can generally be set in the database cluster 105. The data processing method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the database cluster 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the database cluster 105. Accordingly, the data processing device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the database cluster 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the database cluster 105.
[0032] It should be understood that Figure 1The number of first terminal devices, second terminal devices, third terminal devices, networks, database clusters, first nodes, second nodes, and third nodes in the embodiment is merely illustrative. Any number of first terminal devices, second terminal devices, third terminal devices, networks, database clusters, first nodes, second nodes, and third nodes may be provided as required.
[0033] The following will be based on Figure 1 The scene described by Figures 2 to 6 The data processing method of the application embodiment is described in detail.
[0034] Figure 2 A flow chart of a data processing method according to an embodiment of the present application is shown.
[0035] like Figure 2 As shown, the data processing method of this embodiment includes operations S210 to S240.
[0036] In operation S210 , the working status of each of the plurality of working nodes in the cluster within a predetermined historical period is analyzed to determine the processor utilization rate of each of the plurality of working nodes within the predetermined historical period.
[0037] In operation S220, in response to the presence of a first working node among the multiple working nodes whose processor utilization is greater than a preset threshold, the multiple historical tasks are screened according to the historical processor utilization of each of the multiple historical tasks executed by the first working node in a predetermined historical period to obtain a target historical task.
[0038] In operation S230 , a target data set required for processing the target historical task is stored in the first backup node.
[0039] In operation S240 , in response to receiving the newly added task, if it is determined that the newly added task is used to perform a target operation on the target data set, the newly added task is sent to the first backup node, so that the first backup node processes the newly added task.
[0040] In an embodiment of the present application, the cluster may be a database cluster, such as a scale-out cluster, and the database cluster may include multiple different types of nodes, such as working nodes, backup nodes, and scheduling nodes. Working nodes may be used to provide data read and write services, backup nodes may replace failed working nodes in providing read and write services, and scheduling nodes may dispatch tasks received by the database cluster to working nodes or backup nodes that provide read and write services in place of failed working nodes. In a specific embodiment, the data processing method of the present application may be executed by a scheduling node.
[0041] Since there may be no faulty working nodes or few faulty working nodes in the cluster, there are idle backup nodes in the cluster. As a result, when the processor utilization of the working nodes is high, the backup nodes cannot share the load of the working nodes, which leads to a waste of backup node resources. Therefore, idle backup nodes can be used to process tasks that originally need to be processed by working nodes to reduce the processor utilization of working nodes.
[0042] In order to better utilize the idle backup nodes to reduce the processor utilization of the working nodes, we can first determine the first working node with higher processor utilization among multiple working nodes, then analyze the tasks that need to be processed by the first working node, and assign the tasks that consume more processor resources to the idle backup nodes for processing, so as to utilize the idle first backup nodes to share the tasks of the first working node.
[0043] In some embodiments, the first backup node may be determined based on resource configuration information of each of a plurality of idle backup nodes. For example, an idle backup node with a higher resource configuration may be selected as the first backup node.
[0044] When determining the first working node, the working status of multiple working nodes in the cluster within a predetermined historical period can be analyzed, such as performing statistical analysis on the processor utilization of the working nodes at multiple predetermined historical moments in the predetermined historical period, obtaining the processor utilization of each working node within the predetermined historical period, and determining the working node whose processor utilization is greater than a preset threshold as the first working node.
[0045] For example, the scheduling node may send a first acquisition instruction to each working node for acquiring working status data for a predetermined historical period, and the working node returns the working status data for the predetermined historical period to the scheduling node in response to the first acquisition instruction. The scheduling node analyzes the working status of multiple working nodes based on the working status data returned by the working node.
[0046] After the first working node is determined, a target historical task that causes a higher processor utilization rate of the first working node may be determined based on historical tasks executed by the first working node within a predetermined historical period.
[0047] For example, the scheduling node can send a second acquisition instruction to the first worker node for acquiring log data for a predetermined historical period. The first worker node, in response to the second acquisition instruction, returns the log data for the predetermined historical period to the scheduling node. Because the log data includes the historical processor utilization rates of each of the multiple historical tasks executed by the first worker node during the predetermined historical period, the scheduling node can select target historical tasks with high historical processor utilization rates based on the received log data.
[0048] In an embodiment of the present application, since data sets from multiple data tables can be stored in a working node, the high historical processor utilization of the target historical task may be due to the large amount of data or high data complexity of the target data set required to process the target historical task. Therefore, the target data set can be stored in the first backup node to use the first backup node to execute tasks for the target data set.
[0049] Since the tasks for the target data set include executing query operations, calculation operations, and change operations on the target data set, and executing query operations and calculation operations consumes a large amount of processor resources, the query operations and calculation operations can be determined as target operations, and the tasks for executing target operations on the target data set can be assigned to the first backup node for execution, so that the first working node no longer needs to execute the target operations on the target data set.
[0050] After the target dataset has been completely stored on the first backup node, the first backup node is now ready to execute tasks for the target dataset. If the scheduling node determines that the newly added task is for performing the target operation on the target dataset, it can directly send the newly added task to the first backup node for processing.
[0051] According to an embodiment of the present application, when the processor utilization of the first working node is high, the historical tasks executed by the first working node within a predetermined historical period are screened, and the target historical tasks that cause the processor utilization of the first working node to be high are determined. By storing the target data set required for executing the target historical tasks in the first backup node, the first backup node can be used to execute new tasks that perform target operations on the target data set, thereby fully utilizing the idle backup node resources while reducing the resource consumption of the first working node and improving the overall processing efficiency and resource utilization of the cluster.
[0052] According to an embodiment of the present application, in the process of storing the target data set required for processing the target historical task to the first backup node, the data processing method also includes: in response to receiving the first data processing task, parsing the query statement included in the first data processing task, determining the first operation type of the first to-be-executed operation and the first data identifier of the first to-be-executed data set for which the first to-be-executed operation is targeted; based on the first data identifier and the first operation type, determining whether the first data processing task is a first data change task for performing a change operation on the target data set.
[0053] Because a first data change task for performing a change operation on the target dataset may be received during the process of storing the target dataset on the first backup node, the scheduling node cannot directly send the first data processing task to the working node after receiving the first data processing task. This prevents the first working node from directly performing the change operation on the target dataset when the first data processing task is used to perform the change operation on the target dataset, thereby causing data inconsistency between the first backup node and the first working node. Change operations may include, for example, adding data, deleting data, and modifying data.
[0054] When identifying the first data processing task, since users typically represent the operations to be performed using a query statement, the query statement in the first data processing task can be identified to determine the first operation type of the first operation to be performed and the first data identifier of the first data set to be operated on that is targeted by the first operation to be performed. For example, the query statement can be a structured query statement, and the first data identifier can be the table identifier of the data table corresponding to the first data set to be operated on.
[0055] After determining the first operation type and the first data identifier, it can be determined whether the first operation type is a change operation, and whether the first data identifier is the same as the data identifier of the target data table. When the first operation type is a change operation and the first data identifier is the same as the data identifier of the target data table, the first data processing task is determined to be the first data change task for performing a change operation on the target data set.
[0056] According to an embodiment of the present application, by identifying the received first data processing task in the process of storing the target data set required for processing the target historical task in the first backup node, it is possible to avoid changing the target data set in the process of storing the target data set in the first backup node, thereby ensuring data consistency between the first backup node and the first working node.
[0057] According to an embodiment of the present application, the data processing method also includes: in the process of storing the target data set required for processing the target historical task to the first backup node, in response to receiving a first data change task for performing a change operation on the target data set, updating the information of the first data change task and generating a to-be-executed task with added identification information indicating that it is not executed temporarily; sending the to-be-executed task to the first working node, so that after the target data set is completely stored in the first backup node, the first working node updates the first target data to be updated in the to-be-executed task based on the changed data in the to-be-executed task.
[0058] While the target dataset is being stored on the first backup node, the cluster may receive a data change task to modify the target dataset. Because executing the data change task will cause changes to the target dataset on the first working node, to ensure data consistency, the data change task can be temporarily suspended while the target dataset is being stored on the first backup node. After the target dataset is completely stored on the first backup node, modifications to the target dataset are performed and the target dataset stored on the first backup node is updated.
[0059] After receiving the first data change task, the scheduling node may update the first data change task. Specifically, the scheduling node may add a flag indicating that the task is temporarily not executed to the first data change task, obtain the pending task, and send the pending task to the first working node. This causes the first working node to temporarily suspend updating the target dataset after receiving the pending task.
[0060] In an embodiment of the present application, the first working node may send the target dataset stored in its memory to the first backup node via the cluster heartbeat network, so that the target dataset is stored in the memory of the first backup node. After the first working node determines that the target dataset has been completely stored in the first backup node, it executes the pending task to update the first target data.
[0061] According to an embodiment of the present application, by not temporarily executing the first data change task for changing the target data set during the process of storing the target data set to the first backup node, but updating the first target data in the target data set only after the target data set is completely stored to the first backup node, the data consistency between the first backup node and the first working node can be ensured.
[0062] According to an embodiment of the present application, the data processing method also includes: in response to the target data set required for processing the target historical task having been completely stored in the first backup node, generating a second data change task for the first backup node based on the first data change task; sending the second data change task to the first backup node, so that the first backup node updates the first target data stored in the first backup node to the changed data in response to the second data change task.
[0063] After the target data set required for processing the target historical task has been completely stored in the first backup node, the first working node can send a storage completion message to the scheduling node. In response to the storage completion message, the scheduling node can generate a second data change task to change the target data set in the first backup node based on the first data change task.
[0064] According to an embodiment of the present application, after the target data set is completely stored in the first backup node, a second data change task is generated to update the target data set in the first backup node, which not only ensures the data consistency between the first backup node and the first working node, but also ensures the timeliness of the data stored in the first backup node.
[0065] According to an embodiment of the present application, the data processing method further includes: determining the distributed data table to which the target data set belongs; and storing the distributed data table in the first backup node according to a data table identifier of the distributed data table.
[0066] In an embodiment of the present application, a data table can be stored in a distributed manner in multiple worker nodes of a cluster. For example, a distributed data table can include multiple partitions, each partition can correspond to a data set, and each data set is stored in a different worker node. When the cluster receives a data processing task for a distributed data table, the scheduling node needs to simultaneously send the data processing task to multiple worker nodes that store the data sets of the distributed data table, so that the multiple worker nodes can respectively execute the data processing task according to the stored data sets.
[0067] The reason for the high processor utilization of the first working node may be the high complexity of the target data set, such as a large number of fields, which leads to the high complexity of the distributed data table to which the target data set belongs, and thus may lead to the high processor utilization of the working node of the data set storing the distributed data table. Therefore, the data of the distributed data table can be stored in the first backup node, so that the first backup node can be used to perform the target operation on the distributed data table.
[0068] When determining the distributed data table to which the target data set belongs, the distributed data table corresponding to the table identifier can be determined based on the table identifier corresponding to the target data set. Furthermore, the scheduling node can determine other worker nodes storing the data set of the distributed data table based on the table identifier of the distributed data table, and send copy instructions to the other worker nodes. After receiving the copy instructions, the other worker nodes send the data set of the distributed data table stored in each node to the first backup node via the cluster heartbeat network.
[0069] According to an embodiment of the present application, by storing the distributed data table to which the target data set belongs to the first backup node, and utilizing the first backup node to process the target operations on the distributed data table, the first backup node can process the tasks of the associated working nodes of the first working node, thereby improving the overall data processing efficiency of the cluster.
[0070] Figure 3 A schematic diagram of storing a distributed data table in a first backup node according to an embodiment of the present application is shown.
[0071] like Figure 3 As shown, after analyzing the working status of each of Worker Nodes 1, 2, and 3 during a predetermined historical period, the first Worker Node and the target dataset are determined. Worker Node 1 stores Table A dataset 1 and Table B dataset 1, Worker Node 2 stores Table A dataset 2 and Table C dataset 1, and Worker Node 3 stores Table B dataset 2 and Table C dataset 2.
[0072] like Figure 3 As shown, taking working node 3 as the first working node and table B dataset 2 as the target dataset as an example, all tables B belonging to table B dataset 2 need to be stored in the first backup node. Therefore, working node 1 needs to store table B dataset 1 in the first backup node through the node cluster heartbeat network, and working node 3 needs to store table B dataset 2 in the cluster heartbeat network in the first backup node.
[0073] According to an embodiment of the present application, the data processing method also includes: in a case where the target data set required for processing the target historical task has been completely stored in the first backup node, in response to receiving a second data processing task, parsing the query statement included in the second data processing task, determining the second operation type of the second operation to be executed and the second data identifier of the second data set to be operated for the second operation to be executed; based on the second data identifier and the second operation type, determining the identification result for the second data processing task; in a case where it is determined that the identification result represents that the second data processing task is used to perform a change operation on the target data set, sending the second data processing task to the first working node and the first backup node, so that the first working node and the first backup node can be used to process the second data processing task respectively.
[0074] When the target data set required for processing the target historical task has been completely stored on the first backup node, the first backup node is now capable of performing the target operation on the target data set. Therefore, upon receiving the second data processing task, the scheduling node can parse the query statement in the second data processing task and identify the second data identifier and the second operation type of the second data processing task. The second data identifier can, for example, be the table identifier of the data table targeted by the second data processing task.
[0075] When the second operation type is a change operation and the second data identifier is the same as the data table identifier of the distributed data table to which the target data set belongs, it can be determined that the recognition result indicates that the second data processing task is used to perform the change operation on the target data set.
[0076] When the identification result indicates that the second data processing task is used to perform a change operation on the target data set, in order to ensure data consistency and timeliness between the first backup node and the first working node, the second data processing task can be sent to the first working node and the first backup node, so that the first working node and the first backup node respectively execute the second data processing task to update the target data set.
[0077] After completing the second data processing task, the first working node and the first backup node can each send execution results to the scheduling node. If the scheduling node determines that both execution results indicate a successful change, it determines that the second data processing task has been successfully executed and returns an execution result indicating successful execution to the system that initiated the second data processing task.
[0078] According to an embodiment of the present application, when the target data set required for processing the target historical task has been completely stored in the first backup node, the second data processing task for performing the target operation on the target data set is sent to the first backup node and the first working node respectively, thereby ensuring the data consistency and timeliness of the first backup node and the first working node, and improving the data processing efficiency of the cluster.
[0079] Figure 4 A flowchart of processing the second data processing task according to an embodiment of the present application is shown.
[0080] In operation S410 , a second data processing task sent by an application system is received.
[0081] In operation S420, it is determined whether the target data set is synchronized to the first backup node. If it is determined that the target data set is synchronized to the first backup node, operation S430 is performed; otherwise, operation S440 is performed.
[0082] In operation S430 , the second data processing task is sent to the first working node and the first backup node respectively.
[0083] In operation S440 , the second data processing task is sent to the first working node.
[0084] In operation S450 , a modification success notification is sent to the application system.
[0085] According to an embodiment of the present application, a new task is sent to the first backup node so that the first backup node can be used to process the new task, including: sending the system information of the new task and the application system that initiates the new task to the first backup node, so that the first backup node establishes a connection with the application system based on the system information, and sends the processing result of the new task to the application system.
[0086] When the application system needs to use the cluster for data processing, it can initiate a connection to the cluster's master node. When initiating the connection, it includes the system information of the application system, such as the application system's address, connection username and password. After the connection is established, the application system sends the new task to be executed to the master node. After the master node parses the query statement of the new task, if the new task involves the distributed data table to which the target data set belongs, it sends the new task and system information to the first backup node.
[0087] After receiving the system information and the newly added task, the first backup node may establish a connection with the application system according to the system information, so as to directly send the processing result of the newly added task to the application system.
[0088] According to an embodiment of the present application, by synchronously sending system information and newly added tasks to the first backup node, the first backup node can establish a direct connection with the application system. After obtaining the execution result, there is no need to send the execution result to the main node first and then return it to the application system through the main node, which reduces the occupancy of data transmission bandwidth between nodes and improves the processing efficiency of the cluster.
[0089] Figure 5 A schematic diagram of processing a newly added task according to an embodiment of the present application is shown.
[0090] like Figure 5 As shown, using Worker Node 1 as the primary node as an example, the application system establishes a connection with Worker Node 1 by sending system information to Worker Node 1 and then sends a new task after the connection is established. After receiving the new task, Worker Node 1, if it determines that the new task requires execution by the first backup node, sends the new task and system information to the first backup node.
[0091] After receiving the new task and system information, the first backup node can execute the new task to obtain the execution result, and establish a connection with the application system through the system information to return the execution result of the new task directly to the application system without forwarding it to the application system through the working node 1, thereby improving task processing efficiency.
[0092] According to an embodiment of the present application, the data processing method also includes: analyzing the respective working states of the first working node and the first backup node within the target period to obtain the first target processor utilization of the first working node within the target period and the second target processor utilization of the first backup node within the target period; when the difference between the first target processor utilization and the second target processor utilization is greater than a preset difference threshold, screening multiple tasks according to the target processor utilization of each of the multiple tasks executed by the first working node in the target period to obtain the target task; storing the data set required to process the target task to the first backup node, so as to use the first backup node to process the target operation on the data set.
[0093] In an embodiment of the present application, after utilizing the first backup node to share the task of performing the target operation on the target data set, the working status of the first backup node and the first working node can be continuously monitored.
[0094] When the difference between the utilization of the first target processor and the utilization of the second target processor is greater than the preset difference threshold, it indicates that the processor utilization of the first working node is still high and the first backup node still has abundant processor resources. Therefore, the tasks executed by the first working node within the target time period can be analyzed, the target tasks can be further screened out, and the tasks that perform target operations on the data set can be transferred to the first backup node for execution.
[0095] According to an embodiment of the present application, after using the first backup node to process the target operation task for the target data set, the working status of the first backup node and the first working node is monitored, and when the difference between the utilization rate of the first target processor and the second target processor is greater than the preset difference threshold, the first backup node continues to be used to share the tasks that need to be performed by the first working node, thereby improving the overall resource utilization and processing efficiency of the cluster.
[0096] According to an embodiment of the present application, the data processing method also includes: determining multiple idle backup nodes from multiple backup nodes based on the node status of each of the multiple backup nodes in the cluster; and determining the first backup node from the multiple idle backup nodes based on the resource configuration information of each of the multiple idle backup nodes.
[0097] When determining the first backup node, multiple backup nodes in the cluster can be screened based on their node status and resource configuration information. Since tasks with high processor utilization have a lower priority, the first backup node needs to be selected from idle backup nodes based on their node status.
[0098] When determining the first backup node from multiple idle backup nodes, an idle backup node with higher resource configuration may be selected as the first backup node to avoid causing a high processor utilization rate of the first backup node.
[0099] According to an embodiment of the present application, by determining an idle backup node with a higher resource configuration as the first backup node based on the node status and resource configuration information, the overall resource utilization of the cluster can be improved.
[0100] According to an embodiment of the present application, the predetermined historical period includes multiple predetermined historical moments; the working status of each of the multiple working nodes in the cluster within the predetermined historical period is analyzed, and the processor utilization of each of the multiple working nodes within the predetermined historical period is determined, including: obtaining the processor utilization of the working node at multiple predetermined historical moments within the predetermined historical period; based on the average value of the processor utilization at multiple predetermined historical moments, determining the processor utilization of the working node within the predetermined historical period, and obtaining the processor utilization of each of the multiple working nodes.
[0101] When analyzing the working status of the working node in the predetermined historical period, the analysis may be performed based on the processor utilization of the working node at a plurality of predetermined historical moments in the predetermined historical period.
[0102] In an embodiment of the present application, for each working node, the processor utilization of the working node in a predetermined historical period can be determined based on the average value of the processor utilization of the working node at multiple predetermined historical moments.
[0103] According to an embodiment of the present application, by analyzing the processor utilization of the working node at multiple predetermined historical moments, the processor utilization of the working node within the predetermined historical period can accurately characterize the working status of the working node within the predetermined historical period, and then more accurately determine the tasks that need to be shared by the first backup node, thereby improving the processing efficiency of the cluster.
[0104] According to an embodiment of the present application, multiple historical tasks are screened based on the historical processor utilization rates of each of the multiple historical tasks executed by the first working node in a predetermined historical period to obtain a target historical task, including: screening historical tasks for performing a target operation from the multiple historical tasks to obtain at least one candidate historical task, the target operation including at least one of a computing operation and a query operation; determining the target historical task from at least one candidate historical task based on the historical processor utilization rates of each of the at least one candidate historical task.
[0105] When filtering multiple historical tasks, since the target operation usually leads to a high processor utilization of the working node, the multiple historical tasks can be filtered according to the target operation and the historical processor utilization to filter out the target historical task.
[0106] When the target operation is used to filter historical tasks, it can be determined whether the historical task is used to execute the target operation. If the historical task is used to execute the target operation, the historical task is determined to be a candidate task.
[0107] After determining at least one candidate task, the at least one candidate task may be screened according to their respective historical processor utilizations, and a candidate historical task with a higher historical processor utilization may be determined as a target historical task.
[0108] According to an embodiment of the present application, by utilizing target operations and historical processor utilization to implement screening of historical tasks, the target historical tasks that cause the processor utilization of the first working node to be higher can be accurately determined, thereby enabling the processing efficiency of the cluster to be improved when the first backup node is utilized to share tasks related to the target historical tasks.
[0109] According to an embodiment of the present application, a target historical task is determined from at least one candidate historical task based on the historical processor utilization of each of the at least one candidate historical tasks, including: sorting at least one candidate historical task in descending order based on the historical processor utilization of each of the at least one candidate historical task to obtain a candidate historical task sequence; and determining a candidate historical task at a predetermined position in the candidate historical task sequence as the target historical task.
[0110] When screening at least one candidate historical task, at least one candidate historical task can be sorted in descending order according to historical processor utilization to obtain a candidate historical task sequence, and then a candidate task with a higher ranking can be selected from the candidate historical task sequence as a target historical task.
[0111] According to an embodiment of the present application, by sorting at least one candidate historical task in descending order and selecting a top-ranked candidate task from the candidate historical task sequence as the target historical task, the accuracy of determining the target historical task can be improved, thereby improving the processing efficiency of the cluster.
[0112] According to an embodiment of the present application, the data processing method also includes: when there is a second working node that fails among multiple working nodes and there is no idle backup node in the cluster, determining a second backup node from multiple backup nodes; in response to the second working node and the second backup node having the same target storage data, deleting other storage data except the target storage data from the second backup node; mounting the external storage volume of the second working node to the second backup node, so that the second backup node loads other storage data stored in the second working node except the target storage data from the external storage volume, and switching the data processing tasks executed by the second working node to the second backup node for execution.
[0113] Since backup nodes are usually used to replace failed working nodes to perform data processing tasks, in order to ensure the stability of the cluster, the priority of backup nodes to perform data processing tasks instead of failed working nodes should be higher than the priority of processing tasks with high processor utilization.
[0114] When there is a second working node that has failed, a backup node needs to be used to replace the second working node to perform the data processing task. When determining the backup node, it can be selected from multiple idle backup nodes.
[0115] If there are no idle backup nodes in the cluster, the second backup node, which handles tasks with high processor utilization, needs to be used to replace the second working node. Because the second backup node already has data stored when it was used to share tasks with the working node with high processor utilization, the data stored in the second backup node needs to be cleared before it is used to replace the second working node.
[0116] When cleaning the storage data in the second backup node, since the second backup node may be used to share the tasks of the second working node, the second backup node and the second working node may store the same target storage data. Therefore, the target storage data in the second backup node can be retained and other storage data in the second backup node except the target storage data can be deleted.
[0117] After cleaning the second backup node, the other storage data of the second working node can be stored in the second backup node. Since the second working node has failed, data may no longer be normally retrieved from the memory of the second working node. In this case, the external storage volume of the second working node can be mounted to the second backup node, so that the second backup node can load other storage data stored in the second working node except the target storage data from the external storage volume, thereby storing all the data of the second working node in the second backup node.
[0118] After completing data loading, the second backup node can send a loading success instruction to the scheduling node. At this time, the scheduling node can modify the node information of the second backup node and no longer use the second backup node to share tasks with high processor utilization. Instead, the data processing tasks executed by the second working node will be switched to the second backup node for execution.
[0119] According to an embodiment of the present application, by retaining the target storage data shared with the second working node when storing data on the second backup node, the amount of data that needs to be loaded is reduced, the loading time is reduced, and the processing efficiency of the cluster is improved.
[0120] Figure 6 A schematic diagram of a second backup node loading data according to an embodiment of the present application is shown.
[0121] like Figure 6As shown, taking working node 1 as the faulty second working node as an example, since both the second backup node and working node 1 store table B dataset 1, when cleaning the second backup node, the table B dataset 1 stored in the second backup node can be retained.
[0122] like Figure 6 As shown, after the second backup node mounts the external storage volume of the working node 1, the table A data set 1 originally stored in the working node 1 can be stored in the second backup node. At this time, the second backup node stores all the storage data originally stored in the working node 1.
[0123] Based on the above data processing method, this application also provides a data processing device. Figure 7 The device is described in detail.
[0124] Figure 7 The figure shows a structural block diagram of a data processing device according to an embodiment of the present application.
[0125] like Figure 7 As shown, the data processing device 700 of this embodiment includes an analyzing module 710 , a screening module 720 , a storage module 730 and a sending module 740 .
[0126] The analysis module 710 is used to analyze the working status of each of the plurality of working nodes in the cluster within a predetermined historical period, and determine the processor utilization of each of the plurality of working nodes within the predetermined historical period. In one embodiment, the analysis module 710 can be used to perform the operation S210 described above, which will not be repeated here.
[0127] The screening module 720 is configured to, in response to a first work node among the multiple work nodes having a processor utilization greater than a preset threshold, screen the multiple historical tasks based on the historical processor utilization of each of the multiple historical tasks executed by the first work node during a predetermined historical period to obtain a target historical task. In one embodiment, the screening module 720 can be configured to perform operation S220 described above, which will not be further described here.
[0128] The storage module 730 is used to store the target data set required for processing the target historical task to the first backup node, which is determined from multiple idle backup nodes in the cluster. In one embodiment, the storage module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0129] The sending module 740 is configured to, in response to receiving the newly added task and, if it is determined that the newly added task is for performing the target operation on the target dataset, send the newly added task to the first backup node so that the first backup node can process the newly added task. In one embodiment, the sending module 740 can be configured to perform operation S240 described above, which will not be further described here.
[0130] According to an embodiment of the present application, the data processing device 700 further includes a first generating module and a first changing module.
[0131] The first generation module is used to update information of the first data change task in response to receiving a first data change task for performing a change operation on the target data set during the process of storing the target data set required for processing the target historical task on the first backup node, and generate a to-be-executed task with identification information indicating that the task is not yet executed.
[0132] The first change module is used to send the task to be executed to the first working node, so that after the target data set is completely stored in the first backup node, the first working node updates the first target data to be updated in the task to be executed based on the changed data in the task to be executed.
[0133] According to an embodiment of the present application, the data processing device 700 further includes a first parsing module and a change determination module.
[0134] The first parsing module is used to parse the query statement included in the first data processing task in response to receiving the first data processing task, and determine the first operation type of the first operation to be performed and the first data identifier of the first data set to be operated on by the first operation to be performed.
[0135] The change determination module is used to determine, based on the first data identifier and the first operation type, whether the first data processing task is a first data change task for performing a change operation on the target data set.
[0136] According to an embodiment of the present application, the data processing device 700 further includes a second generating module and a second changing module.
[0137] The second generating module is configured to generate a second data change task for the first backup node based on the data change task in response to the target data set required for processing the target historical task being completely stored in the first backup node.
[0138] The second change module is used to send the second data change task to the first backup node, so that the first backup node updates the first target data stored in the first backup node to the changed data in response to the second data change task.
[0139] According to an embodiment of the present application, the sending module 740 includes a sending submodule.
[0140] The sending submodule is used to send the system information of the newly added task and the application system that initiates the newly added task to the first backup node, so that the first backup node can establish a connection with the application system according to the system information and send the processing result of the newly added task to the application system.
[0141] According to an embodiment of the present application, the data processing device 700 further includes a second parsing module, a task identification module, and a change sending module.
[0142] The second parsing module is used to parse the query statement included in the second data processing task in response to receiving the second data processing task when the target data set required for processing the target historical task has been completely stored in the first backup node, and to determine the second operation type of the second to-be-executed operation and the second data identifier of the second to-be-executed data set for which the second to-be-executed operation is targeted.
[0143] The task identification module is used to determine an identification result for the second data processing task based on the second data identifier and the second operation type.
[0144] The change sending module is used to send the second data processing task to the first working node and the first backup node when it is determined that the recognition result represents that the second data processing task is used to perform a change operation on the target data set, so as to use the first working node and the first backup node to process the second data processing task respectively.
[0145] According to an embodiment of the present application, the data processing device 700 further includes a table determination module and a table storage module.
[0146] The table determination module is used to determine the distributed data table to which the target data set belongs.
[0147] The table storage module is used to store the distributed data table in the first backup node according to the data table identifier of the distributed data table.
[0148] According to an embodiment of the present application, the data processing device 700 further includes a backup determination module, a backup deletion module, and a backup loading module.
[0149] The backup determination module is used to determine a second backup node from multiple backup nodes when there is a second working node that fails among the multiple working nodes and there is no idle backup node in the cluster.
[0150] The backup deletion module is configured to delete other storage data except the target storage data from the second backup node in response to the second working node and the second backup node having the same target storage data.
[0151] The backup loading module is used to mount the external storage volume of the second working node to the second backup node, so that the second backup node can load other storage data stored in the second working node except the target storage data from the external storage volume, and switch the data processing tasks executed by the second working node to the second backup node for execution.
[0152] According to an embodiment of the present application, the data processing device 700 further includes an idle determination module and a resource determination module.
[0153] The idle determination module is used to determine a plurality of idle backup nodes from a plurality of backup nodes according to respective node states of a plurality of backup nodes in the cluster.
[0154] The resource determination module is used to determine a first backup node from a plurality of idle backup nodes according to respective resource configuration information of the plurality of idle backup nodes.
[0155] According to an embodiment of the present application, the predetermined historical period includes a plurality of predetermined historical moments; the analysis module 710 includes an acquisition submodule and an analysis submodule.
[0156] The acquisition submodule is used to acquire the processor utilization of the working node at multiple predetermined historical moments within a predetermined historical period.
[0157] The analysis submodule is used to determine the processor utilization of the working node within a predetermined historical period based on the average value of the processor utilization at multiple predetermined historical moments, and obtain the processor utilization of each of the multiple working nodes.
[0158] According to an embodiment of the present application, the screening module 720 includes a first screening submodule and a second screening submodule.
[0159] The first screening submodule is used to screen historical tasks for executing a target operation from a plurality of historical tasks to obtain at least one candidate historical task, where the target operation includes at least one of a computing operation and a query operation.
[0160] The second screening submodule is configured to determine a target historical task from the at least one candidate historical task according to the historical processor utilization of each of the at least one candidate historical task.
[0161] According to an embodiment of the present application, the second screening submodule includes a sorting unit and a determining unit.
[0162] The sorting unit is used to sort the at least one candidate historical task in descending order according to the historical processor utilization of each of the at least one candidate historical task to obtain a candidate historical task sequence.
[0163] The determining unit is configured to determine a candidate historical task at a predetermined position in the candidate historical task sequence as a target historical task.
[0164] According to embodiments of the present application, any multiple modules among the analysis module 710, screening module 720, storage module 730, and sending module 740 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the analysis module 710, screening module 720, storage module 730, and sending module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of the analysis module 710, screening module 720, storage module 730, and sending module 740 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0165] Figure 8 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present application is shown.
[0166] like Figure 8 As shown, an electronic device 800 according to an embodiment of the present application includes a processor 801, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 802 or a program loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0167] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in one or more memories.
[0168] According to an embodiment of the present application, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.
[0169] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0170] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0171] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the data processing method provided in the embodiments of the present application.
[0172] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the processor 801 executes the computer program. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0173] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0174] In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0175] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0177] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0178] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Analyze the working status of each of the plurality of working nodes in the cluster within a predetermined historical period, and determine the processor utilization rate of each of the plurality of working nodes within the predetermined historical period; In response to a first working node having a processor utilization greater than a preset threshold among the plurality of working nodes, the plurality of historical tasks are screened according to respective historical processor utilizations of the plurality of historical tasks executed by the first working node during the predetermined historical period to obtain a target historical task; storing a target data set required for processing the target historical task in a first backup node, where the first backup node is determined from a plurality of idle backup nodes in the cluster; In response to receiving the newly added task, if it is determined that the newly added task is used to perform a target operation on the target data set, the newly added task is sent to the first backup node, so that the first backup node is used to process the newly added task.
2. The method according to claim 1, characterized in that The method further comprises: In the process of storing the target data set required for processing the target historical task in the first backup node, in response to receiving a first data change task for performing a change operation on the target data set, updating information of the first data change task to generate a pending task with identification information indicating that the task is not yet executed; The task to be executed is sent to the first working node, so that after the target data set is completely stored in the first backup node, the first working node updates the first target data to be updated in the task to be executed based on the changed data in the task to be executed.
3. The method according to claim 2, characterized in that The method further comprises: In response to receiving the first data processing task, parsing the query statement included in the first data processing task to determine a first operation type of a first to-be-performed operation and a first data identifier of a first to-be-operated data set targeted by the first to-be-performed operation; Based on the first data identifier and the first operation type, it is determined whether the first data processing task is a first data change task for performing a change operation on the target data set.
4. The method according to claim 2, characterized in that The method further comprises: In response to the target data set required for processing the target historical task being completely stored in the first backup node, generating a second data change task for the first backup node based on the first data change task; The second data change task is sent to the first backup node, so that the first backup node updates the first target data stored in the first backup node to the changed data in response to the second data change task.
5. The method according to claim 1, wherein The sending the newly added task to the first backup node so as to use the first backup node to process the newly added task includes: The newly added task and the system information of the application system that initiated the newly added task are sent to the first backup node, so that the first backup node establishes a connection with the application system according to the system information and sends the processing result of the newly added task to the application system.
6. The method according to claim 1, characterized in that The method further comprises: When the target data set required for processing the target historical task has been completely stored in the first backup node, in response to receiving a second data processing task, parsing a query statement included in the second data processing task to determine a second operation type of a second to-be-executed operation and a second data identifier of a second to-be-operated data set targeted by the second to-be-executed operation; Determining an identification result for the second data processing task based on the second data identifier and the second operation type; When it is determined that the recognition result represents that the second data processing task is used to perform a change operation on the target data set, the second data processing task is sent to the first working node and the first backup node, so that the first working node and the first backup node can respectively process the second data processing task.
7. The method according to claim 1, characterized in that The method further comprises: Determining the distributed data table to which the target data set belongs; According to the data table identifier of the distributed data table, the distributed data table is stored in the first backup node.
8. The method according to claim 1, characterized in that The method further comprises: If a second working node that fails exists among the multiple working nodes and there is no idle backup node in the cluster, determining a second backup node from the multiple backup nodes; In response to the second working node and the second backup node having the same target storage data, deleting other storage data except the target storage data from the second backup node; Mount the external storage volume of the second working node to the second backup node so that the second backup node can load other storage data stored in the second working node except the target storage data from the external storage volume, and switch the data processing tasks executed by the second working node to the second backup node for execution.
9. The method according to claim 1, characterized in that The method further comprises: Determining a plurality of idle backup nodes from the plurality of backup nodes according to respective node states of the plurality of backup nodes in the cluster; The first backup node is determined from the plurality of idle backup nodes according to respective resource configuration information of the plurality of idle backup nodes.
10. The method according to claim 1, characterized in that The predetermined historical period includes a plurality of predetermined historical moments; and the analyzing the working status of each of the plurality of working nodes in the cluster within the predetermined historical period to determine the processor utilization of each of the plurality of working nodes within the predetermined historical period includes: Obtaining processor utilization of the working node at a plurality of predetermined historical moments within the predetermined historical period; Based on an average value of the processor utilizations at the plurality of predetermined historical moments, the processor utilizations of the working nodes within the predetermined historical period are determined to obtain the processor utilizations of the respective working nodes.
11. The method according to claim 1, wherein The filtering of the plurality of historical tasks according to the respective historical processor utilizations of the plurality of historical tasks executed by the first working node in the predetermined historical period to obtain the target historical task includes: Filtering historical tasks for executing the target operation from the plurality of historical tasks to obtain at least one candidate historical task, wherein the target operation includes at least one of a computing operation and a query operation; The target historical task is determined from at least one of the candidate historical tasks according to the historical processor utilization rate of each of the at least one candidate historical tasks.
12. The method according to claim 11, characterized in that Determining the target historical task from at least one of the candidate historical tasks according to the historical processor utilization of each of the at least one candidate historical tasks includes: sorting the at least one candidate historical task in descending order according to the historical processor utilization of each of the at least one candidate historical task to obtain a candidate historical task sequence; A candidate historical task at a predetermined position in the candidate historical task sequence is determined as the target historical task.
13. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Node selection method and node selection device
CN113641453A
Data processing method and device, server, storage medium and program product
CN116737450A
Database data processing method and device, equipment and storage medium
CN120492231A
System and method for crash-consistent incremental backup of cluster storage
US20200104202A1