Task elastic arrangement and dynamic adaptation method for data governance

Through the task elastic orchestration method combined with Lasso algorithm and Raft consensus algorithm, the problem of resource allocation imbalance in data governance is solved, the task operation and exception handling are realized on the optimal hardware node, and task execution efficiency and resource utilization are improved.

CN120407131AActive Publication Date: 2025-08-01YANTAI JIERUI NETWORK TRADING
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510905298.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing data governance process orchestration cannot perceive the dynamic resource requirements of tasks and the changes in the load of computing equipment in real time, resulting in unbalanced resource allocation, inefficient task execution, increasing operational costs, and affecting the timeliness of data processing and the scientific nature of business decisions.

Method used

The Lasso algorithm is used to train the computing model, combine Raft consensus algorithm and heartbeat detection technology to monitor the hardware resource status in real time, perform task preselecting, filtering and preferential sorting, ensure that the task runs on the optimal hardware node, and restart or migrate when the task is abnormal.

Benefits of technology

It realizes the accuracy and timeliness of task resource allocation, reduces idle equipment resources, improves execution efficiency and resource utilization, and meets the enterprise's efficient data governance needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407131A_ABST
    Figure CN120407131A_ABST
Patent Text Reader

Abstract

The invention discloses a data governance-oriented task elastic arrangement and dynamic adaptation method, which comprises the following steps of S1, collecting task node historical records, training a model by using a Lasso algorithm, and analyzing the influence of resource indexes on task duration; s2, expanding a task model, and introducing a new resource index and a timeliness attribute; s3, constructing a resource model to monitor hardware equipment, and ensuring data consistency by using a Raft algorithm; s4, pre-selecting filtering nodes, calculating expected duration ranking based on the model, and selecting an optimal node; s5, monitoring the health state of the task and adopting a retry strategy; and S6, optimizing the model according to the actual operation duration. Through accurate data collection and model training, resource monitoring and task scheduling optimization are combined, the task execution efficiency and the resource utilization rate are improved, and reliable operation of tasks is ensured through a fault-tolerant mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, and particularly to a task elastic orchestration and dynamic adaptation method for data governance. Background Art

[0002] In the process of accelerating the digital transformation of enterprises, the data scale has increased exponentially, rapidly jumping from the GB level to the PB level or even higher. At the same time, data governance covers complex business scenarios such as metadata verification and data encryption, and the resource requirements of different tasks vary significantly. For example, complex data analysis tasks highly rely on CPU computing power, data warehouse update tasks have strict requirements for storage read and write performance, and real-time transaction data processing is extremely sensitive to timeliness.

[0003] Currently, the process orchestration of data governance often uses the graphical component dragging method. Although it simplifies the process construction, the current task scheduling center distributes resources in the task allocation link according to the device number order or evenly. This allocation mode cannot real-time perceive the dynamic resource requirements of tasks and the changes in the load of computing devices, resulting in an imbalance in resource configuration, and the coexistence of idle device resources and low task execution efficiency, which not only increases the enterprise operation cost, but also seriously affects the timeliness of data processing and the scientificity of business decisions. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art that the resource requirements of tasks cannot be accurately matched with the real-time state of computing devices, resulting in low task execution efficiency, serious resource waste, insufficient timeliness of task response, and inability to meet the high-efficiency data governance requirements of enterprises, and to propose a task elastic orchestration and dynamic adaptation method for data governance.

[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions: A task elastic orchestration and dynamic adaptation method for data governance, including the following steps: S1: Data collection and model training: Collect the historical execution records of task nodes, and the collected indicators cover CPU usage, CPU remaining, total CPU, memory usage, memory remaining, total memory, storage rate, hard disk type, bandwidth usage, and task execution duration. Use the Lasso algorithm to analyze and process the collected data, and train a calculation model of CPU, memory, storage, hard disk type, and bandwidth for task duration; S2: Task model expansion: Expand the resource indicators of CPU, memory, hard disk type, external network address, and the timeliness attribute of the task; S3: Resource Monitoring: Build a resource model to store the CPU, memory, storage, disk type, bandwidth usage, and maximum values of each device metric for each task reported in real time by hardware devices. Use the Raft consensus algorithm to ensure strong data consistency across all nodes. S4: Pre-selection, filtering, and optimal sorting: Based on the resource requirements of the nodes in S2 and the hardware resource status, anti-affinity, and resource reservation mechanisms in S3, all nodes that do not meet the requirements are filtered out. Based on the model data extended by S2, the calculation model of S1 is used on the filtered hardware nodes to calculate the expected runtime of each hardware node and prioritize them. Tasks are run on the optimal hardware nodes, and a data model is created to store the relationship between hardware devices and running tasks. S5: Health Check and Fault Tolerance: This section establishes network communication with tasks running in S4, regularly monitors their health status, and records tasks whose runtimes far exceed the expected runtime in S4. It then applies different retry strategies based on the task's health status and the timeliness properties extended by S2. S6: Model optimization: When the number of tasks whose actual running time recorded in S5 far exceeds the expected running time reaches a threshold, the Lasso model is retrained using recent historical data.

[0006] As a further technical solution of the present invention, S1 specifically includes: S11: Create a task collection module to collect the resource usage and task execution time data of historical task execution; S12: For missing values in the data, the mean is used to fill in the missing values, and outliers are removed to ensure data quality. The preprocessed data is divided into training set and test set in a ratio of 7:3; S13: The model is trained based on the Lasso algorithm, and key features are automatically screened by adjusting the regularization parameters. Finally, the model generalization ability is verified using a test set. The impact of various factors on task duration is analyzed based on the feature coefficients to form a calculation model.

[0007] As a further technical solution of the present invention, the S13 specifically includes: S131: Define the loss function as: , in: is the actual task duration of the i-th sample, is the transpose of the eigenvector of the i-th sample, is the feature coefficient vector, n is the number of training set samples, p is the number of features, is the regularization parameter; by adjusting the regularization parameter , so that the model minimizes the loss function on the training set. When it is relatively large, more feature coefficients will be reduced to 0, thus automatically screening out key features; when it is relatively small, the model retains more features; S132: Determine the optimal regularization parameter using the cross-validation method , specifically, divide the training set into k subsets, sequentially select one subset as the validation set, and the remaining subsets as the training subsets. Train the model for each candidate value and calculate the mean squared error on the validation set. Finally, select the value corresponding to the minimum mean squared error as the optimal regularization parameter; S133: Use the test set to validate the trained model, and calculate the coefficient of determination between the predicted task duration and the actual task duration of the model . If the value is greater than the set threshold (such as 0.8), it is determined that the model has good generalization ability; S134: Analyze the influence of each factor on the task duration according to the magnitude and sign of the feature coefficient . The factor with a larger absolute value of the feature coefficient has a greater influence on the task duration. A positive feature coefficient indicates that the factor will extend the task duration, and a negative feature coefficient indicates that the factor will shorten the task duration. Finally, integrate these analysis results to form a calculation model.

[0008] As a further technical solution of the present invention, the specific steps of S2 are as follows: S21: For the CPU type, define its performance index as , through the formula , where is the main frequency of the CPU, is the CPU architecture coefficient, reflecting the performance differences of different CPU architectures, and convert the original CPU type into a standardized performance index; S22: For the memory type, define the memory efficiency index , where is the memory capacity, is the memory read / write speed, reflecting the comprehensive performance of memory capacity and speed; S23: For the hard disk type, define the hard disk efficiency index , where is the write speed, is the read speed, is the storage capacity, used to measure the comprehensive performance of the hard disk; S24: For the external network address resource index, define the network efficiency index , where is the network bandwidth, The packet loss rate comprehensively reflects the transmission capacity and stability of the network; S25: For the timeliness attribute of the task, define the timeliness index ,in is the task completion time, is the task start time, The task deadline is used to measure whether the task execution time is close to the deadline, thereby reflecting the timeliness of the task.

[0009] As a further technical solution of the present invention, S3 specifically includes: S31: Design a standardized data structure to describe resource status, including core attributes such as device identification, indicator type (CPU / memory, etc.), real-time value, and timestamp to store device information; S32: Obtains the CPU, memory, hard disk read / write speed, hard disk type, bandwidth, and latency resource usage data of multiple external network addresses in S2 for each task in real time on the hardware device side, as well as the maximum value of each device indicator, and packages and reports the data in a unified format. S33: Deploy the Raft algorithm on the node that stores the hardware resource status, elect the master node, and establish a heartbeat detection mechanism between nodes to ensure the synchronization of multiple node states and strong consistency of node data.

[0010] As a further technical solution of the present invention, the S33 specifically includes: S331: Master node election of Raft algorithm: Raft algorithm divides nodes into three states: Follower, Candidate and Leader. Initially, all nodes are in Follower state and define timeout parameters. The time interval for a node to transition from the Follower state to the Candidate state. If the Follower node does not receive the heartbeat packet from the Leader node within the specified time, an election is triggered and the node switches to the Candidate state. The Candidate node initiates an election request to other nodes and defines the voting parameters. , when the Candidate node obtains more than half of the votes ( is the total number of nodes), the Leader node is successfully elected; S332: Heartbeat detection mechanism between nodes: The leader node periodically sends heartbeat packets to the follower node, defining the heartbeat interval parameters , if the Follower node is If no leader heartbeat packet is received within the specified time, the leader is considered invalid and the election process is restarted. S333: Data Synchronization and Strong Consistency: Define the set of log entries , each log entry contains an index and a term , the Leader node replicates the new log to the Follower nodes. Through the majority commitment mechanism, it ensures that the log is committed only after it has been successfully replicated on the majority of nodes, guaranteeing strong data consistency; define the majority parameter , when a certain log entry has been successfully replicated by at least nodes, it is considered that the log has been committed.

[0011] As a further technical solution of the present invention, the S4 specifically includes: S41: Build a data model to store operation records, including hardware device markers and running task identifiers; S42: Process the CPU and memory of a single task in the devices reported in S32. For hard real-time tasks, calculate according to the sum of the running resources and the reserved resources, reserve reasonable resources for hard real-time tasks, and calculate the remaining CPU and memory of the device; S43: Filter the device margins obtained in S42 according to the CPU, memory, hard disk type, and external network address required by the task nodes in S2. For hard real-time tasks, calculate the CPU and memory according to the sum of the required running resources and the reserved resources, filter out the hardware devices that do not meet the requirements, and trigger an email alarm if there are no available devices; S44: Perform anti-affinity check on the hardware devices obtained in S43. When the number of hard real-time tasks running on a certain hardware device is greater than a certain threshold of the total number of hard real-time tasks, this node does not participate in the distribution of hard real-time tasks to prevent hard real-time tasks from being concentrated on the same hardware device; S45: Use the calculation model obtained in S13 to calculate the estimated running time for the hardware devices that meet the requirements in S44, and select the hardware device with the shortest running time as the task running carrier.

[0012] As a further technical solution of the present invention, the S5 specifically includes: S51: Build a data model to store abnormal operation records, including task identifiers and the number of abnormalities; S52: Establish a network connection with the running tasks to regularly check the health status and the actual running duration. If the running duration of the task exceeds a certain proportion of the estimated running time obtained in S45, it is saved as an abnormal record in the data model of S51; S53: When the task status obtained according to S52 is a failure, based on the timeliness attribute extended according to S2, directly terminate and mark the failure for hard real-time tasks when a failure occurs, and support 3 retries for soft real-time tasks and batch tasks, and select different nodes for each retry.

[0013] As a further technical solution of the present invention, the specific steps of S6 are as follows: S61: Regularly scan the data model of S51 to obtain the number of task exceptions; S62: When the number of task exceptions exceeds the threshold, re-train the calculation model using the Lasso algorithm according to S1.

[0014] The beneficial effects of the present invention are as follows: 1. By collecting the historical execution records of task nodes and training the calculation model using the Lasso algorithm, accurately analyze the execution of tasks in different hardware and network environments, provide a quantitative basis for task allocation, and change the situation where the traditional scheduling center cannot perceive the dynamic resource requirements of tasks in real time.

[0015] 2. In the process of task allocation, comprehensively consider factors such as task resource requirements, hardware resource status, anti-affinity, and resource reservation mechanism, perform pre-selection filtering and optimal sorting, and run the task on the optimal hardware node to avoid resource configuration imbalance, reduce equipment resource idleness, improve task execution efficiency and resource utilization rate, ensure the supply of key task resources, and solve the disadvantages of resource configuration imbalance and equipment resource idleness existing in the current scheduling mechanism.

[0016] 3. Use heartbeat detection and fault tolerance technology to monitor the task running status in real time. When an exception occurs in the task, be able to trigger restart or migration operations in a timely manner according to the task timeliness attribute to ensure that the task can respond and process in a timely manner, meeting the requirements of enterprise high-efficiency data governance for timeliness. Description of the Drawings

[0017] Figure 1 It is a flowchart of a task elastic orchestration and dynamic adaptation method for data governance proposed by the present invention. Detailed Embodiment

[0018] To make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0019] Please refer to the attached Figure 1 , A task elastic orchestration and dynamic adaptation method for data governance, including the following steps: S1: Data collection and model training: Collect the historical execution records of task nodes. The collected metrics include CPU usage, remaining CPU, total CPU, memory usage, remaining memory, total memory, storage rate, hard disk type, bandwidth usage, and task execution duration. Use the Lasso algorithm to analyze and process the collected data, and train a calculation model for the impact of CPU, memory, storage, hard disk type, and bandwidth on task duration. S11: Create a task collection module to collect data on resource occupancy during historical task execution and task execution duration. S12: Fill in missing values in the data with the mean, and at the same time remove outliers to ensure data quality. Divide the preprocessed data into a training set and a test set in a 7:3 ratio. S13: Train the model based on the Lasso algorithm, and automatically screen key features through regularization parameter adjustment; finally, use the test set to verify the generalization ability of the model, and analyze the impact of each factor on task duration according to the feature coefficients to form a calculation model. S131: Define the loss function as: , where: is the actual task duration of the i-th sample, is the transpose of the feature vector of the i-th sample, is the feature coefficient vector, n is the number of training set samples, p is the number of features, is the regularization parameter; by adjusting the regularization parameter , the model minimizes the loss function on the training set. When is large, more feature coefficients will be shrunk to 0, thus automatically screening out key features; when is small, the model retains more features. S132: Use the cross-validation method to determine the optimal regularization parameter , specifically divide the training set into k subsets, and successively select one of them as the validation set, and the remaining subsets as the training subsets. Train the model for each candidate value and calculate the mean squared error on the validation set. Finally, select the value corresponding to the minimum mean squared error as the optimal regularization parameter. S133: Use the test set to verify the trained model, and calculate the determination coefficient between the predicted task duration and the actual task duration of the model. If value is greater than the set threshold (such as 0.8), it is determined that the model has good generalization ability. S134: According to the feature coefficients Analyze the impact of each factor on the task duration based on the magnitude and sign of the

[0020] S2: Task model expansion: Expand the CPU, memory, hard disk type, external network address resource metrics, and the timeliness attributes of the task; S21: For the CPU type, define its performance metric as , through the formula , where is the CPU main frequency, is the CPU architecture coefficient, reflecting the performance differences of different CPU architectures, and convert the original CPU type into a standardized performance metric; S22: For the memory type, define the memory efficiency metric , where is the memory capacity, is the memory read and write speed, reflecting the comprehensive performance of the memory capacity and speed; S23: For the hard disk type, define the hard disk efficiency metric , where is the write speed, is the read speed, is the storage capacity, used to measure the comprehensive performance of the hard disk; S24: For the external network address resource metrics, define the network efficiency metric , where is the network bandwidth, is the packet loss rate, comprehensively reflecting the transmission ability and stability of the network; S25: For the timeliness attributes of the task, define the timeliness metric , where is the task completion time, is the task start time, is the task deadline, used to measure whether the task execution time is close to the deadline, thus reflecting the timeliness of the task.

[0021] S3: Resource monitoring: Build a resource model to store the CPU, memory, storage, hard disk type, bandwidth usage of each task reported by the hardware device in real time, and the maximum values of each device metric, and use the Raft consensus algorithm to ensure strong consistency of data on all nodes; S31: Design a standardized data structure to describe the resource status, including core attributes such as device identifier, metric type (CPU / memory, etc.), real-time value, and timestamp for storing device information; S32: Obtain the CPU, memory, hard disk read / write speed, hard disk type, bandwidth, and latency resource usage data of each task and the maximum values of various device metrics in real time on the hardware device side, and encapsulate and report the data in a unified format; S33: Deploy the Raft algorithm at the hardware resource status saving node, elect the master node, and establish a heartbeat detection mechanism between nodes to ensure the state synchronization of multiple nodes and the strong consistency of node data; S331: Master node election of the Raft algorithm: The Raft algorithm divides nodes into three states: Follower, Candidate, and Leader. Initially, all nodes are in the Follower state. Define the timeout parameter as the time interval for a node to transition from the Follower state to the Candidate state. If a Follower node does not receive a heartbeat packet from the Leader node within this time, an election is triggered and the node transitions to the Candidate state; The Candidate node sends an election request to other nodes and defines the voting parameter . When the Candidate node obtains more than half of the votes ( is the total number of nodes), it is successfully elected as the Leader node; S332: Heartbeat detection mechanism between nodes: The Leader node periodically sends heartbeat packets to Follower nodes and defines the heartbeat interval parameter . If a Follower node does not receive a Leader heartbeat packet within this time, the Leader is considered to have failed and the election process is restarted; S333: Data synchronization and strong consistency: Define the set of log entries . Each log entry contains an index and a term . The Leader node replicates new logs to Follower nodes. Through the majority commitment mechanism, ensure that the logs are committed only after they are successfully replicated on the majority of nodes, guaranteeing strong data consistency; Define the majority parameter . When a certain log entry is successfully replicated by at least nodes, the log is considered to have been committed.

[0022] S4: Pre-selection filtering and optimal sorting: According to the resource requirements of the nodes in S2, the hardware resource status in S3, as well as the anti-affinity and resource reservation mechanisms, filter out all nodes that do not meet the requirements. Based on the model data expanded in S2, use the calculation model in S1 for the filtered hardware nodes, calculate the expected running duration of each hardware node respectively and perform optimal sorting, run the task on the optimal hardware node, and create a data model to store the relationship between the hardware device and the running task; S41: Build a data model to store the running records, including the hardware device label and the running task identifier; S42: Process the CPU and memory of a single task in the devices reported in S32. For hard real-time tasks, calculate according to the sum of the running resources and the reserved resources, reserve reasonable resources for hard real-time tasks, and calculate the remaining CPU and memory of the device; S43: Screen the device margins obtained in S42 according to the CPU, memory, hard disk type, and external network address required by the task nodes in S2. For hard real-time tasks, the CPU and memory are calculated according to the sum of the required running resources and the reserved resources. Filter out the hardware devices that do not meet the requirements. If there are no available devices, trigger an email alarm; S44: Perform anti-affinity check on the hardware devices obtained in S43. When the number of hard real-time tasks running on a certain hardware device is greater than a certain threshold of the total number of hard real-time tasks, this node does not participate in the distribution of hard real-time tasks to prevent hard real-time tasks from being concentrated on the same hardware device; S45: Use the calculation model obtained in S13 to calculate the expected running time for the hardware devices that meet the requirements in S44, and select the hardware device with the shortest running time as the task running carrier.

[0023] S5: Health detection and fault tolerance guarantee: Establish network communication with the tasks running in S4, regularly monitor the health status, record the tasks whose running time far exceeds the expected running duration in S4, and adopt different retry strategies for the tasks according to the task health status and the timeliness attributes expanded in S2; S51: Build a data model to store abnormal running records, including task identifiers and the number of abnormalities; S52: Establish a network connection with the running tasks to regularly check the health status and the actual running duration. If the running duration of the task exceeds a certain proportion of the expected running time obtained in S45, save it as an abnormal record in the data model of S51; S53: When the task status obtained in S52 is failure, according to the timeliness attributes expanded in S2, directly terminate and mark as failed for hard real-time task failures, and support 3 retries for soft real-time tasks and batch processing tasks, and select different nodes for each retry.

[0024] S6: Model Optimization: When the number of tasks whose actual running time recorded in S5 far exceeds the expected running time reaches the threshold, retrain the Lasso model using recent historical data; S61: Regularly scan the data model in S51 to obtain the number of task anomalies; S62: When the number of task anomalies exceeds the threshold, retrain the calculation model using the Lasso algorithm according to S1.

[0025] Lasso Regression Algorithm: An algorithm that performs feature selection and parameter optimization on CPU resource parameters, memory resource parameters, network environment parameters, disk read and write queue depth, average response time, disk type, etc. by introducing an L1 regularization term into the loss function, and is used to construct a weight calculation model to evaluate the execution of tasks in different hardware and network environments; Raft Consensus Algorithm: An algorithm used to achieve consistency in a distributed system. It is mainly used to reach an agreement on the order of a series of operations among multiple nodes to ensure that the data in the distributed system remains consistent on each node. Even in the face of node failures, network delays, or message losses, it can ensure the normal operation of the system and data consistency, thus providing a reliable foundation for the distributed system.

[0026] Heartbeat Detection: A mechanism used to monitor whether a node is alive. The node will periodically send a special heartbeat packet to the scheduling center as a survival signal at a set time interval (e.g., in this solution, if the node fails to send a signal 3 times within the timeout period, that is, no signal within 1.5 seconds, it is determined as a failure).

[0027] Task Timeliness: Hard Real-Time Tasks: Such tasks must be completed within a fixed time. Soft Real-Time Tasks: Given priority but allowing a certain delay. Batch Processing Tasks: Can be executed during off-peak hours.

[0028] Example 1 The language used is the Java language. A method for elastic scheduling and dynamic adaptation of data governance-oriented tasks includes the following steps: S1: Data Collection and Model Training: Use a Java background development task collection module. This module collects the historical execution records of task nodes from system logs and task execution records through data interaction with the existing data governance task scheduling system; the collected metrics cover CPU usage, remaining amount, total amount, memory usage, remaining amount, total amount, storage rate, hard disk type, bandwidth usage, and task execution duration, etc.; during the data collection process, multi-threading technology is used to improve the data collection efficiency, ensuring that a large amount of historical data can be quickly obtained and saved to the data storage module for subsequent analysis; The collected data was preprocessed using the Java backend. First, code logic was written to remove data indicating task execution failures or task response times exceeding normal values. Next, the mean of various indicators was calculated for each task node. Missing values were then filled with the mean to ensure data integrity and quality. After data cleaning and filling, the preprocessed data was divided into training and test sets in a 7:3 ratio to provide data support for model training and validation. The Lasso algorithm is used to train the training set data, and key features are automatically selected by adjusting the regularization parameters. After training, the generalization ability of the model is verified using the test set, and the impact of various factors on task duration is analyzed based on the feature coefficients, ultimately forming a calculation model. S2: Task model expansion: Based on the existing task model, resource indicators such as CPU, memory, hard disk type, and external network address are added. At the same time, the timeliness attribute of tasks is added, and the distinction between hard real-time tasks, soft real-time tasks, and batch tasks is used to provide a richer basis for subsequent task scheduling and resource allocation. S3: Resource Monitoring: Design a standardized JSON data structure that includes core attributes such as device ID, task ID, CPU usage, total amount, memory usage, total amount, storage speed, hard disk type, bandwidth usage, latency for different external network addresses, and timestamps. Develop a device resource collection module using Java backend. On the hardware device side, collect resource usage data for each task in real time, including CPU, memory, hard disk read / write speed, hard disk type, bandwidth, and latency for multiple external network addresses. Simultaneously, obtain the maximum value of each device metric. After collecting the data, encapsulate it in a unified JSON format and report it. Drawing on the principles of the Raft consensus algorithm, a Java program with master node election, heartbeat detection, and data synchronization functions was developed. The election logic between nodes was implemented in the program. When a node received votes from more than half of the nodes within the specified time, it became the master node. A heartbeat detection mechanism was established between nodes, and the node sent heartbeat packets to other nodes at a set time interval (such as 1 second). If a node did not receive a heartbeat packet from a certain node within three consecutive heartbeat cycles, it was determined that the node might have failed, triggering a data synchronization operation. This mechanism ensures the status synchronization of multiple nodes, ensures strong consistency of node data, and provides accurate resource status information for task scheduling.

[0029] S4: Pre-selection, filtering and optimal sorting: Constructing a data model in the data storage module to store the operation records of hardware device identification and running task identification; In the task scheduling module developed in Java, when performing task allocation for hard real-time tasks, when calculating the remaining capacity of the computing device's CPU and memory, the resources in use are calculated at 120% to reserve reasonable resources for hard real-time tasks; according to the CPU, memory, hard disk type, and external network address required by the task node, the device's remaining capacity is filtered; if the running task is a hard real-time task, the CPU and memory are calculated at 120% of the requirements, and the hardware devices that do not meet the requirements are filtered out; if there are no available devices, the Netty framework is used to trigger the email alert mechanism in the existing alert module through network communication to notify the relevant operation and maintenance personnel to handle it in a timely manner; Continue to perform anti-affinity checks on the hardware devices that meet the requirements after screening in the resource scheduling module, and obtain the number of hard real-time tasks running on each hardware device in the data storage module; when the number of hard real-time tasks on a certain hardware device exceeds 50% of the total number of hard real-time tasks, this node does not participate in the allocation of hard real-time tasks to prevent hard real-time tasks from being concentrated and deployed on the same hardware device; Continue to use the calculation model trained in S1 to calculate the estimated running time for the hardware devices that meet the requirements after anti-affinity checks in the resource scheduling module; by inputting the resource metrics of the hardware devices into the calculation model, obtain the estimated running time of the tasks on each hardware device, and select the hardware device with the shortest running time as the carrier for task operation to achieve efficient task allocation; S5: Health Detection and Fault Tolerance Guarantee: Build a data model in the data storage module to store the abnormal operation records of task identifiers and the number of exceptions; Use the Netty framework to establish a network connection with the running tasks, and regularly check the health status and actual running duration of the tasks; set a scheduled task (such as checking once every 1 second), and obtain the running status information and actual running time of the tasks through the network connection; if the running duration of the task exceeds 50% of the estimated running time calculated in S4, save it as an abnormal record in the abnormal operation record data model of the data storage module; Adopt different strategies according to the health status and timeliness attributes of the tasks. If the task status is failed and it is a hard real-time task, directly terminate and mark it as failed, and use the Netty framework to trigger the email alert mechanism in the existing alert module through network communication to ensure timely notification to relevant personnel; if it is a soft real-time task and a batch processing task, support 3 retries, select different nodes each time, and during the retry process, re-screen and allocate nodes according to the resource status and task requirements to increase the probability of successful task execution; S6: Model Optimization: Use the Java backend to regularly scan the abnormal operation records in the data storage module to obtain the number of abnormal tasks. Set a scheduled task (at 1:00 a.m. every day) to obtain the number of abnormal tasks for each task from the abnormal operation records in the data storage module. When the number of abnormalities in a task exceeds 100, the Lasso algorithm is used to retrain the computing model according to the steps of data collection and model training, and the recent historical task execution records are collected again. Data preprocessing, model training and verification are performed to enable the computing model to adapt to new task execution conditions and resource environment changes, thereby improving the accuracy of task allocation and scheduling.

[0030] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: by collecting the historical execution records of task nodes, using the Lasso algorithm to train the calculation model, accurately analyzing the execution status of tasks in different hardware and network environments, providing a quantitative basis for task allocation, and changing the situation where traditional scheduling centers are unable to perceive the dynamic resource requirements of tasks in real time.

[0031] During the task allocation process, we comprehensively consider factors such as task resource requirements, hardware resource status, anti-affinity, and resource reservation mechanism, perform pre-selection filtering and optimal sorting, and run tasks on the optimal hardware nodes to avoid resource allocation imbalance, reduce idle equipment resources, improve task execution efficiency and resource utilization, ensure the supply of key task resources, and solve the problems of resource allocation imbalance and idle equipment resources in the current scheduling mechanism.

[0032] Use heartbeat detection and fault-tolerant technology to monitor the task running status in real time. When a task encounters an abnormality, it can trigger a restart or migration operation in time according to the task's timeliness attributes, ensuring that the task can be responded and processed in a timely manner, meeting the timeliness requirements of the enterprise's efficient data governance.

[0033] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention is limited to these examples. Within the scope of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0034] The present invention is intended to cover all such substitutions, modifications and variations that fall within the broad scope of the specification. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A task elastic orchestration and dynamic adaptation method for data governance, characterized in that The following steps are involved: S1: Data collection and model training: Collect historical execution records of task nodes, use the Lasso algorithm to analyze and process the collected data, and train a calculation model for the effects of CPU, memory, storage, hard disk type, and bandwidth on task duration; S2: Task model expansion: expand CPU, memory, hard disk type, external network address resource indicators and task timeliness attributes; S3: Resource Monitoring: Build a resource model to store the CPU, memory, storage, disk type, bandwidth usage, and maximum values of each device metric for each task reported in real time by hardware devices. Use the Raft consensus algorithm to ensure strong data consistency across all nodes. S4: Pre-selection, filtering, and optimal sorting: Based on the resource requirements of the nodes in S2 and the hardware resource status, anti-affinity, and resource reservation mechanisms in S3, all nodes that do not meet the requirements are filtered out. Based on the model data extended by S2, the calculation model of S1 is used on the filtered hardware nodes to calculate the expected runtime of each hardware node and prioritize them. Tasks are run on the optimal hardware nodes, and a data model is created to store the relationship between hardware devices and running tasks. S5: Health Check and Fault Tolerance: This section establishes network communication with tasks running in S4, regularly monitors their health status, and records tasks whose runtimes far exceed the expected runtime in S4. It then applies different retry strategies based on the task's health status and the timeliness properties extended by S2. S6: Model optimization: When the number of tasks whose actual running time recorded in S5 far exceeds the expected running time reaches a threshold, the Lasso model is retrained using recent historical data.

2. The task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that Said S1 specifically includes: S11: Create a task collection module to collect the resource usage and task execution time data of historical task execution; S12: For missing values in the data, the mean is used to fill in the missing values, and outliers are removed to ensure data quality. The preprocessed data is divided into training set and test set in a ratio of 7:3; S13: The model is trained based on the Lasso algorithm, and key features are automatically screened by adjusting the regularization parameters. Finally, the model generalization ability is verified using a test set. The impact of various factors on task duration is analyzed based on the feature coefficients to form a calculation model.

3. The task elastic orchestration and dynamic adaptation method for data governance according to claim 2, characterized in that The S13 specifically includes: S131: Define the loss function as: , Wherein: is the actual task duration of the i-th sample, is the transpose of the feature vector of the i-th sample, is the feature coefficient vector, n is the number of samples in the training set, p is the number of features, is the regularization parameter; by adjusting the regularization parameter , the model minimizes the loss function on the training set. When is relatively large, more feature coefficients will be shrunk to 0, thus automatically screening out key features; when is relatively small, the model retains more features; S132: Determine the optimal regularization parameter using the cross - validation method , specifically, divide the training set into k subsets, and successively select one of them as the validation set, and the remaining subsets as the training subsets. Train the model for each candidate value and calculate the mean squared error on the validation set. Finally, select the value corresponding to the minimum mean squared error as the optimal regularization parameter; S133: Use the test set to verify the trained model and calculate the coefficient of determination between the predicted task duration and the actual task duration of the model. , if the value is greater than the set threshold, it is determined that the model has good generalization ability. S134: Analyze the influence of each factor on the task duration according to the magnitude and sign of the characteristic coefficient ​ 4. A task elastic orchestration and dynamic adaptation method for data governance according to claim 3, characterized in that In the S134 : Factors with large absolute values of characteristic coefficients have a greater impact on task duration. A positive characteristic coefficient indicates that the factor will extend task duration, and a negative characteristic coefficient indicates that the factor will shorten task duration. Finally, these analysis results are integrated to form a calculation model.

5. A task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that The S2 specifically includes: S21: For the CPU type, define its performance metric as , through the formula , where is the main frequency of the CPU, is the CPU architecture coefficient, reflecting the performance differences of different CPU architectures, and converting the original CPU type into a standardized performance metric; S22: Define a memory efficiency metric for the memory type , where is the memory capacity, is the memory read / write speed, reflecting the comprehensive performance of memory capacity and speed; S23: Define the hard disk efficiency index for the hard disk type , where is the write speed, is the read speed, is the storage capacity, which is used to measure the comprehensive performance of the hard disk; S24: Define network performance metrics for external network address resource metrics , where is the network bandwidth, is the packet loss rate, comprehensively reflecting the transmission capacity and stability of the network; S25: Define a timeliness metric for the timeliness attribute of the task , where is the task completion time, is the task start time, is the task deadline, which is used to measure whether the task execution time is close to the deadline, thereby reflecting the timeliness of the task.

6. The task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that The S3 specifically includes: S31: Design a standardized data structure to describe resource status, including core attributes such as device identification, indicator type, real-time value, and timestamp to store device information; S32: Obtains the CPU, memory, hard disk read / write speed, hard disk type, bandwidth, and latency resource usage data of multiple external network addresses in S2 for each task in real time on the hardware device side, as well as the maximum value of each device indicator, and packages and reports the data in a unified format. S33: Deploy the Raft algorithm at the hardware resource status saving node, elect the master node, and establish a heartbeat detection mechanism between nodes to ensure the state synchronization of multiple nodes and the strong consistency of node data.

7. A task elastic orchestration and dynamic adaptation method for data governance according to claim 6, characterized in that, The specific steps of S33 include: S331: Leader Node Election in Raft Algorithm: The Raft algorithm divides nodes into three states: Follower, Candidate, and Leader. Initially, all nodes are in the Follower state. Define the timeout parameter as the time interval for a node to transition from the Follower state to the Candidate state. If a Follower node does not receive a heartbeat packet from the Leader node within this time, an election is triggered and the node transitions to the Candidate state. The Candidate node sends election requests to other nodes. Define the voting parameter . When a Candidate node receives more than half of the votes, it is successfully elected as the Leader node; S332: Heartbeat Detection Mechanism between Nodes: The Leader node periodically sends heartbeat packets to the Follower nodes and defines the heartbeat interval parameter , if the Follower node does not receive the Leader heartbeat packet within the specified time, it is considered that the Leader has failed and the election process is restarted; S333: Data Synchronization and Strong Consistency: Define the set of log entries , each log entry contains an index and a term . The Leader node replicates the new log to the Follower nodes. Through the majority commitment mechanism, it ensures that the log is committed only after it has been successfully replicated on the majority of nodes, guaranteeing strong data consistency; define the majority parameter , when a log entry has been successfully replicated by at least nodes, it is considered committed.

8. A task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that The specific steps of S4 include: S41: Build a data model to store operation records, including hardware device tags and running task identifiers; S42: Process the CPU and memory of a single task in the devices reported in S32. For hard real-time tasks, calculate according to the sum of the running resources and the reserved resources, reserve reasonable resources for hard real-time tasks, and calculate the remaining CPU and memory of the devices; S43: Filter the device margins obtained in S42 based on the CPU, memory, hard disk type, and external network address required by the task nodes in S2. For running hard real-time tasks, calculate the CPU and memory according to the sum of the required running resources and the reserved resources, filter out hardware devices that do not meet the requirements, and trigger an email alarm if there are no available devices; S44: Perform anti-affinity checks on the hardware devices obtained in S43. When the number of running hard real-time tasks in a certain hardware device exceeds a certain threshold of the total number of hard real-time tasks, this node does not participate in the allocation of hard real-time tasks; S45: Use the calculation model obtained in S13 to calculate the estimated running time for the hardware devices that meet the requirements in S44, and select the hardware device with the shortest running time as the task running carrier.

9. The task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that The specific steps of S5 include: S51: Build a data model to store abnormal operation records, including task identifiers and the number of abnormalities; S52: Establish a network connection with the running tasks to regularly check the health status and the actual running duration. If the running duration of a task exceeds a certain proportion of the estimated running time obtained in S45, save it as an abnormal record in the data model of S51; S53: When the task status obtained in S52 is a failure, according to the timeliness attributes extended in S2, directly terminate and mark the failure for hard real-time tasks, and support 3 retries for soft real-time tasks and batch processing tasks, and select different nodes for each retry.

10. A task elastic orchestration and dynamic adaptation method for data governance according to claim 1, characterized in that The specific steps of S6 include: S61: Regularly scan the data model of S51 to obtain the number of task abnormalities; S62: When the number of task abnormalities exceeds the threshold, retrain the calculation model using the Lasso algorithm according to S1.

Citation Information

Patent Citations

  • Self-adaptive resource scheduling method for improving cluster reconstruction efficiency

    CN119473590A

  • Multi-task data analysis method and device and storage medium

    CN119847752A

  • Computer data processing method and system

    CN119847862A

  • Virtual machine mirror image static scanning method and system based on hyper-fusion platform

    CN119885168A

  • Remote centralized control distributed data processing and analysis method based on JMS

    CN120128551A