Back-end process engine method, system and device with high concurrent processing capability
Through dynamic hierarchical fault-tolerant processing and predictive elastic scaling processing, combined with LSTM traffic prediction and hash ring virtual node migration technology, the performance bottlenecks and business interruption problems of traditional workflow engines in high concurrency scenarios are solved, and a high availability and fast response process engine system is realized.
Patent Information
- Application Number
- CN202510354709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Traditional relational database-based workflow engines are incompetent in high concurrency and large-scale data processing, and due to rigid fault tolerance mechanisms and configuration updates rely on downtime maintenance, service availability, business interruption and iteration lag.
Dynamic hierarchical fault-tolerant processing and predictive elastic scaling processing are adopted to predict traffic trends through the LSTM traffic prediction model, dynamically adjust the number of nodes, and achieve coordinated optimization of load balancing and resource utilization through hash ring virtual node migration and temporary virtual node insertion mechanisms.
It realizes intelligent migration or hierarchical retry of faulty tasks, compresses the average recovery time of non-critical tasks, reduces the reduction misjudgment rate, and improves service availability and business response speed.
Smart Images

Figure CN120234183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed system architectures, and specifically to a backend process engine method, system, and device with high concurrent processing capabilities. Background Art
[0002] With the development of information technology, the demand of enterprises and organizations for business process automation has been increasing day by day. As a core component to achieve this demand, the workflow engine has been widely used in various industries. However, traditional workflow engines based on relational databases often seem inadequate when facing high concurrency and large-scale data processing.
[0003] Although traditional workflow engines such as Activiti and Flowable can support complex business logics and task assignments, they do not fully consider high-concurrency scenarios in the Internet environment. For example, these systems usually rely on a single database instance for process definition and instance data storage, resulting in obvious performance bottlenecks in high-concurrency situations. In addition, due to the tight coupling between business logic and process control, when business rules change, it often requires modifying software code again, which is unacceptable for enterprises that need to quickly respond to market changes.
[0004] Chinese Patent Invention No. CN117573187A discloses a method for automatically generating different SDKs and example codes based on variable parameters, including the following steps: S1. Define specifications and create an SDK template; S2. Inject variable parameters into the SDK and provide downloads. Step S1 includes the following sub-steps: S11. Define the SDK interface specification and the SDK code structure specification; S12. Write SDK templates in different programming languages according to the definitions, and set variable parameters in the SDK templates; S13. Upload the templates to the template files. This invention improves the standardization, maintainability, and scalability of SDK development for applications, reduces the technical requirements, development complexity, development workload, and running and debugging man-hours of SDK users; can better standardize the SDK development process, improve the standardization and consistency of SDK development in various programming languages, and enhance the code quality of each SDK.
[0005] Most of the above-mentioned and similar process engines have a long node failure recovery time (>10 seconds) due to a rigid fault tolerance mechanism (such as a fixed 5-second retry interval), resulting in a decrease in service availability (SLA < 99.9%). At the same time, due to the characteristic that configuration updates rely on downtime maintenance (such as the lack of BPMN file hot loading), it further causes business interruption (request loss rate > 1%) and iteration lag (hourly update cycle), and thus it is difficult to meet the core requirements of elastic scaling, millisecond-level fault recovery, and real-time hot update of business logic in high-concurrency scenarios. Summary of the Invention
[0006] The purpose of the present invention is to provide a back-end process engine method, system and device with high concurrent processing capabilities to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions: A back-end process engine method with high concurrent processing capabilities, including:
[0008] S1: Dynamic hierarchical fault tolerance processing: According to the type of node failure, match the type of the node failure with the task label, and at the same time obtain the priority of the task label, and process the node failure hierarchically;
[0009] S2: Predictive elastic scaling processing: Through the LSTM traffic prediction model, obtain the traffic trend, determine the number of nodes to be pre-expanded, and migrate the tasks at the faulty node to the healthy node for processing, including:
[0010] Step S2.1: Pre-allocation of expanded nodes: Construct a traffic prediction model through time series data and external features to determine the number of nodes to be pre-expanded, specifically:
[0011]
[0012] Where: is the number of nodes to be pre-expanded, is the expected request rate 15 seconds later predicted at time point is the target utilization rate, is the current number of active nodes;
[0013] S2.2: Hysteretic scaling-down strategy: Determine the scaling-down delay time according to the traffic peak and valley values at the node, and compare the scaling-down delay time with the preset delay time. When the scaling-down delay time is greater than the preset delay time, the tasks corresponding to the scaling-down delay time are dynamically truncated, otherwise, the tasks corresponding to the scaling-down delay time run normally;
[0014] S2.3: Task distribution: Expand the nodes according to the number of pre-expanded nodes and the expected request rate, and allocate the expanded nodes to the constructed hash ring. At the same time, according to the scaling-down delay time and utilization rate corresponding to each node, perform migration processing of node tasks.
[0015] Furthermore, the hierarchical processing of node failures includes:
[0016] S1.1: Determine the failure type: According to the real-time data between each node, determine the failure type of each node, and the failure type includes network jitter, node downtime and resource exhaustion;
[0017] S1.2: Task priority classification: Each task at the node is marked with an integer - type tag. At the same time, according to the tag level corresponding to the task attributes and the system load rate, the tag level corresponding to each task is adjusted. Specifically:
[0018]
[0019] Where: is the adjusted tag level corresponding to the task attributes, is the original tag level corresponding to the task attributes, is the current system load rate, is the load - rate segmentation threshold;
[0020] S1.3: Strategy selection: According to the failure type of the node, the tasks of the node are processed through a retry strategy or a migration strategy. Specifically:
[0021] When the failure type is network jitter, the retry strategy is executed to process the tasks of the node. Otherwise, the migration strategy is executed to process the tasks of the node.
[0022] Furthermore, the failure type of each node is determined, including:
[0023] S1.1.1: Network jitter detection: By setting a sliding window through a preset acquisition number of ICMP Ping packets or application - layer heartbeat packets, the delay standard deviation, delay fluctuation coefficient, and packet loss rate corresponding to the sliding window at each node are obtained. Then, the delay standard deviation, delay fluctuation coefficient, and packet loss rate are compared with the preset delay standard deviation, delay fluctuation coefficient, and packet loss rate. According to the comparison results, it is determined whether the failure type at the node is network jitter. Specifically:
[0024] When the delay standard deviation, delay fluctuation coefficient, and packet loss rate are all greater than the preset delay standard deviation, delay fluctuation coefficient, and packet loss rate respectively, the failure type at the node is network jitter. Otherwise, the failure type at the node is not network jitter;
[0025] S1.1.2: Node downtime detection: The cluster nodes are divided into sending nodes and receiving nodes. The sending nodes broadcast heartbeat signals to the receiving nodes. At the same time, according to the time when the receiving nodes obtain the last heartbeat signal, the timeout threshold for each node to go down is obtained. Then, the timeout threshold is compared with the preset timeout threshold, and according to the comparison results, it is determined whether the failure type at the node is node downtime. Specifically:
[0026] When the timeout threshold for node downtime is greater than the preset timeout threshold, the fault type at the node is node downtime; otherwise, the fault type at the node is not node downtime.
[0027] S1.1.3: Resource exhaustion detection: Obtain a comprehensive score through CPU utilization, memory occupancy, and I / O wait time, compare the comprehensive score with a preset score, and determine whether the fault type at the node is resource exhaustion based on the comparison result. Specifically:
[0028] When the comprehensive score is greater than the preset score, the fault type at the node is resource exhaustion; otherwise, the fault type at the node is not resource exhaustion.
[0029] Furthermore, the tasks of the node are processed through a retry strategy or a migration strategy, including:
[0030] S1.3.1: Retry strategy: Divide the tasks into critical tasks and non-critical tasks according to the adjusted label level, and determine the retry interval according to the baseline parameters corresponding to the divided tasks. Specifically:
[0031]
[0032] Where: is the retry interval corresponding to critical tasks, is the initial retry interval corresponding to critical tasks, is the number of retry times corresponding to critical tasks, is the fixed increment step corresponding to critical tasks, is the retry interval corresponding to non-critical tasks, is the number of backoff times corresponding to non-critical tasks, is the initial backoff interval corresponding to non-critical tasks;
[0033] S1.3.2: Migration strategy: Divide the nodes into healthy nodes and faulty nodes according to the health score corresponding to each node, compare the health score with a preset health score, and determine whether the tasks at the node need to be migrated. Specifically:
[0034] When the health score is not less than the preset health score, the node corresponding to the health score is a healthy node; otherwise, the node corresponding to the health score is a faulty node. At the same time, the execution tasks under the faulty node are migrated to the healthy node for task processing.
[0035] Further, determine a target node from the healthy nodes, and migrate the execution tasks under the faulty node to the target node for processing. The process of determining the target node includes:
[0036] W1: Obtain the matching degree: According to the required CPU and required memory of the tasks in the faulty node, obtain the matching degree between each healthy node and the faulty node. Specifically:
[0037]
[0038] Where: is the matching degree between the healthy node and the faulty node, is the amount of CPU resources required for the task, is the remaining CPU resources of the target node, is the CPU weight coefficient, is the amount of memory resources required for the task, is the remaining memory resources of the target node, is the memory weight coefficient;
[0039] W2: Determine the target node: Compare the matching degrees between all healthy nodes and the faulty node, and determine the maximum matching degree from them. The healthy node corresponding to the maximum matching degree is the final target node.
[0040] Further, perform the migration processing of node tasks, including:
[0041] S2.3.1: Construct a hash ring: Allocate multiple virtual nodes to each physical node, obtain the hash value corresponding to each virtual node through the SHA-1 algorithm, and map the hash value to the hash ring;
[0042] S2.3.2: Influence of node position: Adjust the virtual nodes on the hash ring according to the migration transformation of the physical nodes. The migration ratio of the virtual nodes is specifically:
[0043]
[0044] Where: is the migration ratio of the virtual node, is the original number of physical nodes, is the change number of physical nodes, is the number of virtual nodes;
[0045] S2.3.3: Task routing allocation: Determine the target physical node on the hash ring according to the task hash value and the corresponding virtual node, and at the same time adjust the routing through the temporary virtual node and adjust the weight of the virtual node. Specifically:
[0046] When the physical node is a high-load node, reduce the number of virtual nodes corresponding to the physical node; when the physical node is a low-load node, increase the number of virtual nodes corresponding to the physical node; otherwise, the number of virtual nodes remains unchanged.
[0047] Furthermore, route adjustment is performed through temporary virtual nodes, including:
[0048] S2.3.3.1: Determine the task hash value: Determine the task hash value through the task identifier and the CRC32 algorithm, specifically:
[0049]
[0050] Where: is the task hash value, is the cyclic redundancy check algorithm, is the session identifier, is the modulo operation;
[0051] S2.3.3.2: Determine the target physical node: Compare the task hash value with the hash values corresponding to all virtual nodes, and according to the ascending order of the virtual nodes, determine the first hash value not less than the task hash value. The virtual node corresponding to the first hash value not less than the task hash value is the task virtual node, and the physical node corresponding to the task virtual node is the target physical node;
[0052] S2.3.3.3: Obtain the node utilization rate: Obtain the node utilization rate corresponding to each node through the CPU utilization rate and memory utilization rate of the node, compare the node utilization rate with the preset utilization rate, and perform task processing according to the comparison result, specifically:
[0053] When the node utilization rate is greater than the preset utilization rate, insert a temporary virtual node into the hash ring for route adjustment; otherwise, directly perform task processing at the node.
[0054] Furthermore, compare the node utilization rate with the preset utilization rate. When the node utilization rate is greater than the preset utilization rate, reduce the virtual nodes; otherwise, increase the virtual nodes. The adjustment formula for the number of virtual nodes is specifically:
[0055]
[0056] Where: is the number after reducing the number of virtual nodes, is the number after increasing the number of virtual nodes, is the maximum value of the preset utilization rate, is the minimum value of the preset utilization rate, is the node utilization rate, is the number of virtual nodes.
[0057] A backend process engine system with high concurrent processing ability, characterized in that it uses a backend process engine method with high concurrent processing ability described in any one of the above.
[0058] A backend process engine device with high concurrent processing ability uses a backend process engine method with high concurrent processing ability described in any one of the above.
[0059] Compared with the prior art, the beneficial effects of the present invention are:
[0060] First: Through a three-level fault detection model, namely the network jitter standard deviation threshold, heartbeat timeout determination, and resource scoring formula, the present invention realizes the second-level identification of fault types, and through the task priority dynamic adjustment algorithm, greatly improves the success rate of critical task recovery. At the same time, through exponential backoff retry and resource class fault-triggered migration strategy, the average recovery time of non-critical tasks is greatly compressed;
[0061] Second: Based on the dynamic truncation strategy of traffic peak and valley analysis, the present invention maintains node redundancy for 15 - 30 seconds during the traffic decline stage, thereby preventing mis-scaling due to short-term fluctuations, and greatly reducing the mis-scaling judgment rate;
[0062] Third: Through the dynamic migration of hash ring virtual nodes and the insertion mechanism of temporary virtual nodes, the present invention makes the request loss rate lower than 0.01% during configuration update, and controls the service interruption time within 200ms. Description of the Drawings
[0063] Figure 1 is the flowchart of the backend process engine method of the present invention;
[0064] Figure 2 is the comparison chart of fault recovery time of the present invention (50 tests);
[0065] Figure 3 is the comparison chart of request loss rate under different concurrency levels of the present invention;
[0066] Figure 4 is the comparison chart of resource utilization rate stability of the present invention (4-hour monitoring);
[0067] Figure 5 is the comparison chart of expansion response delay of the present invention (5 tests);
[0068] Figure 6 is the comprehensive performance comparison radar chart of the present invention. Specific Embodiments
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] Most existing process engines have a long node failure recovery time (>10 seconds) due to a rigid fault tolerance mechanism (such as a fixed 5-second retry interval), resulting in a decline in service availability (SLA<99.9%). At the same time, due to the characteristic that configuration updates rely on downtime maintenance (such as the lack of hot loading of BPMN files), it further causes business interruption (request loss rate>1%) and iteration lag (hourly update cycle), and thus it is difficult to meet the core requirements of elastic scaling, millisecond-level fault recovery, and real-time hot update of business logic in high-concurrency scenarios. The technical solution of this application realizes the intelligent migration or hierarchical retry of failed tasks by real-time detecting the node failure type and combining the dynamic adjustment strategy of task priority and the matching of healthy nodes. At the same time, by predicting the traffic trend in the next 15 seconds through the LSTM model, determining the number of pre-expanded nodes, dynamically adjusting the number of nodes, and combining the hysteretic shrinkage strategy and the virtual node hash ring, the coordinated optimization of load balancing and resource utilization is achieved, and finally the technical effects of millisecond-level fault recovery, predictive resource scaling, and highly available task distribution are achieved.
[0071] Embodiment 1
[0072] Refer to Figures 1 - 6 , this embodiment provides a backend process engine method with high-concurrency processing capabilities. The backend process engine method specifically includes the following steps:
[0073] Step S1: Dynamic hierarchical fault tolerance processing. That is, through real-time identification and analysis of faults, determine the type of each fault, and match each fault type with the set task tags. At the same time, set the priority of the task tags, and combine the hybrid backoff algorithm and the health score to migrate tasks, triggering cross-node task migration. Specifically as follows:
[0074] Step S1.1: Determine the fault type. That is, through the acquisition and analysis of real-time data between nodes, determine the fault type of the node, and the fault type includes network jitter, node downtime, and resource exhaustion. Specifically as follows:
[0075] Step S1.1.1: Network jitter detection. That is, monitoring agents are deployed at each node of the cluster, and ICMP Ping packets or application layer heartbeat packets are sent to other nodes at a preset time interval (100 ms). It should be noted that each ICMP Ping packet or application layer heartbeat packet includes round-trip time and packet loss status.
[0076] Furthermore, according to the preset acquisition times of ICMP Ping packets or application layer heartbeat packets, a sliding window is set to obtain the moving average delay and delay standard deviation corresponding to the sliding window. Specifically:
[0077]
[0078] Where: is the moving average delay, is the size of the sliding window, is the round-trip time of the i-th measurement, is the delay standard deviation, is the measurement index.
[0079] Furthermore, according to the obtained moving average delay and delay standard deviation, the delay fluctuation coefficient and packet loss rate corresponding to the sliding window at each node are determined. Specifically:
[0080]
[0081] Where: is the delay fluctuation coefficient, is the packet loss rate, is the moving average delay, is the size of the sliding window, is the delay standard deviation, is the total number of packet losses within the sliding window.
[0082] Specifically, within a preset period (which is specifically set according to actual needs, so it is not specifically elaborated in this embodiment), the delay standard deviation, delay fluctuation coefficient, and packet loss rate corresponding to the sliding window at each node are obtained. At the same time, the obtained delay standard deviation, delay fluctuation coefficient, and packet loss rate are respectively compared with the preset delay standard deviation, preset delay fluctuation coefficient, and preset packet loss rate, and according to the comparison results, it is determined whether the fault type at this node is network jitter. Specifically:
[0083] When the obtained delay standard deviation, delay fluctuation coefficient, and packet loss rate are all greater than the preset delay standard deviation, preset delay fluctuation coefficient, and preset packet loss rate respectively, the fault type at this node is network jitter; otherwise, the fault type at this node is not network jitter.
[0084] In the process of specific implementation, ICMP Ping packets or application layer heartbeat packets within three consecutive cycles are obtained. The measured round-trip times in the third cycle are 250ms, 400ms, and 500ms respectively. The size of the corresponding sliding window is 10, and the number of lost packets is 2 times. That is to say, the delay standard deviation, delay fluctuation coefficient, and packet loss rate corresponding to the third cycle are 104.2ms, 0.27, and 20% respectively. Further, the preset delay standard deviation, preset delay fluctuation coefficient, and preset packet loss rate in this embodiment are 50ms, 0.3, and 5% respectively.
[0085] That is to say, although the obtained delay standard deviation and packet loss rate are both greater than the preset delay standard deviation and preset packet loss rate, the obtained delay fluctuation coefficient is less than the preset delay fluctuation coefficient. Therefore, the fault type at this node is not network jitter.
[0086] Step S1.1.2: Node downtime detection. The cluster nodes are divided into sending nodes and receiving nodes. Each sending node broadcasts a heartbeat signal to other nodes (receiving nodes) in the cluster at a preset interval. The heartbeat signal includes the node ID, timestamp, and load status. At the same time, the receiving node records the time of the last received heartbeat signal.
[0087] Further, based on the time of the last received heartbeat signal obtained at the receiving node, the timeout threshold for each node to go down is obtained. Specifically:
[0088]
[0089] Where: is the timeout threshold for the node to go down, is the sending interval of the heartbeat signal, is the network delay tolerance.
[0090] Specifically, the obtained timeout threshold for the node to go down is compared with the preset timeout threshold, and based on the comparison result, it is determined whether the fault type at this node is node downtime. Specifically:
[0091] When the obtained timeout threshold for the node to go down is greater than the preset timeout threshold, the fault type at this node is node downtime. Otherwise, the fault type at this node is not node downtime.
[0092] In the process of specific implementation, the sending interval of the heartbeat signal is 500ms, and the network delay tolerance is 100ms. Then the timeout threshold for the node to go down in this embodiment is 1.1s. At the same time, the preset timeout threshold in this embodiment is set to 2s. That is to say, the timeout threshold for the node to go down in this embodiment is less than the preset timeout threshold, so the fault type at this node is not node downtime.
[0093] Step S1.1.3: Resource exhaustion detection. That is, by monitoring the real-time usage of CPU, memory, and disk, obtain the CPU utilization rate, memory occupancy rate, and I / O waiting time, and set the corresponding weight according to the actual usage requirements to determine the specific comprehensive score as follows:
[0094]
[0095] Where: is the comprehensive score, is the weight corresponding to the CPU utilization rate, is the weight corresponding to the available memory, is the weight corresponding to the I / O waiting time, is the CPU utilization rate, is the available memory, is the I / O waiting time.
[0096] Specifically, compare the obtained comprehensive score with the preset score, and determine whether the fault type at this node is resource exhaustion according to the comparison result. Specifically:
[0097] When the obtained comprehensive score is greater than the preset score, the fault type at this node is resource exhaustion; otherwise, the fault type at this node is not resource exhaustion.
[0098] Furthermore, by reading the / proc / stat file, extract the cpu line data to determine the CPU idle time and total CPU time within a unit time, and calculate the corresponding CPU utilization rate. Similarly, read the / proc / meminfo file to determine the total system physical memory and available memory, and calculate the corresponding memory occupancy rate. Similarly, read the / proc / diskstats file to determine the average I / O request processing time and the number of I / O operations per second, and calculate the I / O waiting time. Specifically:
[0099]
[0100] Where: is the CPU utilization rate, is the increment of CPU idle time within a unit time, is the increment of total CPU time within a unit time, is the memory occupancy rate, is the available memory, is the total system physical memory, is the I / O waiting time, is the average I / O request processing time, is the number of I / O operations per second, is the maximum I / O wait time.
[0101] In the process of specific implementation, the increment of CPU idle time per unit time is 100, and the increment of total CPU time per unit time is 1000. Therefore, the CPU utilization rate is 90%. Similarly, the available memory is 2GB, and the total physical memory of the system is 16GB. Therefore, the memory occupancy rate is 87.5%. Similarly, the average I / O request processing time is 5000ms, the number of I / O operations per second is 700, and the maximum I / O wait time is 200ms. Therefore, the I / O wait time is 0.04.
[0102] Specifically, the weights corresponding to the CPU utilization rate (0.9), available memory (0.875), and I / O wait time (0.04) in this embodiment are 0.6, 0.2, and 0.2 respectively. Therefore, the comprehensive score is: 0.6 * 0.9 + 0.2 * 0.875 + 0.2 * 0.04 = 0.723.
[0103] Furthermore, the preset score in this embodiment is set to 0.7. That is to say, the obtained comprehensive score is greater than the preset score. Therefore, the fault type at this node is resource exhaustion.
[0104] Step S1.2: Task priority classification. That is, an integer label is assigned to each task, and the value range is set to 1 - 5, specifically:
[0105] Level 1 label: Highest priority, such as real-time payment and core inventory deduction.
[0106] Level 2 label: High priority, such as order status update and transaction compensation.
[0107] Level 3 label: Medium priority, such as asynchronous message notification.
[0108] Level 4 label: Low priority, such as log archiving.
[0109] Level 5 label: Lowest priority, such as non-critical cache preheating.
[0110] Furthermore, according to the set rule matrix, the task attributes of each task are mapped to the corresponding labels. It should be noted that when the current task attribute does not match the corresponding label, the label corresponding to the current task attribute is set to the Level 5 label.
[0111] Specifically, according to the label level corresponding to the task attribute and the system load rate, the label level corresponding to the task attribute is adjusted to obtain the final label level corresponding to the task attribute. It is worth noting that no matter how high the obtained label level is, the highest label level is always a level 5 label. Furthermore, when adjusting the label level corresponding to the task attribute, the adjustment formula is specifically:
[0112]
[0113] in: is the adjusted label level corresponding to the task attribute, is the original label level corresponding to the task attribute, is the current load rate of the system, It is the load rate segmentation threshold.
[0114] In the specific implementation process, the original label level corresponding to the task attribute is 3, the current system load rate is 70%, and the load rate segmentation threshold is 20%. Then the adjusted label level corresponding to the task attribute is 5.
[0115] Step S1.3: Strategy selection. That is, according to the node failure type determined in step S1.1, the corresponding strategy linkage is performed. Furthermore, when the node failure type is network jitter, the task type of the node is divided into critical tasks and non-critical tasks, and the corresponding processing is performed through the set retry strategy to ensure the rapid recovery of critical tasks while avoiding non-critical tasks from increasing the network burden.
[0116] Furthermore, when the node failure type is node downtime or resource exhaustion, the migration strategy is triggered to migrate the node without performing local retry. In other words, local retry cannot solve node-level failures, so it is migrated to healthy nodes for processing.
[0117] Furthermore, when the node failure type determined in step S1.1 includes network jitter, node downtime and / or resource exhaustion, the node failure requires both retry and migration. Specifically, the priority of retry and node migration can be determined according to the following formula, which is:
[0118]
[0119] in: Score the action priority, is the task priority weight, is the retry interval, Rate your health.
[0120] Specifically, compare the action priority scores corresponding to the retry operations and the action priority scores corresponding to the migration operations obtained, and determine the maximum action priority score therefrom. The operation corresponding to the maximum action priority score is the operation to be preferentially executed.
[0121] In this embodiment, according to the node failure type determined in step S1.1 and according to the label level determined in step S1.2, classify the tasks at each node accordingly, and perform specific processing according to the classification results. Specifically as follows:
[0122] Step S1.3.1: Retry policy. That is, according to the size of the task label determined in step S1.2, divide the tasks into critical tasks (label levels 1 and 2) and non-critical tasks (label levels 3, 4, and 5), and obtain the retry intervals of each task.
[0123] Furthermore, according to the baseline parameters (initial retry interval, fixed increment step, and number of retries) obtained for the critical tasks, determine the retry interval corresponding to the critical tasks, specifically:
[0124]
[0125] Where: is the retry interval corresponding to the critical task, is the initial retry interval corresponding to the critical task, is the number of retries corresponding to the critical task, is the fixed increment step corresponding to the critical task.
[0126] In the process of specific implementation, the initial retry interval is set to 100 ms, the fixed increment step is set to 50 ms, and the maximum number of retries is set to 5. That is to say, after the task fails for the first time, it is retried after 100 ms, and at the same time, it is retried at intervals of 50 ms each time until the maximum number of retries is reached. When it still fails after the last retry, the task is marked as permanently failed, and an alarm will be triggered and a log will be recorded here.
[0127] Furthermore, according to the baseline parameters (initial backoff interval, randomization range, and number of backoffs) obtained for the non-critical tasks, determine the retry interval corresponding to the non-critical tasks, specifically:
[0128]
[0129] Where: is the retry interval corresponding to the non-critical task, is the number of backoffs corresponding to the non-critical task, is the initial backoff interval corresponding to the non-critical task.
[0130] In the process of specific implementation, the initial backoff interval is set to 200 ms, and the maximum number of backoffs is set to 8. That is to say, after the task fails for the first time, it is retried at any time within 0 - 200 ms, and at the same time, the random range is exponentially expanded each time (the maximum interval is 51 s) for retry. When the number of retries reaches 8 times and still fails, the task is cancelled and the exception is recorded.
[0131] Step S1.3.2: Migration strategy. That is, according to the failure rate at the node, obtain the health score of each node, and according to the obtained health score, divide the nodes into healthy nodes and faulty nodes. Specifically, compare the obtained health score with the preset health score to divide the nodes, and at the same time determine whether the task at this node needs to be migrated. Specifically:
[0132] When the obtained health score is less than the preset health score, the node corresponding to this health score is a faulty node; otherwise, the node corresponding to the health score is a healthy node. At the same time, migrate the execution tasks under the faulty node to the healthy node for task processing.
[0133] Furthermore, the formula for obtaining the health score is specifically:
[0134]
[0135] Where: is the health score, , , are the weight coefficients, is the node failure rate, is the remaining resource rate, is the average response delay.
[0136] Step S2: Predictive elastic scaling processing. That is, through the LSTM traffic prediction model, obtain the traffic trend, and according to the traffic trend, determine the number of nodes to be pre-expanded. That is to say, according to the failure type of the node, perform corresponding node expansion to migrate the tasks at the faulty node to the healthy node for processing. Specifically as follows:
[0137] Step S2.1: Pre-allocation of expanded nodes. That is, according to the time series data (such as historical traffic data (requests per second) in the past 1 hour) and external features (such as business activities (such as promotions), time periods (such as peak / low), holiday markers), construct a traffic prediction model through the LSTM network structure. Specifically:
[0138]
[0139] Where: For the expected request rate 15 seconds after the predicted future at time point is the prediction function of the LSTM model, For time point −3600 seconds to the historical request volume sequence from the current time point t, For time point −3599 seconds to the historical request volume sequence from the current time point t, is the historical request volume sequence at the current time point t.
[0140] Furthermore, based on the expected request rate obtained from the traffic prediction model, determine the number of nodes to be pre-expanded, specifically:
[0141]
[0142] Where: is the number of nodes to be pre-expanded, For time point the expected request rate 15 seconds after the predicted future, is the target utilization rate, is the current number of active nodes.
[0143] During the specific implementation process, at time point the expected request rate 15 seconds after the predicted future is 9200 RPS. At the same time, the target utilization rate is set to 70%, and the current number of active nodes is 10. Therefore, the number of nodes to be pre-expanded is 4, that is, 4 more nodes need to be pre-expanded currently.
[0144] Step S2.2: Hysteresis scaling-down strategy. That is, based on the obtained peak and valley traffic at the node, determine the corresponding scaling-down delay time, specifically:
[0145]
[0146] Where: is the scaling-down delay time, is the traffic peak, is the traffic valley.
[0147] Furthermore, compare the obtained scaling-down delay time with the preset delay time. When the obtained scaling-down delay time is greater than the preset delay time, the task corresponding to this scaling-down delay time needs to be dynamically truncated. Otherwise, it can run normally.
[0148] Step S2.3: Task distribution. That is, according to the expanded nodes obtained in step S2.1, a hash ring is constructed. At the same time, based on the constructed hash ring and the expected request rate obtained from the traffic prediction model, and according to the number of pre-expanded nodes obtained, node expansion is performed and allocated to the hash ring. Further, according to the shrinkage delay time and utilization rate obtained for each node, migration processing of node tasks is performed. Specifically as follows:
[0149] Step S2.3.1: Construct a hash ring. That is, multiple virtual nodes are allocated to each physical node to ensure uniform load distribution, where each physical node uses the combination of IP address + port + serial number as the unique identifier. At the same time, the hash value corresponding to each virtual node is obtained through the SHA-1 algorithm, and the obtained hash value is mapped to the hash ring, so that all virtual nodes can be sorted in ascending order according to the size of the hash value to form a ring structure.
[0150] Step S2.3.2: Influence of node position. That is, according to the migration transformation of physical nodes, the virtual nodes on the hash ring are adjusted to maintain the correspondence between virtual nodes and physical nodes. The specific migration ratio of virtual nodes is as follows:
[0151]
[0152] Where: is the migration ratio of virtual nodes, is the original number of physical nodes, is the change number of physical nodes, is the number of virtual nodes.
[0153] Step S2.3.3: Task routing and allocation. That is, according to the task hash value and the corresponding virtual node, the target physical node is determined on the hash ring, and routing adjustment is performed through temporary virtual nodes to adjust the virtual node weights. Specifically:
[0154] High-load nodes: Reduce the number of their virtual nodes and lower the probability of new task allocation.
[0155] Low-load nodes: Increase the number of virtual nodes and improve the task processing capacity.
[0156] Further, the method for adjusting the routing is specifically as follows:
[0157] Step S2.3.3.1: Determine the task hash value. That is, use unique fields such as Session ID, request ID, or user ID as the task identifier, and obtain a 32-bit hash value through the CRC32 algorithm to determine the task hash value. Specifically:
[0158]
[0159] Wherein: is the task hash value, is the cyclic redundancy check algorithm, is the session identifier, is the modulo operation.
[0160] Step S2.3.3.2: Determine the target physical node. That is, according to the task hash value determined in Step S2.3.3.1, compare the task hash value with the hash values of all virtual nodes, and determine the first hash value that is not less than the task hash value. The virtual node corresponding to the first hash value that is not less than the task hash value is the task virtual node. Further, according to the mapping relationship between the physical node and the virtual node, the physical node corresponding to the task virtual node is the target physical node.
[0161] Step S2.3.3.3: Obtain the node utilization rate. That is, obtain the node utilization rate corresponding to each node through the CPU utilization rate at the node and the memory utilization rate at the node. Specifically:
[0162]
[0163] Wherein: is the node utilization rate, is the CPU utilization rate at the node, is the memory utilization rate at the node.
[0164] Further, compare the obtained node utilization rate with the preset utilization rate. When the obtained node utilization rate is greater than the preset utilization rate, it is necessary to insert a temporary virtual node in the hash ring for routing adjustment. Otherwise, the task can be directly processed at the node.
[0165] Specifically, compare the obtained node utilization rate with the preset utilization rate. When the obtained node utilization rate is greater than the preset utilization rate, reduce the number of virtual nodes. Otherwise, increase the number of virtual nodes. Specifically, the adjustment formula for the number of virtual nodes is as follows:
[0166]
[0167] Wherein: is the number after reducing the number of virtual nodes, is the number after increasing the number of virtual nodes, is the maximum value of the preset utilization rate, is the minimum value of the preset utilization rate, is the node utilization rate, is the number of virtual nodes.
[0168] In the process of specific implementation, the preset utilization rate is set to 65% - 75%, where 65% is the minimum value of the preset utilization rate and 75% is the maximum value of the preset utilization rate. That is to say, when the node utilization rate is lower than 65%, there is resource waste and the load needs to be increased. When the node utilization rate is between 65% - 75%, it is in an ideal state and no adjustment is required. When the node utilization rate is higher than 75%, there is overload and the load needs to be reduced.
[0169] Furthermore, when the current number of virtual nodes is 1000 and the node utilization rate is 60%, the number of virtual nodes needs to be increased, and the number after the increase is 1077. That is to say, 77 new virtual nodes are added, the task distribution probability is increased by 7.7%, and the utilization rate gradually rises above 65%.
[0170] Reference Figure 2 , Figure 2 is a comparison chart of the fault recovery time (50 tests). It can be seen from it that the median of the traditional scheme is about 12.3 seconds, while the median of the technical scheme of this application is about 0.79 seconds. That is to say, the absolute time is reduced by 11.51 seconds, achieving a 15-fold performance improvement. At the same time, the standard deviation of the traditional scheme is 1.2 seconds, while the standard deviation of the technical scheme of this application is 0.15 seconds. That is to say, the time fluctuation range is reduced by 87.5%, that is, the stability of the system is significantly improved.
[0171] Reference Figure 3 , Figure 3 is a comparison chart of the request loss rate under different concurrency levels. It can be seen from it that in the 5000-concurrency scenario, the loss rate of the traditional scheme is 2.5%, while the loss rate of the technical scheme of this application is 0.12%. That is to say, the absolute difference reaches 2.38 percentage points, and the improvement rate is 95.2%. For the key inflection point, that is, the 2000-concurrency scenario, the loss rate growth rate of the traditional scheme accelerates, increasing from 1.2% to 1.8%, and then increasing to 2.5%. The growth rate drops from 50% to 38%. While the loss rate of the technical scheme of this application still shows a linear growth, that is, increasing from 0.05% to 0.07%, and then increasing to 0.12%. The growth rate increases from 40% to 71%.
[0172] Reference Figure 4 , Figure 4 is a comparison chart of the resource utilization rate stability (4-hour monitoring). It can be seen from it that the resource utilization rate of the traditional scheme fluctuates violently between 30% - 95%, and the peak-to-valley difference reaches 65 percentage points, hitting the 30% idle threshold and the 95% overload threshold many times. While the resource utilization rate of the technical scheme of this application is stable in the target range of 65% - 75%, and the maximum fluctuation range is only 10 percentage points, which is reduced by 84.6% compared with the traditional scheme.
[0173] ReferenceFigure 5 , Figure 5 is the comparison chart of expansion response delay (5 tests). It can be seen from it that the average value of the average delay of the traditional solution is 44.2 seconds, while the average value of the average delay of the technical solution of this application is 14.8 seconds, the absolute time is reduced by 29.4 seconds, and the improvement rate is 66.5%. At the same time, all 5 tests of the traditional solution exceed its nominal threshold of 45 seconds, while the delay of all tests of the technical solution of this application is less than 17 seconds, which is 2.6 times faster than the traditional threshold.
[0174] Reference Figure 6 , Figure 6 is the radar chart of comprehensive performance comparison. It can be seen from it that the fault recovery time has dropped from 12.4 seconds to 0.8 seconds, achieving an acceleration of 15.5 times and realizing sub-second recovery. The request loss rate is 1.8, which is reduced by 36 times. At the same time, the fluctuation range of resource stability is 30, the standard deviation drops from 18.7 to 2.1, and the fluctuation amplitude is reduced by 88.8%. The expansion response delay drops from 45 seconds to 15 seconds, and the response speed is increased by 3 times. The pre-expansion mechanism reduces the peak stack by 71%.
[0175] Furthermore, in this embodiment, a backend process engine system with high concurrency processing ability is also provided. This backend process engine method uses a backend process engine method with high concurrency processing ability described above.
[0176] Furthermore, in this embodiment, a device of a backend process engine with high concurrency processing ability is also provided. This backend process engine device uses a backend process engine method with high concurrency processing ability described above.
[0177] Embodiment 2
[0178] This embodiment provides a backend process engine method with high concurrency processing ability. The specific implementation method is the same as that of Embodiment 1. The difference is that the execution tasks under the faulty node are migrated to a healthy node for task processing. The present invention will be illustrated below in conjunction with the specific implementation manners of this embodiment.
[0179] In this embodiment, during the process of migrating the execution tasks under the faulty node to a healthy node for task processing, a target node needs to be determined from among many healthy nodes. That is to say, the execution tasks under the faulty node are migrated to the target node matching the faulty node for corresponding processing. Specifically as follows:
[0180] Step W1: Obtain the matching degree. That is, according to the task demand CPU and task demand memory in the faulty node, obtain the matching degree between each healthy node and the faulty node. Specifically:
[0181]
[0182] Wherein: is the matching degree between the healthy node and the faulty node, is the CPU resource amount required by the task, is the remaining CPU resource amount of the target node, is the CPU weight coefficient, is the memory resource amount required by the task, is the remaining memory resource amount of the target node, is the memory weight coefficient.
[0183] In the process of specific implementation, the CPU resource amount required by the task is 2 cores, the remaining CPU resource amount of the target node is 4GB, the memory resource amount required by the task is 4 cores, and the remaining memory resource amount of the target node is 8GB. Therefore, the matching degree between the healthy node and the faulty node is 2.0.
[0184] Step W2: Determine the target node. That is, compare the matching degrees between all healthy nodes and faulty nodes, and determine the maximum matching degree from them. The healthy node corresponding to the maximum matching degree is the final target node.
[0185] Embodiment 3
[0186] This embodiment provides a backend process engine method with high concurrency processing ability. The specific implementation method is the same as that of Embodiment 1. The difference is that the hash value is mapped to the hash ring so that all virtual nodes can be sorted in ascending order according to the size of the hash value to form a ring structure. The present invention will be illustrated below in conjunction with the specific implementation manner of this embodiment.
[0187] In this embodiment, the uniformity of the hash ring is ensured in the following manner, specifically as follows:
[0188] Step M1: Multiple hashing. That is, perform multiple hash calculations on each virtual node to generate multiple hash values, and disperse the obtained multiple hash values to different positions on the hash ring to reduce the distribution deviation that may be caused by a single hash function. At the same time, even if the number of physical nodes is small, the positions of virtual nodes can be dispersed through multiple hashing to reduce the hot spot area.
[0189] Step M2: Virtual node redundancy. That is, according to the number of physical nodes, obtain the number of virtual nodes corresponding to each physical node so that each physical node corresponds to the same number of virtual nodes. Further, it can be ensured that even if the number of physical nodes is small, the total number of virtual nodes is still large enough, thereby avoiding sparse distribution on the hash ring.
[0190] Furthermore, the specific formula for determining the number of virtual nodes is:
[0191]
[0192] Wherein: is the number of virtual nodes, is the number of physical nodes.
[0193] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended embodiments and their equivalents.
Claims
1. A backend process engine method with high concurrent processing capability, characterized in that: Included are: S1: Dynamic hierarchical fault-tolerance processing: According to the type of node failure, the type of node failure is matched with the task label, and the priority of the task label is obtained, and the node failure is processed hierarchically; S2: Predictive elastic scaling: The LSTM traffic prediction model is used to obtain traffic trends, determine the number of pre-expansion nodes, and migrate tasks at faulty nodes to healthy nodes for processing, including: Step S2.1: Pre-allocation of expansion nodes: Based on time series data and external features, a traffic prediction model is constructed to determine the number of nodes for pre-expansion, specifically: ; in: is the number of nodes to be pre-expanded, For at time point The expected request rate in the next 15 seconds. is the target utilization rate, is the current number of active nodes; S2.2: Delayed shrinking strategy: Determine the shrinking delay time according to the traffic peak and valley values at the node, and compare the shrinking delay time with the preset delay time. When the shrinking delay time is greater than the preset delay time, the task corresponding to the shrinking delay time is dynamically truncated. Otherwise, the task corresponding to the shrinking delay time runs normally. S2.3: Task distribution: Node expansion is performed according to the number of pre-expanded nodes and the expected request rate, and the expanded nodes are allocated to the constructed hash ring. At the same time, node task migration is performed according to the shrinking delay time and utilization rate corresponding to each node.
2. A backend process engine method with high concurrent processing capability according to claim 1, characterized in that: Node failures are handled in stages, including: S1.1: Determine the fault type: Determine the fault type of each node based on the real-time data between each node. The fault type includes network jitter, node downtime and resource exhaustion; S1.2: Task priority classification: Each task at the node is marked with an integer label, and the label level corresponding to each task is adjusted according to the label level corresponding to the task attribute and the system load rate, specifically: ; in: is the adjusted label level corresponding to the task attribute, is the original label level corresponding to the task attribute, is the current load rate of the system, is the load rate segmentation threshold; S1.3: Strategy selection: According to the failure type of the node, the task of the node is processed through a retry strategy or a migration strategy, specifically: When the fault type is network jitter, the retry strategy is executed to perform task processing of the node; otherwise, the migration strategy is executed to perform task processing of the node.
3. A backend process engine method with high concurrent processing capability according to claim 2, characterized in that: Determine the failure type of each node, including: S1.1.1: Network jitter detection: by obtaining the preset number of ICMP Ping packets or application layer heartbeat packets, a sliding window is set, and the delay standard deviation, delay fluctuation coefficient and packet loss rate corresponding to the sliding window at each node are obtained, and the delay standard deviation, delay fluctuation coefficient and packet loss rate are compared with the preset delay standard deviation, delay fluctuation coefficient and packet loss rate. According to the comparison result, it is determined whether the fault type at the node is network jitter, specifically: When the delay standard deviation, delay fluctuation coefficient and packet loss rate are respectively greater than the preset delay standard deviation, delay fluctuation coefficient and packet loss rate, the fault type at the node is network jitter; otherwise, the fault type at the node is not network jitter; S1.1.2: Node downtime detection: The cluster nodes are divided into sending nodes and receiving nodes, and the sending nodes broadcast heartbeat signals to the receiving nodes. At the same time, according to the time when the receiving nodes obtain the last heartbeat signal, the timeout threshold of each node downtime is obtained, and the timeout threshold is compared with the preset timeout threshold. According to the comparison result, it is determined whether the fault type at the node is node downtime, specifically: When the timeout threshold of the node downtime is greater than the preset timeout threshold, the fault type at the node is node downtime, otherwise, the fault type at the node is not node downtime; S1.1.3: Resource exhaustion detection: Obtain a comprehensive score through CPU utilization, memory occupancy and I / O waiting time, and compare the comprehensive score with the preset score, and determine whether the fault type at the node is resource exhaustion based on the comparison result, specifically: When the comprehensive score is greater than a preset score, the fault type at the node is resource exhaustion; otherwise, the fault type at the node is not resource exhaustion.
4. A backend process engine method with high concurrent processing capability according to claim 2, characterized in that: The tasks of the node are processed through a retry strategy or a migration strategy, including: S1.3.1: Retry strategy: According to the adjusted label level, the task is divided into critical tasks and non-critical tasks, and the retry interval is determined according to the baseline parameters corresponding to the divided tasks, specifically: ; in: is the retry interval corresponding to the critical task, is the initial retry interval corresponding to the critical task, is the number of retries corresponding to the critical task, is the fixed incremental step size corresponding to the key task, is the retry interval corresponding to non-critical tasks, is the number of backoffs corresponding to non-critical tasks, The initial backoff interval corresponding to non-critical tasks; S1.3.2: Migration strategy: According to the health score corresponding to each node, the nodes are divided into healthy nodes and faulty nodes, and the health score is compared with the preset health score to determine whether the task at the node needs to be migrated, specifically: When the health score is not less than the preset health score, the node corresponding to the health score is a healthy node. Otherwise, the node corresponding to the health score is a faulty node. At the same time, the execution tasks under the faulty node are migrated to the healthy node for task processing.
5. A backend process engine method with high concurrent processing capability according to claim 4, characterized in that: A target node is determined from the healthy nodes, and the execution tasks under the faulty node are migrated to the target node for processing. The target node determination process includes: W1: Obtaining the matching degree: Obtaining the matching degree between each of the healthy nodes and the faulty node according to the task requirement CPU and task requirement memory in the faulty node, specifically: ; in: is the matching degree between healthy nodes and faulty nodes, is the amount of CPU resources required by the task, The remaining CPU resources of the target node. is the CPU weight coefficient, is the amount of memory resources required by the task, is the remaining memory resource of the target node, is the memory weight coefficient; W2: Determine the target node: compare the matching degrees between all the healthy nodes and the faulty nodes, and determine the maximum matching degree. The healthy node corresponding to the maximum matching degree is the final target node.
6. A backend process engine method with high concurrent processing capability according to claim 1, characterized in that: Perform node task migration, including: S2.3.1: Build a hash ring: assign multiple virtual nodes to each physical node, obtain the hash value corresponding to each virtual node through the SHA-1 algorithm, and map the hash value to the hash ring; S2.3.2: Influence of node position: According to the migration transformation of the physical node, the virtual nodes on the hash ring are adjusted. The migration ratio of the virtual nodes is specifically as follows: ; in: is the migration ratio of virtual nodes, is the original number of physical nodes, is the number of changes in the physical node, is the number of virtual nodes; S2.3.3: Task routing allocation: According to the task hash value and the corresponding virtual node, the target physical node is determined on the hash ring, and the routing is adjusted through the temporary virtual node, and the weight of the virtual node is adjusted, specifically: When the physical node is a high-load node, the number of virtual nodes corresponding to the physical node is reduced. When the physical node is a low-load node, the number of virtual nodes corresponding to the physical node is increased. Otherwise, the number of virtual nodes remains unchanged.
7. A backend process engine method with high concurrent processing capability according to claim 6, characterized in that: Routing adjustments are made through temporary virtual nodes, including: S2.3.3.1: Determine the task hash value: Determine the task hash value through the task identifier and CRC32 algorithm, specifically: ; in: is the task hash value, is the cyclic redundancy check algorithm, is the session identifier, is the modulo operation; S2.3.3.2: Determine the target physical node: compare the task hash value with the hash values corresponding to all virtual nodes, and determine the first hash value that is not less than the task hash value according to the ascending sorting of the virtual nodes. The virtual node corresponding to the first hash value that is not less than the task hash value is the task virtual node, and the physical node corresponding to the task virtual node is the target physical node; S2.3.3.3: Obtain node utilization: Obtain the node utilization corresponding to each node through the CPU utilization and memory utilization of the node, and compare the node utilization with the preset utilization, and perform task processing based on the comparison result, specifically: When the node utilization is greater than a preset utilization, a temporary virtual node is inserted into the hash ring to perform routing adjustment; otherwise, task processing is performed directly at the node.
8. A backend process engine method with high concurrent processing capability according to claim 7, characterized in that: The node utilization rate is compared with the preset utilization rate. When the node utilization rate is greater than the preset utilization rate, the virtual nodes are reduced. Otherwise, the virtual nodes are increased. The formula for adjusting the number of virtual nodes is specifically: ; in: is the number of virtual nodes after reduction, is the number of virtual nodes after the increase. is the maximum value of the preset utilization rate, is the minimum value of the preset utilization rate, is the node utilization, is the number of virtual nodes.
9. A backend process engine system with high concurrent processing capability, characterized in that: A backend process engine method with high concurrent processing capability as described in any one of claims 1-8 is used.
10. A backend process engine device with high concurrent processing capability, characterized in that: A backend process engine method with high concurrent processing capability as described in any one of claims 1-8 is used.
Citation Information
Patent Citations
Method for automatically generating different SDKs and example codes based on variable parameters
CN117573187A
Cloud computing system capacity expansion method and device based on function computing nodes
CN111897658A
Dynamic capacity expansion and contraction method and device for distributed computing power network service and storage medium
CN117527814A
Distributed automatic expansion method in high-performance computing
CN118838735A
Large model service platform for high-efficiency concurrent computing
CN119149257A