Short message platform multi-node dynamic scheduling and resource optimization method based on load prediction

CN122765601APending Publication Date: 2026-09-15SUZHOU BAIXUN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610875019.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供基于负载预测的短信平台多节点动态调度与资源优化方法,用于解决现有技术难以区分任务的关键程度,也无法感知节点上悬挂的未确认任务对后续调度的潜在压力,导致高价值业务的交付质量无法得到保障的问题;

Benefits of technology

[0014] 1. By collecting the age distribution statistics of unclosed tasks and calculating the virtual compensation load, the scheduler gains a quantitative perception of the scale and aging degree of the hanging backlog within the nodes. This virtual load increases monotonically with the age of the tasks and the risk of timeout. Before the actual occurrence of the receipt storm, the comprehensive scheduling load value of the nodes is raised in advance, so that newly arrived SMS requests are guided to nodes with lighter hanging backlogs. This avoids the requests from continuously gathering on the backlog nodes due to the false idleness of the physical load index. The coefficient of variation of the number of hanging tasks between nodes decreases round by round, and the frequency of timeout closure events converges to below the preset alarm threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122765601A_ABST
    Figure CN122765601A_ABST
Patent Text Reader

Abstract

The application belongs to the message distribution and control technology in the communication network, and is used for solving the problem that the prior art cannot distinguish the key degree of task, cannot perceive the potential pressure of the unconfirmed task hanging on the node on the subsequent scheduling, and cannot guarantee the delivery quality of high-value services, and specifically is a short message platform multi-node dynamic scheduling and resource optimization method based on load prediction, comprising the following steps: acquiring node state data and calculating virtual compensation load; weighting and fusing the physical load index and the virtual compensation load to obtain the comprehensive scheduling load of the node; and distributing the newly arrived short message request to the node with the lowest comprehensive scheduling load according to the comprehensive scheduling load of each node; by collecting the age distribution statistics of the uncompleted task and calculating the virtual compensation load, the scheduler obtains the quantitative perception ability of the size and aging degree of the hanging backlog in the node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to message distribution and control technology in communication networks, specifically a method for dynamic scheduling and resource optimization of multiple nodes in a text messaging platform based on load prediction. Background Technology

[0002] SMS platforms typically employ a multi-node concurrent sending architecture. The scheduler allocates new SMS messages to the node with the lowest current load based on the load metrics reported by each node (such as queue depth, thread utilization, and recent response time). To improve throughput, nodes generally adopt a "submit and release" strategy: after an SMS message is successfully submitted to the operator's gateway, the currently occupied sending slot or connection is immediately released to receive the next SMS message. This strategy works well in scenarios where only sending efficiency is a concern.

[0003] For critical SMS messages requiring clear delivery conclusions, such as financial verification codes and payment confirmations, a "submission successful" message alone does not guarantee that the user has received the message. The operator gateway's "submission successful" response only indicates that the message has been accepted; whether it was ultimately delivered, timed out, or arrived within the validity period requires confirmation from subsequent acknowledgments. When the scheduler continuously receives high submission rates from nodes, it determines that the node is under light load and dispatches more critical SMS messages to it. At this point, the node has accumulated a large number of submitted but unacknowledged tasks. These unconfirmed tasks do not occupy sending slots or directly affect real-time load metrics, but they create a "hanging backlog" within the node. Once subsequent acknowledgments return failures or timeouts in quick succession, the system needs to retry or compensate for these tasks. However, the node may already be occupied by newly dispatched tasks and unable to process them promptly. More subtly, different services have varying degrees of reliance on acknowledgments; verification code tasks will become invalid if delayed beyond their validity period, while marketing tasks have lower timeliness requirements.

[0004] The existing scheduling mechanism has difficulty distinguishing the criticality of tasks and cannot perceive the potential pressure that unconfirmed tasks hanging on nodes will put on subsequent scheduling, resulting in the inability to guarantee the delivery quality of high-value services. Summary of the Invention

[0005] The purpose of this invention is to provide a method for dynamic scheduling and resource optimization of multiple nodes in a text messaging platform based on load prediction, which solves the problem that existing technologies cannot distinguish the criticality of tasks and cannot perceive the potential pressure of unconfirmed tasks hanging on nodes on subsequent scheduling, resulting in the inability to guarantee the delivery quality of high-value services.

[0006] The technical problem to be solved by this invention is: how to provide a method for dynamic scheduling and resource optimization of multiple nodes in a text messaging platform based on load prediction, which can distinguish the criticality of tasks and also perceive the potential pressure of unconfirmed tasks hanging on nodes on subsequent scheduling.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] A method for dynamic scheduling and resource optimization of multiple nodes in an SMS platform based on load prediction includes:

[0009] Obtain node status data: Obtain node status data reported by each SMS sending node. Node status data includes physical load indicators and age distribution statistics of unclosed tasks. Unclosed tasks refer to SMS tasks that have been submitted to the operator but have not yet received the final delivery receipt.

[0010] Calculate the virtual compensation load: Based on the age distribution statistics reported by each node and the dynamic timeout threshold maintained by the node, calculate the virtual compensation load of the node. The dynamic timeout threshold is updated in real time based on the time interval distribution of the most recent successful receipts received by the node.

[0011] Weighted fusion: The physical load index and the virtual unpaid load are weighted and fused to obtain the overall scheduling load of the node;

[0012] Distribute SMS requests: Based on the overall scheduling load of each node, newly arrived SMS requests are distributed to the node with the lowest overall scheduling load.

[0013] The present invention has the following beneficial effects:

[0014] 1. By collecting the age distribution statistics of unclosed tasks and calculating the virtual compensation load, the scheduler gains a quantitative perception of the scale and aging degree of the hanging backlog within the nodes. This virtual load increases monotonically with the age of the tasks and the risk of timeout. Before the actual occurrence of the receipt storm, the comprehensive scheduling load value of the nodes is raised in advance, so that newly arrived SMS requests are guided to nodes with lighter hanging backlogs. This avoids the requests from continuously gathering on the backlog nodes due to the false idleness of the physical load index. The coefficient of variation of the number of hanging tasks between nodes decreases round by round, and the frequency of timeout closure events converges to below the preset alarm threshold.

[0015] 2. The dynamic timeout threshold is updated in real time based on the high quantile value of the time interval distribution of the recent successful receipts of the node. This allows the age band boundary to adaptively follow the latency fluctuations of the operator's receipt link, eliminating the need for manual configuration of a fixed timeout time. Combined with the dynamic adjustment of the basic weights based on the receipt processing history (increasing the weight of the highest age band when the timeout rate exceeds the threshold, and reducing all weights when the proportion of normal receipts in the middle and high age bands exceeds the threshold), the virtual compensation load value can distinguish between two different causes of hanging backlog: "increased overall latency of the receipt link" and "insufficient processing capacity of the node itself". This avoids misjudging node overload due to link jitter and also avoids underestimating the true pressure due to smooth node processing. The overall signal-to-noise ratio of the scheduling load is higher than that of the fixed threshold scheme.

[0016] 3. In the fault succession step, the scheduler transfers the status summary (age distribution statistics and dynamic timeout threshold) of the faulty node to the successor node. The successor node then creates a virtual closed tracking pool based on this and matches receipts during its lifespan. This mechanism ensures that the receipt tracking of suspended tasks on the faulty node is not interrupted due to node failure. The remaining unmatched counts in the virtual pool are converted into compensation events and written to the queue after the lifespan expires, ensuring the continuity of closed tracking. The receipt loss rate of unclosed tasks on the faulty node is reduced to zero, and its suspension pressure is released in an orderly manner in the form of compensation events after the timeout, rather than suddenly impacting the compensation queue.

[0017] 4. The credit pre-screening step performs token bucket-based rate control on the first closure level requests, maintaining an independent quota for each initiator's identity. If the quota is insufficient, the request is directly rejected without entering the node allocation process. This mechanism limits the injection rate of critical SMS messages from the entry point, preventing a sudden surge in traffic from a single initiator from instantly consuming all the acknowledgment processing resources of all nodes. Meanwhile, the second closure level requests are unrestricted and do not track acknowledgments, allowing the system's sending throughput capacity to still be fully utilized by non-critical businesses. The acknowledgment processing resources of critical and non-critical businesses are isolated from each other, and the on-time delivery rate of critical SMS messages does not change with fluctuations in non-critical business traffic. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is the main flowchart of the method in Embodiment 1 of the present invention;

[0020] Figure 2 This is a sub-flowchart of step S3 in Embodiment 1 of the present invention. Detailed Implementation

[0021] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] In the concurrent sending architecture of the SMS platform, the scheduler allocates new requests to the node with the lowest current load based on the load indicators reported by the nodes. Nodes generally adopt the "submit and release" strategy: the sending slot is released immediately after the SMS is successfully submitted to the operator gateway in order to maintain a high apparent throughput. This strategy works normally in scenarios where only sending efficiency is a concern, but it introduces a structural blind spot in observation. Tasks that have been submitted but have not received a final delivery receipt (i.e., unclosed tasks) do not occupy sending slots, nor do they affect processor utilization or thread pool utilization, but they continue to accumulate within the node, forming a hanging backlog.

[0023] The physical load metrics (thread pool utilization, processor utilization) relied upon by the existing scheduling mechanism cannot detect the scale and aging of this suspended backlog. When the scheduler continuously receives low load reports from nodes, it will determine that the node is in an idle state and continue to dispatch new tasks to it. However, the node has actually accumulated hundreds to thousands of unclosed tasks waiting for acknowledgments. When subsequent acknowledgments return failures or timeouts in a concentrated manner, the system needs to perform retry or compensation operations on these tasks. However, at this time, the node has been filled with sending resources by newly dispatched tasks, and compensation tasks are forced to queue up and wait, resulting in a serious delay in the delivery of critical SMS messages (such as financial verification codes and payment confirmations), exceeding the business validity period.

[0024] This problem is exacerbated when tasks of different closure levels are sent together. CAPTCHA tasks will become invalid if delayed beyond their expiration date (typically 60 to 120 seconds), while marketing tasks have lower timeliness requirements. Existing scheduling mechanisms cannot differentiate the criticality of tasks or perceive the proportion of pending tasks of different closure levels on each node, leading to a disconnect between scheduling decisions and business priorities. Specifically, in e-commerce promotional scenarios, CAPTCHA SMS requests surge. The scheduler distributes requests to various nodes according to a physical load balancing strategy. However, some nodes have already accumulated a large number of unreceived marketing tasks. When CAPTCHA tasks are assigned to these nodes, their receipt processing is overwhelmed in the backlog, and many CAPTCHAs expire before receipts arrive. On another node, although the physical load is similar, the pending tasks are mostly handled by younger users, resulting in lower actual processing pressure. Lacking information on age distribution, the scheduler cannot distinguish between these two situations and can only make indiscriminate allocation decisions.

[0025] If the above problems are not resolved, the scheduler will continue to lose its ability to determine the actual processing pressure on nodes; the backlog of tasks not included in the load will cause tasks to continue to accumulate on overloaded nodes, forming a negative feedback loop—the more backlog a node has, the more new tasks it will be assigned, the longer the receipt processing delay will be, the higher the proportion of timeout closure events will be, triggering a large number of invalid retries, further consuming sending resources; at the same time, the mixing of tasks with different closure levels will make it impossible to guarantee the delivery quality of high-priority services, and the increase in verification code timeout rate will directly cause service failures such as user login failure and payment interruption; ultimately, the deviation between the physical load indicators used by the scheduler and the actual available capacity of the nodes will continue to widen, and although the overall system throughput will remain at the nominal value, the effective delivery rate (especially the on-time delivery rate of time-sensitive services) will drop significantly, and the service quality of the SMS platform will not converge to a range acceptable to the business.

[0026] Example 1: As Figure 1-2 As shown, the method for dynamic scheduling and resource optimization of multiple nodes in an SMS platform based on load prediction includes:

[0027] Step S1: Obtain node status data: Obtain the node status data reported by each SMS sending node. The node status data includes physical load indicators and age distribution statistics of unclosed tasks. Unclosed tasks refer to SMS tasks that have been submitted to the operator but have not yet received the final delivery receipt.

[0028] In step S1 of Embodiment 1, it is first necessary to obtain the node status data reported by each SMS sending node. This step is the basic data collection link of the entire scheduling method. Its core is not only to collect traditional physical load indicators, but also to introduce the characterization of the "hanging backlog" status inside the node, that is, the age distribution statistics of unclosed tasks. The following, in conjunction with the specific implementation scheme of this embodiment, will elaborate on the acquisition process, data definition, calculation logic and dynamic update mechanism of step S1.

[0029] Specifically, the SMS platform in this embodiment includes a scheduler and multiple SMS sending nodes; each SMS sending node is responsible for submitting SMS tasks to the corresponding operator gateway and asynchronously receiving the final delivery receipt returned by the operator; the nodes and the scheduler maintain communication through a heartbeat mechanism, and the nodes actively report their current node status data to the scheduler at fixed time intervals (e.g., every 500 milliseconds); the "acquisition" behavior in step S1 refers to the process by which the scheduler receives and parses these reported data.

[0030] In this embodiment, the node status data consists of two main parts: physical load metrics and age distribution statistics of unclosed tasks. The physical load metrics reflect the instantaneous pressure on the node at the real-time computing and transmission levels. To eliminate the bias of a single metric, this embodiment defines the physical load metrics as the weighted average of the node's current transmission thread pool utilization and processor utilization. The processor utilization is the average utilization rate of the node from the last reporting time to the current time, and the thread pool utilization is the instantaneous sampled value at the current time. Both values ​​have been normalized to the [0,1] interval by the node before being included in the calculation. Let the ratio of the number of active threads in the transmission thread pool of node i at time t to the total thread pool size be _____. Processor utilization rate Then the physical load index The calculation formula is:

[0031] ;

[0032] in, The preset weighting coefficient has a value range between 0 and 1. In this embodiment, it is set to [value missing] based on experience. This is to emphasize the direct impact of thread pool utilization on sending capacity; it is worth noting that this physical load metric does not include tasks that have been submitted but have not received a response, which is why subsequent virtual compensation load is introduced to make up for its information blind spot.

[0033] The second part of the node status data, namely the age distribution statistics of unclosed tasks, is the key to load prediction in this embodiment. An "unclosed task" is explicitly defined in this embodiment as an SMS task that a node has successfully submitted to the operator gateway but has not yet received a receipt representing the final delivery status (such as "successfully delivered," "failed to send," "expired," etc.). Each unclosed task enters an "age" timer from the moment it is submitted; the node maintains the submission timestamp of each unclosed task locally and calculates the current age of each task based on the current time each time a status report is generated.

[0034] To efficiently report age information rather than transmitting a detailed task list, this embodiment employs a discretized statistical method based on age bands; specifically, the node first needs to maintain a dynamic timeout threshold. This threshold is not a fixed value, but is updated in real time based on the distribution of time intervals of successful receipts received by the node in the recent period. The node records the arrival time of the most recent N (e.g., N=200) successful receipts, calculates the time interval between adjacent successful receipts, and forms a sample set. Statistical analysis is performed on this set, and the highest quantile value is taken as the basis for the dynamic timeout threshold. Let the p-quantile of the successful receipt time interval distribution be... Where p is a high quantile close to 1, in this embodiment we take ,Right now Dynamic timeout threshold and They are positively correlated, and the specific calculation formula is as follows: ;in This is a multiplier associated with the SMS closure level; in this embodiment, for a typical first-closure level task, Set to 1.5; for subclasses with stricter timeliness requirements (such as payment confirmation SMS), it can be configured to... This allows them to enter higher age groups more quickly; the second closure level task does not track receipts and is not involved in the calculation of this threshold; system settings As an upper limit protection; if the calculated result If the limit is exceeded, the output will be truncated. Meanwhile, the node generates an alarm flag and carries it in the next status report, so that the scheduler can trigger an operation and maintenance notification.

[0035] The dynamic timeout threshold is recalculated every statistical period (e.g., every 10 seconds) to adaptively reflect the current latency characteristics of the carrier's acknowledgment link.

[0036] Obtaining the dynamic timeout threshold Then, the node is divided into multiple age zones according to the age zone boundaries determined by the threshold; let the total number of age zones be K, in this embodiment K=4; the boundaries of the age zones are determined by... Definition of multiples: The first age band (young band) is the age range within the interval. Unclosed tasks within; the second age group (young and middle-aged) is an interval. The third age group (middle age) is a range. The fourth age band (old age band) is a range. The node iterates through all its unclosed tasks, and based on the age range each task falls into, counts the number of tasks within each of the four age bands, denoted as . Thus, the age distribution statistics for unclosed tasks are now expressed as a vector. .

[0037] It is worth noting that tasks with different closure levels may correspond to different timeout threshold multiples. In this embodiment, when an SMS request arrives at the scheduler, it carries a "closure level" field, which is pre-set by the business party based on the criticality of the SMS. For example, verification code SMS messages belong to the "first closure level," requiring strict tracking of receipts; while marketing SMS messages belong to the "second closure level," and their receipt status is not tracked. For tasks of the first closure level, when calculating their age distribution, the node uses the aforementioned division method based on dynamic timeout thresholds, and the dynamic timeout threshold... It directly applies to all first-level closed-loop tasks; however, for certain subclasses with stricter timeliness requirements (such as payment confirmation SMS), a smaller timeout threshold can be specified in the system configuration, for example... This allows them to enter higher age bands more quickly, reflecting their urgency; in this embodiment, an independent age distribution statistic is maintained for each closure level, i.e., in the aforementioned vector. In addition to the existing closure level, an extra dimension is added to identify the corresponding closure level. For simplicity, the following description uses a single closure level as an example; in a real system, multiple levels can be processed in parallel.

[0038] The node will calculate the physical load metrics. and the dynamically adjusted age distribution statistics of unclosed tasks Together with the dynamic timeout threshold on which the current calculation is based The data is encapsulated into a status data message and sent to the scheduler via the network. After receiving the status data of each node, the scheduler completes the acquisition process in step S1. This data will serve as the raw input for subsequent calculation of virtual compensation load and comprehensive scheduling load. Through the above implementation method, step S1 not only collects the real-time physical pressure of the nodes, but also quantifies the invisible backlog of suspended tasks and their aging degree inside the nodes through age band statistics and dynamic threshold mechanism, providing a comprehensive data foundation for load prediction.

[0039] If the number of successful response samples used to calculate percentiles is less than N when the node is started or when... Initialize to the system's default conservative value Each subsequent successful receipt is added to the sample set until the sample size reaches N, at which point percentile calculation is initiated. If no successful receipt is received within two consecutive statistical periods (each period is 10 seconds), The current value will remain unchanged, and the age zoning will still be based on this value until the next period's samples are recovered and recalculated.

[0040] When the age band boundary expands due to an upward adjustment, the virtual uncompensated load may experience a temporary decrease, and the node's overall load ranking in the scheduler will shift forward accordingly; however, the newly assigned tasks will increase the absolute count of unclosed tasks on that node, which will be reported later. The increment will partially offset the effect of boundary expansion; when the physical load of the node increases due to the increase of tasks, the physical load item directly restricts the further reduction of its comprehensive scheduling load, forming negative feedback convergence; therefore, this adaptive mechanism will not cause requests to continuously aggregate to the node with worsening latency.

[0041] Step S2: Fault Handover Step: The scheduler maintains the last reported status summary of each node. The status summary includes the age distribution statistics of unclosed tasks and the tail high quantile of the current successful receipt time interval distribution. When a node fails to report within the timeout period, it is determined to be a faulty node. A successor node is selected from the healthy nodes, and the status summary of the faulty node is sent to the successor node. The successor node creates a virtual closure tracking pool based on this. When it receives a receipt, it first searches in its own real pool. If it does not find it, it matches it in the virtual pool according to age priority and deducts the corresponding count.

[0042] During operation, the scheduler maintains a state summary for each active node. This summary is not simply a cache of the node's last reported raw data, but contains three key fields: the vector of age distribution statistics for unclosed tasks last reported by the node (denoted as...). The dynamic timeout threshold currently used by this node. And the high quantile of the distribution of successful receipt time intervals calculated based on this threshold. ;in and satisfy , Take 1.5; each time the scheduler receives a report from a node, it overwrites the node's state summary with the new data and records the timestamp of this reception.

[0043] When the scheduler detects that a node has exceeded the preset heartbeat timeout period (e.g., no status report received for 1.5 consecutive seconds), it marks the node as a faulty node. At this point, the scheduler does not immediately discard its status digest, but instead selects a replacement node from the remaining healthy nodes. The selection of the replacement node follows a two-layer filtering process: first, it filters nodes that use the same operator channel type as the faulty node, including China Mobile CMPP, China Unicom SGIP, and China Telecom SMGP; if no node with the same channel type exists, it selects the node with the highest channel type compatibility—compatibility is defined by the system's pre-configured channel type conversion table. For example, CMPP and SGIP can interoperate via a protocol adapter, denoted as compatibility level 1, while other types have no compatibility relationship; in the filtering results, the scheduler selects the node with the lowest current overall scheduling load as the replacement node. The value of the overall scheduling load is taken from the result calculated after the node's most recent report.

[0044] Compatibility is determined based on a pre-configured channel type compatibility table. This table uses channel type enumeration values ​​(such as CMPP, SGIP, SMGP) as keys and a list of compatible types and an integer value for the compatibility level as values: 0 indicates completely identical channels, 1 indicates lossless conversion via a protocol adapter, 2 indicates only transcoding but possible partial loss of state mapping, and 3 indicates incompatibility. When selecting, the lowest compatibility level is preferred. If multiple nodes have the same compatibility level, their overall scheduling load is compared.

[0045] After selecting a replacement node, the scheduler will send a state summary of the failed node (including age distribution statistics). and dynamic timeout threshold The faulty node sends a message to the successor node; upon receiving the message, the successor node creates a virtual closed-loop tracking pool in its own memory space; the core data structure of this pool is an integer array of length 4, initially set to the counts of the four age bands reported by the faulty node. This indicates that the successor node needs to track these tasks that have not yet received a response; at the same time, the successor node sets a maximum lifespan for this virtual pool. The calculation formula is: ,in It is the current dynamic timeout threshold that replaces the node itself. Take 1.2.

[0046] The system also sets a minimum survival time threshold for the virtual pool. If the calculated result Then cut off to This lower limit is determined based on the normal arrival window of the carrier's receipts—statistics show that over 99% of normal receipts arrive within 1000ms of submission. This protection ensures that the virtual pool survives at least one complete receipt collection cycle, avoiding issues caused by... An abnormally low value caused the virtual pool to be destroyed prematurely.

[0047] Subsequently, when the successor node receives the carrier's acknowledgment, it performs a double lookup: It searches its own real closed tracking pool by the task ID in the acknowledgment; if a match is found, the tracking is closed normally; otherwise, it moves to the virtual closed tracking pool. The successor node's virtual closed tracking pool does not maintain the specific mapping between the faulty node's task IDs and submission times, but only retains the remaining counts for the four age groups. When a acknowledgment is received but not found in the real pool, the successor node, based on the carrier's acknowledgment's generally first-to-first-delivery timing pattern (acknowledgment delay distribution has short-term stability), starts from the highest age group... The scan begins from the fourth age band, adding the current receipt to the first age band with a remaining count greater than 0, and decrementing the count of that age band by 1. If all age band counts are 0, the receipt is ignored. As an alternative to enhance matching accuracy, when a faulty node submits a task, it can encode the lower 16 bits of the submission timestamp (milliseconds modulo 65536) into the reserved field of the task ID. After receiving the receipt, the successor node parses this field to restore the approximate submission time and determines the age band to which the receipt belongs. The ID collision probability is less than one in ten thousand and does not affect the overall statistical effect.

[0048] During the virtual pool's lifespan, the count for the corresponding age group is deducted for each matching receipt received; when the virtual pool's lifespan exceeds... Then, the successor node checks whether the remaining counts for all four age groups in the array are zero; if not, the remaining counts are converted into timeout compensation events and written to the compensation queue; the conversion method is: the remaining counts of each age group when the virtual pool is destroyed. The data is converted into compensation events and written to the compensation queue in descending order of age band. The conversion rule is as follows: starting from the highest age band (fourth band), the remaining counts of each age band are used to generate the corresponding number of compensation events, and the compensation events of the same age band are arranged consecutively. The higher the age band, the higher its compensation event ranks in the queue. The compensation queue processes data according to this priority order, giving priority to compensating tasks in higher age bands that have not been matched with receipts. After writing is completed, the virtual pool is destroyed and the memory is released.

[0049] For example, the state summary of a faulty node is: To take over the node itself ,but If only two receipts from age band 3 and one receipt from age band 4 are matched within 720ms, the remaining count becomes [20, 15, 6, 2], and then it is converted into a compensation event and written to the queue. The whole process does not depend on the recovery of the faulty node, and achieves seamless succession of the suspended task.

[0050] Step S3: Calculate the virtual compensation load: Based on the age distribution statistics reported by each node and the dynamic timeout threshold maintained by the node, calculate the virtual compensation load of the node. The dynamic timeout threshold is updated in real time based on the time interval distribution of the most recent successful receipts received by the node.

[0051] After a node completes its status data reporting, the scheduler calculates the virtual compensation load for each node; the input to this step is the age distribution statistics of unclosed tasks reported by the node. and the dynamic timeout threshold maintained by this node. The output is a dimensionless virtual unpaid load value. This is used to quantify the processing pressure that the current suspended task of a node may generate in the future.

[0052] The core of the calculation lies in assigning basic weight values ​​to the unclosed tasks for different age groups. Let the basic weights corresponding to the four age groups be... The older the age, the higher the weight; in this embodiment, the initial weight is set to... The weight values ​​reflect the relative magnitude of the compensation cost that needs to be paid after a task times out or fails in that age group; however, these basic weights are not fixed and will be dynamically adjusted by the node based on the recent receipt processing history; the adjustment logic also needs to be reflected in step S3, because the calculation of virtual uncompensated load depends on the adjusted weights.

[0053] The node maintains two sliding window statistics; the first is the percentage of events that were judged as timed out and closed due to exceeding the dynamic timeout threshold among the most recent M closed tasks (M is 500). The second is the percentage of normal receipts belonging to the middle-to-high age group (age group three and age group four) within the same window. The system presets two thresholds: alarm threshold. (Used to determine whether the proportion of time-out closure events has reached the level at which the risk of the highest age group should be amplified), dominance threshold (Used to determine whether the proportion of normal responses from the middle-to-high age group has reached the level at which the weight should be reduced); the adjustment process is carried out in sequence: first determine If the condition is met, then the base weight of the highest age group (fourth age group) will be adjusted. Multiply by the first magnification factor To obtain temporary weights The weights of other age groups remain unchanged; then make a judgment. If true, then all weights (including those that may have been amplified) should be considered. Multiply by the first attenuation coefficient At this point, the weights for each age group become... , , , Note the product of the two coefficients. When both conditions are triggered simultaneously, the highest age band, after being amplified, is attenuated, and its weight returns to its initial value. The other age bands only undergo an attenuation step, with their weights reduced to half of their initial value. This processing allows nodes to maintain their original sensitivity to the timeout risk of the highest age band while moderately reducing the conservative estimate of suspended tasks in ordinary age bands, resulting in a decrease in overall virtual load compared to when no conditions are triggered. This design is based on the following engineering observation: a timeout rate exceeding the threshold indicates latency in the acknowledgment link, but a simultaneous exceedance of the normal acknowledgment rate in the middle and high age bands indicates that the node's ability to handle suspended tasks has not deteriorated synchronously with the increase in timeout events. Therefore, it is unnecessary to maintain the same level of penalty for all age bands; only the highest age band retains a response strength equivalent to its original value, while the other age bands can be appropriately relaxed. The adjusted weight is denoted as... , where k=1,2,3,4.

[0054] The original calculation of virtual unpaid load is a weighted sum of the counts for each age group and their corresponding weights: However, this original value may have inconsistent dimensions due to differences in node throughput capabilities, requiring normalization; this embodiment uses a preset normalization constant. For the conversion, this embodiment uses a globally unified constant. Normalization was performed; this value was estimated based on the maximum number of suspended tasks per node in the system planning: based on the worst-case scenario of a concurrency limit of 2000 tasks per node and extreme concentration of counts in the highest age group. The weighted sum is calculated as follows (with the remainder being 0), and the upper limit is 2000 × 2.0 = 4000. However, this situation rarely occurs in actual operation (it requires all unclosed tasks to time out severely at the same time). Taking a typical value with a relatively balanced distribution under daily high-load scenarios—500 items per age band—the weighted sum is 500 × (0.2 + 0.5 + 1.0 + 2.0) = 1850. This value is rounded up to 2000 as a normalization constant, so that in most scenarios... It falls between 0 and 1, and when extreme cases occur, the normalized value slightly exceeds 1, in conjunction with a global suppression factor. It remains within a controllable range.

[0055] The normalized virtual load is However, directly using the normalized value may still cause a mismatch between the dimensions of the virtual load and the physical load. Therefore, in the subsequent weighted fusion step, the virtual load that actually participates in the fusion is the normalized virtual load multiplied by a global suppression factor. This inhibition factor is used to adjust the degree of influence of virtual load on the overall scheduling load, and its value ranges from 0 to 1. In this embodiment, it is set to 1. Therefore, the final output of the virtual unpaid load in step S4 is: .

[0056] For tasks with different closure levels, this embodiment maintains an independent set of age distribution statistics and a set of dynamic weight adjustment parameters for each level; for example, the first closure level (CAPTCHA type). It might be set to 2.5 to amplify the impact of timeouts, while the second closure level (marketing category) does not track receipts, and its virtual offset load is directly set to zero; when calculating the virtual offset load, the scheduler first selects the corresponding level's statistics and parameters based on the closure level field of the SMS request, calculates them separately, and then sums them up to obtain the total for that node. .

[0057] Step S4: Weighted Fusion: The physical load index and the virtual unpaid load are weighted and fused to obtain the overall scheduling load of the node;

[0058] Step S4 involves merging the physical load metrics obtained in step S1 with the virtual unpaid load calculated in step S3 into a single comprehensive scheduling load value for subsequent request allocation decisions; physical load metrics As defined in step S1, it is the weighted average of the current sending thread pool utilization and processor utilization of node i, i.e. In this embodiment Take 0.6; Virtual unpaid load After step S3, the result is a normalized value multiplied by a global suppression factor. The dimensionless value after that, i.e. ,in , Although the two indicators have the same dimensions, the numerical ranges of physical load and virtual load may differ significantly. Direct addition would cause the larger one to dominate the overall value. Therefore, this embodiment uses additive fusion instead of multiplication. Additive fusion can preserve the independent contribution of the two dimensions and avoid the extreme case where the whole is zero when one dimension is zero in the product form.

[0059] The specific fusion formula is as follows: ;here The global suppression factor and normalization processing are already implicitly included, so no additional coefficients are introduced; the statement that "the overall scheduling load equals the physical load index plus the product of the global suppression factor and the normalized virtual load" is reflected in this embodiment as follows: ,and This is the normalized virtual load; due to the output of step S3 It already includes The factors can be directly added in step S4; this design makes the fusion process transparent to the scheduler and eliminates the need to repeatedly configure the suppression factors.

[0060] For tasks with different closure levels, since virtual load is calculated independently for each level in step S3 (the virtual load for the second closure level is set to zero), different comprehensive scheduling loads will also be obtained during fusion in step S4. When allocating SMS requests, the scheduler compares the comprehensive scheduling load corresponding to the node's closure level based on the request's closure level field. If the request is at the first closure level, the comprehensive scheduling load including virtual load is used. If it is the second closed level, since the virtual load is zero, Degenerate into This ensures that high-criticality requests are more sensitive to dangling backlogs, while low-criticality requests are only concerned with real-time physical pressure.

[0061] After merging, each node corresponds to a floating-point number. The lower the value, the stronger the current overall carrying capacity of the node. The scheduler caches this value and waits for the allocation decision in step S6; the entire weighted fusion process does not introduce additional time windows or smoothing filters to ensure a rapid response to changes in node status.

[0062] Step S5: Credit Pre-check Step: When receiving an SMS request, parse its closure level field; for requests of the first closure level, query the credit management module for the available credit limit corresponding to the initiator's identity. If the limit is insufficient, reject the request; if the limit is sufficient, deduct one credit unit and continue processing; for requests of the second closure level, do not perform credit pre-check and do not track its receipt status.

[0063] Step S5 is executed when the scheduler receives an SMS request, before it is allocated to a specific node. The scheduler parses the closure level field in the request message. This field is an integer value. In this embodiment, the first closure level corresponds to the value 1, and the second closure level corresponds to the value 0. For requests with a closure level of 0, the scheduler skips the credit pre-check, does not query any credit limit information, and does not track the subsequent receipt status for the request. After such requests are submitted to the operator, the node abandons the maintenance of its delivery conclusion. For requests with a closure level of 1, the scheduler calls the credit management module, passing in the initiator's identity identifier (such as the business party's AppKey or signature field).

[0064] The credit management module maintains an independent token bucket data structure for each initiator's identity; the bucket parameters include: capacity. This indicates the maximum number of first-closed-level requests allowed per minute; the current number of tokens. Last timestamp added ; replenishment rate The unit is tokens per second; in this embodiment, it is set as follows: , (That is, 10 tokens are replenished per second, corresponding to 600 tokens per minute); the token bucket update logic is triggered during each credit pre-check: the module obtains the current system time. ,like (If the system clock rollback occurs), then... Reset to , If the current condition remains unchanged, no supplementary operation will be performed this time; the time difference will be recalculated on the next request. Otherwise, the time difference will be calculated. (Unit: seconds); if Then replenish the number of tokens. ,Will Updated to and will Set as ; Judge after completion of supplementation If so, then Decrease by 1 to return sufficient quota to the scheduler; otherwise, return insufficient quota. Requests with insufficient quota are directly rejected by the scheduler and do not enter the subsequent node allocation process.

[0065] A boundary case needs to be handled: when a request is received from an initiator, It is an empty value; at this time the module will Initialize to , Set as The request is allowed after deducting one token; this embodiment does not perform additional timeout or retry tracking for the first closure level request, and the credit pre-check only controls the request admission rate and does not interfere with the node-level receipt processing.

[0066] Example: The token bucket of business A , , , The scheduler received a closure level 1 request at 10:00:05. Seconds, replenish tokens , Not exceeding 600; after deducting 1 Please grant permission; if the same business party receives another request at 10:00:05:500, then... , Seconds, replenish tokens , After deducting 353, it is allowed; if 600 requests arrive consecutively within a short period of time, each request is supplemented with a small amount of tokens, but the total consumption exceeds Eventually, it will appear If the value is less than 1, subsequent requests will be rejected until the token is replenished naturally; for second-closed-level requests, regardless of the credit limit of business party A, they will be allowed directly without modifying the token bucket state.

[0067] Step S6: Distribute SMS requests: Based on the overall scheduling load of each node, distribute newly arrived SMS requests to the node with the lowest overall scheduling load.

[0068] Step S6 is the final stage of scheduling decision-making, occurring after the credit pre-check in step S5 is passed; at this point, the scheduler has obtained the closure level (first or second) of the SMS request and has calculated the corresponding overall scheduling load for each node. For requests at the first closure level, The result is obtained by fusing physical load and virtual compensated load in step S4; for requests at the second closure level, the virtual compensated load is set to zero. Directly equal to physical load The scheduler maintains a global node status table. Each record in the table contains the node identifier, the last reported timestamp, the node health status (healthy or faulty), and the comprehensive scheduling load value of the node for different closure levels. Faulty nodes have been excluded from the allocation scope by the fault succession mechanism in step S2 and do not participate in the allocation of any new requests.

[0069] When a new request arrives, the scheduler iterates through all healthy nodes and reads the node whose closure level matches the current request. Value, record the minimum value among them. and the corresponding set of nodes ;like If there is only one node, the request will be directly assigned to that node; otherwise... The system contains multiple nodes (i.e., nodes with the same overall scheduling load value). The scheduler selects one node by taking the hash value of the node identifier modulo the total number of nodes, or by selecting the first minimum node encountered in this traversal. This embodiment uses the latter to reduce computational overhead: during the traversal, the target node is updated only when a node that is strictly less than the current minimum value is encountered, and is not updated when a node that is equal to the current minimum value is encountered. Therefore, the node finally selected is the one that appears first in the traversal order among all the minimum value nodes.

[0070] The specific implementation of the allocation operation is as follows: The scheduler encapsulates the SMS request into an internal task object and sends it to the receiving queue of the target node through the network connection; after receiving it, the node decides whether to track the receipt based on the closure level of the request: the first closure level task will be recorded by the node in the local unclosed task table, the submission timestamp will be recorded and the age timer will be started; the second closure level task will be directly submitted by the node to the operator gateway and its tracking information will be discarded; the scheduler does not wait for the node to return the processing result and immediately starts processing the next request.

[0071] For example: The scheduler manages three healthy nodes Q1, Q2, and Q3; the current overall scheduling load for the first closure level is as follows: , , A new verification code SMS request arrives (first closure level). After traversal, the minimum value of 0.479 corresponds to nodes Q1 and Q3. Following the traversal order (assuming the node list order is Q1, Q2, Q3), Q1 is initially set as a candidate. When traversing to Q3, since the value is equal, it is not updated, and finally, node Q1 is selected. If the request is a marketing SMS (second closure level), the overall scheduling load of each node degenerates into the physical load, assuming the physical load is... , , If the minimum value of 0.35 is reached, node Q1 is selected. After allocation, the count of unclosed tasks for node Q1 increases, and the age distribution statistics in subsequent state reports will change accordingly, thus affecting the scheduling in the next round. This forms a closed-loop negative feedback; the scheduler does not additionally confirm or retry the allocation result. If a node does not receive a request due to a fault, the fault succession mechanism in step S2 handles the tracking of the allocated task's receipt.

[0072] The values ​​of all thresholds, weighting coefficients, and preset constants in this embodiment can be categorized into three types: engineering experience setting, statistical principle derivation, and system capacity estimation.

[0073] High percentile (p=99.5%): The tail high percentile value (Q99.5%) used to calculate the distribution of successful receipt time intervals. The purpose of selecting an extremely high percentile is to capture the long-tail distribution characteristics of receipt delay, while excluding the severe interference of individual extreme outliers (such as network interruptions) on the threshold, so that the timeout threshold can stably reflect the true tail delay of the operator's link.

[0074] Timeout threshold multiple The multiplier is used to map the statistical tail delay to the actual timeout judgment boundary. Regular tasks are given a 1.5x buffer to avoid premature timeouts caused by normal link jitter. Payment tasks have extremely high timeliness requirements (usually only 60~120 seconds of validity period) and need to be closely close to the actual delay distribution, so it is set to 1.0 to accelerate the entry into the old band.

[0075] Dynamic timeout limit ( As an "upper limit protection" mechanism, if the tail delay increases abnormally (such as the operator channel failure), the timeout threshold will not increase indefinitely. If it exceeds 60 seconds, it will be truncated and an alarm will be triggered, because the receipt delay exceeding this threshold is usually meaningless (verification codes, etc. have become invalid), and compensation logic must be forcibly triggered.

[0076] Initial values ​​( ): When a node starts up or the number of response samples is insufficient (<200 times), the "system preset conservative value" is used, taking a typical initial response window of 3 seconds based on the operator's gateway, which is a conservative estimate in engineering.

[0077] Sample statistics window (N=200 successful receipts): 200 samples are sufficient to statistically support the stability of percentile calculation (avoiding small sample noise), while memory and computational overhead are controllable; the statistical update cycle is 10 seconds, taking into account both rapid response to link changes and computational load.

[0078] Heartbeat timeout threshold (1.5 seconds): A node normally reports at 500 millisecond intervals. If there is no response for 1.5 seconds covering 3 consecutive reporting cycles, it is sufficient to determine that the node is faulty (excluding brief network jitter).

[0079] Base age weight ( The older the task, the higher the compensation cost after timeout or failure; the weight increases exponentially (from 0.2 to 2.0), reflecting the non-linear growth of the deteriorating impact of older tasks on delivery quality; the initial value is based on the system's engineering calibration of the "relative risk of aging degree";

[0080] Historical receipt sliding window (M=500): Take the timeout rate and the normal receipt rate of the most recent 500 closed tasks to ensure that the statistical data is both smooth and real-time, and avoid the drastic oscillation of weights caused by a single fluctuation.

[0081] Alarm threshold ( Timeout rate 15% This value is considered an engineering warning line for abnormality in the receipt link or bottleneck in node processing. If this value is exceeded, the risk of the highest age band needs to be amplified.

[0082] Dominance threshold ( ): 70% of responses came from middle- and older age groups. This indicates that although timeout events exist, the overall processing capacity of nodes for suspended tasks has not deteriorated, so the load forecast for the average age group can be appropriately reduced.

[0083] Bucket capacity ( The replenishment rate (r=10 tokens / second) corresponds to 600 first-level closure requests allowed per minute, or 10 per second. This value is based on the single-business critical SMS injection limit of the system design and belongs to the preset quota at the capacity planning level. It is used to smooth out sudden traffic from the entry point and prevent a single initiator from overwhelming the receipt processing resources.

[0084] This embodiment expands the input to the scheduling decision from a single physical load to a multi-dimensional state space including the age distribution of suspended tasks through the coordinated operation of steps S1 to S6. The age distribution statistics collected in step S1 reveal the aging structure of unclosed tasks within a node. Step S3 calculates the virtual pending load based on the dynamic timeout threshold and the adjusted weights of the receipt history. This value precedes the actual receipt storm in time, allowing the scheduler to sense potential pressure before the physical load increases. After merging the physical and virtual loads in step S4, step S6 allocates requests according to the merged value, prioritizing new tasks to the node with the lowest overall load and avoiding tasks assigned to suspended tasks. Nodes experiencing severe load buildup or delayed receipt processing continue to increase their workload; the fault succession mechanism in step S2 transfers the tracking responsibility of suspended tasks to the successor node when a node fails, preventing the loss of receipt matching opportunities during the failure period; the credit pre-inspection in step S5 controls the injection rate of critical requests from the source, preventing a sudden surge in traffic from a single initiator from overwhelming the receipt processing capacity of each node; under the combined effect of the above mechanisms, the distribution of suspended backlog among nodes tends to be balanced, the tail high quantile of receipt processing delay decreases periodically, the proportion of timeout closure events converges below the preset alarm threshold, and the on-time delivery rate of critical SMS messages moves out of the fluctuation range and stabilizes above the level required by the business.

[0085] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.

[0086] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0087] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for load prediction based short message platform multi-node dynamic scheduling and resource optimization, characterized in that, include: Obtain node status data: Obtain node status data reported by each SMS sending node. The node status data includes physical load indicators and age distribution statistics of unclosed tasks. Unclosed tasks refer to SMS tasks that have been submitted to the operator but have not yet received the final delivery receipt. Calculate the virtual compensation load: Based on the age distribution statistics reported by each node and the dynamic timeout threshold maintained by the node, calculate the virtual compensation load of the node. The dynamic timeout threshold is updated in real time based on the time interval distribution of the most recently received successful receipts by the node. Weighted fusion: The physical load index and the virtual unpaid load are weighted and fused to obtain the overall scheduling load of the node; Distribute SMS requests: Based on the overall scheduling load of each node, newly arriving SMS requests are distributed to the node with the lowest overall scheduling load.

2. The load prediction based short message platform multi-node dynamic scheduling and resource optimization method according to claim 1, characterized in that, The method for obtaining the age distribution statistics of unclosed tasks is as follows: divide the age zone into multiple age zones according to the age zone boundaries determined by the dynamic timeout threshold, and count the number of unclosed tasks in each age zone respectively. The dynamic timeout threshold is positively correlated with the high quantile of the time interval distribution of the successful receipt, and different timeout threshold multiples are corresponding to tasks with different closure levels. 3.The method of claim 1, wherein, During the calculation of virtual unpaid load, different basic weight values ​​are assigned to unclosed tasks of different age groups, with higher weight values ​​for older age groups. Furthermore, the basic weight values ​​are dynamically adjusted based on the recent receipt processing history of the node: when the proportion of recent normal receipts belonging to the middle and high age groups exceeds a preset advantage threshold, all weight values ​​are reduced; when the proportion of recent timeout closure events exceeds a preset alarm threshold, the weight value of the highest age group is increased separately.

4. The load prediction based short message platform multi-node dynamic scheduling and resource optimization method of claim 1, wherein, It also includes a credit pre-check step: when receiving an SMS request, it parses the closure level field; for requests at the first closure level, it queries the credit management module for the available credit limit corresponding to the initiator's identity; if the limit is insufficient, it rejects the request; if the limit is sufficient, it deducts one credit unit and continues processing; for requests at the second closure level, it does not perform a credit pre-check and does not track the receipt status.

5. The load prediction based short message platform multi-node dynamic scheduling and resource optimization method according to claim 4, characterized in that, The credit limit is managed as follows: a token bucket is maintained for each initiator identity, and the bucket capacity is the number of first-closed-level requests allowed per minute; during each credit pre-check, tokens are replenished at a fixed rate based on the difference between the current time and the last replenishment time; if the number of tokens is not less than one unit, the tokens are deducted and allowed, otherwise they are rejected.

6. The load prediction based short message platform multi-node dynamic scheduling and resource optimization method according to claim 1, characterized in that, It also includes a fault succession step: the scheduler maintains the last reported status summary of each node, the status summary includes the age distribution statistics of unclosed tasks and the tail high quantile of the current successful acknowledgment time interval distribution; when a node fails to report within the timeout period, it is determined to be a fault node, a successor node is selected from the healthy nodes, and the status summary of the fault node is sent to the successor node. The successor node creates a virtual closed tracking pool based on this. When it receives a receipt, it first searches in its own real pool. If it does not find the receipt, it matches the virtual pool according to age and priority and deducts the corresponding count.

7. The load prediction based short message platform multi-node dynamic scheduling and resource optimization method according to claim 6, characterized in that, The virtual closed tracking pool maintains the remaining counts for each age group; the successor node sets a maximum survival time for the virtual pool, which is proportional to the tail high quantile of the current successful acknowledgment time interval distribution of the successor node; when the virtual pool survives beyond this time and the remaining count is not cleared, the remaining count is converted into a timeout compensation event, written to the compensation queue, and the virtual pool is destroyed. 8.The load prediction based short message platform multi-node dynamic scheduling and resource optimization method of claim 3, wherein, In the dynamic adjustment of the basic weight value, first determine whether the proportion of timeout closure events exceeds the preset alarm threshold. If it does, first multiply the basic weight value of the highest age group by the first amplification coefficient. Then determine whether the proportion of normal receipts in the middle and high age groups exceeds the preset advantage threshold. If it does, multiply all weight values ​​by the first attenuation coefficient. The product of the first amplification factor and the first attenuation factor is no greater than 1. 9.The load prediction based short message platform multi-node dynamic scheduling and resource optimization method of claim 1, wherein, The physical load metric is the weighted average of the current sending thread pool utilization and processor utilization of the node; the weighted fusion method is as follows: the comprehensive scheduling load is equal to the sum of the physical load metric and the virtual uncompensated load, wherein the virtual uncompensated load is the product of the normalized virtual load and the global suppression factor, and the normalized virtual load is the original virtual load divided by the preset normalization constant. 10.The load prediction based short message platform multi-node dynamic scheduling and resource optimization method of claim 6, wherein, The replacement node is selected as follows: from the healthy nodes, select the node that uses the same operator channel type as the faulty node; if there is no node with the same channel type, select the node with the highest channel type compatibility; from the selection results, select the node with the lowest current overall scheduling load as the replacement node.