Intelligent Scheduling Method for SMS Sending Channels Based on Multi-Source Data Fusion
Through the intelligent scheduling method of SMS sending channel fusion with multi-source data, the neglect of the matching relationship between SMS content characteristics and channel processing capabilities is solved, and the efficient and reliable SMS sending system is realized, especially when processing SMS rich text content, it significantly improves the success rate and timeliness.
Patent Information
- Application Number
- CN202510631297.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing SMS sending channel scheduling method fails to effectively consider the matching relationship between SMS content characteristics and channel processing capabilities, resulting in SMS with different content characteristics showing different sending success rates and delay characteristics on the same channel, and lacking a detailed description of the differences between channels, resulting in the channel selection strategy being too generalized and failing to make full use of the strengths of each channel.
By constructing a multi-source data fusion intelligent scheduling method for SMS sending channels, obtain and clean multi-source data, extract the basic and semantic features of SMS, calculate feature interaction relationships, build content feature intensity vectors and channel affinity vectors, combine business attributes and channel load states, generate a comprehensive affinity scoring matrix, and finally calculate the optimal channel for scheduling.
It improves the success rate and timeliness of SMS sending, and is especially suitable for processing SMS with rich text content, significantly improving the processing accuracy and delivery time of SMS with complex content.
Smart Images

Figure CN120151781B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a communication method, in particular to an intelligent scheduling method for a short message sending channel based on multi-source data fusion. Background Art
[0002] With the development of the mobile Internet and the diversification of short message service application scenarios, short messages continue to play an irreplaceable role as an important information transmission method. Especially in key business scenarios such as financial payment verification, account security, and transaction notifications, the reliability and timeliness of short messages directly affect user experience and business security. Therefore, how to achieve efficient and reliable short message sending in a complex and changeable network environment has become an important technical challenge faced by communication service providers and enterprise users.
[0003] Currently, the commonly used short message channel scheduling methods in the industry are mainly divided into several categories: a static priority strategy based on rules, which sends messages in the order of a preset channel priority list; a simple dynamic scheduling based on historical data statistics, which selects the channel with the best performance by analyzing indicators such as the historical success rate and delay of each channel; a feedback scheduling based on real-time monitoring, which dynamically adjusts the sending strategy according to the current state of the channel. Some more advanced systems have begun to introduce machine learning technologies, such as algorithms like decision trees and random forests, to build channel selection models, considering factors such as operators, regions, and time periods to improve the scheduling accuracy. There are also researchers who have tried to apply reinforcement learning to channel selection, continuously optimizing the scheduling strategy through trial-and-error learning to improve the system's adaptability to environmental changes.
[0004] However, the existing technical solutions still have several technical bottlenecks: First, the existing solutions mainly focus on channel status and performance indicators when making channel selection decisions, ignoring the matching relationship between short message content characteristics and channel processing capabilities, resulting in different sending success rates and delay characteristics for short messages with different content characteristics (such as those containing special characters, very long short messages, or URL links) on the same channel, making it difficult to carry out targeted optimization; Second, the channel has insufficient modeling of the feature intensity differences of feature combinations and lacks a fine-grained description of the differences between channels, resulting in the channel selection strategy being too general and failing to make full use of the strengths of each channel; Third, most existing systems adopt static or simple time-decay data processing mechanisms and do not design adaptive memory and forgetting strategies for the different change rate differences of different feature dimensions, making it difficult to balance the stability of historical data and the real-time response ability. These technical problems seriously limit the performance of the short message sending system in processing diverse content scenarios. Summary of the Invention
[0005] The object of the invention is to provide an intelligent scheduling method for a short message sending channel based on multi-source data fusion, in order to solve at least one technical problem existing in the prior art.
[0006] Technical solution: An intelligent scheduling method for SMS sending channels based on multi-source data fusion, comprising the following steps:
[0007] Obtain the basic data including the original SMS text, perform cleaning, standardization, and spatio-temporal alignment to obtain a multi-source data set;
[0008] Extract the basic features and semantic features of the original SMS text, calculate the feature interaction relationship, and construct a content feature intensity vector;
[0009] Based on the multi-source data set, analyze the historical processing capabilities of the channels, construct a channel differential interaction feature intensity model, perform active detection, and calculate a set of channel affinity vectors;
[0010] Calculate the vector space similarity between the content feature intensity vector and the set of channel affinity vectors, and combine the service attributes and channel load status to obtain a comprehensive affinity score matrix;
[0011] Based on the comprehensive affinity score matrix, use a pre-constructed multi-objective weighting function to calculate the optimal channel and generate a channel scheduling decision.
[0012] Beneficial effects: The present invention effectively improves the success rate and timeliness of SMS sending, and is particularly suitable for processing rich text content SMS. The relevant technical effects will be described in detail below in combination with cases. Description of the drawings
[0013] Figure 1 is the flowchart of the present invention.
[0014] Figure 2 is the flowchart of constructing the content feature intensity vector of the present invention.
[0015] Figure 3 is the flowchart of constructing the set of channel affinity vectors of the present invention.
[0016] Figure 4 is the flowchart of calculating the vector space similarity between the content feature intensity vector and the set of channel affinity vectors to obtain the comprehensive affinity score matrix of the present invention.
[0017] Figure 5 is the flowchart of constructing the feature interaction matrix of the present invention. Detailed implementation manners
[0018] As Figures 1 to 5 shown, an intelligent scheduling method for SMS sending channels based on multi-source data fusion is provided, comprising the following steps:
[0019] S1. Obtain the original data from multiple data sources, clean, standardize, and preliminarily fuse the data to provide a basis for subsequent analysis.
[0020] S11. Read the data generated by the channel status real-time monitoring system, including indicators such as channel availability flags, channel congestion levels, and channel response times. Process the fluctuations through a time window moving average algorithm to obtain the standardized channel status vector C_status.
[0021] S12. Read the data in the SMS sending history record library, including historical SMS content, sending channels, sending results, timestamps, etc. Apply an incremental data acquisition strategy to only obtain new records and output the incremental historical sending record set H_data.
[0022] S13. Read the data such as SMS service types, priority flags, and timeliness requirements provided by the business system. Quantify the business requirements into computable indicators through a standardized interface to generate the business attribute vector B_attr.
[0023] S14. Read the incremental historical sending record set H_data, detect and process missing values, outliers, and duplicate records. Convert the status codes and time formats used by different data sources into a unified standard and output the cleaned historical data set H_clean.
[0024] S15. Read the standardized channel status vector C_status, the cleaned historical data set H_clean, and the business attribute vector B_attr. Align the data based on timestamps and channel IDs, create a data association index, and generate a spatio-temporal aligned multi-source data set D_aligned to provide a unified data view for subsequent analysis.
[0025] S2. Analyze the SMS content, construct a multi-dimensional content feature intensity vector, and quantify the feature intensities of the SMS content in different dimensions.
[0026] S21. Read the original SMS content M_text to be sent, and calculate its basic features respectively through a feature extraction algorithm, including text length, number of segments, character set distribution, proportion of special symbols, number of URLs, etc., and output the basic feature vector F_basic.
[0027] S22. Read the original SMS content M_text, and apply lightweight natural language processing technology to identify the semantic features of the content, including message type recognition (such as verification code, notification, marketing), timeliness expression recognition, importance assessment, etc., and generate the semantic feature vector F_semantic.
[0028] S23. Read the basic feature vector F_basic and the semantic feature vector F_semantic, construct a feature interaction matrix, and apply a context-sensitive interaction intensity calculation method to identify the synergistic effects between features, such as the interaction between URLs and special characters, the interaction between segments and key information, etc., and output the feature interaction matrix F_interaction.
[0029] S24. Read the basic feature vector F_basic, semantic feature vector F_semantic, and feature interaction matrix F_interaction. Map each feature and its interaction effect to a standardized feature intensity space through a feature intensity mapping function to obtain a multi-dimensional content feature intensity vector S, where each dimension represents the sensitivity of a content feature.
[0030] S3. Based on historical transmission data and detection results, construct and maintain a multi-dimensional affinity vector for each channel, which characterizes the processing ability of the channel for different content features.
[0031] S31. Read the spatio-temporally aligned multi-source dataset D_aligned, group it by channel ID, analyze the historical processing ability of each channel for different content features, calculate the conditional success rate and processing delay, and output the channel historical processing ability matrix H_capability.
[0032] S32. Read the data of the channel historical processing ability matrix H_capability and the historical feature interaction matrix F_interaction, analyze the differential feature intensity of each channel for a specific feature combination, construct a channel-specific interaction feature intensity model, and output the channel interaction feature intensity matrix C_sensitivity.
[0033] S33. Based on the uncertain region in the channel interaction feature intensity matrix C_sensitivity, generate a targeted detection SMS set P_messages, send it through each channel and collect the results, update the evaluation of the processing ability of the channel for specific features, and output the detection result dataset P_results.
[0034] S34. Read the channel historical processing ability matrix H_capability, the channel interaction feature intensity matrix C_sensitivity, and the detection result dataset P_results, apply the differential memory and forgetting mechanism, calculate and update the multi-dimensional affinity vector of each channel, and output the channel affinity vector set A, where each vector A_k represents the processing ability of a channel for each content feature.
[0035] S4. Determine the matching degree between specific content and each channel through the matching calculation of the content feature intensity vector and the channel affinity vector.
[0036] S41. Read the content feature intensity vector S of the SMS to be sent and the channel affinity vector set A of all available channels, calculate the similarity of each pair (S, A_k) in the multi-dimensional vector space, and generate the initial affinity score set R_initial.
[0037] S42. Read the initial affinity score set R_initial and the current business attribute vector B_attr. Based on business requirements and the sending context, calculate the dynamic weights for each affinity dimension, adjust the initial scores by weighting, and output the context-weighted affinity score set R_weighted.
[0038] S43. Read the normalized channel status vector C_status, calculate the impact factors of the current load of each channel on affinity matching, pay attention to the performance decay law of high-load channels, and output the load impact factor set L_factors.
[0039] S44. Read the context-weighted affinity score set R_weighted and the load impact factor set L_factors, apply a non-linear combination function to calculate the final content-channel affinity score, and output the comprehensive affinity score matrix R_final, indicating the final matching degree of the SMS to be sent with each available channel.
[0040] S5. Considering factors such as content-channel affinity, channel real-time status, and business requirements comprehensively, formulate an optimal channel scheduling strategy that meets multi-objective constraints.
[0041] S51. Read the business attribute vector B_attr, construct a scheduling objective function according to business requirements, including success rate objectives, latency objectives, cost objectives, load balancing objectives, etc., set the weights of each objective, and obtain the multi-objective weighted function G.
[0042] S52. Read information such as the normalized channel status vector C_status, the service level agreement SLA of the current business, and system capacity limits, construct a set of constraint conditions for scheduling decisions, including channel capacity constraints, latency constraints, and cost constraints, etc., and output the constraint condition set K.
[0043] S53. Read the content feature intensity vector S, design a risk-aware exploration-exploitation balance strategy, adopt a conservative strategy for highly sensitive content and an active exploration strategy for low-sensitive content, and output the exploration rate parameter E.
[0044] S54. Read the comprehensive affinity score matrix R_final, the multi-objective weighted function G, the constraint condition set K, and the exploration rate parameter E, apply a risk-aware multi-objective optimization algorithm, and select the optimal sending channel or channel combination under the premise of meeting the constraint conditions, and output the channel scheduling decision D, including the main channel selection and the backup channel strategy.
[0045] In another embodiment of the present application, the process of content feature intensity feature extraction in S2 is specifically as follows:
[0046] S21. Read the original text content M_text of the SMS to be sent and calculate its basic features through a feature extraction algorithm:
[0047] 1. Calculate the total text length L_total, including counting the number of characters and bytes;
[0048] 2. Analyze the number of segments N_segment required for the SMS and the length distribution of each segment;
[0049] 3. Statistically analyze the character set distribution vector C_dist, including the proportion of categories such as ASCII characters, non-ASCII characters, numbers, punctuation, etc.;
[0050] 4. Identify and count the special symbol set S_special and its position distribution in the text;
[0051] 5. Detect the number of URLs N_url and their features, including URL length, domain name type, and URL complexity;
[0052] 6. Summarize the above features to construct a basic feature vector F_basic = {L_total, N_segment, C_dist, S_special, N_url,...};
[0053] S22. Read the original text content M_text of the SMS and adopt lightweight natural language processing:
[0054] 1. Use rule-based pattern matching and keyword analysis to identify the message type T_type (verification code, notification, marketing, service, etc.);
[0055] 2. Extract the time expressions and timeliness-related vocabulary in the text to evaluate the timeliness score T_urgency;
[0056] 3. Analyze the sentiment tendency and importance marking words of the text to calculate the importance score T_importance;
[0057] 4. Identify the industry feature words in the text and classify them into industry categories T_industry;
[0058] 5. Integrate the above semantic features to generate a semantic feature vector F_semantic = {T_type, T_urgency, T_importance, T_industry,...};
[0059] S23. Read the basic feature vector F_basic and the semantic feature vector F_semantic, and apply a dynamic weight adaptive mechanism for feature interaction and context sensitivity analysis:
[0060] 1. Initialize the feature interaction matrix F_interaction as an n×n zero matrix (n is the total number of features);
[0061] 2. For each pair of features (i, j), perform the following: a. Calculate the basic interaction intensity I_base(i, j) = F[i] × F[j]; b. Extract the historical feature intensity data of the channel pair feature pair (i, j) from the spatio-temporally aligned multi-source dataset D_aligned, and calculate the channel-related interaction weight W_channel(i, j); c. Calculate the time decay factor D_time(i, j) according to the timeliness difference between features i and j; d. Analyze the relative position relationship between features i and j in the text structure, and calculate the position correlation factor P_pos(i, j); e. F_interaction[i, j] = I_base(i, j) × W_channel(i, j) × D_time(i, j) × P_pos(i, j);
[0062] 3. Apply feature interaction recognition specialized for the SMS field: a. Analyze the interaction I_url_special between URLs and special characters, and detect the distribution and encoding method of special characters in URLs; b. Evaluate the interaction I_seg_key between segments and key information, and analyze the distribution of key information in different segments; c. Calculate the interaction I_char_len between character encoding and length, and identify the impact of different encoded characters on length calculation; d. Identify the interaction I_time_content between timeliness markers and content types, and evaluate the performance differences of timeliness in different content types;
[0063] 4. Integrate the identified domain-specific interactions into the feature interaction matrix F_interaction;
[0064] 5. Apply non-linear transformation and normalization processing, and output the final feature interaction matrix F_interaction.
[0065] In order to overcome the limitations of static weights or uniform training weights in traditional feature analysis, the feature interaction weights are dynamically adjusted according to the historical behavior of the channel. By analyzing the sending results of feature combinations in historical data, the sensitivity of the channel to specific feature combinations is extracted, and time decay and position correlation are considered to make the weights more flexible and context-related. In the SMS sending scenario, feature interactions such as special characters and URL combinations, segmentation and key information distribution have a significant impact on the sending results. The dynamic weight mechanism can accurately capture these impacts and optimize channel selection decisions. Actual measurements show that compared with the static weight method, the system's processing accuracy for complex content SMS is improved by 22.5%. In particular, when processing SMS with multiple special feature combinations, the accuracy of channel selection is significantly improved, and it can adapt to changes in the sensitivity of feature combinations in different periods and channels.
[0066] By analyzing the effects of feature interactions in specific locations and environments, rather than focusing only on interactions at the global statistical level. Specifically, by calculating distance functions between feature instances, position distribution pattern scores, and other methods, the influence of the relative position relationship of features in the text structure on the intensity of interaction is evaluated. In the field of SMS sending, the distribution position of features in the text often affects the channel processing capability. For example, when the URL is at the beginning or end of the text, or when special characters are scattered or concentrated, the channel performance varies significantly. Test data shows that after using context-sensitive interaction analysis, the sending success rate under a specific location distribution pattern can be more accurately predicted, with the prediction accuracy improved by about 18.3%. It can capture location-related patterns that are ignored by traditional methods, making channel matching more refined, especially for complex SMS containing multiple URLs or uneven distribution of special characters.
[0067] S24, read the basic feature vector F_basic, the semantic feature vector F_semantic and the feature interaction matrix F_interaction, and perform feature intensity mapping:
[0068] 1. Initialize the content feature intensity vector S to be a zero vector with a dimension of the predefined feature intensity space dimension m;
[0069] 2. For each feature strength dimension k (k=1,2,...,m): a. Identify the basic feature subset F_sub_basic and the semantic feature subset F_sub_semantic related to this dimension b. Extract the feature interaction submatrix F_sub_interaction related to this dimension c. Apply the feature strength mapping function: S[k] = f_map(F_sub_basic, F_sub_semantic, F_sub_interaction);
[0070] 3. Apply regularization processing to each dimension of S to ensure that the value range is within the interval [0, 1].
[0071] 4. Output the final m-dimensional content feature intensity vector S = {s1, s2, ..., s m}, where each dimension represents the sensitivity of a content feature: s1 is the character set feature intensity, indicating the influence degree of the character set characteristics of the short message content on the channel processing; s2: length feature intensity, indicating the influence degree of the short message length and its segmentation characteristics on the channel processing; s3: segmentation feature intensity, indicating the influence degree of the number of short message segments and the segmentation characteristics on the channel processing; s4: link feature intensity, indicating the influence degree of the URL link in the short message on the channel processing; s5: timeliness feature intensity, indicating the influence degree of the timeliness of the short message content on the channel processing; o... (other feature intensity dimensions).
[0072] The calculation process of s1 is as follows: Analyze the Unicode block distribution, calculate the proportion of special characters, and identify rare and error-prone characters.
[0073] The calculation process of s2 is as follows: Consider the relationship between the total length and segmentation; evaluate boundary conditions (close to the segmentation threshold); calculate information entropy and redundancy.
[0074] The calculation process of s3 is as follows: Analyze the number of segments; evaluate the integrity of each segment; calculate the cross-segment semantic coherence;
[0075] The calculation process of s4 is as follows: Identify the number and location of URLs; analyze the length and complexity of URLs; evaluate the domain name type and security level.
[0076] The calculation process of s5 is as follows: Identify timeliness keywords; analyze digital time expressions; evaluate the timeliness relevance of the content.
[0077] In another embodiment of the present application, the construction of the S3 and channel affinity vector is specifically as follows:
[0078] S31. Read the spatio-temporally aligned multi-source dataset D_aligned and analyze the channel processing capabilities:
[0079] 1. Group the data by channel ID to establish the channel historical dataset H_channel;
[0080] 2. Execute for the historical data of each channel c: a. Calculate the conditional success rate P(success|feature) of the channel for different content features; b. Evaluate the delay distribution statistics(delay|feature) of the channel for processing different feature contents; c. Analyze the time series pattern of the channel performance and identify periodic and trend changes pattern(performance);
[0081] 3. Construct the channel history processing capability matrix H_capability, where each row represents a channel and each column represents the processing capability index for specific content features;
[0082] 4. Apply Bayesian correction to handle the data sparsity problem and ensure the reliability of the processing capability evaluation for low-frequency features;
[0083] S32. Read the data of the channel history processing capability matrix H_capability and the historical feature interaction matrix F_interaction, and construct a channel-differentiated interaction feature intensity model:
[0084] 1. Initialize the channel feature interaction feature intensity matrix Sc for each channel c as an n×n identity matrix;
[0085] 2. For each pair of features (i, j), perform the following:
[0086] a. Extract the historical processing data histData(c, i, j) of this channel for the feature pair (i, j) from the channel history dataset H_channel;
[0087] b. Calculate the conditional processing probabilities: p_success_ij = P(success|i exists and j exists); p_success_i = P(success|i exists); p_success_j = P(success|j exists);
[0088] c. Calculate the channel-specific interaction feature intensity: interactionSensitivity = information gain(p_success_ij, p_success_i, p_success_j);
[0089] d. Apply Bayesian correction to process the sparse data: Sc[i, j] = Bayesian correction(interactionSensitivity, number of samples).
[0090] 3. Perform clustering analysis for each channel to identify the channel-specific sensitive interaction pattern Pc;
[0091] 4. Output the channel interaction feature intensity matrix C_sensitivity, which contains the feature interaction feature intensity data of all channels.
[0092] The implementation process is as follows: Obtain the current channel affinity vector A = (a1, a2,..., a n) Obtain the content sensitivity vector S of the newly sent short message; obtain the sending result R (success / failure / delay);
[0093] Calculate the feedback adjustment amount for each dimension, including, for each dimension i, if S[i] > threshold T (high sensitivity), δ i = Calculate the dimension adjustment amount (S[i], R); a ’ i = Update the affinity value (a i , δ i ) Otherwise a ’ i = a i and remain unchanged in the case of low sensitivity.
[0094] Apply non-linear smoothing and normalization, including: A ’ = Smooth and normalize (A ’ ) Apply the forgetting mechanism A ’ = Apply the affinity forgetting mechanism (A ’ , time factor); Update the channel affinity vector library, including:
[0095] Active detection strategy, including: If the detection condition is met, generate a detection short message for a specific dimension; arrange the sending of the detection short message; record the detection task.
[0096] A dedicated interactive recognition algorithm is designed for the content features unique to the short message sending scenario, including URL and special character interaction, segmentation and key information interaction, character encoding and length interaction, time limit marking and content type interaction, etc. It makes up for the deficiencies of the general feature interaction model in the short message field and can capture the feature combination effects rarely concerned in fields such as recommendation systems. In practical applications, this mechanism significantly improves the system's ability to handle boundary cases, such as scenarios where special characters in the URL cause encoding problems and key verification information is truncated across segments. The test results show that after adopting domain-specific interactive recognition, the success rate of short message sending in boundary cases has increased by 25.7%, and the delivery time has been reduced by 38.4%. This improvement is directly attributed to the system's ability to identify and prevent problems that may be caused by specific combinations, select the most suitable channel to handle such special cases in advance, and reduce the number of sending failures and retries.
[0097] In another embodiment of the present application, S33, the active detection strategy is executed, specifically:
[0098] Based on the uncertain area in the channel interaction feature strength matrix C_sensitivity, perform active detection:
[0099] 1. For the feature intensity matrix Sc of each channel c, calculate the uncertainty score uncertainty(i, j) = variance(Sc[i, j]);
[0100] 2. Identify the feature pairs (i, j) with high uncertainty and construct the detection priority queue Q_priority;
[0101] 3. Generate a targeted detection SMS set P_messages according to the detection priority queue, where each detection SMS has a specific feature combination;
[0102] 4. Design a detection execution plan, including detection time, channel allocation, and sample size;
[0103] 5. Execute the detection plan and send detection SMS through each target channel;
[0104] 6. Collect detection results, including sending status, delay, and delivery status, and construct a detection result data set P_results;
[0105] 7. Analyze the detection results and update the evaluation of the processing ability of the channel for specific feature combinations;
[0106] In another embodiment of the present application, S34, affinity vector calculation and update are specifically as follows:
[0107] Read the channel historical processing ability matrix H_capability, the channel interaction feature intensity matrix C_sensitivity, and the detection result data set P_results, and calculate and update the channel affinity vector:
[0108] 1. Initialize the m-dimensional channel affinity vector Ac for each channel c;
[0109] 2. For each affinity dimension k (k = 1, 2,..., m, corresponding to the dimension of the feature intensity vector): a. Extract the historical processing ability data related to this dimension from H_capability b. Extract the interaction feature intensity data related to this dimension from C_sensitivity c. Extract the detection result data related to this dimension from P_results d. Apply the differential memory and forgetting mechanism:
[0110] Calculate the historical data weight w_t = f_decay(t, half-life τ_k) for different time windows;
[0111] Set different half-lives τ_k for different dimensions k to reflect the difference in their change rates;
[0112] Weighted fusion of data from different periods: A_c[k] = Σ(data_t × w_t) / Σw_t;
[0113] 3. Normalize A_c to ensure consistency in the value range;
[0114] 4. Output the channel affinity vector set A, which contains the affinity vectors {A_1, A_2, ..., A_n} of all channels.
[0115] In another embodiment of the present application, S4, content-channel affinity matching, specifically is:
[0116] S41. Read the content feature intensity vector S of the text message to be sent and the channel affinity vector set A of all available channels, and calculate the vector space similarity:
[0117] 1. Perform the following operations on the affinity vector A_c of each channel c: a. Calculate the cosine similarity between S and A_c: cos_sim(S, A_c) = (S•A_c) / (||S|| × ||A_c||) b. Calculate the Euclidean distance between S and A_c: euc_dist(S,A_c) = √(Σ(S[i] - A_c[i]) 2 ) c. Calculate the Mahalanobis distance between S and A_c, considering the correlation between dimensions: mahal_dist(S, A_c) = √((S - A_c) T Σ -1 (S - A_c));
[0118] 2. Combine multiple distance metrics to calculate the comprehensive similarity score: sim_score(S, A_c) = f_combine(cos_sim, euc_dist, mahal_dist);
[0119] 3. Sort the similarity scores of all channels and output the initial affinity score set R_initial = {r_1, r_2, ..., r_n};
[0120] S42. Read the initial affinity score set R_initial and the current service attribute vector B_attr, and apply the dynamic weight adaptive mechanism of feature interaction:
[0121] 1. According to the service type, priority, and timeliness requirements in the service attribute B_attr, calculate the importance weight vector W = {w1, w2, ..., w m};
[0122] 2. Perform the following operations on the initial affinity score \(r_c\) for each channel \(c\): a. Calculate the dimension-weighted affinity score: \(r\) ’ _{c}=\frac{\sum(\text{sim_score_dim}(S[i], A_c[i])\times W[i])}{\sum W[i]}\text{ b. Consider the business-specific weighting factor:}r ’’ _{c}=r ’ _{c}\times f_{\text{business}}(B_{\text{attr}}, c);
[0123] 3. Output the context-weighted affinity score set \(R_{\text{weighted}}=\{r ’’ _1, r ’’ _2,\cdots, r ’’ _n\};
[0124] S43. Read the normalized channel status vector \(C_{\text{status}}\) and calculate the load impact factor:
[0125] 1. Perform the following operations on the current status of each channel \(c\): a. Extract load metrics such as channel congestion \(congestion_c\) and response time \(response_time_c\); b. Based on historical data, establish a load - performance decay model: \(decay_model_c = f_{\text{decay}}(congestion, response_time)\); c. Calculate the performance decay factor under the current load: \(l_c = decay_model_c(congestion_c, response_time_c)\);
[0126] 2. Normalize the load factor to ensure proportional consistency;
[0127] 3. Output the load impact factor set \(L_{\text{factors}}=\{l_1, l_2,\cdots, l_n\};
[0128] S44. Read the context-weighted affinity score set \(R_{\text{weighted}}\) and the load impact factor set \(L_{\text{factors}}\) and calculate the final score:
[0129] 1. Perform the following operations on each channel \(c\): a. Apply the non - linear combination function: \(r_{\text{final}_c}=f_{\text{nonlinear}}(r ’’ _c, l_c)\); b. Adjust the credibility of the score considering data sufficiency: \(r_{\text{final}_c}=r_{\text{final}_c}\times confidence_c\);
[0130] 2. Construct the comprehensive affinity scoring matrix R_final, which includes the scoring of the final matching degree between the SMS to be sent and each available channel.
[0131] In another embodiment of the present application, extract context-sensitive feature interactions, including:
[0132] Extract the basic feature set F = (f1, f2, ..., f n ); Initialize the context-sensitive interaction set C = {};
[0133] Identify the feature context relationship: For each feature instance f in the text i_inst : Extract its context window context = getContext(f i_inst );
[0134] For each feature instance f in the context j_inst : Calculate the context-dependent interaction strength ctxStrength = calculateContextDependentStrength(f i_inst , f j_inst , context);
[0135] Identify the context pattern sensitive to a specific channel, channelSensitivity = identifyChannelSensitivePattern(f i_inst , f j_inst , context); If channelSensitivity > threshold, add (f i , f j , ctxStrength, channelPattern) to C; Aggregate and normalize the interactions in C; Return the context-sensitive feature interaction matrix.
[0136] In another embodiment of the present application, S5. Channel scheduling decision under multi-objective constraints, specifically:
[0137] S51. Read the service attribute vector B_attr and construct a multi-objective function:
[0138] 1. According to the service type and priority, define the following objective functions:
[0139] a. Success rate objective function: G_success(c) = expected success rate(c);
[0140] b. Delay objective function: G_delay(c) = -expected delay(c);
[0141] c. Cost objective function: G_cost(c) = -channel cost(c);
[0142] d. Load balancing objective function: G_balance(c) = -System load imbalance degree(c);
[0143] 2. Determine the weight vector W_goal = {w_success, w_delay, w_cost, w_balance} of each objective according to business requirements;
[0144] 3. Construct a weighted multi-objective function: G(c) = w_success×G_success(c) + w_delay×G_delay(c) + w_cost×G_cost(c) + w_balance×G_balance(c);
[0145] 4. Output the multi-objective weighted function G
[0146] S52. Read information such as the standardized channel status vector C_status, the service level agreement SLA of the current service, and the system capacity limit, and construct constraints:
[0147] 1. Set the delay constraint according to the SLA requirement: delay(c) ≤ SLA_max_delay;
[0148] 2. Set the channel capacity constraint according to the system capacity: load(c) + new_load ≤ capacity(c);
[0149] 3. Set the cost constraint according to the business budget: cost(c) ≤ max_budget;
[0150] 4. Integrate all constraints and output the constraint set K;
[0151] S53. Read the content feature strength vector S and formulate an exploration-exploitation strategy:
[0152] 1. Calculate the overall feature strength score of the content: S_total = Σ(S[i] × w_i);
[0153] 2. Design a risk-aware exploration rate according to the overall feature strength: a. High-sensitivity content (such as verification codes): exploration_rate = low_rate b. Medium-sensitivity content (such as notifications): exploration_rate = medium_rate c. Low-sensitivity content (such as marketing): exploration_rate = high_rate;
[0154] 3. Output the exploration rate parameter E = exploration_rate;
[0155] S54. Read the comprehensive affinity scoring matrix R_final, the multi-objective weighting function G, the constraint set K, and the exploration rate parameter E, and execute the decision-making algorithm:
[0156] 1. Initialize the candidate channel set C_candidates = {};
[0157] 2. Execute the risk-aware exploration-exploitation balance strategy: a. With probability p = 1 - E, select the exploitation strategy: Add the channel with the highest comprehensive score to the candidate set; b. With probability p = E, select the exploration strategy: Randomly select a channel that satisfies the constraints and add it to the candidate set;
[0158] 3. Evaluate the multi-objective function value G(c) for each channel c in the candidate channel set C_candidates;
[0159] 4. Select the channel with the highest G(c) value and that satisfies all constraints K as the primary channel c_primary;
[0160] 5. Determine the backup channel strategy: a. Select the channel with the strongest complementarity to the primary channel from the remaining channels as the backup channel c_backup; b. Design the failover strategy and threshold conditions;
[0161] 6. Output the channel scheduling decision D = {c_primary, c_backup, failover_strategy}.
[0162] Combining content sensitivity with the exploration-exploitation balance, dynamically adjusting the exploration rate according to the sensitivity level of the SMS content, achieving the optimal balance between risk and performance. Adopting a conservative strategy (low exploration rate) for highly sensitive content (such as verification codes), and an active exploration strategy (high exploration rate) for low-sensitive content (such as marketing messages), enabling the system to ensure the stability of critical services while continuously optimizing and discovering new channel capabilities. In the SMS sending system, the importance and timeliness requirements of different types of SMS vary greatly. Verification codes and transaction notifications require high reliability, while marketing messages can accept more attempts and optimizations. Measured data shows that the risk-aware mechanism increases the delivery success rate of verification code SMS to 99.7% (2.3% higher than the unified exploration rate). At the same time, through the active exploration of marketing SMS, the system continuously discovers and utilizes new channel advantages, increasing the overall channel utilization efficiency by 15.8% and reducing the operating cost by 11.2%. It can continuously evolve and optimize while ensuring core services, achieving the best balance between stability and adaptability.
[0163] In a certain scenario, a method for intelligent scheduling of SMS sending channels based on multi-source data fusion is provided, especially a method for achieving precise matching of SMS channels by using the content feature strength and channel affinity (CSTA) model. By mapping the SMS content features and channel processing capabilities to the same vector space and using vector space measurement methods to evaluate the matching degree, the problem that traditional methods ignore the matching relationship between SMS content features and channel processing capabilities is solved.
[0164] In this embodiment, the system first collects raw data from multiple data sources:
[0165] Collection of channel status data: The system obtains real-time channel status data from the channel monitoring interface every 5 minutes, including channel availability markers (Boolean value, 1 means available, 0 means unavailable), channel congestion degree (value range 0 - 100), and channel response time (unit: milliseconds). A sliding average algorithm with a 10-minute time window is used to process fluctuations, and a standardized channel status vector C_status = {c_id, availability, congestion, response_time} is obtained; where c_id is the unique identifier of the channel, availability is the availability marker, congestion is the congestion degree value, and response_time is the average response time value.
[0166] Collection of historical sending data: The system incrementally obtains the sending records of the most recent 24 hours from the SMS sending history database, including information such as SMS content, sending channel, sending result (success / failure), and sending time, and generates an incremental historical sending record set H_data = [{msg_id, msg_content, channel_id, result, timestamp},...]; where msg_id is the unique identifier of the SMS, msg_content is the SMS content, channel_id is the sending channel identifier, result is the sending result (1 means success, 0 means failure), and timestamp is the sending timestamp.
[0167] Collection of business attribute data: Obtain the business attribute information of the SMS from the business system, including business type (1: verification code, 2: notification, 3: marketing, 4: service), priority (1 - 5, 5 is the highest), time requirement (unit: seconds), etc., and generate a business attribute vector B_attr = {msg_id, business_type, priority, time_requirement};
[0168] Clean the incremental historical send record set H_data, including: detecting and handling missing values, filling the missing result values with 0 (failure); marking the records with response_time exceeding 10,000 ms as abnormal; deduplicating based on msg_id and removing duplicate records; uniformly converting timestamps in different time zones to UTC time; generating the cleaned historical data set H_clean.
[0169] Align the standardized channel status vector C_status, the cleaned historical data set H_clean, and the business attribute vector B_attr based on the timestamp and channel ID, create an association index, and generate the spatio-temporal aligned multi-source data set D_aligned to provide a unified data view for subsequent analysis.
[0170] Among them, the content feature intensity vector extraction algorithm quantifies the SMS content features into a standardized feature intensity vector, specifically including:
[0171] Analyze the original SMS content M_text and extract the following basic features:
[0172] Total text length L_total: count the number of characters and bytes;
[0173] Number of segments N_segment: calculate the required number of segments according to the length (each segment length ≤ 70 characters);
[0174] Character set distribution vector C_dist: count the proportion of categories such as ASCII characters, non-ASCII characters, numbers, punctuation, etc.
[0175] Special symbol set S_special: identify and count special symbols (such as emojis, uncommon Unicode characters) and their positions;
[0176] Number of URLs N_url and features: detect URLs and analyze the length, domain type, and complexity;
[0177] Form these features into a basic feature vector F_basic = {L_total, N_segment, C_dist, S_special, N_url};
[0178] Adopt a lightweight natural language processing module to extract semantic features from the SMS content:
[0179] Message type T_type: use keyword matching and pattern recognition to determine the SMS category (verification code, notification, etc.);
[0180] Timeliness score T_urgency: extract time expressions and urgency words to evaluate the timeliness (range 0 - 1);
[0181] Importance score T_importance: Analyze important marking words and tone intensity to evaluate the importance level (in the range of 0 - 1);
[0182] Industry category T_industry: Identify industry characteristic words, such as finance, e-commerce, government affairs, etc.;
[0183] These semantic features constitute the semantic feature vector F_semantic = {T_type, T_urgency, T_importance, T_industry};
[0184] Different from the prior art, this solution not only analyzes single features but also considers the interaction relationships between features. The specific process is as follows:
[0185] 1. Initialize the feature interaction matrix F_interaction as an n×n zero matrix (n is the total number of features);
[0186] 2. For each pair of features (i, j), calculate the basic interaction intensity I_base(i,j) = F[i] × F[j];
[0187] 3. Calculate the channel-related interaction weight W_channel(i,j) = Σ(success_rate(c, i, j)) / N_channels from historical data; where success_rate(c, i, j) represents the processing success rate of channel c for the co-occurrence of features i and j;
[0188] 4. Calculate the time decay factor D_time(i,j) = exp(-|urgency_i - urgency_j| / τ); where τ is the time feature intensity parameter, and urgency_i and urgency_j are the timeliness scores of features i and j;
[0189] 5. Calculate the position correlation factor P_pos(i,j) = 1 / (1 + α × dist(i, j)); where dist(i, j) is the average distance between features i and j in the text, and α is the distance decay parameter;
[0190] 6. Calculate the final interaction intensity F_interaction[i,j] = I_base(i,j) × W_channel(i,j)× D_time(i,j) × P_pos(i,j);
[0191] Perform specialization analysis for the feature combinations unique to the SMS field:
[0192] 1. Interaction analysis of URL and special characters: I_url_special = Σ (the number of occurrences of special characters in the URL × position weight);
[0193] 2. Interaction analysis of segmentation and key information: I_seg_key = Σ (integrity scores of key information across segments);
[0194] 3. Interaction analysis of character encoding and length: I_char_len = Σ (proportion of non - ASCII characters × degree of length approaching the segmentation threshold);
[0195] The results of these specialized interaction analyses are integrated into the feature interaction matrix F_interaction.
[0196] Finally, based on the basic feature vector F_basic, semantic feature vector F_semantic, and feature interaction matrix F_interaction, a standardized 5 - dimensional content feature intensity vector S = {s1, s2, s3, s4, s5} is constructed; where: s1 is the character set feature intensity, representing the influence degree of the character set characteristics of the short message content on channel processing; s2 is the length feature intensity, representing the influence degree of the short message length and its segmentation characteristics on channel processing; s3 is the segmentation feature intensity, representing the influence degree of the number of short message segments and segmentation characteristics on channel processing; s4 is the link feature intensity, representing the influence degree of URL links in the short message on channel processing; s5 is the timeliness feature intensity, representing the influence degree of the timeliness of the short message content on channel processing;
[0197] The feature intensity value of each dimension is calculated through the feature intensity mapping function: s_k = sigmoid(Σ(w_ik × F_basic[i] + w_jk × F_semantic[j]+ w_mnk × F_interaction[m,n])); where w_ik, w_jk, and w_mnk are the contribution weights of each feature to the feature intensity dimension k, and the sigmoid function ensures that the value range is in the interval [0,1].
[0198] Through historical data analysis and active detection, an affinity vector representing the channel processing ability is constructed.
[0199] Historical processing ability analysis: Group the multi - source dataset D_aligned with spatio - temporal alignment by channel ID, and analyze the historical processing ability of each channel:
[0200] 1. Calculate the conditional success rate of the channel for different content features: P(success|feature_i) = number of successful transmissions when feature i exists / total number of transmissions when feature i exists;
[0201] 2. Evaluate the delay distribution of different feature contents processed by the channel: delay_stats(feature_i) = {min_delay, max_delay, avg_delay, std_delay};
[0202] 3. Analyze the time series pattern of the channel performance and identify the periodic changes: pattern(c) = FFT(performance_series(c)); where FFT is the Fast Fourier Transform used to identify the periodic pattern
[0203] 4. Construct the channel historical processing capability matrix: H_capability[c,i] = {P(success|feature_i), delay_stats(feature_i), pattern_strength(i)};
[0204] Construct an interaction feature intensity model for channel differentiation, which is different from the traditional method that uses a unified feature weight for all channels. The specific steps are as follows:
[0205] 1. Initialize the feature interaction intensity matrix Sc for each channel c as an n×n identity matrix;
[0206] 2. Calculate the conditional processing probabilities for each pair of features (i,j): p_success_ij = P(success|i exists and j exists); p_success_i = P(success|i exists); p_success_j = P(success|j exists);
[0207] 3. Calculate the channel-specific interaction feature intensity: interactionSensitivity = log(p_success_ij / (p_success_i × p_success_j));
[0208] Based on the calculation method of mutual information MI, it represents the influence degree of the co-occurrence of features i and j on the success rate.
[0209] 4. Apply Bayesian correction to handle the data sparsity problem: Sc[i,j] = (interactionSensitivity × sample_count + prior) / (sample_count + prior_weight); where prior is the prior feature intensity value, prior_weight is the prior weight, and sample_count is the number of samples;
[0210] 5. Identify the channel-specific sensitive interaction pattern Pc through cluster analysis;
[0211] 6. Integrate the interaction feature intensity data of all channels to generate the channel interaction feature intensity matrix C_sensitivity.
[0212] The implementation process is as follows: For each channel c belonging to C:
[0213] Initialize the channel feature interaction sensitivity matrix Sc = identity matrix (n×n);
[0214] Analyze the differential sensitivity of the channel to feature interaction: For each pair of features (i,j): Extract the historical processing data of the channel for the feature pair histData = Extract historical data (c, i, j);
[0215] Calculate the conditional processing probability; p_success_ij = Calculate the conditional success rate (histData, i exists and j exists); p_success_i = Calculate the conditional success rate (histData, i exists); p_success_j = Calculate the conditional success rate (histData, j exists);
[0216] Calculate the channel-specific interaction sensitivity interactionSensitivity = Calculate the information gain (p_success_ij, p_success_i, p_success_j)
[0217] Apply Bayesian correction to process sparse data Sc[i,j] = Apply Bayesian correction (interactionSensitivity, number of samples);
[0218] Identify the channel-specific sensitive interaction pattern, Pc = Extract the channel feature interaction pattern (Sc); Store the channel feature interaction sensitivity model (c, Sc, Pc).
[0219] Active detection strategy execution: The system executes the active detection strategy based on the uncertain region in the channel interaction feature intensity matrix:
[0220] 1. Calculate the uncertainty score of each element in the feature intensity matrix uncertainty(i,j) = 1 - confidence(Sc[i,j]); where the confidence function calculates the confidence based on the number of samples and the distribution;
[0221] 2. Preferentially select feature pairs (i, j) with high uncertainty to generate probe messages: P_messages = generate_probe_messages(high_uncertainty_features);
[0222] 3. Send the probe messages through each target channel and collect the results to form a probe result dataset P_results;
[0223] Based on historical data and probe results, calculate and update the channel affinity vector: Initialize a 5-dimensional channel affinity vector A_c for each channel c, corresponding one-to-one with the dimension of the content feature intensity vector S: A_c = {a1, a2, a3, a4, a5};
[0224] For each affinity dimension k, apply a differential memory and forgetting mechanism: Set different half-lives τ_k to reflect the change rate of the characteristics of each dimension channel; Calculate the time decay weight w_t = exp(-t / τ_k); where t is the time interval of the data from the current time; Perform weighted fusion of data from different periods A_c[k] = Σ(data_t × w_t) / Σw_t; Normalize the affinity vector to ensure comparability between different channels;
[0225] Map the content and the channel to the same vector space and measure the matching degree through vector distance, specifically:
[0226] Calculate the similarity between the content feature intensity vector and each channel affinity vector:
[0227] Calculate the cosine similarity cos_sim(S, A_c) = (S•A_c) / (||S|| × ||A_c||);
[0228] Calculate the Euclidean distance euc_dist(S, A_c) = √(Σ(S[i] - A_c[i]) 2 )
[0229] Calculate the Mahalanobis distance (considering the correlation between dimensions) mahal_dist(S, A_c) = √((S - A_c) T Σ -1 (S - A_c));
[0230] where Σ is the covariance matrix of the feature intensity dimension;
[0231] 4. Combine multiple distance metrics to calculate the initial affinity score sim_score(S, A_c) = α × cos_sim - β × norm(euc_dist) - γ × norm(mahal_dist); where α, β, and γ are weight parameters, and norm is a normalization function; Dynamically adjust the importance of each feature intensity dimension according to business requirements:
[0232] Calculate the dimension weight vector W[i] = importance(i, B_attr); where the importance function calculates the importance of dimension i based on business type, priority, etc.;
[0233] Apply dimension weighting to calculate the affinity score r ’ _c = Σ(sim_score_dim(S[i], A_c[i]) × W[i]) / ΣW[i];
[0234] Consider the impact of the current load status of the channel on performance:
[0235] Extract the channel congestion degree congestion_c and response time response_time_c;
[0236] Apply the load-performance decay model l_c = 1 - sigmoid(α × congestion_c + β ×(response_time_c / threshold)); where α and β are weight parameters, and threshold is the response time threshold;
[0237] Combine the context-weighted score and the load impact factor:
[0238] Apply the non-linear combination function r_final_c = r ’ _c × (1 - γ × (1 - l_c)); where γ is the load impact degree parameter;
[0239] Consider data sufficiency to adjust the credibility r_final_c = r_final_c × (1 - exp(-sample_count / threshold));
[0240] Construct the comprehensive affinity scoring matrix R_final, which contains the scoring of the final matching degree between the SMS and each channel;
[0241] 6. Channel scheduling decision-making under multi-objective constraints.
[0242] Based on the comprehensive affinity scoring, make the optimal channel scheduling decision under multi-objective constraints.
[0243] Objective function construction: Construct a multi-objective function according to business requirements:
[0244] The success rate objective G_success(c) = predicted_success_rate(c);
[0245] The latency objective G_delay(c) = -predicted_delay(c) / max_acceptable_delay;
[0246] The cost objective G_cost(c) = -cost(c) / max_acceptable_cost;
[0247] The load balancing objective G_balance(c) = -|current_load(c) - average_load| / max_load;
[0248] Construct a weighted multi-objective function G(c) = w_success×G_success(c) + w_delay×G_delay(c) + w_cost×G_cost(c) + w_balance×G_balance(c);
[0249] where w_success, w_delay, w_cost, and w_balance are the weights of each objective, set according to the business type;
[0250] Constraint condition identification: Set constraint conditions according to the system state:
[0251] The latency constraint delay(c) ≤ SLA_max_delay;
[0252] The channel capacity constraint load(c) + new_load ≤ capacity(c);
[0253] The cost constraint cost(c) ≤ max_budget;
[0254] Dynamically adjust the exploration rate according to the content feature intensity:
[0255] Calculate the overall feature intensity score of the content S_total = w1×s1 + w2×s2 + w3×s3 + w4×s4 + w5×s5;
[0256] Design the exploration rate according to the overall feature strength: exploration_rate = baseline_rate × (1 - β× S_total); where baseline_rate is the basic exploration rate and β is the feature strength influence coefficient.
[0257] Set the exploration rate range:
[0258] High-sensitivity content (such as verification codes): exploration_rate ≤ 0.05;
[0259] Medium-sensitivity content (such as notifications): 0.05 < exploration_rate ≤ 0.15;
[0260] Low-sensitivity content (such as marketing): 0.15 < exploration_rate ≤ 0.30;
[0261] Optimal channel selection decision: Execute the channel selection decision algorithm
[0262] Select the exploitation strategy with probability p = 1 - exploration_rate:
[0263] c_exploit = argmax(r_final_c) for all c that meet constraints;
[0264] Select the exploration strategy with probability p = exploration_rate: c_explore = random_select(c) from all c that meet constraints;
[0265] Finally, select the primary channel: c_primary = c_exploit if using exploitation else c_explore;
[0266] Select the channel with the strongest complementarity as the backup channel: c_backup = argmax(complementarity(c, c_primary)) for all c ≠ c_primary that meet constraints; where the complementarity function evaluates the complementarity degree of two channels;
[0267] Generate the channel scheduling decision D = {c_primary, c_backup, failover_strategy}.
[0268] The implementation process is as follows: Extract the content sensitivity vector S; Initialize the best score bestScore = a; Initialize the best channel bestChannel = null;
[0269] For each available channel c: Obtain the channel affinity vector A_k; Obtain the current load status L of the channel; Calculate the base affinity score baseScore = Calculate the affinity score (S, A_k); Apply the load adjustment function adjustedScore = Apply load adjustment (baseScore, L); Apply the channel exploration-exploitation balance strategy finalScore = Apply the exploration-exploitation strategy (adjustedScore, c); If finalScore > bestScore, bestScore = finalScore; Otherwise, bestChannel = c; Record the decision (S, bestChannel, bestScore), and return the optimal channel bestChannel.
[0270] Verification was carried out in the SMS sending system of an e-commerce platform. A total of three types of SMS were tested: verification codes, order notifications, and promotional messages, with 4000 messages of each type. The results are as follows:
[0271] The overall sending success rate increased by 12.6% compared to the traditional rule-based scheduling method and by 8.3% compared to the statistical method; The delay of verification code SMS decreased, and the average delivery time decreased from 2.8 seconds to 1.5 seconds; The system resource consumption decreased, and the usage rate of channel monitoring resources decreased by 23.7%; The special SMS processing ability improved. Among them, the success rate of SMS containing URL links increased by 21.4%, and the success rate of SMS containing special characters increased by 18.9%.
[0272] The effectiveness of the method of the present invention was verified, especially the significant advantages in processing special content SMS.
[0273] In summary, in this embodiment, the main processes of the content feature strength and channel affinity model (CSTA) are as follows: The content feature strength vector extraction algorithm quantifies the SMS content features into a standardized vector; The dynamic weight adaptive mechanism for feature interaction considers the mutual influence between features; The interactive feature strength modeling with channel differentiation constructs a specialized feature strength model for each channel; The differential memory and forgetting mechanism dynamically adjusts the half-life according to the change rates of different dimensions; The exploration-exploitation balance mechanism with risk perception dynamically adjusts the exploration rate based on the content feature strength.
[0274] In another embodiment of the present application, the specific implementation process of calculating the position correlation factor P_pos(i,j) is as follows:
[0275] Extract the instance location information of each feature from the original SMS content. For example, for an SMS containing "Your verification code is 1234, valid within 5 minutes, visit www.example.com to complete registration", the system will identify that the "verification code" feature is located at positions 2-5, the "validity expression" feature is located at positions 14-19, and the "URL" feature is located at positions 23-39.
[0276] Determine the context window for each feature instance, usually taking 10 characters before and after the feature to form the context. Within this window, the system analyzes the positional relationships between feature instances, including distance, order, and nesting.
[0277] Calculate the impact of the positional relationship between features using a distance decay model: distance_impact(i _inst , j _inst ) = exp(-λ × |pos(i _inst ) - pos(j _inst )| / len(text)); where λ is the distance decay coefficient (default value is 2.5), pos() represents the position of the feature instance in the text, and len(text) is the total length of the text.
[0278] Identify the clustering and dispersion patterns of feature instances. The clustering index is obtained by calculating the ratio of the standard deviation to the mean (CV) between feature instances: clustering_score = 1 - min(1, CV / threshold); where threshold is the clustering threshold (default value is 0.5).
[0279] Combine the distance impact factor and the position distribution pattern score to calculate the position correlation factor P_pos(i,j) = α × avg(distance_impact) + (1 - α) × clustering_score; where α is the weight parameter (default value is 0.7).
[0280] In practical applications, this calculation method enables the system to identify the impacts of position patterns such as "verification code is adjacent to the validity period description" and "URL is at the end of the SMS" on channel processing, significantly improving the accuracy of feature interaction analysis.
[0281] In another embodiment of this application, for the feature interaction recognition steps specialized in the SMS field, a dedicated interaction recognition algorithm is designed for the content features unique to SMS. Specifically as follows:
[0282] Analyze the interaction relationship between URLs and special characters in the original SMS content. The system uses regular expressions to extract URLs and then detects the distribution of special characters in the URLs: url_pattern = re.compile(r ’ http[s]?: / / (?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\\(\\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+ ’ );urls =url_pattern.findall(msg_content);
[0283] For each URL, the system counts the number and positions of special characters, and calculates the position influence weight of special characters in the URL special_char_weight_in_url = Σ(w_pos × count(special_char, position)); where w_pos is the position weight factor, and the special characters in the URL parameter part have a higher weight than those in the domain name part. Finally, generate the interaction intensity of URL special characters I_url_special.
[0284] Evaluate the interaction between SMS segmentation and key information. The system first simulates SMS segmentation processing to determine the segmentation boundaries segments = [msg_content[i:i+70] for i in range(0, len(msg_content), 70)];
[0285] Identify the distribution of key information (such as verification codes, links, important instructions, etc.) in each segment, and analyze whether there is a cross-segment situation: key_info_integrity = Σ(completeness(key_info, segment_i) ×importance(key_info)); where completeness() evaluates the integrity of key information in the segment, and importance() represents the importance of key information. Finally, calculate the interaction intensity of segmented key information I_seg_key.
[0286] Analyze the interaction relationship between character encoding and length. First, count the distribution of characters of different encoding types:
[0287] encoding_distribution = {
[0288] ‘ ascii ’ : count_chars(msg_content, ‘ ascii’ ) / len(msg_content),
[0289] ‘ utf8_special ’ : count_chars(msg_content, ‘ utf8_special ’ ) / len(msg_content),
[0290] ‘ emoji ’ : count_chars(msg_content, ‘ emoji ’ ) / len(msg_content)
[0291] };
[0292] Evaluate the impact of these characters on the calculation of the SMS length, especially in cases close to the segmentation threshold. length_border_risk = sigmoid((len(msg_content) % 70) / 10) * encoding_distribution ‘ utf8_special ’ ;
[0293] Generate the character encoding length interaction intensity I_char_len.
[0294] Identify the interaction patterns between the timeliness markers and the content types. The system extracts timeliness expressions (such as "within x minutes", "immediately", etc.) and analyzes their performance differences in different content types: time_content_correlation = correlation(time_urgency_score, content_type_vector); where time_urgency_score is the timeliness urgency score and content_type_vector is the content type vector. Finally, calculate the timeliness content interaction intensity I_time_content.
[0295] Integrating these domain-specific interaction intensities into the feature interaction matrix significantly enhances the domain adaptability of the feature interaction analysis.
[0296] In another embodiment of the present application, the steps for calculating the interaction feature intensity specific to the channel are as follows in the specific implementation process:
[0297] Extract the feature co-occurrence data for each time period from the channel history processing ability matrix. The system sets a sliding time window (default is 7 days) and counts the co-occurrence and individual occurrences of feature i and feature j within the window:
[0298] time_windows = generate_sliding_windows(start_date, end_date, window_size=7, step=1);
[0299] for window in time_windows:
[0300] feature_co_occurrence[window] = extract_co_occurrence(H_capability, window);
[0301] Construct a contingency table of feature co-occurrence and calculate the sample sizes for four cases:
[0302] contingency_table = {
[0303] ‘ n11 ’ : count(i exists and j exists), # Both features exist;
[0304] ‘ n10 ’ : count(i exists and j does not exist), # Feature i exists, feature j does not exist;
[0305] ‘ n01 ’ : count(i does not exist and j exists), # Feature i does not exist, feature j exists;
[0306] ‘ n00 ’ : count(i does not exist and j does not exist) # Both features do not exist;
[0307] }
[0308] Apply mutual information (MI) to calculate the initial interaction information of the feature pair:
[0309] Total sample size N = n11 + n10 + n01 + n00;
[0310] The probability of feature i occurring p_i1 = (n11 + n10) / N;
[0311] The probability \(p_{j1}\) of the occurrence of feature \(j\) is \(p_{j1}=(n_{11} + n_{01}) / N\);
[0312] The probability \(p_{11}\) of the simultaneous occurrence of features \(i\) and \(j\) is \(p_{11}=n_{11} / N\);
[0313] The expected co-occurrence probability \(expected\_p_{11}\) in the independent case is \(expected\_p_{11}=p_{i1}*p_{j1}\);
[0314] The mutual information value \(MI = p_{11}*\log_2(p_{11} / expected\_p_{11})\);
[0315] Calculate the point mutual information (PMI) to evaluate the specific association strength of feature pairs on the channel: \(PMI=\log_2(p_{11} / (p_{i1}*p_{j1}))\);
[0316] Evaluate the statistical significance of feature interaction through the chi-square test: \(\chi^2 = N*(n_{11}*n_{00}-n_{10}*n_{01})^2 / ((n_{11}+n_{10})*(n_{11}+n_{01})*(n_{10}+n_{00})*(n_{01}+n_{00}))\);
[0317] Combining the MI, PMI, and \(\chi\) 2 values, construct an interaction feature strength score \(sensitivity\_score = w1*normalize(MI)+w2*normalize(PMI)+w3*normalize(\chi^2)\); where \(w1\), \(w2\), \(w3\) are weight parameters, and \(normalize()\) is a normalization function.
[0318] Apply Bayesian smoothing correction to mitigate the impact of data sparsity problem: \(smoothed\_score=(sensitivity\_score * sample\_count + prior\_value * prior\_weight) / (sample\_count + prior\_weight)\); where \(prior\_value\) is the prior feature strength value (usually the global average), and \(prior\_weight\) is the prior weight (set according to data reliability).
[0319] Through a refined channel interaction feature strength modeling method, it is possible to accurately capture the unique processing ability of each channel for specific feature combinations, significantly improving the accuracy of channel selection.
[0320] In another embodiment of the present application, the specific implementation process of the differential update of the affinity vector is as follows:
[0321] First, the system analyzes the historical fluctuation frequency and amplitude of each dimension k, and calculates the dimension stability index variations = calculate_variations(dimension_k_history, window_size=30); stability_k = 1 - min(1, variations / max_acceptable_variation); where variations is the historical value change rate of dimension k, and max_acceptable_variation is the maximum acceptable change rate (default value is 0.3).
[0322] The system calculates the state sensitivity index based on the correlation between dimension k and the channel status change correlation_values = []; for status_metric in ‘ congestion ’ , ‘ response_time ’ , ‘ availability ’ :
[0323] correlation_values.append(pearson_correlation(dimension_k_history,channel_status_history[status_metric]));
[0324] sensitivity_k = max(abs(correlation_values));
[0325] The system calculates the trend change rate according to the change trend of dimension k in the recent detection results recent_probes= filter_recent_probes(P_results, days=7);
[0326] trend_variations = calculate_slope(dimension_k_in_probes);
[0327] trend_k = normalize(abs(trend_variations));
[0328] Integrate stability_k, sensitivity_k, and trend_k to determine the change rate characteristic change_rate_k of dimension k: change_rate_k = (1 - stability_k) * 0.5 + sensitivity_k * 0.3 + trend_k * 0.2;
[0329] For dimensions with different change rates, the system assigns different half-lives:
[0330] if change_rate_k >= 0.7: # Rapidly changing dimension;
[0331] τ_k = base_half_life * 0.3 # Short half-life;
[0332] elif change_rate_k >= 0.4: # Moderately changing dimension;
[0333] τ_k = base_half_life * 0.6 # Medium half-life;
[0334] else: # Stable dimension;
[0335] τ_k = base_half_life # Long half-life;
[0336] where base_half_life is the base half-life parameter (default value is 14 days).
[0337] When a significant change is observed in dimension k, dynamically adjust the half-life parameter of this dimension:
[0338] if recent_change_magnitude > significant_change_threshold:
[0339] τ_k = τ_k * adjustment_factor
[0340] where adjustment_factor is the adjustment coefficient (usually between 0.5 and 2).
[0341] Apply a non-linear smoothing function to the affinity vector to ensure that the half-life differences between dimensions do not cause vector instability:
[0342] smoothed_vector = apply_smooth_function(A_c, half_life_differences);
[0343] When calculating the affinity value for dimension k, the system applies a differential memory and forgetting mechanism:
[0344] for data_point in historical_data:
[0345] time_distance = current_time - data_point.timestamp;
[0346] weight = exp(-time_distance / τ_k);
[0347] weighted_sum += data_point.value * weight;
[0348] weight_sum += weight;
[0349] A_c[k] = weighted_sum / weight_sum;
[0350] The differential memory and forgetting mechanism enables the system to assign short half-lives to rapidly changing dimensions (such as timeliness feature intensity) and long half-lives to stable dimensions (such as character set feature intensity), significantly improving the accuracy and adaptability of channel affinity assessment. In the text, feature intensity can be described as sensitivity in some literature, such as the sensitivity to URLs and special characters.
[0351] By setting different half-lives for different affinity dimensions, reflecting the differences in the change rates of the characteristics of each dimension's channels, the problem of the traditional time decay method using a unified decay rate for all characteristics is solved. Analyze the historical fluctuation frequency, state sensitivity, and trend changes of each dimension, and assign short half-lives to rapidly changing dimensions (such as timeliness sensitivity) and long half-lives to stable dimensions (such as character set sensitivity). In SMS channel management, the differences in the change rates of different characteristics are significant. For example, the channel's ability to handle URLs may change rapidly due to policy adjustments, while the ability to handle basic ASCII text is relatively stable. The test results show that after adopting the differential memory and forgetting mechanism, the system's response speed to channel state changes has increased by 42.3%, while maintaining accurate assessment of stable characteristics, enabling the system to quickly adapt to environmental changes while maintaining decision-making stability, and significantly reducing scheduling errors caused by lagging or overly fluctuating channel capacity assessments.
[0352] In summary, in the present invention, a mapping model CSTA of content feature intensity and channel affinity vector space is constructed to map the short message content features and channel processing capabilities to the same vector space, achieving precise matching of short messages and channels. The traditional one-way evaluation (only evaluating the channel status) is transformed into two-way matching (considering both content features and channel capabilities), enabling the system to identify and utilize the processing advantages of different channels for specific content features.
[0353] In practical applications, the CSTA model can preferentially allocate short messages containing URL links to channels with strong URL processing capabilities and allocate multi-segment short messages to channels with more stable segmentation processing, thus significantly improving the sending success rate. The test results show that compared with the traditional scheduling method, the sending success rate of special content short messages (such as those containing special characters, URL links, etc.) has increased by 15%-30%. This effect directly stems from the model's ability to capture the affinity relationship between content and channels, rather than simply relying on the overall performance indicators of the channels.
[0354] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. An intelligent scheduling method for SMS sending channels based on multi-source data fusion, characterized in that, The steps include: Obtain the basic data including the original text of the short message, perform cleaning, standardization, and spatio-temporal alignment to obtain a multi-source dataset; Extract the basic features and semantic features of the original text of the short message, calculate the feature interaction relationship, and construct a content feature strength vector; Based on the multi-source dataset, analyze the historical processing ability of the channel, construct a channel differential interaction feature strength model, perform active detection, and calculate the channel affinity vector set; Calculate the vector space similarity between the content feature strength vector and the channel affinity vector set, and combine the service attributes and the channel load status to obtain a comprehensive affinity scoring matrix; Based on the comprehensive affinity scoring matrix, use a pre-constructed multi-objective weighting function to calculate the optimal channel and generate a channel scheduling decision; The steps for constructing the content feature strength vector include: Extract the basic features of the original text of the short message, including text length, number of segments, character set distribution, special symbol set, and number of URLs, and generate a basic feature vector; Identify the semantic features of the original text of the short message, including message type, timeliness score, importance score, and industry category, and generate a semantic feature vector; Combine the basic feature vector and the semantic feature vector, construct a feature interaction matrix, calculate the dynamic interaction weight, time decay factor, and position correlation factor between features, and analyze domain-specific interactions, including URLs and special characters, segments and key information, to generate a feature interaction matrix; Based on the feature interaction matrix, map each feature and its interaction effect to a standardized feature strength space through a feature strength mapping function to generate a multi-dimensional content feature strength vector; The steps for constructing the channel affinity vector set include: Group the multi-source dataset by channel ID, calculate the conditional success rate and processing delay of each channel for different content features, analyze the time series pattern of channel performance, and generate a channel historical processing ability matrix H_cap; Based on the channel historical processing ability matrix and the historical feature interaction matrix data, calculate the conditional processing probability of each channel for the feature pair, evaluate the interaction feature strength, and identify the sensitive interaction patterns unique to the channel to generate a channel interaction feature strength matrix C_sen; Extract the uncertain region in the channel interaction feature strength matrix to generate a targeted detection short message set, send it through the target channel and collect the results to form a detection result dataset P_res; Based on H_cap, C_sen, and P_res, calculate and update the multi-dimensional affinity vector for each channel to generate a channel affinity vector set; The steps for constructing the feature interaction matrix include: For each pair of features in the basic feature vector and the semantic feature vector, calculate the basic interaction strength; Extract the historical feature strength data of the channel for the feature pair from the multi-source dataset and calculate the channel-related interaction weight; Calculate the time decay factor according to the timeliness difference between feature i and feature j; Analyze the relative position relationship between feature i and feature j in the text structure and calculate the position correlation factor; Multiply the basic interaction strength, the channel interaction weight, the time decay factor, and the position correlation factor to obtain the final interaction strength value of the feature pair, form a feature interaction matrix and perform non-linear transformation and normalization processing; The steps for calculating the position correlation factor include: Extract the instance positions of each feature from the original SMS content and determine its context window; For each feature instance, analyze its positional relationship with other feature instances in the context window; Calculate the distance function between feature instances and evaluate the impact of the positional relationship on the interaction intensity according to the distance attenuation model; Based on the position distribution pattern, identify the clustering and dispersion of feature instances and calculate the position distribution pattern score; Combine the distance influence factor and the position distribution pattern score to generate the position correlation factor of the feature pair.
2. The method according to claim 1, characterized in that, The steps to calculate the vector space similarity between the content feature intensity vector and the channel affinity vector set to obtain the comprehensive affinity score matrix include: Calculate at least two distance metrics for the content feature intensity vector and the affinity vector of each channel to generate an initial affinity score set; According to the service type, priority, and timeliness requirements in the service attribute vector, calculate the importance weight vector for each feature intensity dimension, and perform weighted adjustment on the initial affinity scores to generate a context-weighted affinity score set; Collect channel status data and extract the congestion degree and response time of each channel, and apply the load-performance attenuation model to calculate the performance attenuation factor under the current load to generate a load influence factor set; Combine the context-weighted affinity score set and the load influence factor set through a non-linear combination function, and consider data sufficiency to adjust the score credibility to generate a comprehensive affinity score matrix.
3. The method according to claim 1, characterized in that, Constructing the feature interaction matrix also includes applying feature interaction recognition steps specialized in the SMS field: Analyze the interaction relationship between URLs and special characters in the original SMS content, identify the distribution and encoding method of special characters in the URL, calculate the position influence weight of special characters in the URL, and generate the URL special character interaction intensity; Evaluate the segmentation and key information interaction of the original SMS content, analyze the distribution of key information in different segments, identify the integrity of cross-segment key information, and generate the segment key information interaction intensity; Analyze the impact of different encoded characters on the length calculation in the original SMS content, identify the encoding conversion risk in boundary cases, and generate the character encoding length interaction intensity; Identify the performance differences of timeliness markers in different content types in the original SMS content, calculate the joint distribution of timeliness and content priority, and generate the timeliness content interaction intensity; Integrate the above domain-specific interaction intensities into the feature interaction matrix to enhance the domain adaptability of feature interactions.
4. The method according to claim 1, characterized in that, The steps to generate the channel interaction feature intensity matrix include: Initialize the channel feature interaction intensity matrix Sc for each channel c as an n×n identity matrix; Extract the historical processing data h of the channel for the feature pair from the channel historical processing capacity matrix; Calculate the conditional processing probability p _success_ij , p _success_i and p _success_j , which respectively represent the success rates when features i and j coexist, only i exists, and only j exists; Calculate the channel-specific interaction feature intensity and evaluate the impact degree of the feature pair on the channel processing capacity; Apply the Bayesian correction method to handle the data sparsity problem according to the sample quantity to improve the reliability of the evaluation of low-frequency feature pairs; Perform clustering analysis on the interaction feature intensity matrices of each channel to identify the sensitive interaction patterns unique to the channel; Integrate the interaction feature intensity data of all channels to generate the channel interaction feature intensity matrix.
5. The method according to claim 1, characterized in that The steps to calculate and update the channel affinity vector by applying the differential memory and forgetting mechanism include: Initialize a multi-dimensional channel affinity vector \(A_c\) corresponding to the dimension of the content feature intensity vector for each channel \(c\). For each affinity dimension \(k\), extract data related to this dimension from the channel historical processing capacity matrix, the channel interaction feature intensity matrix, and the detection result dataset. Set different half-lives \(\tau_k\) for each affinity dimension \(k\); calculate the weights \(w_t = f_{decay}(t, \tau_k)\) of historical data in different time windows, where \(f_{decay}\) is a time decay function. Calculate the affinity value of dimension \(k\) by weighted fusion of processing capacity data in different periods to obtain the channel affinity vector. Normalize the channel affinity vector \(A_c\) to ensure that the affinity values between different channels are comparable. Integrate the affinity vectors of all channels to generate a set of channel affinity vectors \(A\).
6. The method according to claim 4, characterized in that The steps for calculating the channel-specific interaction feature intensity include: Extract the frequencies of co-occurrence and individual occurrence of feature \(i\) and feature \(j\) in each time period from the channel historical processing capacity matrix. Construct a four-fold table of binary feature co-occurrence and calculate the number of samples when feature \(i\) and feature \(j\) co-occur, only \(i\) appears, only \(j\) appears, and neither appears. Apply mutual information \(MI\) to calculate the initial interaction information quantity \(MI(i, j)\) of the feature pair \((i, j)\). Combine point mutual information \(PMI\) to evaluate the specific association strength \(PMI(i, j)\) of the feature pair on the channel. Apply the chi-square test to evaluate the statistical significance of feature interactions χ 2 (i,j); Integrate MI, PMI, and χ 2 values to construct an interaction feature intensity score with multi-index fusion; Based on the number of samples and distribution characteristics, apply a Bayesian smoothing correction function to mitigate the impact of data sparsity problems on feature intensity evaluation.
7. The method according to claim 5, characterized in that, The steps for setting different half-lives \(\tau_k\) for each affinity dimension \(k\) include: Analyze the historical fluctuation frequency and amplitude of each dimension \(k\) and calculate the dimension stability index \(sta_k\). Based on the correlation between dimension \(k\) and channel state changes, calculate the state sensitivity index \(sen_k\). According to the change trend of dimension \(k\) in recent detection results, calculate the trend change rate \(tre_k\). Integrate \(sta_k\), \(sen_k\), and \(tre_k\) to determine the change rate characteristics of dimension \(k\). Among them, assign a short half-life to the dimension whose change speed exceeds the threshold, and assign a long half-life to the stable dimension to establish \(\tau_k = f_{assign}(sta_k, sen_k, tre_k)\).
Citation Information
Patent Citations
Short message scheduling method, equipment and storage medium
CN109561403A
Method and device for intelligently and strictly selecting short message channel
CN118972791A