Short message marketing effect optimization method and system based on multi-channel dynamic weight
Through the SMS marketing effect optimization method based on multi-channel dynamic weights, Z-score standardization and LSTM neural network are used to predict conversion rate trends, and gradient optimization is combined to adjust SMS queue diversion. This solves the problems of resource waste and evaluation uniformity under static weight configuration, and achieves intelligent optimization and stability improvement of SMS marketing effects.
Patent Information
- Application Number
- CN202511299677.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In existing SMS marketing systems, multi-channel management based on static weight configuration suffers from response lag and a single evaluation dimension. It is unable to dynamically adjust resource allocation based on real-time conversion effects, resulting in inefficient channels continuously occupying valuable sending resources, and lacks comprehensive consideration of key business indicators such as conversion rate and complaint rate.
A SMS marketing effect optimization method based on multi-channel dynamic weights is adopted. By obtaining real-time performance data, the comprehensive performance score is calculated by weighted summation after Z-score normalization. The LSTM neural network is combined to predict future conversion rate trends. The gradient optimization method is used to calculate the optimal distribution ratio, and the diversion strategy of the SMS queue is adjusted in real time. Dynamic optimization is achieved by combining a smooth transition mechanism.
It achieves dynamic optimization of SMS marketing effects, breaks through the limitations of traditional static configuration, realizes intelligent allocation of channel resources, improves the comprehensiveness and accuracy of evaluation, enhances the foresight of decision-making and the stability of the system, and reduces the cost of manual trial and error.
Smart Images

Figure CN120807020A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of short message marketing, in particular to a short message marketing effect optimization method and system based on multi-channel dynamic weight, which is used for realizing intelligent allocation and effect optimization of short message marketing channel resources. BACKGROUND
[0002] As an important part of digital marketing, short message marketing plays a key role in financial services, e-commerce promotion, customer reach, etc. With the in-depth development of mobile Internet, enterprises have increasingly high requirements for the accuracy and conversion efficiency of short message marketing.
[0003] At present, the short message marketing system generally adopts a multi-channel parallel sending strategy, that is, marketing short messages are sent to target user groups through the interfaces of multiple short message service providers at the same time. Common technical solutions include polling distribution and fixed weight allocation modes, among which the polling distribution uses each channel in a predetermined order, and the fixed weight allocation sets the sending proportion of each channel according to historical experience.
[0004] The most typical one in the prior art is a multi-channel management system based on static weight configuration, which realizes batch distribution of short messages by pre-setting the sending proportion of each channel (such as channel A occupying 60% of the traffic and channel B occupying 40% of the traffic) and combining reach rate monitoring. The technical principle is to divide the short message queue to be sent according to the preset proportion, route it to the corresponding channel interface, and monitor the short message delivery status of each channel. When the reach rate of a certain channel is lower than the threshold, manual intervention adjustment is triggered.
[0005] However, this static configuration-based technical solution has significant response lag and single evaluation dimension defects. The system cannot dynamically adjust the resource allocation strategy according to real-time conversion effect, resulting in the continuous occupation of valuable sending resources by inefficient channels, and the lack of comprehensive consideration of key business indicators such as conversion rate and complaint rate. SUMMARY
[0006] The purpose of the present application is to overcome the deficiencies in the prior art and provide a short message marketing effect optimization method and system based on multi-channel dynamic weight. The method can standardize the real-time performance data of multiple short message channels, calculate the comprehensive performance score of each channel by weighted summation, predict the conversion rate trend change of each channel in the future, calculate the optimal distribution proportion of each channel using the gradient optimization method, and adjust the shunting strategy of the short message queue in real time, thereby realizing dynamic optimization of short message marketing effect.
[0007] In order to achieve the above purpose, the technical solution provided by the present application is as follows: A short message marketing effect optimization method based on multi-channel dynamic weight, comprising: Acquire real-time performance data of multiple short message channels; Standardize the real-time performance data, and map the reach rate, conversion rate, complaint rate and delay index to the interval [0, 1] by using the Z-score standardization method to obtain standardized performance indexes; Based on the standardized performance indexes, calculate the comprehensive performance score of each channel by weighted summation, generate a channel performance benchmark score by using a scoring formula Score = w1 x reach rate + w2 x conversion rate - w3 x complaint rate - w4 x delay index; Collect multi-dimensional performance data of each channel in the past 24 hours, analyze the time sequence characteristics of historical conversion data by using an LSTM neural network, and learn the time sequence pattern by using a three-layer LSTM network structure to predict the conversion rate trend change of each channel in the next 2-4 hours; Based on the conversion rate trend change and the channel performance benchmark score, solve the weight distribution scheme with the goal of maximizing the overall ROI by using a gradient optimization method, set the objective function F(w) = Σ(wi x predicted conversion rate i x weight coefficient), and calculate the optimal distribution ratio of each channel; According to the optimal distribution ratio, real-time adjust the shunt strategy of the short message queue, re-distribute the to-be-sent short messages to each channel queue, and simultaneously start the smooth transition mechanism of weight configuration to realize dynamic optimization of short message marketing effect.
[0008] Preferably, the standardization of the real-time performance data adopts the Z-score standardization method to map the reach rate, conversion rate, complaint rate and delay index to the interval [0, 1] to obtain standardized performance indexes, which includes: Collect the original data of the reach rate, conversion rate, complaint rate and delay index of each short message channel, calculate the mean μ and standard deviation σ of each index in the historical data, and establish a data distribution baseline; Based on the data distribution baseline, apply the Z-score standardization formula Z = (X-μ) / σ to normalize each index, eliminate the influence of the dimension difference of different indexes, and obtain the normalized transformation result; Map the interval of the normalized transformation result, map the Z value to the [0, 1] standard interval by using the Sigmoid function, and obtain the standardized performance indexes.
[0009] Preferably, the calculation of the comprehensive performance score of each channel by weighted summation based on the standardized performance indexes adopts the scoring formula Score = w1 x reach rate + w2 x conversion rate - w3 x complaint rate - w4 x delay index to generate a channel performance benchmark score, which includes: Obtaining a business target configuration parameter, setting a weight coefficient of each index according to a marketing scene, wherein a reach rate weight w1, a conversion rate weight w2, a complaint rate weight w3, and a delay index weight w4 satisfy w1+w2+w3+w4=1, and a weight configuration scheme is established; Based on the weight configuration scheme and the standardized performance index, a scoring formula Score=w1×reach rate+w2×conversion rate-w3×complaint rate-w4×delay index is used for calculation to obtain a real-time performance score of each channel; According to the real-time performance score, a passing line threshold of 60 points and an excellent line threshold of 85 points of channel performance are set to generate the channel performance benchmark score.
[0010] Preferably, the multi-dimensional performance data of each channel in the past 24 hours is collected, the time sequence characteristics of historical conversion data are analyzed through an LSTM neural network, a three-layer LSTM network structure is used for time sequence pattern learning, and the conversion rate trend change of each channel in the future 2-4 hours is predicted, including: Based on the multi-dimensional performance data, the conversion rate moving average values of 3 hours, 6 hours, and 24 hours are calculated, the trend characteristics such as first-order difference and second-order difference of the conversion rate are constructed, and the input feature vector of the prediction model is formed; Based on the input feature vector of the prediction model, a three-layer LSTM network structure is designed, wherein the input layer receives 24 time window feature vectors, the hidden layer captures long and short term dependencies through 128 LSTM units, and a time sequence prediction model is established; The time sequence prediction model is deployed as an online inference service, and the conversion rate trend change of each channel in the future 2-4 hours is predicted in combination with the standardized performance index.
[0011] Preferably, the gradient optimization method is used to solve the weight distribution scheme with the goal of maximizing the total ROI, a target function F(w)=Σ(wi×predicted conversion rate i×weight coefficient) is set, and the optimal distribution ratio of each channel is calculated, including: A target function F(w)=Σ(wi×predicted conversion rate i×weight coefficient) of weight optimization is established, a constraint condition Σwi=1 and wi≥0 is set, and an optimization problem mathematical model is constructed; Based on the optimization problem mathematical model, the Adam optimization algorithm is used for solving, the learning rate is set to 0.01, the maximum iteration number is set to 100 times, and the weight distribution vector converging to the optimal solution is calculated; The weight distribution vector is obtained, compared and analyzed with the current weight configuration, and the optimal distribution ratio is output.
[0012] Preferably, the smooth transition mechanism of the weight configuration is started simultaneously, including: The optimal distribution ratio is obtained by using an exponential decay smoothing algorithm, setting a smoothing coefficient a as 0.3, and establishing a weight smoothing transition formula; Based on the weight smoothing transition formula, the smoothed weight = a x updated weight + (1-a) x current weight is calculated to generate a progressive weight adjustment sequence; According to the progressive weight adjustment sequence, the weight switching operation is performed in steps to avoid the impact of sharp weight changes on system stability.
[0013] Preferably, the method further comprises an A / B testing framework for verifying the effect of the new channel: Randomly select no more than 5% of the user traffic as a test sample, and split the traffic by user ID hash modulo to allocate the test traffic to the channel to be tested, and establish an experimental group and a control group; Collect conversion data of the experimental group and the control group, calculate the confidence interval of conversion rate based on Beta-Binomial conjugate prior distribution, perform Bayesian statistical test, and obtain the Bayesian statistical test result; Based on the Bayesian statistical test result, when the conversion rate difference between the experimental group and the control group reaches statistical significance and the relative improvement exceeds 5%, the new channel that passes the verification is automatically included in the weight distribution pool, and an updated channel weight distribution pool is obtained.
[0014] Preferably, the method further comprises an abnormality detection and emergency handling mechanism: Based on the historical 30-day data, the mean mu and standard deviation sigma of each performance indicator are calculated, and a 3-sigma abnormality detection algorithm is used to establish an abnormality judgment benchmark; Real-time monitoring of the performance indicator data of each channel, when the real-time indicator of a certain channel exceeds the range of mu+3sigma, it is determined as performance abnormality and triggers the alarm mechanism to identify the abnormal channel; Upon detection of the abnormal channel, the weight of the abnormal channel is immediately reduced to 0%, and is proportionally distributed to other channels, while the manual intervention notification mechanism is started to ensure the overall availability of the system.
[0015] Preferably, the method further comprises a channel performance deep mining based on distributed asynchronous exploration: The current channel configuration state of the optimal distribution ratio is taken as the starting point of exploration, the parameters including channel weight, sending period and content template are taken as exploration dimensions, a multi-dimensional decision tree structure containing n channel nodes and depth D is constructed, and a channel configuration exploration tree is generated; Based on the channel configuration exploration tree, k configuration variation branches are generated in parallel, each branch is assigned to a distributed computing node for asynchronous execution, small-scale A / B testing is used to verify the conversion effect of each configuration scheme, and key indicator data including conversion rate and ROI are collected; According to the key indicator data, the regret value of each configuration scheme is calculated, the low-efficiency branch is pruned based on the regret minimization principle, the optimal branch is reserved to continue deep exploration, and the global optimal configuration is converged within at most 2n+O(k 2 2 KD ) times of moving operations.
[0016] Preferably, the generation of the channel configuration exploration tree comprises: The current channel configuration state of the optimal distribution ratio is set as the root node, including the complete parameter set of the weight distribution, sending strategy and content template of all channels, and an exploration starting state is established; Based on the exploration starting state, fine-tuning changes are made to each configuration parameter to generate multiple sub-node branches, wherein each sub-node represents a specific variation of a certain parameter, forming a configuration variation space; The configuration variation space is expanded to a depth D according to the depth-first search strategy, and each leaf node corresponds to a specific configuration scheme, generating the channel configuration exploration tree.
[0017] Preferably, the method further comprises channel stability prediction based on a two-dimensional threshold cellular automaton: The attribute information of the operator type and the geographical area of all short message channels is obtained, each channel is mapped to a two-dimensional grid structure, each grid unit contains a state vector of performance score, load rate, failure probability and neighborhood influence degree, and a channel network topology model is established; Based on the channel network topology model, a threshold conversion rule is designed, when the channel performance benchmark score is lower than 60 points and the continuous low score time exceeds 1 hour, it is determined as a warning state, when the number of fault channels in 8 neighborhoods is greater than or equal to 3 and the current load rate exceeds 80%, the failure probability is increased, and the cellular automaton state update is performed; According to the cellular automaton state update, the system is iteratively evolved for 100 time steps to simulate the propagation and diffusion process of channel failures, calculate the overall stability indicator of the system, and predict the stability trend of the channel network in the next 2-4 hours.
[0018] Preferably, the execution of the cellular automaton state update comprises: The state vector of each channel at the current time is obtained, and state judgment is performed according to the performance threshold rule, neighborhood propagation rule and cascading failure rule, and the state probability distribution of each channel at the next time is calculated; Based on the state probability distribution, the state values of all channels are updated in parallel, the neighborhood influence degree and the failure propagation coefficient are calculated synchronously, the diffusion process of failures in the channel network is simulated, and the failure diffusion process is obtained. The fault diffusion process is quantitatively evaluated, and stability indexes including a health channel proportion and a maximum connected component size are calculated to form a system overall stability score.
[0019] Preferably, the method further comprises a fast approximate minimum spanning tree based path routing optimization: User features including geographical locations of user groups, device types, historical behavior preferences, active time periods, and channel features of channel operators, coverage areas, delay characteristics, and cost structures are extracted to construct a high-dimensional feature space and generate user-channel combination feature vectors; Based on the user-channel combination feature vectors, an FAMST three-stage algorithm is used to construct an optimal connection relationship of the path routing, and through three stages of approximate nearest neighbor graph construction, neural network component connection discovery, and iterative edge refinement, a personalized routing network topology is established. According to the personalized routing network topology, the shortest distance and matching degree to each channel are calculated for each sending task, and the optimal channel is selected in combination with the optimal distribution ratio to realize user-level precise path routing optimization.
[0020] Preferably, the establishment of the personalized routing network topology comprises: Based on the user-channel combination feature vectors, an approximate nearest neighbor graph is constructed using Locality Sensitive Hashing, and the k=10 most similar user-channel combinations are calculated to establish an initial connection network. The initial connection network is input into a neural network model, node features are extracted by an encoder, and connection probabilities between node pairs are calculated using a connection predictor to discover potential high-value connection relationships. The potential high-value connection relationships are iteratively refined by removing the lowest 5% edges, checking network connectivity and adding key connections, and after at most 10 iterations of optimization, the personalized routing network topology is generated.
[0021] A short message marketing effect optimization system based on multi-channel dynamic weights comprises: A data acquisition module is configured to acquire real-time performance data of a plurality of short message channels. A standardization processing module is configured to perform standardization processing on the real-time performance data, and map reach rate, conversion rate, complaint rate, and delay indicators to the [0, 1] interval using a Z-score standardization method to obtain standardized performance indicators. A score calculation module is configured to calculate a comprehensive performance score of each channel by weighted summation based on the standardized performance indicators, and generate a channel performance benchmark score using a score formula Score = w1×reach rate + w2×conversion rate - w3×complaint rate - w4×delay indicator. a prediction module, configured to collect multi-dimensional performance data of each channel in the past 24 hours, analyze time sequence characteristics of historical conversion data through an LSTM neural network, learn time sequence patterns by using a three-layer LSTM network structure, and predict conversion rate trend changes of each channel in the next 2-4 hours; a weight optimization module, configured to solve a weight distribution scheme with the goal of maximizing the overall ROI based on the conversion rate trend changes and the channel performance benchmark score by using a gradient optimization method, set a target function F(w) = Σ(wi x predicted conversion rate i x weight coefficient), and calculate the optimal distribution proportion of each channel; a shunt adjustment module, configured to adjust the shunt strategy of the short message queue in real time according to the optimal distribution proportion, re-distribute the to-be-sent short messages to each channel queue, start a smooth transition mechanism of the weight configuration, and realize dynamic optimization of the short message marketing effect.
[0022] The application has the following advantages: 1. The real-time dynamic weight distribution mechanism breaks through the limitations of traditional static configuration, realizes intelligent allocation of channel resources, and dynamically optimizes the short message marketing effect; 2. The multi-dimensional channel evaluation model is established, the key indicators such as reach rate, conversion rate and complaint rate are comprehensively considered, the disadvantages of single index orientation are avoided, and the comprehensiveness and accuracy of evaluation are improved; 3. The machine learning time sequence prediction capability is integrated, the channel effect trend is predicted based on the LSTM neural network, the forward-looking of decision-making is improved, and the channel performance change can be predicted in advance; 4. The A / B test automation framework is constructed, the new channel is quickly verified and adaptively included, the manual trial and error cost is reduced, and the system iteration efficiency is improved; 5. The abnormality detection and emergency handling mechanism is designed, the stable operation and rapid recovery of the system in the event of channel failure are ensured, and the overall reliability of the system is improved. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 The flowchart of the short message marketing effect optimization method based on multi-channel dynamic weight provided by the embodiment of the present application; Figure 2 The structural diagram of the short message marketing effect optimization system based on multi-channel dynamic weight provided by the embodiment of the present application. DETAILED DESCRIPTION
[0025] The application will be further described below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are intended to be illustrative only and not limiting of the application.
[0026] As shown in the drawings, the short message marketing effect optimization method based on multi-channel dynamic weight provided by the embodiments of the application comprises the following steps: Figure 1 Step S1: acquiring real-time performance data of a plurality of short message channels. Step S2: performing standardization processing on the real-time performance data, and mapping the touch rate, the conversion rate, the complaint rate and the delay index to the interval [0, 1] by using a Z-score standardization method to obtain standardized performance indexes.
[0027] In this embodiment, the performance data of each short message channel is acquired in real time through a data acquisition interface, including key performance indexes such as touch rate, conversion rate, complaint rate and delay index. The touch rate reflects the proportion of short messages successfully delivered to target users, and is an important index for measuring the basic quality of the channel. The conversion rate represents the proportion of users performing expected behaviors (such as clicking links, completing purchases, etc.) after receiving short messages, and is directly related to the ROI and business goal realization of the marketing activities. The complaint rate refers to the proportion of users who complain about receiving short messages, and is an important dimension for measuring user experience and content compliance. The delay index reflects the time delay of short messages from sending to delivery, which is particularly critical for scenarios with high time efficiency requirements (such as verification codes, activity previews, etc.).
[0028] A distributed data acquisition architecture is adopted to acquire real-time performance data from each short message channel through API interfaces or database synchronization, etc. To ensure the real-time and accuracy of data acquisition, an adaptive sampling strategy is implemented: under normal circumstances, a sampling period of 5 minutes is adopted, and when performance fluctuations are detected, the sampling frequency is automatically increased to 1 minute per time to ensure that minor changes in channel performance can be captured. At the same time, the raw data collected is preliminarily cleaned to remove obvious outliers (such as touch rate suddenly dropping to 0%), and interpolation processing is performed on missing data to ensure the data quality for subsequent analysis. In addition, time stamps, batch IDs and other meta information during data acquisition are recorded for subsequent time series analysis and data correlation.
[0029] Step S2: performing standardization processing on the real-time performance data, and mapping the touch rate, the conversion rate, the complaint rate and the delay index to the interval [0, 1] by using a Z-score standardization method to obtain standardized performance indexes.
[0030] The numerical ranges and distribution characteristics of different performance indicators differ significantly, and direct comparison or calculation can lead to inconsistent dimensions. For example, reach rate and conversion rate are usually expressed in percentages, with numerical ranges between 0-100%; while complaint rate is usually a small percentage, possibly within the range of 0-1%; delay indicators may be in units of seconds or milliseconds, with numerical ranges from tens of milliseconds to several seconds. The difference in order of magnitude and distribution can lead to the dominance of indicators with larger values in the final result in the comprehensive score calculation, which cannot objectively reflect the true importance of each indicator.
[0031] The purpose of standardization is to eliminate the dimensional differences between different indicators, so that each indicator can be compared and calculated under a unified standard. Z-score standardization is a commonly used standardization method, which converts the original data into a standard normal distribution based on the mean and standard deviation of the data. After Z-score standardization, all indicators are mapped to the same standard scale, and the numerical value represents the deviation from the mean (in units of standard deviation). Subsequently, the Z value is mapped to the [0, 1] interval through the Sigmoid function, which facilitates subsequent weighted calculation and score analysis.
[0032] After standardization, different performance indicators of each channel can be objectively compared and weighted within a unified numerical range, avoiding the undue influence of dimensional differences in the original data on the evaluation results. This step provides a standardized data basis for subsequent comprehensive performance scoring and is a key pre- step for multi-dimensional channel evaluation.
[0033] Step S3: Based on the standardized performance indicators, the comprehensive performance score of each channel is calculated by weighted summation, using the scoring formula Score = w1×reach rate + w2×conversion rate - w3×complaint rate - w4×delay indicator, to generate a channel performance benchmark score.
[0034] After obtaining the standardized performance indicators, the importance of each indicator needs to be considered comprehensively to generate a comprehensive score that can fully reflect the channel performance. In actual business scenarios, different performance indicators have different degrees of influence on marketing effectiveness, and different weights need to be assigned. For example, in activities with sales conversion as the main goal, the conversion rate indicator may be more important than the reach rate; while in brand promotion activities, low complaint rate and high reach rate may be more valued.
[0035] The comprehensive performance score is calculated by weighted summation, and the scoring formula is: Score = w1 x reach rate + w2 x conversion rate - w3 x complaint rate - w4 x delay index. Among them, w1, w2, w3, w4 represent the weight coefficients of each index, which satisfy the constraint condition w1+w2+w3+w4=1 to ensure the normalization of the scoring result. Since the complaint rate and the delay index are negative indicators (the lower the value, the better), subtraction is used in the formula to ensure that the influence direction of each index on the final score is consistent.
[0036] Through this weighted scoring mechanism, the weight configuration of each index can be dynamically adjusted according to different marketing scenarios and business goals, and a performance score that better meets actual needs can be obtained. For example, for time-sensitive verification code messages, the weight of reach rate and delay index may be increased; for brand-sensitive marketing activities, the weight of complaint rate index may be increased; for sales promotion activities, more attention may be paid to the conversion rate index.
[0037] This multi-dimensional weighted scoring method breaks through the limitations of traditional single-index evaluation and can comprehensively and objectively evaluate the comprehensive performance of each channel, providing a scientific basis for subsequent dynamic weight allocation. At the same time, by setting performance pass lines and excellent lines, the performance of the channel can be managed in stages, realizing the optimal allocation of resources.
[0038] Step S4: Collect multi-dimensional performance data of each channel in the past 24 hours, analyze the time sequence characteristics of historical conversion data through LSTM neural network, and use a three-layer LSTM network structure to learn time sequence patterns to predict the conversion rate trend change of each channel in the next 2-4 hours.
[0039] Channel performance is not static but dynamically changes over time. In addition to the current performance score, it is also necessary to predict the trend of channel performance changes in the future to achieve more forward-looking weight optimization. Especially for the core business indicator of conversion rate, its change is directly related to the ROI and effect of marketing activities. Predicting future conversion rate trends can help anticipate possible performance fluctuations and achieve more stable marketing results.
[0040] In order to achieve high-precision conversion rate trend prediction, the LSTM (Long Short-Term Memory) neural network model is used. LSTM is a special recurrent neural network (RNN) structure designed specifically for processing and predicting time series data. Unlike traditional feedforward neural networks, LSTM has an internal state (cell state) and multiple gating mechanisms (forget gate, input gate, output gate), which can effectively capture long-term dependencies and time sequence patterns in data. These characteristics make LSTM particularly suitable for analyzing and predicting the time series changes of channel performance.
[0041] First, collect multi-dimensional performance data of each channel in the past 24 hours, including conversion rate, reach rate, sending volume, time characteristics, etc. Then, perform feature engineering processing to construct feature vectors that can reflect the trend changes of different time scales. Based on these features, train a three-layer LSTM network model to capture the timing patterns and influencing factors of conversion rate changes. Finally, through the trained model, the conversion rate trend of each channel in the next 2-4 hours can be predicted, providing a forward-looking basis for subsequent weight optimization.
[0042] Through this machine learning-based prediction capability, it is no longer necessary to rely solely on current or historical performance data, but rather to "see" future performance changes, enabling more forward-looking dynamic optimization and further improving the overall effectiveness of SMS marketing.
[0043] Step S5: Based on the conversion rate trend changes and the channel performance benchmark score, use gradient optimization method to solve the weight allocation scheme with the goal of maximizing overall ROI, set the objective function F(w) = Σ(wi × predicted conversion rate i × weight coefficient), calculate the optimal distribution proportion of each channel.
[0044] After collecting real-time performance data of channels, calculating standardized performance scores and predicting future conversion rate trends, it is necessary to determine how to optimally allocate SMS traffic to each channel to maximize overall marketing effectiveness. This is essentially a resource optimization problem that needs to consider the current performance status of the channel, future performance trends and business target requirements.
[0045] Using mathematical optimization methods, the weight allocation problem is formalized as a constrained optimization problem. The objective function is set to maximize the overall ROI, i.e. by optimizing the weight allocation proportion of each channel, the expected overall conversion effect is maximized. The objective function is F(w) = Σ(wi × predicted conversion rate i × weight coefficient), where wi represents the weight proportion allocated to the ith channel, predicted conversion rate i represents the predicted conversion rate of the ith channel in the future period, and weight coefficient i is a parameter related to channel cost and importance.
[0046] The optimization process needs to consider various constraints, including the sum of all weights must be 1, the weight of each channel must be non-negative, and the upper limit of the weight of poor performance channels is limited. To solve this optimization problem, the Adam optimization algorithm is used, which is an adaptive learning rate optimization method based on gradient descent, which can effectively handle non-convex optimization problems, ensuring convergence while providing faster solution speed.
[0047] By gradient optimization solution, the optimal weight distribution ratio of each channel is obtained, which represents the resource allocation scheme that can maximize the ROI under the current state and predicted trend. This dynamic weight calculation method based on mathematical optimization breaks through the limitations of traditional fixed weight or simple rules, and can automatically adjust the optimal allocation strategy according to complex and variable factors, realizing truly intelligent resource allocation.
[0048] Step S6: According to the optimal distribution ratio, real-time adjust the shunt strategy of the SMS queue, re-allocate the to-be-sent SMS to each channel queue, and start the smooth transition mechanism of weight configuration, to realize the dynamic optimization of SMS marketing effect.
[0049] After calculating the optimal distribution ratio, the theoretical weight configuration needs to be converted into actual SMS shunt operation. Direct one-time adjustment to the new weight configuration may cause impact on stability, so a smooth transition mechanism needs to be designed to ensure stable operation during weight adjustment.
[0050] The core of shunt adjustment is to allocate the to-be-sent SMS to each channel's sending queue according to the optimal ratio. Weighted round robin or random sampling is used to realize real-time traffic distribution. For example, for the weight configuration [A:45%, B:35%, C:20%], 45 out of every 100 SMS may be allocated to channel A, 35 to channel B, and 20 to channel C. To avoid the concentration of SMS sending in some channels, fine-grained interleaved allocation is usually implemented to ensure load balancing of each channel.
[0051] At the same time, the smooth transition mechanism of weight configuration is started, and a gradual adjustment strategy is adopted to avoid sudden changes in weight. Smooth transition usually uses exponential decay smoothing algorithm to disperse weight adjustment into multiple small steps, each step only moves a small step towards the target weight until it finally approaches the target configuration. This gradual adjustment allows time to adapt to the new traffic distribution, avoiding instability or performance fluctuations caused by sudden traffic changes.
[0052] Through this dynamic adjustment and smooth transition mechanism, the overall effect of SMS marketing can be continuously optimized while maintaining stable operation and reliability.
[0053] Step S2 specifically includes: Step S21: Collect the original data of the reach rate, conversion rate, complaint rate and delay index of each SMS channel, calculate the mean μ and standard deviation σ of each index in the historical data, and establish the data distribution baseline.
[0054] In this step, the original data of performance indicators of each SMS channel in the past 7 days is collected and sorted. The selection of 7 days as the historical data window is based on practical experience, which can not only contain enough data samples to reflect the general distribution characteristics of the indicators, but also avoid the influence of seasonal or trend changes caused by long time window. For each performance indicator, the historical mean μ and standard deviation σ are calculated respectively as the benchmark parameters for standardization.
[0055] Taking the reach rate indicator as an example, the historical mean calculation formula is: μreach rate = (Σ historical reach rate value) / historical data points. This mean reflects the central tendency of the indicator under normal circumstances. The standard deviation calculation formula is: σreach rate = √[(Σ (historical reach rate value - μreach rate)²) / historical data points], and the standard deviation reflects the fluctuation or dispersion of the indicator.
[0056] For each performance indicator of each channel, the above calculation is performed to generate a parameter matrix containing the mean and standard deviation as the data distribution baseline. This baseline is not only used for subsequent standardization processing, but also serves as a reference for anomaly detection. In order to adapt to the dynamic changes of the indicator distribution, a sliding window mechanism is used to update these parameters regularly, ensuring that the standardization process is based on the latest data distribution characteristics.
[0057] At the same time, the historical quantile information of each indicator (such as 25%, 50%, 75%, 95% quantile) is also saved, which is used to understand the skewness of the data and define the abnormal value, providing more rich statistical reference for subsequent standardization mapping. Through this step, a comprehensive data distribution baseline is established, laying a statistical foundation for subsequent standardization processing.
[0058] Step S22: Based on the data distribution baseline, apply the Z-score standardization formula Z = (X-μ) / σ to each indicator for normalization transformation, eliminating the influence of different indicator dimension differences, and obtain the normalized transformation result.
[0059] In this step, the performance indicator data of each channel collected in real time is normalized by applying the Z-score standardization formula. Z-score standardization is a linear transformation method based on mean and standard deviation, which converts the original data points to the deviation from the mean (in units of standard deviation). The Z-score formula is defined as: Z = (X-μ) / σ, where X is the original data value, μ is the historical mean of the indicator, and σ is the historical standard deviation of the indicator.
[0060] For example, if the current reach rate of a channel is 95%, and the historical mean μ reach rate is 90% and the standard deviation σ reach rate is 3%, the Z value of the current reach rate is calculated as: Z reach rate = (95% - 90%) / 3% = 1.67. This Z value indicates that the current reach rate is 1.67 standard deviations higher than the historical mean, which is a good performance.
[0061] For "the smaller the better" indicators such as complaint rate, the opposite processing method is adopted, and the original formula is adjusted to Z = (μ - X) / σ, to ensure that the larger the Z value represents the better performance, maintaining consistency with other indicators. For example, if the current complaint rate of a channel is 0.05%, and the historical mean μ complaint rate is 0.1% and the standard deviation σ complaint rate is 0.03%, the Z value of the current complaint rate is calculated as: Z complaint rate = (0.1% - 0.05%) / 0.03% = 1.67, indicating that the current complaint rate is 1.67 standard deviations lower than the historical mean, which is a good performance.
[0062] By Z-score standardization, indicators of different dimensions and distribution ranges are uniformly converted to Z values of standard normal distribution, eliminating the dimensional differences between different indicators, so that each indicator can be directly compared and calculated. This normalization transformation preserves the relative relationship and statistical characteristics of the original data, while providing a unified measurement standard, laying the foundation for subsequent performance scoring.
[0063] Step S23: Interval mapping is performed on the normalization transformation result, and a Sigmoid function is used to map the Z value to the [0, 1] standard interval to obtain the standardized performance indicator.
[0064] The result of Z-score standardization is theoretically a variable that follows a standard normal distribution, with a value range of (-∞, +∞). In order to facilitate subsequent weighted calculation and result interpretation, the Z value needs to be further mapped to a bounded interval. In this step, the Sigmoid function is used to map the Z value to the (0, 1) interval. The Sigmoid function is a commonly used S-shaped curve function, defined as: f(Z) = 1 / (1 + e (-Z) ).
[0065] The sigmoid function has the following properties: when Z tends to negative infinity, f(Z) tends to 0; when Z tends to positive infinity, f(Z) tends to 1; when Z=0, f(Z)=0.5. This means that the same performance as the historical average will get a standardized score of 0.5, and the performance score will be higher than 0.5 if it is better than the average, and lower than 0.5 if it is worse than the average. Due to the non-linear characteristics of the sigmoid function, extreme Z values (such as ±3 or more) will be compressed to the interval close to 0 or 1, avoiding the excessive influence of extreme values on subsequent calculations.
[0066] For example, the above reach rate is 1.67, and the sigmoid function is used to calculate: f(1.67) = 1 / (1 +e (-1.67) ) ≈ 0.84. This 0.84 standardized score indicates that the reach rate of the channel is better than 84% of the historical data.
[0067] For "the smaller the value, the better" indicators such as complaint rate and delay indicators, after applying the adjusted Z-score formula, the sigmoid function can be directly used for mapping, maintaining consistent scoring logic (the larger the value, the better the performance).
[0068] By mapping the Z values of each indicator to the standardized performance indicators in the (0, 1) interval through the sigmoid function, these indicators have a unified numerical range and interpretation method, making it easy to calculate and score later. At the same time, this mapping method also makes the scoring results easier to understand and compare, providing a standardized indicator system for multi-dimensional evaluation of channel performance.
[0069] Step S3 specifically includes: Step S31: Obtain business target configuration parameters, set the weight coefficients of each indicator according to the marketing scenario, wherein the reach rate weight w1, the conversion rate weight w2, the complaint rate weight w3, and the delay indicator weight w4 satisfy w1+w2+w3+w4=1, and a weight configuration scheme is established.
[0070] In this step, according to the characteristics of the current business target and marketing scenario, the weight coefficients of each performance indicator are set. The weight configuration reflects the relative importance of each indicator in a specific business scenario, directly affecting the comprehensive performance score of the channel. In order to adapt to different types of marketing activities, multiple weight configuration schemes are preset, and marketing personnel can customize according to specific needs.
[0071] For sales promotion marketing activities (such as promotional offers, time-limited discounts, etc.), a higher conversion rate weight is set, configured as: w1=0.3 (reach rate), w2=0.5 (conversion rate), w3=0.15 (complaint rate), and w4=0.05 (delay index). The main goal of such activities is to achieve sales conversion, so the conversion rate index weight is the highest, accounting for 50%; At the same time, ensure the basic reach quality and user experience, allocate 30% weight to the reach rate and 15% weight to the complaint rate; Since the timeliness requirement is relatively low, the delay index only accounts for 5% of the weight.
[0072] For brand maintenance marketing activities (such as brand promotion, member care, etc.), more attention is paid to user experience and brand image, configured as: w1=0.25 (reach rate), w2=0.3 (conversion rate), w3=0.35 (complaint rate), and w4=0.1 (delay index). Such activities need to strictly control the complaint rate to avoid negative impact on the brand image, so the complaint rate index weight is the highest, accounting for 35%; At the same time, maintain a certain conversion effect and reach quality, allocate 30% weight to the conversion rate and 25% weight to the reach rate; The weight of the delay index is 10% to ensure that the information is delivered in a timely manner.
[0073] For emergency notification marketing activities (such as verification code, reminders, etc.), high reach rate and low delay are emphasized, configured as: w1=0.5 (reach rate), w2=0.2 (conversion rate), w3=0.1 (complaint rate), and w4=0.2 (delay index). The most important thing for such activities is to ensure that information is quickly and accurately delivered to users, so the reach rate index weight is the highest, accounting for 50%; At the same time, the delay index is also important, accounting for 20% of the weight; The conversion rate and the complaint rate are relatively secondary, accounting for 20% and 10% of the weight, respectively.
[0074] Through the configuration management interface, business personnel can adjust these preset schemes or create new weight configurations according to actual needs. All configuration schemes must meet the constraint condition that the sum of the weights is 1 to ensure the standardization and comparability of the scoring results. Through this flexible weight configuration mechanism, the channel performance score that best meets the business goals can be generated for different marketing scenarios.
[0075] Step S32: Based on the weight configuration scheme and the standardized performance index, the scoring formula Score = w1 x reach rate + w2 x conversion rate - w3 x complaint rate - w4 x delay index is substituted to calculate the real-time performance score of each channel.
[0076] In this step, the normalized performance indicators obtained in step S23 are combined with the weight coefficients set in step S31, and the comprehensive performance score of each channel is calculated by weighted summation. During the calculation process, first ensure that all normalized indicators have been mapped to the [0, 1] interval, and for "the smaller the value, the better" indicators (such as complaint rate, delay indicator) have been appropriately converted to make all indicators consistent in direction (the larger the value, the better the performance).
[0077] The specific score calculation formula is: Score = w1 x normalized reach rate + w2 x normalized conversion rate + w3 x (1- normalized complaint rate) + w4 x (1- normalized delay indicator). Note that for complaint rate and delay indicator, since they have been processed in the standardization stage (through Z = (μ-X) / σ and Sigmoid function), the normalized values can be used directly here, without further conversion of 1-x. Therefore, the actual formula used is simplified to: Score = w1 x normalized reach rate + w2 x normalized conversion rate + w3 x normalized complaint rate + w4 x normalized delay indicator.
[0078] To make the final score result more intuitive, multiply the result of weighted summation by 100 to convert it to a score range of 0-100. For example, a channel has the following normalized indicators: reach rate 0.85, conversion rate 0.72, complaint rate 0.93 (converted to the larger the value the better), delay indicator 0.88 (converted to the larger the value the better), and under the weight configuration of sales promotion type marketing activities (w1=0.3, w2=0.5, w3=0.15, w4=0.05), its performance score is calculated as: Score = (0.3 x 0.85 + 0.5 x 0.72 + 0.15 x 0.93 + 0.05 x 0.88) x 100 = 79.6 points.
[0079] Perform the above calculation for all channels to obtain a set of score results reflecting the current comprehensive performance of each channel. These scores take into account multiple performance dimensions and their relative importance, and can comprehensively and objectively reflect the overall performance of each channel in a specific business scenario. The score results are updated in real time and dynamically change with the collection of new performance data, providing the latest decision basis for subsequent weight optimization and shunt adjustment.
[0080] Step S33: According to the real-time performance score, set the pass line threshold of channel performance to 60 points and the excellent line threshold to 85 points, and generate the channel performance benchmark score.
[0081] In this step, two key performance score thresholds are set according to business needs and practical experience: passing line threshold 60 points and excellent line threshold 85 points. These two thresholds divide the channel performance into three levels: not passing (<60 points), passing (60-85 points), and excellent (>85 points). The setting of thresholds considers industry standards, historical data analysis and business needs, aiming to establish an objective channel performance grading standard.
[0082] The passing line threshold of 60 points represents the minimum acceptable level of channel performance. Channels below this threshold are considered to have poor performance and will be subject to a series of measures to restrict their use, including: reducing their weight allocation proportion (usually no more than 10%), triggering performance warning notifications, automatically starting backup channels, etc. In extreme cases (such as consecutive periods below the passing line), the weight may be reduced to 0%, completely suspending the use of the channel, and notifying the operator for manual intervention and channel status check.
[0083] The excellent line threshold of 85 points represents an excellent level of channel performance. Channels exceeding this threshold are considered high-quality channels and will be given priority in increasing their weight allocation to improve overall marketing effectiveness. For channels with excellent performance, their configuration parameters and usage scenarios are also recorded to form best practice cases, providing references for future channel selection and optimization.
[0084] Based on these two thresholds and real-time performance scores, performance benchmark scores for each channel are generated, which are comprehensive evaluation results including scores and levels. For example, if a channel's real-time performance score is 79.6 points, its performance benchmark score is "79.6 points (passing)". These benchmark scores are not only used for subsequent weight optimization calculations, but also serve as key indicators for channel performance monitoring dashboards, providing a visual display of the current status of each channel.
[0085] By setting clear performance thresholds and grading standards, the standardized management of channel performance is achieved, providing objective decision-making basis for dynamic weight adjustment, and also facilitating operators to quickly identify problem channels and high-quality channels, improving overall operational efficiency.
[0086] Step S4 specifically includes: Step S41: Based on the multi-dimensional performance data, calculate the conversion rate sliding average of 3 hours, 6 hours and 24 hours, construct trend features such as first-order difference and second-order difference of conversion rate, and form input feature vectors of the prediction model.
[0087] In this step, the performance data of each channel collected in the past 24 hours is processed through feature engineering to extract and construct feature vectors that can reflect the timing characteristics, providing high-quality input data for the LSTM prediction model. Feature engineering is the key to the prediction effect of machine learning, and good feature design can significantly improve the prediction accuracy of the model.
[0088] First, the moving average of different time windows is calculated to capture trend information at different time scales. Specifically, it includes: 3-hour moving average: reflects short-term trend changes and is sensitive to recent fluctuations; 6-hour moving average: reflects medium-term trends and smooths the impact of short-term fluctuations; 24-hour moving average: reflects long-term trends and represents the average level of a complete day cycle.
[0089] These three different time scale moving averages together constitute a multi-scale representation of conversion rate changes, which can help the model understand the relationship between short-term fluctuations, medium-term trends and long-term baseline.
[0090] In addition to moving averages, difference features of conversion rates are also calculated to capture speed and acceleration information: First-order difference: the difference between the current value and the previous time value, reflecting the speed of change or the slope of the trend; Second-order difference: the difference of the first-order difference, reflecting the acceleration of change or the curvature of the trend.
[0091] These difference features can help the model identify dynamic characteristics of conversion rate changes, such as rising trends, falling trends, accelerating changes or decelerating changes, etc.
[0092] In addition, a series of auxiliary features are integrated to enrich the input information of the model: Time features: hour, day of the week, workday / weekend label, holiday label, etc.; Channel load features: current sending volume, queue length, load, etc.; Historical periodic patterns: historical conversion rate data at the same time period; Special event labels: marketing activities, maintenance, policy changes, etc.
[0093] For categorical features (such as day of the week), One-Hot encoding is used to convert them into numerical representations; for numerical features, normalization is performed to ensure the scale of each feature is consistent. Finally, a high-dimensional feature vector is generated for each channel at each time point, containing rich time series information and context features, providing comprehensive input data for the LSTM model to support high-precision conversion rate trend prediction.
[0094] Step S42: Based on the input feature vector of the prediction model, a three-layer LSTM network structure is designed, where the input layer receives the 24-time window feature vector, the hidden layer captures long and short-term dependencies through 128 LSTM units, and a time series prediction model is established.
[0095] In this step, a three-layer LSTM neural network model is constructed based on the feature vectors generated in step S41, which is used to capture the complex patterns and dependencies of the conversion rate time series data. LSTM (Long Short-Term Memory) is a special recurrent neural network structure designed to handle sequence data, which can effectively solve the gradient vanishing problem faced by traditional RNNs and capture long-term dependencies in data.
[0096] The designed three-layer LSTM network structure is as follows: Input layer: receives 24 consecutive time window (1 hour per window) feature vectors. The feature vector dimension of each time window is n (contains all the above features), so the shape of the input tensor is [batch_size, 24, n]. The input layer sends these features to the subsequent LSTM layer for processing.
[0097] First hidden layer: contains 128 LSTM units, which is the main feature extraction layer of the network. Each LSTM unit contains a memory cell and three gating mechanisms (forget gate, input gate, and output gate), which can learn when to remember or forget information. The main task of the first layer of LSTM is to extract basic time series patterns from the original feature sequence, such as short-term trends, periodic changes, etc. The output dimension of this layer is [batch_size, 24, 128], which retains the complete time series information.
[0098] Second hidden layer: contains 64 LSTM units, which further extract higher-level time series features. The second layer of LSTM learns more complex time series patterns, such as the interaction between different features, medium-term trend changes, etc., based on the basic features extracted by the first layer. The output dimension of this layer is [batch_size, 24, 64].
[0099] Third hidden layer: contains 32 LSTM units, which is the final feature fusion layer of the network. The third layer of LSTM integrates the features extracted by the first two layers to capture the highest level of time series patterns, such as long-term trends, special event impacts, etc. To meet the needs of multi-step prediction, this layer returns the output at the last time, with a dimension of [batch_size, 32].
[0100] Output layer: contains 4 neurons, which predict the conversion rate values for the next 1 hour, 2 hours, 3 hours, and 4 hours, respectively. The output layer uses a fully connected structure to map the output of the third layer of LSTM to the predicted values. To ensure the reasonableness of the prediction results, the output layer uses the Sigmoid activation function to constrain the predicted values within the range of [0, 1], which conforms to the actual meaning of the conversion rate.
[0101] The network training adopts mean square error (MSE) as the loss function, uses the Adam optimizer for parameter updating, and sets the initial learning rate to 0.001. To prevent overfitting, a Dropout layer with a dropout rate of 0.2 is added between the LSTM layers, and an early stopping strategy is used, which stops training when the validation set loss does not improve for 5 consecutive epochs. During training, the mini-batch method with a batch size of 32 is used, and the maximum number of training rounds is 100.
[0102] Through this deep LSTM network structure, various time series patterns and dependencies in conversion rate data can be effectively captured, enabling high-precision trend prediction. The model not only learns obvious seasonal and periodic patterns, but also adapts to sudden events and abnormal changes, providing reliable prediction basis for subsequent weight optimization.
[0103] Step S43: Deploy the time series prediction model as an online inference service, and predict the conversion rate trend changes of each channel in the next 2-4 hours based on the standardized performance indicators.
[0104] In this step, the trained LSTM model is deployed as an online inference service to realize real-time prediction of future conversion rates for each channel. The model deployment adopts a lightweight service architecture to ensure low latency and high availability in the prediction process, supporting weight optimization decisions every 15 minutes.
[0105] The deployment process includes the following key links: First, the trained LSTM model parameters are exported and converted into a format suitable for online inference. To improve inference efficiency, model quantization technology is used to convert 32-bit floating-point parameters to 8-bit integer representations, significantly reducing model size and improving calculation speed while maintaining prediction accuracy.
[0106] Second, a prediction service cluster with automatic scaling capability is constructed, using containerized deployment to ensure service flexibility and reliability. The prediction service receives real-time feature data, performs model inference, and returns the conversion rate prediction values at future time points. To handle sudden prediction request peaks, a request queue and priority mechanism are implemented to ensure that critical channel prediction tasks are processed first.
[0107] In the actual prediction process, the latest 24-hour performance data of each channel is collected, and the same feature engineering processing as in the training phase is performed to generate the input feature vector required by the model. Then, the feature vector is fed into the LSTM model to obtain the conversion rate prediction values for the next 1-4 hours. The prediction results will undergo post-processing steps, including reasonableness checks (to ensure that the prediction values are within a reasonable range) and smoothing processing (to avoid sharp fluctuations in prediction values).
[0108] The visualization of the prediction results is also implemented to help the operation personnel intuitively understand the future trends of each channel. The visualization content includes trend graphs, prediction confidence intervals, historical prediction accuracy, and other information to help the operation personnel judge the reliability of the prediction and develop response strategies.
[0109] In addition to directly predicting the absolute value of the conversion rate, the trend of the conversion rate is also calculated, and each channel is classified into "rising trend", "stable trend", and "declining trend". This trend classification is more intuitive and facilitates adjustment decisions by the weight optimization algorithm. For example, for channels with a predicted conversion rate showing a clear upward trend, their weight distribution may be increased in advance; for channels with a significantly declining predicted conversion rate, their weight proportion may be appropriately reduced.
[0110] By deploying the LSTM prediction model as an online service, real-time prediction of future channel performance is achieved, providing a forward-looking basis for dynamic weight optimization and significantly improving the ability to predict and respond to changes in channel performance.
[0111] Step S5 specifically includes: Step S51: Establish a target function F(w) = Σ(wi × predicted conversion rate i × weight coefficient) for weight optimization, set the constraint conditions Σwi = 1 and wi ≥ 0, and construct an optimization problem mathematical model.
[0112] In this step, the channel weight distribution problem is formalized as a mathematical optimization model, with the optimization objective and constraint conditions clearly defined. The objective function is set to maximize the overall ROI, i.e., by optimizing the weight distribution of each channel, the overall conversion effect is maximized.
[0113] The objective function is defined as the weighted sum of each channel weight, its predicted conversion rate, and the weight coefficient: F(w) = Σ(wi × predicted conversion rate i × weight coefficient i) where wi represents the weight proportion allocated to the ith channel, predicted conversion rate i represents the average conversion rate of the ith channel in the next 2-4 hours predicted by the LSTM model in step S4, and weight coefficient i is a tuning parameter related to channel characteristics, considering factors such as channel cost, quality stability, business importance, etc. The calculation formula of the weight coefficient is: weight coefficient i = base coefficient i × quality factor i × cost factor i The base coefficient reflects the basic importance of the channel and is usually set to 1; the quality factor is calculated based on the performance benchmark score of the channel, and the higher the score, the larger the quality factor; the cost factor is inversely proportional to the sending cost of the channel, and the lower the cost, the larger the cost factor. In this way, while pursuing high conversion rates, the quality stability and cost-effectiveness of the channel are also considered.
[0114] The constraints of the optimization model include: 1. The sum of all channel weights is 1: Σwi = 1; This ensures a reasonable weight distribution and the total weight is always 100%.
[0115] The weight of each channel is non-negative: wi ≥ 0, ; This is a physical constraint, weights cannot be negative.
[0116] 3. The upper limit of the weight of channels with performance scores below the passing score is: wi ≤ 0.1, if score i < 60; In order to control risks, the maximum weight of channels with poor performance is limited to no more than 10%.
[0117] 4. Weight change limit: |wi - current weight i| ≤ 0.2, ; In order to avoid the impact of drastic changes in weights on stability, the weight change in each adjustment shall not exceed 20 percentage points.
[0118] 5. Minimum weight guarantee: wi ≥ 0.05, if the current weight i>0; In order to maintain the active status of the channel and meet data collection needs, the enabled channel retains a minimum weight of 5%.
[0119] These constraints ensure that the weight optimization results not only maximize theoretical ROI but also ensure practical feasibility and stability. The entire optimization problem can be expressed as: maximize F(w) = Σ(wi × predicted conversion rate i × weight coefficient i), satisfying the above five constraints.
[0120] Step S52: Based on the mathematical model of the optimization problem, the Adam optimization algorithm is used to solve it, the learning rate is set to 0.01, the maximum number of iterations is 100, and the weight distribution vector that converges to the optimal solution is calculated.
[0121] In this step, the Adam (Adaptive Moment Estimation) optimization algorithm is used to solve the above optimization problem. The Adam algorithm combines the advantages of the momentum method and RMSProp, can adaptively adjust the learning rate of each parameter, and is suitable for handling non-convex optimization problems.
[0122] When implementing Adam optimization, the following specific parameters are set: Learning rate α: 0.01, controls the step size of parameter updates; Exponential decay rate of first moment estimation β1: 0.9, control the degree of influence of momentum; Exponential decay rate of second moment estimation β2: 0.999, control the smoothing degree of gradient variance; Numerical stability constant ε: 10 -8 , prevent division by zero error; Maximum number of iterations: 100, control the amount of calculation of the algorithm; Convergence condition: the change of objective function value in the last 5 iterations is less than 10 -4 .
[0123] The optimization process starts from the current weight configuration and gradually approaches the optimal solution through iterative updates. In each iteration, the gradient of the objective function with respect to each weight parameter is calculated, and then the Adam update rule is applied to adjust the weight values. To handle the constraints, the projected gradient descent method is used to project the weights into the feasible region after each update, ensuring that all constraints are satisfied.
[0124] In specific implementation, first, handle the weight non-negative and weight upper and lower limit constraints, limit each channel weight in the allowed range through clipping operation; then through normalization operation to ensure the sum of all weights is 1. This projection method is simple and efficient, which can ensure the satisfaction of the constraint conditions while not affecting the convergence of the algorithm.
[0125] Through the Adam optimization algorithm, the weight optimization problem can be efficiently solved, and the weight distribution scheme that can maximize the overall ROI under the current conditions can be found. The algorithm usually converges to the optimal solution or an approximate optimal solution within 20-50 iterations, with a calculation time of 100 milliseconds, meeting the real-time decision-making requirements.
[0126] Step S53: Obtain the weight distribution vector, compare and analyze with the current weight configuration, and output the optimal distribution ratio.
[0127] In this step, the weight distribution vector obtained by the optimization algorithm is compared and analyzed in detail with the current weight configuration being used, and a weight adjustment suggestion and explanation report are generated, and finally the optimal distribution ratio is output as the basis for the next operation.
[0128] First, calculate the difference between the new and old weight configurations, including absolute difference (percentage point change) and relative difference (percentage change). For channels with large weight changes (absolute difference more than 5 percentage points or relative difference more than 20%), special marking and analysis of the change reason are performed, mainly from the following aspects: Performance score change: whether the comprehensive performance score of the channel has changed significantly, such as upgrading from "qualified" to "excellent" or downgrading from "qualified" to "unqualified".
[0129] Predicted Conversion Rate Trend: Whether the future conversion rate predicted by the LSTM model is rising, stable, or falling, and the strength of the trend.
[0130] Weight Coefficient Adjustment: Whether the cost, quality, or other factors of the channel have changed, leading to an adjustment of the weight coefficient.
[0131] Constraint Impact: Whether the optimal solution has been affected by constraints such as minimum weight guarantee, maximum change amplitude limit, etc.
[0132] For example, the following analysis report may be generated: "Channel A's weight increased from 25% to 35% (+10 percentage points, a 40% relative increase). Main reasons: 1) Performance score increased from 78 to 87, reaching an excellent level; 2) Predicted future 3-hour conversion rate showed a significant upward trend, with an expected growth of 15%; 3) Cost factor remained stable." "Channel B's weight decreased from 30% to 20% (-10 percentage points, a 33% relative decrease). Main reasons: 1) Performance score decreased from 65 to 58, below the passing line; 2) Predicted future conversion rate showed a downward trend; 3) Affected by the minimum weight guarantee constraint, it did not decrease further." After completing the detailed analysis, the final optimal distribution ratio is output, which is a complete configuration containing all channels and their corresponding weights. For example, for three channels A, B, and C, the optimal distribution ratio may be represented as [A:45%, B:35%, C:20%], meaning that 45% of the short messages are allocated to channel A, 35% to channel B, and 20% to channel C.
[0133] This optimal distribution ratio not only serves as the direct basis for subsequent shunt adjustments, but also is recorded in the log for historical tracking and effect analysis. The effectiveness of historical weight adjustment is reviewed regularly (such as every week) to evaluate the accuracy and effectiveness of the weight optimization algorithm, and to continuously improve the optimization model and parameter settings.
[0134] Through detailed comparative analysis and clear weight change explanation, not only data-driven decision results are provided, but also the explainability and transparency of the decision-making process are enhanced, facilitating the understanding and supervision of automated decisions by operation personnel.
[0135] The smooth transition mechanism for simultaneously starting the weight configuration includes: Step S61: Obtain the optimal distribution ratio, use the exponential decay smoothing algorithm, set the smoothing coefficient α to 0.3, and establish the weight smoothing transition formula.
[0136] In this step, the exponential decay smoothing algorithm is used to achieve smooth transition of the weight, avoiding the impact of sudden changes in weight configuration on stability. The smooth transition formula is defined as: Smoothed weight = a * updated weight + (1-a) * current weight where a is the smoothing coefficient, taking a value of 0.3, indicating that each adjustment only advances 30% of the distance to the target weight. In this way, the optimal weight configuration can be gradually approached rather than being adjusted all at once.
[0137] Step S62: Based on the weight smoothing transition formula, calculate the smoothed weight = a * updated weight + (1-a) * current weight, generate a gradual weight adjustment sequence.
[0138] In this step, according to the smoothing transition formula, the smoothed weight value of each channel is calculated. For multiple adjustment periods, a gradual weight adjustment sequence is generated, allowing the weight configuration to gradually approach the optimal solution.
[0139] For example, assume that the current weight of channel A is 0.3 and the optimal weight is 0.45, then the smoothed weight sequence is: 1st adjustment: 0.3 + 0.3 * (0.45 - 0.3) = 0.345; 2nd adjustment: 0.345 + 0.3 * (0.45 - 0.345) = 0.3765; 3rd adjustment: 0.3765 + 0.3 * (0.45 - 0.3765) = 0.39855.
[0140] Through multiple small adjustments, the weight of channel A is gradually increased from 0.3 to close to 0.45.
[0141] Step S63: Perform weight switching operation in steps according to the gradual weight adjustment sequence to avoid the impact of sudden weight changes on system stability.
[0142] In this step, the system performs the actual weight switching operation according to the calculated gradual weight adjustment sequence. The specific implementation is as follows: Perform weight adjustment every 15 minutes, update the configuration according to the next weight value in the sequence; For each adjustment, the system monitors changes in key indicators (such as system load, queue length, channel response time); If abnormal fluctuations (indicator changes exceeding the preset threshold) are detected, pause the adjustment and revert to the last stable configuration; When the difference between the weight configuration and the target weight is less than the preset threshold (e.g. 1%), the adjustment process is complete.
[0143] Through this step-by-step switching strategy, a smooth transition to the new weight configuration can be achieved while maintaining stable system operation.
[0144] In this embodiment, in addition to the above basic process, the method further includes the following extended functions: 1. A / B test framework for new channel effect verification: Randomly select no more than 5% of user traffic as test samples, and achieve traffic splitting through user ID hash modulo method, allocate test traffic to the channel to be tested, establish experimental group and control group; Collect conversion data of the experimental group and the control group, calculate the confidence interval of conversion rate based on Beta-Binomial conjugate prior distribution, perform Bayesian statistical test, and obtain Bayesian statistical test result; Based on the Bayesian statistical test result, when the conversion rate difference between the experimental group and the control group reaches statistical significance and the relative improvement exceeds 5%, automatically include the verified new channel into the weight distribution pool to obtain the updated channel weight distribution pool.
[0145] Specifically, the A / B test framework is a key mechanism for the system to identify and verify high-performance channels, and is particularly suitable for evaluating the effectiveness of newly accessed channels. Traditional methods usually rely on experience judgment or simple historical data comparison, which is difficult to accurately quantify the statistical significance of channel effect difference and is easily affected by sample bias and random fluctuations. The system designs a scientific and rigorous A / B test framework based on Bayesian statistical principles, which realizes accurate evaluation and automatic decision of new channel performance.
[0146] The framework first uses a scientific traffic splitting method to randomly select no more than 5% of user traffic as test samples. This proportion is carefully designed to provide sufficient sample size to ensure statistical reliability while not significantly affecting overall business effectiveness, even if the test channel performs poorly. Traffic splitting is achieved through user ID hash modulo method, with the specific formula: user ID hash value % 100<5. This consistent allocation based on user ID ensures that all messages for the same user are routed to the same group (experimental group or control group), avoiding inconsistencies in user experience.
[0147] Allocate test traffic to the channel to be tested to establish an experimental group, while keeping the remaining traffic using the existing optimal channel as a control group. To ensure the fairness of the test, the system will control multiple key variables, including: time distribution (ensure that the experimental group and the control group have consistent traffic proportions at each time), message type distribution (ensure that the two groups have the same proportion of marketing messages, notification messages, etc.), and user attribute distribution (ensure that the two groups have similar distribution in user region, device type, etc.). This multi-dimensional balanced allocation greatly reduces potential confounding factors and improves the reliability of test results.
[0148] During the test process, conversion data is collected for both the experimental and control groups, including complete funnel data such as send volume, reach volume, click volume, conversion volume, etc. Real-time stream processing architecture is used for data collection to ensure that test results can be reflected in decision-making in a timely manner. The system not only records absolute values, but also calculates conversion rates and relative change rates at each link to comprehensively evaluate channel performance.
[0149] One of the innovations of this system is the use of Bayesian statistical methods based on Beta-Binomial conjugate prior distribution for test evaluation, which is more suitable for marketing scenarios than traditional frequency hypothesis testing. The core idea of the Beta-Binomial model is to consider conversion rate as a probability distribution rather than a single point estimate, which can better express the uncertainty of the estimate and the integration of prior knowledge. Specifically, the system uses Beta(α, β) distribution as the prior distribution of conversion rate, where the α and β parameters are set based on historical data or domain knowledge. For example, for general SMS marketing, Beta(3, 97) can be used as the prior, reflecting an expected conversion rate of about 3%.
[0150] After collecting experimental data, the posterior distribution is calculated using the Bayesian updating rule: if the experimental group observes k conversions with a total sample size of n, the posterior distribution is Beta(α+k, β+n-k). Similarly, the posterior distribution of the control group is also calculated in a similar manner. The system further calculates the probability that the conversion rate of the experimental group is higher than that of the control group (referred to as the Bayesian factor), which directly quantifies the confidence that the new channel is better than the existing channel.
[0151] One of the significant advantages of Bayesian statistical testing is the ability to directly interpret the business significance of the results, such as "the probability that the new channel is better than the existing channel is 95%, with an expected conversion rate increase of 5.2%", which is more intuitive and more instructive to decision-making than the traditional p-value. The system sets a double decision standard: a statistical significance standard (the probability that the new channel is better than the existing channel is more than 95%) and a business significance standard (the relative increase is more than 5%), only when both conditions are met, the new channel will be considered for inclusion in the formal resource pool.
[0152] For new channels that pass the test, they will not be allocated a large amount of traffic recklessly, but will adopt a gradual weight allocation strategy. The new channel first obtains a conservative initial weight (usually 5-10%), and then adjusts it gradually based on its continuous performance. The system will monitor the performance stability of the new channel, including the consistency of performance in different time periods and different user groups, to ensure that its advantages are universally applicable rather than accidental results in specific scenarios.
[0153] The A / B test framework also includes automated test management functions, allowing for parallel testing of multiple channels, intelligent allocation of test resources, and prioritization of the most promising channels. The system dynamically adjusts test duration based on preliminary results, ending testing early for channels that are clearly superior or inferior, and concentrating resources on borderline cases that require more data.
[0154] Through this scientifically sound A / B test framework, high-performance channels can be continuously discovered and verified, and channel resource pools can be continuously optimized, achieving long-term stable improvement of SMS marketing effectiveness. This framework not only meets the requirements of business agility, allowing for quick verification of new channel effectiveness, but also ensures the scientific nature of decision-making, avoiding the risks that may arise from subjective judgment and experience-based decision-making.
[0155] 2. Abnormality detection and emergency handling mechanism: Based on historical 30-day data, the mean μ and standard deviation σ of each performance indicator are calculated, and a 3-sigma abnormality detection algorithm is used to establish an abnormality judgment benchmark. Real-time monitoring of performance indicator data for each channel, with the real-time indicators of a channel exceeding the range of μ ± 3σ being determined as performance abnormalities and triggering an alarm mechanism to identify abnormal channels. Upon detection of the abnormal channel, the weight of the abnormal channel is immediately reduced to 0%, and is proportionally allocated to other channels, while an artificial intervention notification mechanism is initiated to ensure overall system availability.
[0156] Specifically, SMS channel performance is not always stable and can be affected by network fluctuations, operator policy adjustments, holiday traffic surges, and other factors. To ensure stable operation of the system in the face of various abnormal situations, the system has designed a comprehensive abnormality detection and emergency handling mechanism, achieving a complete closed loop from abnormal early warning to automatic recovery.
[0157] The abnormality detection mechanism is based on in-depth analysis of historical data to establish a normal behavior model for each performance indicator. The system uses the last 30 days of complete data as the basis to calculate the mean μ and standard deviation σ of each performance indicator (including reach rate, conversion rate, complaint rate, and delay indicators), and to build a dynamically updated abnormality judgment benchmark. The 30-day time window is the best choice after repeated verification, as it can include sufficient data samples to ensure statistical reliability while adapting to seasonal changes and long-term trends in business.
[0158] The 3-sigma anomaly detection algorithm is used as the basic decision mechanism, which is a classic method based on the assumption of normal distribution, suitable for the fluctuation characteristics of most performance indicators. According to this algorithm, when the real-time value of a certain indicator exceeds μ ± 3σ, it is judged as abnormal. This standard means that under the assumption of normal distribution, the theoretical probability of abnormal events is only 0.27%, effectively balancing sensitivity and specificity.
[0159] To adapt to the characteristics of different indicators, the system has made several enhancements to the basic algorithm: For non-normal distribution indicators (such as complaint rate, which usually presents a right-skewed distribution), the system applies Box-Cox transformation or logarithmic transformation to convert the data to approximately normal distribution before applying the 3-sigma rule.
[0160] For indicators with obvious time patterns (such as conversion rate differences between weekdays and weekends), the system establishes a time-based benchmark model to ensure that the anomaly detection takes into account the time factor.
[0161] For indicators of different importance, different detection sensitivities are set. Key indicators (such as reach rate) use a stricter 2.5-sigma standard, while secondary indicators may use a more relaxed 3.5-sigma standard.
[0162] In addition to static threshold detection, a variety of advanced anomaly detection algorithms are implemented: Rate of change detection: Monitor the first derivative (rate of change) and second derivative (rate of acceleration) of the indicator to capture sudden and sharp changes, even if the absolute value is still within the normal range.
[0163] Seasonal adjustment: Use time series decomposition techniques to decompose the data into trend, seasonality, and residual components, and apply anomaly detection to the residual components to exclude the influence of normal seasonal fluctuations.
[0164] Multi-dimensional correlation detection: Monitor the correlation patterns between multiple indicators to find abnormal patterns that violate historical correlations, such as a decrease in reach rate usually accompanied by an increase in delay, but if only the reach rate decreases while the delay is normal, it may indicate a different type of problem.
[0165] Cluster anomaly detection: Perform multi-dimensional cluster analysis on channel performance to identify abnormal points that deviate from normal clusters and discover complex multi-dimensional abnormal patterns.
[0166] Real-time monitoring of performance indicator data for each channel, the anomaly detection engine executes at a high frequency of once per minute to ensure timely detection of problems. When performance anomalies are detected, the system triggers a hierarchical alarm mechanism: Level 1 alarm (observation level): A single indicator slightly exceeds the threshold for a short period of time, the system records the event and increases the monitoring frequency, but does not take intervention measures.
[0167] Secondary alert (warning level): Single indicator consistently exceeds threshold or multiple indicators show minor anomalies, system sends warning notification to operations team and prepares emergency resources.
[0168] Tertiary alert (intervention level): Critical indicators show severe anomalies or multiple indicators show significant anomalies, system automatically initiates emergency handling process and sends high-priority notification to technical lead.
[0169] When a channel is determined to have a severe anomaly (tertiary alert), an emergency handling process is immediately triggered. The primary measure is to quickly reduce the weight of the abnormal channel to 0%, completely isolating the problem channel to prevent further impact on business. At the same time, the system reassigns the traffic originally allocated to the abnormal channel to other healthy channels according to the pre-set emergency distribution strategy. The redistribution uses a weighted algorithm based on health and capacity to ensure that no channel suddenly bears excessive load.
[0170] For example, assume channel A (weight 30%) is detected as abnormal and its traffic needs to be redistributed to channel B (current weight 40%) and channel C (current weight 30%). Simple proportional distribution would result in B obtaining 40 / (40+30)×30%=17.1% increment and C obtaining 30 / (40+30)×30%=12.9% increment. However, if B's capacity utilization rate has reached 80% while C's is only 50%, the system will consider the capacity factor and may allocate more traffic (e.g. 20%) to C and only 10% to B, avoiding B approaching the capacity limit.
[0171] Emergency handling is not limited to traffic redistribution, the system also initiates a series of safeguard measures: Automatic retry mechanism: For messages that have entered the abnormal channel queue but have not been sent, the system will automatically re-route them to healthy channels, ensuring that messages are not lost.
[0172] Message priority adjustment: In resource-constrained situations, the system will increase the priority of critical business messages (such as verification codes, order notifications) to ensure that important business is not affected.
[0173] Flow control protection: In extreme cases, if the capacity of healthy channels cannot accommodate all traffic, the system will initiate intelligent flow control, temporarily suspending the sending of marketing messages with lower priority according to pre-set business priority.
[0174] Channel recovery detection: The system will periodically send a small number of probe messages (usually no more than 1% of total traffic) to abnormal channels to monitor their recovery, and once performance is detected to return to normal levels, a gradual weight recovery process will be initiated.
[0175] The manual intervention notification mechanism is an important supplement to emergency handling, ensuring that there is supervision and intervention by human experts in addition to automated measures. The system sends detailed abnormality reports through various channels (email, SMS, instant messaging, phone), including abnormal indicators, abnormality degree, possible cause analysis, and measures taken by the system. The report is designed in layers, with both summary information that management can understand and detailed diagnostic data that technical personnel need.
[0176] A complete life cycle management of abnormal events is also established, recording the entire process from discovery to resolution, including detection time, alarm level, impact range, handling measures, recovery time, and root cause analysis. These records are not only used for post-audit, but also as training data for machine learning models, continuously improving the accuracy of abnormality detection and the efficiency of handling.
[0177] Through this complete abnormality detection and emergency handling mechanism, the system can maintain high flexibility and resilience in the face of various abnormal situations, minimizing the impact of abnormal events on business, and ensuring the overall availability and stability of the SMS marketing system. Even in the most adverse conditions, the system can quickly adjust resource allocation, prioritize critical business, and achieve smooth transition of business continuity.
[0178] 3. Channel performance deep mining based on distributed asynchronous exploration: The current channel configuration state of the optimal distribution ratio is taken as the exploration starting point, and parameters including channel weight, sending period, and content template are taken as exploration dimensions to build a multi-dimensional decision tree structure containing n channel nodes and depth D, generating a channel configuration exploration tree; Based on the channel configuration exploration tree, k configuration variation branches are generated in parallel, and each branch is assigned to a distributed computing node for asynchronous execution. Small-scale A / B testing is used to verify the conversion effect of each configuration scheme, and key indicator data including conversion rate and ROI are collected; According to the key indicator data, the regret value of each configuration scheme is calculated, and inefficient branches are pruned based on the regret minimization principle. The optimal branch is retained for further exploration. Within a maximum of 2n+O(k²2 kD ) moves, the global optimal configuration is converged.
[0179] Specifically, traditional channel weight optimization is usually limited to fine-tuning of existing configurations, making it difficult to find potential performance breakthroughs and innovative combinations. To overcome this limitation, the system designs a channel performance deep mining mechanism based on distributed asynchronous exploration, combining heuristic search and parallel computing to explore the optimal configuration in a wider parameter space and tap the potential upper limit of channel performance.
[0180] The mechanism first acquires the current channel configuration state as the starting point of exploration, and expands the exploration dimension to multiple key parameters, far beyond the traditional single weight adjustment: Channel weight: the proportion of traffic allocation for each channel, which is the most basic optimization dimension.
[0181] Sending period: the sending priority and weight adjustment of different time periods, considering the time difference of user activity and channel performance.
[0182] Content template: the format, length, personalization degree and presentation of messages, which have a significant impact on conversion rate.
[0183] User segmentation: targeted sending strategy based on user characteristics such as activity, purchasing power and region.
[0184] Sending frequency: message frequency control strategy for different user groups.
[0185] Channel combination: multi-channel coordination strategy in specific business scenarios.
[0186] A multi-dimensional decision tree structure containing n channel nodes and depth D is constructed, called channel configuration exploration tree. Each node represents a complete set of configuration parameters, and the connection between nodes represents the parameter adjustment path. The root node of the tree is the current configuration, each layer represents an optimization dimension, and the maximum depth D of the tree is usually set to 6-8, ensuring the breadth and depth of exploration.
[0187] Based on this exploration tree, k configuration variation branches (usually k=5-10) are generated in parallel, each representing a possible optimization direction. The variation generation adopts a combination of multiple strategies: Gradient-guided variation: calculate the gradient of the parameter on the objective function based on historical data, and generate variation along the gradient direction.
[0188] Random perturbation variation: add random perturbation based on the current configuration to explore the local space.
[0189] Cross variation: combine the parameters of multiple historical good configurations to obtain potential excellent characteristics.
[0190] Genetic algorithm variation: borrow the variation operation of genetic algorithm and introduce moderate randomness.
[0191] Expert rule variation: generate targeted variation based on pre-set rules based on domain knowledge.
[0192] These branches are assigned to different nodes of a distributed computing cluster for asynchronous execution and evaluation. The system employs containerization technology and job queue management to ensure efficient utilization of computing resources and reliable execution of tasks. Each computing node independently executes the assigned configuration scheme evaluation without waiting for other nodes, significantly improving exploration efficiency.
[0193] To evaluate the effectiveness of each configuration scheme, a small-scale A / B testing method is used, allocating 1-2% of traffic to each scheme for experimentation. This small-scale testing provides statistically significant results while minimizing trial risks. Key indicator data including conversion rates, ROI, and user feedback are collected to comprehensively evaluate the effectiveness of the configuration scheme.
[0194] During the exploration process, a multi-armed bandit strategy combining Thompson sampling and Upper Confidence Bound (UCB) algorithm is used to dynamically adjust resource allocation among branches. Branches with better performance receive more testing traffic and computing resources, while underperforming branches are eliminated early, achieving optimal resource allocation.
[0195] For each configuration scheme, its regret value (the difference from the currently known optimal scheme) is calculated, and branch pruning is performed based on the regret minimization principle. This method ensures the efficiency of exploration and avoids excessive investment in obviously inferior solutions. In theory, the system can converge to the global optimal configuration within at most 2n+O(k 2 2 KD ) moves, where n is the number of channels, k is the number of parallel branches, and D is the exploration depth.
[0196] The exploration process is not one-time, but a normal mechanism that runs continuously. The system will regularly (such as every week) start a new round of deep exploration to adapt to changes in business environment and user behavior. Each round of exploration retains the excellent results of the previous round as a starting point, achieving knowledge accumulation and inheritance.
[0197] To balance the relationship between exploration and utilization, the system uses an ε-decay strategy: initial exploration accounts for a larger proportion (such as 20% of traffic for exploration), and as time passes and performance improves, the exploration proportion gradually decreases (possibly down to 5%), allocating more resources to utilizing known excellent configurations.
[0198] A key innovation of this mechanism is the asynchronous parallel experimental evaluation system, breaking through the efficiency bottleneck of traditional sequential experiments. In traditional methods, only one configuration change can be tested at a time, and sufficient data must be accumulated before making a judgment; while this system can evaluate multiple configuration schemes simultaneously, significantly accelerating the optimization iteration speed.
[0199] The automatic analysis and knowledge extraction function of the experimental results is also implemented, and the universal optimization rules and patterns are summarized from the successful configuration schemes, such as "the conversion rate of personalized content sent to young users at 22:00-2:00 is significantly higher than the average level". These rules are added to the knowledge base to guide future optimization direction.
[0200] Through this deep mining mechanism based on distributed asynchronous exploration, the system can break through the limitation of local optimization and find innovative configuration combinations in a wider parameter space to mine the potential upper limit of channel performance. Practice has proved that this method can bring an additional 10-20% performance improvement compared to traditional single-dimensional optimization, especially in complex and variable business environments.
[0201] The method comprises the following steps: The current channel configuration state of the optimal distribution ratio is set as the root node, which contains the complete parameter set of weight distribution, sending strategy and content template of all channels, and an exploration starting state is established; Based on the exploration starting state, fine-tuning changes are made to each configuration parameter to generate a plurality of sub-node branches, wherein each sub-node represents a specific variation of a certain parameter, forming a configuration variation space; The configuration variation space is expanded to a depth D according to a depth-first search strategy, and each leaf node corresponds to a specific configuration scheme, thereby generating the channel configuration exploration tree.
[0202] Specifically, the current channel configuration state of the optimal distribution ratio is set as the root node, which contains the complete parameter set of weight distribution, sending strategy and content template of all channels, and an exploration starting state is established. This step first takes the best configuration of the current system as the starting point of exploration, ensuring that the exploration process is based on an effective basis. The root node contains complete configuration information, not limited to channel weight distribution, but also including sending priority strategies for each period, targeting rules for each user group, content template parameters in different scenarios, etc. The system will comprehensively record these parameters, including their value range, sensitivity and mutual dependence, to provide constraint conditions and guidance direction for subsequent variation exploration. Based on the exploration starting state, fine-tune changes are made to each configuration parameter to generate multiple sub-node branches, where each sub-node represents a specific variation of a certain parameter, forming a configuration variation space. In this phase, the system generates direct child nodes of the root node using various variation strategies. For continuous parameters (such as channel weights), the system generates three variations: up, down, and no change; for discrete parameters (such as content template types), it attempts all possible alternative values. The variation amplitude is dynamically adjusted according to parameter sensitivity, with important parameters having smaller variation steps and less important parameters having larger adjustment spaces. The system also considers the coupling relationship between parameters to ensure that the modified configuration still meets business constraints, such as total weight being 100%, minimum guarantee for key channels, etc. Through this targeted parameter fine-tuning, the system can fully explore potential optimization points in the parameter space while maintaining reasonable configurations. The configuration variation space is expanded to a depth of D according to the depth-first search strategy, with each leaf node corresponding to a specific configuration scheme, generating the channel configuration exploration tree. The system uses a depth-first search (DFS) strategy to expand the exploration tree layer by layer until the maximum preset depth D is reached. During the expansion process, the system applies similar variation rules to each intermediate node as the root node to generate child nodes in the next layer. As the depth of the tree increases, the cumulative variation effect produces schemes that differ greatly from the initial configuration, allowing exploration to cover a broader configuration space. To control the size and computational complexity of the tree, the system implements an intelligent pruning mechanism to terminate expansion of obviously suboptimal branches. Pruning criteria include historical performance of similar configurations, domain knowledge rules, and preliminary evaluation results. The final exploration tree typically contains hundreds to thousands of leaf nodes, each representing a complete candidate configuration scheme, waiting for further evaluation and verification by the system. This structured exploration method ensures both the breadth and depth of the search, and improves exploration efficiency through heuristic rules, enabling the discovery of innovative optimization configurations within reasonable computational resource constraints.
[0203] 4. Channel stability prediction based on two-dimensional threshold cellular automata: Obtain attribute information of the operator type and geographic area including all short message channels, map each channel to a two-dimensional grid structure, each grid cell contains a state vector of performance score, load rate, failure probability, and neighborhood influence degree, and establish a channel network topology model; Based on the channel network topology model, design threshold conversion rules: when the channel performance benchmark score is less than 60 and the continuous low score time exceeds 1 hour, it is determined as a warning state; when the number of failed channels in the 8-neighborhood is greater than or equal to 3 and the current load rate exceeds 80%, the failure probability is increased, and the cellular automaton state is updated; According to the cellular automaton state update, iterate evolution 100 time steps, simulate the propagation and diffusion process of channel failure, calculate the overall stability index of the system, and predict the stability trend of the channel network in the next 2-4 hours.
[0204] Specifically, the stability of the SMS channel is not only related to the performance of a single channel, but also involves the systemic risk of the entire channel network. There is a complex interdependence and influence relationship between channels. The failure of a channel can propagate and affect other channels through various mechanisms, ultimately leading to systemic risk. Traditional independent monitoring methods are difficult to predict this network effect, so the system innovatively introduces a channel stability prediction mechanism based on a two-dimensional threshold cellular automaton, which can simulate the dynamic process of failure propagation and provide early warning of system risk.
[0205] Cellular automata (Cellular Automata) is a discrete dynamic system model composed of simple rules but complex behavior cells (cells), widely used in complex system simulation. The system applies it to channel stability prediction and builds a two-dimensional network model reflecting the relationship between channels.
[0206] First, obtain the attribute information of all SMS channels, including operator type (such as mobile, China Unicom, China Telecom), geographic region (such as East China, South China, North China), service provider, capacity level, etc. These attributes are the basis for building the channel network topology, determining the "proximity" and potential influence between channels.
[0207] Based on these attributes, the system maps each channel to a two-dimensional grid structure to form a channel network topology model. The mapping process uses multidimensional scaling (MDS) algorithm to convert the similarity between channels into distance in two-dimensional space, ensuring that channels with similar attributes are close in the grid. For example, channels from the same operator and the same region will be mapped to adjacent positions, reflecting their potential high correlation.
[0208] In this two-dimensional grid structure, each grid cell (representing a channel) contains a state vector describing the key characteristics of its current state: Performance score: Reflects the current overall performance level of the channel, calculated by the aforementioned scoring model.
[0209] Load rate: The proportion of the current channel load relative to its maximum capacity.
[0210] Failure probability: Estimated probability of channel failure, calculated based on historical failure data and current state.
[0211] Neighborhood influence: The potential influence of channel failure on surrounding channels, related to channel importance and connectivity.
[0212] Threshold-based state transition rules are designed for the cellular automaton to simulate the dynamic evolution of channel states over time and environmental changes. The core transition rules include: Warning state determination: When the channel performance benchmark score is below 60 points and the consecutive low score time exceeds 1 hour, the channel enters a warning state, and the failure probability begins to rise.
[0213] Failure propagation rule: When the number of failed channels in the 8-neighborhood (Moore neighborhood) is greater than or equal to 3 and the current load rate exceeds 80%, the failure probability of the central channel significantly increases. This reflects the cascading failure effect under high load conditions.
[0214] Recovery rule: When the neighborhood failure pressure of the channel decreases and the system takes intervention measures (such as reducing load), the failure probability begins to decrease, and the channel has the opportunity to recover to a normal state.
[0215] Threshold adjustment rule: Over time, the state threshold of the channel is dynamically adjusted according to the overall system condition, simulating the adaptive behavior of the system.
[0216] With the state vector and transition rules, the system performs the state update process of the cellular automaton. At each time step, the states of all channels are updated simultaneously according to the transition rules and the current neighborhood conditions. This parallel update mechanism can capture the emergent behavior and nonlinear dynamics in complex systems.
[0217] Usually, iterate evolution for 100 time steps (each step represents about 2-3 minutes in the actual system), simulate the propagation and diffusion process of channel failures. During the evolution process, the system records the change trajectory of key indicators, including: Failure channel proportion: The percentage of channels in a failure state in the total number of channels.
[0218] System connectivity: The graph theory connectivity index of the network, reflecting the overall health of the channel network.
[0219] Critical channel state: The state change of the core channel that is particularly important to the system.
[0220] Clustering coefficient: Reflects the degree of aggregation of channel failures, high clustering means that failures are concentrated in a specific area.
[0221] Based on these evolution trajectories, the system calculates a series of stability indicators to assess the overall stability trend of the channel network in the next 2-4 hours: System vulnerability index: Overall risk measure based on the distribution of channel failure probabilities and network structure.
[0222] Cascade failure risk: Estimated probability of large-scale cascading failure in the system.
[0223] Recovery Elasticity Index: Estimation of the system's ability to recover from local failures.
[0224] Stable Interval Prediction: Expected duration of the system's maintenance in a stable state.
[0225] These indicators not only provide quantitative predictions of future stability but also identify vulnerable points and potential risk sources within the system. For example, a possible early warning could be: "In the next 3 hours, there is a 20% risk of cascading failures in the telecommunications channels in the East China region. It is recommended to reduce the load in advance and activate backup channels."
[0226] The prediction results are directly fed back to the weight optimization module, affecting subsequent weight allocation decisions. When the system predicts an increase in instability in a certain region or type of channel, it actively adjusts the weights of related channels, diverting traffic to more stable regions and proactively reducing risks.
[0227] To improve prediction accuracy, an adaptive learning mechanism is used, constantly adjusting the parameters and rules of the cellular automaton based on actual observed failure propagation patterns. This includes calibration of state transition thresholds, adjustment of neighborhood influence weights, and introduction of new transition rules. Through this closed-loop learning, the system can gradually improve the simulation accuracy of the dynamic characteristics of specific channel networks.
[0228] Another innovative point of this prediction mechanism is the integration with external data sources, which automatically acquire and integrate various external factors that may affect channel stability, such as: Major event calendar: Holidays, promotional activities, and other events that may cause traffic surges.
[0229] Operator maintenance plan: Announced network maintenance or upgrade activities.
[0230] Weather anomaly data: Warning of severe weather that may affect communication infrastructure.
[0231] Internet health status: Global or regional network congestion or anomaly indicators.
[0232] These external data serve as adjustment factors for the cellular automaton, affecting the probability and speed of state transitions, making the prediction model more realistic.
[0233] Through this two-dimensional threshold cellular automaton-based channel stability prediction mechanism, it can start from the micro-channel state, simulate and predict the macro-network stability changes, and realize the transformation from passive response to active prevention. This ability is of great value to ensure the stable operation of the system in a complex and changing environment, especially during peak periods and special periods. Predictive risk management can effectively prevent potential systemic failures and ensure business continuity.
[0234] The executing the cellular automaton state update comprises: Obtaining a state vector of each channel at a current time, judging the state according to a performance threshold rule, a neighborhood propagation rule and a cascade failure rule, and calculating a state probability distribution of each channel at a next time; Parallelly updating state values of all channels based on the state probability distribution, synchronously calculating neighborhood influence degrees and failure propagation coefficients, simulating a failure diffusion process in the channel network, and obtaining the failure diffusion process; Quantitatively evaluating the failure diffusion process, calculating stability indexes including a health channel proportion and a maximum connected component size, and forming a system overall stability score.
[0235] Specifically, the state vector of each channel at the current time is obtained, and state judgment is performed according to the performance threshold rule, neighborhood propagation rule and cascade failure rule, and the state probability distribution of each channel at the next time is calculated. The system first extracts the latest state vector of each channel from the monitoring database, including performance score, load rate, failure probability and neighborhood influence degree of four key dimensions. For each channel, three core rules are applied to determine the state evolution: the performance threshold rule adjusts the failure probability according to the comparison result of the current performance score of the channel and the preset threshold (usually 60 points), the longer the time of continuous low threshold, the faster the failure probability rises; the neighborhood propagation rule investigates the state of the 8 adjacent channels (molecular neighborhood) around the channel, when the number of failed channels in the neighborhood reaches a certain threshold (usually 3) and the current channel load rate exceeds 80%, the failure probability of the channel is significantly increased; the cascade failure rule simulates the chain reaction under high load conditions, when a channel enters the failure state, the traffic it carries will be redistributed to other channels, increasing the load rate and failure risk of these channels. The system integrates the influence of the three rules to calculate the probability distribution of each channel in the next time, which may be in a healthy, warning, mild failure or serious failure state. Based on the state probability distribution, the state values of all channels are updated in parallel, the neighborhood influence degree and failure propagation coefficient are calculated synchronously, the diffusion process of the fault in the channel network is simulated, and the fault diffusion process is obtained. The system uses the Monte Carlo method to determine the specific state of all channels at the next time according to the calculated state probability distribution. This parallel updating mechanism can accurately simulate the complex dynamics of multiple channels evolving simultaneously in the real world. After the state update, the system recalculates the neighborhood influence degree of each channel, which reflects the potential impact strength of the channel failure on the surrounding channels, and is positively correlated with the importance (load size) of the channel, the connection degree (the degree of association with other channels) and the current state (failure degree). At the same time, the system also calculates the failure propagation coefficient, which measures the diffusion speed and range of the failure in the network, which is affected by factors such as network topology, channel similarity and load distribution. Through continuous iteration and update of multiple time steps, the system generates a complete fault diffusion spatiotemporal evolution sequence, which intuitively shows how the fault propagates, diffuses, aggregates or subsides in the channel network, forming a comprehensive simulation of the dynamic stability of the system. The fault diffusion process is quantitatively evaluated, and stability indicators including the proportion of healthy channels and the size of the largest connected component are calculated to form the overall stability score of the system. The fault diffusion process obtained by simulation is analyzed quantitatively from multiple dimensions, and key stability indicators are extracted. The proportion of healthy channels reflects the overall health status of the system, and the calculation method is the number of channels in the healthy state divided by the total number of channels, and the trend of this indicator over time can reveal whether the system is developing towards stability or instability.The maximum connected component size is based on graph theory analysis, treating the channel network as a graph structure, with healthy channels as nodes and connections between channels as edges. The size of the largest connected subgraph in the network is calculated, reflecting the functional integrity and service continuity of the system. In addition, the system also calculates auxiliary indicators such as fault clustering coefficient (measuring the concentration of faults), key channel stability (the degree of state stability of core channels), and recovery elasticity (the ability of the system to recover from local faults). Finally, the system generates an overall stability score of 0-100 by weighting and fusing these indicators, and sets warning thresholds based on historical data and expert experience. When the stability score falls below a certain threshold (usually 65), the corresponding level of warning mechanism is triggered, guiding the operation team to take preventive measures in advance to prevent potential systemic risks.
[0236] 5. Channel routing optimization based on fast approximate minimum spanning tree: Extract user characteristics such as user group's geographic location, device type, historical behavior preference, active period, and channel characteristics such as operator type, coverage area, delay characteristics, and cost structure, construct a high-dimensional feature space, and generate user-channel combination feature vectors; Based on the user-channel combination feature vectors, use the FAMST three-stage algorithm to construct the optimal connection relationship of channel routing, and through the three stages of approximate nearest neighbor graph construction, neural network component connection discovery, and iterative edge refinement, establish a personalized routing network topology; According to the personalized routing network topology, calculate the shortest distance and matching degree of each sending task to each channel, select the optimal channel combined with the optimal distribution ratio, and realize user-level precise channel routing optimization.
[0237] Specifically, traditional SMS channel selection usually uses weight-based random allocation or simple polling, ignoring the matching degree between user characteristics and channel characteristics, resulting in poor resource allocation efficiency. This system innovatively introduces a channel routing optimization mechanism based on fast approximate minimum spanning tree (FAMST), achieving user-level precise channel matching and significantly improving the relevance and effectiveness of SMS sending.
[0238] This mechanism first constructs a high-dimensional feature space to comprehensively capture multi-dimensional features of users and channels. User characteristics include: Geographic location: the province, city, and region where the user is located, which directly affects the network connection quality with different operator channels.
[0239] Device type: the brand, model, and operating system version of the user's mobile phone, which affects the display effect and interactive experience of messages.
[0240] Historical Behavior Preferences: User's historical open rates, click rates, conversion rates, and response patterns to different types of messages.
[0241] Active Periods: User's typical active time periods, such as morning, afternoon, evening, or late night, reflecting user's usage habits.
[0242] User Value: User's comprehensive value score calculated based on dimensions such as consumption ability, activity level, loyalty, etc.
[0243] Sensitivity: User's sensitivity to message frequency, past unsubscribe or complaint behavior, etc.
[0244] Channel features include: Operator Type: The telecom operator (such as Mobile, Unicom, Telecom) or third-party service provider that the channel belongs to.
[0245] Coverage Area: The geographical coverage of the channel and its performance in different regions.
[0246] Latency Characteristics: The latency performance of the channel under different time periods and different loads.
[0247] Content Compatibility: The channel's support capabilities for different message formats, lengths, and rich media content.
[0248] Period Performance: The performance fluctuation pattern of the channel in different time periods of the day.
[0249] Cost Structure: The charging method of the channel and the cost changes under different conditions.
[0250] Based on these features, the system generates user-channel combination feature vectors for evaluating the potential matching degree of each user-channel combination. These feature vectors contain cross combinations, weight adjustments, and nonlinear transformations of the original features, forming a high-dimensional matching space.
[0251] In this feature space, the system uses the FAMST (Fast Approximate Minimum Spanning Tree) three-stage algorithm to construct the optimal connection relationship of the channel routing. The FAMST algorithm is an efficient approximation version of the traditional minimum spanning tree algorithm, especially suitable for handling large-scale data set connection optimization problems. Its three key stages include: Approximate Nearest Neighbor Graph Construction: First, use the Local Sensitivity Hashing (LSH) algorithm to quickly construct the approximate nearest neighbor graph in the feature space. LSH realizes O(n log n) time complexity of approximate nearest neighbor search by mapping high-dimensional data to low-dimensional signatures, greatly improving processing speed. In this step, find the most similar k channels (usually k=3-5) in the feature space for each user, forming a preliminary candidate set.
[0252] Neural network component connection discovery: Evaluate the quality scores of candidate user-channel connections using a trained deep neural network model. This model takes a user-channel combination feature vector as input and outputs a match score between 0 and 1, reflecting the expected effectiveness of the connection. The neural network model is trained on historical sending data, capturing complex nonlinear relationships and hidden patterns. Based on these scores, the best connections are retained, forming a directed weighted graph.
[0253] Iterative edge refinement: Continuously refine and optimize connection relationships through an iterative process. In each iteration, the system removes the edges with the lowest weights and attempts to add new potential connections, evaluating whether they can improve overall effectiveness. This process is similar to the Kruskal algorithm for constructing a minimum spanning tree, but uses heuristic rules and early stopping strategies to significantly improve efficiency. Iteration continues until convergence conditions are met or the maximum number of iterations is reached.
[0254] Through these three stages, an optimized personalized routing network topology is finally constructed, determining the best channel selection strategy for each user and each message type. This topology considers both local user-channel matching degrees and global load balancing and cost efficiency.
[0255] In actual channel routing decisions, the system considers three key factors: Matching degree: The similarity score of the user-channel combination in the feature space, reflecting the degree of adaptation of the channel to the specific user.
[0256] Shortest distance: The path distance from the sending task to each channel in the routing network topology, reflecting the efficiency of routing.
[0257] Weight constraint: Needs to meet the global optimal distribution ratio calculated in step S5, ensuring that the overall traffic distribution conforms to the optimization results.
[0258] A weighted multi-objective optimization method is used to balance these three factors to select the optimal channel for each sending task. The specific decision function is: Score(user, channel) = w1 × matching degree + w2 × (1 / shortest distance) + w3 × weight satisfaction degree; Where the weight satisfaction degree measures whether selecting this channel helps achieve the target weight distribution. For example, if the actual usage rate of a channel is lower than its target weight, the weight satisfaction degree of selecting that channel is higher.
[0259] To improve system efficiency, the routing decision adopts a multi-level cache and batch processing strategy. The system caches the routing results of similar users and applies the same routing strategy to users with the same characteristics in batches, significantly reducing the computational overhead. At the same time, an incremental update mechanism is implemented, which only recalculates the routing strategy when the user characteristics or channel performance changes significantly.
[0260] Another innovation of this mechanism is the adaptive learning capability. Continuous collection of the effect data of each sending, including the reach state, user response and final conversion, constantly updates and refines the user-channel matching model. This closed-loop learning enables the routing strategy to adapt to changes in user behavior and fluctuations in channel performance, achieving continuous optimization.
[0261] Real-time response capability is also implemented, which can dynamically adjust routing decisions according to the current system load and channel state. For example, when detecting a sudden increase in real-time delay of a channel, the system temporarily reduces the matching score of that channel, diverting traffic to other stable performance channels until the abnormal situation is resolved.
[0262] Through this channel routing optimization mechanism based on fast approximation of minimum spanning tree, a leap from coarse-grained channel-level allocation to fine-grained user-level precise matching is achieved, significantly improving the targeting and effectiveness of SMS marketing. Practice has proved that compared with traditional random allocation, this precise matching mechanism can improve the conversion rate by 15-25%, while reducing the complaint rate by 10-15%, achieving double improvement of marketing effect and user experience.
[0263] Among them, the establishment of the personalized routing network topology includes: Based on the user-channel combination feature vector, an approximate nearest neighbor graph is constructed using Locality Sensitive Hashing, and the k=10 most similar user-channel combinations are calculated to establish an initial connection network; The initial connection network is input into a neural network model, the node features are extracted by the encoder, and the connection probability between node pairs is calculated using the connection predictor to discover potential high-value connection relationships; Iterative edge refinement is performed on the potential high-value connection relationships, removing the lowest 5% edges, checking network connectivity and adding key connections, and after a maximum of 10 iterations of optimization, the personalized routing network topology is generated.
[0264] Specifically, based on the user-channel combination feature vector, an approximate nearest neighbor graph is constructed using Locality Sensitive Hashing, and k=10 most similar user-channel combinations are calculated to establish an initial connection network. This step first maps the user-channel combinations in the high-dimensional feature space to low-dimensional hash signatures, and quickly identifies potential neighbor relationships by comparing the similarity of hash signatures. The system uses a multi-hash table strategy to generate multiple hash tables using different hash function families, improving the accuracy and recall rate of neighbor search. For each user-channel combination, the system retrieves the k=10 most similar combinations in the hash space to construct the initial connection relationship. This LSH-based approximate search reduces the computational complexity from the traditional O(n2) to O(n log n), significantly improving the large-scale data processing capability. The initial connection network forms a sparse graph structure, laying the foundation for subsequent fine optimization. The initial connection network is input into a neural network model, node features are extracted by an encoder, and connection probabilities between node pairs are calculated using a connection predictor to discover potential high-value connection relationships. In this stage, the system deploys a graph neural network model containing two core components: a feature encoder and a connection predictor. The feature encoder uses a multi-layer graph convolutional network (GCN) structure that can capture local topological information and feature representations of nodes, converting the original feature vector into an embedded representation containing structural information. The connection predictor calculates the connection probability based on the embedded representation of node pairs, using a multi-layer perceptron structure to output a connection strength score between 0 and 1. This model is supervised by historical sending data and can identify potential high-value connections missed in the initial connection network and filter out low-quality connections introduced by LSH. The system calculates the prediction score for all possible connections, retains connections with scores above a threshold (usually set to 0.7), and forms an enhanced connection network. The potential high-value connection relationships are iteratively refined, the lowest 5% edges are removed, the network connectivity is checked and key connections are added, and the personalized routing network topology is generated after a maximum of 10 iterations of optimization. Iterative edge refinement is a fine-tuning process, and the system performs three key operations in each iteration: first, the lowest 5% edges in the current network are removed, which usually represent inefficient or redundant connections; second, the network connectivity after removing the edges is checked to ensure that no node is isolated, and key connections are added as necessary to maintain the basic structure of the network; finally, the overall performance indicators of the network are re-evaluated, including average path length, clustering coefficient, and load balancing, etc. The optimization progress is monitored by the objective function, and the optimization is stopped when the performance improvement of two consecutive iterations is less than the preset threshold (0.5%) or the maximum number of iterations (10) is reached. The final personalized routing network topology not only retains high-value connections, but also has good structural characteristics, which can support efficient and accurate message routing decisions. The topology structure is updated regularly (usually every 24 hours) to adapt to the dynamic changes of user behavior and channel performance.
[0265] As shown in Figure 2 The embodiment of the application also provides a multi-channel dynamic weight based SMS marketing effect optimization system, which comprises: A data acquisition module 701 is configured to acquire real-time performance data of a plurality of SMS channels; A standardization processing module 702 is configured to perform standardization processing on the real-time performance data, and map the reach rate, the conversion rate, the complaint rate and the delay indicator to the [0, 1] interval by using a Z-score standardization method to obtain standardized performance indicators; A score calculation module 703 is configured to calculate a comprehensive performance score of each channel by weighted summation based on the standardized performance indicators, and generate a channel performance benchmark score by using a score formula Score = w1 x reach rate + w2 x conversion rate - w3 x complaint rate - w4 x delay indicator; A prediction module 704 is configured to collect multi-dimensional performance data of each channel in the past 24 hours, analyze the time sequence characteristics of historical conversion data by using an LSTM neural network, learn the time sequence pattern by using a three-layer LSTM network structure, and predict the conversion rate trend change of each channel in the next 2-4 hours; A weight optimization module 705 is configured to solve a weight distribution scheme with the goal of maximizing the overall ROI based on the conversion rate trend change and the channel performance benchmark score by using a gradient optimization method, set an objective function F(w) = Σ(wi x predicted conversion rate i x weight coefficient), and calculate the optimal distribution proportion of each channel; The shunt adjustment module 706 is configured to adjust the shunt strategy of the short message queue in real time according to the optimal distribution ratio, and re-allocate the to-be-sent short messages to the channel queues, while starting a smooth transition mechanism of the weight configuration, so as to realize dynamic optimization of the short message marketing effect.
[0266] The functions of the system modules in the embodiment correspond to the foregoing method steps, and the short message marketing effect optimization based on the multi-channel dynamic weight is realized through cooperative work. The system can automatically complete a series of operations such as data acquisition, standardized processing, score calculation, trend prediction, weight optimization and shunt adjustment, reduce the demand for manual intervention, and improve the overall effect of short message marketing.
[0267] In summary, the short message marketing effect optimization method and system based on the multi-channel dynamic weight provided by the present application realize intelligent allocation and effect optimization of short message marketing channel resources through technical means such as a real-time dynamic weight allocation mechanism, a multi-dimensional channel evaluation model, a machine learning time series prediction capability, an A / B test automation framework, and an abnormality detection and emergency handling mechanism. The method can overcome the limitations of traditional static configuration, improve the accuracy and conversion efficiency of short message marketing, and at the same time guarantee the stable operation ability of the system when facing channel failures.
[0268] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. The present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for optimizing SMS marketing effects based on multi-channel dynamic weights, characterized in that: include: Get real-time performance data for multiple SMS channels; Standardize the real-time performance data and use the Z-score standardization method to map the reach rate, conversion rate, complaint rate, and latency indicators to the [0, 1] interval to obtain standardized performance indicators; Based on the standardized performance indicators, the comprehensive performance score of each channel is calculated through weighted summation. The scoring formula Score = w1 × reach rate + w2 × conversion rate - w3 × complaint rate - w4 × latency index is used to generate a channel performance benchmark score. Collect multi-dimensional performance data from each channel over the past 24 hours, analyze the temporal characteristics of historical conversion data using an LSTM neural network, and use a three-layer LSTM network structure to learn temporal patterns and predict conversion rate trends for each channel over the next 2-4 hours. Based on the conversion rate trend change and the channel performance benchmark score, a gradient optimization method is used to solve the weight distribution scheme with the goal of maximizing the overall ROI, setting the objective function F(w)=Σ(wi×predicted conversion rate i×weight coefficient) to calculate the optimal distribution ratio of each channel; According to the optimal distribution ratio, the diversion strategy of the SMS queue is adjusted in real time, and the SMS messages to be sent are redistributed to the queues of each channel. At the same time, a smooth transition mechanism of the weight configuration is started to achieve dynamic optimization of the SMS marketing effect.
2. The method according to claim 1, characterized in that The real-time performance data is standardized, and the reach rate, conversion rate, complaint rate, and delay indicators are mapped to the [0, 1] interval using the Z-score standardization method to obtain standardized performance indicators, including: Collect raw data on reach, conversion, complaint, and latency metrics for each SMS channel, calculate the mean μ and standard deviation σ of each metric in historical data, and establish a data distribution baseline. Based on the data distribution baseline, the Z-score normalization formula Z = (X-μ) / σ is applied to normalize each indicator to eliminate the influence of the dimensional differences of different indicators and obtain the normalized transformation results; The normalized transformation result is interval mapped, and the Z value is mapped to the standard interval of [0, 1] using the Sigmoid function to obtain the standardized performance index.
3. The method according to claim 1, characterized in that Based on the standardized performance indicators, the comprehensive performance score of each channel is calculated by weighted summation, and the scoring formula Score = w1 × reach rate + w2 × conversion rate - w3 × complaint rate - w4 × latency index is used to generate a channel performance benchmark score, including: Obtain business target configuration parameters and set weight coefficients for each indicator based on the marketing scenario. The reach rate weight w1, conversion rate weight w2, complaint rate weight w3, and latency indicator weight w4 must satisfy w1+w2+w3+w4=1. Establish a weight configuration plan. Based on the weight configuration scheme and the standardized performance indicators, the real-time performance score of each channel is calculated by substituting the scoring formula Score = w1 × reach rate + w2 × conversion rate - w3 × complaint rate - w4 × latency indicator; According to the real-time performance score, a passing threshold of 60 points and an excellent threshold of 85 points for channel performance are set to generate the channel performance benchmark score.
4. The method according to claim 1, wherein The method collects multi-dimensional performance data of each channel in the past 24 hours, analyzes the time series characteristics of historical conversion data through an LSTM neural network, and uses a three-layer LSTM network structure to learn time series patterns to predict the conversion rate trend changes of each channel in the next 2-4 hours, including: Based on the multi-dimensional performance data, calculate the sliding average of the conversion rate over 3 hours, 6 hours, and 24 hours, construct trend features such as the first-order difference and second-order difference of the conversion rate, and form the input feature vector of the prediction model; Based on the input feature vector of the prediction model, a three-layer LSTM network structure is designed, in which the input layer receives the feature vectors of 24 time windows, and the hidden layer captures long-term and short-term dependencies through 128 LSTM units to establish a time series prediction model; The time series prediction model is deployed as an online inference service, and combined with the standardized performance indicators, the conversion rate trend changes of each channel in the next 2-4 hours are predicted.
5. The method according to claim 1, wherein The gradient optimization method is used to solve the weight distribution scheme with the goal of maximizing the overall ROI, setting the objective function F(w)=Σ(wi×predicted conversion rate i×weight coefficient) and calculating the optimal distribution ratio of each channel, including: Establish the objective function of weight optimization F(w)=Σ(wi×predicted conversion rate i×weight coefficient), set the constraints Σwi=1 and wi≥0, and construct the mathematical model of the optimization problem; Based on the mathematical model of the optimization problem, the Adam optimization algorithm is used to solve it, the learning rate is set to 0.01, the maximum number of iterations is 100, and the weight distribution vector that converges to the optimal solution is calculated; Obtain the weight distribution vector, compare and analyze it with the current weight configuration, and output the optimal distribution ratio.
6. The method according to claim 1, characterized in that The smooth transition mechanism of simultaneously starting weight configuration includes: Obtain the optimal distribution ratio, adopt an exponential decay smoothing algorithm, set the smoothing coefficient α to 0.3, and establish a weight smoothing transition formula; Based on the weight smoothing transition formula, the smoothed weight = α × updated weight + (1-α) × current weight is calculated to generate a progressive weight adjustment sequence; According to the progressive weight adjustment sequence, the weight switching operation is performed in steps to avoid the impact of rapid weight changes on system stability.
7. The method according to claim 1, characterized in that It also includes an A / B testing framework for validating the effectiveness of new channels: Randomly select no more than 5% of user traffic as test samples, split the traffic by taking the user ID hash modulus, and assign the test traffic to the channel to be tested to establish experimental and control groups; Collecting the conversion data of the experimental group and the control group, calculating the confidence interval of the conversion rate based on the Beta-Binomial conjugate prior distribution, performing a Bayesian statistical test, and obtaining a Bayesian statistical test result; Based on the Bayesian statistical test results, when the difference in conversion rate between the experimental group and the control group reaches statistical significance and the relative improvement exceeds 5%, the new channels that have passed the verification are automatically included in the weight allocation pool to obtain an updated channel weight allocation pool.
8. The method according to claim 1, characterized in that It also includes anomaly detection and emergency response mechanisms: Based on 30 days of historical data, the mean μ and standard deviation σ of each performance indicator are calculated, and the 3-sigma anomaly detection algorithm is used to establish an anomaly determination benchmark; Monitor the performance indicator data of each channel in real time. When the real-time indicator of a channel exceeds the range of μ±3σ, it is judged as a performance abnormality and an alarm mechanism is triggered to identify the abnormal channel; When the abnormal channel is detected, the weight of the abnormal channel is immediately reduced to 0% and allocated to other channels in proportion. At the same time, the manual intervention notification mechanism is activated to ensure the overall availability of the system.
9. The method according to claim 1, characterized in that It also includes in-depth exploration of channel performance based on distributed asynchronous exploration: Obtaining the current channel configuration state of the optimal distribution ratio as the exploration starting point, using parameters including channel weight, sending period, and content template as exploration dimensions, constructing a multidimensional decision tree structure including n channel nodes and a depth of D, and generating a channel configuration exploration tree; Based on the channel configuration exploration tree, k configuration variation branches are generated in parallel, each branch is assigned to a distributed computing node for asynchronous execution, and a small-scale A / B test is used to verify the conversion effect of each configuration solution, collecting key indicator data including conversion rate and ROI; According to the key indicator data, the regret value of each configuration scheme is calculated, and the inefficient branches are pruned based on the regret minimization principle, and the optimal branches are retained for further in-depth exploration. kD ) moves to converge to the global optimal configuration.
10. The method according to claim 9, characterized in that The generation channel configuration exploration tree includes: The current channel configuration state of the optimal distribution ratio is set as the root node, which includes the complete parameter set of weight distribution, sending strategy, and content template of all channels, and establishes the exploration starting state; Based on the exploration starting state, each configuration parameter is fine-tuned to generate multiple child node branches, where each child node represents a specific variation of a parameter, forming a configuration variation space; The configuration variation space is expanded to a depth D according to a depth-first search strategy, and each leaf node corresponds to a specific configuration scheme, thereby generating the channel configuration exploration tree.
11. The method according to claim 1, characterized in that Also included is channel stability prediction based on two-dimensional threshold cellular automata: Obtain attribute information including operator type and geographical area of all SMS channels, map each channel to a two-dimensional grid structure, and establish a channel network topology model by including a state vector of performance score, load rate, failure probability, and neighborhood influence for each grid cell. Based on the channel network topology model, a threshold conversion rule is designed. When the channel performance benchmark score is lower than 60 points and the continuous low score time exceeds 1 hour, it is determined to be in a warning state. When the number of faulty channels in the 8-neighborhood is greater than or equal to 3 and the current load rate exceeds 80%, the fault probability is increased and the cellular automation state update is executed; According to the cellular automaton state update, iterative evolution is performed for 100 time steps to simulate the propagation and diffusion process of channel faults, calculate the overall stability index of the system, and predict the stability trend of the channel network in the next 2-4 hours.
12. The method according to claim 11, characterized in that The execution of the cellular automation state update includes: Obtain the state vector of each channel at the current moment, make state judgments according to the performance threshold rule, neighborhood propagation rule, and cascading failure rule, and calculate the state probability distribution of each channel at the next moment; Based on the state probability distribution, the state values of all channels are updated in parallel, the neighborhood influence and fault propagation coefficient are calculated synchronously, and the fault diffusion process in the channel network is simulated to obtain the fault diffusion process; The fault diffusion process is quantitatively evaluated, and stability indicators including the proportion of healthy channels and the size of the largest connected component are calculated to form an overall stability score of the system.
13. The method according to claim 1, wherein Also included is channel routing optimization based on a fast approximate minimum spanning tree: Extract user characteristics such as the user group's geographic location, device type, historical behavioral preferences, and active time periods, as well as channel characteristics such as the channel's operator type, coverage area, latency characteristics, and cost structure, construct a high-dimensional feature space, and generate a user-channel combination feature vector. Based on the user-channel combination feature vector, the FAMST three-stage algorithm is used to construct the optimal connection relationship of the channel routing, and a personalized routing network topology is established through three stages: approximate nearest neighbor graph construction, neural network component connection discovery, and iterative edge refinement. According to the personalized routing network topology, the shortest distance and matching degree to each channel are calculated for each sending task, and the optimal channel is selected in combination with the optimal distribution ratio to achieve user-level precise channel routing optimization.
14. The method according to claim 13, wherein: The establishing of a personalized routing network topology includes: Based on the user-channel combination feature vector, an approximate nearest neighbor graph is constructed using Locality Sensitive Hashing, and the k=10 most similar user-channel combinations are calculated to establish an initial connection network. Inputting the initial connection network into a neural network model, extracting node features through an encoder, and using a connection predictor to calculate the connection probability between node pairs to discover potential high-value connection relationships; Iterative edge refinement is performed on the potential high-value connection relationships, the 5% edges with the lowest scores are removed, network connectivity is checked and key connections are added, and after a maximum of 10 iterative optimizations, the personalized routing network topology is generated.
15. A SMS marketing effect optimization system based on multi-channel dynamic weights, characterized in that: include: Data collection module, used to obtain real-time performance data of multiple SMS channels; A standardization processing module is used to standardize the real-time performance data and map the reach rate, conversion rate, complaint rate, and latency indicators to the [0, 1] interval using a Z-score standardization method to obtain standardized performance indicators; A scoring calculation module is configured to calculate the comprehensive performance score of each channel by weighted summation based on the standardized performance indicators, using the scoring formula Score = w1 × reach rate + w2 × conversion rate - w3 × complaint rate - w4 × latency index to generate a channel performance benchmark score; The prediction module collects multi-dimensional performance data from each channel over the past 24 hours, analyzes the temporal characteristics of historical conversion data through an LSTM neural network, and uses a three-layer LSTM network structure to learn temporal patterns to predict conversion rate trends for each channel over the next 2-4 hours. A weight optimization module is used to solve a weight distribution scheme with the goal of maximizing the overall ROI based on the conversion rate trend change and the channel performance benchmark score using a gradient optimization method, set an objective function F(w)=Σ(wi×predicted conversion rate i×weight coefficient), and calculate the optimal distribution ratio of each channel; The diversion adjustment module is used to adjust the diversion strategy of the SMS queue in real time according to the optimal distribution ratio, redistribute the SMS to be sent to the queues of each channel, and at the same time start the smooth transition mechanism of the weight configuration to achieve dynamic optimization of the SMS marketing effect.
Citation Information
Patent Citations
Short message distribution method and system
CN107347182A
Resource operation data prediction method, prediction model training method and device
CN114529007A
Method for reasoning state of mobile phone number by mining short message receipt data
CN116361357A
Short message channel fault processing method based on dynamic routing adjustment
CN120075859A
Short message channel intelligent recommendation system and method established based on intelligent marketing system
CN120186567A