Dynamic Channel Configuration System and Method for Multi-Area SMS Services

By adopting dynamic channel configuration system and layered reinforcement learning model in multi-regional SMS services, channel selection is optimized, and short-sighted problems of long-term channel selection in the existing technology are solved, achieving more efficient SMS delivery and user satisfaction.

CN119697604BActive Publication Date: 2025-05-27SIMBA NETWORK TECH (NANJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510214822.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing technology lacks comprehensive consideration of long-term channel selection in SMS services in multiple regions, resulting in the future transmission effect being affected by current channel selection.

Method used

The dynamic channel configuration system and method of multi-regional SMS services are adopted to define the agent's action space and return function by obtaining the comprehensive quality indicators of each channel, including average delay, success rate, cost and price advantage coefficients, and optimize channel cluster selection and channel selection strategies using a hierarchical reinforcement learning model.

Benefits of technology

The dynamic channel configuration of SMS services in multiple regions has been realized, taking into account factors such as delay, success rate, and cost, which has improved SMS delivery rate and timeliness and improved user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697604B_ABST
    Figure CN119697604B_ABST
Patent Text Reader

Abstract

The present application discloses a dynamic channel configuration system and method for multi-region SMS services, which relates to the technical field of information transmission. For each channel, its average delay, SMS sending success rate, and average cost for every 10,000 successfully sent test SMS are obtained, the price advantage coefficient is calculated, and the comprehensive quality index of each channel is calculated; based on the comprehensive quality index, the first-layer action space and the second-layer action space are defined, the state s is defined as the business attribute feature vector of the SMS sending task, the action a of the agent is defined, the first-layer reward function and the second-layer reward function are defined; a hierarchical reinforcement learning model is adopted, the channel cluster selection value function is defined, the channel cluster selection strategy objective function is designed, the channel selection Q-value function is defined, the channel selection strategy objective function is designed, and the hierarchical reinforcement learning model is optimized; based on the hierarchical reinforcement learning model, channels are distributed for real-time SMS sending tasks, realizing the dynamic channel configuration of multi-region SMS services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information transmission technology, and particularly to a dynamic channel configuration system and method for multi-region SMS services. Background Art

[0002] To meet the needs of multi-region SMS services, SMS service providers usually connect SMS channels in multiple regions and dynamically select the optimal channel according to the location of the target mobile phone number to send SMS.

[0003] The patent application with the publication number CN111246406A discloses a method, system, storage medium, and terminal device for sending SMS. The method includes obtaining interface documents corresponding to each SMS channel, and respectively configuring channel information corresponding to each SMS channel according to each interface document, where the channel information includes the request method corresponding to the SMS channel and the correspondence between the channel fields in the SMS channel and the system source fields; when receiving a user's SMS sending request, determining a target SMS channel from the SMS channels according to the SMS sending request, and determining the target user and the source field content corresponding to the SMS sending request; converting the system source fields into target channel fields corresponding to the target SMS channel according to the correspondence, and determining the source field content corresponding to the system source fields as the target field content corresponding to the target channel fields; obtaining the target request method corresponding to the target SMS channel, and using the target request method and the target field content to send SMS to the target user.

[0004] Traditional methods usually select channels based on the effect of a single send. This single decision-making idea lacks consideration of long-term benefits and does not take into account the impact that the current channel selection may have on future send effects. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the prior art to some extent. For this reason, an object of this application is to propose a dynamic channel configuration system and method for multi-region SMS services, which realizes the dynamic channel configuration of multi-region SMS services.

[0006] One aspect of this application provides a dynamic channel configuration method for multi-region SMS services, including:

[0007] Step S100: For each channel, obtain its average delay, SMS send success rate, and average cost per 10,000 test SMS sent successfully, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel;

[0008] Step S200: Define the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index. Define the state s as the business attribute feature vector of the SMS sending task. Define the action a of the agent as selecting a channel cluster from the first-layer action space and a channel from the second-layer action space in state s. Define the first-layer reward function and the second-layer reward function;

[0009] Step S300: Adopt a hierarchical reinforcement learning model. The hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy. Define the channel cluster selection value function and design the channel cluster selection strategy objective function. Define the channel selection Q-value function and design the channel selection strategy objective function to optimize the hierarchical reinforcement learning model;

[0010] Step S400: Obtain the real-time SMS sending task, distribute channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and perform the sending;

[0011] For each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 successfully sent test SMS. Calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels. The specific method for calculating the comprehensive quality index of each channel is as follows:

[0012] Step S110: Randomly send test SMS on the channel, record the time difference from sending each test SMS to receiving the receipt, obtain the delay sample set, and calculate the arithmetic mean of the delay sample set to obtain the average delay of the channel ;

[0013] Step S120: Take the maximum value of the average delays of all channels as the delay normalization denominator ;

[0014] Step S130: Preset a time window , within the time window ct, count the total number of test SMS sent on the channel, denoted as , count the number of successfully sent test SMS as , calculate the ratio between the number of successfully sent test SMS and the total number of sent test SMS to obtain the SMS sending success rate of the channel ;

[0015] Step S140: Obtain the reachability adjustment coefficient of the area where the channel is located ;

[0016] Step S150: Obtain the unit price from the price list of the channel, denoted as , and calculate the average cost per 10,000 successfully sent test SMS of the channel ;

[0017] Step S160: Obtain the average unit price of other channels in the same region as the channel, calculate the average unit price of other channels and the unit price of each channel , and calculate the price advantage coefficient between them ; ;

[0018] Step S170: For each channel, define a comprehensive quality index function, and calculate the comprehensive quality index of each channel based on the calculated average delay, delay normalization denominator, SMS sending success rate, reachability adjustment coefficient, average cost, and price advantage coefficient ;

[0019] Define the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index. Define the state s as the business attribute feature vector of the SMS sending task. The specific method of defining the action a of the agent as selecting a channel cluster from the first-layer action space and a channel from the second-layer action space in the state s is as follows:

[0020] Step S210: Form a channel set with all channels;

[0021] Step S220: Use the comprehensive quality index of each channel as a data point, and randomly select the comprehensive quality indexes of M channels as the initial M clustering centers;

[0022] Step S230: For each channel, calculate its Euclidean distance from the M clustering centers, assign each channel to the channel cluster where the nearest clustering center is located. For each channel cluster, when the channels are all assigned, update its clustering center to the mean value of the comprehensive quality indexes of all channels within the channel cluster. Repeat the assignment of each channel to the channel cluster where the updated clustering center is located until all clustering centers and the data points within the channel clusters where the clustering centers are located no longer change, obtaining K channel clusters; K = M;

[0023] Step S240: Define two layers of action spaces. The first-layer action space represents the selection of channel clusters, and define the K channel clusters as the first-layer action space , where is the Kth channel cluster; the second-layer action space represents selecting a channel within the channel cluster , where, represents the jth channel in the kth channel cluster ; j ≤ J; k ≤ K; ;

[0024] Step S250: Define the state s as the business attribute feature vector of the SMS sending task. The business attribute feature vector includes the SMS sending volume Target area , Sending time window ;

[0025] Step S260: Define action a as the agent selecting a channel cluster from the first layer action space in state s , and from the channel cluster Select a channel Execute SMS sending;

[0026] Step S270: Define a hierarchical action selection strategy for the agent to select action a in state s ,in, Indicates the selection of channel cluster in state s The probability of Indicates that in state s, from the channel cluster Select channel probability;

[0027] The calculation method of the first-layer reward function is:

[0028] Step S280: For the first layer action space, define the first layer reward function as , get the channel cluster The average comprehensive quality index of all channels in , get the channel cluster used in state s Average cost , based on channel clusters The mean of the average delay of the internal channels and the mean of the SMS sending success rate are used to calculate the channel cluster selected under state s. Synergy , select the channel cluster under calculation state s In the target area Matching Rewards , based on the average comprehensive quality index, average cost, synergy effect, and matching reward, calculate the first-layer reward function, and define the cumulative reward of the first-layer action space based on the first-layer reward function ;

[0029] The calculation method of the second-layer reward function is:

[0030] Step S290: For the second-layer action space, define the second-layer reward function as , construct the current state-action triplet according to the current state, the selected channel cluster and the channel, construct the historical state-action triplet sequence according to the state of the historical N time steps, the selected channel cluster and the channel, and calculate the similarity between the current state-action triplet and the historical state-action triplet sequence , get the channel continuity reward, calculate the channel selection under state s In the target area resource occupancy penalty , based on the channel clusters in the channels comprehensive quality index , the average cost of using the channel in state s , the synergy effect of selecting the channel cluster in state s , the channel continuity reward and the resource occupancy penalty, calculate the second-layer return function, and define the cumulative return of the second-layer action space based on the second-layer return function ; ; ;

[0031] The adopted hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define a channel cluster selection value function and design a channel cluster selection strategy objective function, define a channel selection Q-value function and design a channel selection strategy objective function, and the specific method for optimizing the hierarchical reinforcement learning model is as follows:

[0032] Step S310: Adopt a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, and define the channel cluster selection value function of the hierarchical reinforcement learning model , representing the long-term value of selecting the channel cluster in state s, and design a channel cluster selection strategy objective function based on the channel cluster selection value function ; ;

[0033] Step S320: Define the channel selection Q-value function , representing the long-term value of selecting the channel from the channel cluster in state s , and design a channel selection strategy objective function based on the channel selection Q-value function; ;

[0034] Step S330: Construct the loss function of the hierarchical reinforcement learning model, use the stochastic gradient descent algorithm to minimize the loss function, use the policy gradient method to maximize the channel cluster selection strategy objective function, and update the parameters of the channel cluster selection strategy , use the policy gradient method to maximize the channel selection strategy objective function, and update the parameters of the channel selection strategy , and obtain the optimized hierarchical reinforcement learning model;

[0035] The specific method for obtaining the real-time SMS sending task, distributing channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and performing the sending is as follows:

[0036] Step S410: Obtain a real-time SMS sending task and extract the business attribute feature vector of the real-time SMS sending task;

[0037] Step S420: Use the business attribute feature vector of the real-time SMS sending task as the state s and input it into the hierarchical reinforcement learning model. Make a decision on the first-layer action space through the channel cluster selection strategy, calculate the probability of selecting each channel cluster in state s, and output the channel cluster selected by the first-layer action space by the channel cluster selection strategy;

[0038] Step S430: Input the state s and the selected channel cluster into the channel selection strategy, calculate the probability of selecting each channel in state s and the selected channel cluster, make a decision on the second-layer action space through the channel selection strategy, and output the channel selected by the second-layer action space;

[0039] Step S440: Distribute the real-time SMS sending task to the selected channel for sending.

[0040] One aspect of the present application provides a dynamic channel configuration system for multi-region SMS services, including:

[0041] A channel quality calculation module, for each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 test SMS sent successfully, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel;

[0042] An agent definition module, based on the comprehensive quality index, define the first-layer action space and the second-layer action space of the agent, define the state s as the business attribute feature vector of the SMS sending task, define the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space in state s, and define the first-layer reward function and the second-layer reward function;

[0043] A reinforcement learning optimization module, using a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define the channel cluster selection value function and design the channel cluster selection strategy objective function, define the channel selection Q-value function and design the channel selection strategy objective function, and optimize the hierarchical reinforcement learning model;

[0044] A real-time SMS sending module, obtain a real-time SMS sending task, distribute channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and perform the sending.

[0045] One aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it realizes the steps in the dynamic channel configuration method for multi-region SMS services.

[0046] One aspect of the present application provides a readable storage medium storing a computer program adapted to be loaded by a processor to execute steps in a method for dynamic channel configuration of multi-region SMS services.

[0047] The dynamic channel configuration system and method for multi-region SMS services proposed in the present application have the following advantages compared with the prior art:

[0048] The calculation of the comprehensive quality index of the present application fully considers key factors such as channel delay, success rate, cost, etc., as well as external influences such as regional differences and market competition, making the evaluation results closer to the actual business needs.

[0049] The present application upgrades the original channel-level action space to a channel cluster level, greatly reducing the dimension of the action space and simplifying the exploration difficulty of the reinforcement learning algorithm. Channels within the same cluster have similar performance, and selecting channels within the same cluster has little impact on the sending effect. This hierarchical action space design is more in line with the characteristics of the actual task.

[0050] In the first-layer and second-layer reward functions defined in the present application, not only the comprehensive quality and sending cost of the channel are considered, but also multiple influencing factors such as synergy effect, matching degree, and continuity are considered for reward or punishment, guiding the reinforcement learning algorithm to weigh various influencing factors and balance short-term and long-term effects.

[0051] The present application uses a hierarchical reinforcement learning model to model and optimize the channel selection process, defines a channel cluster selection value function and a channel selection Q-value function, and models the original problem as a hierarchical Markov decision process of making decisions at the cluster level first and then at the channel level. Channel cluster selection considers long-term rewards, and channel selection considers local optimality. The entire model takes into account both long-term benefits and local optimality, overcoming the short-sightedness of single decision-making in traditional methods.

[0052] The present application applies the trained hierarchical reinforcement learning model to the actual SMS sending task, dynamically adjusts the channel combination according to the specific requirements of each task, realizes personalized optimization at the task granularity, improves the delivery rate and timeliness of SMS, and enhances user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of the method for dynamic channel configuration of multi-region SMS services provided by the present application;

[0054] Figure 2 is a flowchart of the calculation method for the comprehensive quality index provided by the present application;

[0055] Figure 3Flowchart of the construction process of the hierarchical reinforcement learning model provided by this application;

[0056] Figure 4 Functional module diagram of the dynamic channel configuration system for multi-region SMS services provided by this application. Detailed implementation manners

[0057] To better understand this application, various aspects of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0058] In the drawings, for ease of illustration, the sizes, dimensions, and shapes of the elements have been slightly adjusted. The drawings are only examples and are not drawn to an exact scale. As used herein, terms such as "substantially", "about", and similar terms are used as terms indicating approximation, rather than terms indicating degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art. Additionally, in this application, the order of description of the various step processes does not necessarily represent the order in which these processes occur in actual operation, unless otherwise clearly specified or derivable from the context.

[0059] It should also be understood that expressions such as "comprises", "comprising", "has", "including", and / or "including having" are open-ended rather than closed-ended expressions in this specification, which means that there are the stated features, elements, and / or components, but do not exclude the existence of one or more other features, elements, components, and / or their combinations. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features, rather than just individual elements in the list. In addition, when describing embodiments of this application, the use of "may" means "one or more embodiments of this application". And the term "exemplary" is intended to refer to an example or illustration.

[0060] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as commonly understood by a person of ordinary skill in the art to which this application belongs. It should also be understood that, unless clearly stated in this application, words defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense.

[0061] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the accompanying drawings and in combination with the embodiments.

[0062] Embodiment 1

[0063] As Figure 1 shown, the dynamic channel configuration method for multi-region SMS services provided by this application includes:

[0064] Step S100: For each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 successfully sent test SMS, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel;

[0065] The specific method for obtaining the average delay, SMS sending success rate, and average cost per 10,000 successfully sent test SMS for each channel, calculating the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculating the comprehensive quality index of each channel is as follows:

[0066] Step S110: Randomly send test SMS on the channel, record the time difference from sending each test SMS to receiving the receipt, obtain the delay sample set, and calculate the arithmetic mean of the delay sample set to obtain the average delay of the channel ;

[0067] Step S120: Take the maximum value of the average delays of all channels as the delay normalization denominator ;

[0068] Step S130: Preset a time window , within the time window ct, count the total number of test SMS sent on the channel, denoted as , count the number of successfully sent test SMS as , calculate the ratio between the number of successfully sent test SMS and the total number of sent test SMS to obtain the SMS sending success rate of the channel ;

[0069] The calculation formula for the SMS sending success rate of the channel is: ;

[0070] Step S140: Obtain the reachability adjustment coefficient of the region where the channel is located according to expert experience scoring ;

[0071] The reachability adjustment coefficient is used to reflect the impact of factors such as the maturity of the communication infrastructure and control measures in a region on the success rate of a channel. Based on factors such as the communication infrastructure status and control measures in the region where the channel is located, those skilled in the art set the reachability adjustment coefficient for the region where the channel is located according to experience, and the range is [0, 1];

[0072] Step S150: Obtain the unit price from the price list of the channel and denote it as , and calculate the average cost for the channel to successfully send 10,000 test text messages ;

[0073] The calculation formula for the average cost of the channel to successfully send 10,000 test text messages is: ;

[0074] Step S160: Obtain the average unit price of other channels in the same region as the channel , and calculate the price advantage coefficient between the average unit price of other channels and the unit price of each channel ; ;

[0075] The calculation formula for the price advantage coefficient is: ;

[0076] Step S170: For each channel, define a comprehensive quality index function, and calculate the comprehensive quality index of each channel based on the calculated average delay, delay normalization denominator, short message sending success rate, reachability adjustment coefficient, average cost, and price advantage coefficient ;

[0077] The function expression of the comprehensive quality index is: , where α is the delay penalty coefficient, is the smoothing term, β is the success rate reward coefficient, and γ is the cost-benefit elasticity coefficient;

[0078] The values of the delay penalty coefficient, smoothing term, success rate reward coefficient, and cost-benefit elasticity coefficient are set by those skilled in the art according to experience. Among them, the delay penalty coefficient, success rate reward coefficient, and cost-benefit elasticity coefficient are all greater than 1, and the smoothing term is a small positive number close to 0, which is used to avoid abnormal situations where the denominator is 0;

[0079] The comprehensive quality index includes three indicators: delay, success rate, and cost-benefit;

[0080] As Figure 2 shown, it is the flowchart of the calculation method for the comprehensive quality index provided by this application;

[0081] In the above step S100, by obtaining various performance indicators of the channels and introducing the reachability adjustment coefficient and the price advantage coefficient, the comprehensive quality of each channel can be comprehensively and objectively evaluated, providing a reliable data basis for subsequent channel clustering and intelligent selection.

[0082] Step S200: Define the first-layer action space and the second-layer action space of the intelligent agent based on the comprehensive quality index. Define the state s as the business attribute feature vector of the SMS sending task. Define the action a of the intelligent agent as selecting a channel cluster from the first-layer action space and a channel from the second-layer action space in the state s. Define the first-layer reward function and the second-layer reward function.

[0083] The specific method of defining the first-layer action space and the second-layer action space of the intelligent agent based on the comprehensive quality index, defining the state s as the business attribute feature vector of the SMS sending task, and defining the action a of the intelligent agent as selecting a channel cluster from the first-layer action space and a channel from the second-layer action space in the state s is as follows:

[0084] Step S210: Form a channel set with all channels.

[0085] Step S220: Take the comprehensive quality index of each channel as a data point, and randomly select the comprehensive quality indexes of M channels as the initial M clustering centers.

[0086] The M is the number of clusters, and the value of M is set by those skilled in the art according to experience.

[0087] Step S230: For each channel, calculate its Euclidean distance from the M clustering centers, assign each channel to the channel cluster where the nearest clustering center is located. For each channel cluster, when the channels are all assigned, update its clustering center to the mean value of the comprehensive quality indexes of all channels in the channel cluster. Repeat the assignment of each channel to the channel cluster where the updated clustering center is located until all clustering centers and the data points in the channel clusters where the clustering centers are located no longer change, obtaining K channel clusters; K = M.

[0088] Step S240: Define two layers of action spaces. The first-layer action space represents the selection of channel clusters, and define the K channel clusters as the first-layer action space , where is the Kth channel cluster; the second-layer action space represents selecting a channel in the channel cluster , where, represents the jth channel in the kth channel cluster ; j ≤ J; k ≤ K.

[0089] ​Step S250: Define the state s as the business attribute feature vector of the SMS sending task, and the business attribute feature vector includes the SMS sending volume , the target area , the sending time window ;

[0090] Step S260: Define the action a as that the agent selects a channel cluster from the first-layer action space in the state s , and selects a channel from the channel cluster to execute SMS sending;

[0091] Step S270: Define the hierarchical action selection strategy for the agent to select the action a in the state s , where represents the probability of selecting the channel cluster in the state s, represents the probability of selecting the channel from the channel cluster in the state s; The comprehensive quality index is calculated based on the average delay, the delay normalization denominator, the SMS sending success rate, the reachability adjustment coefficient, the average cost, and the price advantage coefficient.

[0092] The calculation method of the first-layer reward function is as follows:

[0093] Step S280: For the first-layer action space, define the first-layer reward function as , obtain the average comprehensive quality index of all channels in the channel cluster , obtain the average cost of using the channel cluster in the state s, calculate the synergy of selecting the channel cluster in the state s based on the mean of the average delay and the mean of the SMS sending success rate of the channels within the channel cluster , calculate the matching degree reward of selecting the channel cluster in the target area , calculate the first-layer reward function based on the average comprehensive quality index, the average cost, the synergy, and the matching degree reward, and define the cumulative reward of the first-layer action space based on the first-layer reward function;

[0094] The calculation formula for the synergy of selecting the channel cluster in the state s is: , where J represents the total number of channels in the channel cluster , ​Denote the channel cluster The average delay of the j-th channel in Denote the channel cluster The success rate of SMS sending of the j-th channel in 、 Are the weight coefficients of the average delay and the success rate of SMS sending respectively, which are set by those skilled in the art according to experience;

[0095] The selected channel cluster in the state s In the target area The calculation formula of the matching degree reward is: , where Denote the channel cluster The historical performance score in the target area Denote the channel cluster The correlation with the target area;

[0096] The channel cluster The calculation method of the historical performance score in the target area is: count the total number of SMS sent by the channel cluster within the historical time, count the number of successfully sent SMS in the channel cluster , calculate the ratio of the number of successfully sent SMS to the total number of sent SMS to obtain the success rate of SMS sending of the channel cluster in the target area , calculate the mean value of the average delay of all channels in the channel cluster in the target area to obtain the average delay of the channel cluster in the target area , count the maximum success rate of SMS sending among all channel clusters in different target areas , count the minimum average delay of all channel clusters , calculate the historical performance score of the channel cluster in the target area;

[0097] The channel cluster The calculation formula of the historical performance score in the target area is: ;

[0098] The channel cluster The calculation formula of the correlation with the target area is: , where Denote the channel cluster The coverage rate in the target area Denote the maximum coverage rate among all channel clusters;

[0099] The function expression of the first-layer reward function is: , where is the weight coefficient of the synergy effect, which is set by those skilled in the art according to experience;

[0100] The calculation formula for the cumulative return of the first-layer action space is: , where t represents the time step, Y represents the discount factor, represents the state at time step t, represents the channel cluster selected at time step t, represents the state-action sequence selected by the channel cluster in a text message sending task;

[0101] The calculation method of the second-layer reward function is:

[0102] Step S290: For the second-layer action space, define the second-layer reward function as , construct the current state-action triple according to the current state, the selected channel cluster and the channel, construct the historical state-action triple sequence according to the states, the selected channel clusters and the channels in the historical N time steps, and calculate the similarity between the current state-action triple and the historical state-action triple sequence , obtain the channel continuity reward, and calculate the resource occupancy penalty for selecting channel in the target area , based on the channel cluster in the channel comprehensive quality index , the average cost of using channel in state s , the synergy effect of selecting channel cluster in state s , the channel continuity reward and the resource occupancy penalty, calculate the second-layer reward function, and define the cumulative return of the second-layer action space based on the second-layer reward function ;

[0103] The calculation formula for the resource occupancy penalty is: , where represents the resource occupancy rate of channel in the target area, which is statistically obtained through the actual communication resource occupancy of this channel in the target area, is the penalty coefficient of the resource occupancy rate, which is set by those skilled in the art according to experience;

[0104] The channel The calculation formula for the resource occupancy rate in the target area is: , where represents the amount of resources currently used by channel in the target area, Represents the total communication resources available in the target area;

[0105] The calculation formula of the second-layer return function is: , where is the weight coefficient of similarity;

[0106] The similarity between the current state-action triple and the historical state-action triple sequence is the cosine similarity;

[0107] The calculation formula of the cumulative return of the second-layer action space is: , where is the channel selected at time step t;

[0108] The goal of the agent is to maximize the cumulative return of the first action space and the cumulative return of the second-layer action space of the SMS sending task;

[0109] As Figure 3 shown, it is the flow chart of the construction process of the hierarchical reinforcement learning model provided by this application;

[0110] The above step S200 uses the clustering method to group the channels, forming multiple channel clusters, which can take into account local optimization and global optimization when selecting channels, ensuring the complementarity of channels within the same cluster and enabling balanced scheduling between clusters, improving the overall performance of the system, and providing a basis for the establishment of the subsequent hierarchical reinforcement learning model.

[0111] Step S300: Adopt a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define a channel cluster selection value function and design a channel cluster selection strategy objective function, define a channel selection Q-value function and design a channel selection strategy objective function, and optimize the hierarchical reinforcement learning model;

[0112] The specific method of adopting a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define a channel cluster selection value function and design a channel cluster selection strategy objective function, define a channel selection Q-value function and design a channel selection strategy objective function, and optimize the hierarchical reinforcement learning model is:

[0113] Step S310: Adopt a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define the channel cluster selection value function of the hierarchical reinforcement learning model , which represents the long-term value of selecting channel cluster in state s, and design a channel cluster selection strategy objective function based on the channel cluster selection value function;

[0114] The channel cluster selection value function has the following functional expression: V (s, a k )= E [ R ( τ A ) s 0 = s , a 0 = a k , Π (a|s) ] , where represents the starting state, represents the channel cluster selected in state ; E [ ∙ ] represents the expected cumulative return obtained under the given state and action;

[0115] The functional expression of the channel cluster selection policy objective function is: J ( θ A )= E s~D [ V (s, a k )] , where represents the parameter of the channel cluster selection policy, D represents the state distribution, and the channel cluster selection policy objective function represents the expected cumulative return generated by the channel cluster selection policy on the state distribution D;

[0116] Step S320: Define the channel selection Q-value function , which represents the long-term value of selecting channel from channel cluster in state s, and design the channel selection policy objective function based on the channel selection Q-value function;

[0117] The functional expression of the channel selection Q-value function is: Q ( s , a k , a kj )= E [ R ( τ a ) s 0 = s , a 0 = a k , a k0 = a kj , Π (a|s) ] , where represents the status in the channel cluster selected in;

[0118] The functional expression of the channel selection strategy objective function is: J ( θ a )= E s~D, a k ~ π A [ Q ( s , a k , a kj )] , where represents the channel cluster selection strategy, is the parameter of the channel selection strategy;

[0119] Step S330: Construct the loss function of the hierarchical reinforcement learning model, use the stochastic gradient descent algorithm to minimize the loss function, use the policy gradient method to maximize the channel cluster selection strategy objective function, and update the parameters of the channel cluster selection strategy , use the policy gradient method to maximize the channel selection strategy objective function, and update the parameters of the channel selection strategy , to obtain the optimized hierarchical reinforcement learning model;

[0120] The loss function The calculation formula of is: ;

[0121] Using the hierarchical reinforcement learning model in the above step S300 to model and optimize the channel selection process can autonomously learn and extract experience from massive data, continuously improve the channel cluster selection strategy and the channel selection strategy, and make the system have self - adaptability and robustness.

[0122] Step S400: Obtain the real-time SMS sending task, distribute channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and perform the sending;

[0123] The specific method for obtaining the real-time SMS sending task, distributing channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and performing the sending is as follows:

[0124] Step S410: Obtain the real-time SMS sending task, and extract the business attribute feature vector of the real-time SMS sending task;

[0125] Step S420: Take the business attribute feature vector of the real-time SMS sending task as the state s and input it into the hierarchical reinforcement learning model. Make a decision on the first-layer action space through the channel cluster selection strategy, calculate the probability of selecting each channel cluster in the state s, and output the channel cluster selected by the first-layer action space by the channel cluster selection strategy;

[0126] Step S430: Input the state s and the selected channel cluster into the channel selection strategy, calculate the probability of selecting each channel in the state s and the selected channel cluster, make a decision on the second-layer action space through the channel selection strategy, and output the channel selected by the second-layer action space;

[0127] Step S440: Distribute the real-time SMS sending task to the selected channel for sending.

[0128] The above Step S400 applies the trained hierarchical reinforcement learning model to the actual SMS sending task, which can automatically complete the intelligent selection and distribution of channels, greatly improving the working efficiency of the system and reducing the need for manual intervention.

[0129] Embodiment 2

[0130] As Figure 4 shown, the dynamic channel configuration system for multi-region SMS services provided by this application includes:

[0131] Channel quality calculation module. For each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 test SMS sent successfully, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel;

[0132] Agent definition module. Define the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index, define the state s as the business attribute feature vector of the SMS sending task, define the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space in the state s, and define the first-layer reward function and the second-layer reward function;

[0133] The reinforcement learning optimization module adopts a hierarchical reinforcement learning model, which includes a channel cluster selection strategy and a channel selection strategy. Define the channel cluster selection value function and design the channel cluster selection strategy objective function, define the channel selection Q-value function and design the channel selection strategy objective function, and optimize the hierarchical reinforcement learning model.

[0134] The real-time SMS sending module obtains real-time SMS sending tasks, distributes channels for the real-time SMS sending tasks based on the hierarchical reinforcement learning model, and sends them.

[0135] Embodiment 3

[0136] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute the dynamic channel configuration method for multi-region SMS services as described above.

[0137] The method or system according to the embodiment of the present application can also be implemented by means of the architecture of the electronic device disclosed in the present application. The electronic device may include a bus, one or more CPUs, a read-only memory (ROM), a random access memory (RAM), a communication port connected to the network, input / output components, a hard disk, etc. The storage device in the electronic device, such as ROM or hard disk, can store the dynamic channel configuration method for multi-region SMS services provided by the present application. The dynamic channel configuration method for multi-region SMS services may include, for example: for each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 test SMS sent successfully, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel; define the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index, define the state s as the business attribute feature vector of the SMS sending task, define the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space in the state s, and define the first-layer reward function and the second-layer reward function; adopt a hierarchical reinforcement learning model, which includes a channel cluster selection strategy and a channel selection strategy, define the channel cluster selection value function and design the channel cluster selection strategy objective function, define the channel selection Q-value function and design the channel selection strategy objective function, and optimize the hierarchical reinforcement learning model; obtain real-time SMS sending tasks, distribute channels for the real-time SMS sending tasks based on the hierarchical reinforcement learning model, and send them. Further, the electronic device may also include a user interface.

[0138] Embodiment 4

[0139] The present application also discloses a readable storage medium. Computer-readable instructions are stored on the readable storage medium. When the computer-readable instructions are run by a processor, the dynamic channel configuration method of the multi-region SMS service disclosed with reference to the present application can be executed. The readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0140] In addition, according to an embodiment of the present application, the process disclosed in the present application can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided in the present application. For example: for each channel, obtain its average delay, SMS sending success rate, and average cost per 10,000 successfully sent test SMSs, calculate the price advantage coefficient between the unit price of this channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel; based on the comprehensive quality index, define the first-layer action space and the second-layer action space of the agent, define the state s as the business attribute feature vector of the SMS sending task, define the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space in the state s, define the first-layer reward function and the second-layer reward function; adopt a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, define the channel cluster selection value function and design the channel cluster selection strategy objective function, define the channel selection Q-value function and design the channel selection strategy objective function, and optimize the hierarchical reinforcement learning model; obtain a real-time SMS sending task, distribute channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and perform the sending. When this computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.

[0141] The method, apparatus, and device of the present application may be implemented in many ways. For example, the method, apparatus, and device of the present application may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present application are not limited to the above specifically described order unless otherwise specifically stated. In addition, in some embodiments, the present application may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0142] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are the same as the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.

[0143] As described above in the specific embodiments, the purpose, technical solutions and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dynamic channel configuration method for a multi-region SMS service, characterized in that: include: For each channel, obtain its average latency, SMS sending success rate, and average cost of successfully sending 10,000 test SMS messages, calculate the price advantage coefficient between the unit price of the channel and the average unit price of other channels, and calculate the comprehensive quality index of each channel; Based on the comprehensive quality index, the first-layer action space and the second-layer action space of the agent are defined. The state s is defined as the business attribute feature vector of the SMS sending task. The action a of the agent is defined as selecting a channel cluster from the first-layer action space and a channel from the second-layer action space under the state s. The first-layer reward function and the second-layer reward function are defined. Adopting a hierarchical reinforcement learning model, the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, defining a channel cluster selection value function and designing a channel cluster selection strategy objective function, defining a channel selection Q value function and designing a channel selection strategy objective function, and optimizing the hierarchical reinforcement learning model; Obtain real-time SMS sending tasks, distribute channels for real-time SMS sending tasks based on the hierarchical reinforcement learning model, and send them.

2. The method for dynamic channel configuration of multi-regional SMS service as claimed in claim 1, characterized in that: The specific method of defining the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index, defining the state s as the business attribute feature vector of the SMS sending task, and defining the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space under the state s is as follows: All channels form a channel set; Taking the comprehensive quality index of each channel as the data point, randomly select the comprehensive quality indexes of M channels as the initial M cluster centers; For each channel, calculate its Euclidean distance with M cluster centers, assign each channel to the channel cluster with the nearest cluster center, and for each channel cluster, update its cluster center to the mean of the comprehensive quality index of all channels in the channel cluster when the channel allocation is completed, and repeatedly assign each channel to the channel cluster with the updated cluster center until all the cluster centers and the data points in the channel clusters where the cluster centers are located no longer change, and obtain K channel clusters; K = M; Define two layers of action space. The first layer of action space represents the selection of channel clusters. Defined as the first-level action space ,in is the Kth channel cluster; the second-level action space represents the channel cluster Select a channel ,in, represents the kth channel cluster The Jth channel in; j≤J; k≤K; Define state s as the business attribute feature vector of the SMS sending task, which includes the amount of SMS sent Target area , Sending time window ; Action a is defined as the agent selecting a channel cluster from the first-layer action space in state s. , and from the channel cluster Select a channel Execute SMS sending; Define a hierarchical action selection strategy for the agent to select action a in state s ,in, Indicates the selection of channel cluster in state s The probability of Indicates that in state s, from the channel cluster Select channel probability; The comprehensive quality index is calculated based on the average delay, the normalized denominator of the delay, the success rate of SMS sending, the reachability adjustment coefficient, the average cost and the price advantage coefficient.

3. The method for dynamic channel configuration of multi-regional SMS service as claimed in claim 2, characterized in that: The calculation method of the first-layer reward function is: For the first-layer action space, the first-layer reward function is defined as , get the channel cluster The average comprehensive quality index of all channels in , get the channel cluster used in state s Average cost , based on channel clusters The mean of the average delay of the internal channels and the mean of the SMS sending success rate are used to calculate the channel cluster selected under state s. Synergy , select the channel cluster under calculation state s In the target area Matching Rewards , based on the average comprehensive quality index, average cost, synergy effect, and matching reward, calculate the first-layer reward function, and define the cumulative reward of the first-layer action space based on the first-layer reward function .

4. The method for dynamic channel configuration of multi-regional SMS service as claimed in claim 3, characterized in that: The calculation method of the second-layer reward function is: For the second-layer action space, the second-layer reward function is defined as , construct the current state-action triplet according to the current state, the selected channel cluster and the channel, construct the historical state-action triplet sequence according to the state of the historical N time steps, the selected channel cluster and the channel, and calculate the similarity between the current state-action triplet and the historical state-action triplet sequence , get the channel continuity reward, calculate the channel selection under state s In the target area Resource usage penalty , based on channel clusters Middle Channel Comprehensive quality indicators , use channels in status s Average cost , select channel cluster in state s Synergy , channel continuity reward and resource occupation penalty, calculate the second-layer reward function, and define the cumulative reward of the second-layer action space based on the second-layer reward function .

5. The method for dynamic channel configuration of multi-regional SMS service as claimed in claim 4, characterized in that: The hierarchical reinforcement learning model is adopted, and the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, defines a channel cluster selection value function and designs a channel cluster selection strategy objective function, defines a channel selection Q value function and designs a channel selection strategy objective function, and the specific method for optimizing the hierarchical reinforcement learning model is: A hierarchical reinforcement learning model is used, wherein the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, and a channel cluster selection value function of the hierarchical reinforcement learning model is defined. , indicating the selection of channel clusters in state s The long-term value of the channel cluster selection strategy is designed based on the channel cluster selection value function. ; Define the channel selection Q value function , which means that in state s, from the channel cluster Select channel The long-term value of the channel selection strategy is designed based on the channel selection Q value function; Construct the loss function of the hierarchical reinforcement learning model, use the stochastic gradient descent algorithm to minimize the loss function, use the policy gradient method to maximize the channel cluster selection strategy objective function, and update the parameters of the channel cluster selection strategy , use the policy gradient method to maximize the channel selection strategy objective function and update the parameters of the channel selection strategy , and obtain the optimized hierarchical reinforcement learning model.

6. The method for dynamic channel configuration of multi-regional SMS service as claimed in claim 5, characterized in that: The specific method of obtaining the real-time SMS sending task, distributing the real-time SMS sending task channels based on the hierarchical reinforcement learning model, and sending the task is as follows: Obtaining a real-time SMS sending task, and extracting a business attribute feature vector of the real-time SMS sending task; The business attribute feature vector of the real-time SMS sending task is input into the hierarchical reinforcement learning model as the state s. The channel cluster selection strategy is used to make decisions on the first-layer action space, and the probability of selecting each channel cluster under the state s is calculated. The channel cluster selection strategy outputs the channel cluster selected in the first-layer action space. Input the state s and the selected channel cluster into the channel selection strategy, calculate the probability of selecting each channel under the state s and the selected channel cluster, make decisions on the second-layer action space through the channel selection strategy, and output the channel selected in the second-layer action space; Distribute real-time SMS sending tasks to the selected channels for sending.

7. A dynamic channel configuration system for multi-regional SMS services, which is implemented based on the dynamic channel configuration method for multi-regional SMS services according to any one of claims 1 to 6, characterized in that: include: The channel quality calculation module obtains the average delay, SMS sending success rate, and average cost of successfully sending 10,000 test SMS messages for each channel, calculates the price advantage coefficient between the unit price of the channel and the average unit price of other channels, and calculates the comprehensive quality index of each channel; The agent definition module defines the first-layer action space and the second-layer action space of the agent based on the comprehensive quality index, defines the state s as the business attribute feature vector of the SMS sending task, defines the action a of the agent as selecting a channel cluster from the first-layer action space and selecting a channel from the second-layer action space under the state s, and defines the first-layer reward function and the second-layer reward function; A reinforcement learning optimization module adopts a hierarchical reinforcement learning model, wherein the hierarchical reinforcement learning model includes a channel cluster selection strategy and a channel selection strategy, defines a channel cluster selection value function and designs a channel cluster selection strategy objective function, defines a channel selection Q value function and designs a channel selection strategy objective function, and optimizes the hierarchical reinforcement learning model; The real-time SMS sending module obtains the real-time SMS sending task, distributes the channels for the real-time SMS sending task based on the hierarchical reinforcement learning model, and sends the task.

8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the dynamic channel configuration method for multi-regional SMS service as described in any one of claims 1 to 6 are implemented.

9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the dynamic channel configuration method for multi-region SMS service as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Short message sending method and system, storage medium and terminal equipment

    CN111246406A

  • Method and device for capturing space target and storage medium

    CN111687840A

  • Short message channel determination method and related device

    CN117835169A