Short message sending optimization system based on big data

By using big data-based user profile feature filtering and dynamic traffic allocation, a spatiotemporal testing matrix is ​​constructed, which solves the shortcomings of existing technologies in user grouping and traffic allocation, and achieves accurate matching of SMS sending strategies and improved marketing effectiveness.

CN120980459AActive Publication Date: 2025-11-18BEIJING XUNYIN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511126679.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing technologies, when optimizing SMS sending based on big data, fail to deeply consider the intrinsic relationship between user profile characteristics and SMS response rates. This results in difficulties in achieving the required homogeneity and heterogeneity in user groups, a lack of scientific mapping in traffic allocation, and a lack of precision in the strategy decision feedback module, making it impossible to implement accurate SMS sending strategies.

Method used

By filtering user profile features, using Pearson correlation coefficient to accurately divide homogeneous groups, and rigorously verifying internal homogeneity and inter-group heterogeneity, a spatiotemporal test matrix is ​​constructed, traffic is dynamically allocated, and iterative optimization is carried out in conjunction with the strategy decision feedback module to ensure data representativeness and scientific validity.

Benefits of technology

This enabled precise matching of SMS sending strategies to the needs of different user groups, improving overall marketing effectiveness and increasing user response and conversion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980459A_ABST
    Figure CN120980459A_ABST
Patent Text Reader

Abstract

The invention discloses a short message sending optimization system based on big data, and relates to the technical field of big data marketing, and the system comprises a compliance pre-verification module, a multi-version content generation module, a space-time test matrix construction module, a dynamic flow distribution module, and a strategy decision feedback module. The method comprises the following steps of: screening user portrait features, accurately selecting features with high correlation with a short message response rate by applying a Pearson's correlation coefficient, dividing a target user group into K homogeneous groups, strictly verifying internal homogeneity and inter-group heterogeneity of the groups, and on the basis, mapping each group user to a test unit matrix according to a proportion; and the consistency between the sample feature distribution and the target group is verified by calculating the feature statistics, so that the sample of the test unit can be ensured to represent the target group to the maximum extent, and the data on which the subsequent strategy decision feedback module depends is more representative and scientific.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data marketing, in particular to a short message sending optimization system based on big data. BACKGROUND

[0002] In today's digital wave, the communication methods between enterprises and users are increasingly diverse, and as a basic and important information reaching means, short message still plays a key role in many fields, whether it is an e-commerce platform pushing promotional activities to users, a financial institution informing customers of account changes, or an educational institution sending course notifications. Short message has become an effective channel for information dissemination due to its immediacy and convenience. With the rise of big data technology, how to optimize short message sending with massive data to improve user response rate and marketing effect has become a key problem that enterprises need to solve.

[0003] However, the current short message sending optimization based on big data has obvious deficiencies. Many systems fail to consider the internal relationship between user portrait features and short message response rate when grouping users, and simply group them according to a few surface features, resulting in difficulty in meeting the homogeneity and heterogeneity of grouping, and inability to accurately distinguish the reaction differences of different user groups to short messages. In the traffic allocation link, there is a lack of scientific mapping and verification mechanism, and the test unit sample feature distribution is often out of touch with the target group, making the subsequent strategies based on test samples lack of universality and effectiveness. Moreover, the existing strategy decision feedback module often ignores the analysis of the complex interaction effects between content, time and user grouping, making it difficult to develop accurate and effective short message sending strategies for different user groups.

[0004] In summary, it is of great practical significance to develop a short message sending optimization system based on big data that can accurately divide user groups, scientifically allocate traffic, and deeply analyze strategy effects to achieve precise feedback optimization. SUMMARY

[0005] The short message sending optimization system based on big data aims to make up for the shortcomings of the prior art, and the system is used for user grouping and traffic allocation, selects features related to short message response rate through user portrait feature screening and Pearson correlation coefficient, divides target user groups into K homogeneous groups, strictly verifies the internal homogeneity and inter-group heterogeneity of the groups, maps each group of users to a test unit matrix at a proportion of 5% on this basis, verifies the consistency of sample feature distribution and target group through calculation of feature statistics, can ensure that the test unit sample represents the target group to the greatest extent, provides a reliable data basis for subsequent strategy making, makes the data relied on by the subsequent strategy decision feedback module more representative and scientific, and finally makes the short message sending strategy based on these data more suitable for the needs of different user groups, and improves the overall short message marketing effect.

[0006] The short message sending optimization system based on big data aims to make up for the shortcomings of the prior art, and the system is used for user grouping and traffic allocation, selects features related to short message response rate through user portrait feature screening and Pearson correlation coefficient, divides target user groups into K homogeneous groups, strictly verifies the internal homogeneity and inter-group heterogeneity of the groups, maps each group of users to a test unit matrix at a proportion of 5% on this basis, verifies the consistency of sample feature distribution and target group through calculation of feature statistics, can ensure that the test unit sample represents the target group to the greatest extent, provides a reliable data basis for subsequent strategy making, makes the data relied on by the subsequent strategy decision feedback module more representative and scientific, and finally makes the short message sending strategy based on these data more suitable for the needs of different user groups, and improves the overall short message marketing effect. The compliance pre-verification module is used for connecting a compliance database, and performing compliance pre-verification on all versions of short message content, intercepting short message content versions containing forbidden words and false propaganda through real-time semantic analysis, and outputting a compliance content pool. The multi-version content generation module generates N short message content versions based on user historical behavior data, and inputs the N short message content versions to the compliance pre-verification module for filtering to generate a final test content set. ; The space-time test matrix construction module performs Cartesian product combination on the test content set C and a sending time set to construct a test unit matrix , wherein each unit represents content sent at time ; The dynamic traffic allocation module divides user groups into K groups according to user portrait features, and maps each group to different test units of the CT matrix at a preset proportion, to ensure balanced distribution of test sample of each unit. The strategy decision feedback module quantitatively analyzes conversion rate data of each unit, decouples the independent effects of content factors and time factors, outputs an optimal content-time combination strategy, and feeds back to the multi-version content generation module for iterative optimization.

[0007] Further, the compliance pre-verification module comprises: Compliance database interface unit: for real-time docking of compliance database containing advertising law forbidden word library, industry standard clause library, false propaganda feature library, and supporting automatic updating and synchronization of the database; Semantic analysis engine: adopting bidirectional long short-term memory network Bi-LSTM to perform word segmentation processing on the input short message content version, extract semantic feature vectors, and calculate the violation confidence score of the risk phrase by calculating with the feature vectors in the compliance database , identify potential forbidden words and false propaganda expressions in the content; The violation confidence score , wherein is the number of key phrases extracted in the short message content, is the Bi-LSTM output feature vector of the i-th key phrase, is the standard feature vector of the i-th violation feature in the compliance database, is the cosine similarity function, is the weight of the violation word in the historical violation case, is the adjustable weight coefficient, and the default value , represents taking the maximum value of the similarity with all violation features.

[0008] Further, the compliance pre-checking module further includes a hierarchical interception unit, which performs hierarchical operations according to the violation confidence score: When , it is determined that there is a high-risk violation, and an alarm is immediately generated to trigger the manual review process; When , it is determined that there is a general violation, and the short message content is rewritten, and the confidence score is recalculated after rewriting; When , the content is allowed to enter the compliance content pool; The hierarchical interception unit records in detail the content version of each round of pre-checking, interception result, violation confidence score, processing method and time stamp, forms a traceable compliance checking log, and feeds back the high-frequency appearing violation features to the compliance database.

[0009] Further, the multi-version content generation module includes: User feature extraction unit: for obtaining user multi-dimensional data from a big data platform and constructing a user feature vector wherein is the historical behavior feature, the user feature vector adopts weighted fusion to obtain a user comprehensive feature score ​​in, The feature weights are those that satisfy the following conditions: ; Content Template Library: Stores scenario-based content templates categorized by industry scenarios. Each template contains fixed text segments and variable placeholders, and is associated with speech style tags. Generative AI Engine: Employs a lightweight generative model, taking user feature vectors as input. Combine content templates to generate personalized content candidates; Version diversity verification: Based on the basic content output by the generative AI engine, an initial version is generated by synonym replacement, sentence transformation, and addition / removal of redundant information. The semantic similarity between the two initial versions is then calculated. The output value range is [0, 1]. N target versions are selected from the initial version to ensure that the average similarity between versions satisfies [0, 1]. ,in, Indicates from The number of combinations of selecting two elements from n elements without order. For diversity threshold, Representing the Each version of the SMS content Representing the For each SMS content version, the iteration count is automatically increased if the filtering result does not meet the threshold. N personalized SMS content versions that have passed diversity validation are output and input into the compliance pre-validation module for filtering. After removing non-compliant versions, the final test content set is generated. ,in and To ensure the validity of the test, This represents the number of items retained after filtering by the compliance pre-verification module. Each version of the SMS message content.

[0010] Furthermore, the specific process by which the generative AI engine generates personalized content candidates is as follows: Template matching: Score users based on their overall characteristics Match the templates with the applicable feature ranges of the template library and select the three candidate templates with the highest matching degree; Variable population: User features are associated with template variables based on entity linking technology. This entity linking technology performs type parsing on template variable placeholders, establishes a variable type library, and assigns a unique ID to each variable type from the user feature vector. Extract entities that match the type, construct an entity pool, and calculate the semantic association between entities and variables using the following formula: ,in, Candidate entities in the entity pool For template variables, For BERT embedding vectors of entities / variables, To populate the variable with the entity that appears most frequently in the user's historical behavior, select the entity with the highest Rel value. Style Adjustment: Based on the wording style tags in the user tags, the template text is rewritten to adapt to the user and obtain personalized content candidates.

[0011] Furthermore, the spatiotemporal test matrix construction module includes: Time Set Generation Unit: Used to extract user activity time features from user historical behavior data and generate a set of sending times. ; Matrix building unit: the final set of test content With the set of sending times Combine them to construct a test unit matrix. Each unit Content In time The sending strategy; The output of the spatiotemporal test matrix construction module is a dynamically adjusted test unit matrix. Each unit A corresponding content-time combination strategy is input into the dynamic traffic allocation module for sample allocation.

[0012] Furthermore, the dynamic traffic allocation module includes: a user grouping unit, a traffic mapping unit, and a traffic balance verification unit, used to allocate user groups to different units of the test unit matrix, wherein: User grouping unit: Based on user profile features, from user feature vectors The target user group was divided into K homogeneous groups based on features highly correlated with SMS response rate. The correlation coefficient of SMS response rate was Pearson correlation coefficient. Filter features, calculate the internal homogeneity and inter-group heterogeneity of each group, and verify the effectiveness of the grouping. The internal homogeneity should satisfy the variance of the SMS response rate of users within the group ≤ 0.1, and the inter-group heterogeneity should satisfy the difference in the mean of the SMS response rate between groups ≥ 0.2. Traffic mapping unit: Maps users from each group to the test unit matrix at a ratio of 5%. Different units; Balance Validation Unit: This unit verifies the consistency between the sample feature distribution of each test unit and the target population. It calculates the feature statistics of each test unit's samples and evaluates the degree of difference between the sample feature distribution and the target population. This degree of difference is obtained by calculating the standard deviation between the test unit samples and the target population. When the degree of difference... Then adjust the traffic allocation ratio and reallocate the samples; The output of the dynamic traffic allocation module is a uniformly distributed test sample, with each test unit... A set of homogeneous user samples are input into the strategy decision feedback module for effect analysis.

[0013] Further, the strategy decision feedback module comprises: A grouping response analysis unit: for the K homogeneous user groups divided by the dynamic traffic distribution module, respectively statistics the response data of each test unit in different groups, the response data includes opening rate , conversion rate , the opening rate represents the proportion of users who open the short message after the content is sent at time , the conversion rate represents the proportion of users who complete the target behavior after opening the short message in the kth user group, a grouping-content-time three-dimensional response matrix is constructed, and the value of each cell in the matrix is ; A cause-effect decoupling unit: separate the independent effects of content and time on user response, and calculate a comprehensive effect score , the content-independent effect only considers the effect of content on the kth group, excluding the interference of time factors, and when calculating, the time is fixed as a constant , and the conversion rate of the content in this period is calculated, that is , the time-independent effect only considers the effect of time on the kth group, excluding the interference of content factors, and when calculating, the content is fixed as the content with the highest historical response rate of the group users , and the conversion rate of the fixed content at time is calculated, that is , and the comprehensive effect score ; A grouping optimal strategy screening unit: for the kth group, the of all test units are sorted from high to low, and the top 10% of test units are selected as the optimal strategy, which is fed back to the multi-version content generation module for iterative optimization and guidance to generate content containing the feature.

[0014] Compared with the prior art, the short message sending optimization system based on big data has the following beneficial effects: One, the present application is in the user group and traffic distribution, through the user portrait feature screening, using Pearson correlation coefficient to select the features with high correlation with short message response rate, and the target user group is divided into K homogeneous groups, and the internal homogeneity and the heterogeneity between groups are strictly verified, on the basis of this, each group user is mapped to the test unit matrix according to 5% proportion, and the consistency of sample feature distribution and target group is verified by calculating the feature statistics, which can ensure that the test unit sample represents the target group to the greatest extent, provide reliable data basis for subsequent strategy making, make the data relied on by subsequent strategy decision feedback module more representative and scientific, and finally make the short message sending strategy based on these data more suitable for different user group needs, and improve the overall short message marketing effect.

[0015] Two, the present application constructs a grouping-content-time three-dimensional response matrix through the grouping response analysis unit, comprehensively records the response of different groups to different content and sending time, the cause and effect decoupling unit accurately separates the independent influence of content and time factors on user response, and then calculates the comprehensive effect score, the strategy screening unit screens and verifies the optimal content-time combination of each group according to the comprehensive effect score, converts these optimal strategies into rules, and feeds back to the multi-version content generation module, which can continuously improve the accuracy and effectiveness of the short message sending strategy, help enterprises better use short messages for precision marketing, and improve user response rate and conversion rate.

[0016] Other advantages, objects, and features of the present application will be in part apparent and in part pointed out hereinafter in the specification, and in part will be observed by persons skilled in the art upon examination of the following specification, or can be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, without paying creative labor, other drawings can also be obtained from these drawings; Figure 1 The operation flow chart of the short message sending optimization system based on big data; Figure 2 The module composition diagram of the short message sending optimization system based on big data; Figure 3 The working principle diagram of the generative AI engine in embodiment one. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1 This embodiment aims to explain in detail the working principle of a big data-based SMS sending optimization system, such as... Figure 2 As shown, this system achieves precise optimization of SMS sending through the collaborative work of a compliance pre-verification module, a multi-version content generation module, a spatiotemporal test matrix construction module, a dynamic traffic allocation module, and a strategy decision feedback module. Starting with compliance verification of SMS content, the system generates multiple versions of personalized content, constructs a test matrix based on the time dimension, conducts tests through scientific user grouping and traffic allocation, and finally determines the optimal sending strategy based on feedback data and continuously iterates, effectively improving the response rate and conversion rate of SMS marketing while ensuring the compliance of SMS content.

[0020] In its implementation, the compliance pre-verification module interfaces with a compliance database to pre-verify the compliance of all versions of SMS content, blocking content containing prohibited words and false advertising, and outputting a compliance content pool to provide a legal and compliant content foundation for subsequent SMS sending. This module mainly consists of a compliance database interface unit, a semantic analysis engine, and a tiered interception unit. The compliance database interface unit interfaces with the compliance database in real time to ensure the system can obtain the latest compliance standards promptly. The compliance database contains key content such as a prohibited advertising word library, an industry standard clause library, and a false advertising feature library. To ensure the timeliness of the database information, the interface unit supports automatic database updates. The system employs a new synchronization mechanism. When industry standards change, the database updates rapidly, and the interface unit detects and synchronizes these changes to the system in real time, ensuring that content verification always adheres to the latest compliance requirements. The semantic analysis engine uses a Bidirectional Long Short-Term Memory (Bi-LSTM) network to process the input SMS content. First, the SMS content is segmented into meaningful words or phrases. Then, Bi-LSTM extracts semantic feature vectors from these words or phrases, accurately representing the semantic information of the text. To determine whether the SMS content violates regulations, a violation confidence score for the risk phrases needs to be calculated. The calculation formula is as follows: ,in, The number of key phrases extracted from the text message content. For the first The Bi-LSTM output feature vector of each key phrase For the first in the compliance database a standard feature vector of the individual violation feature, is a cosine similarity function, used to calculate the similarity degree of two feature vectors, the value closer to 1 indicates that the semantics of the two are more similar, is a violation word is the weight in the historical violation case, which reflects the importance of the violation word in the violation judgment; and is an adjustable weight coefficient, and The default value is used to balance the influence of semantic similarity and violation word weight in the violation confidence score; represents taking the maximum value of the similarity of all violation features, that is, for each key phrase, find the most similar violation feature, and take the calculation result of the most similar feature as the contribution of the key phrase to the violation confidence score. The violation confidence score calculated by this formula , can comprehensively consider the semantic similarity between the key phrases in the short message content and the violation features and the historical weight of the violation words, so as to accurately identify the potential prohibited words and false propaganda expressions in the content; The hierarchical interception unit executes hierarchical operation according to the violation confidence score calculated by the semantic analysis engine : when , it is determined that there is a high-risk violation, at this time, the system will immediately intercept the short message content version and generate an alarm to trigger the manual review process, because the score reaches this threshold, indicating that the short message content has a high possibility of serious violation, which needs to be further audited in detail by manual, to avoid the sending of violation content, when , it is determined that there is a general violation, and the system will automatically rewrite the short message content. Recalculate the confidence score after rewriting, because the violation degree of this type is relatively light, and the automatic rewriting of the system may make it meet the compliance standard, and the recalculation of the score is to confirm whether the rewritten content meets the requirements, , it means that the compliance of the short message content is high, and the content is allowed to enter the compliance content pool for subsequent testing and sending. At the same time, the hierarchical interception unit will record in detail each round of pre-checking content version, interception result, violation confidence score, processing method and time stamp, forming a traceable compliance verification log. These logs not only provide a basis for subsequent compliance audit, but also feedback the high-frequency violation features to the compliance database, further enrich and perfect the compliance database, and improve the compliance verification capability of the system.

[0021] The multi-version content generation module generates a plurality of short message content versions with diversity based on user historical behavior data and inputs to a compliance pre-check module for filtering, and finally generates a test content set. The module is composed of a user feature extraction unit, a content template library, a generative AI engine and a version diversity verification unit. The user feature extraction unit obtains user multi-dimensional data from a big data platform, which covers user's basic information, consumption habits, browsing records, historical purchase behavior and other aspects. Based on these data, a user feature vector is constructed , wherein is the historical behavior feature. In order to comprehensively evaluate the user features, a weighted fusion method is used to obtain the user comprehensive feature score, and the formula is: , wherein is the feature weight and satisfies The determination of the feature weight is based on the influence degree of each feature on the short message response rate. The greater the influence of the feature on the short message response rate, the higher the weight, so as to ensure that the user comprehensive feature score can accurately reflect the potential response of the user to the short message; The content template library stores the scenario content templates classified by industry scenarios. Each template contains fixed text segments and variable placeholders, and is associated with a speech style label. Different industry scenarios (such as e-commerce promotion, financial notification, education course promotion, etc.) have different templates. The fixed text segment ensures the professionalism and basic framework of the short message content, and the variable placeholder provides space for personalized content filling. The speech style label helps subsequent style adjustment according to user features; The generative AI engine adopts a lightweight generative model, inputs the user feature vector and the content template to generate personalized content candidates, as shown in Figure 3 , the specific process is: matching the user comprehensive feature score with the applicable feature interval of the template library. Each template has its corresponding applicable feature interval. The three templates with the highest matching degree are selected by matching to provide a basis for subsequent content generation. Based on entity linking technology, the user features are associated with the template variables. First, the type of the template variable placeholder is analyzed, a variable type library is established, and a unique ID is assigned to each variable type. Then, the type-matched entities are extracted from the user feature vector to construct an entity pool. In order to select the most suitable entity to fill the variable, the semantic correlation degree between the entity and the variable needs to be calculated, and the formula is: , wherein is the candidate entity in the entity pool, is the template variable, is the BERT embedding vector of the entity / variable. The BERT model can convert entities and variables into vectors with semantic information, For the frequency of the entity in the user's historical behavior, the entity with the highest Rel value is selected to fill the variable, making the filled content more in line with the user's personalized needs. According to the style of the user's label, the template text is adapted and rewritten. The initial version is generated by using synonym replacement, sentence conversion, and redundant information addition and subtraction on the basis of the output of the generative AI engine. To ensure sufficient diversity between versions, the semantic similarity of two initial versions is calculated , with an output value range of [0, 1]. The closer the value is to 0, the greater the semantic difference between the two versions. From the initial version, a target version is selected , so that the average similarity between versions meets: , where represents the number of combinations of selecting 2 elements from elements without order, is the diversity threshold, represents the th short message content version, represents the th short message content version, and the formula ensures that the selected versions have sufficient diversity. When the screening result does not meet the threshold, the system will automatically increase the number of iterations, regenerate and screen the versions until personalized short message content versions that pass the diversity check are output. These versions are input into the compliance pre-check module for filtering, and the final test content set is generated after removing the illegal versions , where and to ensure test effectiveness, represent the th short message content version retained after filtering by the compliance pre-check module.

[0022] The space-time test matrix construction module combines the test content set with the sending time set to construct a test unit matrix, providing a structural framework for subsequent traffic allocation and strategy testing. This module includes a time set generation unit and a matrix construction unit. The time set generation unit extracts user active time features from user historical behavior data. User historical behavior data, such as the time of opening a short message and the time of clicking a link, can reflect the user's active situation at different times. By analyzing and mining these data, a sending time set is generated. These time points are relatively active time periods for the user, and sending short messages at these times can achieve higher attention. The matrix construction unit combines the final test content set with the sending time set Carry out Cartesian product combination, which means that each content is paired with each time point to build a test cell matrix , where each cell represents a content at a time sending strategy, the output of the matrix building unit is a dynamically adjusted test cell matrix , each cell corresponds to a content-time combination strategy, which is input to the dynamic traffic allocation module for sample allocation, providing specific policy options for subsequent testing.

[0023] The dynamic traffic allocation module divides the user group into appropriate groups according to the user portrait features, and maps each group to different test cells of the CT matrix in a pre-set proportion, ensuring balanced distribution of test samples in each cell. The module includes a user grouping unit, a traffic mapping unit, and a balance verification unit. The user grouping unit selects features with high correlation to the SMS response rate from the user feature vector based on user portrait features. SMS response rate is an important indicator of SMS effectiveness, and selecting features with high correlation to this indicator can improve the accuracy of grouping. Here, the Pearson correlation coefficient is used to screen features. The Pearson correlation coefficient can reflect the degree of linear correlation between two variables. A value greater than 0.5 indicates a strong positive correlation between the feature and the SMS response rate. Based on the selected features, the target user group is divided into homogeneous groups. To verify the effectiveness of the grouping, the internal homogeneity and inter-group heterogeneity of each group need to be calculated. Internal homogeneity requires that the SMS response rate variance of users within the same group be similar, and inter-group heterogeneity requires that the difference between the SMS response rates of different groups be significant, ensuring that users in different groups respond differently to SMS. Such grouping can provide a basis for subsequent precise strategy formulation. The traffic mapping unit maps users in each group to different cells of the test cell matrix at a rate of 5%. Selecting a 5% rate takes into account both the representativeness of the sample and the cost of testing. It ensures that the sample has a certain size to reflect the overall situation, while not increasing excessive costs and resource consumption due to a large sample size. The balance verification unit verifies the consistency of the sample feature distribution of each test cell with the target group by calculating the feature statistics (such as mean, variance, etc.) of each test cell sample. It assesses the difference between the sample feature distribution and the target group. The difference is obtained by calculating the standard deviation of the test cell sample and the target group. When the difference If the difference is less than 0.05, it indicates a significant deviation between the sample characteristic distribution and the target group, requiring adjustment of the traffic allocation ratio and reallocation of the samples. If the difference is less than 0.05, the sample distribution is considered balanced and effective, and the output of the dynamic traffic allocation module is a balanced distribution of test samples, with each test unit... A set of homogeneous user samples will be input into the strategy decision feedback module for effect analysis, providing data support for strategy optimization.

[0024] The strategy decision feedback module quantitatively analyzes the conversion rate data of each test unit to determine the optimal content-time combination strategy and feeds it back to relevant modules for iterative optimization. This module includes a group response analysis unit, a causal decoupling unit, and a strategy filtering unit. The group response analysis unit is specifically designed for the dynamic traffic allocation module. For each homogeneous user group, the response data of each test unit in different groups is statistically analyzed, including the open rate. and conversion rate Open rate Indicates the first In each user group, the content In time Percentage of users who opened the message after sending it, conversion rate Indicates the first Within each user group, the percentage of users who completed the target action after opening the SMS message was analyzed. Based on this data, a three-dimensional response matrix of group, content, and time was constructed, where the value of each cell in the matrix is... This matrix comprehensively records the responses of different groups to different content and time factors, providing a detailed data foundation for subsequent analysis. The causal decoupling unit separates the independent effects of content factors and time factors, and calculates the comprehensive effect score. Regarding the content independence effect Considering only the content For the The impact of grouping should be considered, excluding the interference of time factors. Time should be taken into account during calculation. Fixed as a constant (This constant is based on the first) The optimal time for determining the historical active time of grouped users, statistical content. The conversion rate during this period, i.e. Regarding the time-independent effect Considering only time For the The impact of grouping, excluding content-related interference, and taking content into account during calculation. Fixed to the content with the highest historical response rate for this user group. Statistics on this fixed content over time The conversion rate, i.e. Overall effect score The weighting of 0.6 and 0.4 is determined based on the degree of influence of content and time factors on the final conversion rate. The strategy screening unit [is used for] the [specific parameters]. Grouping all test units into groups The top 10% of test units are selected as the optimal strategy, sorted from highest to lowest. This selection is designed to ensure the effectiveness of the strategy while retaining a number of alternative solutions for flexible adjustments based on actual circumstances. The selected optimal strategies not only include combinations of content and time but also cover corresponding user group characteristics, enabling precise matching of the needs of different user groups. These optimal strategies are fed back to the multi-version content generation module to guide subsequent content generation iterations and optimizations. Simultaneously, the strategy selection unit continuously tracks and verifies the selected optimal strategies, monitoring their actual effectiveness in real time during subsequent SMS sending. If a significant drop in conversion rate is detected, the strategy selection process is restarted to ensure that the output strategies remain optimal.

[0025] In summary, this embodiment details the complete operation process of a big data-based SMS sending optimization system. From the compliance pre-verification module ensuring the legality and compliance of SMS content through multi-dimensional compliance checks, to the multi-version content generation module generating diverse and personalized content based on user characteristics, to the spatiotemporal test matrix construction module constructing a content-time combination strategy matrix, to the dynamic traffic allocation module achieving scientific user grouping and balanced sample distribution, and finally, the strategy decision feedback module accurately analyzing and outputting the optimal strategy, a complete optimization system is formed. This system, by combining user profiles, historical behavioral data, and other multi-dimensional information, achieves intelligent operation of the entire SMS sending process, from content generation to strategy optimization.

[0026] Example 2 This embodiment provides the operation process of a big data-based SMS sending optimization system for intelligent reconciliation and verification of financial transactions, such as... Figure 1 As shown, the specific steps of this process are as follows: (1) User profile construction and content generation Obtain user historical behavior data from big data platforms; Construct multi-dimensional user feature vectors (such as consumption frequency and active time periods); Calculate the user's comprehensive feature score (weighted fusion of key features); Match the three templates with the highest compatibility in the content template library; Populate personalized variables (based on user characteristics and semantic relevance); Generate basic SMS content and rewrite it to adapt the style; Generate N initial versions by synonym replacement / sentence transformation Filter semantic difference up-to-standard versions (average similarity ≤ 0.3) (2) Content compliance verification Real-time docking with compliance database (forbidden word library, industry standard library); Use Bi-LSTM model to extract semantic features; Calculate the confidence score of key phrases and violation library; Perform hierarchical interception: ≥0.8 points: high-risk interception and trigger manual review; 0.3 ~ 0.8 points: re-score after automatic rewriting; <0.3 points: include in the compliance content pool; Record verification logs and update violation feature library; (3) Construction of spatio-temporal strategy matrix Analyze user active period to generate sending time set; Take the compliance content pool and time point as Cartesian product; Construct a two-dimensional test matrix (row=content version, column=sending period); Each matrix unit corresponds to a specific content+time combination strategy; (4) User grouping and traffic allocation Select features strongly related to SMS response (Pearson coefficient > 0.5); Divide K homogeneous user groups (intra-group response rate variance ≤ 0.1, inter-group difference ≥ 0.2); Extract 5% of users from each group; Map samples uniformly to space-time matrix units; Verify sample distribution consistency (feature difference < 0.05); (5) Strategy implementation and effect feedback Send differentiated SMS according to matrix units; Collect three-dimensional response data (group × content × time); Separate independent effects: Fixed period statistics content conversion rate; Fixed content statistics period conversion rate; Calculate the comprehensive effect score (content weight 60% + time weight 40%); Select the top 10% optimal strategies for each group; Feedback to the content generation module for iterative optimization.

[0027] It will be apparent to those skilled in the art that the application is not limited to the details of the above-exemplified embodiments and that the present application can be implemented in other particular forms without departing from the spirit or essential characteristics of the present application. The embodiments should therefore be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the above description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. No reference signs in the claims should be considered as limiting the scope of the claims with respect to the figures of the patent document.

Claims

1. A big data based SMS sending optimization system characterized in that, The components of the system include: compliance pre-check module, multi-version content generation module, space-time test matrix construction module, dynamic traffic distribution module, policy decision feedback module; The compliance pre-check module is used for connecting a compliance database, pre-checking the compliance of all versions of short message content, intercepting short message content versions containing forbidden words and false propaganda through real-time semantic analysis, and outputting a compliance content pool; The multi-version content generation module generates N short message content versions based on user historical behavior data and inputs them to the compliance pre-check module for filtering to generate a final test content set ; The spatio-temporal test matrix construction module constructs a test content set C and a sending time set Performs Cartesian product combination to construct a test unit matrix Wherein each unit Indicates content At a time Sending strategy; The dynamic traffic distribution module divides user groups into K groups according to user portrait features, and maps each group to different test units of the CT matrix according to a preset proportion, to ensure balanced distribution of test samples in each unit; The policy decision feedback module decouples the independent effects of content factors and time factors by quantitatively analyzing the conversion rate data of each unit, outputs the optimal content-time combination strategy, and feeds back to the multi-version content generation module for iterative optimization.

2. The big data based SMS sending optimization system of claim 1, wherein, The compliance pre-check module includes: The compliance database interface unit is used for real-time connection with a compliance database containing an advertising law forbidden word library, an industry specification clause library, and a false propaganda feature library, and supports automatic updating and synchronization of the database; The semantic analysis engine: adopts bidirectional long short-term memory network Bi-LSTM to carry out word segmentation processing on the input short message content version, extracts semantic feature vectors, and calculates the violation confidence score of the risk phrase by calculating with the feature vectors in the compliance database , identify potential banned words and false advertising expressions in the content; the violation confidence score wherein, is the number of key phrases extracted from the SMS content, is the Bi-LSTM output feature vector of the th key phrase, is the standard feature vector of the th violation feature in the compliance database, is the cosine similarity function, is the weight of the violation word in the historical violation cases, is the adjustable weight coefficient, and the default value , denotes taking the maximum value of the similarity with all violation features.

3. The big data based SMS sending optimization system of claim 2, wherein, The compliance pre-check module further includes a hierarchical interception unit, which performs hierarchical operations according to the violation confidence score: When a high-risk violation is determined to exist, the violation is immediately intercepted and an alert is generated, triggering a manual review process. When a general violation is determined to exist, the SMS content is rewritten, and the confidence score is recalculated after the rewriting. When the content is allowed to enter the pool of compliant content; The hierarchical interception unit records in detail the content version, interception result, violation confidence score, processing method and timestamp of each round of pre-checking, forms a traceable compliance verification log, and feeds back high-frequency appearing violation features to the compliance database.

4. The big data based SMS sending optimization system of claim 1, wherein, The multi-version content generation module includes: User feature extraction unit: for obtaining user multi-dimensional data from big data platform, constructing user feature vector Wherein, is the historical behavior feature, the user feature vector adopts weighted fusion, and a user comprehensive feature score is obtained Wherein, is the feature weight and satisfies ; The content template library stores scenario-based content templates classified by industry scenarios, each template contains fixed text segments and variable placeholders and is associated with a speech style label; Generative AI engine: Lightweight generative model that takes in a user feature vector and generates a personalized content candidate Version diversity check: for the basic content output by the generative AI engine, generate initial versions by synonym replacement, sentence transformation, and redundant information addition and subtraction, and calculate the semantic similarity of the two initial versions The output value range is [01], and the target version is selected from the initial version The average similarity between versions meets , wherein represents the number of combinations of selecting 2 elements from elements in random order, is the diversity threshold, represents the th short message content version, represents the th short message content version, and when the screening result does not meet the threshold, the iteration number is automatically increased, N personalized short message content versions that have passed the diversity check are output, input into the compliance pre-check module for filtering, and after removing the illegal versions, the final test content set is generated , wherein and to ensure test effectiveness, represents the th short message content version that remains after filtering by the compliance pre-check module.

5. The big data based SMS sending optimization system of claim 4, wherein, The specific process of the generative AI engine to generate personalized content candidates is: Template matching: Score users based on their overall characteristics Match the templates with the applicable feature ranges of the template library and select the three candidate templates with the highest matching degree; Variable filling: based on entity linking technology, the user features are associated with template variables, the entity linking technology performs type analysis on template variable placeholders, establishes a variable type library, and assigns a unique ID to each variable type from the user feature vector The type-matched entity is extracted from the user historical behavior, the entity pool is constructed, and the semantic correlation degree between the entity and the variable is calculated, and the formula is: Wherein, is the candidate entity in the entity pool, is the template variable, is the BERT embedding vector of the entity / variable, is the frequency of the entity appearing in the user historical behavior, and the entity with the highest Rel value is selected to fill the variable; Style adjustment: according to the speech style label in the user label, the template text is adaptively rewritten to obtain personalized content candidates.

6. The big data based SMS sending optimization system of claim 1, wherein, The space-time test matrix construction module includes: time set generating unit: used for extracting user active time features from user historical behavior data, generating sending time set ; Matrix building unit: combines the final test content sets with the transmission time sets to build a test unit matrix where each unit represents a content at a time of transmission strategy; The output of the spatio-temporal test matrix construction module is a dynamically adjusted test cell matrix wherein each cell corresponding to one content-time combination strategy, is input to the dynamic traffic allocation module for sample allocation.

7. The big data based SMS sending optimization system of claim 1, wherein, The dynamic traffic distribution module includes a user grouping unit, a traffic mapping unit, and an equalization verification unit, which are used to distribute user groups to different units of the test unit matrix, wherein: The user grouping unit: according to the user portrait feature, the user feature vector The target user group is divided into K homogeneous groups according to the features with high correlation with the short message response rate, and the short message response rate correlation adopts Pearson correlation coefficient The screening feature, the internal homogeneity of each group and the inter-group heterogeneity are calculated, and the effectiveness of the grouping is verified, the internal homogeneity should meet the short message response rate variance of the users in the group ≤0.1, and the inter-group heterogeneity should meet the short message response rate mean difference between groups ≥0.2; Traffic mapping unit: map users per packet to the test unit matrix in a 5% proportion different units; The balance verification unit verifies the consistency of the sample feature distribution of each test unit with the target group, calculates the feature statistics of each test unit sample, evaluates the difference degree of the sample feature distribution with the target group, and obtains the difference degree by calculating the standard deviation of the test unit sample and the target group. When the difference degree is less than a preset threshold, the sample is considered to be consistent with the target group, and the sample is considered to be inconsistent with the target group when the difference degree is greater than the threshold. The flow distribution ratio is adjusted, and the sample is redistributed. The output of the dynamic flow distribution module is the evenly distributed test samples, and each test unit The corresponding set of homogeneous user samples is input into the policy decision feedback module for effect analysis.

8. The big data based SMS sending optimization system of claim 1, wherein, The policy decision feedback module includes: Packet response analysis unit: for k homogeneous user groups divided by dynamic traffic distribution module, respectively, statistics of each test unit in different groups of response data, the response data includes opening rate , conversion rate , the opening rate Indicates that in the kth user group, the proportion of users who open the short message after the content Is sent at time , the conversion rate Indicates that in the kth user group, the proportion of users who complete the target behavior after opening the short message, a three-dimensional response matrix of grouping-content-time is constructed, and the value of each cell in the matrix is ; Causal decoupling unit: separate the independent effects of content and time on user response, compute the combined effect score , for content independent effect , only consider content , for the kth group, exclude time factor interference, calculate when time , is fixed as a constant , statistics content , conversion rate in this period, that is , time independent effect , only consider time , for the kth group, exclude content factor interference, calculate when content , is fixed as the content with the highest response rate in the user history of this group , statistics conversion rate of the fixed content in time , the combined effect score ;​ Strategy filtering unit: For the k-th group, select all test units... Sort the test units from highest to lowest and select the top 10% as the optimal strategy. Feedback is then fed back to the multi-version content generation module for iterative optimization, guiding it to generate content containing this feature.

Citation Information

Patent Citations

  • Customized management evaluation system for short message service module

    CN117156466A

  • Short message content analysis method and system based on artificial intelligence

    CN118394945A

  • Short message information optimization processing method and system based on user operation feedback

    CN118764892A

  • Short message sending test method, system, equipment and medium

    CN119922499A