A method and system for generating intelligent SMS templates based on a large language AI model
By extracting a cultural background distribution matrix from social media and device logs, establishing a taboo word avoidance rule library, analyzing expression habits and symbolic semantics, and optimizing SMS template generation methods, the problem of insufficient cultural adaptation in cross-cultural SMS communication is solved, and the accuracy and acceptance of communication are improved.
Patent Information
- Application Number
- CN202510868904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing SMS template generation methods based on large language models are unable to flexibly respond to dynamic changes in the target audience's cultural background, resulting in insufficient cultural adaptation and affecting the accuracy and acceptability of cross-cultural SMS communication.
By extracting the cultural background distribution matrix from social media interaction data and device usage logs, a taboo word avoidance rule library is pre-established, expression habits preferences and symbolic semantics are analyzed, SMS templates are generated and optimized, acceptance is evaluated and templates are regenerated until the preset standards are met.
It significantly improves the accuracy and acceptance of cross-cultural SMS communication, enhances audience participation, and ensures the applicability and security of SMS templates in different cultural contexts.
Smart Images

Figure CN120373279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a method and system for generating intelligent SMS templates based on a large language AI model. Background Art
[0002] Intelligent SMS template generation technology based on large language models plays a crucial role in modern communications. Its core value lies in improving the efficiency and acceptance of information transmission. Especially in the context of increasingly frequent cross-cultural communication, cultural sensitivity has become a key indicator of technological excellence. With the acceleration of globalization, SMS templates must not only meet functional requirements but also adapt to diverse cultural contexts to ensure emotional resonance and smooth understanding among target audiences. However, research and application in this area still face significant challenges, particularly in the dynamic and accurate cultural adaptation. Breaking through existing bottlenecks is urgently needed to address complex scenarios. Current SMS template generation methods based on large language models often rely on static cultural rule bases or universal language patterns, making them inflexible to the dynamic changes in the target audience's cultural background. These methods often exhibit limited adaptability to regional cultural differences, such as ignoring specific cultural taboos, failing to accurately reflect expression preferences, or providing a superficial understanding of cultural symbols. This results in low acceptance of generated templates among certain cultural groups. Existing solutions are limited by the lack of adaptive cultural adjustment mechanisms, making it difficult to update relevant strategies in response to changes in cultural background distribution.
[0003] Therefore, how to dynamically adjust the cultural adaptation layer to optimize the retention ratio of culturally relevant expressions and localized translation strategies, and achieve adaptive updates in cultural taboo detection, expression habit preferences, and depth of symbolic understanding, has become a key issue in improving the overall cultural acceptance of SMS templates. Solving this problem will directly determine the actual effectiveness and wide applicability of the technology in global communications, but currently existing technologies cannot solve this problem. Summary of the Invention
[0004] To address the aforementioned issues in the prior art, the present invention aims to provide a method and system for generating intelligent SMS templates based on a large language AI model. This method can effectively improve the accuracy and acceptance of cross-cultural SMS communication and enhance audience engagement.
[0005] The method for generating an intelligent SMS template based on a large language AI model described in the present invention comprises the following steps:
[0006] S1. Extract cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs to generate a cultural background distribution matrix.
[0007] S2. Pre-establishing a taboo word avoidance rule library, identifying regional cultural characteristics in the cultural background distribution matrix, obtaining taboo matching results, and updating the cultural background distribution matrix to avoid taboo words;
[0008] S3. Based on the updated cultural background distribution matrix, extract the target audience's habitual adaptation of address and speaking speed and rhythm preferences, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table;
[0009] S4. Generating a preliminary SMS template based on the expression habit preference matrix and the symbol semantic mapping table, optimizing the preliminary SMS template, and outputting an adjusted template set that meets message length control requirements;
[0010] S5. Analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than a preset acceptance threshold, re-acquire the cross-cultural communication frequency and generate a new template set.
[0011] S6. Evaluate the text alignment of the new template set, combine the taboo word avoidance results and the symbol semantic mapping table, analyze the fit between the sentence preference analysis results and the word sentiment tendency, and obtain an information clarity score;
[0012] S7. Screen candidate SMS templates based on the information clarity score, optimize visual element embedding and interactive button layout, analyze the audience's engagement potential based on speech speed and rhythm preferences, and determine the target SMS template.
[0013] Preferably, the step S1 specifically includes:
[0014] Collect message data posted by users, perform word segmentation on the message data, extract language usage feature data, and generate a language sentence structure matrix based on the language usage feature data;
[0015] If the message data contains a mixture of languages, a classification algorithm is used to divide the content in different languages, and cross-cultural communication frequency data is extracted from the division results to generate a cultural interaction intensity matrix;
[0016] Obtain user click behavior records and ad page visit data, perform clustering operations to extract user interaction time series features, and determine user activity indicators;
[0017] Calculate the message text length distribution based on the language sentence structure matrix and generate message length control parameters based on user activity indicators;
[0018] A neural network is used to perform sentiment polarity analysis on message data, extract the sentiment expression feature vector, and combine it with the cultural interaction intensity matrix and message length control parameters to generate a cultural background distribution matrix that includes lexical sentiment tendency and message length control.
[0019] Preferably, the step S2 specifically includes:
[0020] Extracting regional cultural feature data from the cultural background distribution matrix to generate regional language distribution feature data, and processing the regional language distribution feature data using a neural network model to generate a regional cultural feature vector;
[0021] According to the regional cultural feature vector, a taboo word rule set of the corresponding region is obtained from a taboo word avoidance rule library, and taboo words in the input text are identified to generate a taboo word frequency statistical matrix;
[0022] If the taboo word frequency in the taboo word frequency statistical matrix exceeds a preset sensitivity threshold, a replacement word list is extracted from the taboo word avoidance rule library, an optimal replacement word is selected using a semantic similarity algorithm, and optimized text content is generated;
[0023] The cultural background distribution matrix is updated according to the word frequency distribution, language usage characteristics and cultural sensitivity score of the optimized text content.
[0024] Preferably, the step S3 specifically includes:
[0025] The target area is divided according to the updated cultural background distribution matrix. The speech segments in the target area are analyzed, the phoneme sequences and their duration and pause characteristics are extracted, the speech rate feature vector is generated, and a speech rate rhythm indicator matrix is constructed.
[0026] Classify the target region's address corpus, construct an address relationship map, and generate address adaptation vectors based on the closeness of the relationship between people;
[0027] By parsing the text collection of the target area, extracting sentence structure, rhetoric and modal particle features, and generating an expression habit preference matrix;
[0028] Convolutional neural networks are used to extract the semantic features of cultural symbols in text collections, build a regional symbol library, and generate a symbol semantic mapping table through semantic quantification.
[0029] Preferably, the step S4 specifically includes:
[0030] Extracting semantic features from the symbol semantic mapping table to generate a semantic vector, and combining the language style parameters in the expression habit preference matrix to generate an initial template text, and forming a preliminary SMS template group by segmenting and combining semantic units;
[0031] Construct rhetorical feature vectors based on rhetorical device statistics, calculate the applicability probability of rhetorical devices in different scenarios, optimize the rhetorical ratio and generate a rhetorical optimization template;
[0032] Obtain the display area size of the target user's terminal device, calculate the font size and paragraph spacing, generate typesetting parameters, and generate a display optimization template based on color contrast optimization;
[0033] If the template text exceeds the length limit, a semantic-preserving compression algorithm is used to control the length, and a template set is generated based on the layout and color optimization results.
[0034] Preferably, the step S5 specifically includes:
[0035] Obtain the adjusted template set and analyze its language conciseness, calculating the number of characters in a single sentence, the sentence nesting level, and the proportion of modifiers to generate a language conciseness vector.
[0036] Extract the size, position, and spacing parameters of the button layout in the template set, and calculate the layout adaptation score based on the interface design specifications;
[0037] The comprehensive acceptance evaluation value is calculated by combining the language conciseness vector and the layout adaptation score. If the acceptance is lower than the preset acceptance threshold, the cultural background distribution matrix is updated by analyzing the user conversation records and a new template set is generated.
[0038] Preferably, the step S6 specifically includes:
[0039] Obtain the text alignment and spacing features of the new template set, calculate the adaptation ratio of line spacing to font size, generate typesetting reference vectors, quantify the font display effect, and generate a font display parameter table;
[0040] Check the new template content according to the taboo word avoidance rules and the symbol semantic mapping table to generate text compliance data;
[0041] Obtain the language structure features of the new template, generate sentence feature vectors, and use deep neural networks to analyze the sentiment attributes of words to generate a sentiment tendency mapping matrix;
[0042] A weighted calculation is performed on the font display parameters, the text compliance data, and the sentiment tendency mapping matrix to generate a template candidate scoring table, and high-scoring templates are screened for information integrity verification to generate an information clarity score.
[0043] Preferably, the step S7 specifically includes:
[0044] sorting candidate SMS templates according to the information clarity scores, extracting visual layout attributes of high-scoring templates to generate visual structure vectors, and optimizing the layout of decorative icons to generate a visual element layout diagram;
[0045] Allocate space for button positions, calculate the ratio of trigger area to screen size, set safe spacing, and generate a button layout specification table;
[0046] Collect user reading speed and scrolling frequency data, construct a reading habit feature vector, extract speech rhythm features to generate a speech rate preference matrix, and optimize the text rhythm to generate a rhythm adaptation vector;
[0047] The visual element layout diagram, the button layout specification table and the rhythm adaptation vector are comprehensively scored, the weight distribution is calculated, the template comprehensive score is generated, and the highest-scoring template is selected as the target SMS template.
[0048] The present invention also proposes an intelligent SMS template generation system based on a large language AI model, comprising:
[0049] The cross-cultural communication analysis module extracts cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs, generating a cultural background distribution matrix that includes lexical sentiment tendencies and message length controls;
[0050] The taboo word avoidance module is used to pre-establish a taboo word avoidance rule library, identify regional cultural characteristics in the cultural background distribution matrix, scan taboo words based on regional cultural characteristics, obtain taboo matching results, and update the cultural background distribution matrix based on the matching degree and frequency of occurrence to avoid taboo words;
[0051] The expression habit analysis module is used to extract the target audience's habitual adaptation and speaking speed and rhythm preferences based on the updated cultural background distribution matrix, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table;
[0052] The SMS template generation module is used to generate preliminary SMS templates based on the symbol semantic mapping table and the expression habit preference matrix, optimize the proportion of rhetorical devices selected in the preliminary SMS templates, adjust text color contrast and layout spacing based on screen size adaptation data, and output a set of adjusted templates that meet the message length control;
[0053] The template optimization module is used to analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than the preset acceptance threshold, the cross-cultural communication frequency is re-acquired, the cultural background distribution matrix is updated, and a new template set is generated;
[0054] An acceptance evaluation module, used to evaluate the text alignment of the new template set, combine the results of taboo word avoidance and the symbol semantic mapping table, screen candidate SMS templates, analyze the fit between the sentence pattern preference analysis result and the lexical sentiment tendency, and obtain an information clarity score;
[0055] An SMS template screening module, used to screen candidate SMS templates according to the information clarity score, optimize the embedding of visual elements and the layout of interactive buttons, analyze the engagement potential of the audience under the preference of speech rate and rhythm, and determine the target SMS template.
[0056] A method and system for generating intelligent SMS templates based on a large language AI model according to the present invention has the following advantages:
[0057] The method for generating an intelligent SMS template based on a large language AI model of the present invention generates a cultural background distribution matrix by extracting cross-cultural communication frequencies, sentence pattern preferences, and screen size adaptation data from social media interaction data and device usage logs, and can accurately grasp the cultural characteristics, language habits, and device adaptation requirements of the target audience, providing comprehensive and practical basic data for subsequent template generation; using a pre-established taboo word avoidance rule library to identify regional cultural characteristics in the cultural background distribution matrix and avoid taboo words can effectively avoid misunderstandings or offenses caused by cultural differences, and improve the applicability and security of SMS templates in different cultural backgrounds; according to the updated cultural background distribution matrix, extracting appellation habit adaptation and speech rate and rhythm preferences to generate an expression habit preference matrix, and analyzing the deep meaning of cultural symbol semantics to obtain a symbol semantic mapping table can make the SMS template more conform to the expression habits and cultural understanding of the target audience in terms of language expression and symbol use, enhancing the accuracy of information transmission; generating a preliminary SMS template based on the expression habit preference matrix and the symbol semantic mapping table, and optimizing the output to adjust the template set that meets the message length control can ensure that the SMS template meets the actual usage requirements in terms of content expression and format specification, and avoid affecting the reading experience due to long information or improper format; analyzing the acceptance of the adjusted template set in terms of language conciseness and interactive button layout, and if it does not meet the standard, re-obtaining data to generate a new template set helps to continuously optimize the template quality, improve the acceptance degree and usage satisfaction of users for the SMS template; evaluating the text alignment of the new template set, combining various factors to analyze the fit between the sentence pattern preference and the lexical sentiment tendency to obtain an information clarity score can comprehensively measure the information transmission effect of the SMS template from multiple dimensions and provide a quantitative basis for template screening; screening candidate templates according to the information clarity score, optimizing the visual elements and interactive button layout, and analyzing the audience engagement potential to determine the target SMS template can generate high-quality SMS templates that perform well in terms of content, vision, and interactive experience, and meet the usage habits and needs of the audience. The intelligent SMS template generation method can effectively improve the accuracy and acceptance of cross-cultural SMS communication and enhance the audience engagement.
[0058] The intelligent SMS template generation system based on the large language AI model of the present invention systematically and comprehensively generates intelligent SMS templates that are in line with audience characteristics, avoid cultural risks, have high-quality content, and have a good interactive experience through a series of processes such as data collection, cultural risk avoidance, expression habit adaptation, template generation optimization, acceptance evaluation and screening, thereby significantly improving the professionalism and practicality of SMS template generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a flow chart of a method for generating intelligent SMS templates based on a large language AI model described in the present invention. DETAILED DESCRIPTION
[0060] like Figure 1 As shown, the method for generating an intelligent SMS template based on a large language AI model according to the present invention includes the following steps:
[0061] S1. Extract cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs to generate a cultural background distribution matrix.
[0062] S2. Pre-establish a taboo word avoidance rule library, identify regional cultural characteristics in the cultural background distribution matrix, obtain taboo matching results, and update the cultural background distribution matrix to avoid taboo words;
[0063] S3. Based on the updated cultural background distribution matrix, extract the target audience's habitual adaptation of address and speaking speed and rhythm preferences, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table;
[0064] S4. Generating a preliminary SMS template based on the expression habit preference matrix and the symbol semantic mapping table, optimizing the preliminary SMS template, and outputting an adjusted template set that meets the message length control;
[0065] S5. Analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than a preset acceptance threshold, re-acquire the cross-cultural communication frequency and generate a new template set.
[0066] S6. Evaluate the text alignment of the new template set, combine the taboo word avoidance results and the symbol semantic mapping table, analyze the fit between the sentence preference analysis results and the word sentiment tendency, and obtain the information clarity score;
[0067] S7. Filter candidate SMS templates based on message clarity scores, optimize visual element embedding and interactive button layout, analyze the audience's engagement potential based on speech speed and rhythm preferences, and determine the target SMS template.
[0068] Furthermore, in this embodiment, step S1 specifically includes:
[0069] Collect message data posted by users, perform word segmentation on the message data, extract language usage feature data, and generate a language sentence structure matrix based on the language usage feature data;
[0070] If the message data contains mixed use of multiple languages, a classification algorithm is used to divide the content in different languages, and the cross-cultural communication frequency data is extracted from the division results to generate a cultural interaction intensity matrix;
[0071] Obtain user click behavior records and ad page visit data, perform clustering operations to extract user interaction time series features, and determine user activity indicators;
[0072] Calculate the message text length distribution based on the language sentence structure matrix and generate message length control parameters based on user activity indicators;
[0073] A neural network is used to analyze the sentiment polarity of message data, extracting the sentiment expression feature vector. This is combined with the cultural interaction intensity matrix and message length control parameters to generate a cultural background distribution matrix that includes lexical sentiment tendency and message length control.
[0074] Specifically, user message data is captured through the social media application program interface, the message data is segmented, language usage feature data is obtained from the segmentation results, and a language sentence structure matrix is generated based on the language usage feature data;
[0075] If multiple languages are mixed in the message data, a random forest classifier is used to classify the content in different languages. The classifier input features include word frequency statistical vectors, language identifiers, and character encoding types. The cross-cultural communication frequency data is obtained from the classification results, and a cultural interaction intensity matrix is generated based on the cross-cultural communication frequency data.
[0076] Obtain user device screen resolution parameters and application usage time records through mobile terminal applications, perform clustering operations on the usage time records according to time periods, obtain device interaction time series data based on the clustering results, and extract user activity indicators from the device interaction time series data;
[0077] Calculate the message text length distribution based on the language sentence structure matrix, adaptively adjust the message length based on the user activity index, and generate the message length control parameter;
[0078] A convolutional neural network is used to identify the sentiment polarity of message data. The network input layer receives a sequence of word vectors, extracts sentiment features through three convolutional layers, and obtains the sentiment expression feature vector from the sentiment polarity identification results.
[0079] Generate a multidimensional cultural background distribution matrix based on the cultural interaction intensity matrix, message length control parameters, and emotional expression feature vectors. The matrix dimensions include region, language, emotion, and interaction frequency.
[0080] Here is an example:
[0081] Based on basic style specifications, fill in or modify elements in the style to serve as the final display template, such as a red envelope template or a multi-image and text template. This function can be integrated into the software system of the communication platform or implemented as a standalone application. The specific implementation method can be determined according to the actual scenario.
[0082] This embodiment does not impose too many restrictions on the activation method of the generation function. The user can activate the function through interface operation or trigger it through remote commands. For example, the communication platform can set a template generation button on the message editing interface, and the user can activate the function by clicking the button, or send a command through a third-party device to remotely activate the generation function.
[0083] After the function is launched, the target audience's social media interaction data and device usage logs are collected to determine the cultural adaptation basis for the generated template. Social media interaction data includes the text content posted by users, interaction frequency and timestamps. Device usage logs include screen resolution, application usage time and interaction timing.
[0084] By analyzing the above data, we extract cross-cultural communication frequency, sentence preference, and screen adaptation information to generate a cultural background distribution matrix. This embodiment does not impose too many restrictions on the specific method of data collection, which can be set by technical personnel according to actual needs.
[0085] Specifically, the user-posted message data is collected in real time through the API interface of the social media platform, for example, 1,000 messages containing text content, publishers, and timestamps are collected;
[0086] Word segmentation uses a word segmentation tool to segment messages, generate a word list, and extract language usage characteristics, including word frequency, sentence length, and lexical diversity. For example, analysis results show that each message contains an average of 12 words, and common words include "thank you" and "share";
[0087] Based on the above features, a language sentence structure matrix is generated to record the distribution of sentence types. For example, declarative sentences account for 65% and interrogative sentences account for 25%. This matrix provides a basis for sentence preference for subsequent cultural adaptation.
[0088] When the message data involves mixed use of multiple languages, a random forest classifier is used to classify the language content;
[0089] The input features of the classifier include word frequency statistics vectors, language identifiers, and character encoding types. For example, after analyzing 1,000 messages, the classifier output shows that 800 are mainly Chinese and 150 are a mixture of Chinese and English.
[0090] Cross-cultural communication frequency data is extracted from the segmentation results. The number of multilingual message interactions is counted to generate a cultural interaction intensity matrix. This matrix records the frequency of interactions between users of different languages. For example, the number of daily interactions between Chinese and English users is 40 times, reflecting the activeness of cross-cultural communication.
[0091] Tracking points are set up during the template creation process via mobile terminal applications (i.e., embedding tracking statistics SDK) to collect user click behavior data, including click location, click frequency, and click time. At the same time, access behaviors such as ad page loading, exposure, clicks, and dwell time can also be monitored to comprehensively evaluate user interaction with the application interface.
[0092] For example, when collecting operational data from 100 users, it was found that 60 of them clicked on interface elements at least five times a day, visited ad pages at least once, or spent more than 30 seconds on them. Cluster analysis of the temporal distribution of user click behavior divided user activity into morning, noon, and evening hours. The results showed that 65% of users performed high-frequency click operations primarily in the evening.
[0093] Based on the above data, we extract the temporal features of user interactions and define the following user activity indicators:
[0094] The number of effective clicks per day should be no less than 5;
[0095] Visit the ad page at least once a day, or stay on the ad page for more than 30 seconds in total;
[0096] Users who meet one of the above two conditions are considered active users, accounting for 60% of the total. This activity indicator can be used to optimize message push strategies, interface layout design, and advertising effectiveness evaluation;
[0097] Analyze the language sentence structure matrix and calculate the message text length distribution. For example, statistics show that 75% of messages are between 8 and 18 words long.
[0098] Based on user activity indicators, the message length is adaptively adjusted and a message length control parameter is generated. It is recommended that the template length be controlled within 15 words to improve reading efficiency. This parameter ensures the display effect and user acceptance of the template on different devices.
[0099] A convolutional neural network is used to analyze the sentiment polarity of message data. The network inputs a sequence of word vectors and extracts sentiment features through multi-layer convolution. For example, after analyzing 1,000 messages, 55% were positive and 35% were neutral.
[0100] The emotion expression feature vector records the emotional intensity of each message, for example, a positive emotion score of 0.75. Combining the cultural interaction intensity matrix, message length control parameters, and the emotion expression feature vector, a multidimensional cultural background distribution matrix is generated. This matrix includes four dimensions: region, language, emotion, and interaction frequency. For example, it shows that users in a certain region mainly speak Chinese, have a high proportion of positive emotions, and have a daily interaction frequency of 90 times, reflecting their social activity and emotional tendencies.
[0101] In this embodiment, the generation of the cultural background distribution matrix can dynamically capture the cultural characteristics of the target audience, provide data support for the subsequent avoidance of taboo words and adaptation of expression habits, the message length control parameter optimizes the reading efficiency of the template, the emotional expression feature vector enhances the emotional resonance of the template, and the cultural interaction intensity matrix improves the accuracy of cross-cultural communication.
[0102] Furthermore, in this embodiment, step S2 specifically includes:
[0103] Extracting regional cultural feature data from the cultural background distribution matrix to generate regional language distribution feature data, and using a neural network model to process the regional language distribution feature data to generate a regional cultural feature vector;
[0104] According to the regional cultural feature vector, the taboo word rule set of the corresponding region is obtained from the taboo word avoidance rule library, and the taboo words in the input text are identified to generate a taboo word frequency statistical matrix;
[0105] If the frequency of taboo words in the taboo word frequency statistical matrix exceeds the preset sensitivity threshold, a replacement word list is extracted from the taboo word avoidance rule library, and the semantic similarity algorithm is used to select the optimal replacement word, and the optimized text content is generated;
[0106] Update the cultural background distribution matrix based on the word frequency distribution, language usage characteristics, and cultural sensitivity scores of the optimized text content;
[0107] Specifically, a taboo word avoidance rule base is established based on regional cultural characteristic data. The rule base includes word type identification, usage scenario restrictions, and a replacement suggestion word list. Regional language usage characteristics are extracted from the cultural background distribution matrix to generate regional language distribution characteristic data.
[0108] The target region's cultural characteristics are identified using regional language distribution data. A deep neural network is used to calculate the cultural sensitivity score of each region. The network input layer includes three dimensions: language usage frequency, cultural differences, and regional distribution characteristics. The output layer generates a regional cultural feature vector.
[0109] Obtain the taboo word rule set for the corresponding region from the avoidance rule library based on the regional cultural feature vector. Use a text scanning tool to identify taboo words in the input text. Use a word frequency statistics method to calculate the frequency of taboo words and generate a taboo word frequency statistics matrix.
[0110] Filter the taboo word frequency statistical matrix according to the preset sensitivity threshold to generate a list of high-frequency taboo words, extract the corresponding replacement word list from the avoidance rule library, and select the optimal replacement word based on the semantic similarity algorithm;
[0111] The natural language processor replaces taboo words in the input text. The replacement operation selects words based on contextual semantic relevance to generate optimized text content.
[0112] The cultural background distribution matrix is updated based on the word frequency distribution, language usage characteristics, and cultural sensitivity scores in the optimized text content. The updated content includes data on three dimensions: regional language usage frequency, cultural differences, and the probability of sensitive word occurrence.
[0113] Here is an example:
[0114] In this embodiment, when generating a culturally adapted SMS template, it is necessary to ensure that the input text content conforms to the cultural norms of the target audience;
[0115] Therefore, we use the pre-built taboo word avoidance rule library and combine it with the regional cultural characteristics in the cultural background distribution matrix to scan and optimize the taboo words in the input text.
[0116] The taboo word avoidance rule library includes word sensitivity classification, usage scenario restrictions, and replacement suggestions. It aims to identify words that may cause cultural conflicts and provide appropriate alternatives.
[0117] By analyzing the cultural background distribution matrix, regional cultural characteristics, such as language usage habits and cultural sensitivities, are extracted to generate taboo matching results. The matrix is then updated based on the matching degree and word frequency distribution. This embodiment does not impose excessive restrictions on the specific algorithm for identifying taboo words, and technicians can choose an appropriate implementation method based on actual scenarios.
[0118] Specifically, regional cultural characteristic data is extracted from the cultural background distribution matrix, including language usage frequency, cultural differences, and regional distribution characteristics. For example, analysis shows that users in region A prefer short sentences, with an average sentence length of 14 characters, indicating a high degree of cultural differences.
[0119] Generate regional language distribution feature data, record common vocabulary and sentence patterns, and process this data using a deep neural network. The network input layer receives language usage frequency, cultural differences, and regional distribution characteristics, and the output layer generates a regional cultural feature vector. For example, for users in region A, the feature vector may reflect a mixture of English and the local language, with a cultural sensitivity score of 0.82, indicating that culturally sensitive words should be avoided. This vector provides an accurate basis for subsequent taboo word screening.
[0120] Using regional cultural feature vectors, we extract the taboo word rule set for the corresponding region from the taboo word avoidance rule library;
[0121] The rule set includes word type identifiers, such as "sensitive" or "taboo", and usage scenario restrictions, such as "forbidden for public messaging";
[0122] The text scanning tool scans the input text word by word to identify taboo words and their context. For example, after scanning 1,000 input texts, it was found that the frequency of "a certain event" in region A was 4 times per 1,000 words.
[0123] Using word frequency statistics, we generate a taboo word frequency matrix to record the location, frequency, and context of taboo words, ensuring the accuracy of subsequent replacement operations. This matrix provides data support for sensitive word management.
[0124] Filter the taboo word frequency statistical matrix, set the sensitivity threshold to 0.65, and only retain taboo words with a frequency higher than the threshold to generate a high-frequency taboo word list. For example, the list shows that the frequency of "conflict" in texts in a certain region is 6 times per thousand words;
[0125] Extract the replacement word list from the rule base and select the optimal replacement word based on the semantic similarity algorithm, for example, replacing "conflict" with "disagreement";
[0126] The semantic similarity algorithm calculates the cosine distance between word vectors to ensure that the replacement word is semantically close to the original word and fits the context;
[0127] The natural language processor performs replacement operations to generate optimized text content. For example, if a user inputs "the event caused a conflict," the replacement will generate "the event caused a disagreement," maintaining sentence fluency and semantic coherence.
[0128] Analyze the word frequency distribution and language usage characteristics of the optimized text content and calculate the cultural sensitivity score. For example, the probability of sensitive words appearing in the replaced text drops from 4% to 0.8%, indicating that the circumvention rules are effective.
[0129] Based on the above data, the cultural background distribution matrix was updated. The updated content includes regional language usage frequency, cultural differences, and the probability of sensitive word occurrence. For example, the matrix shows that the probability of sensitive words in a certain region has decreased, the language usage frequency has tended to be simpler, and the cultural differences have been slightly adjusted.
[0130] The updated matrix provides a dynamic reference for subsequent template generation, ensuring that content continues to adapt to the cultural needs of the target audience;
[0131] This embodiment significantly improves the cultural sensitivity and acceptability of SMS templates by dynamically updating the matrix, and reduces the risk of misunderstanding caused by cultural differences.
[0132] Furthermore, in this embodiment, step S3 specifically includes:
[0133] The target area is divided according to the updated cultural background distribution matrix. The speech segments in the target area are analyzed, the phoneme sequences and their duration and pause characteristics are extracted, the speech rate feature vector is generated, and a speech rate rhythm indicator matrix is constructed.
[0134] Classify the target region's address corpus, construct an address relationship map, and generate address adaptation vectors based on the closeness of the relationship between people;
[0135] By parsing the text collection of the target area, extracting sentence structure, rhetoric and modal particle features, and generating an expression habit preference matrix;
[0136] A convolutional neural network is used to extract the semantic features of cultural symbols in text collections, build a regional symbol library, and generate a symbol semantic mapping table through semantic quantification;
[0137] Specifically, the target area is divided according to the cultural background distribution matrix, the phoneme sequence in the speech segment is extracted through the speech recognizer, the duration and pause position of the phoneme sequence are calculated, and the speech rate feature vector is generated;
[0138] The speech segments are clustered into rhythmic patterns based on the speech rate feature vectors, and a speech rate and rhythm index matrix is constructed through the cluster centers. The matrix dimensions include speech rate, pause ratio, and syllable duration.
[0139] Support vector machine is used to classify the appellation corpus in the target area. The feature vector includes the appellation part-of-speech tag, contextual relationship identifier and usage scenario tag. The appellation relationship map is constructed from the classification results.
[0140] Calculate the closeness of relationships based on the title relationship graph, build a set of title usage rules, and generate a title adaptation vector from the rule set. The vector dimensions include identity level, relationship distance, and formality of the occasion.
[0141] Parse the text in the target area through a natural language processor, extract language expression features, including sentence structure, rhetorical devices, and the use of modal particles, and generate an expression habit preference matrix;
[0142] Use a convolutional neural network to identify cultural symbols in the text. The network input is the word segmentation result and the词性标注序列 (词性 tagging sequence), extract the context semantic features of the symbols, and generate a cultural symbol feature vector;
[0143] Construct a regional symbol library based on the cultural symbol feature vector, perform semantic quantization calculation on the symbols, and the calculation dimensions include cultural attributes, emotional tendencies, and usage frequencies, and generate a symbol semantic mapping table;
[0144] Examples are as follows: [[ID=IO]]
[0145] In this embodiment, use the updated cultural background distribution matrix to analyze the language expression habits and cultural symbol preferences of the target audience, and ensure that the generated SMS template conforms to the regional cultural norms in terms of appellation selection, speech rate rhythm, and symbol use;
[0146] The cultural background distribution matrix divides the cultural boundaries of the target area through dimensions such as language usage frequency and cultural sensitivity;
[0147] Extract appellation habits, speech rate rhythm preferences, and cultural symbol semantic features, generate an expression habit preference matrix and a symbol semantic mapping table. These data structures provide a precise cultural adaptation basis for template generation, significantly improving user acceptance and interaction experience. This embodiment does not overly limit the specific analysis method, and a suitable implementation method can be selected according to the actual scenario;
[0148] Specifically, collect a set of voice segments from the target area, such as 1000 voice messages uploaded by users. Each voice contains expressions with a specific cultural background;
[0149] The speech recognizer decomposes the voice into a phoneme sequence. For example, decompose "您好" into phonemes / nin / and / hao / , calculate that / nin / lasts for 0.35 seconds, / hao / lasts for 0.45 seconds, and the pause is 0.15 seconds, and generate a speech rate feature vector [0.35, 0.45, 0.15]. This vector reflects that the speech rate is relatively slow and is suitable for formal scenarios;
[0150] Use the K-means clustering algorithm to cluster the speech rate feature vectors into two categories: fast rhythm and slow rhythm;
[0151] The clustering center generates a speech rate rhythm index matrix, including a speech rate of 1.8 phonemes per second, a pause ratio of 18%, and a syllable duration of 0.4 seconds. This matrix provides a basis for rhythm optimization in the template voice playback scenario. For example, ensure that a fast rhythm template is used in a fast communication scenario;
[0152] It should be noted that the "词性标注序列" in item is a literal translation. In the field of natural language processing, there is a more accurate term for it, such as "词性标注结果" or "词性标注信息". You can adjust it according to the actual situation.Collect the appellation corpus of the target area, such as a set of 5000 appellation data labeled with word nature, context relationship and usage scenario;
[0153] The support vector machine uses the appellation word nature, context identifier and scenario marker as input features to classify the usage context of appellations. For example, "teacher" is labeled as respectful in an educational scenario and affectionate in a family scenario;
[0154] The classification result generates an appellation relationship graph, showing that the relationship between "teacher - student" is highly intimate and suitable for formal greetings;
[0155] Calculate the degree of intimacy of the relationship between people according to the graph, and generate an appellation adaptation vector, which includes identity level, relationship distance and formality of the occasion. For example, the vector of the appellation "friend" in a leisure scenario is [1, 0.9, 0.3], indicating equal identity, close relationship and informal occasion. This vector guides the selection of appropriate appellations when generating templates, enhancing the user's sense of intimacy;
[0156] Analyze the text corpus of the target area, such as 10000 social media texts, and extract features of sentence structure, rhetorical devices and usage of modal particles;
[0157] The natural language processor recognizes that texts in region D mostly start with "May I ask", the usage rate of the modal particle "la" is 12%, prefer short sentences, and the average sentence length is 10 characters;
[0158] Generate a preference matrix for expression habits, record the distribution of sentence types, such as declarative sentences accounting for 60% and interrogative sentences accounting for 30%, and rhetorical device preferences, such as the usage rate of metaphor being 8%. This matrix provides a basis for generating templates that conform to the regional expression habits. For example, in region A, use an affectionate tone and concise sentences first to ensure that the template content is close to the user's daily communication style;
[0159] Use a convolutional neural network to process the word segmentation results and word nature annotation sequences of the text corpus, and extract the context semantic features of cultural symbols. For example, the symbol "dragon" is associated with "authority" in the texts of region A, and the network generates a feature vector [authority, positive, 0.12], indicating its positive emotional tendency and usage frequency of 12%;
[0160] Construct a regional symbol library according to the feature vector, including the cultural attributes and semantic backgrounds of the symbols. The symbol library generates a symbol semantic mapping table through semantic quantization, and the quantization dimensions include cultural attributes, emotional tendency and usage frequency. For example, the cultural attribute of the "dragon" symbol is traditional, the emotional tendency is positive, and the usage frequency is 0.12;
[0161] The mapping table provides a basis for symbol selection in template generation. For example, in a formal publicity scenario, "dragon" is preferred to enhance cultural resonance, while avoiding sensitive symbols to reduce the risk of cultural conflicts.
[0162] Furthermore, in this embodiment, step S4 specifically includes:
[0163] Extract semantic features from the symbolic semantic mapping table to generate semantic vectors. Combined with the language style parameters in the expression habit preference matrix, the initial template text is generated. The initial SMS template group is formed by segmenting and combining semantic units.
[0164] Construct rhetorical feature vectors based on rhetorical device statistics, calculate the applicability probability of rhetorical devices in different scenarios, optimize the rhetorical ratio and generate a rhetorical optimization template;
[0165] Obtain the display area size of the target user's terminal device, calculate the font size and paragraph spacing, generate typesetting parameters, and generate a display optimization template based on color contrast optimization;
[0166] If the template text exceeds the length limit, a semantic-preserving compression algorithm is used to control the length, and a template set is generated based on the layout and color optimization results;
[0167] Specifically, a text generator is used to extract semantic feature tags based on the symbolic semantic mapping table to generate a fixed-length semantic vector. The language style parameters of the target region are obtained through the expression habit preference matrix to generate the initial template text.
[0168] The initial template text is divided into semantic units, the number of characters and expression complexity of each semantic unit are calculated, and the semantic units are combined according to the preset message length threshold to generate a preliminary SMS template group;
[0169] Based on the statistical data of rhetorical devices, a rhetorical feature vector is established. A support vector machine is used to calculate the applicability probability of various rhetorical devices in different scenarios. The combination of rhetorical devices with the highest applicability probability is selected to generate a rhetorical optimization template.
[0170] The screen resolution parameters and display area size are read through the text typesetting processor, the optimal font size is calculated according to the text display area, and a font display parameter table is generated;
[0171] According to the font display parameter table, the rhetoric optimization template is segmented into paragraphs, the minimum display spacing between paragraphs is calculated, and the paragraph layout vector is generated;
[0172] The color mapping component is used to calculate the brightness contrast value between the foreground color and the background color of the text, and a color scheme is selected according to the preset visual comfort threshold to generate a color mapping table.
[0173] The template layout is optimized based on the paragraph layout vector and color mapping table. The text that exceeds the length limit is processed using a semantic-preserving compression algorithm to generate a template set that meets the length limit.
[0174] The semantic symbol mapping table is used to store the correspondence between symbols and their semantic features, and the expression habit preference matrix reflects the language style parameters of the target area;
[0175] Furthermore, a text processor is used to scan the rhetoric devices in the preliminary SMS template, and the frequency of occurrence of each type of rhetoric devices in the text is counted to generate a rhetoric device frequency vector;
[0176] The target audience's rhetorical preference parameters are extracted based on the cultural background distribution matrix, and the random forest algorithm is used to optimize the rhetorical device frequency vector to generate a rhetorical device target ratio table;
[0177] The text content is reorganized through the rhetorical device mapper, and the usage ratio of various rhetoric devices is adjusted according to the rhetorical device target ratio table to generate rhetorically optimized text;
[0178] Collect the target audience's terminal screen resolution data through mobile device applications, extract the display area size from the device parameter library, and generate a screen specification parameter table;
[0179] Calculate the optimal font size for different screen sizes based on text display standards, set the text size and line spacing ratio, and generate a typesetting benchmark table;
[0180] The color mapping component is used to calculate the brightness difference between the foreground and background colors of the text, and the best color combination is selected according to the preset visual threshold to generate a color contrast matrix;
[0181] Adjust the format of the rhetorically optimized text according to the typesetting benchmark table and color contrast matrix, calculate the paragraph spacing adaptively according to the display area, and generate a display-optimized template;
[0182] The length of the display optimization template is controlled by a semantic compression algorithm, the core semantic information is retained, and an adjusted template set is generated;
[0183] Here is an example:
[0184] In this embodiment, a preliminary SMS template is generated using a symbol semantic mapping table and an expression habit preference matrix to ensure that the template content fits the target audience's language style and cultural symbol preferences;
[0185] By analyzing the applicability of rhetoric, the display characteristics of device screens, and message length limitations, the template is optimized in multiple dimensions.
[0186] The symbolic semantic mapping table records the correspondence between cultural symbols and their semantic features. For example, "lotus" is associated with "purity" in region D, and its emotional tendency is positive;
[0187] The expression habit preference matrix reflects the language style of the target region. For example, region D prefers short sentences and friendly particles.
[0188] By optimizing rhetoric, adjusting layout, and adapting colors, we generate template sets that meet cultural backgrounds and device display requirements, improving user reading experience and information delivery efficiency.
[0189] This embodiment does not impose too many restrictions on the specific optimization algorithm, and technicians can choose a suitable implementation method according to the actual scenario;
[0190] Specifically, a text generator is used to extract semantic features from the symbol semantic mapping table. For example, the semantic vector of "lotus" [purity, positive, 0.18] is extracted to represent its emotional tendency and usage frequency.
[0191] Combined with the language style parameters in the expression habit preference matrix, for example, Region A prefers a sentence length of 10 characters and the modal particle "la," an initial template text is generated, such as "The lotus is in full bloom, may you remain as pure as before."
[0192] To ensure the template length is appropriate, the initial text is segmented into semantic units, such as "Lotus blossoms" and "May you be as pure as before," which contain 4 and 9 characters respectively and have low expression complexity.
[0193] Based on the preset message length threshold of 12 characters, semantic units are combined to generate a preliminary SMS template group, such as "Lotus blossoms, pure as ever." This template group retains the core semantics, is suitable for SMS scenarios, and improves user acceptance.
[0194] By analyzing the corpus, we generated statistical data on rhetorical devices. For example, in the blessing scenes in region C, metaphors accounted for 60%, personification 25%, and parallelism 15%. Based on this data, we constructed a rhetorical feature vector [0.6, 0.25, 0.15].
[0195] A support vector machine was used to calculate the applicability of various rhetoric figures using scene labels and cultural preferences as input. For example, the probability of metaphor in blessing scenes was 0.68.
[0196] Select the rhetorical combination with the highest probability, such as metaphor 50%, personification 30%, and parallelism 20%, and generate a rhetorical optimization template;
[0197] The original template, "Lotus blossoms, pure as ever," was adjusted to "Heart like a lotus, fragrance dancing, may you be pure," incorporating personification and parallelism to enhance the expressive appeal. This optimization ensures the template aligns with regional cultural preferences and enhances emotional resonance.
[0198] Obtain the display area size of the target user's terminal device, combine it with the display parameters provided by different mobile phone manufacturers, including visible width, default font size, and system-level UI scaling information, calculate the appropriate font size and paragraph spacing, and generate layout parameters adapted to the target user's terminal device;
[0199] Typesetting parameters are used to intelligently segment and adjust the layout of rhetorically optimized templates to improve reading comfort and visual consistency. For example, by collecting display parameters of the devices used by the target user group, it was found that 80% of user devices come from manufacturers A and B. Their system default font sizes are 14 pixels and 16 pixels, respectively, and the screen viewable width accounts for approximately 85%;
[0200] Based on the above parameter differences, the text layout processor calculates the adapted font sizes as 15 pixels and 17 pixels respectively, and uniformly sets the line spacing ratio to 1.4, thereby generating a differentiated typesetting benchmark table;
[0201] The rhetoric optimization template is divided into paragraphs and the paragraph spacing is calculated to be 9 pixels to ensure clear readability on small screen devices;
[0202] The color mapping component calculates the brightness contrast between the text foreground color and the background color. For example, if the contrast between a white foreground color and a dark blue background color reaches the visual comfort threshold of 0.7, a color contrast matrix is generated.
[0203] Generate display-optimized templates based on the layout benchmark table and color matrix. For example, adjust the template text to high-contrast colors and compact layout to adapt to the display requirements of various devices and enhance the user's visual experience.
[0204] For templates that exceed the length limit, such as "Heart is like a lotus, its fragrance dances, may you be pure and happy every day," a semantic-preserving compression algorithm is used to analyze the weights of semantic units, retain the core information, and generate the compressed template "Heart is like a lotus, may you be pure and happy."
[0205] The algorithm ensures the semantic integrity of the compressed text by calculating the similarity of word vectors and context relevance, and combines typesetting parameters and color contrast optimization to generate a final template set that is adapted to different devices and scenarios. For example, the template set contains blessing text messages with high contrast and simple typesetting, and the length is controlled within 12 characters to meet the needs of fast reading. This set significantly improves the applicability of templates in cross-cultural and multi-device scenarios, and enhances user interaction efficiency.
[0206] Furthermore, in this embodiment, step S5 specifically includes:
[0207] Obtain the adjusted template set and analyze its language conciseness, calculating the number of characters in a single sentence, the sentence nesting level, and the proportion of modifiers to generate a language conciseness vector.
[0208] Extract the size, position, and spacing parameters of the button layout in the template set, and calculate the layout adaptation score based on the interface design specifications;
[0209] The language simplicity vector and layout adaptation score are combined to calculate the comprehensive acceptance evaluation value. If the acceptance is lower than the preset acceptance threshold, the cultural background distribution matrix is updated by analyzing the user conversation records and a new template set is generated.
[0210] Specifically, a text evaluator is used to measure the language conciseness of the adjusted template set, calculating three indicators: the number of characters in a single sentence, the sentence nesting level, and the proportion of modifiers, to generate a language conciseness vector;
[0211] The interactive interface analyzer extracts button layout parameters, including button size, position coordinates, and spacing ratio, and calculates the layout adaptation score according to the interface layout specifications;
[0212] Calculate the comprehensive acceptance evaluation value based on the language simplicity vector and layout adaptation score. If the evaluation value is lower than the preset adaptation threshold, the template optimization process is triggered;
[0213] Collect text conversation records from the target user group through social media data interfaces, use deep learning neural networks to identify multilingual switching points, and calculate the number of language switches per unit time;
[0214] The interaction frequency parameters in the cultural background distribution matrix are updated based on the language switching frequency data, and the cultural weights of each region are recalculated using a support vector machine.
[0215] The natural language processor restructures the text, adjusts the language expression according to the updated cultural weight, and generates optimized text content;
[0216] Use the interactive layout optimizer to adjust the button layout based on the target device's operating habits, calculate the optimal button spacing and position parameters, and generate a new template set;
[0217] Here is an example:
[0218] In this embodiment, the adjusted template set is evaluated for language simplicity and interactive button layout to verify whether it meets the cultural and operational needs of the target audience;
[0219] The template set contains text content and the corresponding interactive interface layout. The language expression must be clear and concise while ensuring that the interface operation is intuitive and efficient.
[0220] By quantifying language simplicity, evaluating button layout adaptability, and dynamically updating the cultural background distribution matrix based on user conversation data, we can generate a template collection that better meets regional needs.
[0221] This embodiment does not impose too many restrictions on the specific evaluation algorithm or data collection method. Technicians can flexibly choose the implementation method according to the actual scenario;
[0222] Specifically, a text evaluator is used to scan text in a template set, such as the template "May you bloom like a lotus and be forever happy";
[0223] The evaluator calculated that the number of characters in a single sentence is 9, the sentence nesting level is 1 (no complex clauses), and the proportion of modifiers is 22% ("happiness" is a modifier). These indicators generate a language conciseness vector of [9, 1, 0.22], reflecting that the text is concise and clear, suitable for quick reading scenarios;
[0224] The evaluator is based on natural language processing technology and uses a pre-trained model to extract sentence features to ensure indicator accuracy. This vector provides the data foundation for subsequent acceptance evaluation, ensuring that the template language is easy to understand and reducing the reading burden for users.
[0225] Extract button layout parameters from the interactive interface of the template collection. For example, the "Send" button is 85x42 pixels in size, located at the bottom center of the screen at coordinates [540, 1580], and has an 11% margin to the edge.
[0226] The interactive interface analyzer verifies whether buttons meet the minimum touch area and intuitive operation requirements according to interface design specifications;
[0227] The calculation results show a layout adaptation score of 0.92, indicating that the button design is efficient and meets ergonomic standards. This score reflects the user-friendliness of the interface on different devices and provides a key basis for comprehensive acceptance evaluation.
[0228] The language conciseness vector score of 0.85 and the layout adaptation score of 0.92 were weighted and calculated to obtain a comprehensive acceptance evaluation value of 0.87;
[0229] If the comprehensive acceptance evaluation value is lower than the preset acceptance evaluation threshold of 0.9, the text conversation records of the target user group are collected through the social media data interface, such as 5,000 blessing text message records in region A;
[0230] Using deep learning neural networks, we identify multi-language switching points in records and calculate the number of language switches per unit time. For example, we can switch between Chinese and English 1.8 times per minute, reflecting users' preference for a dynamic language environment.
[0231] Update the interaction frequency parameter of the cultural background distribution matrix based on the number of switches. For example, adjust the interaction frequency of region A to 0.28.
[0232] The regional cultural weights were recalculated using a support vector machine to generate a weight vector [simple 0.45, friendly 0.35, multilingual 0.2];
[0233] The natural language processor restructures the template text based on the weights. For example, "May you bloom like a lotus, and happiness be with you forever" is changed to "Likealotus, happiness follows you." The character count is limited to 11, and English expressions are incorporated to accommodate multilingual preferences.
[0234] The interactive layout optimizer adjusts the button layout according to the operating habits of the target device, increases the button size to 90x48 pixels, and adjusts the spacing ratio to 13% to improve touch accuracy. The new template set integrates the optimized text and interface layout, significantly enhancing regional adaptability and user interaction experience.
[0235] Furthermore, in this embodiment, step S6 specifically includes:
[0236] Obtain the text alignment and spacing features of the new template set, calculate the adaptation ratio of line spacing to font size, generate typesetting reference vectors, quantify the font display effect, and generate a font display parameter table;
[0237] Check the new template content according to the taboo word avoidance rules and symbol semantic mapping table to generate text compliance data;
[0238] Obtain the language structure features of the new template, generate sentence feature vectors, and use deep neural networks to analyze the sentiment attributes of words to generate a sentiment tendency mapping matrix;
[0239] Perform weighted calculations on font display parameters, text compliance data, and sentiment mapping matrix to generate a template candidate scoring table. High-scoring templates are then screened for information integrity verification to generate an information clarity score.
[0240] Specifically, a typesetting evaluator is used to extract alignment feature data of the template text, including line indentation value, double-justification parameters and paragraph spacing ratio, calculate the ratio of text line spacing to font size, and generate a typesetting reference vector;
[0241] Quantitatively evaluate the font display effect based on the typesetting benchmark vector, calculate the text and title font size ratio, font weight parameters, and line height scaling factor, and generate a font display parameter table;
[0242] Use a text scanner to check the template content according to the taboo word avoidance rules, extract the cultural symbol usage norms from the symbol semantic mapping table, and generate text compliance data;
[0243] The sentence analyzer extracts the language structure features in the template, including sentence length distribution, subordination hierarchy, and modifier position, and generates a sentence feature vector.
[0244] The language preference matching degree of the text is calculated based on the sentence feature vector, and the emotional attribute tags of the vocabulary are identified using a deep neural network to generate an emotional tendency mapping matrix.
[0245] A text scorer is used to perform weighted calculation on the font display parameter table, text compliance data, and sentiment mapping matrix to generate a template candidate score table;
[0246] According to the template candidate scoring table, the template groups with the highest scores are screened out, and the information integrity of the screened templates is verified to generate an information clarity score;
[0247] The typesetting evaluator is used to extract the alignment feature data of the template text and analyze the standardization of the text in visual presentation;
[0248] Here is an example:
[0249] In this embodiment, a multi-dimensional evaluation is conducted on the new template collection, covering the visual layout of the text, the cultural compliance of the content, and the emotional adaptability of the language, in order to select the SMS template that best meets the needs of the target audience;
[0250] The evaluation process combines typographic feature analysis, compliance checks for taboo words and cultural symbols, and sentiment analysis of sentence patterns and vocabulary to generate a comprehensive score to guide template optimization.
[0251] The final candidate template group output achieves the best balance between information clarity and user acceptance, and is suitable for cross-cultural communication scenarios;
[0252] This embodiment does not impose too many restrictions on the specific evaluation algorithm or data processing method. Technicians can choose an appropriate implementation method according to the actual scenario.
[0253] Specifically, a typesetting evaluator is used to analyze the visual presentation features of the template text, such as extracting the line indentation value of 2 characters, the justification setting, and the paragraph spacing of 1.6 times the line spacing of the business invitation template;
[0254] The evaluator calculates a line spacing to font size ratio of 1.3:1 to ensure readability on small-screen devices. It generates a typographic baseline vector [2, 1, 1.6, 1.3], reflecting the degree of normalization of alignment and spacing. Based on this vector, the font display effect is quantified, with the title font size set to 15 pixels and the body text size to 11 pixels. The ratio is 1.36:1, the font weight is medium-thick, and the line height scaling factor is 1.5.
[0255] The generated font display parameter table ensures that the template is clear and eye-catching on different devices, optimizing the user reading experience;
[0256] Use a text scanner to check the compliance of template content, such as scanning the business SMS template "Please join the meeting and guide the future";
[0257] The scanner, based on the taboo word avoidance rules, confirms that there are no sensitive words such as "urgent", and extracts cultural symbol specifications from the symbol semantic mapping table. For example, it verifies that "→" represents guidance in Region A, which conforms to user habits.
[0258] Generate text compliance data with a score of 0.95, indicating that the content avoids the risk of cultural conflicts.
[0259] Use a sentence parser to extract the sentence length, subordination level, and modifier position of the template. For example, the template "Please come to the meeting and guide the future" has a sentence length of 12 characters, a subordination level of 1, and the modifier placed at the end of the sentence.
[0260] Generate a sentence pattern feature vector [12, 1, 1] to reflect a concise language structure. Use a deep neural network with the sentence pattern feature vector and the word segmentation result as inputs to identify the emotional attributes of words. For example, "come to" is marked as a positive and polite emotion, and "guide" is marked as a positive and motivating emotion.
[0261] Generate an emotional tendency mapping matrix to record the emotional scores of words. For example, [come to: 0.8, guide: 0.7], ensuring that the tone of the template fits the preferences of the target audience and enhancing emotional resonance.
[0262] Use a text scoring device to integrate multi-dimensional evaluation data. For example, the font display parameter table score is 0.9, the text compliance is 0.95, and the emotional tendency mapping matrix score is 0.85. Calculate the comprehensive score of 0.92 through weighted calculation.
[0263] Generate a candidate template scoring table to rank the quality of templates. For example, the template "Please come to the meeting and guide the future" has a score of 0.92 and ranks among the top.
[0264] Verify the information integrity of the high-score templates to ensure that they contain key information. For example, "Meeting time: October 15th, 15:00, Location: Conference Room B".
[0265] Generate an information clarity score through user testing. For example, the optimized template score is 0.93, which is better than the initial template's 0.75. This score reflects the comprehensive performance of the template in terms of vision, compliance, and emotion, significantly improving the effectiveness of cross-cultural communication.
[0266] Furthermore, in this embodiment, step S7 specifically includes:
[0267] Sort the candidate SMS templates according to the information clarity score, extract the visual layout attributes of the high-score templates, generate a visual structure vector, and optimize the layout of the decorative icons to generate a visual element layout diagram.
[0268] Allocate space for the button positions, calculate the ratio of the trigger area to the screen size, set a safety distance, and generate a button layout specification table.
[0269] Collect user reading speed and scrolling frequency data, construct a reading habit feature vector, extract speech rhythm features to generate a speech rate preference matrix, and optimize the text rhythm to generate a rhythm adaptation vector;
[0270] Comprehensively score the visual element layout diagram, button layout specification table, and rhythm adaptation vector, calculate the weight distribution, generate a comprehensive template score, and select the template with the highest score as the target SMS template;
[0271] Specifically, a scoring filter is used to sort candidate SMS templates in descending order based on their message clarity scores. The visual layout attributes of the top-ranked templates, including image-text ratio, element distribution density, and white space ratio, are extracted to generate a visual structure vector.
[0272] Layout and position the decorative icons based on the visual structure vector, calculate the ratio between the icon size and the display area, set the minimum interval threshold between icons, and generate a visual element layout diagram;
[0273] Use the interactive layout tool to allocate space for button positions, calculate the ratio of button trigger area to screen size, set the safe distance between adjacent buttons, and generate a button layout specification table;
[0274] The target audience's reading speed and scrolling frequency are recorded through the user behavior collector. The reading habit feature vector is constructed based on the user interaction trajectory to generate a user behavior dataset.
[0275] Extract speech rhythm features based on user behavior datasets, including single-word dwell time, inter-sentence pause time, and reading coherence indicators, and generate a speech rate preference matrix.
[0276] A deep neural network is used to calculate the degree of match between speech rate preference and template rhythm, optimize the rhythm of text paragraph structure, and generate a rhythm adaptation vector.
[0277] The template evaluator comprehensively scores the visual layout, button specifications, and speech speed adaptation data. A support vector machine is used to calculate the weight distribution of each indicator to generate a comprehensive template score.
[0278] According to the template comprehensive score, the template with the highest score is selected as the target SMS template;
[0279] Furthermore, a graphic-text evaluator is used to count the proportion of visual elements in the candidate templates, including the area of text areas, icon elements, and blank areas, to generate an element distribution vector.
[0280] Calculate the image and text spacing parameters based on the element distribution vector, use the image processor to calibrate the icon size according to the preset display specifications, dynamically adjust the spacing between adjacent visual elements, and generate visual layout specifications;
[0281] The interactive layout optimizer reads screen display parameters, calculates the ratio between the button trigger area and the display area, sets the button spacing threshold according to operational safety standards, and generates a button layout plan.
[0282] A user behavior collector is used to record reading operation trajectories, including dwell time distribution, sliding rate changes, and operation response delay, to construct a user interaction feature matrix.
[0283] Extract speech rate and rhythm parameters based on the user interaction feature matrix, use a deep neural network to calculate the reading coherence index of the text paragraph, and generate speech rate adaptation data;
[0284] The layout evaluator performs multi-dimensional quantification on visual specifications, button schemes, and speech rate data, sets evaluation dimension weight coefficients, and generates a layout score vector.
[0285] Use support vector machine to perform feature matching on layout score vectors, select the optimal layout scheme according to preset optimization rules, and generate a final template scheme;
[0286] Use the image-text evaluator to count the proportion of visual elements in the candidate template and generate an element distribution vector;
[0287] Here is an example:
[0288] In this embodiment, the optimal template is selected by comprehensively evaluating the information clarity, visual presentation, and interaction efficiency of candidate SMS templates, ensuring that it conforms to the language habits of the target audience and provides an efficient operation experience in cross-cultural communication.
[0289] Analyze the distribution of visual elements, the usability of button layout, and the rhythm preferences of user reading behavior, and generate target templates based on a multi-dimensional scoring mechanism;
[0290] The final template strikes a balance between visual appeal, ease of use, and content adaptability, significantly improving user engagement and information delivery efficiency.
[0291] This embodiment does not impose too many restrictions on the specific screening algorithm or user data analysis method. Technicians can choose an appropriate implementation method according to the actual scenario.
[0292] Specifically, a scoring filter is used to sort candidate templates in descending order according to their information clarity scores, for example, to filter out templates with scores higher than 0.9;
[0293] We extracted visual layout attributes from the high-scoring template, including an image-text ratio of 65%, an element density of 0.8 (number of elements per square centimeter), and a white space ratio of 12%. We generated a visual structure vector [0.65, 0.8, 0.12]. Based on this vector, we optimized the layout of the decorative icons and calculated their size to be 4% of the screen width, proportional to the display area.
[0294] The minimum spacing threshold between icons is set to 1.2 times the icon width, for example, 24 pixels, to ensure visual clarity. The generated visual element layout diagram standardizes icon position and spacing, improving the visual comfort of the template on small-screen devices and enhancing the intuitiveness of information transmission.
[0295] Analyze the template's interactive interface using the interactive layout tool. For example, on a 1080-pixel screen, set the primary button width to 400 pixels, which accounts for approximately 37% of the screen width.
[0296] The minimum trigger area is 48x48 pixels to meet touch requirements. The safe distance between adjacent buttons is set to 0.7 times the button width, for example 28 pixels, to avoid accidental touches. The width of the secondary button is 75% of the primary button to form a visual hierarchy.
[0297] The generated button layout specification table records button coordinates, sizes, and spacing, such as the primary button coordinates [540,1600], to ensure intuitive and efficient operation. This specification table optimizes the user interaction experience and is particularly suitable for quick response scenarios.
[0298] Use user behavior collectors to record the target audience's reading behavior. For example, data collected from 1,000 users on business text message reading showed that the dwell time on a single word was 0.25 seconds, the pause between sentences was 0.6 seconds, and the scrolling speed was 180 pixels per second during quick browsing, dropping to 90 pixels per second during careful reading.
[0299] Construct a reading habit feature vector [0.25, 0.6, 180, 90] to reflect the user's attention pattern, and extract speech rhythm features to generate a speech rate preference matrix. Record the comfortable reading speed as 190 words per minute and the paragraph length as 12 words.
[0300] A deep neural network is used to calculate the matching degree between the matrix and the template rhythm. For example, a paragraph is adjusted to 15 words or less, with a pause rhythm of 0.7 seconds, and a rhythm adaptation vector [190, 12, 0.7] is generated. This vector optimizes the text structure, improving reading fluency and information absorption efficiency.
[0301] A template evaluator was used to integrate multi-dimensional data, with the weights of visual layout set at 35%, button specification at 30%, and rhythm adaptation at 35%. For example, a template with a visual layout score of 8.8, button specification at 9.2, and rhythm adaptation at 8.6 had a total score of 8.86.
[0302] The scoring vectors were feature-matched using a support vector machine, and the overall performance of candidate templates was compared to obtain the highest-scoring template. For example, the slogan "Please attend the meeting on October 20th and create a better future together" with an optimized layout scored 8.9 and was selected for its clear visual layout, convenient button operation, and adaptive reading rhythm. This target template fully adapted to the target audience in terms of visuals, interaction, and content, improving user engagement and communication effectiveness in business scenarios.
[0303] The present invention also proposes an intelligent SMS template generation system based on a large language AI model, comprising:
[0304] The cross-cultural communication analysis module extracts cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs, generating a cultural background distribution matrix that includes lexical sentiment tendencies and message length controls;
[0305] The taboo word avoidance module is used to pre-establish a taboo word avoidance rule library, identify regional cultural characteristics in the cultural background distribution matrix, scan taboo words based on regional cultural characteristics, obtain taboo matching results, and update the cultural background distribution matrix based on the matching degree and frequency of occurrence to avoid taboo words;
[0306] The expression habit analysis module is used to extract the target audience's habitual adaptation and speaking speed and rhythm preferences based on the updated cultural background distribution matrix, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table;
[0307] The SMS template generation module is used to generate preliminary SMS templates based on the symbol semantic mapping table and the expression habit preference matrix, optimize the proportion of rhetorical devices selected in the preliminary SMS templates, adjust text color contrast and layout spacing based on screen size adaptation data, and output a set of adjusted templates that meet the message length control;
[0308] The template optimization module is used to analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than the preset acceptance threshold, the cross-cultural communication frequency is re-acquired, the cultural background distribution matrix is updated, and a new template set is generated;
[0309] The acceptance evaluation module is used to evaluate the text alignment of the new template set, combine the taboo word avoidance results and the symbol semantic mapping table, screen candidate SMS templates, analyze the fit between the sentence preference analysis results and the word sentiment tendency, and obtain the information clarity score;
[0310] The SMS template screening module is used to screen candidate SMS templates based on information clarity scores, optimize visual element embedding and interactive button layout, analyze the audience's engagement potential based on speech speed and rhythm preferences, and determine the target SMS template.
[0311] In the description of the present invention, it should be understood that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the device or element referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention.
[0312] Those skilled in the art can make various other corresponding changes and deformations based on the technical solutions and concepts described above, and all of these changes and deformations should fall within the scope of protection of the claims of the present invention.
Claims
1. A method for generating intelligent SMS templates based on a large language AI model, characterized in that: The following steps are involved: S1. Extract cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs to generate a cultural background distribution matrix. Specifically, it includes: Collect message data posted by users, perform word segmentation on the message data, extract language usage feature data, and generate a language sentence structure matrix based on the language usage feature data; If the message data contains a mixture of languages, a classification algorithm is used to divide the content in different languages, and cross-cultural communication frequency data is extracted from the division results to generate a cultural interaction intensity matrix; Obtain user click behavior records and ad page visit data, perform clustering operations to extract user interaction time series features, and determine user activity indicators; Calculate the message text length distribution based on the language sentence structure matrix and generate message length control parameters based on user activity indicators; A neural network is used to analyze the sentiment polarity of message data, extracting the sentiment expression feature vector. This is combined with the cultural interaction intensity matrix and message length control parameters to generate a cultural background distribution matrix that includes lexical sentiment tendency and message length control. S2. Pre-establishing a taboo word avoidance rule library, identifying regional cultural characteristics in the cultural background distribution matrix, obtaining taboo matching results, and updating the cultural background distribution matrix to avoid taboo words; S3. Based on the updated cultural background distribution matrix, extract the target audience's habitual adaptation of address and speaking speed and rhythm preferences, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table; S4. Generating a preliminary SMS template based on the expression habit preference matrix and the symbol semantic mapping table, optimizing the preliminary SMS template, and outputting an adjusted template set that meets message length control requirements; S5. Analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than a preset acceptance threshold, re-acquire the cross-cultural communication frequency and generate a new template set. S6. Evaluate the text alignment of the new template set, combine the taboo word avoidance results and the symbol semantic mapping table, analyze the fit between the sentence preference analysis results and the word sentiment tendency, and obtain an information clarity score; S7. Screen candidate SMS templates based on the information clarity score, optimize visual element embedding and interactive button layout, analyze the audience's engagement potential based on speech speed and rhythm preferences, and determine the target SMS template.
2. The method for generating an intelligent SMS template based on a large language AI model according to claim 1 is characterized in that: The step S2 specifically includes: Extracting regional cultural feature data from the cultural background distribution matrix to generate regional language distribution feature data, and processing the regional language distribution feature data using a neural network model to generate a regional cultural feature vector; According to the regional cultural feature vector, a taboo word rule set of the corresponding region is obtained from a taboo word avoidance rule library, and taboo words in the input text are identified to generate a taboo word frequency statistical matrix; If the taboo word frequency in the taboo word frequency statistical matrix exceeds a preset sensitivity threshold, a replacement word list is extracted from the taboo word avoidance rule library, an optimal replacement word is selected using a semantic similarity algorithm, and optimized text content is generated; The cultural background distribution matrix is updated according to the word frequency distribution, language usage characteristics and cultural sensitivity score of the optimized text content.
3. The method for generating an intelligent SMS template based on a large language AI model according to claim 1 is characterized in that: The step S3 specifically includes: The target area is divided according to the updated cultural background distribution matrix. The speech segments in the target area are analyzed, the phoneme sequences and their duration and pause characteristics are extracted, the speech rate feature vector is generated, and a speech rate rhythm indicator matrix is constructed. Classify the target region's address corpus, construct an address relationship map, and generate address adaptation vectors based on the closeness of the relationship between people; By parsing the text collection of the target area, extracting sentence structure, rhetoric and modal particle features, and generating an expression habit preference matrix; Convolutional neural networks are used to extract the semantic features of cultural symbols in text collections, build a regional symbol library, and generate a symbol semantic mapping table through semantic quantification.
4. The method for generating an intelligent SMS template based on a large language AI model according to claim 1, characterized in that: The step S4 specifically includes: Extracting semantic features from the symbol semantic mapping table to generate a semantic vector, and combining the language style parameters in the expression habit preference matrix to generate an initial template text, and forming a preliminary SMS template group by segmenting and combining semantic units; Construct rhetorical feature vectors based on rhetorical device statistics, calculate the applicability probability of rhetorical devices in different scenarios, optimize the rhetorical ratio and generate a rhetorical optimization template; Obtain the display area size of the target user's terminal device, calculate the font size and paragraph spacing, generate typesetting parameters, and generate a display optimization template based on color contrast optimization; If the template text exceeds the length limit, a semantic-preserving compression algorithm is used to control the length, and a template set is generated based on the layout and color optimization results.
5. The method for generating an intelligent SMS template based on a large language AI model according to claim 1, characterized in that: The step S5 specifically includes: Obtain the adjusted template set and analyze its language conciseness, calculating the number of characters in a single sentence, the sentence nesting level, and the proportion of modifiers to generate a language conciseness vector. Extract the size, position, and spacing parameters of the button layout in the template set, and calculate the layout adaptation score based on the interface design specifications; The comprehensive acceptance evaluation value is calculated by combining the language conciseness vector and the layout adaptation score. If the acceptance is lower than the preset acceptance threshold, the cultural background distribution matrix is updated by analyzing the user conversation records and a new template set is generated.
6. The method for generating an intelligent SMS template based on a large language AI model according to claim 2, characterized in that: The step S6 specifically includes: Obtain the text alignment and spacing features of the new template set, calculate the adaptation ratio of line spacing to font size, generate typesetting reference vectors, quantify the font display effect, and generate a font display parameter table; Check the new template content according to the taboo word avoidance rules and the symbol semantic mapping table to generate text compliance data; Obtain the language structure features of the new template, generate sentence feature vectors, and use deep neural networks to analyze the sentiment attributes of words to generate a sentiment tendency mapping matrix; A weighted calculation is performed on the font display parameters, the text compliance data, and the sentiment tendency mapping matrix to generate a template candidate scoring table, and high-scoring templates are screened for information integrity verification to generate an information clarity score.
7. The method for generating an intelligent SMS template based on a large language AI model according to claim 2, characterized in that: The step S7 specifically includes: sorting candidate SMS templates according to the information clarity scores, extracting visual layout attributes of high-scoring templates to generate visual structure vectors, and optimizing the layout of decorative icons to generate a visual element layout diagram; Allocate space for button positions, calculate the ratio of trigger area to screen size, set safe spacing, and generate a button layout specification table; Collect user reading speed and scrolling frequency data, construct a reading habit feature vector, extract speech rhythm features to generate a speech rate preference matrix, and optimize the text rhythm to generate a rhythm adaptation vector; The visual element layout diagram, the button layout specification table and the rhythm adaptation vector are comprehensively scored, the weight distribution is calculated, the template comprehensive score is generated, and the highest-scoring template is selected as the target SMS template.
8. An intelligent SMS template generation system based on a large language AI model, used to execute the method according to any one of claims 1 to 7, characterized in that: include: The cross-cultural communication analysis module extracts cross-cultural communication frequency, sentence preference analysis results, and screen size adaptation data from the target audience's social media interaction data and device usage logs, generating a cultural background distribution matrix that includes lexical sentiment tendencies and message length controls; The taboo word avoidance module is used to pre-establish a taboo word avoidance rule library, identify regional cultural characteristics in the cultural background distribution matrix, scan taboo words based on regional cultural characteristics, obtain taboo matching results, and update the cultural background distribution matrix based on the matching degree and frequency of occurrence to avoid taboo words; The expression habit analysis module is used to extract the target audience's habitual adaptation and speaking speed and rhythm preferences based on the updated cultural background distribution matrix, generate an expression habit preference matrix, analyze the deep meaning of cultural symbol semantics, and obtain a symbol semantic mapping table; The SMS template generation module is used to generate preliminary SMS templates based on the symbol semantic mapping table and the expression habit preference matrix, optimize the proportion of rhetorical devices selected in the preliminary SMS templates, adjust text color contrast and layout spacing based on screen size adaptation data, and output a set of adjusted templates that meet the message length control; The template optimization module is used to analyze the acceptability of the adjusted template set in terms of language simplicity and interactive button layout. If the acceptability is lower than the preset acceptance threshold, the cross-cultural communication frequency is re-acquired, the cultural background distribution matrix is updated, and a new template set is generated; The acceptance evaluation module is used to evaluate the text alignment of the new template set, combine the taboo word avoidance results and the symbol semantic mapping table, screen candidate SMS templates, analyze the fit between the sentence preference analysis results and the word sentiment tendency, and obtain the information clarity score; The SMS template screening module is used to screen candidate SMS templates based on information clarity scores, optimize visual element embedding and interactive button layout, analyze the audience's engagement potential based on speech speed and rhythm preferences, and determine the target SMS template.
Citation Information
Patent Citations
Generative large model-based malicious short message variant word restoration method
CN117010328A
Computer intelligent writing system driven by artificial intelligence
CN119150807A