Short message content compliance generation method and system fusing semantic large model
Through the SMS content compliance generation method that integrates semantic big models, the intelligence and compliance problems of SMS content review are solved, and the SMS generation is generated dynamically identifying and risk avoidance is realized, which improves compliance and user experience.
Patent Information
- Application Number
- CN202510543567.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the existing technology, SMS content review relies on manual experience and rule databases, making it difficult to accurately identify the risks of semantic ambiguity and complex scenarios, and the static rule database is difficult to adapt to regulatory changes, resulting in lagging compliance reviews and being unable to dynamically identify and avoid compliance risks.
The SMS content compliance generation method is adopted with a fusion semantic model, and SMS content that meets user needs and compliance requirements is generated through rough identification, in-depth semantic analysis, multi-dimensional compliance comparison and optimization and adjustment.
It realizes intelligent generation of SMS content, improves compliance and user experience, dynamically adapts to regulatory changes, and accurately identifys and avoids compliance risks.
Smart Images

Figure CN120409460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and particularly to a method and system for generating compliant SMS content integrating a semantic large model. Background Art
[0002] As an important information transmission method in communication services, SMS is widely used in multiple industries such as commerce, finance, healthcare, and government affairs. In these industries, SMS content usually involves sensitive content such as user privacy, financial information, and medical data. Therefore, ensuring the compliance of SMS content is crucial. Currently, the review of SMS content mainly relies on the matching of manual experience and rule libraries. When faced with the need to process a large volume of SMS, manual review is inefficient and there are subjective judgment biases. It is particularly difficult to accurately identify risks in complex scenarios such as semantic ambiguity and context correlation. For example, in financial industry marketing SMS, induced expressions such as "high returns with no risk" may avoid detection by the basic rule library through simple keyword replacement, and manual review is prone to missed judgments due to insufficient depth of semantic understanding; in medical SMS, hidden expressions involving efficacy commitments are often difficult to be captured by traditional automated systems due to complex context logic. At the same time, the regulatory requirements for communication services are dynamically adjusted with industry policies. For example, the list of prohibited words for Internet financial advertisements is updated frequently, and traditional static rule libraries are difficult to adapt in a timely manner, resulting in compliance reviews lagging behind regulatory changes and increasing the operational risks of enterprises. In addition, there are often contradictions between users' requirements for the expression effect of SMS (such as the attractiveness of marketing copy and the clarity of notification content) and compliance requirements. Without an intelligent dynamic balance mechanism, it is easy to cause over-review and damage service efficiency. Therefore, how to improve the intelligent level of content generation and user experience while ensuring the compliance of SMS content has become a technical problem to be solved urgently. Summary of the Invention
[0003] This application provides a method and system for generating compliant SMS content integrating a semantic large model, aiming to solve the technical problem that the generation of SMS content in communication services relies on templates and cannot dynamically identify and avoid compliance risks according to user needs.
[0004] In the first aspect disclosed in this application, a method for generating compliant SMS content integrating a semantic large model is provided. The method includes: performing rough identification on the SMS content to be processed by a service provider to obtain the SMS type and user needs; inputting the SMS content into a semantic large model for in-depth semantic parsing and outputting SMS semantic parsing information; performing multi-dimensional compliance comparison on the SMS semantic parsing information according to the SMS type and user needs to identify risk description content; and optimizing and adjusting the risk description content according to the SMS type and user needs according to the semantic parsing information of the risk content to generate the target SMS content.
[0005] Another aspect disclosed in this application provides a system for generating compliant SMS content integrating a semantic large model, which includes: a rough recognition unit: performing rough recognition on the SMS content to be processed by the service provider to obtain the SMS type and user requirements; a deep semantic parsing unit: inputting the SMS content into the semantic large model for deep semantic parsing and outputting SMS semantic parsing information; a multi-dimensional compliance comparison unit: performing multi-dimensional compliance comparison on the SMS semantic parsing information according to the SMS type and user requirements to identify risk description content; an optimization and adjustment unit: optimizing and adjusting the risk description content according to the SMS type and user requirements according to the semantic parsing information of the risk content to generate the target SMS content.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: The above method for generating compliant SMS content integrating a semantic large model first conducts a preliminary analysis on the SMS content to be processed submitted by the service provider to identify which category the SMS belongs to (such as marketing, notification, etc.) and the specific expression requirements of the user. Subsequently, the SMS text is input into the trained semantic large model for in-depth semantic understanding to extract the semantic information expressed in the SMS. Then, based on the identified SMS type and user requirements, multi-dimensional compliance comparison is performed on the parsed semantic information to detect whether there are risky expressions that may violate industry norms or laws and regulations. Finally, for the identified risk content, targeted content optimization and adjustment are carried out in combination with the semantic information to ensure that while meeting the user's expression intention, compliant, safe, and sendable SMS content is generated.
[0007] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented in accordance with the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically illustrates the specific embodiments of this application. Brief Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] Figure 1 It is a flowchart of the method for generating compliant SMS content integrating a semantic large model in one embodiment.
[0010] Figure 2 It is an architecture diagram of the system for generating compliant SMS content integrating a semantic large model in one embodiment.
[0011] Explanation of the accompanying symbols: coarse recognition unit 11, deep semantic analysis unit 12, multi-dimensional compliance comparison unit 13, optimization and adjustment unit 14. DETAILED DESCRIPTION
[0012] The embodiments of the present application provide a method and system for generating SMS content compliance that integrates a semantic big model, thereby solving the technical problem that SMS content generation in communication services relies on templates and cannot dynamically identify and avoid compliance risks based on user needs.
[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0015] Example 1, as Figure 1 As shown, the present application provides a method for generating compliant SMS content by integrating a semantic large model, the method comprising: Roughly identify the content of SMS messages to be processed by the service provider to obtain the SMS type and user needs.
[0016] In an embodiment of the present application, the original SMS content submitted by the service provider is preliminarily classified and the intent is identified. Specifically, the system first receives the SMS text to be processed by the service provider and extracts key features therein, such as keywords, sentence structures, etc. Subsequently, combined with the service scenarios corresponding to the SMS (such as finance, medical care, government affairs and other industries) and the purpose of the SMS (such as marketing promotion, user notifications, service reminders, etc.), a rough identification is performed through the trained logistic regression model, that is, preliminary classification and intent identification, and the specific type of the SMS and the user needs it reflects are output. In this way, clear scenario labels and user intent basis can be provided for subsequent semantic analysis and compliance processing, which helps to improve the pertinence and accuracy of compliance processing.
[0017] Furthermore, this application provides a rough identification of the content of SMS messages to be processed by the service provider to obtain the SMS type and user needs, including: Obtain the industry types and SMS demand types for SMS compliance review; based on the industry types and SMS demand types, perform data tagging to construct training data; use the training data to train a logistic regression model to construct a semantic rough recognition channel, where the logistic regression model is used to identify and output the SMS type and user demand, the SMS type corresponds to the industry type, and the user demand corresponds to the SMS demand type.
[0018] Preferably, obtain the industry types and SMS demand types for SMS compliance review based on the current communication service compliance standards, industry supervision regulations, and SMS service access scope, etc. Among them, the industry type is the industry classification sorted out according to the objects and content sources of SMS platform services. For example, finance, medical, government affairs, education, e-commerce, transportation, etc.; the SMS demand type is classified according to the functional position of SMS in the business process, such as marketing type (such as advertising promotion), notification type (such as verification code, bill reminder), service type (such as progress query, appointment service), etc. Subsequently, collect multiple SMS samples, and label each SMS sample with the industry type and SMS demand type to construct training data. Then, perform text preprocessing on the training data, that is, perform operations such as word segmentation, stop word removal, and special character cleaning on the SMS samples, and then use TF-IDF (term frequency-inverse document frequency), bag-of-words model, etc. to convert the preprocessed SMS samples into numerical feature vectors so that the logistic regression model can receive and learn. After that, select the multinomial logistic regression model as the core classifier, construct a semantic rough recognition channel, and initialize key parameters, such as task type (multinomial), optimizer (lbfgs), regularization strength (1.0), etc. Then, divide the training data into a training set and a validation set, use the numerical feature vectors as input features, and use the corresponding labels (industry type, SMS demand type) as target variables to train the semantic rough recognition channel. During the training process, the semantic rough recognition channel iteratively optimizes its weight parameters to minimize the loss function (cross-entropy loss) between the prediction result and the true label, thereby gradually improving the classification accuracy of the channel until convergence or reaching the preset maximum number of iterations. After the training is completed, use the validation set to evaluate the model, calculate indicators such as classification accuracy, recall rate, and F1-score to ensure the generalization ability of the model on unseen data, and further adjust the model hyperparameters (such as regularization parameters) according to the evaluation results. Finally, achieve the accurate classification of the industry type and demand type of SMS. Among them, the industry type identified by the semantic rough recognition channel will be output as the SMS type, and the identified SMS demand type will be output as the user demand, which is used to provide a pre-input condition for subsequent in-depth semantic analysis and compliance judgment, and help perform more accurate semantic analysis and risk comparison.
[0019] Furthermore, the present application provides that the industry types include at least: financial industry, medical industry, and government announcements, and the SMS demand types include at least marketing, notifications, and services.
[0020] The optional "Industry Type" refers to the business area covered by the SMS message, reflecting the background or business scope of the service provider to which the SMS message belongs. This includes at least the financial industry, the medical industry, and government announcements. The financial industry encompasses banking, insurance, securities, and lending services, such as credit card repayment reminders, loan approval notifications, and investment and financial management promotions. The medical industry encompasses hospitals, physical examination centers, and health management platforms, such as appointment registration notifications, physical examination report reminders, and vaccination reminders. Government announcements are notification messages sent by government agencies or government service platforms, such as community announcements, epidemic prevention campaigns, and policy change notifications. The SMS demand type refers to the actual use or intent of the SMS message, used to determine the scenario or purpose for the message. This includes at least marketing, notification, and service. Marketing refers to SMS messages used to promote products or services, such as merchant promotions, limited-time discounts, and membership benefits. Notifications convey important information to users, such as verification codes, bill reminders, and successful appointment notifications. Service refers to SMS messages used to provide auxiliary service support, such as customer service follow-up visits, satisfaction surveys, and system upgrade notifications. By analyzing SMS text, we can identify the industry type and demand type to which it belongs, thereby providing accurate contextual information for subsequent semantic analysis and compliance processing.
[0021] The SMS content is input into the semantic big model for deep semantic analysis, and SMS semantic analysis information is output.
[0022] In one embodiment, after obtaining a piece of SMS content to be processed, a semantic big model will be used to conduct an in-depth understanding and semantic structure analysis of the SMS content. Specifically, the pre-processed SMS text will be input into the constructed semantic big model. This semantic big model is usually built based on a deep neural network architecture and includes a multi-layer structure such as an input layer, an encoding layer, a context understanding layer, and an output layer. By training this semantic big model, a deep semantic parsing channel can be built. This deep semantic parsing channel can not only identify the explicit information in the SMS, but also deeply explore the implicit semantic relationships, providing accurate semantic support for subsequent compliance comparison and content optimization. Finally, the deep semantic parsing channel will perform a deep semantic analysis of the input SMS content based on the learned knowledge, and output SMS semantic parsing information to help determine whether the content is compliant and which parts may be at risk, thereby improving the intelligence level and compliance assurance capabilities of SMS content generation in communication services.
[0023] Further, the present application provides for inputting the SMS content into a semantic large model for in-depth semantic parsing and outputting SMS semantic parsing information, including: Construct a large model framework, including an input layer, an encoder layer, a context understanding layer, a semantic parsing layer, and an output layer; train and converge the large model framework through the collected training data set to obtain the semantic large model, and build an in-depth semantic parsing channel based on the semantic large model for in-depth semantic parsing of the SMS content and outputting the SMS semantic parsing information.
[0024] Optionally, the semantic large model framework designed and built based on a deep neural network includes five core structural layers: an input layer, an encoder layer, a context understanding layer, a semantic parsing layer, and an output layer. Among them, the input layer: responsible for receiving the preprocessed SMS text; the encoder layer supports the multi-head attention mechanism, enabling the model to simultaneously focus on information at different positions in the text; the context understanding layer can determine whether a word has a risk meaning in different contexts, providing a basis for subsequent processing; the semantic parsing layer further deepens the understanding of the overall semantics of the SMS, extracting deep semantic features, including logical relationships, expression intentions, emotional tendencies, etc.; the output layer integrates the features extracted by the previous layers and outputs the semantic parsing information of the SMS for downstream compliance comparison and optimization processing. After completing the model structure design, it enters the model training stage. Specifically, a training data set covering multi-industry and multi-type SMS will be collected and constructed, semantic annotation will be performed on each SMS, such as marking the positions of sensitive words, semantic relationship labels, potential risk levels, etc., and then the labeled training data will be input into the constructed large model framework, and supervised learning will be used for training. In this process, the training data will be input into the model for forward propagation, and through layer-by-layer transmission, the predicted SMS semantic parsing information will be calculated, and then the error between the predicted result and the true label will be backpropagated through a loss function such as cross-entropy, and the model parameters will be iteratively adjusted in combination with an optimization algorithm (such as Adam or SGD) until the loss function converges, the model performance is stable, and the accuracy reaches the set standard. When the training is completed and verified by the validation set, the trained semantic large model will be deployed as an in-depth semantic parsing channel. In actual applications, this channel receives new SMS texts, performs operations such as semantic decomposition, context evaluation, and risk identification, and outputs structured SMS semantic parsing information to provide semantic support for subsequent multi-dimensional compliance comparison and content optimization.
[0025] Further, the present application provides for constructing a large model framework, including: The input layer is used to process text embeddings; the encoding layer adopts the self-attention mechanism to capture the dependencies of the text, and allows the model to simultaneously focus on different parts of the text through multi-head attention; the context understanding layer is used to detect sensitive words and identify information related to local context and sensitive words; the semantic parsing layer uses multiple layers of encoders to extract deep semantic features and further analyze the deep association relationship between sensitive words and global context; the output layer is used to integrate sensitive words, deep semantic features and context association relationships and output the semantic parsing information of the short message.
[0026] Optionally, after the short message text undergoes preprocessing (such as removing stop words, word segmentation, and cleaning special characters), it is input into the input layer of the model. The task of the input layer is to convert the original text information into a numerical vector form recognizable by the model. The specific method is the same as described above, thus forming a complete text vector sequence. This sequence retains the basic semantic features of the vocabulary and is the basis for subsequent in-depth analysis.
[0027] The encoding layer is used to model the dependencies between words and adopts the self-attention mechanism (Self-Attention) to capture the relative importance of each word in the short message in the context. Through the multi-head attention mechanism (Multi-HeadAttention), the model can learn the multi-dimensional relationships between words in parallel from multiple subspaces. For example, syntactic structures, syntactic dependencies, entity associations, etc., enhancing the model's ability to understand semantic diversity. The output encoding result is a set of context-related word vectors, representing the semantic position of each word in the entire short message text.
[0028] After obtaining the context-aware encoding result, the context understanding layer is responsible for the detection and analysis of sensitive words. By quickly matching the short message content with a preset risk word library, potential risky vocabulary is located. Subsequently, centered on these sensitive words, local context is identified based on the dependencies output by the encoding layer, helping the model determine whether the sensitive words truly have a risky meaning in a specific context and avoiding misjudgment caused by "lexical ambiguity".
[0029] Based on the context understanding result, the semantic parsing layer further constructs the deep semantic structure of the short message. Semantic parsing is usually composed of multiple layers of encoders stacked together to perform more in-depth abstraction and feature extraction on the entire short message content. Through the semantic parsing layer, the model can output highly abstract semantic labels, such as risk descriptions, compliance expressions, ambiguous intentions, etc., providing a structural basis for subsequent processing.
[0030] Finally, the output layer integrates the information from the encoding layer, context understanding layer, and semantic parsing layer to form a set of structured output content, including the detected sensitive words and their positions in the text, the contextual semantic environment and dependency relationships of the sensitive words, the deep semantic tags of the entire SMS (such as intent, expression risk level), risk intensity scores, or judgments on the types of potential violation risks, etc. These output results constitute the final SMS semantic parsing information and serve as the input basis for subsequent multi-dimensional compliance comparison and content optimization and adjustment.
[0031] In summary, this semantic large model has a clear hierarchical structure, from low-level vocabulary understanding to high-level semantic reasoning, ensuring that SMS content can be comprehensively parsed at different levels, effectively identifying potential compliance risks, and providing a solid foundation for intelligent compliance generation.
[0032] Furthermore, this application provides a method for inputting the SMS content into a semantic large model for in-depth semantic parsing and outputting SMS semantic parsing information, including: Preprocess the SMS content, including cleaning and removing irrelevant characters, and input it into the semantic large model; the input layer converts the preprocessed SMS content into an embedding vector; the encoding layer captures the dependency relationships between various words in the SMS content based on the self-attention mechanism and assigns each word a weight representing its importance in the current context; the context understanding layer matches the SMS content with a list of sensitive words to identify sensitive words, and analyzes the surrounding local context information of the detected sensitive words based on the dependency relationships output by the encoding layer to determine the meaning and riskiness of the sensitive words in the current context; based on the identified sensitive words and local context relationships, the semantic parsing layer uses a multi-layer encoder structure to extract the deep semantic features of the SMS content and further analyzes the deep association relationships between the sensitive words in the global context; connect the identified sensitive words, local context relationships, deep semantic features, and deep association relationships to output the SMS semantic parsing information.
[0033] Optionally, first, clean and preprocess the input SMS content. This step mainly includes removing irrelevant characters (such as punctuation marks, special symbols), and standardizing the text format (such as unifying case, removing stop words). The processed text is convenient for subsequent analysis and ensures the purity of the information. On this basis, through operations such as word segmentation, the text is converted into a form that can be processed by the model and is ready to be input into the semantic large model for analysis.
[0034] The preprocessed SMS content is input into the input layer. The main function of this layer is to convert the text content into a numerical embedding vector. In this way, each word in the SMS is represented as a high-dimensional vector, and the model can understand the semantics of the words based on these vectors. The output of this stage is a complete sequence of SMS vectors, representing the overall information structure of the SMS.
[0035] At the encoding layer, by employing self-attention and multi-head attention mechanisms, the model is able to capture the dependencies between individual words in the text message content. The self-attention mechanism allows the model to dynamically adjust the weight of each word based on its relationship to other words when processing it. Specifically, the model calculates each word's relevance to other words using the "Query (query vector) - Key (key vector)" score in the attention mechanism and generates a set of weights that reflect the importance of different words in the current context. Through multi-head attention, the model allows simultaneous focus on different parts of the text from multiple different perspectives. Each attention head independently calculates a set of weights, allowing the model to understand different semantic levels of the text message in parallel across multiple subspaces. This mechanism effectively improves the model's ability to understand the complex semantic structure of the text, ensuring that the importance of each word in the text message is accurately assessed in its current context.
[0036] At the context understanding layer, the model matches the text message content with a list of sensitive words extracted from a risk vocabulary to identify potentially sensitive words within the message. Once a sensitive word is identified, the model further analyzes its local context within the text based on the dependency relationships output by the encoding layer. Specifically, the model centers around these sensitive words and extracts surrounding words within their context. Using a sliding window mechanism, it analyzes the dependencies and co-occurrence frequencies between these words to determine the specific meaning of the sensitive words in the current context and their potential risk. This step is a critical risk identification stage, ensuring that the model can identify potentially illegal content within the text.
[0037] The semantic parsing layer utilizes a multi-layered encoder structure to conduct a deeper semantic analysis of the relationship between identified sensitive words and the local context, extracting the deep semantic features of the message. Specifically, this layer not only focuses on the local context but also further analyzes the relationship between sensitive words and the global context. Through these multi-layered encoders, the model can deeply understand the semantic structure of the entire message and infer the specific meaning and risks of sensitive words in the global context. For example, it can determine whether words like "transfer," "winning a lottery," and "privacy" are in the context of real services or marketing solicitations. The output of this layer is a highly abstract semantic feature that comprehensively reflects the potential intent and compliance risks of the message content.
[0038] Finally, at the output layer, the model integrates the identified sensitive words, local context relationships, deep semantic features, and global association relationships to generate structured SMS semantic parsing information, including the sensitive words and their positions in the text, local context information (semantic associations between sensitive words and their surrounding words), deep semantic features (the intent, sentiment, potential risks, etc. of the SMS), and the risk assessment of sensitive words in the global context. Through this information, compliance judgment can be made on the SMS content, providing a basis for subsequent risk comparison and optimization adjustment, thereby achieving accurate semantic parsing and compliance analysis of the SMS content.
[0039] Furthermore, the present application provides a method for determining the meaning and risk of sensitive words in the current context, including: Configuring the size of the moving window according to the weight of the sensitive word; selecting the local context of the sensitive word based on the moving window, identifying the co-occurrence frequency of the sensitive word and directly associated words within the size range of the moving window, analyzing the lexical grammar relationship according to the dependency relationship between words, and identifying the meaning and risk of the directly associated word combination; adjusting the window size according to the risk within the moving window and the similarity of the words at the window edge, and based on the adjusted moving window, identifying the indirect association between the sensitive word and the intermediate words, and analyzing the meaning and risk of the indirectly associated words according to the dependency relationship between words; comprehensively analyzing the meaning and risk of the directly associated word combination and the meaning and risk of the indirectly associated words, and outputting the meaning and risk of the sensitive word in the current context.
[0040] Optionally, first determine the size of the moving window for analysis. The window can be of a fixed size, for example, 5 words before and after, or a dynamically adjusted size, depending on the specific context and the importance of the sensitive word. In the initial processing, a fixed-size window is selected for local text analysis. In subsequent steps, the size of the moving window can be dynamically adjusted according to the weight and semantic importance of the sensitive word (the calculation method is: dividing the weight of the sensitive word by the standard weight, and then multiplying the quotient by the fixed moving window size). For example, the more complex or important the context of the sensitive word is, the window range may be appropriately expanded to ensure capturing more context information.
[0041] After determining the moving window size, the first-order correlation analysis begins. At this stage, the model starts to perform local context selection within this window, checking the co-occurrence frequency between the sensitive word and directly related words (i.e., words that are grammatically closely connected to the sensitive word). That is, it calculates the frequency of the co-occurrence of the sensitive word and each relevant word (such as adjectives, adverbs, or verbs, etc.). For example, if the sensitive word is "loan", the model will focus on other words related to "loan" within the window, such as "approval", "interest rate", "amount", or "application", etc. By calculating the frequency of the co-occurrence of these words and "loan" within the window, it can evaluate whether these words frequently co-occur with the sensitive word, thereby determining the semantic association strength between them. Subsequently, the model will analyze based on the dependency relationships between the words to further understand the grammatical structure of these directly related words. The model will identify the grammatical dependency relationships between the sensitive word and other words through syntactic analysis (such as dependency parsing). For example, in "loan application", there is a dependency relationship between "loan" as a noun and "application", indicating that "application" is a further description of the sensitive word "loan". By analyzing these dependency relationships, the model can determine the roles and meanings of each word at the grammatical level, thereby more accurately understanding their semantic connections with the sensitive word. After identifying the co-occurrence frequency of the sensitive word and directly related words and their grammatical dependency relationships, a semantic analysis will be performed on the combinations of these directly related words to judge their meanings in the current context. For example, the combination of the words "loan" and "approval" may indicate a formal loan approval with lower risk; while the combination of "loan" and "high interest rate" may indicate that the loan has a higher risk. Based on the context meaning of the combination, the risk coefficient of the co-occurrence of this combination is obtained, and the risk coefficient is multiplied by the corresponding co-occurrence frequency to obtain the riskiness of the combination of directly related words.
[0042] After the first-order analysis, it enters the second-order correlation analysis stage. In this stage, the model will dynamically adjust the window size according to the similarity of the words at the window edges. Among them, the edge words refer to the words located at the window boundaries, which may affect the semantic connection and risk assessment between the sensitive word and other words within the window; the similarity is obtained by calculating the cosine similarity between the word vectors of the window edge words and the sensitive word. Based on the similarity analysis, the window size will be adjusted adaptively. If the model detects that the importance or relevance of some edge words is high, the window will be expanded; if the edge words are irrelevant or have a large semantic difference, the window will be shrunk. In this way, the model can flexibly select an appropriate window range to capture more meaningful context information. Subsequently, within the adjusted moving window, a similar method as above will be used to analyze the dependency relationships and risks to obtain the meanings and riskiness of indirectly related words.
[0043] Finally, compare the risks of all the obtained directly associated word combinations with the corresponding preset risk thresholds. If it is greater than or equal to the risk threshold, extract the meaning and corresponding risk of this directly associated word combination. Similarly, perform the same operation on the indirectly associated words, and extract the meaning and risks of the indirectly associated words that are greater than or equal to the corresponding risk thresholds. By dividing all the extracted meanings into sensitive words, obtain the meaning of each sensitive word in the current context, and then perform weighted calculation on the risks belonging to the same sensitive word to obtain the risk of the sensitive word in the current context. Generally speaking, this process ensures the dynamic identification of sensitive words in the SMS content and the accurate judgment of their meanings and potential risks in a specific context through moving window, association analysis, and risk assessment.
[0044] Perform multi-dimensional compliance comparison on the SMS semantic parsing information according to the SMS type and user requirements to identify the risk description content.
[0045] In one embodiment, first, according to the SMS type and user requirements, determine which category it belongs to. This step is crucial because different types of SMS need to follow different compliance rules. For example, marketing SMS needs to comply with the relevant provisions of the Advertising Law, while notification SMS may need to pay attention to issues such as privacy protection and data security. Subsequently, based on the SMS type and user requirements, compare the SMS semantic parsing information with preset compliance rules such as industry standards, laws and regulations, and company policies. These rules can be in the form of textual descriptions or model-based rule libraries, covering multiple aspects such as the expression mode of SMS, the compliance of words used, and the disclosure of sensitive information. In addition, also check the grammar and format in the SMS semantic parsing information and analyze the potential risk descriptions therein. For example, marketing SMS may involve misleading publicity, false advertising, or unauthorized promotion content; service SMS may involve the leakage of user privacy or notifications without the consent of the user. Through multi-dimensional comparison, identify whether there are risk points in the SMS content that violate the compliance requirements. During the comparison process, specific risk description content can be identified, such as non-compliant advertising slogans (exaggerated marketing statements, false preferential information, etc.), overly exposed sensitive information (leakage of personal information, unauthorized promotions, etc.), marketing behaviors without clear consent (advertising or promotion activities without user authorization), and other compliance loopholes (the SMS does not provide an unsubscribe method or customer service contact information in the prescribed format, etc.). Through this process, not only can it be determined whether it is compliant according to the SMS type and user requirements, but also potential risk content can be accurately identified in the SMS, providing a basis for subsequent compliance review and optimization.
[0046] Optimize and adjust the risk description content according to the SMS type and user requirements according to the semantic parsing information of the risk content to generate the target SMS content.
[0047] In one embodiment, first, obtain the corresponding industry compliance constraints according to the type of the short message (such as marketing, notification, service, etc.), and analyze the expression characteristics of the short message according to the user's needs to ensure that the information transmission meets the user's expectations. Subsequently, use these industry compliance constraints and user need characteristics to construct a fitness evaluation function to quantify the matching degree between the short message content and the compliance standards and user needs. Based on this evaluation function, the short message content will be adjusted to obtain the short message content that best meets the compliance requirements and user needs, thereby generating the final target short message content. This target short message content complies with relevant laws, regulations and industry standards, avoids potential legal risks, better meets the user's needs, and ensures that the user's experience and rights and interests are not damaged.
[0048] Furthermore, the present application provides for optimizing and adjusting the risk description content according to the short message type and user needs according to the semantic parsing information of the risk content to generate the target short message content, including: Obtain industry compliance constraints according to the short message type; obtain user statement expression need characteristics according to the user needs; construct a fitness evaluation function according to the industry compliance constraints and user statement expression need characteristics, and construct an optimization strategy analysis space based on the fitness evaluation function to search for the short message content with the largest evaluation result that meets the industry compliance constraints and user statement expression need characteristics, and generate the target short message content.
[0049] Optionally, determine the industry compliance constraints associated with the SMS based on the SMS type. Different industries (such as finance, healthcare, government affairs, etc.) have different compliance requirements for SMS content. For example, the financial industry requires clear risk warnings to users, the healthcare industry requires avoiding false propaganda, and government announcements require information accuracy and transparency. Therefore, it is necessary to obtain corresponding compliance rules according to the industry type of the SMS. According to the user's needs, obtain the user's statement expression requirement characteristics and identify specific requirements therein, such as coherence, advertising promotion effect, etc. For example, advertising SMS needs to be fluent and attractive, avoid long or abrupt sentences, ensure clear and persuasive information transmission, and should highlight key information such as discounts and promotions, and clearly display the selling points that attract users (such as discounts, gifts, etc.); government affairs SMS requires accurate information content, avoiding ambiguity and misunderstanding. These user requirement characteristics can help the system understand the tone, style, and key content that the user hopes to convey, and ensure meeting user expectations when generating SMS. Subsequently, based on the industry compliance constraints and user statement expression requirement characteristics, a fitness evaluation function will be constructed. The main goal of this function is to measure the performance of the SMS content in meeting industry compliance requirements and user requirement expression, and it is constructed by weighting the compliance score, coherence score, and advertising effect score. Among them, the compliance score is obtained by performing the aforementioned sensitive word recognition operation on the optimized SMS content to obtain the number of sensitive words, then dividing this number of sensitive words by the maximum allowable sensitive threshold of the industry, and using 1 minus this quotient; the coherence score is obtained by evaluating the optimized SMS content by inputting it into an existing language model (such as GPT, GPT-3 of OpenAI, etc.); the advertising effect score is obtained by dividing the density of marketing keywords by the recommended density threshold, then comparing the calculated quotient with 1, and taking the smaller one. After that, based on the fitness evaluation function, an optimization strategy analysis space will be constructed. This analysis space aims to search for SMS content that meets industry compliance constraints and maximizes user requirement expression by considering various possible SMS content generation strategies, such as various expression methods (different tones, information organizational structures, etc.), compliance adjustment (adjusting the wording of the SMS content or adding necessary compliance statements), effect optimization (adjusting the order of advertising information or highlighting selling points), etc.By searching in this space, optimized SMS content is obtained, and the optimized SMS content is evaluated using the constructed fitness evaluation function to calculate the fitness value of the optimized SMS content. Then, this fitness is compared with the preset fitness to determine whether the fitness is greater than the preset fitness. If not, the optimized SMS content is discarded and new SMS content is generated again from the optimization strategy analysis space. Otherwise, the current optimized SMS content is used as the target SMS content. This target SMS content not only meets the compliance requirements of the industry but also maximally satisfies the needs of users, and can improve the intelligent level and compliance guarantee ability of SMS content generation in communication services.
[0050] Furthermore, the present application provides for generating the target SMS content, and further includes: Establish a multi-dimensional risk word library, and synchronously update the multi-dimensional risk word library according to the sensitive words of the SMS type and the update frequency of the compliance rules; combine the semantic rough recognition channel and the deep semantic parsing channel to construct a bilingual semantic understanding channel, and integrate the bilingual semantic understanding channel with the optimization strategy analysis space and the multi-dimensional risk word library to construct an SMS integrated generation module; perform integrated processing of semantic parsing, strategy optimization, and compliance conversion on the SMS content to be processed through the SMS integrated generation module to generate the target SMS content.
[0051] Optionally, a multi-dimensional risk word library is established in advance. This word library is used to store all sensitive words and risk words related to the SMS content. This word library will be updated according to the sensitive words of the SMS type and the update frequency of compliance rules, that is, it will be updated regularly according to the sensitive words of the new SMS type and the new industry compliance rules. The update frequency depends on the changes in industry regulations and the addition or deletion of sensitive words. For example, some advertising words may be disabled as "misleading advertisements", so it is necessary to keep track of relevant regulations in real time and update the sensitive word library. In these ways, a comprehensive and dynamically updated multi-dimensional risk word library is maintained, which can help subsequent compliance checks and risk assessments. Subsequently, the semantic rough recognition channel and the deep semantic parsing channel are combined to form a bilingual semantic understanding channel. The semantic rough recognition channel preliminarily classifies the SMS content through a multinomial logistic regression model to identify the industry type and user demand type of the SMS; the deep semantic parsing channel uses a deep neural network model to perform more refined semantic parsing on the SMS content, capture the complex semantic structure and context dependence in the text, and identify potential risks and compliance issues. Through this bilingual semantic understanding channel, the system can analyze the compliance and risks of the SMS from both the industry level and the text level to ensure that the SMS content complies with industry norms and there are no implicit violations semantically. After that, the bilingual semantic understanding channel is integrated with the optimization strategy analysis space and the multi-dimensional risk word library to form a unified SMS integrated generation module. The optimization strategy analysis space is used to perform multi-dimensional optimization on the SMS content according to user needs and industry compliance requirements; the multi-dimensional risk word library will detect sensitive words in the SMS in real time to ensure that the SMS does not contain illegal words. Through this SMS integrated generation module, the system can automatically optimize and adjust the SMS content after semantic parsing, risk assessment, and compliance detection. Then, with the support of the SMS integrated generation module, operations such as semantic parsing, strategy optimization, and compliance conversion are performed on the SMS content to be processed, so as to generate target SMS content that meets all industry compliance requirements and can accurately express user needs, ensuring its best performance in terms of compliance and marketing effects, thereby improving the efficiency and accuracy of SMS generation.
[0052] In summary, the embodiments of the present application at least have the following technical effects: In the embodiment of the present application, the text content to be processed by the service provider is first roughly recognized to obtain the text message type and user requirements; subsequently, the text content is input into the semantic large model for in-depth semantic analysis, and the text message semantic analysis information is output; then, the text message semantic analysis information is compared for compliance in multiple dimensions according to the text message type and user requirements to identify the risk description content; finally, the risk description content is optimized and adjusted according to the text message type and user requirements according to the semantic analysis information of the risk content to generate the target text message content. These technical effects together solve the technical problem that the generation of text message content in communication services depends on templates and cannot dynamically identify and avoid compliance risks according to user requirements, and achieve the technical effect of improving the intelligent level and compliance guarantee ability of text message content generation in communication services through in-depth semantic analysis and multi-dimensional risk comparison by the semantic large model.
[0053] Embodiment 2, based on the same inventive concept as the method for generating compliant text message content by integrating the semantic large model in the foregoing embodiment, as Figure 2 shown, the present application provides a system for generating compliant text message content by integrating the semantic large model, and the system includes: a rough recognition unit 11: roughly recognize the text message content to be processed by the service provider to obtain the text message type and user requirements; a deep semantic analysis unit 12: input the text message content into the semantic large model for in-depth semantic analysis, and output the text message semantic analysis information; a multi-dimensional compliance comparison unit 13: compare the text message semantic analysis information for compliance in multiple dimensions according to the text message type and user requirements to identify the risk description content; an optimization and adjustment unit 14: optimize and adjust the risk description content according to the text message type and user requirements according to the semantic analysis information of the risk content to generate the target text message content.
[0054] Furthermore, the rough recognition unit 11 is further configured to execute the following method: Obtain the industry type and text message demand type of text message compliance review; perform data tagging according to the industry type and text message demand type to construct training data; use the training data to train a logistic regression model to construct a semantic rough recognition channel, where the logistic regression model is used to identify and output the text message type and user requirements, and the text message type corresponds to the industry type, and the user requirements correspond to the text message demand type.
[0055] Furthermore, the rough recognition unit 11 is further configured to execute the following method: The industry type at least includes: the financial industry, the medical industry, and government announcements, and the text message demand type at least includes marketing, notification, and service.
[0056] Furthermore, the deep semantic analysis unit 12 is further configured to execute the following method: Build a large model framework, including an input layer, an encoder layer, a context understanding layer, a semantic parsing layer, and an output layer; train and converge the large model framework through the collected training dataset to obtain the semantic large model, and build a deep semantic parsing channel based on the semantic large model for deep semantic parsing of the SMS content to output the SMS semantic parsing information.
[0057] Further, the deep semantic parsing unit 12 is further configured to execute the following method: The input layer is used to process text embedding; the encoding layer uses a self-attention mechanism to capture the dependencies of the text, and allows the model to simultaneously focus on different parts of the text through multi-head attention; the context understanding layer is used to detect sensitive words and identify information related to local context and sensitive words; the semantic parsing layer uses multiple layers of encoders to extract deep semantic features and further analyze the deep association relationship between sensitive words and the global context; the output layer is used to integrate the sensitive words, deep semantic features, and context association relationship to output the SMS semantic parsing information.
[0058] Further, the deep semantic parsing unit 12 is further configured to execute the following method: Preprocess the SMS content, including cleaning and removing irrelevant characters, and input it into the semantic large model; the input layer converts the preprocessed SMS content into an embedding vector; the encoding layer captures the dependencies between each word in the SMS content based on the self-attention mechanism and assigns a weight to each word to represent its importance in the current context; the context understanding layer matches the SMS content with a list of sensitive words to identify sensitive words, and analyzes the surrounding local context information of the detected sensitive words based on the dependencies output by the encoding layer to determine the meaning and risk of the sensitive words in the current context; based on the identified sensitive words and local context relationships, the semantic parsing layer uses a multi-layer encoder structure to extract the deep semantic features of the SMS content and further analyze the deep association relationship between the sensitive words in the global context; connect the identified sensitive words, local context relationships, deep semantic features, and deep association relationships to output the SMS semantic parsing information.
[0059] Further, the deep semantic parsing unit 12 is further configured to execute the following method: Configure the moving window size according to the weight of sensitive words; perform local context selection of sensitive words based on the moving window, identify the co-occurrence frequency of sensitive words and directly associated words within the moving window size, analyze the lexical grammar relationship according to the dependency relationship between words, and identify the meaning and risk of the combination of directly associated words; adjust the window size according to the risk within the moving window and the similarity of the edge words of the window, and based on the adjusted moving window, identify the indirect association between the sensitive word and the intermediate word, and analyze the meaning and risk of the indirectly associated words according to the dependency relationship between words; comprehensively analyze the meaning and risk of the combination of directly associated words and the meaning and risk of indirectly associated words, and output the meaning and risk of the sensitive word in the current context.
[0060] Further, the optimization and adjustment unit 14 is further configured to execute the following method: Obtain industry compliance constraints according to the type of the short message; obtain the user statement expression requirement characteristics according to the user requirements; construct a fitness evaluation function according to the industry compliance constraints and the user statement expression requirement characteristics, and construct an optimization strategy analysis space based on the fitness evaluation function for searching for the short message content that meets the industry compliance constraints and has the largest evaluation result of the user statement expression requirement characteristics, and generate the target short message content.
[0061] Further, the optimization and adjustment unit 14 is further configured to execute the following method: Establish a multi-dimensional risk word library, and synchronously update the multi-dimensional risk word library according to the sensitive words and compliance rule update frequencies of the short message type; combine the semantic rough recognition channel and the deep semantic parsing channel to construct a bilingual semantic understanding channel, integrate the bilingual semantic understanding channel with the optimization strategy analysis space and the multi-dimensional risk word library, and construct a short message integrated generation module; perform integrated processing of semantic parsing, strategy optimization, and compliance conversion on the short message content to be processed through the short message integrated generation module, and generate the target short message content.
[0062] It should be noted that the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0063] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0064] This specification and the accompanying drawings are merely illustrative of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A method for generating compliant SMS content by integrating a semantic large model, characterized in that, Including: Coarsely identify the SMS content to be processed by the service provider to obtain the SMS type and user requirements; Input the SMS content into a semantic large model for in-depth semantic parsing and output SMS semantic parsing information; Perform multi-dimensional compliance comparison on the SMS semantic parsing information according to the SMS type and user requirements to identify risk description content; Optimize and adjust the risk description content according to the SMS type and user requirements according to the semantic parsing information of the risk content to generate the target SMS content.
2. The method for generating compliant SMS content by integrating a semantic large model according to claim 1, wherein Coarsely identify the SMS content to be processed by the service provider to obtain the SMS type and user requirements, including: Obtain the industry type and SMS requirement type for SMS compliance review; Perform data tagging according to the industry type and SMS requirement type to construct training data; Use the training data to train a logistic regression model to construct a coarse semantic recognition channel, where the logistic regression model is used to identify and output the SMS type and user requirements, and the SMS type corresponds to the industry type and the user requirements correspond to the SMS requirement type.
3. The method for generating compliant SMS content by integrating a semantic large model according to claim 2, wherein, The industry type includes at least: financial industry, medical industry, government announcements, and the SMS requirement type includes at least marketing, notification, and service.
4. The method for generating compliant short message content by integrating a semantic large model according to claim 2, wherein, Input the SMS content into a semantic large model for in-depth semantic parsing and output SMS semantic parsing information, including: Construct a large model framework, including an input layer, an encoder layer, a context understanding layer, a semantic parsing layer, and an output layer; Train and converge the large model framework through the collected training data set to obtain the semantic large model, and build an in-depth semantic parsing channel based on the semantic large model for in-depth semantic parsing of the SMS content and output the SMS semantic parsing information.
5. The method for generating compliant short message content by integrating a semantic large model according to claim 4, characterized in that, Construct a large model framework, including: The input layer is used to process text embedding; The encoding layer uses a self-attention mechanism to capture the dependencies of the text, and allows the model to simultaneously focus on different parts of the text through multi-head attention; The context understanding layer is used to detect sensitive words and identify information related to local context and sensitive words; The semantic parsing layer uses multiple layers of encoders to extract deep semantic features and further analyze the deep association relationship between sensitive words and global context; The output layer is used to integrate sensitive words, deep semantic features, and context association relationships and output SMS semantic parsing information.
6. The method for generating compliant SMS content by integrating a semantic large model according to claim 5, wherein Input the SMS content into a semantic large model for in-depth semantic parsing and output SMS semantic parsing information, including: Preprocess the SMS content, including cleaning and removing irrelevant characters, and input it into the semantic large model; The input layer converts the preprocessed SMS content into an embedding vector; Through the encoding layer, capture the dependencies between each word in the SMS content based on the self-attention mechanism and assign a weight to each word to represent its importance in the current context; The context understanding layer matches the SMS content with a list of sensitive words to identify sensitive words, and analyzes the surrounding local context information of the detected sensitive words based on the dependencies output by the encoding layer to determine the meaning and risk of the sensitive words in the current context; Based on the identified sensitive words and local context relationships, the deep semantic features of the SMS content are extracted through the semantic parsing layer using a multi-layer encoder structure, and the deep correlation relationships between sensitive words in the global context are further analyzed; Connect the identified sensitive words, local context relationships, deep semantic features, and deep correlation relationships, and output the semantic parsing information of the SMS.
7. The method for generating compliant short message content by integrating a semantic large model according to claim 6, wherein Determine the meaning and risk of sensitive words in the current context, including: Configure the mobile window size according to the weight of the sensitive word; Perform local context selection of sensitive words based on the mobile window, identify the co-occurrence frequency of sensitive words and directly related words within the mobile window size, analyze the lexical grammar relationship according to the dependency relationship between words, and identify the meaning and risk of directly related word combinations; Adjust the window size according to the risk within the mobile window and the similarity of the edge words of the window. Based on the adjusted mobile window, identify the indirect association between sensitive words and intermediate words, and analyze the meaning and risk of indirectly related words according to the dependency relationship between words; Comprehensively analyze the meaning and risk of directly related word combinations and the meaning and risk of indirectly related words, and output the meaning and risk of the sensitive word in the current context.
8. The method for generating compliant short message content by integrating a semantic large model according to claim 4, wherein Optimize and adjust the risk description content according to the SMS type and user requirements according to the semantic parsing information of the risk content to generate the target SMS content, including: Obtain industry compliance constraints according to the SMS type; Obtain the user statement expression requirement characteristics according to the user requirements; Construct a fitness evaluation function according to the industry compliance constraints and user statement expression requirement characteristics, and construct an optimization strategy analysis space based on the fitness evaluation function to search for the SMS content with the largest evaluation result that meets the industry compliance constraints and user statement expression requirement characteristics, and generate the target SMS content.
9. The method for generating compliant SMS content by integrating a semantic large model according to claim 8, wherein Generating the target SMS content also includes: Establish a multi-dimensional risk word library, and synchronously update the multi-dimensional risk word library according to the sensitive words of the SMS type and the update frequency of compliance rules; Combine the semantic rough recognition channel and the deep semantic parsing channel to construct a bilingual semantic understanding channel, and integrate the bilingual semantic understanding channel, the optimization strategy analysis space, and the multi-dimensional risk word library to construct an SMS integrated generation module; Perform integrated processing of semantic parsing, strategy optimization, and compliance conversion on the SMS content to be processed through the SMS integrated generation module to generate the target SMS content.
10. A short message content compliance generation system integrating a semantic large model, characterized in that, The system is used to execute the SMS content compliance generation method of the fusion semantic large model according to any one of claims 1-9, including: Coarse recognition unit: Coarsely recognize the SMS content to be processed by the service provider to obtain the SMS type and user requirements; Deep semantic parsing unit: Input the SMS content into the semantic large model for deep semantic parsing, and output the semantic parsing information of the SMS; Multi-dimensional compliance comparison unit: Perform multi-dimensional compliance comparison on the semantic parsing information of the SMS according to the SMS type and user requirements, and identify the risk description content; Optimization and adjustment unit: Optimize and adjust the risk description content according to the SMS type and user requirements according to the semantic parsing information of the risk content to generate the target SMS content.
Citation Information
Patent Citations
Big data analysis method applied to information security field
CN109660526A
Mobile terminal voice analysis system
CN112750426A
Method and device for intelligently and strictly selecting short message channel
CN118972791A
Intelligent financial question-answering system realized based on large language model
CN119441404A
Cited By
Text short message auditing method and device, computer equipment and storage medium
CN121030299A
Information security processing method and system for online content
CN121309143A
Robustness generative engine optimization method based on adversarial perception and security constraint
CN121562720A
Processing method and device for intelligent short message generation task
CN122242472A
A method and device for processing an intelligent short message generation task
CN122242472B