An artificial intelligence-based short message script sensitive word recognition analysis method

By constructing a cross-modal semantic alignment loss function and a dynamic sensitive word avoidance strategy, the problem of visual element recognition in rich media content of marketing SMS systems is solved, achieving efficient sensitive word recognition and compliant generation, and improving marketing effectiveness and system automation level.

CN121072550BActive Publication Date: 2026-02-17GUANGDONG BOJIN INFORMATION TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511598656.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-17
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

Existing marketing SMS systems struggle to identify potentially sensitive semantics within visual elements when processing rich media content, resulting in low compliance rates. Furthermore, multimodal semantic analysis technology lacks dynamic semantic consistency correction capabilities, making it prone to misjudging sensitive words and impacting marketing effectiveness.

Method used

An artificial intelligence-based approach is adopted, which extracts high-dimensional semantic feature vectors of SMS text through a pre-trained language model and extracts semantic representation vectors of images through a convolutional neural network. A cross-modal semantic alignment loss function is constructed to generate optimized text expression after semantic deviation correction. Dynamic sensitive word identification and avoidance strategies are then implemented in conjunction with a pre-set sensitive word library.

Benefits of technology

It significantly reduced the false positive and false negative rates of sensitive words, improved multimodal semantic consistency, increased compliance pass rate and marketing effectiveness, reduced manual review costs, and enhanced marketing ROI.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121072550B_ABST
    Figure CN121072550B_ABST
Patent Text Reader

Abstract

The application relates to an artificial intelligence-based SMS script sensitive word recognition analysis method, which aims at the technical problem of realizing rich media display and text content synthesis under compliance requirements of marketing SMS, proposes a deep semantic modeling and cross-modal semantic alignment mechanism based on multi-modal data, combines a pre-training language model and a convolutional neural network to respectively extract high-dimensional semantic features of text and images, and optimizes semantic consistency through an attention mechanism and a contrast loss function. Further, the system introduces a dynamic sensitive word recognition and avoidance strategy, effectively filters sensitive expressions, and generates compliance semantic equivalent content through a text reconstruction model when necessary. Finally, the content and the image are fused in format to generate a rich media SMS that can be sent in compliance. The scheme improves the semantic accuracy and content compliance of multi-modal information fusion, ensuring that marketing SMS meets legal compliance requirements and has good user appeal and communication effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of "multimodal artificial intelligence and semantic analysis technology", and in particular to an artificial intelligence-based method for identifying and analyzing sensitive words in SMS text messages. Background Technology

[0002] In the current field of intelligent compliance generation and sensitive word identification for marketing SMS messages, mainstream technologies are gradually moving towards a combination of deep learning models and multimodal semantic analysis. Existing marketing SMS systems mostly employ keyword dictionaries, regular expression matching, and simple contextual semantic rules to detect sensitive words in the text. These solutions can meet basic compliance requirements when processing plain text or single-modal content. However, with the widespread use of rich media marketing SMS messages (such as those containing text, images, and card-style advertisements), sensitive word identification and text compliance review face more complex multimodal fusion and semantic understanding challenges.

[0003] Currently, the following are the technological advancements and shortcomings in the field of multimodal content marketing:

[0004] 1) Existing SMS sensitive word detection systems mostly focus on the text dimension, limiting sensitive word detection to plain text rule matching, NER (Named Entity Recognition), and simple semantic similarity calculation. This makes it difficult to handle potential sensitive semantics contained in visual elements such as images and cards. For example, over-rendering of product promotional images, vulgar elements, or suggestive visual symbols are often not identified in a timely manner, leading to a decline in overall copywriting review capabilities and a low compliance rate.

[0005] 2) While multimodal marketing copy generation and compliance optimization technologies have incorporated text-image fusion models, semantic alignment discrepancies can easily arise between text and visual modalities in real-world deployment scenarios. Typical examples include text using compliant vocabulary while accompanying images suggest sensitive or illegal content, or promotional copy failing to align with the brand image, leading to misjudgments and blocking during SMS review, thus impacting user experience and campaign effectiveness.

[0006] 3) Mainstream multimodal semantic analysis technologies often employ static loss functions, lacking dynamic semantic consistency correction capabilities. They cannot automatically adjust copy structure or strategies based on actual SMS content and the user's delivery environment. Especially in highly sensitive areas such as e-commerce marketing and brand promotion, simple modal fusion lacks semantic bias measurement and intelligent avoidance strategies, leading to content being misjudged as illegal or sensitive, impacting marketing ROI and compliance approval rates. Furthermore, the industry lacks a systematic patented solution to the problem of "false detection of sensitive words caused by semantic mismatch between modalities."

[0007] 4) Existing replacement strategies mostly rely on sensitive word dictionary lookups and static synonym replacements, which are insufficient to handle the requirements of compliant expression in complex contexts, user profiles, and visual scenarios, and lack the ability to proactively generate and optimize structures. Sensitive word avoidance often remains at the word or phrase level, failing to perform semantic equivalence transfer or overall copy reconstruction, resulting in reduced appeal of the final marketing content or loss of marketing intent. Summary of the Invention

[0008] This application provides an artificial intelligence-based method for identifying and analyzing sensitive words in SMS text messages, aiming to solve one of the problems or issues of the existing technology mentioned in the background section.

[0009] This application provides an artificial intelligence-based method for identifying and analyzing sensitive words in SMS text messages, specifically including:

[0010] S1: Obtain the text content of the marketing SMS to be sent and the associated rich media image materials as multimodal input data.

[0011] S2: Perform semantic embedding processing on the SMS text and use a pre-trained language model to extract high-dimensional semantic feature vectors.

[0012] S3: Perform visual feature encoding on the rich media image and extract the image semantic representation vector using a convolutional neural network.

[0013] S4: Based on the text semantic feature vector and the image semantic representation vector, construct a cross-modal semantic alignment loss function to quantify semantic differences.

[0014] S5: Input the cross-modal semantic alignment loss function into the multimodal semantic consistency optimization model to generate the optimized text expression after semantic bias correction.

[0015] S6: Based on the optimized text expression after semantic deviation correction, and combined with the preset sensitive word library, a dynamic sensitive word identification and avoidance strategy is generated.

[0016] S7: The text content generated by the avoidance strategy is merged with the image materials to generate a rich media card, and then output to the SMS sending module.

[0017] S8: During the generation of rich media cards, if sensitive semantic residues still exist after semantic deviation correction, the text reconstruction mechanism is triggered to regenerate compliant semantically equivalent expressions.

[0018] This application provides an artificial intelligence-based method for identifying and analyzing sensitive words in SMS text messages, which has the following beneficial effects:

[0019] 1) Enhanced multimodal semantic consistency:

[0020] This invention is the first to deeply integrate cross-modal semantic feature alignment and mutual calibration mechanisms in the field of sensitive word avoidance. Through a series of modules such as semantic embedding (e.g., BERT text embedding and convolutional neural network visual encoding), semantic alignment loss function, attention-weighted fusion, and modality invariance constraints, it can effectively quantify and proactively eliminate the differences in expression and understanding between text and images. This method ensures the semantic consistency between SMS text and rich media images, significantly reduces the false positive and false negative rates of sensitive words, and solves the problem that existing technologies struggle to handle the complex contexts of rich media fusion.

[0021] 2) Dynamic sensitive word avoidance and intelligent expression optimization:

[0022] The system not only identifies common sensitive words but also utilizes semantic similarity calculation, dynamic threshold determination, and context-related replacement algorithms to effectively generate more compliant and marketing-oriented text. Addressing the risks of semantic residue and terminal interception, it can automatically trigger text reconstruction, using a semantically preservative sequence generation model and syntactic restructuring to ensure that the replaced text avoids sensitive points while maintaining its original marketing appeal. This dynamic, intelligent, and context-aware avoidance strategy is significantly superior to existing solutions that rely on fixed sensitive word libraries or word-level replacement, improving business applicability and automation intelligence.

[0023] 3) Significantly improved recognition accuracy and efficiency:

[0024] By using algorithms such as cosine similarity, contrastive learning loss, and maximum mean difference (MMD) to globally optimize the semantics of text and images, the false positive rate of cross-modal sensitive words is significantly reduced, and the overall compliance pass rate is greatly improved. All aspects of the system (text parsing, image processing, multimodal fusion, sensitive word identification and avoidance) can be processed in batches and efficiently. Compared with traditional manual proofreading or repeated rule correction methods, the time cost of automatic review and generation of every 10,000 SMS messages is significantly reduced, greatly improving the efficiency of marketing operations.

[0025] 4) Reduce compliance risks and improve productivity and resource utilization efficiency:

[0026] This invention significantly reduces the risks of terminal interception and brand complaints, minimizes campaign failures and customer churn caused by misjudgments of sensitive keywords, improves SMS delivery rates and user click-through rates, and achieves a higher marketing ROI. Simultaneously, batch automatic evaluation and replacement greatly reduce the manpower and time consumption of manual review and content correction, empowering content production and marketing teams and improving overall productivity and management efficiency. Attached Figure Description

[0027] Appendix Figure 1 This is the main flowchart of a method for identifying and analyzing sensitive words in SMS text messages based on artificial intelligence.

[0028] Appendix Figure 2This is a sub-flowchart of a method for identifying and analyzing sensitive words in SMS text messages based on artificial intelligence.

[0029] Appendix Figure 3 This is another sub-flowchart of a method for identifying and analyzing sensitive words in SMS text messages based on artificial intelligence. Detailed Implementation

[0030] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0031] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0032] As attached Figure 1 As shown, this application provides an artificial intelligence-based method for identifying and analyzing sensitive words in SMS text messages, specifically including:

[0033] S1: Obtain the text content of the marketing SMS to be sent and the associated rich media image materials as multimodal input data.

[0034] S2: Perform semantic embedding processing on the SMS text and use a pre-trained language model to extract high-dimensional semantic feature vectors.

[0035] S3: Perform visual feature encoding on the rich media image and extract the image semantic representation vector using a convolutional neural network.

[0036] S4: Based on the text semantic feature vector and the image semantic representation vector, construct a cross-modal semantic alignment loss function to quantify semantic differences.

[0037] S5: Input the cross-modal semantic alignment loss function into the multimodal semantic consistency optimization model to generate the optimized text expression after semantic bias correction.

[0038] S6: Based on the optimized text expression after semantic deviation correction, and combined with the preset sensitive word library, a dynamic sensitive word identification and avoidance strategy is generated.

[0039] S7: The text content generated by the avoidance strategy is merged with the image materials to generate a rich media card, and then output to the SMS sending module.

[0040] S8: During the generation of rich media cards, if sensitive semantic residues still exist after semantic deviation correction, the text reconstruction mechanism is triggered to regenerate compliant semantically equivalent expressions.

[0041] Step S1: Obtain the text content of the marketing SMS message to be sent and the associated rich media image materials as multimodal input data. Specifically, this includes:

[0042] S1.1: Perform structured parsing of the marketing copy text content from the SMS sending queue, extract the body text fields and target user tag information to obtain the raw text input for semantic modeling.

[0043] For marketing copy text content from SMS sending queues, a text parsing method based on regular expressions and context-free grammar (CFG) (parameters: regular expression template set, CFG production rule set) is used to decompose the message data structure and classify the fields.

[0044] Furthermore, by using the Named Entity Recognition (NER) algorithm (parameters: BiLSTM-CRF model, pre-trained word vector dimension 300), the algorithm accurately locates key semantic fragments such as brand name, product category, and promotion time in the parsed text, and obtains a structured set of semantic tags.

[0045] Furthermore, a key-value pair generation algorithm based on a field mapping table is adopted (parameters: the field mapping table includes text field mapping and user tag mapping rules) to map the recognition results into standardized text field information and target user tag information, and generate the corresponding metadata structure.

[0046] Furthermore, the semantic consistency test algorithm (parameters: window size = 5, similarity threshold = 0.85) is used to verify the semantic relevance between the text fields and user tags, and output the verification confidence score matrix.

[0047] By using structured data encapsulation processing, the text fields that have passed the consistency check are combined with the target user tag information and encapsulated into raw text input data objects, thereby achieving high-quality input generation for multimodal semantic modeling.

[0048] For example, in JSON-formatted marketing copy data from an SMS sending queue, the "content" field contains the complete promotional copy (approximately 120 characters long), and the "user_tags" field contains several user profile tags (such as "young women" and "sports enthusiasts"). The promotional validity period is extracted from the copy using a regular expression template \d{4}-\d{2}-\d{2}. The product keyword "running shoes" and brand "Brand X" are extracted using a NER model. A field mapping table is used to map "running shoes" to category tag ID 103, "Brand X" to brand tag ID 205, and "young women" to user tag ID U12. During the contextual semantic consistency check, the word embedding cosine similarity calculation method is used. The average similarity between the text fields and user tags is 0.88, which is greater than the threshold of 0.85, thus the consistency is deemed successful. In the final output of the original text input object, the body field is {"category_id":103, "brand_id":205, "promotion_date":"2024-06-15"}, and the user tag field is {"user_tag_ids":[U12]}, which provides accurate and standardized input for subsequent multimodal semantic embedding.

[0049] S1.2: Based on the marketing copy text content, identify its associated rich media image resource identifiers, including but not limited to image URLs, local storage paths, or cloud resource IDs, in order to locate the image input source used for multimodal fusion.

[0050] Based on the original text input data object obtained by structured parsing, a resource association retrieval algorithm based on relational mapping dictionary (parameters: mapping dictionary contains text keyword-image resource identifier pairs, matching similarity threshold = 0.85) is used to retrieve the corresponding rich media image resource identifier set from fields such as brand, product category, and promotion theme.

[0051] Furthermore, a fuzzy string matching algorithm (parameter: Levenshtein edit distance threshold = 2) is used to calculate the similarity between keywords in the text that contain mixed Chinese / English characters, differences between simplified and traditional characters, or similar spellings and entries in the resource mapping dictionary, and to obtain an expanded list of candidate resource identifiers for matching.

[0052] Furthermore, by using a contextual semantic similarity calculation method (parameters: cosine similarity calculated based on BERT embedding space, similarity threshold = 0.90), the semantic matching degree between candidate resource identifiers and the original text context is verified, and a set of accurate resource identifiers that conform to contextual constraints is generated.

[0053] Furthermore, the set of resource identifiers that have passed semantic verification is input into the Uniform Resource Locator (URL) parsing module. A multi-protocol parsing algorithm (parameters: supports HTTP / HTTPS, file URI, cloud storage API protocol) is used to identify the type of resource identifiers and standardize the access path, so as to obtain image URLs, local storage paths or cloud resource IDs in standard format.

[0054] By using a multi-source identifier normalization processing method (parameters: unified encoding format UTF-8, path separator standardization rules), the result of the previous step is transformed into a list of image input sources that can be directly called by the multimodal data loading module, thereby achieving accurate positioning and preparation for calling image resources.

[0055] For example, in the original text input data object containing the fields {"category_id":103, "brand_id":205, "promotion_date":"2024-06-15"}, the resource mapping dictionary specifies that category_id=103 corresponds to the keyword "running shoes", and brand_id=205 corresponds to the keyword "X brand". Their corresponding HTTP protocol image URLs are http: / / image.cdn / 103.jpg and http: / / image.cdn / 205.jpg, respectively. Edit distance calculation shows that the edit distance between the text "running shoes" and the resource dictionary entry for "running shoes" is 1, which is less than the threshold of 2, so the mapping result is retained. Next, based on BERT semantic embedding, similarity is calculated, and the similarity between "X brand" and the context "X brand running shoes special offer" is 0.94, which is greater than the threshold of 0.90, so this mapping pair is retained. Finally, the unified resource resolution module identifies http: / / image.cdn / 103.jpg as an HTTP protocol type, and the path is a standardized URL; the cloud resource ID 205_20240615 is resolved to the standard HTTPS path https: / / cloud.storage / img / 205_20240615.png through the API, and outputs a list of image input sources ["http: / / image.cdn / 103.jpg", "https: / / cloud.storage / img / 205_20240615.png"], providing standard input for loading multimodal image data.

[0056] S1.3: Obtain the raw pixel data of the rich media image material through HTTP protocol or local file system call to form an image input tensor that can be used for visual feature extraction.

[0057] S1.4: Perform image preprocessing operations on the original pixel data, including size normalization, color space conversion and noise filtering, to obtain a standardized image input matrix with a unified format.

[0058] S1.5: The standardized image input matrix is ​​encapsulated with the corresponding marketing copy text content through multimodal data alignment to construct a structured multimodal input sample for subsequent semantic embedding and visual encoding.

[0059] Step S2: Semantic embedding processing is performed on the SMS text, and a high-dimensional semantic feature vector is extracted using a pre-trained language model. Specifically, this includes:

[0060] S2.1: Preprocess the obtained original text content of the marketing SMS, including word segmentation, removal of stop words and standardization, to obtain a standardized text semantic unit sequence.

[0061] The original text content of the obtained marketing SMS messages is processed using a Chinese word segmentation algorithm based on Bidirectional Maximum Matching (BiMM) (parameter: forward matching dictionary size ≥ 10). 5 records, reverse matching dictionary size ≥ 10 5 (6 records, maximum matching length = 6) to achieve segmentation of continuous Chinese character sequences, splitting the original text into basic semantically independent word units.

[0062] Furthermore, by using a stop word filtering algorithm (parameter: stop word list size ≥ 5000 entries, covering general function words, function words and marketing noise words), words that have no actual semantic contribution in the word segmentation results are removed, and word sequence with high semantic information density is obtained, so as to reduce noise interference in the subsequent semantic modeling process.

[0063] Furthermore, the Unicode standardization algorithm (parameters: standardization form NFKC, character set covering Chinese full-width, half-width, uppercase, lowercase, traditional and simplified Chinese conversion rules) is adopted to achieve unified encoding conversion of heterogeneous characters, different encoding forms and uppercase and lowercase differences in the word sequence, and generate intermediate text representation with consistent character encoding.

[0064] By using a hash mapping index generation algorithm (parameter: the hash function seed value is fixed at 2024 to avoid encoding inconsistencies between multiple runs), the standardized word sequence is transformed into a unique index ID sequence, constructing a standardized text semantic unit sequence, thereby achieving consistency and efficiency in word vector search and encoding in the subsequent BERT embedding model.

[0065] Exemplarily, for the original marketing SMS text "X-brand running shoes are on sale at a special price for a limited time, only sold for 299 yuan, the event lasts until 2024-06-15!", under the BiMM Chinese word segmentation algorithm, the word segmentation sequence ["X-brand", "running shoes", "limited time", "special price", "only sold", "299", "yuan", "event", "until", "2024-06-15"] is obtained; by using the stop word filtering algorithm, the word items without semantic contribution "only sold", "event", "until" are deleted, and the effective word tokens ["X-brand", "running shoes", "limited time", "special price", "299", "yuan", "2024-06-15"] are obtained; by using the Unicode normalization algorithm, the possible full-width number "299" is converted to the half-width "299", and the traditional Chinese character "優惠" is replaced with the simplified Chinese character "优惠"; based on the regular rules, "299 yuan" is mapped to " <price>This maps "2024-06-15" to "". <date>, and get the sequence after placeholder replacement ["Brand X", "Running Shoes", "Limited Time Offer", "Special Offer", " <price> "," <date>"];Finally, generate a fixed-length index ID sequence through hash mapping, such as [100234, 100876, 100543, 100321, 100001, 100002], as the standardized text semantic unit sequence input BERT embedding processing, realize the clean, unified and quickly indexed semantic model input, and improve the modeling consistency and generalization ability of cross-sample scripts.

[0066] S2.2: Based on the standardized text semantic unit sequence, a BERT pre-training language model is used to perform forward propagation calculation of the embedding layer to obtain an initial word-level semantic embedding vector matrix.

[0067] S2.3: Perform context-aware Transformer encoding processing on the initial word-level semantic embedding vector matrix, and fuse context semantic dependencies based on the self-attention mechanism to generate context-enhanced semantic embedding representation.

[0068] S2.4: Based on the context-enhanced semantic embedding representation, use the pooling operation to extract the text global semantic feature vector to obtain high-dimensional semantic feature expression.

[0069] S2.5: Perform normalization processing on the high-dimensional semantic feature expression, and generate a standardized semantic feature vector based on the L2 norm standardization method to serve as the input basis for the subsequent cross-modal semantic alignment module.

[0070] The step S3: performing visual feature coding on the rich media image, using a convolutional neural network to extract image semantic representation vector. As shown in Figure 2 , specifically comprising:

[0071] S3.1: Perform image preprocessing operation on the rich media image, including size normalization, color space conversion and noise suppression, to obtain standardized image input of uniform format, providing basic image data for subsequent feature extraction.

[0072] S3.2: Based on the standardized image input, perform multi-level convolution operation on the image using a pre-trained convolutional neural network model to extract local texture features and global structure features of the image, and generate an original convolutional feature map.

[0073] On the standardized image input, a pre-trained convolutional neural network model (parameters: network structure ResNet-50, input size = 224x224 pixels, convolution kernel size = {3x3, 5x5}, stride = {1, 2}, padding method = same) is used to perform first layer convolution operation, realize the capture of low-level edge features and basic texture patterns, and generate an initial convolutional response map.

[0074] Further, higher-level local texture information and regional shape features are extracted layer by layer through the second to fourth convolution modules (each module includes a convolution layer, a batch normalization layer and a ReLU activation layer, the convolution kernel size is fixed at 3x3, and the number of channels increases according to {64, 128, 256}), and the spatial resolution is preserved for fine-grained feature analysis.

[0075] Further, a cross-channel feature fusion bottleneck residual unit (parameters: 1x1 dimension reduction convolution + 3x3 convolution + 1x1 dimension increase convolution, residual direct connection) is used to reuse features and optimize gradient flow for the intermediate feature map, reduce the risk of gradient vanishing in deep training, and enhance the semantic representation ability.

[0076] Further, a multi-scale convolution kernel combination strategy (convolution kernel size = {1x1, 3x3, 5x5} parallel calculation) is used to capture context information in different receptive field ranges in a single level, realize early integration of global structure features, and generate composite feature mapping that adapts to multiple image semantics.

[0077] Through forward propagation of the whole network, the original convolution feature map with local texture details and global layout patterns is obtained, which is used as the input of the subsequent pooling and high-dimensional semantic mapping steps to realize comprehensive encoding of multi-level semantic information.

[0078] For example, for a standardized marketing product image with an input size of 224x224x3, ResNet-50 is used as the pre-trained convolutional neural network, the first convolution kernel size is 7x7, the stride is 2, and the output feature map size is 112x112x64; the second to fourth convolution modules output feature maps with sizes of 56x56x128, 28x28x256 and 14x14x512 respectively. When performing bottleneck residual unit processing, the number of channels is reduced to 1 / 4 through 1x1 convolution, then 3x3 convolution is used to extract context features, and finally 1x1 convolution is used to restore the original number of channels to realize efficient expression of feature channels. In multi-scale convolution kernel combination operation, 1x1, 3x3 and 5x5 convolution kernels are used in parallel for the same input feature, and the outputs are concatenated in the channel dimension to obtain a feature map with a comprehensive receptive field, and the number of channels after concatenation is increased to 3 times the original number. Under this process, the system can accurately extract key visual patterns such as shoe texture, brand logo geometry and overall color layout, and after subsequent pooling processing, the matching accuracy is improved by about 12%, providing high-quality visual features for cross-modal semantic alignment.

[0079] S3.3: Perform a pooling operation on the original convolution feature map to reduce the feature dimension and preserve key visual semantic information through a max pooling algorithm, and obtain a reduced image feature representation.

[0080] For the original convolutional feature map, a two-dimensional max pooling algorithm (parameters: pooling kernel size = 2×2, stride = 2, boundary padding method = valid) is used to select the maximum value of the feature response in adjacent local regions, retaining the strongest response signal in each local region to enhance key visual patterns. This processing method effectively reduces the spatial resolution of the feature map while reducing redundant feature information, forming a preliminary dimensionality-reduced image feature map.

[0081] Furthermore, a max-pooling algorithm optimized based on the channel attention mechanism (parameters: channel weight update function = Softmax, normalization coefficient ε = ...) is used. This mechanism dynamically weights the importance of each feature channel during dimensionality reduction, ensuring that channels with high semantic relevance are prominently represented in the pooling results. This approach preserves visual features highly correlated with the target semantics while reducing dimensionality, thus improving cross-modal semantic alignment.

[0082] Furthermore, a block-adaptive max pooling method is adopted (parameter: output block size = {7×7}) to map input feature maps of different sizes to output tensors of uniform size, ensuring that the feature dimensions can be uniformly input to subsequent fully connected layers when processing image data of different original resolutions, thus achieving batch processing compatibility.

[0083] Furthermore, by combining edge-preserving filtering with pooling preprocessing (parameters: filter radius r=2, spatial weight σs=1.5, intensity weight σr=0.5), low-frequency background interference is weakened before compressing spatial resolution, while preserving the intensity of high-frequency features at object edges. This measure effectively prevents the key target contours from being weakened after max pooling.

[0084] Through the above chained processing, the original convolutional feature map is transformed into an image feature representation with reduced dimensionality and highly concentrated semantic information, achieving the technical effect of reducing computational load while preserving discriminative visual patterns.

[0085] For example, for an original convolutional feature map with an input size of 14×14×512, a 2D max pooling (2×2 kernel, stride 2) is first applied to obtain a feature tensor with an output size of 7×7×512. During channel attention weighting, Softmax regularization ensures that the sum of channel weights is 1; for example, the top 10 most important channels have an aggregated weight of 0.35 to highlight local product features relevant to the marketing theme. Under block adaptive max pooling with a unified size mapping, feature maps from 28×28×256 and 14×14×512 are all output as 7×7×C tensors, achieving batch unified input. During edge-preserving filtering, high-frequency response values ​​(such as the edge of the shoe logo) are reduced by less than 2%, effectively avoiding the attenuation of important textures. After this step, the feature dimension is reduced by approximately 75%, and the cross-modal semantic alignment module shows an average similarity improvement of approximately 8% in accuracy evaluation, verifying the effectiveness of this dimensionality reduction pooling strategy.

[0086] S3.4: Based on the reduced-dimensional image feature representation, a fully connected layer is used for high-dimensional semantic mapping to generate a fixed-dimensional image semantic embedding vector as a high-level semantic abstraction of the image content.

[0087] Based on the reduced-dimensional image feature representation, a fully connected layer high-dimensional semantic mapping algorithm (parameters: input dimension = 7×7×C, output dimension = 1024, weight initialization method = He Normal, activation function = ReLU) is used to transform the two-dimensional spatial feature tensor into a one-dimensional high-dimensional feature vector for global semantic integration.

[0088] Furthermore, through a batch normalization algorithm (parameters: momentum coefficient = 0.9, ε = ...), ... This stabilizes the numerical distribution of the output features of the fully connected layer, alleviates the vanishing and exploding gradient phenomena, and improves the convergence speed and generalization ability of the network during the training phase.

[0089] Furthermore, the Dropout random deactivation algorithm (parameter: deactivation probability p=0.5) is adopted to randomly mask the output of some neurons in the feature mapping vector layer, thereby reducing the risk of overfitting the model to some feature patterns and balancing the weight contribution of features from different channels.

[0090] Furthermore, combining the linear transformation matrix With bias vector Affine mapping formula:

[0091]

[0092] Achieve input feature Weighted combination generates high-dimensional embedding vectors ,in The input feature dimension, To output the feature dimension, These are the weighting coefficients. This is a bias term.

[0093] Furthermore, L2 regularization penalty terms are applied. The weight matrix is ​​constrained to control the magnitude of the weights to prevent overfitting and maintain robustness in the high-dimensional semantic space. This is the regularization coefficient.

[0094] Through the above fully connected layer mapping and regularization processing, the dimensionality-reduced image feature representation is transformed into a fixed-dimensional image semantic embedding vector, realizing a high-level abstract representation of the image content and forming input features that can directly participate in cross-modal semantic alignment operations.

[0095] For example, for a dimensionality-reduced feature map with an input dimension of 7×7×512 ( =512), first flatten it into a one-dimensional vector of length 25088, and input it into a fully connected layer (output dimension = 1024), with the weights initialized as a He Normal distribution (mean 0, variance 512). In the batch normalization step, the output is standardized using the mini-batch mean and variance, and the expressive power of the feature distribution is restored through trainable scaling and offset parameters. In the Dropout process, the output feature elements are randomly set to zero with a probability of p=0.5 to reduce the network's tendency to depend on a certain feature subset. Finally, the fully connected weight matrix constrained by L2 regularization (λ=0.01) provides a stable, dense image semantic embedding vector with global awareness. In the cross-modal similarity calculation task, compared with the original feature vector without fully connected mapping, the image-text matching accuracy is improved by about 9%, verifying the effectiveness of this high-dimensional semantic mapping strategy.

[0096] S3.5: Perform normalization processing on the image semantic embedding vector, and use the L2 normalization method to standardize the vector to obtain an image semantic representation vector that can be used for cross-modal semantic alignment, as the input of the subsequent semantic alignment module.

[0097] Step S4: Based on the text semantic feature vector and the image semantic representation vector, construct a cross-modal semantic alignment loss function to quantify semantic differences. For example... Figure 3 As shown, it specifically includes:

[0098] S4.1: Perform cross-modal similarity calculation on the text semantic feature vector and the image semantic representation vector, and use the cosine similarity algorithm to measure the degree of semantic matching between different modalities to obtain the cross-modal semantic similarity matrix.

[0099] Based on the acquired text semantic feature vector and image semantic representation vector, the cosine similarity calculation method (parameter: similarity measurement function = CosineSimilarity, numerical range = [-1,1]) is used to measure the directional consistency of the two modal vectors in the high-dimensional semantic space, which serves as the initial quantitative basis for the degree of cross-modal matching.

[0100] Furthermore, the cosine of the angle between two vectors is calculated using vector dot product and norm normalization, with the following formula:

[0101]

[0102] in, This is a text semantic feature vector. It is an image semantic representation vector. For vector dimensions.

[0103] Furthermore, a cross-modal semantic similarity matrix is ​​constructed through batch similarity calculation (parameters: batch size = B, dimension = d). , where matrix elements Indicates the first The first text sample and the second The cosine similarity value of each image sample.

[0104] Furthermore, a numerical stabilization method (parameters: stabilization function = Min-Max Scaling, target interval = [0,1]) is used to normalize each element of the similarity matrix, eliminating the impact of numerical distribution differences between different batches on subsequent loss calculations.

[0105] Furthermore, combining sample pairing and annotation information (parameter: annotation matrix) , =1 indicates semantic relevance. The similarity matrix is ​​masked and weighted to reduce the weight of unpaired items in the downstream loss calculation, thereby strengthening the semantic alignment contribution of paired samples.

[0106] By using the above algorithm, the feature vectors between text and image modalities are mapped into a unified cross-modal semantic similarity matrix, enabling a quantifiable description of the degree of multimodal matching and providing accurate basic data for the construction and optimization of subsequent semantic alignment loss.

[0107] For example, for a text-image input with a batch size of 16, where both the text feature vector and the image feature vector have a dimension of 1024, when calculating the similarity using the cosine similarity formula, the dot product operation of the numerator partial vectors is first performed to obtain an intermediate result matrix of length 16×16. For instance, in... , The value at that position is 518.42; then calculate the L2 norm of the corresponding vectors, such as... =22.77, =23.18, dividing the numerator value by the norm product yields a cosine similarity value of 0.9987. After batch calculations, an initial similarity matrix is ​​obtained. After Min-Max normalization to the [0,1] interval, the value at the above position is updated to 0.9812; combined with the annotation matrix The value is multiplied by a weight of 1 (because it is a real paired sample), while unpaired items are multiplied by a weight of 0.2 to reduce interference. The resulting cross-modal semantic similarity matrix shows higher discriminativeness and stability in the subsequent contrastive learning loss input, improving the accuracy and robustness of cross-modal alignment loss calculation in sensitive word avoidance scenarios.

[0108] S4.2: Based on the cross-modal semantic similarity matrix, construct an initial semantic alignment loss term, and use a contrastive learning loss function to quantify the semantic deviation between the text and image modalities to generate cross-modal semantic alignment error values.

[0109] Based on the cross-modal semantic similarity matrix, a contrastive learning loss function (parameters: loss form = InfoNCE Loss, temperature parameter τ = 0.07) is used to achieve a quantitative measurement of the semantic deviation between text and image modalities.

[0110] Furthermore, by comparing and optimizing the similarity difference between paired and unpaired samples, the similarity difference between the true paired item and all candidate items is calculated, and an exponential scaling function is used to suppress the excessive influence of large similarity differences on gradient updates.

[0111] S4.3: Perform weighted fusion processing on the cross-modal semantic alignment error value, and combine the attention mechanism to dynamically assign weights to the multimodal semantic features to generate a weighted semantic alignment loss value.

[0112] Based on the cross-modal semantic alignment error value, a multimodal feature weighted fusion algorithm is adopted (parameters: fusion strategy = attention-weighted fusion, learning rate = ...). With a regularization coefficient of 0.01, dynamic weight allocation of text semantic features and image semantic features is achieved, thereby enhancing the flexibility and accuracy of intermodal semantic matching.

[0113] Furthermore, a self-attention weight calculation module is constructed (parameters: weight calculation function = Scaled Dot-Product Attention, scaling factor = ...). The similarity contribution between each text feature vector and image feature vector is evaluated separately, and an inter-modal interaction weight matrix is ​​generated to highlight modal pairings with high semantic contribution.

[0114] Furthermore, the attention weights are calculated using the following formula:

[0115]

[0116] in, For the first Each text feature vector For the first Image feature vectors, For feature dimension, The total number of samples.

[0117] Furthermore, the attention weight matrix and the cross-modal semantic alignment error value are multiplied element-wise to achieve weighted correction of the semantic alignment error value and generate a preliminary weighted semantic alignment loss matrix.

[0118] Furthermore, a multi-head attention mechanism (parameters: number of heads = 8, dimension per head = 128) is adopted to compute multiple sets of weighted results in parallel. The outputs of different attention heads are concatenated on the feature dimension and then input into the linear transformation layer to integrate the weighted effects of different subspaces and obtain the fused weighted semantic alignment loss value.

[0119] Through the aforementioned weighted fusion and attention allocation mechanism, the initial cross-modal semantic alignment error value is transformed into a weighted semantic alignment loss index that focuses on optimizing high-contribution features, thereby improving the technical effect of cross-modal matching robustness and alignment accuracy in sensitive scenarios.

[0120] S4.4: Based on the weighted semantic alignment loss value, a modality invariance constraint term is introduced, and the maximum mean difference (MMD) algorithm is used to measure the difference between different modal distributions in order to further optimize cross-modal semantic consistency.

[0121] S4.5: Jointly optimize the weighted semantic alignment loss value with the modality invariance constraint term to construct the final cross-modal semantic alignment loss function, so as to output a unified multimodal semantic deviation quantification index.

[0122] Step S5: Input the cross-modal semantic alignment loss function into the multimodal semantic consistency optimization model to generate the semantically bias-corrected optimized text representation. Specifically, this includes:

[0123] S5.1: Input the cross-modal semantic alignment loss function into the multimodal semantic consistency optimization model to initialize the semantic deviation between the text semantic feature vector and the image semantic representation vector, so as to obtain the initial semantic deviation metric.

[0124] S5.2: Based on the initial semantic deviation metric, the gradient backpropagation algorithm is used to iteratively update the text generation parameters in the multimodal semantic consistency optimization model in order to reduce the semantic differences between text and images.

[0125] Based on the initial semantic deviation metric, the gradient backpropagation algorithm (parameters: optimization objective function = cross-modal semantic alignment loss function, optimizer type = Adam, initial learning rate = 1e-4, momentum parameters β1 = 0.9, β2 = 0.999) is used to realize the gradient calculation and update of text generation parameters in the multimodal semantic consistency optimization model.

[0126] Furthermore, by using an automatic differentiation mechanism to backpropagate the gradient information of the cross-modal semantic alignment loss function in the computational graph, the weight parameters acting on each layer of the text generation network are obtained. gradient .

[0127] Furthermore, the parameter update formula is adopted:

[0128]

[0129] in, For the first The parameter values ​​for the next iteration. The current learning rate, The loss gradient of the corresponding layer parameters is superimposed on the weight update to reduce cross-modal semantic differences.

[0130] Furthermore, a dynamic learning rate adjustment strategy (parameters: decay factor = 0.95, adjustment step size = 2000 steps) is used to gradually reduce the learning rate during the iteration process to balance the model's convergence speed and stability. Gradient clipping (parameters: maximum norm = 5.0) is used to prevent the negative impact of gradient explosion on parameter optimization.

[0131] Furthermore, by adding an L2 regularization term By constraining the size of the weight parameters within a reasonable range, the risk of overfitting can be mitigated, and the consistency of multimodal representation can be maintained.

[0132] Through the aforementioned gradient backpropagation and parameter iterative update process, the weight parameters of the model's text generation part are optimized in the direction of reducing the semantic differences between text and images, thereby achieving semantic consistency enhancement under multimodal conditions.

[0133] For example, for the initial semantic deviation metric Given a batch of text-image samples, assume the text generation network contains three Transformer encoding layers, with a total number of parameters. Driven by the Adam optimizer, the first layer weight matrix... The corresponding gradient norm is 1.37, and the learning rate is... At that time, the update volume was After updating the formula with applied parameters, An element in the dataset is updated from 0.25614 to 0.2560063. After applying dynamic learning rate decay, the learning rate decreases to [value missing] at the 4000th iteration. As the update step size of the same element decreases accordingly, the convergence curve tends to be stable.

[0134] S5.3: An attention mechanism is introduced during the parameter update process to focus on correcting semantic units in the text that are highly related to the semantics of the image, thereby generating semantically aligned enhanced text feature representations.

[0135] In the multimodal semantic consistency optimization model after parameter update, a self-attention mechanism is adopted (parameters: attention type = Scaled Dot-Product Attention, scaling factor). ) Calculate the correlation score between the text semantic feature vector and the image semantic representation vector to achieve refined quantification of the semantic correlation between modalities.

[0136] S5.4: Based on semantic alignment-enhanced text feature representation, a sequence-to-sequence generation model is used to generate preliminary optimized text representations to form candidate text representations with enhanced semantic consistency.

[0137] S5.5: Perform semantic consistency verification on the candidate text optimization expression. If the semantic consistency score is higher than the set threshold, output the semantically bias-corrected text optimization expression; otherwise, return to the gradient update step for further optimization.

[0138] Step S6: Based on the semantic deviation-corrected text optimization expression, and combined with a preset sensitive word library, a dynamic sensitive word identification and avoidance strategy is generated. Specifically, this includes:

[0139] S6.1: Perform lexical and syntactic structural analysis on the optimized expression of the text after semantic deviation correction, and use a context-aware sensitive word matching rule set to perform preliminary sensitive semantic screening in order to generate a preliminary sensitive word candidate set.

[0140] S6.2: Perform semantic similarity calculation on the candidate terms in the preliminary sensitive word candidate set. Based on the semantic anchor points of the predefined sensitive word library in the BERT semantic vector space, determine the semantic association strength between the candidate words and the sensitive semantics to generate semantic matching score results.

[0141] Based on the candidate terms in the preliminary sensitive word candidate set, the BERT semantic embedding method (parameters: model version = BERT - Base, uncased, embedding dimension = 768) is used to generate context - aware high - dimensional semantic vector representations for each candidate word and the word terms in the predefined sensitive word library respectively, realizing the accurate embedding of vocabulary in the semantic space.

[0142] Furthermore, through vector normalization processing (method: L2 - norm standardization), the above - mentioned semantic vectors are mapped to the unit hypersphere to ensure that the subsequent similarity calculation is not affected by the difference in vector modulus lengths.

[0143] Furthermore, the cosine similarity calculation method (formula: ) is used to calculate the semantic proximity between any candidate word semantic vector and each sensitive word semantic anchor vector in the sensitive word library to obtain a two - way similarity quantification result.

[0144] Furthermore, through the maximum similarity aggregation strategy, the maximum value of the cosine similarity values between each candidate word and all sensitive words in the sensitive word library is taken as the semantic association strength index of this candidate word , reflecting the degree to which this candidate word is closest to the sensitive semantics.

[0145] Furthermore, is subjected to min - max normalization processing.

[0146] Through the above - mentioned semantic embedding, similarity measurement and normalization scoring processing methods, the context semantics of candidate words are accurately mapped to a quantified sensitive semantic proximity index, realizing a robust evaluation of the sensitive word semantic matching degree in a dynamic context.

[0147] Exemplarily, in the e - commerce preferential text message, the preliminary sensitive word candidate set contains 3 word terms:

Flash Sale

Limited Time

Free Order

Flash Sale

Free Order

Flash Sale

Violent Price Slashing

Ultra - Fast Purchase

Act Immediately

Flash Sale

Limited Time

Free Order

[0148] S6.3: Construct a sensitive word identification decision model based on semantic matching score results, and use a threshold dynamic adjustment mechanism to determine whether candidate words belong to sensitive words in the current context, so as to generate a dynamically identified sensitive word set.

[0149] S6.4: For each sensitive word in the dynamically identified sensitive word set, a semantic equivalent substitution strategy is executed to generate multiple compliant expression candidates based on a pre-trained synonym semantic generation model, so as to generate an evasion candidate expression set.

[0150] S6.5: Based on the overall semantic coherence constraint mechanism of the copy, the candidate expressions in the avoidance candidate expression set are evaluated for contextual consistency, and the compliant expression with the best semantic match and consistent style is selected for replacement to generate the text content optimized by the avoidance strategy.

[0151] Step S7: The text content generated by the avoidance strategy is merged with image materials to generate a rich media card, which is then output to the SMS sending module. Specifically, this includes:

[0152] S7.1: Based on the text content generated by the avoidance strategy, extract its structured semantic expression and keyword tags to obtain the text semantic input for rich media card layout modeling.

[0153] S7.2: Execute a card template matching algorithm on the text semantic input, and select an appropriate rich media card style template based on the target user tag information to form a visual presentation structure that conforms to user preferences.

[0154] S7.3: Visually align the standardized image input matrix of the rich media image material with the selected card style template, and use image cropping and scaling algorithms to generate image display units that fit the template size, so as to ensure the semantic integrity and visual appeal of the image content in the card.

[0155] A pixel-level mapping relationship is established between the standardized image input matrix of the rich media image material and the visual placeholder unit of the selected card style template. A feature point matching algorithm (parameters: feature extraction operator = ORB, matching strategy = Hamming distance nearest neighbor) is used to achieve the initial positioning of the main image region in the template space.

[0156] Furthermore, by employing a scaling algorithm (parameters: interpolation method = bicubic interpolation, target size = template placeholder size), the original image is matched to the template size requirement while maintaining the aspect ratio without distortion, resulting in a pre-scaled image data matrix. .

[0157] Furthermore, a center-cropping algorithm is employed (parameters: cropping window size = template placeholder unit width and height, cropping center = image geometric center), from... Extract the core visual region required by the template and generate a cropped image matrix. Ensure that the main content is preserved intact.

[0158] Furthermore, using an image edge-preserving smoothing algorithm (parameters: kernel size = 3×3, filter type = bilateral filter, σ = 50), the image edge is further smoothed. Edge smoothing is performed to reduce jagged edges and edge artifacts caused by scaling and cropping, and a smoothed image matrix is ​​generated. .

[0159] Furthermore, a color histogram matching model is constructed (parameters: color space = Lab, matching algorithm = histogram cumulative distribution mapping), which will... The color distribution is adjusted to match the main color of the template, generating an image display unit with adjusted colors. .

[0160] Through the above-mentioned multi-level visual alignment and optimization processing, the standardized image input matrix of the previous step is transformed into a high-quality image display unit that conforms to the card template size and visual style, so as to achieve semantic integrity and visual appeal of rich media cards when displayed on the terminal.

[0161] For example, in an e-commerce marketing scenario, the obtained standardized image input matrix size is 800×600 pixels, the card-style template image placeholder unit size is 400×300 pixels, and the ORB feature extraction is used to detect the bounding box coordinates of the main image region (50,40)-(750,560). The scaling algorithm calculates the aspect ratio to be 1.33, and the matching ratio with the template is 1.33. Therefore, the width and height are scaled to 400×300 according to the aspect ratio to generate... The center crop window is set to 400×300 pixels, and the centered area is directly cropped to obtain the desired result. After bilateral filtering (σ=50), the average gradient decrease of edge artifact pixels was 12%. After Lab color space histogram matching, the color difference ΔE decreased from 9.2 to 1.3, forming... .final After being embedded into the card template, in display tests on 50 Android devices, the image display integrity rate was 100%, the main color consistency rate was over 95%, and the user click-through rate increased by about 12% compared to the unoptimized version.

[0162] S7.4: Based on the text semantic expression and image display unit, execute the multimodal content fusion rendering engine, and use the HTML5+CSS3 technology stack to build a structured rich media card document to generate interactive card content that can be parsed and displayed by the SMS client.

[0163] S7.5: Encapsulate the interactive card content into a multimedia message data packet conforming to the MMS protocol specification, and output it to the SMS sending module through the SMS gateway interface to complete the compliant sending process of rich media SMS.

[0164] Step S8: During the rich media card generation process, if it is determined that sensitive semantic residues still exist after semantic deviation correction, a text reconstruction mechanism is triggered to regenerate a compliant semantically equivalent expression. Specifically, this includes:

[0165] S8.1: Perform sensitive semantic residue detection on the optimized expression of the text after semantic deviation correction, and calculate the sensitivity score of the current text based on the preset sensitive semantic threshold model to determine whether there are still semantic residues that may trigger the terminal interception mechanism.

[0166] S8.2: If the sensitivity score is higher than the preset threshold, a sensitive semantic residual identifier signal is generated as the start condition for triggering the text reconstruction mechanism, so as to start the subsequent compliant semantic equivalent expression generation process.

[0167] S8.3: Based on the aforementioned sensitive semantic residual identifier signal, invoke the semantically preserving text reconstruction model. This model adopts a sequence-to-sequence generation framework based on the Transformer architecture to perform local semantic replacement and structural rearrangement on the current text in order to generate semantically equivalent but differently expressed alternative text.

[0168] Based on the aforementioned sensitive semantic residual identifier signal, a semantically preserving text reconstruction model is invoked. This model employs a sequence-to-sequence generation framework based on the Transformer architecture to perform local semantic replacement and structural rearrangement on the current text. A multi-head self-attention mechanism (parameters: number of attention heads = 8, hidden layer dimension = 512) is used to model the global dependencies of the input text and generate an initial context-dependent encoding vector set. .

[0169] Furthermore, through a cross-attention mechanism (parameters: number of attention heads = 8, attention projection dimension = 512), the following is achieved: Embedded with sensitive word location information Alignment is performed to enhance the model's ability to focus on potentially sensitive semantic regions and to obtain an enhanced context-encoded vector set. .

[0170] Furthermore, a gating unit (parameters: update gate activation function = sigmoid, reset gate activation function = tanh) is used to... Perform semantic substitution control and generate a substitution control matrix. By combining with the original word embedding sequence Element-wise weighted calculation:

[0171]

[0172] in, For the replaced embedded sequence, For the sequence of compliant equivalence word embeddings generated by the decoder, This indicates the Hadamard product operation.

[0173] Furthermore, a sequence generation process based on a Transformer decoder is employed (parameters: number of layers = 6, feedforward network dimension = 2048) to drive... Generate new text sequences based on semantic coherence. Furthermore, positional encoding is introduced during the generation process to preserve word order information.

[0174] Furthermore, the structural rearrangement module (algorithm: dependency parsing tree-based reordering, parameter: rearrangement threshold = 0.75) is used to... Optimize the syntax to ensure that the replaced copy maintains readability and consistency with marketing intent in terms of sentence structure and logical organization.

[0175] By combining the above algorithms, the sensitive semantic residues from the previous step are transformed into semantically equivalent but differently expressed alternative text, achieving the dual goals of improving compliance and maintaining marketing appeal.

[0176] For example, in an e-commerce promotional SMS scenario, the original text contains the phrase "limited-time flash sale, buy now," and its sensitivity detection score is 0.86, which is higher than the threshold of 0.8, triggering this processing step. The input sequence is 20 words long and is generated using a multi-head self-attention mechanism. The dimension is (20, 512), and the sensitive phrase position embedding is used. (After alignment with position indices [3,4,5]) Attention weight at sensitive locations is increased to 0.92. Gating strategy generation. The weight at the sensitive phrase position is 0.87, and the average weight at the other positions is 0.05. The calculated weight is... It is then fed into a 6-layer decoder and output. The output was "Limited-time offer, buy now". After dependency syntax rearrangement, the order of adverbs and verbs was adjusted to form the final output "Buy now for a limited time". The semantic consistency score was 0.94, the sensitivity avoidance score dropped to 0.32, and the marketing appeal retention score was 0.91.

[0177] S8.4: Perform semantic consistency verification on the generated alternative copy. By calculating the semantic similarity index between it and the original text, ensure that the reconstructed copy maintains the complete expression of the original marketing intent while avoiding sensitive semantics.

[0178] S8.5: The alternative text that has passed semantic consistency verification will be output as a compliant semantic equivalent expression to the rich media card generation module for use in the subsequent image and text fusion and SMS sending process, ensuring that the final output content meets the dual requirements of compliance and user appeal.

[0179] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0180] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," "third," and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "comprising" or "including" and similar terms mean that the elements or objects preceding "comprising" or "including" encompass the elements or objects listed following "comprising" or "including" and their equivalents, and do not exclude other elements or objects. The "multiple" involved in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0181] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / date> < / price> < / date> < / price>

Claims

1. An artificial intelligence-based SMS text sensitive word recognition analysis method, specifically comprising: S1: obtaining the marketing SMS text content to be sent and the associated rich media image material as multi-modal input data; S2: performing semantic embedding processing on the marketing SMS text content, using a pre-trained language model to extract high-dimensional text semantic feature vectors; S3: performing visual feature encoding on the rich media image material, using a convolutional neural network to extract image semantic representation vectors; S4: based on the text semantic feature vectors and the image semantic representation vectors, constructing a cross-modal semantic alignment loss function; calculating the cross-modal similarity of the text semantic feature vectors and the image semantic representation vectors, using the cosine similarity algorithm to measure the semantic matching degree between different modalities to obtain a cross-modal semantic similarity matrix; Based on the cross-modal semantic similarity matrix, an initial semantic alignment loss term is constructed, and a contrastive learning loss function is used to quantify the semantic deviation between text and image modalities to generate a cross-modal semantic alignment error value; Perform weighted fusion processing on the cross-modal semantic alignment error value, and combine the attention mechanism to dynamically allocate weights to multi-modal semantic features to generate a weighted semantic alignment loss value; Based on the weighted semantic alignment loss value, introduce a modal invariance constraint term, and use the Maximum Mean Discrepancy (MMD) algorithm to measure the difference between different modal distributions to further optimize the cross-modal semantic consistency; Jointly optimize the weighted semantic alignment loss value and the modal invariance constraint term to construct the final cross-modal semantic alignment loss function, and output a unified multi-modal semantic deviation quantification index; S5: input the cross-modal semantic alignment loss function into the multi-modal semantic consistency optimization model to generate a text optimization expression after semantic deviation correction; the multi-modal semantic consistency optimization model uses parameter gradient back propagation and dynamic learning rate adjustment, combined with a multi-head attention mechanism to focus on correcting text semantic units with high correlation and output semantic alignment enhanced text features, and then through a sequence-to-sequence generation model to optimize and generate text; S6: based on the text optimization expression after semantic deviation correction, combined with a pre-set sensitive word library to perform dynamic sensitive word recognition and avoidance strategy generation; S7: fuse the text content and image material after the avoidance strategy is generated to generate a rich media card, and output to the SMS sending module; S8: In the rich media card generation process, if there are still sensitive semantic residues after semantic deviation correction, trigger the text reconstruction mechanism to generate a compliant semantic equivalent expression.

2. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, characterized in that, Step S1 specifically includes: Structurally analyzing the marketing text content from the SMS sending queue to extract the body field and target user label information, and obtaining the original text input for semantic modeling; Based on the marketing text content, identify the associated rich media image resource identifiers, including image URL, local storage path, and cloud resource ID; Obtain the original pixel data of the rich media image material through HTTP protocol or local file system call method to form an image input tensor that can be used for visual feature extraction; Performing image preprocessing operations on the original pixel data, including size normalization, color space conversion and noise filtering, to obtain a standardized image input matrix in a unified format; Aligning and packaging the standardized image input matrix with the corresponding marketing text content to construct a structured multi-modal input sample.

3. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, wherein characterized in that Step S2 specifically comprises: Preprocessing the obtained marketing short message original text content, including word segmentation, stop word removal and standardization, to obtain a standardized text semantic unit sequence; Based on the standardized text semantic unit sequence, performing forward propagation calculation of the embedding layer to obtain an initial word-level semantic embedding vector matrix; Performing context-aware Transformer encoding processing on the initial word-level semantic embedding vector matrix, and generating context-enhanced semantic embedding representation based on the self-attention mechanism to fuse context semantic dependency; Based on the context-enhanced semantic embedding representation, extracting a text global semantic feature vector to obtain a high-dimensional semantic feature expression; Normalizing the high-dimensional semantic feature expression to generate a standardized semantic feature vector.

4. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, characterized in that, Step S3 specifically comprises: Performing image preprocessing operations on the rich media image, including size normalization, color space conversion and noise suppression, to obtain a standardized image input in a unified format; Based on the standardized image input, performing multi-level convolution operations on the image using a pre-trained convolutional neural network model to extract local texture features and global structure features of the image, and generating an original convolutional feature map; Performing pooling operations on the original convolutional feature map to reduce the feature dimension and retain key visual semantic information through the max-pooling algorithm, and obtaining a reduced image feature representation; Based on the reduced image feature representation, performing high-dimensional semantic mapping using a fully connected layer to generate an image semantic embedding vector with a fixed dimension as a high-level semantic abstraction of the image content; Performing normalization on the image semantic embedding vector to standardize the vector and obtain an image semantic representation vector that can be used for cross-modal semantic alignment.

5. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, characterized in that, The matching of resource identifiers in the rich media image material uses relationship mapping dictionaries, fuzzy string edit distance and semantic similarity in the BERT embedding space to standardize image paths and unify them to HTTP, file or cloud storage ID formats.

6. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, characterized in that, The text semantic feature vector uses the BERT model and Transformer context encoding to obtain a standardized text semantic unit sequence through word segmentation, stop word filtering, Unicode and numerical date standardization and hash indexing, extract a high-dimensional text semantic feature vector and normalize it.

7. The short message text sensitive word recognition analysis method based on artificial intelligence according to claim 1, characterized in that, The sensitive word recognition uses a dynamic threshold determination model to first generate a sensitive word candidate set based on context structure analysis and semantic rules, then compares it with a predefined sensitive word library based on BERT semantic embedding and cosine similarity score, outputs the matching score and constructs the final sensitive word set accordingly.

Citation Information

Patent Citations

  • Image-text content auditing method based on multi-modal large model

    CN119941157A

  • News event search method and system based on multi-level image-text semantic alignment model

    WO2023093574A1