Method for identifying advertorial advertisement in social media
By constructing a soft article advertisement recognition model, using the bidirectional encoder characterization method, multi-head self-attention mechanism and multi-scale convolution operation, the characteristics of tweet data are extracted and integrated, and the problem of difficult soft article advertisements in social media is solved, and the accurate identification and management of soft article advertisements is achieved.
Patent Information
- Application Number
- CN202510226023.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to accurately identify soft article advertisements in social media. The soft article advertisements are highly concealed and have a non-fixed length, making it difficult to effectively detect and supervise through the prior art.
A soft article advertisement recognition model is constructed, including the input layer, local feature extraction layer, global feature extraction layer, feature fusion layer and output layer. Through the bidirectional encoder characterization method, multi-head self-attention mechanism, long and short time recording network and multi-scale convolution operation, the features of tweet data are extracted and fused, and finally the soft article advertisement recognition is recognized through a linear classifier.
It realizes accurate identification of soft article advertisements in social media, improves the detection ability of soft article advertisements with high concealment and irregular length, and enhances the platform's ability to manage and supervise advertising content.
Smart Images

Figure CN120144770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the application field of social media content analysis and advertising monitoring, and particularly relates to a method for identifying soft-text advertisements in social media. Background Art
[0002] Currently, advertising is one of the main sources of income for social media platforms, providing financial support and development funds for the platforms and promoting the continuous development and innovation of social media platforms. The number of advertisements on social media platforms is increasing continuously. When users browse content, they will be interfered by a large number of advertisements, which may lead to a decrease in user satisfaction with the platform and even cause users to leave the platform. Therefore, social media platforms usually review and supervise the content of the advertisements placed and limit the number of advertisements. However, a large number of illegal advertisements are placed in the user interaction area through short texts, which not only affect the user experience but may also contain malicious links that cause property losses to users. Advertisements on social media usually appear in two ways, namely hard advertisements and soft advertisements. Hard advertisements directly and clearly convey product information and purchase appeals to the audience, highlighting the clear brand name and distinct product features, helping users quickly identify the advertisement source and brand information, and covering a large number of audiences in a short time through high-frequency exposure, increasing users' familiarity with the advertisement and making it easier to be recognized and recalled during potential purchase decisions. Soft advertisements refer to integrating advertisement information indirectly and dispersedly into the text, packaging it as personal experiences or experience sharing, usually only retaining the necessary information to identify the product, highlighting the subjective experience of product use, and obtaining users' trust through other content such as identities and cases, thereby indirectly changing users' perception of advertisement information and then influencing users' decisions.
[0003] In the prior art, social media is mainly a user communication platform for information acquisition and content creation through text and pictures. The direct promotion method of hard advertisements easily interrupts the user's browsing experience and causes users' disgust and resistance. The platform can detect and supervise these hard advertisement texts with obvious features better.
[0004] However, malicious soft-text advertisements can interfere with users' judgments through false experiences and exaggerated effects of recommenders. Different from the obvious features of hard advertisements, since the expression methods of soft-text advertisements on social media are similar to users' daily content, they have high concealment and variable lengths, usually mainly short texts. The prior art is often used to identify hard advertisement texts with obvious features and cannot judge soft-text advertisements.
[0005] Therefore, there is an urgent need for an identification method to accurately identify soft-text advertisements in social media. Summary of the Invention
[0006] Based on this, it is necessary to provide a method for identifying soft - text advertisements in social media for the above - mentioned technical problems.
[0007] This specification adopts the following technical solutions:
[0008] This specification provides a method for identifying soft - text advertisements in social media, including:
[0009] Construct a soft - text advertisement recognition model; the soft - text advertisement recognition model includes: an input layer, a local feature extraction layer, a global feature extraction layer, a feature fusion layer, and an output layer;
[0010] Obtain the pre - processed tweet data;
[0011] Input the pre - processed tweet data into the soft - text advertisement recognition model. In the input layer, use the bidirectional encoder representation model to perform word segmentation, encoding, and context modeling on the pre - processed tweet data to obtain the per - word hidden features and pooling features of the tweet data;
[0012] In the local feature extraction layer, generate fixed - position encodings through sine and cosine functions, add them to the per - word hidden features, and through a multi - layer encoder, extract the semantic features in the per - word hidden features with fixed - position encodings added and retain the dimension of the per - word hidden features; expand the pooling features by one dimension, add them to the semantic features and input them into a long short - term memory network for temporal encoding, output temporal features, and through a max - pooling operation, extract the maximum value of the temporal features to obtain deep - level local features;
[0013] In the global feature extraction layer, perform multi - scale convolution operations on the per - word hidden features through multiple convolutional kernels of different sizes, extract local features of different scales and splice them to obtain multi - scale global features;
[0014] In the feature fusion layer, splice the deep - level local features and the multi - scale global features in the feature dimension to obtain complete features;
[0015] In the output layer, input the complete features into a linear classifier to map them to the number of dimensions of the target classification, and generate the classification result corresponding to the tweet data, that is, the recognition result of the soft - text advertisement in the tweet data.
[0016] Preferably, the obtaining of the pre - processed tweet data specifically includes:
[0017] Batch - collect tweet data through web crawlers or by calling the API interfaces provided by social media platforms;
[0018] Clean the data in the tweet data that includes HTML tags, URL links, @ user mentions, hashtag tags, extra spaces, punctuation marks, and non - language characters;
[0019] Unify the case of letters in the tweet data;
[0020] Standardize the dates and times in the tweet data;
[0021] Remove stop words from the tweet data; the stop words include: function words, adverbs, and conjunctions;
[0022] Retain emoticons related to sentiment, currency symbols, and unit symbols in the tweet data;
[0023] Parse the hyperlinks in the tweet data and filter out non-text information.
[0024] Preferably, obtaining the per-word hidden features and pooled features of the tweet data specifically includes:
[0025] Input the preprocessed tweet data into a pre-trained bidirectional encoder representation model;
[0026] Use the tokenizer of the bidirectional encoder representation model to tokenize the text in the preprocessed tweet data into sub-word units, map them to the input positions in the vocabulary, and add markers; generate an attention mask to indicate the positions of valid sub-word units, and obtain the tokenization result;
[0027] Input the input positions and the attention mask into the multi-layer Transformer of the bidirectional encoder representation model for encoding and context modeling to obtain per-word hidden features and pooled features.
[0028] Preferably,
[0029] The per-word hidden feature is a three-dimensional vector, which includes the hidden state representation of each sub-word unit;
[0030] The pooled feature is a two-dimensional vector, which is a global representation generated based on the hidden states of the markers.
[0031] Preferably, the multi-layer encoder includes: a multi-head attention mechanism and a feed-forward neural network; extracting the semantic features in the per-word hidden features with fixed positional encoding through the multi-layer encoder specifically includes:
[0032] Through the multi-head attention mechanism of the multi-layer encoder, divide the per-word hidden features with fixed positional encoding into multiple attention heads; and calculate the dependency relationships between each attention head and other attention heads;
[0033] Concatenate the results of the multiple attention heads and use a linear transformation to map back to the original dimension;
[0034] Perform training on the concatenated result through residual connection and normalization operations to obtain the trained features;
[0035] Through a feedforward neural network, two linear transformation functions and the ReLU activation function are used to further extract non-linear features, and semantic features in the per-word hidden features with fixed-position encoding added are obtained.
[0036] Preferably, through the max-pooling operation, the maximum value of the temporal features is extracted to obtain deep local features, which specifically includes:
[0037] The temporal features output by the long short-term memory network are converted into vectors of a fixed length, and the maximum value of each feature is extracted along the length dimension of the temporal features to obtain deep local features.
[0038] Preferably, the global feature extraction layer includes: a plurality of convolutional kernels of different sizes, the ReLU activation function, and a max-pooling layer; the obtaining of the multi-scale global features specifically includes:
[0039] Through the plurality of convolutional kernels of different sizes, non-linear features are introduced using the ReLU activation function, and the feature length is reduced through the max-pooling layer to obtain local features of different scales;
[0040] The shapes of the local features of different scales output by each convolutional kernel are flattened, and the flattened local features are concatenated to obtain multi-scale global features.
[0041] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0042] In a method for soft article advertisement recognition in social media provided in this specification, a soft article advertisement recognition model is constructed. After adding position encoding to the input tweet data, feature extraction is performed through two layers of multi-head self-attention mechanisms; after adding the features generated by each layer of the Transformer encoder to the pooled features generated by the Bidirectional Encoder Representations from Transformers (BERT) model, further encoding processing is performed through a long short-term memory network (LSTM); after performing the above operations multiple times, the max-pooling method is used to obtain local features; at the same time, the features output by the BERT model are subjected to multi-scale convolution processing using different convolutional kernels, and the multi-scale features are concatenated to form global multi-scale features. Finally, the obtained local features and global features are concatenated and input into a linear classifier (Linear) for final classification determination to achieve accurate recognition of soft article advertisements.
[0043] In summary, through the constructed soft article advertisement recognition model, the present invention realizes the stable training and accelerated convergence of semantic features in tweet data through the feed-forward network and multi-head attention mechanism in the local feature extraction layer, combined with residual connection and layer normalization operations. Further, the LSTM network is used to deeply perceive the semantic features of tweet data, retain fine context dependencies in the time series dimension, and adopt max pooling operation to extract the maximum value of each feature along the sequence length dimension, obtaining the most important local information in tweet data; through the global feature extraction layer, multi-scale convolutional kernels are used to perform convolutional operations on semantic features, efficiently extract local features of different scales and splice them, and the obtained multi-scale global features contain more comprehensive semantic information; through the feature fusion layer, both fine-grained local features and generalized global information are utilized to provide a richer semantic representation for the recognition of soft article advertisements in tweet data, thereby realizing the accurate recognition of soft article advertisements in tweet data. Description of the Drawings
[0044] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0045] Figure 1 It is a schematic flowchart of a method for recognizing soft article advertisements in social media provided in this specification;
[0046] Figure 2 It is a schematic structural diagram of a method model for recognizing soft article advertisements in social media provided in this specification. Detailed Embodiments
[0047] To make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] The following will describe in detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0049] Figure 1 It is a schematic flowchart of a XX method in this specification, which specifically includes the following steps:
[0050] S101: Construct a soft article advertisement recognition model; the soft article advertisement recognition model includes: an input layer, a local feature extraction layer, a global feature extraction layer, a feature fusion layer, and an output layer.
[0051] S102: Obtain the preprocessed tweet data, including:
[0052] Batch collect tweet data through web crawlers or by calling the API interfaces provided by social media platforms;
[0053] The tweet data includes various types of topics and sources;
[0054] Perform data preprocessing on the tweet data, specifically including:
[0055] Clean the tweet data to remove HTML tags, URL links, @user mentions, hashtag tags, extra spaces, punctuation marks, and non-verbal characters;
[0056] Normalize the text in the tweet data, including: unifying the letter case and standardizing the dates and times;
[0057] Remove the high-frequency words in the tweet data that contribute little to the semantics;
[0058] Selectively retain the emotional emojis, other emotion-related symbols, currency symbols, and unit symbols with high semantic value in the tweet data.
[0059] S103: Input the preprocessed tweet data into the soft article advertisement recognition model. In the input layer, use the bidirectional encoder representation model to tokenize, encode, and perform context modeling on the preprocessed tweet data to obtain the per-word hidden features and pooling features of the tweet data; in the local feature extraction layer, generate fixed-position encodings through sine and cosine functions, add them to the per-word hidden features, and through multiple layers of encoders, extract the semantic features in the per-word hidden features with the fixed-position encodings added and retain the dimension of the per-word hidden features; expand the pooling features by one dimension, add them to the semantic features and input them into a long short-term memory network for temporal encoding, output the temporal features, and through max pooling operation, extract the maximum value of the temporal features to obtain deep local features; in the global feature extraction layer, perform multi-scale convolution operations on the per-word hidden features through multiple convolutional kernels of different sizes, extract local features of different scales and splice them to obtain multi-scale global features; in the feature fusion layer, splice the deep local features and the multi-scale global features in the feature dimension to obtain complete features; in the output layer, input the complete features into a linear classifier to map them to the number of dimensions of the target classification and generate the classification result corresponding to the tweet data;
[0060] Among them, the gradient backpropagation of the per-word hidden features and pooling features of the tweet data is prohibited;
[0061] Obtain the per-word hidden features and pooling features of the tweet data, including:
[0062] Input the preprocessed tweet data into the pre-trained BERT model;
[0063] The BERT model uses its tokenizer to tokenize the text in the preprocessed tweet data into sub-word units, map them to input_ids in the vocabulary, and add markers; at the same time, generate attention_mask to indicate the positions of valid sub-word units and obtain the tokenization result;
[0064] Input input_ids and attention_mask into the multi-layer Transformer of the BERT model for encoding and perform context modeling to obtain per-word hidden features and pooled features;
[0065] The per-word hidden feature is a three-dimensional vector that includes the hidden state representation of each sub-word unit;
[0066] The pooled feature is a two-dimensional vector, which is a global representation generated based on the hidden state of the marker;
[0067] The multi-layer encoder includes: a multi-head attention mechanism and a feed-forward neural network;
[0068] The multi-head attention mechanism of the multi-layer encoder divides the per-word hidden feature with added fixed-position encoding into multiple attention heads; and calculates the dependency relationships between each attention head and other attention heads;
[0069] Concatenate the results of multiple attention heads and map them back to the original dimension using a linear transformation;
[0070] Perform training on the concatenated result through residual connection and normalization operations to obtain the trained features;
[0071] Through the feed-forward neural network, use its two linear transformation functions and the ReLU activation function to further extract non-linear features and obtain the semantic features in the per-word hidden feature with added fixed-position encoding;
[0072] Convert the temporal features output by the LSTM into a fixed-length vector, and extract the maximum value of each feature along the length dimension of the temporal features to obtain deep local features.
[0073] The multiple convolution kernels of different sizes introduce non-linear features using the ReLU activation function and reduce the feature length through the max-pooling layer to obtain multi-scale local features of different granularities;
[0074] Flatten the shapes of the multi-scale local features output by each convolution kernel, and concatenate each flattened multi-scale local feature to obtain multi-scale global features.
[0075] S104: Input the tweet data to be recognized into the soft - article advertisement recognition model to recognize the soft - article advertisements that disguise as personal experiences to induce consumption in the tweet data to be recognized.
[0076] Based on the above steps, in this embodiment, a large number of tweet data are collected in batches through web crawlers or by calling the API interfaces provided by social media platforms, ensuring that the data covers diverse topics and sources to obtain representative tweet samples.
[0077] The collected data is preliminarily screened and parsed to extract the text content and the included hyperlink information in the tweets. The system automatically labels the tweets containing product web page links as soft - article advertisement texts, while the tweets without product web page links are labeled as non - soft - article advertisement texts. All the automatic labeling results are manually reviewed and corrected by professionals to ensure the accuracy and consistency of the labeling.
[0078] For the data cleaning of short social media texts, a series of targeted methods usually need to be adopted to improve the data quality and adapt to downstream tasks. Clean the text of the collected tweet data, remove HTML tags, special characters, duplicate content, and redundant whitespace to ensure the purity and consistency of the data; then, perform word - segmentation operations to break the tweet text into independent words or phrases; next, perform normalization processing on the text, including unifying case and replacing synonyms; subsequently, parse the hyperlinks in the tweets and filter out non - text - related information, retaining the valid links pointing to specific web pages; finally, construct a word - frequency matrix or convert the text into a vectorized form (such as TF - IDF or word embeddings) to adapt to the input requirements of the model. At the same time, split the dataset into a training set, a validation set, and a test set to ensure the balance and randomness of the data distribution and provide high - quality input for subsequent model training.
[0079] First, remove the noise in the data, including cleaning HTML tags, URL links, @ user mentions (such as @username), topic tags (such as #hashtag), extra spaces, punctuation marks, and non-language characters (such as emoticons or special symbols, depending on the task requirements). Although these elements are common in social media texts, they may be interference items for many natural language processing tasks. Secondly, perform text normalization, such as unifying the letter case (usually converting to lowercase), and standardizing the date, time, and other content as needed. Subsequently, you can choose to remove stop words, such as "的" and "了", which are high-frequency words but have little contribution to semantics, to reduce data noise. Finally, it is also particularly important to process special symbols, such as selectively retaining emoticons and other emotion-related symbols, because these symbols may have strong semantic value in tasks such as sentiment analysis, while currency symbols, unit symbols, etc. can be cleaned or replaced according to specific needs. Through the above steps, the cleaned social media short text can express semantics more clearly and provide more valuable data input for subsequent analysis or model training.
[0080] Build a recognition model, including:
[0081] The input text (content) is preprocessed and passed to the BERT model, and the output of the model is extracted for downstream tasks. First, the function uses BERT's tokenizer to encode the input text, including tokenizing the text into sub-word units, mapping them to input_ids in the vocabulary, and adding special tags such as [CLS] and [SEP]. At the same time, an attention_mask is generated to indicate valid Token positions, with the padding part being 0. The generated input data is then converted into a PyTorch tensor and transferred to the specified GPU device for processing in the model.
[0082] After ensuring that the length of the input sequence does not exceed the maximum supported length of BERT, which is 512, the tokenized results context (input_ids) and attention_mask are passed to the BERT model. The BERT model performs context modeling on the input sequence through multiple layers of Transformer encoding and generates two main outputs: last_hidden_state and pooler_output. Among them, last_hidden_state contains the hidden state representation of each Token, which is a three-dimensional tensor with a shape of [batch_size, seq_length, hidden_size], capturing the context semantic information of each Token; pooler_output is a two-dimensional tensor with a shape of [batch_size, hidden_size], which is a global representation generated based on the hidden state of the special token [CLS].
[0083] Finally, these tensors are detached from the BERT output, gradient backpropagation is prohibited, and they are transferred from the GPU to the CPU and converted to NumPy arrays for convenient subsequent processing. Finally, the two output vectors of last_hidden_state and pooler_output are returned, providing a basic representation for subsequent tasks.
[0084] To enable the model to perceive the sequential relationship of the vocabulary in the input sequence, a fixed positional encoding (Positional Encoding) generated by sine and cosine functions is added to last_hidden_state. The generation method of the positional encoding ensures the uniqueness of the encoding for each position, while having good distinguishability for the relative distances between different positions. This addition operation does not change the shape of the input (still [batch_size, seq_length, hidden_size]), but explicitly incorporates the positional information, enabling the Transformer to more effectively capture the sequential dependencies in the sequence through the subsequent self-attention mechanism. The positional encoding is simple and efficient, without training parameters, and can well make up for the lack of awareness of the sequence position in the Transformer architecture.
[0085] The model then further extracts semantic features through multiple Encoder layers. Each Encoder consists of two parts: the Multi-Head Attention mechanism and the Position-wise FeedForward neural network. The Multi-Head Attention mechanism divides the input features into multiple attention heads, calculates the dependencies between each Token and other Tokens, then concatenates the results of multiple heads and maps them back to the original dimension through a linear transformation. Subsequently, through residual connections and layer normalization operations, the training is stabilized and the convergence is accelerated. The FeedForward network further extracts non-linear features through two linear transformations and ReLU activation, and also uses residual connections and layer normalization to enhance stability. These encoder layers enable the model to efficiently capture semantic information and retain the feature dimensions of the input [batch_size, seq_length, hidden_size]. During the process of extracting features through the encoder layers, the model also incorporates the global semantic information of pooler_output. Specifically, pooler_output is expanded by one dimension, added to the output of the encoder, and then passed to the LSTM layer. LSTM is a time series model that can effectively capture the temporal dependencies in the sequence, complementing the global modeling ability of Transformer. This fusion method enables the model to perceive deep semantic features at each step while retaining fine-grained context dependencies in the sequence dimension. The output shape of LSTM is [batch_size, seq_length, hidden_size], providing richer information for subsequent local feature extraction.
[0086] It is achieved through the MaxPool operation. The pooling operation takes the maximum value of each feature along the sequence length dimension (seq_length), thereby extracting the most important feature information. This process converts the sequence features output by LSTM into a fixed-length vector with the shape [batch_size, hidden_size]. The advantage of pooling is that it can significantly reduce the dimension of the data while retaining the most prominent features in the input sequence. In addition, the pooling operation is robust to the sequence length and can handle input sequences of different lengths, providing a stable local feature representation for subsequent global feature extraction and classification tasks.
[0087] To extract the global features of a sentence, the model convolves the last_hidden_state using multiple convolutional kernels of different sizes (filter_sizes=(2,3,4)). Each convolutional kernel captures local features at different granularities, such as the correlations between consecutive 2 Tokens, 3 Tokens, or 4 Tokens. The output of the convolution is introduced with non-linearity through the ReLU activation function and further reduces the sequence length through the max pooling layer. After the output shapes of each convolutional kernel are flattened, all the outputs are concatenated into a global feature representation with the shape [batch_size, num_filters*len(filter_sizes)]. The advantage of convolution is that it can efficiently extract local information at different granularities and has a strong ability to model spatial structures. By combining multiple convolutional kernel sizes, the model can capture more comprehensive global features.
[0088] After obtaining the local and global features, the model concatenates the two in the feature dimension to form a complete feature representation. Specifically, the local feature (out_local) and the global feature (out_global) are [batch_size, hidden_size] and [batch_size, num_filters*len(filter_sizes)] respectively, and the shape of the concatenated feature is [batch_size, hidden_size+num_filters*len(filter_sizes)]. Through this fusion, the model can utilize both fine-grained local information and generalized global information to provide a richer semantic representation for downstream classification tasks.
[0089] Finally, the concatenated features pass through a fully connected layer that maps the features to the number of dimensions num_classes for the target classification. The output shape of this layer is [batch_size, num_classes], and each value represents the prediction score for the corresponding class. The role of the fully connected layer is to compress the extracted high-dimensional features into the classification space and generate the corresponding classification results for each input sentence. By further reducing the risk of overfitting through the Dropout layer, the model can more stably adapt to the training and test data.
[0090] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity in description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
Claims
1. A method for identifying soft-text advertisements in social media, characterized in that: include: Build a soft copy advertising recognition model; The soft-text advertisement recognition model comprises: an input layer, a local feature extraction layer, a global feature extraction layer, a feature fusion layer and an output layer; Get the preprocessed tweet data; The preprocessed tweet data is input into the soft copy advertisement recognition model. At the input layer, the preprocessed tweet data is segmented, encoded, and context modeled using a bidirectional encoder representation model to obtain word-by-word hidden features and pooling features of the tweet data. In the local feature extraction layer, a fixed position code is generated by a sine function and a cosine function, and added to the word-by-word hidden feature. The semantic feature in the word-by-word hidden feature with the fixed position code is extracted through a multi-layer encoder, and the dimension of the word-by-word hidden feature is retained; the pooled feature is expanded by one dimension, added to the semantic feature and input into a long short-term memory network for temporal coding, and the temporal feature is output. The maximum value of the temporal feature is extracted through a maximum pooling operation to obtain a deep local feature; In the global feature extraction layer, a multi-scale convolution operation is performed on the word-by-word hidden features through a plurality of convolution kernels of different sizes, local features of different scales are extracted and concatenated to obtain multi-scale global features; In the feature fusion layer, the deep local features and the multi-scale global features are spliced in the feature dimension to obtain complete features; In the output layer, the complete features are input into the linear classifier to map them to the number of dimensions of the target classification, generating the classification results corresponding to the tweet data, that is, the recognition results of soft-text advertisements in the tweet data.
2. A method for identifying soft-text advertisements in social media as claimed in claim 1, characterized in that: The obtaining of the preprocessed tweet data specifically includes: Collect tweet data in batches by crawling the web or calling the API interface provided by the social media platform; Clean the tweet data including HTML tags, URL links, @ user mentions, hashtags, extra spaces, punctuation, and non-language characters; Unify the capitalization of letters in tweet data; Standardize the date and time in the tweet data; Remove stop words from tweet data; the stop words include: function words, adverbs and conjunctions; Parse the hyperlinks in tweet data and filter out non-text information. Keep emotional emoticons, currency symbols, and unit symbols in tweet data.
3. A method for identifying soft-text advertisements in social media as claimed in claim 1, characterized in that: The step of obtaining the word-by-word hidden features and pooled features of tweet data specifically includes: Input the preprocessed tweet data into the pre-trained bidirectional encoder representation model; The word segmenter of the bidirectional encoder representation model is used to segment the text in the preprocessed tweet data into subword units, map them to the input positions in the vocabulary and add tags; an attention mask is generated to indicate the valid subword unit positions to obtain the word segmentation results; The input position and attention mask are input into the multi-layer Transformer of the bidirectional encoder representation model for encoding and context modeling to obtain word-by-word hidden features and pooling features.
4. A method for identifying soft-text advertisements in social media as claimed in claim 4, characterized in that: The word-by-word hidden feature is a three-dimensional vector, which includes a hidden state representation of each sub-word unit; The pooled feature is a two-dimensional vector, which is a global representation generated based on the hidden state of the marker.
5. A method for identifying soft-text advertisements in social media according to claim 1, characterized in that: The multi-layer encoder includes: a multi-head attention mechanism and a feedforward neural network; the semantic features in the word-by-word hidden features added with fixed position encoding are extracted through the multi-layer encoder, specifically including: Through the multi-head attention mechanism of the multi-layer encoder, the word-by-word hidden features with fixed position encoding are divided into multiple attention heads; and the dependency relationship between each attention head and other attention heads is calculated; Concatenate the results of multiple attention heads and map them back to the original dimension using linear transformation; The concatenated results are trained through residual connection and normalization operations to obtain trained features; Through the feedforward neural network, the two linear transformation functions and the activation function ReLU are used to further extract nonlinear features, and the semantic features in the word-by-word hidden features with fixed position encoding are obtained.
6. A method for identifying soft-text advertisements in social media as claimed in claim 1, characterized in that: The maximum value of the temporal feature is extracted through the maximum pooling operation to obtain deep local features, which specifically includes: The time series features output by the long short-term memory network are converted into a vector of fixed length, and the maximum value of each feature is extracted along the length dimension of the time series features to obtain deep local features.
7. A method for identifying soft-text advertisements in social media as claimed in claim 1, characterized in that: The global feature extraction layer includes: a plurality of convolution kernels of different sizes, an activation function ReLU and a maximum pooling layer; the multi-scale global feature acquisition specifically includes: Nonlinear features are introduced by using the activation function ReLU through the multiple convolution kernels of different sizes, and the feature length is reduced through the maximum pooling layer to obtain local features of different scales; The shapes of local features of different scales output by each convolution kernel are flattened, and the flattened local features are concatenated to obtain multi-scale global features.
Citation Information
Cited By
Abstract content generation method and device based on improved Bert
CN121188300A