A Method and System for Joint Detection of User-Generated Content Target Stance Based on Encoding-Decoding Structure

By directly predicting the goals and positions of user-generated content based on the method based on the codec structure, the problems of manual label dependence and error cascade in the prior art are solved, and more efficient and accurate position detection is achieved.

CN118916739BActive Publication Date: 2025-06-03HARBIN INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410944911.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-06-03
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

The prior art relies on manual labeling of target information in the position detection of user-generated content, resulting in high labor costs, and the target recognition error in the two-stage method will affect the performance of position detection and produce error cascade.

Method used

The user-generated content target position joint detection method based on the codec structure is adopted, and the context and emotional information of the text are obtained through the sequence sub-encoder and the emotional sub-encoder, and the self-attention mechanism and cross-attention mechanism are used for feature interaction and fusion, directly predicting the target and position, avoiding manual annotation and error cascade.

Benefits of technology

The need to manually mark target information reduces the manual dependence of position detection, avoids error cascade, and improves the accuracy and feasibility of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118916739B_ABST
    Figure CN118916739B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for jointly detecting the target stance of user-generated content based on an encoding and decoding structure, and relates to the stance detection of social media. The present invention solves the problem of manual dependence in the stance detection task and eliminates the error cascade phenomenon. Technical key points: The preprocessed social media text data is input into an encoder, which is composed of a sequence encoder and a fine-tuned sentiment encoder. The self-attention mechanism is used between query vectors to dynamically calculate the degree of association between each query vector and other query vectors, so as to better capture the dependence relationship between different query vectors. Then, the sequence features output by the encoder are input into a decoder, and a cross-attention mechanism is performed with the query vectors. All query vectors fused with sequence features are input into a target-stance aggregation layer. The aggregated query vectors and the sentiment features output by the encoder are input into a target stance pair decoding layer. First, weights are assigned to the query vectors and sentiment features through an attention mechanism, and then the two features are concatenated to obtain a final feature vector. The final feature representation is input into a decoder composed of two fully connected neural networks to output the prediction results of the target and the stance. The present invention is applied to social network analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of stance detection in social media and social network data processing, and particularly relates to a method and system for jointly detecting the target stance of user-generated content. Background Art

[0002] Nowadays, the rise of social media platforms has changed people's lifestyles, enabling people to obtain information, share opinions, and communicate feelings more quickly and conveniently, promoting the rapid spread of information and the wide sharing of knowledge. Stance detection for social media can help relevant personnel better understand and analyze public opinion, views, and attitudes, so as to make more informed decisions in the fields of politics, business, society, etc. By analyzing the stance of the public towards a certain topic, it can help the government understand public opinion and formulate more reasonable policies; help enterprises understand consumers' views on products or services, and conduct market positioning and marketing strategy formulation; help media and news agencies understand the public's reaction to news events and improve the accuracy and quality of reports. In summary, stance detection technology has important research significance and application value in real-world scenarios.

[0003] For the task of stance detection for social media, scholars have carried out a large number of related studies. Yuan et al. [1] introduced the human stance reasoning process as task knowledge into stance detection calculations, enabling the model to effectively filter redundant features of text data itself and thus rely more on target features. For tweets related to climate change, Upadhyaya et al. [2] introduced fine-grained emotion recognition and aggression recognition as auxiliary tasks, and used emojis in social media as multimodal information to provide richer semantic information. Ko et al. [3] adopted a hierarchical attention network to learn the relationship between three different levels of semantic information, and incorporated external knowledge of the real world into the process of political stance prediction by constructing a political knowledge graph and knowledge encoding, effectively improving the performance of the model. On the basis of traditional stance detection research, Li et al. [4] proposed a two-stage stance detection method for unlabeled target information. First, in the first stage, the target faced by social media text is identified, and then in the second stage, stance detection is carried out on the basis of the first stage. This research greatly reduces the labor cost in stance detection and provides a new idea for the field of stance detection.

[0004] The prior art with the document number CN118070774A provides a method and system for detecting the stance of user-generated content based on target information recognition. It aims to solve the problems that existing methods for detecting or recognizing the stance of user-generated content require a large amount of labor costs to label target information, and the only similar methods often need to use a large amount of data to train or fine-tune the model in the target recognition stage, resulting in the quality of sample data directly affecting the performance and accuracy of target recognition. The technical key points are as follows: First, extract representative keywords from the given social media text; then calculate the similarity between the keywords and specific targets in the target set through cosine similarity, and determine the target object targeted by the text according to the similarity; finally, based on the identified target object, use the multi-task BERTweet model to detect the stance relationship between the text and the target object. The proposed method for detecting the stance of user-generated content based on target information recognition can effectively reduce labor costs, thereby improving the feasibility and practicality of the stance detection method in actual applications.

[0005] The prior art with the document number CN116992873A discloses a multi-target stance detection method and system based on contrastive learning and consistency detection. This prior art enables the model to learn more feature information of the targets and strengthen the connection of semantic information between the targets, enabling the targets to assist each other in detecting their own stances. In addition, in this prior art, BERT is fine-tuned and BiLSTM is embedded as the encoder to make more full use of the semantic information between hidden contexts. This prior art also takes joint training as a multi-task learning method, allowing the model to share specific domain information based on the dataset. It solves the problems of noisy datasets, isolated targets, and insufficient specific domain information in multi-target stance detection. It not only improves the accuracy of performance in this task, but also has reference significance for other multi-target text classification tasks with similar problems.

[0006] The prior art with the document number CN114330360A discloses a stance detection method for a specific target. This prior art uses a deep network to extract the semantic features of sentences and fully considers target features during stance detection to achieve the interaction between target features and sentence features. The model uses a densely connected BiLSTM network and a nested LSTM network to extract the semantic features of sentences, which can capture the deep semantic information of sentences while solving the problems of gradient disappearance and long-term dependence; it uses an attention mechanism to obtain the importance of a specific target for each part of the sentence, thereby obtaining a sentence vector representation incorporating specific target information to help the model fully consider the given specific target during stance detection.

[0007] Although the above research and the existing technologies have achieved varying degrees of breakthroughs in the stance detection task, there are still the following deficiencies: (1) Traditional stance detection methods often rely on both social media texts and manually annotated target information. Before performing detection calculations, manual annotation of the text to be detected is required, resulting in excessively high labor costs; (2) The current target-stance detection methods for user-generated content usually adopt a two-stage method. Although this method reduces the labor cost of the stance detection task, the errors generated in the first-stage target recognition task will have a negative impact on the performance of the second-stage stance detection task, thereby causing an error cascade phenomenon. Summary of the Invention

[0008] The technical problem to be solved by the present invention is:

[0009] In view of the above problems, the present invention proposes a method and system for jointly detecting the target stance of user-generated content based on an encoder-decoder structure. This method not only solves the problem of manual dependence in the stance detection task but also eliminates the error cascade phenomenon, effectively improving the feasibility and accuracy of the stance detection method in practical applications.

[0010] The technical solution adopted by the present invention to solve the above technical problem is:

[0011] 1. A method for jointly detecting the target stance of user-generated content based on an encoder-decoder structure, characterized in that the implementation process of the method is as follows:

[0012] Step 1. Preprocess the user-generated content data:

[0013] The data preprocessing is divided into the following two steps: (1) Data cleaning: used to eliminate network elements such as emojis and web links contained in the user-generated content; (2) Conversion of online jargon: Convert common online jargon in social media texts into written language using a predefined abbreviation dictionary.

[0014] Step 2. Fine-tune the BERT model using the preprocessed sentiment analysis dataset. The dataset has two sentiment labels: "positive" (positive sentiment) and "negative" (negative sentiment); the dataset includes a training set, a validation set, and a test set, and the data format includes a data number, an input text sequence, and the true sentiment label corresponding to the text.

[0015] Step 3: Input the preprocessed social media text data into the encoder, which consists of a sequence sub-encoder and a fine-tuned sentiment sub-encoder, both based on the BERT model. The sequence sub-encoder is used to obtain the sequence features formed by the text context information, while the sentiment sub-encoder is used to obtain the sentiment information implicit in the text sequence, providing richer semantic features for subsequent calculations.

[0016] Step 4: Randomly initialize m query embedding vectors. First, in the query embedding vector internal attention layer, use the self-attention mechanism for the query embedding vectors, enabling effective interaction and information transfer between the query embedding vectors to capture the dependencies between different query embedding vectors. The query vector, key vector, and value vector in the self-attention mechanism all come from the query embedding vectors themselves.

[0017] Then, in the tweet feature attention layer, input the sequence features output by the encoder into the decoder and perform cross-attention mechanism with the interacted query embedding vectors to promote the interaction and fusion of the two types of features.

[0018] In the cross-attention mechanism, use the query embedding vectors as the query vectors and the sequence features as the key vectors and value vectors.

[0019] Step 5: Input all the query vectors fused with sequence features into the target-stance aggregation layer. The main purpose of this layer is to make the model focus on the features that are more important or relevant to the current task, thereby improving the prediction performance of the model.

[0020] Assign corresponding weights to different query vectors and perform weighted summation on all query vectors to obtain a vector representation that combines all query features.

[0021] Step 6: Input the aggregated query vectors and the sentiment features output by the encoder into the target stance pair decoding layer. First, assign weights to the query vectors and sentiment features, and then concatenate the two types of features to obtain the final feature vector. Input the final feature representation into the decoder composed of two fully connected neural networks to output the prediction results of the target and stance.

[0022] Furthermore, in Step 2, use the preprocessed SST-2 sentiment analysis dataset to fine-tune the BERT model. This dataset has two sentiment labels: "positive" (positive sentiment) and "negative" (negative sentiment). The training set, validation set, and test set of this dataset contain 67349, 872, and 1821 data respectively. The specific data format includes: data number idx, input text sequence sentence, and the true sentiment label label corresponding to this text.

[0023] Furthermore, in step three, the encoder consists of a sequence sub-encoder and a fine-tuned sentiment sub-encoder, and the specific formula is as follows:

[0024] h sent = BERT sent (T)

[0025] h seq = BERT seq (T)

[0026] where h sent and h seq represent sequence features and sentiment features respectively, T = {w 1 , w 2 , w 3 ,..., w n} represents the input text sequence, and w i (1 ≤ i ≤ n) represents the words in the sequence.

[0027] Furthermore, in step four, the specific formula is as follows:

[0028]

[0029] q r = Attention(q org , q org , q org )

[0030] q = Attention(q r , h seq , h seq )

[0031] where Attention represents the attention mechanism, q org represents the randomly initialized original query embedding vector, q r represents the output vector of the internal attention layer of the query embedding vector, and q represents the output vector of the tweet feature attention layer; Q, K, V represent the query parameter, key parameter, and value parameter in the attention mechanism respectively; d k represents the vector dimension of the key parameter.

[0032] Furthermore, in step five, the specific formula is as follows:

[0033] α = softmax(w q tanh(q))

[0034] q all = qα T

[0035] where α represents the weight of the final query vector, qall It represents the vector representation that aggregates all query features.

[0036] Furthermore, in step six, the specific formula is as follows:

[0037]

[0038] Among them, pair represents the final prediction result that includes the target and the stance, w o represents the weights of different feature vectors, and FCN t and FCN s respectively represent the fully connected decoding layers for the target and the stance.

[0039] A user-generated content target-stance joint detection system based on an encoding-decoding structure, the system has program modules corresponding to the steps of the technical solution, and when running, executes the steps in the user-generated content target-stance joint detection method based on the encoding-decoding structure.

[0040] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the user-generated content target-stance joint detection method based on the encoding-decoding structure when called by a processor.

[0041] The present invention has the following beneficial technical effects:

[0042] The user-generated content target-stance joint detection method based on the encoding-decoding structure of the present invention realizes the stance detection of social media texts without manual annotation of the target, solves the dependence on manual labor in the stance detection task, avoids the generation of error cascading phenomena, and improves the accuracy of stance detection. Currently, the target-stance detection method for user-generated content usually adopts a two-stage method. First, in the first stage, the target object targeted by the user-generated content is identified, and then based on the identified target object, the stance detection is carried out in the second stage. Although this two-stage method effectively alleviates the dependence on manual annotation of target information in the stance detection task, the error generated in the first-stage target recognition task will directly affect the performance of the second-stage stance detection task, thus generating obvious error cascading phenomena. No one in the prior art has discovered and proposed such a technical problem that urgently needs to be solved. The present invention discovers the technical problems objectively existing in the prior art and the reasons for generating the technical problems. In view of the above problems and the reasons for their generation, the present invention specifically proposes a user-generated content target-stance joint detection method based on the encoding-decoding structure, which models the target information and stance information hidden in the user-generated content in an end-to-end manner, not only avoids error cascading, but also effectively improves the performance of target-stance detection, thereby further improving the feasibility, accuracy and practicality of the stance detection method in practical applications.

[0043] The present invention is applied to social network analysis, and more specifically, to the detection of the stance of social media texts without manual annotation of targets and related research. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a block diagram of the calculation process of the method of the present invention.

[0045] Figure 2 is an architecture diagram of a target stance joint detection model based on an encoder-decoder structure.

[0046] Figure 3 is a screenshot of an SST-2 data sample.

[0047] Figure 4 is a comparison diagram of target recognition results. DETAILED DESCRIPTION OF THE INVENTION

[0048] Combined with the attached Figures 1-4 , the implementation and technical effect verification of a method for jointly detecting the target stance of user-generated content based on an encoder-decoder structure according to the present invention are described as follows:

[0049] 1. User-generated content on social media platforms is full of network elements, usually does not follow traditional grammar rules, and shows obvious colloquial characteristics. Therefore, it needs to be preprocessed before calculating the user-generated content data. The data preprocessing can be divided into the following two steps: (1) Data cleaning: This step aims to eliminate network elements such as emojis and network links contained in user-generated content; (2) Conversion of Internet terms: Use a predefined abbreviation dictionary to convert common Internet terms in social media texts into written terms. For example, "4U" is converted to "For you"; "2nyt" is converted to "tonight".

[0050] 2. Use the preprocessed SST-2 sentiment analysis data set to fine-tune the BERT model. This data set has two sentiment labels: "positive" (positive sentiment) and "negative" (negative sentiment). The training set, validation set, and test set of this data set contain 67,349, 872, and 1,821 data respectively. The specific data format is as Figure 3 shown. Among them, 'idx' represents the data number,'sentence' represents the input text sequence, and 'label' represents the true sentiment label corresponding to this text.

[0051] 3. Input the preprocessed social media text data into the encoder, which consists of a sequence sub-encoder and an emotion sub-encoder. The sequence encoder is used to obtain the sequence features formed by the text context information, while the emotion encoder is used to obtain the emotion information hidden in the text sequence, providing richer semantic features for subsequent calculations. Both of these two sub-encoders are constructed based on BERT. Among them, the sequence encoder directly uses the pre-trained BERT model and uses the output vectors of all words as sequence features. The emotion encoder uses the BERT model fine-tuned on the SST-2 dataset in step 2 and uses the output vector of the [CLS] symbol as the emotion feature. The specific formula is as follows:

[0052] h sent = BERT sent (T)

[0053] h seq = BERT seq (T)

[0054] Among them, h sent and h seq represent the sequence feature and the emotion feature respectively, T = {w 1 , w 2 , w 3 ,..., w n} represents the input text sequence, and w i (1 ≤ i ≤ n) represents the word in the sequence.

[0055] 4. Randomly initialize m query embedding vectors. First, in the internal attention layer of the query embedding vectors, use the self-attention mechanism for the query embedding vectors to enable effective interaction and information transfer between the query embedding vectors, so as to better capture the dependencies between different query embedding vectors. The query vector, key vector, and value vector in the self-attention mechanism all come from the query embedding vectors themselves. Then, in the tweet feature attention layer, input the sequence features output by the encoder into the decoder and perform cross-attention mechanism with the interacted query embedding vectors to promote the interaction and fusion of the two features. In the cross-attention mechanism, use the query embedding vectors as the query vector, and the sequence features as the key vector and value vector. The specific formula is as follows:

[0056]

[0057] q r = Attention(q org , q org , q org )

[0058] q = Attention(q r , hseq , h seq )

[0059] Among them, Attention represents the attention mechanism, and q org represents the original query embedding vector initialized randomly, and q r represents the output vector of the internal attention layer of the query embedding vector, and q represents the output vector of the tweet feature attention layer; d k represents the vector dimension of the key parameter.

[0060] Q, K, and V represent the query parameter, key parameter, and value parameter in the attention mechanism respectively. q r and q are equivalent to assigning values to Q, K, and V and calling the attention function.

[0061] 5. Input all the query vectors that have incorporated sequence features into the target - stance aggregation layer. The main purpose of this layer is to make the model focus on the features that are more important or relevant to the current task, thereby improving the prediction performance of the model. First, assign corresponding weights to different query vectors and perform a weighted sum of all query vectors to obtain a vector representation that synthesizes all query features. This method enhances the positive impact of important query vectors while weakening the negative impact of irrelevant query vectors. The specific formula is as follows:

[0062] α = softmax(w q tanh(q))

[0063] q all = qα T

[0064] Among them, α represents the weight of the final query vector, and q all represents the vector representation that aggregates all query features;

[0065] w q is the weight of all query vectors, which is a trainable parameter and is initialized randomly; tanh(q) represents the activation function, and tanh is a common activation function.

[0066] 6. Input the aggregated query vector and the sentiment feature output by the encoder into the target stance pair decoding layer. First, assign weights to the query vector and the sentiment feature, and then concatenate the two features to obtain the final feature vector. Input the final feature representation into the decoder composed of two fully - connected neural networks to output the prediction results of the target and the stance. The specific formula is as follows:

[0067]

[0068] Among them, pair represents the final prediction result that includes the target and the stance, and w oRepresents the weights of different feature vectors, FCN t and FCN s respectively represent the fully connected decoding layers for the target and stance.

[0069] 7. To verify the effectiveness and superiority of the user-generated content stance detection method based on target information recognition proposed in the present invention, experiments are conducted using the dataset provided in reference [4]. This dataset is obtained by merging four commonly used stance detection datasets: SemEval-2016, AM, COVID-19, and P-stance. This dataset contains a total of 18 target categories, as shown in Table 1 specifically; and 4 stance categories: 'Dummy Stance', 'FAVOR', 'NONE', 'AGAINST', where 'Dummy Stance' is not predicted. To better fit the situation in real social networks, the 'Unrelated' target category is also added to this dataset to simulate the stance detection process in reality.

[0070] The experimental results are shown in Table 2, Figure 4 As shown. Since the proposed method is based on deep learning, in order to more comprehensively verify the model effect, three different seeds are selected in the experiment, and the final experimental results are the average of the three experimental results. The F1 value and accuracy are used as evaluation metrics in the target-stance detection experiment. The micro-averaged F1 is used as the evaluation metric in the target recognition experiment. It can be seen from the experimental results that the present method has achieved the best experimental results in both the target-stance detection task and the target recognition task, and has improved to varying degrees compared with other comparison models.

[0071] Table 1: Dataset target categories

[0072]

[0073] Table 2: Experimental results of target-stance detection

[0074]

[0075] 8. It can be seen from the above experimental results that the method proposed in the present invention has increased the F1 value and accuracy by 4.18% and 3.68% respectively compared with the best baseline model in the target-stance detection task, and has increased by 6.83% in the target recognition task. This fully demonstrates the superiority of the method proposed in the present invention.

[0076] In summary, through simulation experiments and practical applications, the method of the present invention has verified the claimed technical effects. The proposed method not only solves the problem of manual dependence in the stance detection task but also eliminates the error cascade phenomenon, effectively improving the feasibility and accuracy of the stance detection method in practical applications.

[0077] The algorithm (method) proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on the algorithm.

[0078] Based on the algorithm (method) proposed by the present invention, a user-generated content target stance joint detection system based on an encoding-decoding structure is developed using a programming language. The system has program modules corresponding to the steps of the above technical solution and executes the steps in the above-mentioned user-generated content target stance joint detection method based on an encoding-decoding structure when running.

[0079] The computer program of the developed system (software) is stored on a computer-readable storage medium. The computer program is configured to implement the steps of the above-mentioned user-generated content target stance joint detection method when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.

[0080] Based on the above method, a user-generated content target stance joint detection device can also be developed. The detection device includes at least one processor and a memory communicatively connected to the at least one processor. Among them, the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned user-generated content target stance joint detection method based on an encoding-decoding structure, realize real-time analysis of social networks, and realize stance detection of social media texts without manual annotation of targets. The user-generated content target stance joint detection device serves as the terminal product of the application of the present invention.

[0081] Various embodiments of the systems and technologies described herein can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0082] The computational procedures (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor, and these computational procedures can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disks, optical disks, memories, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0083] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this application can be achieved, and all are within the protection scope of the present invention.

[0084] The details of the references cited in the present invention are as follows:

[0085] [1] Yuan J, Zhao Y, Lu Y, et al. Ssr: Utilizing simplified stance reasoning process for robust stance detection[C] / / Proceedings of the 29th International Conference on Computational Linguistics. 2022: 6846 - 6858.

[0086] [2] Upadhyaya A, Fisichella M, Nejdl W. A multi-task model for emotion and offensive aided stance detection of climate change tweets[C] / / Proceedings of the ACM Web Conference 2023. 2023: 3948 - 3958.

[0087] [3] Ko Y, Ryu S, Han S, et al. KHAN: knowledge-aware hierarchical attention networks for accurate political stance prediction[C] / / Proceedings of the ACM Web Conference 2023. 2023: 1572-1583.

[0088] [4] Li Y, Garg K, Caragea C. A New Direction in Stance Detection: Target-Stance Extraction in the Wild[C] / / Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics(Volume 1: Long Papers). 2023: 10071-10085。

Claims

1. A method for joint detection of user-generated content target stance based on codec structure, characterized in that: The implementation process of the method is: Step 1: Preprocess user-generated content data: Data preprocessing is divided into the following two steps: (1) Data cleaning: used to remove emoticons and web link network elements contained in user-generated content; (2) Internet term conversion: using a predefined abbreviation dictionary to convert common Internet terms in social media texts into written terms; Step 2: Fine-tune the BERT model using a preprocessed sentiment analysis dataset, where the dataset has two sentiment labels: "positive" and "negative". The dataset includes a training set, a validation set, and a test set. The data format includes a data number, an input text sequence, and the true sentiment label corresponding to the text. Step 3: Input the preprocessed social media text data into the encoder, which consists of a sequence sub-encoder and a fine-tuned sentiment sub-encoder. Both the sequence sub-encoder and the sentiment sub-encoder are based on the BERT model. The sequence sub-encoder is used to obtain sequence features composed of text context information, while the sentiment sub-encoder is used to obtain the sentiment information implicit in the text sequence, providing richer semantic features for subsequent calculations. Step 4: Randomly initialize m query embedding vectors. First, in the internal attention layer of the query embedding vector, a self-attention mechanism is used on the query embedding vector so that the query embedding vectors can interact and transfer information effectively to capture the dependency between different query embedding vectors. The query vector, key vector, and value vector in the self-attention mechanism all come from the query embedding vector itself. Then, in the tweet feature attention layer, the sequence features output by the encoder are input into the decoder, and a cross-attention mechanism is performed with the query embedding vector after interaction to promote the interaction and fusion of the two features. In the cross-attention mechanism, the query embedding vector is used as the query vector, and the sequence features are used as the key vector and value vector; Step 5: All query vectors fused with sequence features are input into the target-stance aggregation layer. The main purpose of this layer is to make the model focus on features that are more important or relevant to the current task, thereby improving the prediction performance of the model. Assign corresponding weights to different query vectors and perform weighted summation on all query vectors to obtain a vector representation that integrates all query features; Step 6: Input the aggregated query vector and the sentiment features output by the encoder into the target stance pair decoding layer. First, assign weights to the query vector and the sentiment features, then concatenate the two features to obtain the final feature vector. The final feature representation is input into a decoder composed of two fully connected neural networks to output the prediction results of the target and stance.

2. The method for joint detection of user-generated content target stance based on codec structure according to claim 1, characterized in that: In step 2, the BERT model is fine-tuned using the preprocessed SST-2 sentiment analysis dataset, which has two sentiment labels: "positive" and "negative". The training set, validation set, and test set of this dataset contain 67,349, 872, and 1,821 data items, respectively. The specific data format includes: idx represents the data number, sentence represents the input text sequence, and label represents the true sentiment label corresponding to the text.

3. A method for joint detection of user-generated content target stance based on codec structure according to claim 1 or 2, characterized in that: In step 3, the encoder is composed of a sequence sub-encoder and a fine-tuned emotion sub-encoder. The specific formula is as follows: h sent =BERT sent (T) h seq =BERT seq (T) Among them, h sent and h seq Represent the sequence features and sentiment features respectively, T={w1,w2,w3,...,w n } represents the input text sequence, w i Represents the words in the sequence, 1≤i≤n.

4. The method for joint detection of user-generated content target stance based on codec structure according to claim 3, characterized in that: In step 4, the specific formula is as follows: q r =Attention(q org ,q org ,q org ) q=Attention(q r ,h seq ,h seq ) Among them, Attention represents the attention mechanism, q org represents the randomly initialized original query embedding vector, q r represents the output vector of the attention layer inside the query embedding vector, q represents the output vector of the tweet feature attention layer; Q, K, V represent the query parameter, key parameter, and value parameter in the attention mechanism respectively; d k The dimensions of the vector representing the key parameters.

5. The method for joint detection of user-generated content target stance based on codec structure according to claim 4, characterized in that In step five, the specific formula is as follows: α=softmax(w q tanh(q)) q all =qα T Among them, α represents the final query vector weight, q all Represents a vector representation that aggregates all query features.

6. The method for joint detection of user-generated content target stance based on codec structure according to claim 5, characterized in that In step six, the specific formula is as follows: Among them, pair represents the final prediction result including the target and the position, w o Represents the weights of different feature vectors, FCN t and FCN s Represent the fully connected decoding layers for target and stance, respectively.

7. A user-generated content target stance joint detection system based on a codec structure, characterized by: The system has a program module corresponding to the steps of any one of claims 1 to 6, and executes the steps in the method for joint detection of target stance of user-generated content based on codec structure when running.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the method for joint detection of target stance of user-generated content based on a codec structure according to any one of claims 1 to 6 when called by a processor.

Citation Information

Patent Citations

  • Standard detection method for specific target

    CN114330360A

  • Multi-target vertical field detection method and system based on comparative learning and consistency detection

    CN116992873A

  • User generated content standing detection method and system based on target information identification

    CN118070774A

  • Standard detection method based on multi-task learning

    CN114638195A

  • Network security protection method and network security protection platform based on big data

    CN115296870A