Cross-language and cross-target stance detection method based on collaborative guidance of dual expert models
Through the collaborative guidance of dual expert models, cross-language and cross-target training datasets are used to optimize the model loss weight, which solves the problem of insufficient utilization of unlabeled data in cross-language stance detection and achieves higher accuracy and generalization ability.
Patent Information
- Application Number
- CN202510820057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing methods fail to fully utilize unlabeled data in the target language and find it difficult to effectively align the semantic associations between the target and the text, resulting in insufficient generalization capabilities for cross-language and cross-target tasks.
A method based on collaborative guidance of dual expert models is adopted. By constructing cross-language and cross-target training datasets, cross-language expert models and cross-target expert models are designed. The models are optimized using prompt fine-tuning, consistency learning and contrastive learning. Combined with the unsupervised contrastive learning mechanism, loss weights are dynamically allocated, and multiple losses are integrated to improve model performance.
The accuracy, robustness and generalization ability of cross-lingual stance detection tasks have been improved, especially in data-scarce situations, and the adaptability and predictive ability of the model in multilingual and multi-target environments have been significantly improved.
Smart Images

Figure CN120354887B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-language and cross-target stance detection method based on collaborative guidance of dual expert models, and belongs to the technical field of natural language processing. Background Art
[0002] Stance detection is a key task in text mining and social media analysis, aiming to automatically identify users' attitudes toward specific targets (e.g., "support" or "oppose"). In recent years, with the growth of social media and user-generated content, stance detection has become particularly important in applications such as sentiment analysis and public opinion monitoring. While significant progress has been made in single-language stance detection, especially in research based on English datasets, cross-lingual stance detection remains a significant challenge due to the lack of sufficient annotated data in low-resource languages. Existing methods typically rely on language alignment techniques or pre-trained models (such as the mBERT model and XLM-RoBERTa) to bridge language differences. However, these methods are limited in their ability to leverage unlabeled data in the target language, particularly in terms of target category diversity and the effective use of unlabeled data, and thus fail to significantly improve the model's generalization ability across different languages and targets.
[0003] For example, Zhang et al. (2020) proposed a method for stance detection based on pre-trained cross-lingual models (such as XLM-R), using multilingual data for joint training. However, this method performs poorly on low-resource languages. In another study, Li et al. (2021) proposed a cross-lingual stance detection method based on language alignment, using alignment techniques to align the semantic spaces of the source and target languages. However, this method is not very effective when using unlabeled data in the target language. To enhance the robustness of cross-lingual models, pre-trained models such as BERT (Devlin et al., 2019) and mBERT (Pires et al., 2019) have been widely used in cross-lingual tasks. Although these models have improved transfer learning between different languages, most studies have focused solely on knowledge transfer between the source and target languages, lacking in-depth exploration of specific category information in the target language.
[0004] In addition, with the in-depth study of the problem of cross-lingual stance detection, some researchers have begun to pay attention to the potential semantic relationship between target categories. For example, Wang et al. (2022) proposed a stance detection method based on graph neural network (GNN), which captures the semantic connection between entities by constructing an entity relationship graph, thereby improving the cross-lingual reasoning ability of the model. In another study, Yuan et al. (2022) combined pre-training models and contrastive learning to improve the effect of stance detection by learning the relationship information between categories. Although these methods have improved the accuracy of cross-lingual stance detection to a certain extent, most of the existing research has ignored how to effectively utilize the unlabeled data of the target language, and how to effectively capture and optimize the semantic relationship between target categories in the target language.
[0005] Tie.
[0006] Therefore, to address the shortcomings of existing methods, the present invention proposes a cross-language and cross-target stance detection method based on the guidance of a dual-expert model. Summary of the Invention
[0007] The technical problem solved by the present invention is: the present invention provides a cross-language and cross-target stance detection method based on the collaborative guidance of a dual-expert model, which is used to solve the problem that existing methods fail to fully utilize unlabeled data in the target language and have difficulty in effectively aligning the semantic association between the target and the text. In particular, in the case of data scarcity, the generalization ability of cross-language and cross-target tasks cannot be effectively improved. The present invention improves the accuracy, robustness and generalization ability of cross-language stance detection tasks.
[0008] The technical solution of the present invention is a cross-language and cross-target stance detection method based on collaborative guidance of dual expert models, the method comprising:
[0009] Step 1: Build cross-language and cross-target training datasets, pre-process the source language training data and target language unlabeled data, and generate the input required for model training;
[0010] Step 2: Design a cross-lingual expert model and use prompt fine-tuning and consistency learning to optimize the adaptability of the mBERT model in cross-lingual tasks;
[0011] Step 3: Build a cross-target expert model to generate high-quality target representations through target category information mining and contrastive learning-based optimization methods;
[0012] Step 4: Utilize unlabeled data in the target language and optimize the representation capability of the collaborative module through an unsupervised comparative learning mechanism.
[0013] Step 5. Determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference. Under the dual-expert model framework, the cross-language expert model and the cross-target expert model are jointly used to guide the collaborative module for training, dynamically allocate loss weights, integrate the losses of the cross-language expert model and the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0014] Furthermore, the Step 1 includes:
[0015] Step 1.1: Build a source language training dataset ,in represents the target of the source language, represents the text in the source language, Indicates the source language stance label;
[0016] Step 1.2: Build an unlabeled dataset in the target language ,in represents the target language, Represents text in the target language. Unlabeled data does not contain stance labels.
[0017] Step 1.3: Perform text cleaning, word segmentation, and standardization on the source language training data, and introduce translation or alignment technology for the target language unlabeled data to generate prompt templates for cross-language tasks;
[0018] Step 1.4: Embed the target text pair generated by mBERT encoding into the predefined prompt template, complete the template instantiation by replacing the [target] and [text] placeholders in the template, and finally form a structured input.
[0019] Furthermore, the Step 2 includes:
[0020] Step 2.1: Use the mBERT model as the basis for the cross-lingual expert model and fine-tune the model using the prompt template of the cross-lingual task so that the model can capture the semantic features of the target language data.
[0021] Step 2.2, introduce the consistency learning mechanism, build the source language monolingual template and the target language cross-language template based on the prompt template of the cross-language task; establish the signal of cross-language semantic alignment through the consistency of the [mask] prediction probability distribution of the source language monolingual template and the target language cross-language template on the same sample; optimize the mBERT model parameters by using the mutual supervision between the source language representation of the monolingual template and the target language representation of the cross-language template; thus derive the cross-entropy loss of the monolingual template and cross-language template cross entropy , single language template cross entropy loss and cross-language template cross entropy It is expressed as follows:
[0022] (1);
[0023] (2);
[0024] in represents the true label of the sample, Indicates the prediction result obtained after the single language prompt template, It represents the prediction result obtained after the cross-language language prompt template, and N represents the number of samples;
[0025] Step 2.3: Use Kullback-Leibler divergence to calculate the distribution difference between the source language and the target language and integrate the cross entropy loss of the single language prompt template and the cross-language prompt template to obtain the cross-language expert model loss. , cross-lingual expert model loss Expressed as:
[0026] (3) ;
[0027] in, represents the distribution consistency loss, Hyperparameters for loss balancing.
[0028] Furthermore, the Step 3 includes:
[0029] Step 3.1: Using the source language data, we analyze the semantic association and stance consistency between the target and the text, and then use the mBERT model to initially learn the semantic representation of the target entity, thus generating a preliminary representation of the target.
[0030] Step 3.2: Use the k-means clustering method to classify the target into multiple target categories based on the similarity matrix of the target's preliminary representation. The target category is a label or category that classifies the target entity according to its semantic attribute, domain or type in the stance detection task.
[0031] Step 3.3. Calculate the category center representation for each target category and optimize the target representation using intra-class constraints and inter-class constraints to make the target representation within the target category more compact and the target representation between target categories more discriminative.
[0032] Step 3.4: Further optimize the target category through contrastive learning method, generalize the target category information to unseen targets in the target language, and obtain high-quality target representation. as follows:
[0033] (4);
[0034] in, Text in the input source language and target language goals The matching function, is the sample size;
[0035] Step 3.5: By combining contrastive learning and using cross-entropy loss to optimize the cross-target expert model, the stance detection performance of unseen targets is achieved. The cross-target expert model loss is as follows:
[0036] (5);
[0037] Where, It is The true labels of samples, is the prediction result of the cross-target expert model, Represents the cross entropy loss between the true label and the prediction results of the cross-target expert model.
[0038] Furthermore, the Step 4 includes:
[0039] Step 4.1: Using unlabeled data in the target language, extract embedding vectors of text and target through the encoder, and further optimize the target representation through contrastive learning;
[0040] Step 4.2: Use contrastive learning loss Optimize the feature representation capabilities of collaborative modules;
[0041] (6);
[0042] in, To enter text in the target language and the target language The matching function, is the sample size.
[0043] Furthermore, the Step 5 includes:
[0044] Step 5.1. Determine the confidence of the two expert models based on the absolute value of the predicted probability difference. Under the dual expert model framework, the cross-language expert model and the cross-target expert model are combined to guide the collaborative module to train, and the loss weight is dynamically allocated to obtain the cross-language expert model sample weighting coefficient. and cross-target expert model sample weighting coefficient , and The calculation formula is as follows:
[0045] (7);
[0046] (8);
[0047] in, and are the two output probabilities of the cross-lingual expert model, and are the two output probabilities of the cross-target expert model, is a small constant that prevents the denominator from being zero;
[0048] Step 5.2: Use the target language unlabeled data to perform stance detection through the collaborative module, and combine the output of the cross-language expert model and the cross-target expert model to perform joint training through weighted loss and contrastive learning. as follows:
[0049] (9);
[0050] in, represents the cross-target expert model loss, represents the cross-lingual expert model loss, represents the contrastive learning loss of the collaborative module in Step 4, is the weight of contrastive learning loss.
[0051] The present invention also provides a cross-language and cross-target stance detection system based on dual-expert model collaborative guidance, the system comprising: a module for executing the cross-language and cross-target stance detection method based on dual-expert model collaborative guidance.
[0052] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the cross-language and cross-target stance detection method based on collaborative guidance of dual expert models.
[0053] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the cross-language and cross-target stance detection method based on collaborative guidance of dual expert models is implemented.
[0054] The beneficial effects of the present invention are:
[0055] 1. This paper generates high-quality target representations through target category information mining and an optimization method based on contrastive learning to improve the model's generalization ability for unseen targets;
[0056] 2. This invention further optimizes the representation capability of the collaborative module through an unsupervised contrastive learning mechanism, enhancing the model's ability to identify complex target positions;
[0057] 3. Experimental results on the multilingual stance dataset X-STANCE show that the proposed method outperforms traditional methods across various benchmark datasets and task settings, particularly in cross-lingual tasks.
[0058] 4. By combining cross-language and cross-target expert guidance mechanisms, the present invention not only enhances the knowledge transfer between the source language and the target language, but also makes full use of the unlabeled data in the target language to improve the accuracy, robustness and generalization ability of the cross-language stance detection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Schematic diagram of the framework of the cross-language and cross-target stance detection method based on collaborative guidance of dual expert models in the present invention. DETAILED DESCRIPTION
[0060] Example 1: Figure 1 As shown, the following method proposed in this embodiment was experimented on the X-stance dataset. This paper uses the publicly available X-stance dataset (Vamvas and Sennrich, 2020) for experiments. This dataset focuses on Swiss political issues and contains two languages, German and French, with German as the source language and French as the target language. Each sample contains a question from a voter and an answer from a candidate, with positions categorized as "support" or "opposition." Based on the X-stance dataset, this paper constructs two sub-datasets:
[0061] (1) The Politics (P) dataset covers the fields of “foreign policy” and “immigration”, involving 31 different targets, including 7064 German samples and 2582 French samples;
[0062] (2) The Social (S) dataset covers the “society” and “security” domains, involving 32 different targets, including 7,362 German samples and 2,467 French samples.
[0063] In order to evaluate the effectiveness of this method, the present invention sets three different target configurations:
[0064] (1) All: All targets in the source and target languages are identical;
[0065] (2) Partial: Some targets in the source and target languages are the same, and 50% of the targets in the two languages are randomly selected to overlap;
[0066] (3) None: The targets of the source language and the target language do not overlap at all. 50% of the targets in the source language and the other 50% of the targets in the target language are randomly selected.
[0067] In this way, six different experimental settings (i.e., combinations of two sub-datasets with three target configurations) were formed to evaluate the effectiveness of this method. The statistics of the datasets are shown in Tables 1 and 2:
[0068] Table 1 shows the statistics of the political field data set
[0069]
[0070] Table 2 shows the statistics of the social domain dataset
[0071]
[0072] A cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model, the method comprising:
[0073] Step 1: Build cross-language and cross-target training datasets, pre-process the source language training data and target language unlabeled data, and generate the input required for model training;
[0074] Furthermore, the Step 1 includes:
[0075] Step 1.1: Build a source language training dataset ,in represents the target of the source language, represents the text in the source language, Indicates the source language stance label;
[0076] Step 1.2: Build an unlabeled dataset in the target language ,in represents the target language, Represents text in the target language. Unlabeled data does not contain stance labels.
[0077] Step 1.3: Perform text cleaning, word segmentation, and standardization on the source language training data, and introduce translation or alignment technology for the target language unlabeled data to generate prompt templates for cross-language tasks;
[0078] Step 1.4: Embed the target text pair generated by mBERT encoding into the predefined prompt template, complete the template instantiation by replacing the [target] and [text] placeholders in the template, and finally form a structured input.
[0079] Step 2: Design a cross-lingual expert model and use prompt fine-tuning and consistency learning to optimize the adaptability of the mBERT model in cross-lingual tasks;
[0080] Furthermore, the Step 2 includes:
[0081] Step 2.1: Use the mBERT model as the basis for the cross-lingual expert model and fine-tune the model using the prompt template of the cross-lingual task so that the model can capture the semantic features of the target language data.
[0082] Step 2.2, introduce the consistency learning mechanism, build the source language monolingual template and the target language cross-language template based on the prompt template of the cross-language task; establish the signal of cross-language semantic alignment through the consistency of the [mask] prediction probability distribution of the source language monolingual template and the target language cross-language template on the same sample; optimize the mBERT model parameters by using the mutual supervision between the source language representation of the monolingual template and the target language representation of the cross-language template; thus derive the cross-entropy loss of the monolingual template and cross-language template cross entropy , single language template cross entropy loss and cross-language template cross entropy It is expressed as follows:
[0083] (1);
[0084] (2);
[0085] in represents the true label of the sample, Indicates the prediction result obtained after the single language prompt template, It represents the prediction result obtained after the cross-language language prompt template, and N represents the number of samples;
[0086] Step 2.3. Use Kullback-Leibler divergence to calculate the distribution difference between the source language and the target language and integrate the cross entropy loss of the single language prompt template and the cross-language prompt template to finally obtain the cross-language expert model loss. , cross-lingual expert model loss Expressed as:
[0087] (3) ;
[0088] in, represents the distribution consistency loss, Hyperparameters for loss balancing.
[0089] Step 3: Build a cross-target expert model. By mining target category information and optimizing it with contrastive learning, we generate high-quality target representations to improve the model’s generalization ability for unseen targets.
[0090] Furthermore, the Step 3 includes:
[0091] Step 3.1: Using the source language data, we analyze the semantic association and stance consistency between the target and the text, and then use the mBERT model to initially learn the semantic representation of the target entity, thus generating a preliminary representation of the target.
[0092] Step 3.2: Use the k-means clustering method to classify the target into multiple target categories based on the similarity matrix of the target's preliminary representation. The target category is a label or category that classifies the target entity according to its semantic attribute, domain or type in the stance detection task.
[0093] Step 3.3. Calculate the category center representation for each target category and optimize the target representation using intra-class constraints and inter-class constraints to make the target representation within the target category more compact and the target representation between target categories more discriminative.
[0094] Step 3.4: Further optimize the target category through contrastive learning method, generalize the target category information to unseen targets in the target language, and obtain high-quality target representation. as follows:
[0095] (4);
[0096] in, Text in the input source language and the target language The matching function, is the sample size;
[0097] Step 3.5: By combining contrastive learning and using cross-entropy loss to optimize the cross-target expert model, the stance detection performance of unseen targets is achieved. The cross-target expert model loss is as follows:
[0098] (5);
[0099] Where, It is The true labels of samples, is the prediction result of the cross-target expert model, Represents the cross entropy loss between the true label and the prediction results of the cross-target expert model.
[0100] Step 4: Utilize unlabeled data in the target language and optimize the representation capabilities of the collaborative module (CoModule) through an unsupervised contrastive learning mechanism to enhance the model's ability to recognize complex target positions.
[0101] Furthermore, the Step 4 includes:
[0102] Step 4.1: Using unlabeled data in the target language, extract embedding vectors of text and target through the encoder, and further optimize the target representation through contrastive learning;
[0103] Step 4.2: Use contrastive learning loss Optimize the feature representation capability of the collaborative module (CoModule);
[0104] (6);
[0105] in, To enter text in the target language and the target language The matching function, is the sample size.
[0106] Step 5. Determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference. Under the dual-expert model framework, the cross-language expert model and the cross-target expert model are jointly used to guide the collaborative module for training, dynamically allocate loss weights, integrate the losses of the cross-language expert model and the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0107] Furthermore, the Step 5 includes:
[0108] Step 5.1. Determine the confidence of the two expert models based on the absolute value of the predicted probability difference. Under the dual expert model framework, the cross-language expert model and the cross-target expert model are combined to guide the collaborative module to train, and the loss weight is dynamically allocated to obtain the cross-language expert model sample weighting coefficient. and cross-target expert model sample weighting coefficient , and The calculation formula is as follows:
[0109] (7);
[0110] (8);
[0111] in, and are the two output probabilities of the cross-lingual expert model, and are the two output probabilities of the cross-target expert model, is a small constant that prevents the denominator from being zero;
[0112] Step 5.2: Use the target language unlabeled data to perform stance detection through the collaborative module, and combine the output of the cross-language expert model and the cross-target expert model to perform joint training through weighted loss and contrastive learning. as follows:
[0113] (9);
[0114] in, represents the cross-target expert model loss, represents the cross-lingual expert model loss, represents the contrastive learning loss of the collaborative module in Step 4, is the weight of contrastive learning loss.
[0115] The present invention also provides a cross-language and cross-target stance detection system based on collaborative guidance of dual expert models, the system comprising:
[0116] The dataset acquisition module is used to build cross-language and cross-target training datasets, preprocess the source language training data and target language unlabeled data, and generate the input required for model training;
[0117] A cross-lingual expert model design module, which optimizes the adaptability of the mBERT model in cross-lingual tasks using hint fine-tuning and consistency learning;
[0118] A cross-target expert model building module is used to generate high-quality target representations through target category information mining and contrastive learning-based optimization methods;
[0119] The collaborative module optimization module is used to optimize the representation ability of the collaborative module through an unsupervised contrastive learning mechanism using unlabeled data in the target language;
[0120] The joint training and stance detection module is used to determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference. Under the dual-expert model framework, the cross-language expert model and the cross-target expert model are jointly guided to train the collaborative module, dynamically allocate loss weights, integrate the losses of the cross-language expert model and the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0121] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the cross-language and cross-target stance detection method based on collaborative guidance of dual expert models.
[0122] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the cross-language and cross-target stance detection method based on collaborative guidance of dual expert models is implemented.
[0123] The dual-expert collaborative framework designed by the present invention combines the cross-language expert model and the cross-target expert model to jointly guide the training of the collaborative module, thereby effectively improving the performance of the model on the target language task. Secondly, the cross-language expert model and the cross-target expert model are used to obtain preliminary stance information from the perspective of language and target respectively, and on this basis, more accurate stance predictions are derived. Finally, for the unlabeled data of the target language, the performance of the stance detection model is further optimized through a dual knowledge distillation mechanism. Specifically, the model optimizes the collaborative relationship between cross-language and cross-target experts by designing a loss function, thereby enhancing the predictive ability of the target language task. The method of the present invention solves the problem that existing methods fail to fully utilize shared knowledge between languages and targets in cross-language and cross-target tasks, thereby improving the adaptability and accuracy of the stance detection model in multi-language and multi-target environments. Comparative experimental results show that the method of the present invention has achieved significant performance improvement in cross-language and cross-target stance detection tasks.
[0124] This paper uses the F1 score as the primary evaluation metric. The F1 score combines the model's precision and recall, comprehensively measuring the model's predictive power across different categories. A higher F1 score indicates that the model is able to better balance correct classifications and missed classifications when processing samples, resulting in stronger overall performance.
[0125] To verify the effectiveness of the proposed model, we selected several single-language methods and baseline methods for cross-language tasks for comparison. The single-language methods and baseline methods for cross-language tasks are described below:
[0126] BiCond (Augenstein et al., 2016): This method incorporates target information into text representations through a bidirectional conditional LSTM (BiLSTM), aiming to improve target-related stance detection performance. By introducing contextual information about the target, the model can better capture the semantic information related to the target.
[0127] TAN (Du et al., 2017): This method uses an attention mechanism to extract target-related representations from text, enhancing the model's ability to perceive the target. This method uses adaptive attention weights to focus on parts closely related to the target, thereby improving performance on cross-target tasks.
[0128] CrossNet (Xu et al., 2018): uses a self-attention mechanism to learn target-independent representations, specifically for cross-target stance detection tasks. It captures global information in the text through the self-attention mechanism and decouples target information from the influence of other irrelevant targets to enhance the model's generalization ability.
[0129] JointCL (Liang et al., 2022b): This method uses a target-aware prototype graph contrastive learning method to model the zero-shot stance detection task. This method constructs a target-related graph structure and uses contrastive learning to enable the model to perform effective stance detection without target samples.
[0130] In addition, the proposed method is compared with some models specifically designed for cross-language tasks, including:
[0131] CLKD (Xu and Yang, 2017): This method trains a classifier based on labeled source language data and transfers the source language knowledge to the target language through knowledge distillation. It uses an alignment mechanism between the source and target language models to reduce the differences between the source and target languages, thereby improving the stance detection performance of the target language.
[0132] ADAN (Chen et al., 2018): This method aligns the representations of the source and target languages through adversarial training, aiming to reduce language differences in cross-lingual tasks. This method improves the performance of cross-lingual models by aligning the embedding spaces of the source and target languages using a generative adversarial network (GAN).
[0133] mBERT-FT (Devlin et al., 2019): This method improves cross-lingual performance by fine-tuning the mBERT model on the target task. This method leverages the mBERT model’s multilingual capabilities and adapts it to a specific stance detection task through fine-tuning.
[0134] mBERT-PT: This method uses stance templates from the source language to perform prompt tuning on the mBERT model to improve the accuracy of stance detection in the target language. This method enables the mBERT model to better understand the target information in the target language by performing prompt tuning on the target language task.
[0135] CCSD (Zhang et al., 2023): This approach uses a dual knowledge distillation framework involving cross-language and cross-target experts to extract language-specific knowledge from source language data to address stance detection in unlabeled target language data. This approach enables knowledge transfer between the source and target languages and is particularly well-suited for cross-lingual tasks in low-resource languages. By comparing these baselines, we can comprehensively evaluate the effectiveness of our proposed model across various tasks and language settings.
[0136] The method of the present invention is compared with the baseline model on the X-STANCE dataset. In order to verify the effectiveness of the model proposed by the present invention, the method of the present invention is compared with the baseline model. The specific results are shown in Table 3 and Table 4:
[0137] (1) Overall performance: On both datasets (“Politics-All” (politics - the same goal) and “Society-All” (society - the same goal)), our method achieves the best F1 score in all experimental settings, indicating that our method has a strong advantage in the cross-lingual stance detection task.
[0138] (2) Comparison with monolingual baseline methods: Compared with traditional monolingual methods, such as JointCL, our method achieves significant improvements in all target settings. In particular, in the "Politics-All" and "Society-All" settings, the F1 scores are 5.12% and 3.39% higher than the best-performing baseline models, respectively, highlighting the role of cross-language transfer in promoting task performance.
[0139] (3) JointCL performs significantly better than attention-based TAN and CrossNet when handling target inconsistency by constructing prototype graphs, which shows that the strategy of modeling the target inconsistency problem by mining and utilizing high-level features is effective.
[0140] (4) Comparison with cross-lingual methods: In terms of cross-lingual methods, the fine-tuning and hint tuning strategies of the mBERT model have achieved significant results, demonstrating the superior adaptability of pre-trained language models in handling cross-lingual tasks. Although CLKD transfers information from source language data to the target model through knowledge distillation, it still shows strong performance in the absence of labeled data in the target language, further verifying its effectiveness in cross-lingual environments.
[0141] Table 3 shows the experimental results on the X-STANCE political domain dataset.
[0142]
[0143] Table 4 shows the experimental results on the X-STANCE social domain dataset
[0144]
[0145] Advantages of the method of the present invention: Compared with other cross-language task models, such as CCSD, although the latter also adopts a cross-language and cross-target dual knowledge distillation framework, the method of the present invention has achieved more significant performance improvements by introducing collaborative modules and fine-tuning different target settings. The method of the present invention has shown great advantages in dealing with the consistency of different targets. Under the "Politics-All" and "Society-All" settings, the F1 scores are improved by 5.12% and 3.39% respectively compared with the baseline method, showing the strong potential of the model in cross-language stance detection. Through the above analysis, it can be seen that the method of the present invention has shown better performance than traditional methods under different benchmark datasets and task settings, especially in cross-language tasks.
[0146] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A cross-language and cross-target stance detection method based on collaborative guidance of dual-expert models, characterized by: The method comprises: Step 1: Build cross-language and cross-target training datasets, pre-process the source language training data and target language unlabeled data, and generate the input required for model training; Step 2: Design a cross-lingual expert model and use prompt fine-tuning and consistency learning to optimize the adaptability of the mBERT model in cross-lingual tasks; Step 3: Build a cross-target expert model to generate high-quality target representations through target category information mining and contrastive learning-based optimization methods; Step 4: Utilize unlabeled data in the target language and optimize the representation capability of the collaborative module through an unsupervised comparative learning mechanism. Step 5: Determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference. Within the dual-expert model framework, the cross-language expert model and the cross-target expert model are combined to guide the collaborative module training. Loss weights are dynamically allocated, and the losses of the cross-language expert model, the cross-target expert model, and the contrastive learning loss are integrated. The overall performance of the model is improved by optimizing the total loss. Step 2 includes: Step 2.1: Use the mBERT model as the basis for the cross-lingual expert model and fine-tune the model using the prompt template of the cross-lingual task so that the model can capture the semantic features of the target language data. Step 2.2, introduce the consistency learning mechanism, build the source language monolingual template and the target language cross-language template based on the prompt template of the cross-language task; establish the signal of cross-language semantic alignment through the consistency of the [mask] prediction probability distribution of the source language monolingual template and the target language cross-language template on the same sample; optimize the mBERT model parameters by using the mutual supervision between the source language representation of the monolingual template and the target language representation of the cross-language template; thus derive the cross-entropy loss of the monolingual template and cross-language template cross entropy , single language template cross entropy loss and cross-language template cross entropy It is expressed as follows: (1); (2); in represents the true label of the sample, Indicates the prediction result obtained after the single language prompt template, It represents the prediction result obtained after the cross-language language prompt template, and N represents the number of samples; Step 2.3: Use Kullback-Leibler divergence to calculate the distribution difference between the source language and the target language and integrate the cross entropy loss of the single language prompt template and the cross-language prompt template to obtain the cross-language expert model loss. , cross-lingual expert model loss Expressed as: (3) ; in, represents the distribution consistency loss, Balancing hyperparameters for loss; Step 3 includes: Step 3.1: Using the source language data, we analyze the semantic association and stance consistency between the target and the text, and then use the mBERT model to initially learn the semantic representation of the target entity, thus generating a preliminary representation of the target. Step 3.2: Use the k-means clustering method to classify the target into multiple target categories based on the similarity matrix of the target's preliminary representation. The target category is a label or category that classifies the target entity according to its semantic attribute, domain or type in the stance detection task. Step 3.
3. Calculate the category center representation for each target category and optimize the target representation using intra-class constraints and inter-class constraints to make the target representation within the target category more compact and the target representation between target categories more discriminative. Step 3.4: Further optimize the target category through contrastive learning method, generalize the target category information to unseen targets in the target language, and obtain high-quality target representation. as follows: (4); in, Text in the input source language and target language goals The matching function, is the sample size; Step 3.5: By combining contrastive learning and using cross-entropy loss to optimize the cross-target expert model, the stance detection performance of unseen targets is achieved. The cross-target expert model loss is as follows: (5); Where, It is The true labels of samples, is the prediction result of the cross-target expert model, Represents the cross entropy loss between the true label and the prediction results of the cross-target expert model.
2. The cross-language and cross-target stance detection method based on dual-expert model collaborative guidance according to claim 1, characterized in that: Step 1 includes: Step 1.1: Build a source language training dataset ,in represents the target of the source language, represents the text in the source language, Indicates the source language stance label; Step 1.2: Build an unlabeled dataset in the target language ,in represents the target language, Represents text in the target language. Unlabeled data does not contain stance labels. Step 1.3: Perform text cleaning, word segmentation, and standardization on the source language training data, and introduce translation or alignment technology for the target language unlabeled data to generate prompt templates for cross-language tasks; Step 1.4: Embed the target text pair generated by mBERT encoding into the predefined prompt template, complete the template instantiation by replacing the [target] and [text] placeholders in the template, and finally form a structured input.
3. The cross-language and cross-target stance detection method based on dual-expert model collaborative guidance according to claim 1, characterized in that: Step 4 includes: Step 4.1: Using unlabeled data in the target language, extract embedding vectors of text and target through the encoder, and further optimize the target representation through contrastive learning; Step 4.2: Use contrastive learning loss Optimize the feature representation capabilities of collaborative modules; (6); in, To enter text in the target language and target language goals The matching function, is the sample size.
4. The cross-language and cross-target stance detection method based on dual-expert model collaborative guidance according to claim 1, characterized in that: Step 5 includes: Step 5.
1. Determine the confidence of the two expert models based on the absolute value of the predicted probability difference. Under the dual expert model framework, the cross-language expert model and the cross-target expert model are combined to guide the collaborative module to train, and the loss weight is dynamically allocated to obtain the cross-language expert model sample weighting coefficient. and cross-target expert model sample weighting coefficient , and The calculation formula is as follows: (7); (8); in, and are the two output probabilities of the cross-lingual expert model, and are the two output probabilities of the cross-target expert model, is a small constant that prevents the denominator from being zero; Step 5.2: Use the target language unlabeled data to perform stance detection through the collaborative module, and combine the output of the cross-language expert model and the cross-target expert model to perform joint training through weighted loss and contrastive learning. as follows: (9); in, represents the cross-target expert model loss, represents the cross-lingual expert model loss, represents the contrastive learning loss of the collaborative module in Step 4, is the weight of contrastive learning loss.
5. A cross-language and cross-target stance detection system based on collaborative guidance of dual expert models, characterized by: The system includes: a module for executing the cross-language and cross-target stance detection method based on dual-expert model collaborative guidance as described in any one of claims 1 to 4.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the cross-language and cross-target stance detection method based on dual-expert model collaborative guidance is implemented as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-language and cross-target stance detection method based on dual-expert model collaborative guidance is implemented as described in any one of claims 1 to 4.