Cross-language cross-target standing field detection method based on collaborative guidance of double expert models
Through the method of collaborative guidance of dual expert models, cross-language and cross-objective training data sets are constructed, cross-language and cross-objective expert models are designed, and the target language has no label data is used to optimize the model performance. The problem of insufficient semantic association utilization in cross-language position detection is solved, and efficient generalization of cross-language and cross-objective tasks is achieved.
Patent Information
- Application Number
- CN202510820057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing methods fail to make full use of labelless data in the target language and make it difficult to effectively align semantic associations between targets and text, resulting in insufficient generalization capabilities for cross-language and cross-target tasks.
Using a method based on the collaborative guidance of the dual expert model, a cross-language and cross-objective training data set is constructed, a cross-language expert model and a cross-objective expert model are designed, and the model is optimized through prompt fine-tuning, consistent learning and contrast learning. The target language has label-free data, dynamically allocate loss weights, and integrate losses to improve model performance.
The accuracy, robustness and generalization capabilities of cross-language position detection tasks are improved, especially in the case of data scarcity, and the adaptability of the model in multilingual and multi-objective environments is significantly improved.
Smart Images

Figure CN120354887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross - language cross - target stance detection method based on collaborative guidance of a dual - expert model, belonging to the technical field of natural language processing. Background Art
[0002] Stance Detection is a key task in text mining and social media analysis, aiming to automatically identify a user's attitude towards a specific target (such as "support" or "oppose"). In recent years, with the increase in social media and user - generated content, stance detection has become particularly important in application scenarios such as sentiment analysis and public opinion monitoring. Although significant progress has been made in monolingual stance detection, especially in research based on English datasets, cross - language stance detection still faces major challenges due to the lack of sufficient labeled data in low - resource languages. Existing methods usually rely on language alignment techniques or pre - trained models (such as the mBERT model, XLM - RoBERTa, etc.) to bridge language differences, but these methods have limited performance in utilizing unlabeled data in the target language, especially in terms of the diversity of target categories and the effective utilization of unlabeled data, and cannot significantly improve the generalization ability of the model across different languages and targets.
[0003] For example, Zhang et al. (2020) proposed an idea of stance detection by using a pre - trained cross - language model (such as XLM - R) and conducting joint training with multilingual data, but this method has poor performance in low - resource languages. Another study, Li et al. (2021) proposed a cross - language stance detection method based on language alignment, which uses alignment techniques to align the semantic spaces of the source language and the target language, but this method has poor performance in utilizing unlabeled data in the target language. To enhance the robustness of cross - language models, pre - trained models such as BERT (Devlin et al., 2019) and the mBERT model (Pires et al., 2019) have been widely applied to cross - language tasks. Although the transfer learning between different languages of these models has been improved, most studies only focus on the knowledge transfer between the source language and the target language, lacking in - depth mining of the specific category information of the target language.
[0004] In addition, with the in-depth study of cross-lingual stance detection problems, some researchers have begun to pay attention to the potential semantic relationships between target categories. For example, Wang et al. (2022) proposed a stance detection method based on graph neural networks (GNNs), which constructs an entity relationship graph to capture the semantic connections between entities and enhances the cross-lingual reasoning ability of the model. Another study, Yuan et al. (2022), combined pre-trained models and contrastive learning to improve the effectiveness of stance detection by learning the relationship information between categories. Although these methods have improved the accuracy of cross-lingual stance detection to a certain extent, most existing studies have ignored how to effectively utilize unlabeled data in the target language and how to effectively capture and optimize the semantic relationships between target categories in the target language.
[0005] Therefore, aiming at the deficiencies of existing methods, the present invention proposes a cross-lingual cross-target stance detection method guided by a dual-expert model. Summary of the Invention
[0006] The technical problem solved by the present invention is: The present invention provides a cross-lingual cross-target stance detection method guided by a dual-expert model to solve the problem that existing methods fail to fully utilize unlabeled data in the target language and are difficult to effectively align the semantic associations between targets and texts, especially in the case of scarce data, where the generalization ability of cross-lingual and cross-target tasks cannot be effectively improved. The present invention improves the accuracy, robustness, and generalization ability of the cross-lingual stance detection task.
[0007] The technical solution of the present invention is: A cross-lingual cross-target stance detection method guided by a dual-expert model, the method comprising:
[0008] Step1, construct a cross-lingual and cross-target training data set, preprocess the source language training data and the target language unlabeled data to generate the input required for model training;
[0009] Step2, design a cross-lingual expert model, and use prompt tuning and consistency learning to optimize the adaptability of the mBERT model in cross-lingual tasks;
[0010] Step3, construct a cross-target expert model, and generate high-quality target representations through target category information mining and an optimization method based on contrastive learning;
[0011] Step4, utilize the target language unlabeled data to optimize the representation ability of the collaborative module through an unsupervised contrastive learning mechanism;
[0012] Step 5. Determine the confidence levels of the two expert models by calculating the absolute value of the predicted probability difference; under the dual-expert model framework, jointly guide the collaborative module for training using the cross-lingual expert model and the cross-target expert model, dynamically allocate loss weights, integrate the losses of the cross-lingual expert model, the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0013] Further, the said Step 1 includes:
[0014] Step 1.1. Construct a source language training dataset , where represents the target of the source language, represents the text of the source language, represents the source language stance label;
[0015] Step 1.2. Construct a target language unlabeled dataset , where represents the target of the target language, represents the text of the target language, and the unlabeled data does not contain stance labels;
[0016] Step 1.3. Clean, tokenize, and standardize the source language training data, and introduce translation or alignment techniques for the target language unlabeled data to generate a prompt template for the cross-lingual task;
[0017] Step 1.4. Embed the target text pairs generated by mBERT encoding into the predefined prompt template, complete the template instantiation by replacing the [target] and [text] placeholders in the template, and finally form a structured input.
[0018] Further, the said Step 2 includes:
[0019] Step 2.1. Use the mBERT model as the basis for the cross-lingual expert model, and fine-tune the model using the prompt template for the cross-lingual task to enable the model to capture the semantic features of the target language data;
[0020] Step 2.2. Introduce a consistency learning mechanism, and based on the prompt template for the cross-lingual task, construct a source language monolingual template and a target language cross-lingual template; establish a signal for cross-lingual semantic alignment through the consistency of the [mask] prediction probability distributions of the source language monolingual template and the target language cross-lingual template on the same samples; utilize the mutual supervision between the source language representation of the monolingual template and the target language representation of the cross-lingual template to optimize the mBERT model parameters; thereby obtaining the cross-entropy loss of the monolingual template and the cross-entropy of the cross-lingual template , the cross-entropy loss of the monolingual template and the cross-entropy of the cross-lingual template It is expressed as follows:
[0021] (1);
[0022] (2);
[0023] where represents the true label of the sample, represents the prediction result obtained after the monolingual prompt template, represents the prediction result obtained after the cross - language prompt template, and N represents the number of samples;
[0024] Step2.3. Calculate the distribution difference between the source language and the target language using the Kullback - Leibler divergence, and integrate the cross - entropy losses of the monolingual prompt template and the cross - language prompt template. Finally, obtain the cross - language expert model loss , the cross - language expert model loss is expressed as:
[0025] (3);
[0026] where represents the distribution consistency loss, is the loss balance hyperparameter.
[0027] Furthermore, the said Step3 includes:
[0028] Step3.1. Use the source - language data. By analyzing the semantic association and stance consistency between the target and the text, and then through the mBERT model, obtain the semantic representation of the target entity obtained by preliminary learning, that is, generate the preliminary representation of the target;
[0029] Step3.2. Adopt the k - means clustering method. According to the similarity matrix of the preliminary representation of the target, divide the target into multiple target categories, where the target category is a label or category for classifying the target entity according to its semantic attributes, fields, or types in the stance detection task;
[0030] Step3.3. Calculate the category - center representation for each target category, and use the intra - class constraint and inter - class constraint to optimize the target representation, so as to make the target representations within the target category more compact and the target representations between target categories more distinguishable;
[0031] Step3.4. Through the contrastive learning method, further optimize the target category, and generalize the target category information to the unseen targets in the target language to obtain high - quality target representations. The contrastive learning loss is as follows:
[0032] (4);
[0033] Among them, is the text of the source language input and the target of the target language matching function, is the number of samples;
[0034] Step3.5. Optimize the cross-target expert model by combining contrastive learning and using cross-entropy loss to achieve the stance detection performance for unseen targets. The loss of the cross-target expert model is as follows:
[0035] (5);
[0036] In the formula, is the true label of the th sample, is the prediction result of the cross-target expert model, represents the cross-entropy loss between the true label and the prediction result of the cross-target expert model.
[0037] Furthermore, the said Step4 includes:
[0038] Step4.1. Use the unlabeled data in the target language, extract the embedding vectors of the text and the target through the encoder, and further optimize the target representation through contrastive learning;
[0039] Step4.2. Use the contrastive learning loss to optimize the feature representation ability of the collaborative module;
[0040] (6);
[0041] Among them, is the text of the input target language and the target of the target language matching function, is the number of samples.
[0042] Furthermore, the said Step5 includes:
[0043] Step5.1. Determine the confidence of the two expert models based on the absolute value of the prediction probability difference. Under the dual-expert model framework, jointly guide the collaborative module to train the cross-language expert model and the cross-target expert model, dynamically allocate the loss weights, and thus obtain the sample weighting coefficient of the cross-language expert model and the sample weighting coefficient and The calculation formulas of are as follows:
[0044] (7);
[0045] (8);
[0046] Among them, and are two output probabilities of the cross - language expert model, and are two output probabilities of the cross - target expert model, is a small constant to prevent the denominator from being zero;
[0047] Step5.2. Use the unlabeled data in the target language to perform stance detection through the collaborative module, and combine the outputs of the cross - language expert model and the cross - target expert model, and perform joint training through weighted loss and contrast learning. The comprehensive loss is as follows:
[0048] (9);
[0049] Among them, represents the loss of the cross - target expert model, represents the loss of the cross - language expert model, represents the contrast learning loss of the collaborative module in Step4, is the weight of the contrast learning loss.
[0050] The present invention also provides a cross - language and cross - target stance detection system based on the collaborative guidance of a dual - expert model. The system includes a module for executing the cross - language and cross - target stance detection method based on the collaborative guidance of the dual - expert model.
[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the cross - language and cross - target stance detection method based on the collaborative guidance of the dual - expert model.
[0052] The present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the cross - language and cross - target stance detection method based on the collaborative guidance of the dual - expert model.
[0053] The beneficial effects of the present invention are:
[0054] 1. Through the mining of target category information and the optimization method based on contrast learning, the present invention generates high - quality target representations to improve the generalization ability of the model for unseen targets;
[0055] 2. The present invention further optimizes the representation ability of the collaborative module through an unsupervised contrastive learning mechanism, enhancing the model's ability to identify complex target stances.
[0056] 3. The experimental results of the present invention on the multilingual stance dataset X-STANCE show that the present invention exhibits superior performance compared to traditional methods under different benchmark datasets and task settings, especially having a greater advantage in cross-lingual tasks.
[0057] 4. By combining cross-lingual and cross-target expert guidance mechanisms, the present invention enhances knowledge transfer between the source language and the target language while making full use of unlabeled data in the target language to improve the accuracy, robustness, and generalization ability of cross-lingual stance detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic diagram of the framework of the cross-lingual and cross-target stance detection method based on the collaborative guidance of a dual-expert model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Example 1: As Figure 1 shown, the following method proposed in this example was experimented on the X-stance dataset. The present invention used the publicly available X-stance dataset (Vamvas and Sennrich, 2020) for experiments. This dataset focuses on Swiss political issues and contains two languages, German and French, where German is the source language and French is the target language. Each sample contains a voter's question and a candidate's answer, and the stance can be divided into "support" or "oppose". Based on the X-stance dataset, the present invention established two sub-datasets:
[0060] (1) Political (P) dataset, covering the fields of "foreign policy" and "immigration", involving 31 different targets, including 7064 German samples and 2582 French samples;
[0061] (2) Social (S) dataset, covering the fields of "society" and "security", involving 32 different targets, including 7362 German samples and 2467 French samples.
[0062] To evaluate the effectiveness of this method, the present invention set three different target configurations:
[0063] (1) All: All targets in the source language and the target language are exactly the same;
[0064] (2) Partial: Some targets in the source language and the target language are the same, and 50% of the targets in the two languages are randomly selected to overlap;
[0065] (3) None: The targets of the source language and the target language do not overlap at all. 50% of the targets of the source language are randomly selected and combined with the other 50% of the targets of the target language.
[0066] In this way, six different experimental settings (i.e., the combination of 2 sub-datasets and 3 target configurations) are formed to evaluate the effectiveness of this method. The statistical information of the datasets is shown in Tables 1 and 2 as follows:
[0067] Table 1 shows the statistics of the political domain dataset
[0068] Table 2 shows the statistics of the social domain dataset
[0069] A cross-language and cross-target stance detection method based on the collaborative guidance of a dual-expert model, the method includes:
[0070] Step1. Construct a cross-language and cross-target training dataset, preprocess the source language training data and the target language unlabeled data to generate the input required for model training;
[0071] Further, the Step1 includes:
[0072] Step1.1. Construct a source language training dataset , where represents the target of the source language, represents the text of the source language, represents the source language stance label;
[0073] Step1.2. Construct a target language unlabeled dataset , where represents the target of the target language, represents the text of the target language, and the unlabeled data does not contain stance labels;
[0074] Step1.3. Clean, tokenize, and standardize the source language training data, and introduce translation or alignment techniques for the target language unlabeled data to generate a prompt template for the cross-language task;
[0075] Step1.4. Embed the target text pairs generated by mBERT encoding into the predefined prompt template, and complete the template instantiation by replacing the [target] and [text] placeholders in the template, and finally form a structured input.
[0076] Step2. Design a cross-language expert model, and use prompt fine-tuning and consistency learning to optimize the adaptation ability of the mBERT model in cross-language tasks;
[0077] Furthermore, Step 2 includes:
[0078] Step 2.1: Use the mBERT model as the basis of the cross - language expert model, and fine - tune the model using the prompt template of the cross - language task, so that the model can capture the semantic features of the target - language data;
[0079] Step 2.2: Introduce a consistency learning mechanism. Based on the prompt template of the cross - language task, construct a source - language monolingual template and a target - language cross - language template; establish a signal for cross - language semantic alignment through the consistency of the [mask] prediction probability distributions of the source - language monolingual template and the target - language cross - language template on the same samples; use the mutual supervision between the source - language representation of the monolingual template and the target - language representation of the cross - language template to optimize the parameters of the mBERT model; thus obtain the cross - entropy loss of the monolingual template and the cross - entropy of the cross - language template The cross - entropy loss of the monolingual template and the cross - entropy of the cross - language template are expressed as follows:
[0080] (1);
[0081] (2);
[0082] where represents the true label of the sample, represents the prediction result obtained after passing through the monolingual prompt template, represents the prediction result obtained after passing through the cross - language prompt template, and N represents the number of samples;
[0083] Step 2.3: Use the Kullback - Leibler divergence to calculate the distribution difference between the source language and the target language and integrate the cross - entropy losses of the monolingual prompt template and the cross - language prompt template, and finally obtain the cross - language expert model loss The cross - language expert model loss is expressed as:
[0084] (3);
[0085] where, represents the distribution consistency loss, is the loss balance hyperparameter.
[0086] Step 3: Construct a cross - target expert model, and through target - category information mining and an optimization method based on contrastive learning, generate high - quality target representations to improve the generalization ability of the model for unseen targets;
[0087] Furthermore, Step 3 includes:
[0088] Step 3.1: Using the source language data, through analyzing the semantic association and stance consistency between the target and the text, and then obtaining the semantic representation of the target entity initially learned by the mBERT model, that is, generating the initial representation of the target;
[0089] Step 3.2: Adopting the k-means clustering method, according to the similarity matrix of the initial representation of the target, dividing the target into multiple target categories, where the target category is a label or category for classifying the target entity according to its semantic attributes, domain, or type in the stance detection task;
[0090] Step 3.3: Calculating the category center representation for each target category, and using the intra-class constraint and inter-class constraint to optimize the target representation, which is used to make the target representations within the target category closer and the target representations between target categories more distinguishable;
[0091] Step 3.4: Through the contrastive learning method, further optimizing the target category, generalizing the target category information to the unseen targets in the target language, and obtaining a high-quality target representation, the contrastive learning loss is as follows:
[0092] (4);
[0093] where, is the text of the input source language and the target in the target language 's matching function, is the number of samples;
[0094] Step 3.5: By combining contrastive learning and using the cross-entropy loss to optimize the cross-target expert model to achieve the stance detection performance for unseen targets, the cross-target expert model loss is as follows:
[0095] (5);
[0096] In the formula, is the true label of the th sample, is the prediction result of the cross-target expert model, represents the cross-entropy loss between the true label and the prediction result of the cross-target expert model.
[0097] Step 4: Using the unlabeled data in the target language, through the unsupervised contrastive learning mechanism, optimizing the representation ability of the collaborative module (CoModule) to enhance the model's ability to identify complex target stances;
[0098] Furthermore, Step 4 includes:
[0099] Step 4.1: Use the unlabeled data in the target language, extract the embedding vectors of the text and the target through the encoder, and further optimize the target representation through contrastive learning;
[0100] Step 4.2: Use the contrastive learning loss to optimize the feature representation ability of the collaborative module (CoModule);
[0101] (6);
[0102] Among them, is the matching function of the input text in the target language and the target in the target language , is the number of samples.
[0103] Step 5: Determine the confidence levels of the two expert models by calculating the absolute value of the predicted probability difference; in the dual-expert model framework, jointly guide the collaborative module for training by the cross-language expert model and the cross-target expert model, dynamically allocate loss weights, integrate the losses of the cross-language expert model, the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0104] Furthermore, Step 5 includes:
[0105] Step 5.1: Determine the confidence levels of the two expert models based on the absolute value of the predicted probability difference. In the dual-expert model framework, jointly guide the collaborative module for training by the cross-language expert model and the cross-target expert model, dynamically allocate loss weights, so as to obtain the sample weighting coefficient of the cross-language expert model and the sample weighting coefficient of the cross-target expert model . The calculation formulas of
[0106] (7);
[0107] (8);
[0108] Among them, and are the two output probabilities of the cross-language expert model, and are the two output probabilities of the cross-target expert model, is a small constant to prevent the denominator from being zero;
[0109] Step5.2. Use the unlabeled data in the target language to perform stance detection through the collaborative module, and combine the outputs of the cross-language expert model and the cross-target expert model. Conduct joint training through weighted loss and contrastive learning, and the comprehensive loss is as follows:
[0110] (9);
[0111] wherein, represents the loss of the cross-target expert model, represents the loss of the cross-language expert model, represents the contrastive learning loss of the collaborative module in Step4, is the weight of the contrastive learning loss.
[0112] The present invention also provides a cross-language and cross-target stance detection system based on the collaborative guidance of a dual-expert model. The system includes:
[0113] A dataset acquisition module, which is used to construct a cross-language and cross-target training dataset, preprocess the source language training data and the unlabeled data in the target language, and generate the input required for model training;
[0114] A cross-language expert model design module, which is used to optimize the adaptability of the mBERT model in cross-language tasks by using prompt fine-tuning and consistency learning;
[0115] A cross-target expert model construction module, which is used to generate high-quality target representations through target category information mining and an optimization method based on contrastive learning;
[0116] A collaborative module optimization module, which is used to optimize the representation ability of the collaborative module by using the unlabeled data in the target language through an unsupervised contrastive learning mechanism;
[0117] A joint training and stance detection module, which is used to determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference; under the framework of the dual-expert model, jointly guide the collaborative module to train the cross-language expert model and the cross-target expert model, dynamically allocate loss weights, integrate the losses of the cross-language expert model, the cross-target expert model, and the contrastive learning loss, and improve the overall performance of the model by optimizing the total loss.
[0118] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the cross-language and cross-target stance detection method based on the collaborative guidance of the dual-expert model.
[0119] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the cross-language and cross-target stance detection method based on the collaborative guidance of the dual-expert model is implemented.
[0120] The dual-expert collaborative framework designed by the present invention combines a cross-language expert model and a cross-target expert model to jointly guide the collaborative module for training, thereby effectively improving the performance of the model in target language tasks. Secondly, the cross-language expert model and the cross-target expert model are used to obtain preliminary stance information from the perspectives of language and target respectively, and on this basis, more accurate stance predictions are derived. Finally, for the unlabeled data in the target language, through a dual knowledge distillation mechanism, the performance of the stance detection model is further optimized. Specifically, the model designs a loss function to optimize the collaborative relationship between the cross-language and cross-target experts, enhancing the prediction ability of the target language task. The method of the present invention solves the problem that the existing methods fail to fully utilize the shared knowledge between languages and targets in cross-language and cross-target tasks, thereby improving the adaptability and accuracy of the stance detection model in a multi-language and multi-target environment. The comparative experiment results show that the method of the present invention has achieved significant performance improvement in cross-language and cross-target stance detection tasks.
[0121] The present invention uses the F1 score as the main evaluation index. The F1 score combines the precision and recall of the model and can comprehensively measure the prediction ability of the model in different categories. The higher the F1 score, the better the model can balance correct classification and missed classification when processing samples, thus showing stronger overall performance.
[0122] To verify the effectiveness of the model proposed by the present invention, we selected several single-language methods and baseline methods for cross-language tasks for comparison. The single-language methods and baseline methods for cross-language tasks are introduced as follows:
[0123] BiCond (Augenstein et al., 2016): Incorporates target information into the text representation through a bidirectional conditional LSTM (BiLSTM), aiming to improve the target-related stance detection performance. By introducing the context information of the target, this method enables the model to better capture the target-related semantic information.
[0124] TAN (Du et al., 2017): Adopts an attention mechanism to extract target-related representations from the text, enhancing the model's perception ability of the target. This method uses adaptive attention weights to focus on the parts closely related to the target, thereby improving the performance in cross-target tasks.
[0125] CrossNet (Xu et al., 2018): It uses the self-attention mechanism to learn target-agnostic representations, especially for cross-target stance detection tasks. It captures global information in the text through the self-attention mechanism and decouples the target information from the influence of other irrelevant targets to enhance the generalization ability of the model.
[0126] JointCL (Liang et al., 2022b): It adopts a target-aware prototype graph contrastive learning method to model the zero-shot stance detection task. This method constructs a target-related graph structure and, through contrastive learning, enables the model to perform effective stance detection without target samples.
[0127] In addition, the method of the present invention is compared with some models specifically for cross-language tasks, including:
[0128] CLKD (Xu and Yang, 2017): This method trains a classifier based on labeled source language data and transfers the knowledge of the source language to the target language through knowledge distillation. It adopts an alignment mechanism between the source language model and the target language model to reduce the differences between the source language and the target language, thereby improving the stance detection performance of the target language.
[0129] ADAN (Chen et al., 2018): It aligns the representations of the source language and the target language through an adversarial training method, aiming to reduce language differences in cross-language tasks. This method aligns the embedding spaces of the source language and the target language through a generative adversarial network (GAN), improving the performance of the cross-language model.
[0130] mBERT model - FT (Devlin et al., 2019): It improves the performance in cross-language tasks by fine-tuning the mBERT model on the target task. This method utilizes the powerful capabilities of the mBERT model in a multilingual environment and fine-tunes it to adapt to a specific stance detection task.
[0131] mBERT model - PT: It uses the stance templates of the source language to perform prompt tuning on the mBERT model to improve the accuracy of target language stance detection. This method performs prompt tuning on the target language task, enabling the mBERT model to better understand the target information in the target language.
[0132] CCSD (Zhang et al., 2023): Through a dual knowledge distillation framework of cross - language experts and cross - target experts, it extracts language - target - related knowledge from source - language data to solve the stance detection problem of unlabeled data in the target language. This method can achieve knowledge transfer between the source language and the target language, and is particularly suitable for cross - language tasks in low - resource languages. Through the comparison with the above baseline methods, we can comprehensively evaluate the effectiveness of the model proposed in the present invention under different tasks and different language settings.
[0133] The method of the present invention was compared with the baseline model on the X - STANCE dataset. To verify the effectiveness of the model proposed in the present invention, the method of the present invention was compared with the baseline model, and the specific results are shown in Tables 3 and 4 as follows:
[0134] (1) Overall performance: On the two datasets ("Politics - All" (Politics - Identical Targets) and "Society - All" (Society - Identical Targets)), the method of the present invention showed the best F1 scores under all experimental settings, indicating that the method of the present invention has strong advantages in cross - language stance detection tasks.
[0135] (2) Comparison with monolingual baseline methods: Compared with traditional monolingual methods such as JointCL, the method of the present invention achieved significant improvements under all target settings. Especially in the "Politics - All" and "Society - All" settings, the F1 scores were 5.12% and 3.39% higher than the best - performing baseline models respectively, highlighting the promoting effect of cross - language transfer on task performance.
[0136] (3) When dealing with target inconsistency by constructing a prototype graph, JointCL performed significantly better than the attention - based TAN and CrossNet, indicating that the strategy of modeling target - inconsistency problems by mining and utilizing high - level features is effective.
[0137] (4) Comparison with cross - language methods: In terms of cross - language methods, both the fine - tuning and prompt - tuning strategies of the mBERT model achieved remarkable results, demonstrating the superior adaptability of pre - trained language models in handling cross - language tasks. Although CLKD transfers the information of source - language data to the target model through knowledge distillation, it still showed strong performance in the absence of target - language annotation data, further verifying its effectiveness in cross - language environments.
[0138] Table 3 shows the experimental results on the X - STANCE political domain dataset
[0139] Table 4 shows the experimental results on the X-STANCE social domain dataset
[0140] Advantages of the method of the present invention: Compared with models for other cross-lingual tasks, such as CCSD, although the latter also adopts a dual knowledge distillation framework for cross-lingual and cross-target, the method of the present invention achieves a more significant performance improvement by introducing a collaborative module and fine-tuning for different targets. The method of the present invention particularly shows great advantages in dealing with different target consistencies. Under the "Politics-All" and "Society-All" settings, the F1 scores are respectively improved by 5.12% and 3.39% compared with the baseline method, demonstrating the strong potential of the model in cross-lingual stance detection. Through the above analysis, it can be seen that the method of the present invention shows better performance than traditional methods under different benchmark datasets and task settings, especially having great advantages in cross-lingual tasks.
[0141] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A cross - language and cross - target stance detection method based on collaborative guidance of a dual - expert model, characterized in that: The method includes: Step 1: Construct a cross - language and cross - target training dataset, preprocess the source - language training data and target - language unlabeled data to generate the input required for model training. Step 2: Design a cross - language expert model, and use prompt fine - tuning and consistency learning to optimize the adaptability of the mBERT model in cross - language tasks. Step 3: Construct a cross - target expert model, and generate high - quality target representations through target - category information mining and an optimization method based on contrastive learning. Step 4: Use the target - language unlabeled data to optimize the representation ability of the collaborative module through an unsupervised contrastive learning mechanism. Step 5: Determine the confidence of the two expert models by calculating the absolute value of the predicted probability difference; under the dual - expert - model framework, jointly guide the collaborative module to train the cross - language expert model and the cross - target expert model, dynamically allocate loss weights, integrate the losses of the cross - language expert model, the cross - target expert model, and the contrastive - learning loss, and improve the overall performance of the model by optimizing the total loss.
2. The cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model according to claim 1, wherein: The said Step 1 includes: Step 1.
1. Construct a source language training dataset , where represents the target of the source language, represents the text of the source language, represents the source language stance label; Step1.
2. Construct a target-language dataset without labels , where represents the target in the target language, represents the text in the target language, and the unlabeled data does not contain stance labels; Step 1.3: Perform text cleaning, tokenization, and normalization on the source - language training data, and introduce translation or alignment techniques for the target - language unlabeled data to generate a prompt template for cross - language tasks. Step 1.4: Embed the target text pairs generated by mBERT encoding into the predefined prompt template, complete the template instantiation by replacing the placeholders [target] and [text] in the template, and finally form a structured input.
3. The cross-language cross-target stance detection method based on collaborative guidance of a dual-expert model according to claim 1, wherein: The said Step 2 includes: Step 2.1: Use the mBERT model as the basis of the cross - language expert model, and fine - tune the model using the prompt template for cross - language tasks to enable the model to capture the semantic features of the target - language data. Step 2.2: Introduce a consistency learning mechanism. Based on the prompt templates for cross-lingual tasks, construct source-language monolingual templates and target-language cross-lingual templates; establish signals for cross-lingual semantic alignment by the consistency of the [mask] prediction probability distributions of the source-language monolingual templates and the target-language cross-lingual templates on the same samples; utilize the mutual supervision between the source-language representations of the monolingual templates and the target-language representations of the cross-lingual templates to optimize the mBERT model parameters; thereby obtaining the cross-entropy loss of the monolingual templates and the cross-entropy of the cross-lingual templates , the cross-entropy loss of the monolingual templates and the cross-entropy of the cross-lingual templates are expressed as follows: (1); (2); wherein represents the true label of the sample, represents the prediction result obtained after the single - language prompt template, represents the prediction result obtained after the cross - language prompt template, and N represents the number of samples; Step 2.3: Calculate the distribution difference between the source language and the target language using the Kullback-Leibler divergence, and integrate the cross-entropy losses of the monolingual prompt template and the cross-lingual prompt template. Finally, obtain the cross-lingual expert model loss , the cross-lingual expert model loss is expressed as: (3) ; Among them, represents the distribution consistency loss, is the loss balance hyperparameter.
4. The cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model according to claim 1, wherein: The said Step 3 includes: Step 3.1: Use the source - language data, analyze the semantic association and stance consistency between the target and the text, and then obtain the semantic representation of the target entity initially learned by the mBERT model, that is, generate a preliminary representation of the target. Step 3.2: Adopt the k - means clustering method, and divide the targets into multiple target categories according to the similarity matrix of the preliminary representations of the targets, where the target category is a label or category for classifying the target entity according to its semantic attributes, domain, or type in the stance - detection task. Step 3.3: Calculate the class - center representation for each target category, and use intra - class constraints and inter - class constraints to optimize the target representation, so as to make the target representations within the target category closer and the target representations between target categories more discriminative. Step 3.
4. Further optimize the target categories through a contrastive learning method, generalize the target category information to unseen targets in the target language, and obtain high-quality target representations. The contrastive learning loss is as follows: (4); Among them, is the text of the source language input and the target of the target language matching function, is the number of samples; Step 3.5: Optimize the cross - target expert model by combining contrastive learning and using cross - entropy loss to achieve the stance - detection performance for unseen targets. The loss of the cross - target expert model is as follows: (5); wherein, is the true label of the th sample, is the prediction result of the cross-target expert model, represents the cross-entropy loss between the true label and the prediction result of the cross-target expert model.
5. The cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model according to claim 1, characterized in that: The said Step 4 includes: Step 4.1: Use the target - language unlabeled data, extract the embedding vectors of the text and the target through an encoder, and further optimize the target representation through contrastive learning. Step4.2: Use contrastive learning loss Optimize the feature representation ability of the collaborative module; (6); Among them, is the text in the input target language and the target of the target language matching function, is the number of samples.
6. The cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model according to claim 1, characterized in that: The said Step 5 includes: Step 5.
1. Determine the confidence levels of the two expert models based on the absolute value of the predicted probability difference. Under the dual-expert model framework, jointly train the cross-language expert model and the cross-target expert model to guide the collaborative module, and dynamically allocate the loss weights to obtain the sample weighting coefficients of the cross-language expert model and the sample weighting coefficients of the cross-target expert model , and The calculation formulas are as follows: (7); (8); Among them, and are two output probabilities of the cross - language expert model, and are two output probabilities of the cross - target expert model, is a small constant to prevent the denominator from being zero; Step 5.2: Use the unlabeled data in the target language to perform stance detection through the collaboration module, and combine the outputs of the cross-language expert model and the cross-target expert model for joint training through weighted loss and contrastive learning, and the comprehensive loss is as follows: (9); Among them, represents the loss of the cross-target expert model, represents the loss of the cross-language expert model, represents the contrastive learning loss of the collaborative module in Step4, is the weight of the contrastive learning loss.
7. A cross - language and cross - target stance detection system based on collaborative guidance of a dual - expert model, characterized in that, The system includes: a module for executing the cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model as described in any one of claims 1 to 6.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cross-language and cross-target stance detection method based on collaborative guidance of a dual-expert model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Zero-sample text field detection method based on mixed contrast learning and generative data enhancement
CN115758159A
Unknown target vertical field detection method and device based on graph contrast learning
CN116257632A
Zero sample vertical field detection device and method based on multilevel knowledge selection
CN116628196A
Low-resource Chinese-crossing cross-language abstract generation method based on multi-stage fine tuning
CN117493546A
Social media public opinion identification method based on cross-language transfer learning algorithm framework
CN118069843A