Zero sample cross-language text classification method based on attention adaptive migration
By using the attention adaptive transfer method in cross-language text classification, the importance of important words is extracted and modeled, and through pseudo-label generation and student model training, the problem of scarcity of target language annotation data and neglect of important words is solved, and efficient cross-language text classification is achieved.
Patent Information
- Application Number
- CN202510126979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
Existing cross-language text classification methods are difficult to effectively classify when target language labeling data is scarce, and traditional methods ignore the dynamic evaluation of important words, resulting in the model being unable to fully capture semantic information and adapt to the contextual characteristics of the target language during cross-language migration.
A zero-sample cross-language text classification method based on attention adaptive migration is proposed. By extracting seed words from the source language and performing importance modeling, an importance matrix is generated; a teacher model is used to map seed words to the target language, and through pseudo-label generation and student model training, the text classification model of the target language is gradually optimized.
Without the need for a large amount of labeling data, the text classification performance of the target language is significantly improved, and the words that play an important role in classification are effectively paid attention to, solving the problem of insufficient attention to vocabulary importance and semantic differences in cross-language migration in traditional methods.
Smart Images

Figure CN120067329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a zero-shot cross-lingual text classification method based on attention adaptive transfer, belonging to the technical field of natural language processing. Background Art
[0002] The cross-lingual text classification task refers to training a model only relying on source language annotated data and migrating it to the target language, which requires a certain scale of annotated data in the target language.
[0003] The main bottleneck faced by existing supervised learning methods in cross-lingual text classification is the high cost of obtaining target language annotated data. Especially for low-resource languages, the annotated data is extremely scarce, which makes it impossible for many languages to effectively utilize these methods. In addition, zero-shot research methods usually assume that the importance of all words for the classification task is equal, ignoring the necessity of dynamically evaluating and emphasizing key words. Especially in multilingual texts, certain words may play a decisive role in the classification task, and ignoring these key words may seriously affect the accuracy of model classification because it fails to fully utilize the most informative part of the text for decision-making. Traditional methods lack targeted modeling capabilities in this regard, resulting in the model being unable to fully capture the hierarchy and difference of semantic information. More importantly, these methods often cannot adapt to the context characteristics of the target language in the zero-shot setting and cannot effectively handle semantic shifts caused by language structure and cultural background differences. To solve these problems, the present invention proposes a zero-shot cross-lingual text classification method based on attention adaptive transfer, namely AATC, for cross-lingual text classification. Summary of the Invention
[0004] The technical problem solved by the present invention is: The present invention provides a zero-shot cross-lingual text classification method based on attention adaptive transfer to address the problems of scarce target language annotated data and traditional cross-lingual text classification methods ignoring important words; without a large amount of annotated data, it significantly improves the text classification performance of the target language, and at the same time effectively focuses on the words that play an important role in classification by modeling the importance of seed words. The present invention has achieved good results in the cross-lingual text classification task.
[0005] The technical solution of the present invention is: A zero-shot cross-lingual text classification method based on attention adaptive transfer, the method comprising:
[0006] Step1. Extract seed words on the source language based on the word distribution characteristics of the text, and model their importance to generate an importance matrix; train a teacher model using a large amount of source language annotated data, and further emphasize key information through the word probability distribution of the seed words during the training process;
[0007] Step 2. Map the seed words captured by the teacher model to the target language through the cross-bilingual dictionary mapping relationship;
[0008] Step 3. Use the teacher model to generate pseudo-labels for the unlabeled data containing the seed words in the target language, and use the pseudo-labels as the initial training data for the student model to train the student model;
[0009] Step 4. The student model further predicts the unlabeled documents in the target language, generates new labeled data and expands the training set, and finally obtains the student model for the target language classification task through iterative optimization.
[0010] As a further solution of the present invention, the Step 1 includes:
[0011] Step 1.1. Given the source language document Its bag-of-words encoding is The weight matrix W trained on the source language document dataset D S The classifier predicts the category If the word Increases the probability through the weight W kc > 0 Then select the word As the seed word; the set of all seed words is The seed word set is larger than the translation budget B, so only the most important seed words are left through sparse regularization Seed word The calculation formula is expressed as:
[0012]
[0013] Where L is the loss function, R sparse (·) is the sparse regularizer, that is, the L1 norm, λ B Represents the hyperparameter used to control R sparse ; y i Represents the true label of the i-th sample, Represents finding the weight matrix W that minimizes the entire expression, Represents the predicted value of the i-th sample, N represents the sum of the loss functions, K represents the total number of categories, |V S | Represents the length of the source language vocabulary;
[0014] Step 1.2. For the set of seed words, use the pre-trained GloVe 300-dimensional word embeddings for importance modeling; for the GloVe word embedding vectors of all seed words, calculate the dot product between them to obtain a self-attention score matrix A representing the similarity between words; the self-attention score matrix A is expressed as:
[0015] A = E·E Τ (2)
[0016] Among them, the dimension of matrix A is n×n, where n is the number of seed words, and E represents the matrix composed of GloVe embeddings of all seed words in G S ;
[0017] Step1.3. Use the Softmax function to normalize each row of the self-attention score matrix A to obtain the self-attention weight matrix. This process ensures that the sum of the attention weights of each word to other words is equal to 1, thereby obtaining the probability distribution Q ij :
[0018]
[0019] Among them, Q ij represents the attention weight between the i-th word and the j-th word, A ij represents the original correlation between the i-th word and the j-th word, A ik represents the original correlation between the i-th word and the k-th word, and n represents the number of seed words.
[0020] As a further solution of the present invention, the said Step2 includes:
[0021] Use the MUSE bilingual dictionary to obtain the translation of the seed words extracted from the source language corresponding to the target language; transfer the seed word weights of the source language to the corresponding vocabulary in the target language to initialize a teacher classifier.
[0022] As a further solution of the present invention, the said Step3 includes:
[0023] Step3.1. Use the teacher classifier to predict the pseudo-label q for the unlabeled document in the target language containing seed words; the calculation formula of the pseudo-label q j is expressed as: j
[0024]
[0025] Among them, is the bag-of-words encoding of the unlabeled document in the target language , w' k Τ is the k-th row of; thus obtaining indicating the probability that the j-th sample corresponds to the k-th category;
[0026] Step3.2. Next, train the student model using the context of the seed words For the unlabeled document in the target language The predicted probability r of the student model j :
[0027]
[0028] where f T represents the prediction function with weight parameter θ, and the prediction function f T adopts a pre-trained Transformer classifier or a pre-trained monolingual Bert classifier; represents the probability that the j-th sample corresponds to the K-th category;
[0029] The student is trained through the "distillation" objective:
[0030]
[0031] where H(q, r) = -∑ k q k log 2 r k is the cross-entropy between the teacher's prediction and the student's prediction, R(·) is the sparse regularizer, i.e., the L2 norm, and the hyperparameter λ is used to control R.
[0032] As a further aspect of the present invention, the Step 4 includes:
[0033] The student makes predictions on all unlabeled documents in D T to obtain a larger Use D” T to train the student until convergence.
[0034] The present invention also provides a zero-shot cross-lingual text classification system based on attention adaptive transfer, and the system includes: a module for executing the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0036] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and characterized in that when the computer program is executed by a processor, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0037] The present invention also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer.
[0038] The beneficial effects of the present invention are as follows:
[0039] 1. Aiming at the problem that important words are ignored during cross-lingual transfer in cross-lingual text classification, the present invention effectively focuses on the words that play an important role in classification by modeling the importance of seed words, and solves the problem that traditional methods pay insufficient attention to the importance of vocabulary and semantic differences during cross-lingual transfer;
[0040] 2. Aiming at the problem of scarce data in low-resource languages, the present invention combines cross-lingual alignment and pseudo-label generation strategies to significantly improve the text classification performance of the target language without a large amount of labeled data;
[0041] 3. The present invention uses the attention mechanism to dynamically model the importance of seed words, combines the bilingual dictionary alignment relationship, and fully captures the vocabulary and its context information, so as to achieve efficient cross-lingual transfer and text classification. Description of the Drawings
[0042] Figure 1 is the flowchart in the present invention;
[0043] Figure 2 is the zero-shot cross-lingual text classification method based on attention adaptive transfer proposed by the present invention. Detailed Embodiments
[0044] Embodiment 1: As Figure 1 - Figure 2 shown, a zero-shot cross-lingual text classification method based on attention adaptive transfer, the method includes:
[0045] Step1. Extract seed words on the source language based on the word distribution characteristics of the text, and model their importance to generate an importance matrix; use a large-scale source language labeled data to train a teacher model, and further emphasize key information through the word probability distribution of the seed words during the training process;
[0046] As a further solution of the present invention, the Step1 includes:
[0047] Step1.1. Given a source language document whose bag-of-words encoding is the weight matrix W trained on the source language document dataset D S the classifier predicts the category If the word increases the probability through the weight W kc > 0 Then select the word as the seed word; the set of all seed words is The seed word set is larger than the translation budget B, so only the most important seed words are left through sparse regularization Seed word The calculation formula is expressed as:
[0048]
[0049] where L is the loss function, and R sparse (·) is the sparse regularizer, that is, the L1 norm, and λ B represents the hyperparameter used to control R sparse ; y i represents the true label of the i-th sample, represents finding the weight matrix W that minimizes the entire expression, represents the predicted value of the i-th sample, N represents the sum of the loss functions, K represents the total number of categories, and |V S | represents the length of the source language vocabulary;
[0050] Step1.2. For the seed word set, use the pre-trained GloVe 300-dimensional word embeddings for importance modeling; for the GloVe word embedding vectors of all seed words, calculate the dot product between them to obtain a self-attention score matrix A representing the similarity between words; the self-attention score matrix A is expressed as:
[0051] A = E·E Τ (2)
[0052] where the dimension of matrix A is n×n, n is the number of seed words, and E represents the matrix composed of the GloVe embeddings of all seed words in G S ;
[0053] Step1.3. Use the Softmax function to normalize each row of the self-attention score matrix A to obtain the self-attention weight matrix. This process ensures that the sum of the attention weights of each word to other words is equal to 1, thereby obtaining the probability distribution Q ij :
[0054]
[0055] where Q ij represents the attention weight between the i-th word and the j-th word, A ij represents the original correlation between the i-th word and the j-th word, A ik represents the original correlation between the i-th word and the k-th word, and n represents the number of seed words.
[0056] Step 2. Map the seed words captured by the teacher model to the target language through the cross-bilingual dictionary mapping relationship;
[0057] As a further solution of the present invention, the Step 2 includes:
[0058] Use the MUSE bilingual dictionary to obtain the translation of the seed words extracted from the source language corresponding to the target language; transfer the weights of the seed words in the source language to the corresponding vocabulary in the target language for initializing a teacher classifier.
[0059] Step 3. Use the teacher model to generate pseudo-labels for the unlabeled data in the target language containing the seed words, and use the pseudo-labels as the initial training data of the student model to train the student model;
[0060] As a further solution of the present invention, the Step 3 includes:
[0061] Step 3.1. Use the teacher classifier to predict the pseudo-label q for the unlabeled documents in the target language containing the seed words j ; The calculation formula of the pseudo-label q j is expressed as:
[0062]
[0063] where is the bag-of-words encoding of the unlabeled document in the target language , w' k Τ is the k-th row of; thus obtaining representing the probability that the j-th sample corresponds to the k-th category;
[0064] Step 3.2. Next, use the context of the seed words to train the student model for the unlabeled documents in the target language the student model predicts the probability r j :
[0065]
[0066] where f T represents the prediction function with weight parameters θ, and the prediction function f T adopts a pre-trained Transformer classifier or a pre-trained monolingual Bert classifier. In this embodiment, a pre-trained monolingual Bert classifier is adopted, which captures the word composition of a specific language; represents the probability that the j-th sample corresponds to the K-th category;
[0067] The student is trained with the "distillation" objective:
[0068]
[0069] where H(q,r) = -∑ k q k log 2 r k is the cross-entropy between the teacher's prediction and the student's prediction, R(·) is the sparsity regularizer, i.e., the L2 norm, and the hyperparameter λ is used to control R.
[0070] Step4. The student model further predicts the unannotated documents in the target language, generates new annotated data and expands the training set, and finally obtains the student model for the target language classification task through iterative optimization. As a further solution of the present invention, the Step4 includes:
[0071] The student predicts all the unannotated documents in the target language in D T to obtain a larger Use D” T to train the student until convergence.
[0072] The present invention also provides a zero-shot cross-lingual text classification system based on attention adaptive transfer, and the system includes: a module for executing the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0073] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0074] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and characterized in that when the computer program is executed by a processor, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0075] The present invention also provides a computer program product, including a computer program, and characterized in that when the computer program is executed by a processor, it implements the zero-shot cross-lingual text classification method based on attention adaptive transfer as described above.
[0076] To illustrate the effect of the present invention, the present invention has carried out experimental verification.
[0077] The experiment uses accuracy to evaluate the classification performance of the model. The specific formula is as follows:
[0078]
[0079] The higher the accuracy rate, the better the classification effect of the model.
[0080] The experiments were conducted on the public datasets MLDoc and CLS. Tables 1 and 2 respectively conducted experimental tests on the effectiveness of each module in the method of the present invention on the MLDoc dataset and the CLS dataset. Among them, AATC - seed dictionary is the accuracy rate after not using the MUSE dictionary; AATC - co - training is the accuracy rate after not using co - training; AATC - self - attention is the accuracy rate after not using self - attention.
[0081] Table 1 is the ablation experiment on the MLDoc dataset
[0082]
[0083] Table 2 is the ablation experiment on the CLS dataset
[0084]
[0085] It can be found from Tables 1 and 2 that after removing the seed dictionary, the accuracy rates decreased significantly by 41.8% and 14.8% respectively on the two datasets, indicating that the seed words play a key role in cross - language transfer. After removing the co - training module, the accuracy rates decreased by 18.8% and 31.1% respectively on the two datasets. This shows that co - training not only strengthens the guiding role of the teacher model on the student model, but also realizes more accurate modeling of the target - language text through iterative optimization. Co - training is especially useful in sparse data because co - training significantly reduces the negative impact of unlabeled data in the target language on the final classification performance. After removing the self - attention mechanism, the accuracy rates decreased by 1.8% and 2.6% respectively on the two datasets. Although the decline is relatively small, combined with the overall results, it can be seen that the core value of the self - attention mechanism lies in enhancing the model's semantic capture ability of the seed words. Its contribution is particularly significant in tasks where the source language and the target language have large differences. The interaction between each module is also an important factor affecting the model performance. For example, the combination of the seed dictionary and the attention mechanism can more accurately identify and weight important words; while co - training ensures that this information can be effectively transmitted between the teacher and student models, jointly promoting the overall performance of the model.
[0086] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above - mentioned embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A zero-shot cross-language text classification method based on adaptive attention transfer, characterized by: The method comprises: Step 1: Extract seed words from the source language based on the word distribution characteristics of the text, model their importance, and generate an importance matrix; use large-scale source language annotation data to train the teacher model, and further emphasize key information through the word probability distribution of seed words during the training process; Step 2: Map the seed words captured by the teacher model to the target language through the cross-bilingual dictionary mapping relationship; Step 3: Use the teacher model to generate pseudo labels for the unlabeled data containing seed words in the target language, use the pseudo labels as the initial training data of the student model, and train the student model; Step 4: The student model further predicts the unlabeled documents in the target language, generates new labeled data and expands the training set. Through iterative optimization, the student model for the target language classification task is finally obtained.
2. The zero-shot cross-language text classification method based on attention adaptive migration according to claim 1, characterized in that: The Step 1 includes: Step 1.1: Given a source language document Its bag-of-words encoding is In the source language document dataset D S The weight matrix W obtained from the training, the classifier predicts the category If the word By weight W kc >0 increases the probability Then the word Select as the seed word; the set of all seed words is The seed word set is larger than the translation budget B, so only the most important seed words are kept through sparse regularization Seed Word The calculation formula is expressed as: in L is the loss function, R sparse (·) is the sparse regularizer, i.e., the L1 norm, λ B Represents a hyperparameter used to control R sparse ;y i represents the true label of the i-th sample, It means to find the minimum weight matrix W of the entire expression. represents the predicted value of the i-th sample, N represents the sum of the loss functions, K represents the total number of categories, |V S | indicates the length of the source language vocabulary; Step 1.2: For the seed word set, use the pre-trained GloVe 300-dimensional word embedding for importance modeling; for the GloVe word embedding vectors of all seed words, calculate the dot product between them to obtain a self-attention score matrix A representing the similarity between words; the self-attention score matrix A is expressed as: A=E·E Τ (2) Among them, the dimension of matrix A is n×n, n is the number of seed words, and E represents G S The matrix composed of the GloVe embeddings of all seed words in; Step 1.3, use the Softmax function to normalize each row of the self-attention score matrix A to obtain the self-attention weight matrix. This process ensures that the sum of the attention weights of each word for other words is equal to 1, thereby obtaining the probability distribution Q ij : Among them, Q ij represents the attention weight between the i-th word and the j-th word, A ij represents the original correlation between the i-th word and the j-th word, A ik represents the original correlation between the i-th word and the k-th word, and n represents the number of seed words.
3. The zero-shot cross-language text classification method based on attention adaptive migration according to claim 1, characterized in that: The Step 2 includes: The MUSE bilingual dictionary is used to obtain the translation of the seed words extracted from the source language into the target language; the weights of the seed words in the source language are transferred to the corresponding vocabulary in the target language to initialize a teacher classifier.
4. The zero-shot cross-language text classification method based on attention adaptive migration according to claim 1, characterized in that: The Step 3 includes: Step 3.1: Use the teacher classifier to classify unlabeled documents in the target language containing seed words Perform pseudo label q j prediction; pseudo label q j The calculation formula is expressed as: in, is an unannotated document in the target language Bag-of-words encoding, w' k Τ yes The kth row of Indicates the probability that the j-th sample corresponds to the k-th category; Step 3.2: Next, train the student model using the context of the seed word For unannotated documents in the target language The student model predicts the probability r j : Among them, f T Represents a prediction function with weight parameter θ, prediction function f T Use a pre-trained Transformer classifier or a pre-trained single-language Bert classifier; Indicates the probability that the j-th sample corresponds to the K-th category; Students are trained with the "distillation" objective: Where H(q,r)=-∑ k q k log2r k is the cross entropy between the teacher's prediction and the student's prediction, R(·) is the sparse regularizer, i.e., the L2 norm, and the hyperparameter λ is used to control R.
5. The zero-shot cross-language text classification method based on attention adaptive migration according to claim 1, characterized in that: The Step 4 includes: Students on D T All unlabeled documents in the target language are predicted to obtain a larger Use D' T 'Train students to convergence.
6. A zero-shot cross-language text classification system based on attention adaptive transfer, characterized in that: The system comprises: a module for executing a zero-shot cross-language text classification method based on attention adaptive migration as claimed in any one of claims 1 to 5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, a zero-sample cross-language text classification method based on attention adaptive migration is implemented as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the zero-shot cross-language text classification method based on attention adaptive migration as described in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the zero-shot cross-language text classification method based on attention adaptive migration as described in any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Cross-language text classification and processing method and system based on deep transfer learning
CN121478978A
A Method and System for Cross-Language Text Classification and Processing Based on Deep Transfer Learning
CN121478978B