Hen-more industry text classification method and system based on prompt learning and adaptive loss weighting

Through the method based on prompt learning and adaptive loss weighting, the problems of language differences and data scarcity in cross-border industrial text classification are solved, efficient classification in Chinese and Vietnamese text classification tasks are achieved, and the adaptability and accuracy of the model are improved.

CN120336534APending Publication Date: 2025-07-18KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510442868.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems of language differences between different languages, imbalance of corpus resources and scarcity of labeled data in cross-border industrial text classification, which leads to poor performance in Chinese and Vietnamese text classification tasks, especially in the case of scarce data and imbalance of languages, which is difficult to effectively classify in the scenarios of small sample learning with scarce data and imbalance of languages.

Method used

Using a method based on prompt learning and adaptive loss weighting, a generalized prompt template is designed by constructing the Hanyue cross-border industrial text classification dataset, a text pair is constructed and a vocabulary mapper is expanded, and a pre-trained language model is optimized with dynamic mixed loss functions to realize cross-language semantic alignment and correlation estimation, alleviating the problem of language imbalance.

Benefits of technology

It significantly improves the performance of the model in cross-border industrial text classification, especially in the scenario of small samples, improves the generalization ability and robustness of the model, and is suitable for scenarios where data scarcity and language imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336534A_ABST
    Figure CN120336534A_ABST
Patent Text Reader

Abstract

The invention relates to a Han-Cross industry text classification method and system based on prompt learning and adaptive loss weighting, and belongs to the technical field of natural language processing. The method comprises the following steps: designing and constructing a universal prompt template; the method comprises the following steps: recombining a Han-Cross cross-border industry text classification data set, namely converting an original single sample into paired samples; in a few-sample and multi-language scene, related vocabularies are adopted as external knowledge resources, and vocabularies most related to the mapping labels are retrieved from the related vocabularies; expanding the vocabulary mapper by introducing synonyms and associated vocabularies; adopting a dynamic mixed loss function and applying the dynamic mixed loss function to a pre-training language model to optimize a few-sample classification task; and classifying Chinese and Vietnamese cross-border industry texts by using the optimized pre-training language model. The method shows a remarkable effect in Chinese and Vietnamese industry text classification tasks, and is particularly suitable for a few-sample learning scene with data scarcity and language imbalance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a Chinese-Vietnamese industrial text classification method and system based on prompt learning and adaptive loss weighting, belonging to the technical field of natural language processing. Background Art

[0002] Text classification is a key task in Natural Language Processing (NLP), and is widely applied in fields such as sentiment analysis, spam detection, opinion analysis, etc.

[0003] The existing research on industrial text classification mainly focuses on the optimization of classification models in a single-language scenario. However, in a cross-border industrial scenario, due to the high cost of annotating industrial text data in Southeast Asian languages (taking Vietnamese as an example) and the scarcity of data resources, traditional deep learning models based on supervised training face significant challenges when processing cross-border industrial data, and the classification performance is limited. The complexity of Vietnamese text annotation mainly manifests in two aspects: on the one hand, annotators need to have the reading and comprehension ability of Vietnamese; on the other hand, annotators also need to master the professional knowledge of the industrial division field. This dual requirement leads to high annotation costs and difficulty in finding annotators who meet the requirements. Although machine translation technology can convert Vietnamese text into Chinese for unified processing, the current translation effect of Southeast Asian languages is still not mature enough, especially in scenarios involving enterprise information and professional terms, where the translation error rate is relatively high and it is difficult to meet the requirements of high-precision industrial text classification. Some existing research has improved the model's understanding ability of samples by introducing external knowledge or label information, but these methods are not only difficult to apply to scenarios lacking label data, but also lack cross-language knowledge transfer, so they are not applicable to cross-border industrial text classification. Therefore, there is an urgent need for a multilingual classification method that can process Chinese and Vietnamese texts under the condition of limited annotated data.

[0004] However, the challenges faced by existing cross-border industrial text classification include:

[0005] Language differences and uneven corpus resources between different languages. In the process of aligning overseas industrial text information to Chinese norms, cross-language knowledge transfer and semantic alignment are required.

[0006] Lack of labeled data. On the side of relatively low-resource languages, industrial text annotation data resources are scarce and the annotation cost is high. Summary of the Invention

[0007] To address the challenges faced in cross - border industrial text classification, the present invention provides a Chinese - Vietnamese industrial text classification method and system based on prompt learning and adaptive loss weighting, which is used to solve the problems of language differences between different languages, data imbalance between languages, and scarcity of labeled data in cross - border industrial text classification. The present invention has shown remarkable effects in Chinese - Vietnamese industrial text classification tasks, and is particularly suitable for few - shot learning scenarios with data scarcity and language imbalance.

[0008] The technical solution of the present invention is: a Chinese - Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting. The steps of the method are as follows:

[0009] Step1. Construction and pre - processing of the Chinese - Vietnamese cross - border industrial text classification dataset: Construct a Chinese - Vietnamese cross - border industrial text classification dataset through manual collection, annotation, and verification, and perform pre - processing.

[0010] Step2. Design of the prompt template: Design and construct a generalized prompt template for constructing input examples for industrial text classification tasks.

[0011] Step3. Construction of text pairs: Re - organize the Chinese - Vietnamese cross - border industrial text classification dataset, that is, convert the original single - sample into paired samples, which is used to transform the original multi - class classification problem into a classification task based on correlation estimation.

[0012] Step4. Construction of the Verbalizer: In few - shot and multilingual scenarios, use relevant vocabulary as an external knowledge resource to retrieve the vocabulary most relevant to the mapped label from the relevant vocabulary; expand the Verbalizer by introducing synonyms and related vocabulary.

[0013] Step5. Adopt a dynamic hybrid loss function and apply it to the pre - trained language model to optimize the few - shot text classification task; the optimized pre - trained language model classifies Chinese and Vietnamese cross - border industrial texts in few - shot learning scenarios with data scarcity and language imbalance.

[0014] Furthermore, Step1 includes:

[0015] Step1.1. The data sources in the Chinese - Vietnamese cross - border industrial text classification dataset cover nine major industry fields, namely "Agriculture, Forestry, Animal Husbandry, and Fishery", "Mining", "Manufacturing", "Transportation, Warehousing, and Postal Services", "Finance", "Real Estate", "Education", "Health and Social Work", and "Culture, Sports, and Entertainment"; the specific data comes from Chinese and Vietnamese industrial news websites, which are all publicly available Internet data; the core fields of the collected industrial news include news id, title, abstract, body content, and industry category.

[0016] Step1.2. Adopt a fine-grained data collection strategy to construct a complete category dataset; in the data preprocessing stage, perform structured processing on the original text, and the structured processing includes information extraction, field standardization, missing data processing, and manual verification; specifically:

[0017] Information extraction: Extract specific information from web pages or original documents, and the specific information includes: title, text, publication time, source;

[0018] Field standardization includes: unifying field names, standardizing time formats, cleaning HTML tags, and handling garbled characters;

[0019] Missing data processing and manual verification include: checking whether fields are missing or filled in incorrectly; if there are no ready-made labels, classify manually or semi-automatically through a fine-grained strategy; use manual means to proofread the field content to ensure accuracy;

[0020] Step1.3. In terms of data division, divide the Chinese-Vietnamese cross-border industrial text classification dataset into a training set and a test set in a ratio of 8:2; the test set undergoes double quality control: first, automatically filter duplicate and low-quality texts, and then perform manual verification and supplementary annotation to ensure the accuracy of the annotation of industry category labels; all text data is stored using UTF-8 encoding.

[0021] Furthermore, the said Step2 includes:

[0022] (1). Design and construct a generalized prompt template for constructing input examples for industrial text classification tasks; the prompt input x p is represented in the form of:

[0023] x p =[text1][SEP] Are the topics relevant?They are[MASK] .text2]

[0024] Among them, the underlined part represents the prompt template, the MASK mark is used for the pre-trained language model to predict the mask position, text1 and text2 respectively represent the two sample text segments in the text pair, and SEP is the separator mark of the pre-trained language model, which is used to separate different parts of the input text. In the prompt template [text1][SEP]Are the topics relevant? They are[MASK].[text2], SEP separates the text segment text1 from the prompt template and text2, helping the pre-trained language model to identify the boundaries of different text segments and enhancing semantic understanding;

[0025] (2) In terms of implementation details, add Segment tags to [text1] and [text2] respectively to distinguish different text segments; do not add Segment tags to the prompt template; fill in the [MASK] token to enable the pre-trained language model to predict whether the text pair belongs to the same category.

[0026] Further, the said Step3 includes:

[0027] Given a cross-border industrial text classification dataset with multiple languages and few samples including Chinese and Vietnamese Use to represent the training dataset, and use to represent the test sample set; the steps to construct the text pairs of the training data are as follows:

[0028] Step3.1. Given the training dataset Each sample in the training dataset is represented as d=(x d , y d ), where x d represents the sample text and y d represents the label corresponding to the sample text;

[0029] Pair the samples in pairs to generate paired samples where d i and d j are both samples in the training dataset, and i, j are sample indices; each pair of samples includes two text segments and mark the correlation label in the form of the same class and different class pairings, so that the pre-trained language model can learn the correlation features between classes from the multi-language input; where respectively represent the text contents of samples d i and d j ;

[0030] When constructing text pairs of the same class, the text pairs of the same class can have both texts in Chinese or both texts in Vietnamese, or one text in Chinese and one text in Vietnamese;

[0031] The constructed training dataset is as follows:

[0032]

[0033] where p(·,·) is a prompt function that fills the text segment into the prompt template to generate the input of the pre-trained language model; the label indicates whether the text pair belongs to the same category. If it belongs to the same category, then y ij = 0; otherwise, y ij = 1; Denotes the union generated by pairwise pairing of all samples in the training dataset ; Refers to the training dataset after being processed by the prompt template, including text pairs and their relevance labels Denotes the label of the i-th and j-th samples

[0034] Step 3.2. In the few-shot setting, select k samples from each category, where k / 2 samples are Chinese texts and the remaining k - k / 2 samples are Vietnamese texts; evenly divide the Chinese and Vietnamese samples to enable the pre-trained language model to more fully learn the category relationships between the two languages and reduce the interference of a single language on the training of the pre-trained language model.

[0035] Furthermore, the said Step 4 includes:

[0036] Step 4.1. Use relevant vocabulary as an external knowledge resource and retrieve the n most relevant vocabulary to the mapped label from it; expand the vocabulary mapper by introducing more synonyms and related vocabulary

[0037] Step 4.2. Use the vocabulary mapper to define the mapping relationship between the set of label words V in the vocabulary and the label space Y, represented by the function f(·): V → Y; based on the pre-trained language model, the vocabulary mapper is used to indirectly infer the category to which the text belongs by filling the vocabulary in the [MASK] position; specifically, for the prompt input x p , the probability distribution of the pre-trained language model over the set of label words is expressed as:

[0038] P(y ∈ Y|x p ) = P M ([MASK] = v ∈ V y |x p )

[0039] where y represents the target category label, x p represents the prompt template, i.e., the content of the filled prompt template; P(y ∈ Y|x p ) represents the probability distribution that the target category label y belongs to the label space Y under the condition of the given prompt input x p ; V y represents the subset of vocabulary mapped to the target category label y, and the mapping relationship is given by f(·); v ∈ V y |x p represents the label word v in the subset of label words V p belonging to the target category label y under the condition of the given x y ; P M ([MASK] = v ∈ V y |xp ) represents the probability that the pre-trained language model predicts that the [MASK] position is filled with the label word v under the condition of the given input text x p of the probability.

[0040] Further, the Step5 includes:

[0041] Step5.1. Reconstruct the few-shot text classification task into a relevance estimation task of text pairs;

[0042] Step5.2. The pre-trained language model adopts the pre-training task of the masked language model. Let P(·; θ) be a masked language model parameterized by the parameter θ, and Pv ocab (·; θ) is the output word probability of the masked language model at the [MASK] position; the optimization objective is defined as follows:

[0043]

[0044] where, represents the input text pair after being processed by the prompt template; represents the relevance label of the text pair; represents the loss function; represents optimizing the parameter θ to minimize the loss; represents the optimized model parameter; is the dynamic mixed loss function of the masked language model, specifically a weighted combination of cross-entropy loss, label smoothing loss and focal loss; φ(·) represents the probability distribution on the label categories. The correct label position of the input sample is set to 1, and the rest are set to 0; f(·) represents a predefined, task-general label mapping vocabulary, which is used to map the output word probability P vocab (·; θ) to the binary classification probability distribution f cls (·; θ). Specifically, f(·) classifies the logits parameters corresponding to the output words {relevant, similar, consistent} at the [MASK] position as the prediction scores of label 1, and at the same time classifies the logits parameters corresponding to the output words {irrelevant, inconsistent, different} at the [MASK] position as the prediction scores of label 0. Logits are the raw unprocessed scores or scores of the output layer of the masked language model.

[0045] Further, the Step5.2 includes:

[0046] Step5.2.1. Use cross-entropy loss to provide basic classification ability; cross-entropy loss is defined as follows:

[0047]

[0048] where y k represents the true label distribution of the k-th class, k represents the class index, and P vocab (·; θ) is the output word probability of the masked language model at the [MASK] position, represents the predicted probability of the model for the k-th class, and N is the number of classes;

[0049] Step5.2.2. Introduce focal loss to assign higher weights to difficult-to-classify samples and enhance the attention of the masked language model to minority classes; focal loss balances the language imbalance problem by assigning higher weights to difficult-to-classify samples. The focal loss is defined as follows:

[0050]

[0051] where p t represents the predicted probability of the masked language model for the target class, ɑ is the sample balance factor, and γ is the focusing factor, which is used to control the degree of attention to difficult samples;

[0052] Step5.2.3. Adopt the label smoothing strategy to smooth the "one-hot" label distribution to weaken the overfitting of the masked language model to a single class; the label smoothing loss reduces the overfitting of the masked language model to noisy labels by converting the true label distribution into a smoothed distribution. The label smoothing loss is defined as follows:

[0053]

[0054] where is the smoothed label distribution, satisfying where β is the smoothing coefficient;

[0055] Step5.2.4. Define the dynamic hybrid loss as follows:

[0056]

[0057] where is the cross-entropy loss, which is used to provide basic classification ability; is the focal loss, which is particularly important for optimizing difficult-to-classify samples; is the label smoothing loss, which is used to alleviate overfitting; σ ce 、σ fl 、σ ls are learnable parameters, representing the uncertainty of the corresponding loss;

[0058] During the training process, the masked language model dynamically adjusts the loss impact through the weight term 1 / 2exp(σ l ) When σ l increases, the corresponding loss weight decays exponentially, indicating that the masked language model believes that the task has a high degree of uncertainty (such as noise interference or difficult optimization), thus reducing its interference; while σ l itself serves as a regularization term, constraining the parameter growth through linear penalty, forcing the model to balance between reducing the loss weight and suppressing the value of σ l , and avoiding a single loss term from dominating the optimization direction due to excessive gradient contribution. Through joint optimization of and σ l , the model autonomously achieves a dynamic balance of "downweighting high-uncertainty losses and focusing on low-uncertainty losses" without manual intervention, and can adapt to task complexity and data distribution changes;

[0059] The dynamic mixed loss combines the advantages of multiple loss functions and solves the bottleneck of traditional combination methods through an adaptive weighting mechanism. Taking the cross-entropy loss as the basic objective for classification, it not only utilizes the classification ability of cross-entropy, but also improves the generalization ability of the model in few-shot and multilingual scenarios through focal loss and label smoothing. After the model is optimized, the prompt model is used as a correlation measurement method in the inference stage. Since the prompt model can input two text samples at a time, it can utilize the cross-sample and cross-language interaction information to more accurately estimate the correlation and achieve efficient and accurate classification.

[0060] After optimizing the Chinese-Vietnamese cross-border industrial text classification model, when there are new query samples with unknown labels that need to be classified, the Chinese-Vietnamese cross-border industrial text classification model converts the classification task into a correlation measurement problem by calculating the correlation scores between the query samples and all training sample pairs, and finally selects the most relevant category as the prediction result based on all classification scores.

[0061] The present invention also provides a Chinese-Vietnamese industrial text classification system based on prompt learning and adaptive loss weighting, and the system includes: a module for executing the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0062] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0063] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0064] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0065] The beneficial effects of the present invention are as follows:

[0066] 1. Aiming at the few-shot classification scenario of cross-border industrial texts, the present invention proposes a prompt learning method based on semantic relatedness measurement, uses the prompt learning template as a tool for cross-lingual semantic alignment, effectively adapts to the cross-border industrial scenario by expanding the text pair construction strategy (supporting both monolingual and cross-lingual inputs) and enriching the semantic coverage of the Verbalizer;

[0067] 2. Aiming at the problem of unbalanced languages in cross-border industrial text classification, the present invention designs a dynamic hybrid loss function, combines cross-entropy, focal loss, and label smoothing loss, and alleviates the language imbalance problem through dynamic weighting, enhancing the generalization ability and robustness of the model;

[0068] 3. By constructing a cross-lingual prompt template, the present invention utilizes the implicit alignment ability of the pre-trained model to capture common features between languages, and innovatively reconstructs the multi-lingual classification task into a text pair correlation estimation task in a cross-semantic space; at the same time, in order to better solve the problem of language imbalance, a dynamic hybrid loss function is designed, integrating cross-entropy, focal loss, and label smoothing loss, and the loss contribution degree is dynamically adjusted through learnable weight parameters, which not only improves the robustness of the model for classifying language-imbalanced samples, but also effectively alleviates the overfitting problem.

[0069] 4. The present invention collects and constructs a Chinese-Vietnamese industrial text classification dataset, and uses a large number of experiments to prove the effectiveness of the method proposed by the present invention, and the effect is significantly better than the existing mainstream methods. Description of the Drawings

[0070] Figure 1 It is the overall framework diagram of the method of the present invention. Detailed Embodiments

[0071] Embodiment 1: In actual situations, due to problems such as language differences between different languages, data imbalance between languages, and scarcity of labeled data in cross-border industrial texts, especially more prominent in low-resource languages, the difficulty of cross-border industrial data classification is increased. Aiming at this problem, the present invention proposes a Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting, which significantly improves the classification performance of the model in cross-border scenarios, such as Figure 1As shown, a Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting. The present invention alleviates the problem of data scarcity based on the prompt learning framework, and utilizes the prior knowledge of the pre-trained model to enhance the few-shot learning ability; secondly, it realizes knowledge transfer and semantic alignment in the semantic space by constructing cross-lingual text pairs; at the same time, it innovatively designs a dynamic hybrid loss function, which multi-objectively optimizes the cross-entropy loss, focal loss, and label smoothing loss, and dynamically adjusts the weights of each loss term based on the uncertainty weighting mechanism - the cross-entropy loss guarantees the basic classification ability; the focal loss strengthens the attention to difficult-to-classify samples; label smoothing effectively suppresses the overfitting risk. This method shows remarkable effects in Chinese and Vietnamese industrial text classification tasks, and is especially suitable for few-shot learning scenarios with data scarcity and language imbalance; the steps of the method are as follows:

[0072] Step1. Construction and preprocessing of the Chinese-Vietnamese cross-border industrial text classification dataset: Since the Chinese standard classification system has insufficient coverage of Vietnamese data and lacks a standardized Vietnamese industrial classification dataset. Therefore, the present invention constructs a Chinese-Vietnamese cross-border industrial text classification dataset through manual collection, annotation, and verification, with a total of more than 130,000 pieces, providing data support for the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0073] The said Step1 includes:

[0074] Step1.1. The data sources in the Chinese-Vietnamese cross-border industrial text classification dataset cover nine major industry fields: "Agriculture, Forestry, Animal Husbandry, and Fishery", "Mining", "Manufacturing", "Transportation, Warehousing, and Postal Services", "Finance", "Real Estate", "Education", "Health and Social Work", "Culture, Sports, and Entertainment"; the specific data comes from Chinese and Vietnamese industrial news websites, all of which are publicly available Internet data (such as platforms like Huanqiu.com and VNA); the core fields of the collected industrial news include news id, title, abstract, body content, and industry category; the detailed field information is shown in Table 1. During the data collection process, a total of 130,539 pieces of industrial news data were obtained, including 113,781 pieces of Chinese data and 16,758 pieces of Vietnamese data, covering multiple industries and multiple languages, with a wide range of sources and representativeness;

[0075] Table 1 Specific field information of the dataset

[0076]

[0077] Step1.2. Aiming at the problem that some Chinese and Vietnamese category data lack ready-made complete labels, the present invention adopts a fine-grained data collection strategy to construct a complete category dataset;

[0078] Taking "agriculture, forestry, animal husbandry, and fishery" as an example, by collecting relevant data on agriculture, forestry, animal husbandry, fishery, etc. respectively, a complete category dataset is constructed. In the data preprocessing stage, the original text is structured, and the standardized fields include 8 core fields such as news id, title, abstract, and body content, and manual verification is carried out to ensure the integrity and accuracy of the fields;

[0079] Table 2 Dataset statistics

[0080]

[0081] In the data preprocessing stage, the original text is structured. The structuring process includes information extraction, field standardization, missing data processing, and manual verification; specifically:

[0082] Information extraction: Extract specific information from web pages or original documents. The specific information includes: title, body, publication time, source;

[0083] Field standardization includes: unifying field naming, standardizing time formats, cleaning HTML tags, and handling garbled characters;

[0084] Missing data processing and manual verification include: checking whether fields are missing or filled in incorrectly; if there are no ready-made labels, manually or semi-automatically classify through a fine-grained strategy; use manual means to proofread field content to ensure accuracy;

[0085] Step1.3. In terms of data division, the Chinese-Vietnamese cross-border industrial text classification dataset is divided into a training set and a test set in a ratio of 8:2; the test set undergoes double quality control: first, automatically filter duplicate and low-quality texts, and then conduct manual verification and supplementary annotation to ensure the accuracy of the annotation of industry category labels; all text data is stored using UTF-8 encoding.

[0086] Step2. Prompt template design: To adapt to the prompt learning mechanism, design and construct a general prompt template for constructing input examples for industrial text classification tasks;

[0087] Furthermore, the said Step2 includes:

[0088] (1). Design and construct a general prompt template for constructing input examples for industrial text classification tasks; the prompt input x p is represented in the form of:

[0089] x p =[text1][SEP] Are the topics relevant?They are[MASK] .[text2]

[0090] Among them, the underlined part represents the prompt template, the MASK token is used for the pre-trained language model to predict the masked position, text1 and text2 respectively represent the two sample text segments in the text pair, and SEP is the separator token of the pre-trained language model, which is used to separate different parts of the input text. In the prompt template [text1][SEP]Are the topics relevant?They are[MASK].[text2], SEP separates the text segment text1 from the prompt template and text2, helping the pre-trained language model to identify the boundaries of different text segments and enhancing semantic understanding;

[0091] (2) In terms of implementation details, Segment tokens are added to [text1] and [text2] respectively to distinguish different text segments, and the maximum length of each text segment is 250; the prompt template is not marked with Segment tokens; by filling in the [MASK] token, the pre-trained language model can be used to predict whether the text pair belongs to the same category. The goal of this template is to guide the model to identify the semantic relationship between text pairs, and its core design lies in filling in the [MASK] token so that the model can predict whether the text pair belongs to the same category. For example, the outputs'relevant, consistent, similar' indicate the same category, and the outputs 'irrelevant, inconsistent, different' indicate different categories. With this structure, the template realizes the unification of the input formats of different language texts, and at the same time effectively enhances the model's ability in cross-language semantic alignment, so that it can more deeply understand and match the cross-language semantic similarity during the learning process.

[0092] Step3. Text pair construction: In order to make full use of the advantages of prompt learning, the Chinese-Vietnamese cross-border industrial text classification dataset is reorganized, that is, the original single sample is converted into a paired sample, so as to convert the original multi-class classification problem into a classification task based on relevance estimation; this data reorganization strategy can significantly improve the performance of the model in the few-shot scenario, especially in the cross-border industrial text classification task, which helps the model to effectively capture the subtle differences between categories in different language environments.

[0093] Furthermore, the said Step3 includes:

[0094] Given a Chinese-Vietnamese multi-language few-shot Chinese-Vietnamese cross-border industrial text classification dataset Use to represent the training dataset, and use to represent the test sample set; the steps to construct the text pairs of the training data are as follows:

[0095] Step3.1. Given the training dataset Each sample in the training dataset is represented as d = (x d , y d ), where x d represents the sample text and y d represents the label corresponding to the sample text;

[0096] Pair the samples in pairs to generate paired samples where d i and d j are both samples in the training dataset, and i and j are sample indices; each pair of samples includes two text segments and mark the correlation labels in the form of same-class and different-class pairings, so that the pre-trained language model can learn the correlation features between classes from multi-language inputs; where respectively represent the text contents of samples d i and d j ;

[0097] When constructing text pairs of the same class, the text pairs of the same class can have both texts in Chinese or both texts in Vietnamese, or one text in Chinese and one text in Vietnamese;

[0098] The constructed training dataset is as follows:

[0099]

[0100] where p(·,·) is a prompting function that fills the text segment into the prompting template to generate the input of the pre-trained language model; the label indicates whether the text pair belongs to the same class. If it belongs to the same class, then y ij = 0; otherwise, y ij = 1; represents the union generated by pairing all samples in the training dataset ; refers to the training dataset processed by the prompting template, including text pairs and their correlation labels, represents the labels of the i-th and j-th samples;

[0101] Step 3.2. In the few-shot setting, to ensure the balance of data among different languages, k samples are selected from each category, where k / 2 samples are Chinese texts and the remaining k - k / 2 samples are Vietnamese texts. This sample division strategy can ensure that the model learns category features evenly in a multilingual environment and reduce the impact of language bias on the model. If only a few samples are randomly selected, it may lead to data bias towards a specific language, thereby affecting the model's general judgment of categories. Therefore, the Chinese and Vietnamese samples are evenly divided to enable the pre-trained language model to more fully learn the category relationships between the two languages and reduce the interference of a single language on the training of the pre-trained language model. In the Chinese-Vietnamese text pairs of the same category, the model can gradually identify the semantic features that remain unchanged across categories; while in the text pairs of different categories, the model can learn the differential features between categories. This method can not only alleviate the interference of language-specific features on model training but also prompt the model to more deeply explore the essential features of categories.

[0102] Step 4. Construction of the Lexical Mapper: In prompt learning, the Verbalizer (Lexical Mapper) is a key component that maps the words predicted by the model to category labels, and the quality of its design directly affects the classification performance. In the few-shot and multilingual scenarios, expanding the Verbalizer is crucial for improving the classification performance. The present invention uses relevant words as external knowledge resources and retrieves 15 words most relevant to the mapping labels therefrom, and expands the Lexical Mapper by introducing synonyms and related words; thereby enriching its semantic coverage. The expanded Verbalizer can more accurately capture semantic differences and effectively improve the performance of the model in multilingual text classification tasks.

[0103] Further, the said Step 4 includes:

[0104] Step 4.1. Use relevant words as external knowledge resources and retrieve 15 words most relevant to the mapping labels therefrom; expand the Lexical Mapper by introducing more synonyms and related words; and enrich its semantic coverage. The expanded Verbalizer (denoted as V Ext ) can more accurately capture semantic differences and effectively improve the performance of the model in multilingual text classification tasks;

[0105] Step 4.2. Define the mapping relationship between the set of label words V in the vocabulary and the label space Y using the Lexical Mapper, represented by the function f(·): V → Y; based on the pre-trained language model, the Lexical Mapper is used to indirectly infer the category to which the text belongs by filling in the words at the [MASK] positions; specifically, for the prompt input x p , the probability distribution of the pre-trained language model over the set of label words is expressed as:

[0106] P(y ∈ Y|x p ) = P M ([MASK] = v ∈ V y |x p )

[0107] where y represents the target class label, and x p represents the prompt template, that is, the content of the filled prompt template; P(y ∈ Y|x p ) represents the probability distribution that the target class label y belongs to the label space Y under the condition of the given prompt input x p ; V y represents the subset of vocabulary words mapped to the target class label y, and the mapping relationship is given by f(·); v ∈ V y |x p represents the label word v in the subset of vocabulary words V that belongs to the target class label y under the condition of the given x p ; P y ([MASK] = v ∈ V M |x y ) represents the probability that the pre-trained language model predicts that the [MASK] position is filled with the label word v under the condition of the given input text x p . p Through the mapping relationship f(·), the text classification task can be transformed into a label word probability prediction problem. For example, expanding the original V = {relevant} to V = {consistent, similar, aligned, pertinent, related,...} enables the Verbalizer to express the category semantics more comprehensively.

[0108] Step 5. In the few-shot classification task of cross-border industry data, the data usually comes from real Internet information and is inevitably affected by label noise, cross-languages, and sample sparsity. A single cross-entropy loss shows certain limitations in such tasks, mainly reflected in its tendency to optimize the dominant classes, which results in poor performance on minority classes or difficult-to-classify samples. In addition, the label noise in Internet data stems from annotation subjectivity and language expression differences, further increasing the learning difficulty of the model. To solve these problems, the present invention adopts a dynamic hybrid loss function (Dynamic Hybrid Loss, DHL) and applies it to the pre-trained language model to optimize the few-shot text classification task; the optimized pre-trained language model classifies Chinese and Vietnamese cross-border industry texts in the few-shot learning scenario of data scarcity and language imbalance.

[0109] Furthermore, the said Step 5 includes:

[0110] Furthermore, the said Step 5 includes:

[0111] Step 5.1. Reconstruct the few-shot text classification task into a text pair relevance estimation task; compared with traditional soft label design methods, it reduces the need for a specific task vocabulary. Through this task reconstruction, our optimization objective is more consistent with the pre-training objective of the pre-trained language model (PLM).

[0112] Step 5.2. The pre-trained language model adopts the pre-training task of the masked language model. Let P(·; θ) be a masked language model parameterized by the parameter θ, and P vocab (·; θ) be the output word probability of the masked language model at the [MASK] position; define the optimization objective as follows:

[0113]

[0114] where, represents the input text pair after being processed by the prompt template; represents the relevance label of the text pair; represents the loss function; represents optimizing the parameter θ to minimize the loss; represents the optimized model parameters; is the dynamic mixed loss function of the masked language model, specifically a weighted combination of cross-entropy loss, label smoothing loss, and focal loss; φ(·) represents the probability distribution over label classes, with the correct label position of the input sample set to 1 and the rest set to 0; f(·) represents a predefined, task-general label mapping vocabulary used to map the output word probability P vocab (·; θ) to the binary classification probability distribution f cls (·; θ). Specifically, f(·) assigns the logits parameters corresponding to the output words {relevant, similar, consistent} at the [MASK] position to the prediction score of label 1, and at the same time assigns the logits parameters corresponding to the output words {irrelevant, inconsistent, different} at the [MASK] position to the prediction score of label 0. Logits are the raw, unprocessed scores or scores of the output layer of the masked language model.

[0115] In the field of multi-task learning (MTL), Kendall et al. (2018) proposed a method of adaptive loss weighting, which dynamically adjusts the weights of loss terms for different tasks based on the uncertainty of the tasks. This research is of great significance in deep learning, especially in scenarios where multiple objectives are optimized simultaneously, significantly improving the robustness and convergence of the model. Its core idea is to automatically allocate learning resources according to the importance and difficulty of the tasks by learning the weighting parameters of the loss terms. This theory has been widely verified in complex task scenarios, such as multi-task image processing, object detection, etc.

[0116] In Kendall's method, the optimization objective is as follows:

[0117]

[0118] where σ i is the uncertainty parameter of task i and also dynamically adjusts the weight of the loss term . In this way, the model can pay more attention to tasks with high uncertainty and greater difficulty, while reducing the interference of unimportant tasks on the training process. In the hybrid loss design of the present invention, we draw on the adaptive weighting method proposed by Kendall et al. to dynamically adjust the weights of three loss functions (Cross-Entropy, Label Smoothing, Focal Loss).

[0119] Furthermore, the said Step5.2 includes:

[0120] Step5.2.1. Use cross-entropy loss to provide basic classification ability; the cross-entropy loss is defined as follows:

[0121]

[0122] where y k represents the true label distribution of the k-th class, k represents the class index, P vocab (·; θ) is the output word probability of the masked language model at the [MASK] position, represents the predicted probability of the model for the k-th class, and N is the number of classes; although the cross-entropy loss is effective in conventional classification tasks, its generalization ability is weak in few-shot and multi-language scenarios, and the experimental results are unstable, with large fluctuations;

[0123] Step5.2.2. To address the issues of data noise and cross - language, Focal Loss is introduced to assign higher weights to difficult - to - classify samples, enhancing the masked language model's attention to minority classes; Lin et al. (2017) have verified the effectiveness of Focal Loss on imbalanced data, providing a theoretical basis for its application in few - shot classification tasks. Focal Loss balances the language imbalance problem by assigning higher weights to difficult - to - classify samples. Focal Loss is defined as follows:

[0124]

[0125] where p t represents the prediction probability of the masked language model for the target class, α is the sample balancing factor, and γ is the focusing factor, which is used to control the degree of attention to difficult samples;

[0126] Step5.2.3. Label Smoothing Loss (Label Smoothing Loss, ): To mitigate the interference of label noise on the model, the present invention simultaneously adopts the Label Smoothing strategy to smooth the "one - hot" label distribution, weakening the model's over - fitting to a single class. Szegedy et al. (2016) pointed out that label smoothing can not only improve the model's robustness but also reduce its sensitivity to incorrect labels. Label Smoothing Loss reduces the masked language model's over - fitting to noisy labels by converting the true label distribution into a smoothed distribution. Label Smoothing Loss is defined as follows:

[0127]

[0128] where is the smoothed label distribution, satisfying where β is the smoothing coefficient;

[0129] Step5.2.4. Define the dynamic hybrid loss as follows:

[0130]

[0131] where is the cross - entropy loss, which is used to provide basic classification ability; is the Focal Loss, which is particularly important for optimizing difficult - to - classify samples; is the Label Smoothing Loss, which is used to mitigate over - fitting; σ ce 、σ fl 、σ ls are learnable parameters, representing the uncertainty of the corresponding loss; during training, the masked language model uses the weight term 1 / 2exp(σ l)Dynamic adjustment of loss impact: When σ l increases, the corresponding loss weight decays exponentially, indicating that the masked language model considers the task to have a high degree of uncertainty (such as noise interference or difficult optimization), thus reducing its interference; while σ l itself acts as a regularization term, constraining the parameter growth through linear penalties, forcing the model to balance between reducing the loss weight and suppressing the σ l value, and preventing a single loss term from dominating the optimization direction due to excessive gradient contribution. By jointly optimizing and σ l , the model autonomously achieves a dynamic balance of "downweighting losses with high uncertainty and focusing on losses with low uncertainty", adapting to task complexity and data distribution changes without manual intervention;

[0132] After optimizing the Chinese-Vietnamese cross-border industrial text classification model, when there are query samples with new unknown labels that need to be classified, the Chinese-Vietnamese cross-border industrial text classification model converts the classification task into a relevance measurement problem by calculating the relevance scores between the query samples and all training sample pairs, and finally selects the most relevant category as the prediction result based on all classification scores.

[0133] The dynamic hybrid loss combines the advantages of multiple loss functions and solves the bottleneck of traditional combination methods through an adaptive weighting mechanism. Taking the cross-entropy loss as the basic objective for classification, it not only utilizes the classification ability of cross-entropy, but also improves the generalization of the model in few-shot and multilingual scenarios through focal loss and label smoothing. After optimizing the model, the prompt model is used as a relevance measurement method during the inference stage. Since the prompt model can input two text samples at a time, it can utilize the cross-sample and cross-language interaction information to more accurately estimate the relevance and achieve efficient and accurate classification.

[0134] The present invention also provides a Chinese-Vietnamese industrial text classification system based on prompt learning and adaptive loss weighting, and the system includes: a module for executing the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0135] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0136] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0137] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting.

[0138] The present invention uses Precision, Recall, and F1-score as evaluation metrics to measure the performance of industrial text classification. The overall evaluation metric is the average Precision, Recall, and F1 value of all categories. By using the method of weighted average, the comprehensive performance of the model on various categories can be measured more comprehensively, ensuring that the evaluation results can reflect the classification effects of all categories. A higher Precision value and Recall value can improve the F1-score, indicating better performance.

[0139] To verify the effectiveness of the model proposed by the present invention, the method of the present invention is compared with the following several baseline methods, specifically as follows:

[0140] mBERT: Use mBERT to generate 768-dimensional sentence embeddings and map them to specific categories through a linear layer to improve the accuracy of text classification.

[0141] PET: Manually construct prompt words using templates, and guide the pre-trained language model to understand the task through fill-in-the-blank phrases to optimize the classification effect.

[0142] PBML: Combine the prompt mechanism with the meta-learning framework, reduce the dependence on large-scale data through label word learning and template learning, and improve the few-shot classification performance.

[0143] MetricPrompt: Model text classification as a relevance estimation task, use the prompt model to measure the matching degree between the input and the label, and optimize the model using cross-entropy loss.

[0144] KPT: Introduce external knowledge to expand the search space of label words, and optimize the selection of label words through the pre-trained model to improve the stability and classification accuracy of prompt-tuning.

[0145] Experiment 1: The method of the present invention is effectively compared with the baseline model on the Chinese-Vietnamese industrial text classification dataset, and the experimental results are shown in Table 3.

[0146] Table 3 shows the experimental results under different few-shot settings

[0147]

[0148]

[0149] As can be seen from the experimental results in Table 3, in the few-shot scenario, the method (Ours) proposed in the present invention outperforms the baseline models in terms of different sample sizes and evaluation metrics, demonstrating strong performance advantages. Specifically, under the three settings of 2-shot, 4-shot, and 8-shot, the average F1 values of the method of the present invention reached 75.11, 80.35, and 83.75 respectively, which were 11.65, 6.11, and 4.01 percentage points higher than those of the relatively better-performing baseline model MetricPrompt. It is worth noting that the performance of the method of the present invention under the 4-shot condition has exceeded the highest level of all baseline models under the 8-shot setting. This result indicates that the present invention can significantly improve the text classification performance and achieve better classification effects in the extremely low-sample scenario.

[0150] In the experimental design, the present invention performed uniform sampling of languages on the training data, dividing the samples into equal amounts of Chinese and Vietnamese, enabling the model to learn the category features of the two languages more evenly. This design effectively reduces the deviation problem caused by a single language in model learning and further improves the classification effect. Especially under the 2-shot setting, the F1 value of the method Ours of the present invention reached 75.11, significantly higher than that of other baseline models, 19.38 percentage points higher than the baseline model mBERT, and significantly better than the best baseline model KPT (+3.39%). This fully demonstrates the key role of the language balance sampling strategy in the extremely low-resource scenario.

[0151] For the Vietnamese task, the present invention uses the pre-trained model mT5-base with a larger number of parameters as the baseline comparison model, conducts supervised training on the complete training set, and conducts a comparative analysis of the training effects of the method proposed in the present invention and the method in the 8-shot scenario. As shown in the experimental data in Table 4, the model based on few-shot training of the method of the present invention shows significant advantages in multi-dimensional evaluation metrics.

[0152] Table 4 shows the performance comparison between the mT5 model and the method proposed in the present invention in the 8-shot setting for the Vietnamese task

[0153]

[0154] Experiment 2: Ablation experiment

[0155] To evaluate the impact of Verbalizer expansion on performance, the present invention conducted ablation experiments under different few-shot settings, and the experimental results are shown in Table 5:

[0156] Table 5 is the ablation experiment of Verbalizer

[0157]

[0158]

[0159] As can be seen from Table 5, the extended Verbalizer is particularly important for improving the model performance. In the 2-shot scenario with extremely low sample sizes, the average F1 value using Verbalizer extension (Ours) is 2.34% higher than that without Verbalizer extension (Ours w / o V Ext ), indicating that with the additional relevant semantic information provided by Verbalizer extension, the model can capture features more comprehensively and further improve the classification performance. In scenarios with 4-shot and above, the model's dependence on the extended information decreases, suggesting that the basic semantic features can support the model's decision-making when the sample size increases.

[0160] Meanwhile, to further verify the effectiveness of the adaptive loss weighting method proposed in this paper , the present invention designed ablation experiments under the 2-shot setting. By randomly sampling the data (80 texts each time, allowing for the dominance of a single language, such as all Vietnamese samples), the performance of this method under extremely few samples and language imbalance conditions was verified. The experimental results are shown in Table 6:

[0161] Table 6 shows the ablation experiments in the 2-shot few-shot scenario

[0162]

[0163] As can be seen from the experimental results in Table 6, in the 2-shot scenario, using the method of the complete loss combination (Ours), the F1 value reaches 65.83, significantly better than the cases where each loss is removed individually. In terms of the F1 value, compared with only removing the cross-entropy loss (Ours w / o ), the F1 value of the complete loss combination is increased by approximately 3.85 percentage points; compared with 64.52 of removing the focal loss (Ours w / o ) and 63.79 of removing the label smoothing loss (Ours w / o ), they are increased by 1.31 and 2.04 percentage points respectively. The same trend is also shown in the Precision and Recall metrics, and the complete loss combination leads other variants with 68.97 and 66.11 respectively. These results verify the effectiveness of the multi-loss collaborative optimization strategy.

[0164] It is worth noting that removing any single loss term will lead to a performance decline. Among them, the absence of cross-entropy loss results in the largest decrease in the F1 value (a relative decrease of 6.21%), which indicates that cross-entropy loss plays a key role in improving the basic classification performance of the model. In contrast, removing focal loss and label smoothing loss respectively leads to a 2.03% and 3.20% relative decrease in the F1 value compared to the complete loss combination, which indicates that the two have a complementary effect in alleviating sample imbalance and preventing overfitting.

[0165] The above results show that the adaptive hybrid loss method can effectively adapt to the characteristics of the multi-language few-shot scenario. By dynamically adjusting the weights of different loss terms during training, it significantly alleviates the negative impact of the uneven language distribution among multiple languages on the model performance. This dynamic weighting strategy makes the model more balanced during the optimization process, can make more full use of the shared information between languages, and at the same time avoid biases caused by uneven language distribution, thereby improving the overall performance of the model in the industrial text classification task and achieving higher accuracy and robustness.

[0166] The specific implementation manners of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above implementation manners. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting, characterized in that: The method includes: Step1. Construction and preprocessing of the Chinese-Vietnamese cross-border industrial text classification dataset: Construct a Chinese-Vietnamese cross-border industrial text classification dataset through manual collection, annotation, and verification, and perform preprocessing; Step2. Prompt template design: Design and construct a generalized prompt template for constructing input examples for industrial text classification tasks; Step3. Text pair construction: Reorganize the Chinese-Vietnamese cross-border industrial text classification dataset, that is, convert the original single sample into a paired sample, for converting the original multi-class classification problem into a classification task based on relevance estimation; Step4. Lexical mapper construction: In the few-shot and multilingual scenarios, use relevant vocabulary as an external knowledge resource to retrieve the vocabulary most relevant to the mapping label from the relevant vocabulary; expand the lexical mapper by introducing synonyms and related vocabulary; Step5. Adopt a dynamic hybrid loss function and apply it to the pre-trained language model to optimize the few-shot text classification task; the optimized pre-trained language model classifies Chinese and Vietnamese cross-border industrial texts in the few-shot learning scenarios of scarce data and language imbalance.

2. The Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to claim 1, wherein: The said Step1 includes: Step1.

1. The data sources in the Chinese-Vietnamese cross-border industrial text classification dataset cover nine major industry fields including "Agriculture, Forestry, Animal Husbandry, and Fishery", "Mining", "Manufacturing", "Transportation, Warehousing, and Post", "Finance", "Real Estate", "Education", "Health and Social Work", and "Culture, Sports, and Entertainment"; the specific data comes from Chinese and Vietnamese industrial news websites, all of which are publicly available Internet data; the core fields of the collected industrial news include news id, title, abstract, body content, and industry category; Step1.

2. Adopt a fine-grained data collection strategy to construct a complete category dataset; in the data preprocessing stage, perform structured processing on the original text, and the structured processing includes information extraction, field standardization, missing data processing, and manual verification; specifically: Information extraction: Extract specific information from web pages or original documents, and the specific information includes: title, body, release time, source; Field standardization includes: unified field naming, time format standardization, cleaning HTML tags, and garbled characters; Missing data processing and manual verification include: checking whether the fields are missing and whether there are filling errors; if there are no ready-made labels, classify manually or semi-automatically through a fine-grained strategy; use manual means to proofread the field content to ensure accuracy; Step1.

3. In terms of data division, divide the Chinese-Vietnamese cross-border industrial text classification dataset into a training set and a test set according to a ratio of 8:2; the test set undergoes double quality control: first, automatically filter duplicate and low-quality texts, and then perform manual verification and supplementary annotation to ensure the annotation accuracy of industry category labels; all text data is stored using UTF-8 encoding.

3. The Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to claim 1, characterized in that: The said Step2 includes: (1) Design and construct a generalized prompt template for constructing input examples for industrial text classification tasks; the prompt input x p is represented in the form of: x p = [text1][SEP] Are the topics relevant?They are[MASK] .[text2] Among them, the underlined part represents the prompt template, the MASK token is used for the pre-trained language model to predict the masked position, text1 and text2 respectively represent the two sample text segments in the text pair, and SEP is the separator token of the pre-trained language model, which is used to separate different parts of the input text. In the prompt template [text1][SEP]Are the topics relevant? They are[MASK].[text2], SEP separates the text segment text1 from the prompt template and text2, helping the pre-trained language model to identify the boundaries of different text segments and enhancing semantic understanding; (2) In terms of implementation details, Segment tokens are added to [text1] and [text2] respectively to distinguish different text segments; the prompt template is not marked with Segment tokens; by filling in the [MASK] token, the pre-trained language model can predict whether the text pair belongs to the same category.

4. The Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to claim 1, wherein: The said Step3 includes: Given a cross - border industrial text classification dataset of Chinese - Vietnamese with multilingual and few - shot samples Use to represent the training dataset, and use to represent the test sample set; the steps to construct text pairs of the training data are as follows: Step 3.1: Given a training dataset Each sample in the training dataset is represented as d = (x d , y d ), where x d represents the sample text and y d represents the label corresponding to the sample text; Pair the samples in pairs to generate paired samples where d i , d j are both samples in the training dataset, and i and j are sample indices; each pair of samples includes two text segments and mark the correlation labels in the form of same-class and different-class pairings, enabling the pre-trained language model to learn the correlation features between classes from multi-language inputs; where respectively represent the text contents of samples d i , d j ; When constructing text pairs of the same category, the text pairs of the same category can have both texts in Chinese or both texts in Vietnamese, or one text in Chinese and one text in Vietnamese; The constructed training dataset is as follows: where p(·, ·) is a prompting function that fills the text segment into the prompting template to generate the input of the pre-trained language model; the label indicates whether the text pair belongs to the same category. If it belongs to the same category, then y ij = 0; otherwise, represents the union generated by pairwise pairing all samples in the training dataset ; refers to the training dataset processed by the prompting template, which contains text pairs and their relevance labels, indicating the label of the i-th and j-th samples; Step3.

2. In the few-shot setting, k samples are selected from each category, where k / 2 samples are Chinese texts and the remaining k - k / 2 samples are Vietnamese texts; the Chinese and Vietnamese samples are evenly divided to enable the pre-trained language model to more fully learn the category relationship between the two languages and reduce the interference of a single language on the training of the pre-trained language model.

5. The Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to claim 1, wherein: The said Step4 includes: Step4.

1. Use relevant vocabulary as an external knowledge resource to retrieve the n most relevant vocabulary to the mapped label from it; expand the vocabulary mapper by introducing more synonyms and related vocabulary; Step4.

2. Define the mapping relationship between the set of labeled words V in the vocabulary and the label space Y using a lexical mapper, represented by the function f(·): V → Y; based on the pre-trained language model, the lexical mapper is used to indirectly infer the category to which the text belongs by filling in the vocabulary at the [MASK] position; specifically, for the prompted input x after being processed by the prompt template p , the probability distribution of the pre-trained language model over the set of labeled words is expressed as: P(y ∈ Y|x p ) = P M ([MASK] = v ∈ V y |x p ) Among them, y represents the target class label, and x p represents the prompt template, that is, the content of the filled prompt template; P(y ∈ Y|x p ) represents the probability distribution that the target class label y belongs to the label space Y under the condition of the given prompt input x p ; V y represents the subset of vocabulary mapped to the target class label y, and the mapping relationship is given by f(·); v ∈ V y |x p represents the subset of label words V p of the target class label y under the condition of the given x y ; P M ([MASK] = v ∈ V y |x p ) represents the probability that the pre-trained language model predicts that the [MASK] position is filled with the label word v under the condition of the given input text x p .

6. The Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to claim 1, wherein: The said Step5 includes: Step5.

1. Reconstruct the few-shot text classification task into a relevance estimation task of text pairs; Step5.

2. The pre-trained language model adopts the pre-training task of the masked language model. Let P(·; θ) be a masked language model parameterized by the parameter θ, and P vocab (·; θ) be the output word probability of the masked language model at the [MASK] position; the optimization objective is defined as follows: Among them, represents the input text pair after being processed by the prompt template; represents the relevance label of the text pair; represents the loss function; represents optimizing the parameter θ to minimize the loss; represents the optimized model parameter; is the dynamic mixed loss function of the masked language model, specifically, it is a weighted combination of cross-entropy loss, label smoothing loss, and focal loss; φ(·) represents the probability distribution over label classes, with the correct label position of the input sample set to 1 and the rest set to 0; f(·) represents a predefined, task-general label mapping vocabulary used to map the output word probability P vocab (·; θ) to the binary classification probability distribution f cls (·; θ), specifically, f(·) assigns the logits parameter corresponding to the output word at the [MASK] position to the prediction score of label 1, and at the same time assigns the logits parameter corresponding to the output word at the [MASK] position to the prediction score of label 0. Logits are the raw, unprocessed scores or scores of the output layer of the masked language model.

7. The method for classifying Chinese-Vietnamese industrial texts based on prompt learning and adaptive loss weighting according to claim 6, characterized in that: The said Step5.2 includes: Step 5.2.1: Use cross-entropy loss to provide basic classification ability; the cross-entropy loss is defined as follows: Among them, y k represents the true label distribution of the k-th category, k represents the category index, and P vocab (·; θ) is the output word probability of the masked language model at the [MASK] position, represents the predicted probability of the model for the k-th category, and N is the number of categories; Step5.2.

2. Introduce focal loss to assign higher weights to difficult-to-classify samples and enhance the masked language model's attention to minority classes. Focal loss balances the language imbalance problem by assigning higher weights to difficult-to-classify samples. The focal loss is defined as follows: Among them, p t represents the prediction probability of the masked language model for the target category, α is the sample balancing factor, and γ is the focusing factor, which is used to control the degree of attention to difficult samples; Step5.2.

3. Adopt a label smoothing strategy to smooth the "one-hot" label distribution to weaken the overfitting of the masked language model to a single category; the label smoothing loss reduces the overfitting of the masked language model to noisy labels by converting the true label distribution into a smoothed distribution. The label smoothing loss is defined as follows: Among them, is the smoothed label distribution, satisfying where β is the smoothing coefficient; Step5.2.

4. Define the dynamic mixed loss as follows: Among them, is the cross-entropy loss, which is used to provide basic classification ability; is the focal loss, which is particularly important for optimizing difficult-to-classify samples; is the label smoothing loss, which is used to alleviate overfitting; σ ce σ, fl σ, ls are learnable parameters, representing the uncertainty of the corresponding loss; After optimizing the Chinese-Vietnamese cross-border industrial text classification model, when there are new query samples with unknown labels that need to be classified, the Chinese-Vietnamese cross-border industrial text classification model converts the classification task into a relevance measurement problem by calculating the relevance scores between the query samples and all training sample pairs, and finally selects the most relevant category as the prediction result based on all classification scores.

8. A Chinese-Vietnamese industrial text classification system based on prompt learning and adaptive loss weighting, characterized in that, The said system includes: a module for executing the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting as described in any one of claims 1 to 7.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the Chinese-Vietnamese industrial text classification method based on prompt learning and adaptive loss weighting according to any one of claims 1 to 7.

Citation Information

Cited By

  • MAML-based cross-border public opinion analysis model, method and equipment and storage medium

    CN121031609A

  • MAML-based cross-border public opinion analysis model, method, device and storage medium

    CN121031609B

  • Large language model interpretation generation method and system based on fine tuning and joint tasks

    CN121257553A

  • Fine-tuning and joint task-based large language model explanation generation method and system

    CN121257553B

  • Few-sample hierarchical text classification method based on pre-training language model

    CN121501999A