Large model text classification method and system based on alignment strategy
By adopting the large-model text classification method with the alignment strategy in the text classification method, the label propagation and alignment enhancement is used to use language prompts and text semantic graphs for label propagation and alignment enhancement, the problem of low text classification accuracy when samples are scarce is solved, and high accuracy classification is achieved in the case of scarce samples.
Patent Information
- Application Number
- CN202510060342.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The existing text classification methods have severely affected the accuracy of model classification in the case of scarcity of samples and lack effective solutions.
The large-model text classification method based on alignment strategy is adopted, and the text classification task is converted into text completion problems in natural language prompts by constructing language prompts, and the label propagation and alignment enhancement is used to improve the classification accuracy of the model.
In the case of scarcity of samples, the accuracy of model classification is improved, and through the post-training alignment strategy, it is possible to make full use of label knowledge without introducing additional training costs to maintain the robustness of the pre-trained language model.
Smart Images

Figure CN119474390B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a large-model text classification method and system based on an alignment strategy. Background Art
[0002] With the rapid development of Internet technology, user-generated content on various social media has experienced explosive growth. In-depth mining and analysis of these fragmented contents can provide a data basis for important applications such as recommendation systems and public opinion analysis.
[0003] Most existing text classification methods adopt the paradigm of "training from scratch" based on given data with clear labels, or the paradigm of fine-tuning based on pre-trained language models for specific tasks. However, both methods rely on a large number of labeled samples to optimize model weights, which will seriously affect the accuracy of model classification when samples are scarce.
[0004] There is currently no effective solution to the problem in related technologies that the accuracy of model classification is affected when samples are scarce. Summary of the invention
[0005] Based on this, it is necessary to provide a large-model text classification method and system based on an alignment strategy that can improve the model classification accuracy when samples are scarce, in response to the above technical problems.
[0006] In the first aspect, a large model text classification method based on an alignment strategy is provided in this embodiment, including:
[0007] Construct language prompts based on the text to be classified;
[0008] Obtaining an output vector based on the pre-trained language model and the language prompt;
[0009] Determine a probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; the probability distribution matrix includes predicted classification labels;
[0010] Label propagation is performed based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain a text classification result of the text to be classified.
[0011] In some of the embodiments, obtaining the output vector based on the pre-trained language model and the language prompt includes:
[0012] Presetting candidate words of the pre-trained language model;
[0013] The language prompt is input into the pre-trained language model to obtain an output vector of the pre-trained language model for each of the candidate words.
[0014] In some of the embodiments, it also includes:
[0015] Calculate the loss function value based on the preset loss function and the true classification label of the training sample;
[0016] The output vector of the pre-trained language model is adjusted with minimizing the loss function value as an optimization goal.
[0017] In some embodiments, determining the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector includes:
[0018] Determine a probability score and a predicted classification label for each answer in the candidate word according to the output vector and the candidate word;
[0019] The probability distribution matrix is determined based on the probability scores and the predicted classification labels.
[0020] In some of the embodiments, it also includes:
[0021] Adjusting the predicted classification labels in the probability distribution matrix based on a preset threshold and the probability scores in the probability distribution matrix;
[0022] According to the true classification labels of the training samples, the probability distribution matrix is aligned and enhanced.
[0023] In some of the embodiments, it also includes:
[0024] Determine text nodes and word nodes according to the text to be classified;
[0025] Based on the sliding window, determining the number of occurrences of a word and the number of co-occurrences of two words in the text to be classified;
[0026] Calculate the connection weight between the word nodes and the edge weight between the text node and the word node according to the word occurrence count and the two word co-occurrence count;
[0027] The text semantic graph is constructed based on the connection weights and the edge weights.
[0028] In some embodiments, the label propagation is performed based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified, including:
[0029] Based on the connection weights and edge weights in the text semantic graph, label propagation is performed in combination with the probability distribution matrix after alignment enhancement;
[0030] The text classification result of the text to be classified is determined according to the updated predicted classification labels of the text nodes in the text semantic graph.
[0031] In the second aspect, a large model text classification system based on an alignment strategy is provided in this embodiment, including:
[0032] A prompt building module, used to build language prompts based on the text to be classified;
[0033] A model output module, used to obtain an output vector based on the pre-trained language model and the language prompt;
[0034] A probability distribution calculation module, used to determine the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; the probability distribution matrix includes predicted classification labels;
[0035] The text classification reasoning module is used to perform label propagation based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified.
[0036] In a third aspect, a computer device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the large model text classification method based on the alignment strategy described in the first aspect is implemented.
[0037] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, the large model text classification method based on the alignment strategy described in the first aspect is implemented.
[0038] Compared with the related art, the large model text classification method and system based on alignment strategy provided in this embodiment constructs language prompts based on the text to be classified; obtains an output vector based on the pre-trained language model and the language prompts; determines the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; the probability distribution matrix includes predicted classification labels; and performs label propagation based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified. Through this embodiment, language prompts are constructed based on the text to be classified, the text classification task is converted into a task-oriented text completion problem in natural language prompts, and the text semantic graph is used for label propagation, and the predicted text classification results of the pre-trained language model are aligned and enhanced, which can improve the accuracy of model classification when samples are scarce.
[0039] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0041] Figure 1 is a hardware structure block diagram of a terminal of a large model text classification method based on an alignment strategy in an embodiment;
[0042] Figure 2 is a flowchart of a large model text classification method based on an alignment strategy in one embodiment;
[0043] Figure 3 is a schematic diagram of a process of constructing a text semantic graph in an embodiment;
[0044] Figure 4 is a flowchart of a large model text classification method based on an alignment strategy in another embodiment;
[0045] Figure 5 The structure diagram of a large model text classification system based on an alignment strategy in one embodiment.
[0046] In the figure: 102, processor; 104, memory; 106, transmission device; 108, input and output device; 10, prompt construction module; 20, model output module; 30, probability distribution calculation module; 40, text classification reasoning module. DETAILED DESCRIPTION
[0047] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0048] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0049] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 : is a hardware structure block diagram of a terminal of the large model text classification method based on the alignment strategy of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0050] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the large model text classification method based on the alignment strategy in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0051] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0052] With the rapid development of Internet technology, user-generated content on various social media has experienced explosive growth. In-depth mining and analysis of these fragmented contents can provide a data basis for important applications such as recommendation systems and public opinion analysis.
[0053] Most existing text classification methods adopt the paradigm of "training from scratch" based on given data with clear labels, or the paradigm of fine-tuning based on a pre-trained language model for a specific task. However, the weight optimization of the former method only relies on supervised training data. When labels are scarce, this method has generalization problems. The latter method is task-specific and stacks classification heads on pre-trained language models. Although methods under this paradigm benefit from pre-trained grammatical and semantic knowledge, the optimization of auxiliary layers still relies on abundant high-quality tags, and fine-tuning pre-trained language models has been shown to distort high-quality pre-trained features and impair the robustness of the model. Both methods rely on a large number of labeled samples for model weight optimization, which will seriously affect the accuracy of model classification when samples are scarce.
[0054] In this embodiment, a large model text classification method based on an alignment strategy is provided. Figure 2 is a flowchart of the large model text classification method based on the alignment strategy in this embodiment. Figure 2As shown, the method comprises the following steps:
[0055] Step S201, constructing a language prompt based on the text to be classified.
[0056] Specifically, based on prompt learning, the text classification task is converted into a task-oriented text completion problem in natural language prompts. For a given text to be classified, add the [MASK] tag before the text to be classified to construct a language prompt containing the mask token. The following is a given text to be classified , construct the language prompt corresponding to the mask token One way:
[0057] ;
[0058] Among them, i represents the i-th text to be classified.
[0059] For example, news texts to be classified can be classified into specific news types, such as entertainment news, sports news, current affairs news, and economic news, through text classification. By constructing language prompts such as "What type of news is the following content?", the text classification task can be converted into a question-and-answer type of text completion problem.
[0060] Step S202, obtaining an output vector based on the pre-trained language model and the language prompt.
[0061] Specifically, for the above language prompts, the pre-trained language model continues to write based on the language prompts, so as to use the pre-trained language model as a classifier under the small sample task of text completion problem. The text classification prompts of the pre-trained language model are pre-set to allow the answer vocabulary, and each word in the vocabulary is used as a candidate word for the pre-trained language model. Exemplarily, in the text classification of news, the candidate words can be different language forms of different news types such as entertainment news, sports news, current affairs news and economic news. In the sentiment analysis of the text, the candidate words can be different sentiment words such as happy, happy, good reviews, satisfied, bad reviews, sad, hate, general, etc. Among them, the pre-trained language model includes but is not limited to open source models such as LLaMA (LargeLanguage Model Meta AI, large language model meta AI), GPT (Generative Pre-trainedTransformer, generative pre-trained transformer).
[0062] Based on the pre-trained language model, we can get the output vector of the language prompt for each candidate word, which contains the vector representation of the model's completion content.
[0063] Step S203, determining the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; the probability distribution matrix includes the predicted classification label.
[0064] Specifically, the output vector of each candidate word of the pre-trained language model is converted into the conditional probability distribution of all candidate words through the softmax (normalized exponential) function. This conditional probability distribution can form a probability distribution matrix, which includes the probability score of each answer in the candidate word among all candidate words for a text to be classified, and the corresponding predicted classification label determined according to the probability score. For example, in the text classification of news, the classification labels include but are not limited to news types such as entertainment, sports, current affairs and economy, and in the sentiment analysis of text, the classification labels include but are not limited to positive, negative, neutral, etc.
[0065] Step S204, label propagation is performed based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain a text classification result of the text to be classified.
[0066] Specifically, a text semantic graph is constructed based on the text to be classified. The graph contains word nodes and text nodes, and both local semantic information and global semantic information are considered. The connection between word nodes and specific text nodes provides local semantic context information, and different texts are connected through word nodes to provide global semantic context information. A post-training alignment strategy is adopted. An alignment component based on the text semantic graph is introduced after the pre-trained language model is trained. This can make full use of label knowledge without introducing additional trainable modules. In addition, the real classification label is introduced on the basis of the above predicted classification label to obtain the probability distribution matrix after alignment enhancement. Label propagation is performed based on the text semantic graph and the probability distribution matrix after alignment enhancement to obtain the predicted classification label of the text to be classified as the text classification result.
[0067] Through the above steps, language prompts are constructed based on the text to be classified, the text classification task is converted into a task-oriented text completion problem in natural language prompts, the pre-trained language model is used as a classifier under a small sample task, and the semantic relationship in the text semantic graph is further used for label propagation. The post-training alignment strategy is adopted to align and enhance the predicted text classification results of the pre-trained language model to alleviate the dependence on sample labels. Compared with the method in the prior art that relies on a large number of labeled samples to optimize the model weights, this embodiment can improve the accuracy of model classification when samples are scarce. In addition, the post-training alignment strategy can fully utilize label knowledge without introducing additional training costs, and will not affect the pre-training characteristics and robustness of the pre-trained language model.
[0068] In some embodiments, the output vector is obtained based on the pre-trained language model and the language prompt in step S202, including:
[0069] Pre-set candidate words for the pre-trained language model; input the language prompt into the pre-trained language model to obtain an output vector of each candidate word of the pre-trained language model.
[0070] Specifically, a vocabulary list of allowed answers for the text classification prompt of the pre-trained language model is pre-set, and each word in the vocabulary list is used as a candidate word of the pre-trained language model, and the candidate words are respectively mapped to different classification labels. For example, in the text classification of news, the candidate words can be different language forms of different news types such as entertainment news, sports news, current affairs news, and economic news. In the sentiment analysis of the text, the candidate words can be different sentiment words such as happy, glad, good review, satisfied, bad review, sad, hate, and general.
[0071] The following are input language tips , the output vector of the pre-trained language model M for each candidate word v A representation of:
[0072] ;
[0073] Among them, i represents the i-th text to be classified.
[0074] The output space of the pre-trained language model M is a subset of the vocabulary, which can be mapped to classification labels separately.
[0075] By using the pre-trained language model as a classifier in this embodiment, the output vector of the model for each candidate word is obtained, and in subsequent steps, the classification label can be mapped according to the output vector.
[0076] In some of the embodiments, the following steps are also included:
[0077] Based on the preset loss function and the true classification labels of the training samples, the loss function value is calculated; with minimizing the loss function value as the optimization goal, the output vector of the pre-trained language model is adjusted.
[0078] Specifically, for a given pre-trained language model and text prompt, the output vector of the pre-trained language model is adjusted using the true classification label of the training sample, without introducing additional training tasks. The training sample includes a text prompt, an answer, and the corresponding true classification label. The text prompt of the training sample is input into the pre-trained language model to obtain the probability score and predicted classification label of each answer in the candidate word, and the loss function value is calculated. The probability score of the correct answer token y+ is maximized and the probability score of the incorrect answer token y- is penalized by minimizing the loss function value, and the output vector of the pre-trained language model is adjusted. Specifically, a logit (logistic regression) loss function that can suppress the corresponding incorrect answer token can be used. The following is an expression of the loss function L:
[0079] ;
[0080] in, represents the i-th text prompt, n represents a total of n text prompts, Indicates text prompt The probability score of the correct answer token y+, Indicates text prompt The probability score of the incorrect answer token y-; BCE stands for Binary Cross Entropy.
[0081] By constructing the loss function in this embodiment, the output vector of the pre-trained language model can be adjusted without introducing additional training tasks.
[0082] In some embodiments, the above step S203 determines the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector, including the following steps:
[0083] Based on the output vector and the candidate words, the probability score and predicted classification label of each answer in the candidate words are determined; based on the probability score and predicted classification label, the probability distribution matrix is determined.
[0084] Specifically, the output vector of each candidate word of the pre-trained language model is converted into the conditional probability distribution of all candidate words through the softmax function. This conditional probability distribution can form a probability distribution matrix, which includes the probability scores of each answer of the model among all candidate words for each language prompt. The corresponding predicted classification label can be further determined based on the probability score.
[0085] The following are tips for a given language , the probability score of answer y A way to calculate:
[0086] ;
[0087] Among them, y represents the answer of the model; v represents the candidate word, there are V candidate words in total; exp represents the exponential function.
[0088] In some embodiments, based on a preset threshold and the probability score in the probability distribution matrix, the predicted classification label in the probability distribution matrix is adjusted; and according to the true classification label of the training sample, the probability distribution matrix is aligned and enhanced.
[0089] After obtaining the probability distribution matrix, the predicted classification labels therein are processed. Specifically, according to the probability score and the preset threshold, the output vectors with probability scores greater than or equal to the preset threshold are assigned unique hot labels to adjust the predicted classification labels in the probability distribution matrix. Exemplarily, the preset threshold is 0.6. For output vectors with probability scores greater than or equal to 0.6, the probability score is adjusted to 1, and the predicted classification labels are adjusted accordingly. According to the adjusted predicted classification labels and the true classification labels of the training samples, a probability distribution matrix after alignment enhancement processing is constructed, wherein the probability score corresponding to the true classification label is 1. In some embodiments, it is also possible to first construct a probability distribution matrix after alignment enhancement processing according to the predicted classification labels and the true classification labels of the training samples, and then adjust the predicted classification labels in the probability distribution matrix. The specific order is not limited.
[0090] Through this embodiment, the probability distribution matrix is determined according to the output vector of the model and the alignment enhancement processing is performed, so that the alignment strategy after training can be implemented in the subsequent steps.
[0091] In some of these embodiments, Figure 3 is a schematic diagram of the process of constructing a text semantic graph in this embodiment, such as Figure 3 As shown in Figure 1, constructing a text semantic graph includes the following steps:
[0092] Step S301, determining text nodes and word nodes according to the text to be classified.
[0093] Step S302, based on the sliding window, determining the number of occurrences of a word and the number of co-occurrences of two words in the text to be classified.
[0094] Step S303, calculating the connection weights between word nodes and the edge weights between text nodes and word nodes according to the word occurrence times and the two-word co-occurrence times.
[0095] Step S304: construct a text semantic graph based on the connection weights and edge weights.
[0096] Specifically, the text to be classified and the individual words in it are used as text nodes and word nodes of the text semantic graph. A sliding window of length L is used to cut the text to be classified into small segments, and the number of times a single word appears in the window and the number of times two words co-occur is counted, and the word frequency and the co-occurrence frequency of two words are further calculated. The following is a calculation formula for the word frequency P (a) and the co-occurrence frequency of two words P (a, b):
[0097] ;
[0098] ;
[0099] Among them, a and b represent words, count represents the number of occurrences of the word; M represents the number of sliding windows.
[0100] The connection weight PMI (Pointwise Mutual Information) between word nodes is calculated based on the word frequency and the two co-occurrence frequencies. The connection weight PMI (a, b) is calculated as follows:
[0101] ;
[0102] Among them, P(a,b) represents the co-occurrence frequency of two words, P(a) and P(b) represent the word frequencies of the two co-occurring words respectively; log represents the logarithmic function.
[0103] The edge weight between the text node and the word node is determined by using the TF-IDF (Term Frequency-Inverse Document Frequency) of the words in the text. The edge weight TF-IDF (w, d) is calculated as follows:
[0104] ;
[0105] ;
[0106] ;
[0107] Among them, d represents the document, w represents the word, and N represents the total number of documents; represents the number of occurrences of word w in document d, represents the total number of occurrences of all words in document d; N(w) represents the number of documents containing word w; TF(d,w) represents the word frequency of word w in document d, and IDF(w) represents the inverse document frequency of word w.
[0108] Based on the connection weight and edge weight, a text semantic graph is constructed. The purpose of TF-IDF is to evaluate the importance of a word to a specific document in a document collection. The higher the TF-IDF value, the more important the word is to the document (text node).
[0109] By constructing a text semantic graph in this embodiment, local semantic information and global semantic information are considered at the same time. The connection between word nodes and specific text nodes provides local semantic context information, and different texts are connected through word nodes to provide global semantic context information, so as to enhance the reasoning results of the pre-trained language model by using the semantic relationship in the text semantic graph in the future.
[0110] In some embodiments, the step S204 performs label propagation based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain a text classification result of the text to be classified, including the following steps:
[0111] Based on the connection weights and edge weights in the text semantic graph, label propagation is performed in combination with the probability distribution matrix after alignment enhancement; the text classification result of the text to be classified is determined according to the predicted classification label after the text node in the text semantic graph.
[0112] Specifically, the following are the text classification results A way of expressing:
[0113] ;
[0114] Among them, A represents the text semantic graph; k represents the number of hops of label propagation; H represents the probability distribution matrix after alignment enhancement.
[0115] Define the rules for label propagation, run k-hop label propagation in the text semantic graph, start from the text node, and propagate the probability distribution along the edge of the graph to the adjacent word nodes according to the label propagation rules. For each word node, update its own probability distribution according to the probability distribution of the text node it is connected to. Then, propagate the updated word node probability distribution back to the text node, or propagate to other text nodes that are not directly connected (through the word node as an intermediary). Repeat this process k times, and each propagation updates the probability distribution of the node according to the current label. After k-hop propagation, for each text node, it will have an updated probability distribution based on its neighbors (including directly and indirectly connected word nodes and other text nodes). You can choose to aggregate these probabilities (for example, by taking the average, maximum value, or weighted average) to obtain the final corresponding predicted classification label. According to the aggregated classification label probability, select the category with the highest probability as the final classification result of the text.
[0116] In this embodiment, the semantic relationship in the text semantic graph is used for label propagation, and a post-training alignment strategy is adopted to align and enhance the predicted text classification results of the pre-trained language model, thereby alleviating the dependence on sample labels. This makes it possible to fully utilize label knowledge without introducing additional training costs and improve the accuracy of model classification.
[0117] The present embodiment is described and illustrated below through preferred embodiments.
[0118] Figure 4 is a flowchart of the large model text classification method based on the alignment strategy of this embodiment. Figure 4 As shown, the method comprises the following steps:
[0119] Step S401: construct a language prompt based on the text to be classified.
[0120] Step S402: input the language prompt into the pre-trained language model to obtain an output vector of each candidate word of the pre-trained language model.
[0121] Step S403, based on the preset loss function and the real classification label of the training sample, the loss function value is calculated; and the output vector of the pre-trained language model is adjusted with minimizing the loss function value as the optimization goal.
[0122] Step S404, based on the output vector and the candidate words, determine the probability score and predicted classification label of each answer in the candidate words, and construct a probability distribution matrix.
[0123] Step S405, based on a preset threshold and the probability score in the probability distribution matrix, the predicted classification label in the probability distribution matrix is adjusted; and according to the true classification label of the training sample, the probability distribution matrix is aligned and enhanced.
[0124] Step S406, label propagation is performed based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain a text classification result of the text to be classified.
[0125] Through prompt-based learning in this embodiment, the pre-trained language model is used as a classifier for small sample tasks, and the text classification task is expressed as a task-oriented text completion problem in natural language prompts. On this basis, the semantic relationship between texts is used to align and enhance the inference results of the pre-trained large model without introducing additional training costs, which alleviates the dependence on labeled data and improves the accuracy of text classification when samples are scarce.
[0126] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0127] In this embodiment, a large model text classification system based on an alignment strategy is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. The terms "module", "unit", "sub-unit", etc. used below may implement a combination of software and / or hardware for a predetermined function. Although the system described in the following embodiments is preferably implemented in software, the implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0128] Figure 5 is a structural block diagram of the large model text classification system based on the alignment strategy of this embodiment, such as Figure 5 As shown, the system includes:
[0129] A prompt construction module 10, for constructing a language prompt based on the text to be classified;
[0130] A model output module 20, for obtaining an output vector based on a pre-trained language model and a language prompt;
[0131] The probability distribution calculation module 30 is used to determine the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; the probability distribution matrix includes the predicted classification label;
[0132] The text classification reasoning module 40 is used to perform label propagation based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified.
[0133] Through the system provided by this embodiment, language prompts are constructed based on the text to be classified, the text classification task is converted into a task-oriented text completion problem in a natural language prompt, the pre-trained language model is used as a classifier under a small sample task, and the semantic relationship in the text semantic graph is further used for label propagation. A post-training alignment strategy is adopted to align and enhance the predicted text classification results of the pre-trained language model to alleviate the dependence on sample labels. Compared with the method in the prior art that relies on a large number of labeled samples to optimize the model weights, this embodiment can improve the accuracy of model classification when samples are scarce. In addition, the post-training alignment strategy can fully utilize label knowledge without introducing additional training costs, and will not affect the pre-training characteristics and robustness of the pre-trained language model.
[0134] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0135] In this embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0136] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0137] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.
[0138] In addition, in combination with the large model text classification method based on alignment strategy provided in the above embodiments, a storage medium can also be provided in this embodiment to implement. The storage medium stores a computer program; when the computer program is executed by the processor, any large model text classification method based on alignment strategy in the above embodiments is implemented.
[0139] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0140] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0141] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0142] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0143] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A large model text classification method based on alignment strategy, characterized in that: include: Construct language prompts based on the text to be classified; Based on the pre-trained language model and the language prompt, obtaining an output vector; Determine a probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; The probability distribution matrix includes predicted classification labels; Label propagation is performed based on the pre-constructed text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified; wherein, the following steps are included: Based on the connection weights and edge weights in the text semantic graph, label propagation is performed in combination with the probability distribution matrix after alignment enhancement; wherein, the probability distribution matrix after alignment enhancement processing is constructed according to the adjusted predicted classification labels in the probability distribution matrix and the true classification labels of the training samples; According to the updated predicted classification labels of the text nodes in the text semantic graph, the text classification results of the text to be classified are determined; the following is the text classification result A way of expressing: ; Wherein, A represents the text semantic graph; k represents the number of hops of label propagation; H represents the probability distribution matrix after alignment enhancement; k-hop label propagation is run in the text semantic graph, starting from the text node, and according to the pre-defined label propagation rules, the probability distribution is propagated along the edge of the graph to the adjacent word nodes; for each word node, its own probability distribution is updated according to the probability distribution of the text node connected to it; the updated word node probability distribution is propagated back to the text node, or propagated to other text nodes that are not directly connected; after k-hop propagation, for each text node, the updated probability distribution is aggregated to obtain the final corresponding predicted classification label; according to the probability of the aggregated predicted classification label, the category with the highest probability is selected as the text classification result.
2. The method according to claim 1, characterized in that The step of obtaining an output vector based on the pre-trained language model and the language prompt includes: Presetting candidate words of the pre-trained language model; The language prompt is input into the pre-trained language model to obtain an output vector of the pre-trained language model for each of the candidate words.
3. The method according to any one of claim 1 or claim 2, characterized in that: Also includes: Calculate the loss function value based on the preset loss function and the true classification label of the training sample; The output vector of the pre-trained language model is adjusted with minimizing the loss function value as an optimization goal.
4. The method according to claim 1, characterized in that Determining the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector includes: Determine a probability score and a predicted classification label for each answer in the candidate word according to the output vector and the candidate word; The probability distribution matrix is determined based on the probability scores and the predicted classification labels.
5. The method according to claim 3, characterized in that: Also includes: Adjusting the predicted classification labels in the probability distribution matrix based on a preset threshold and the probability scores in the probability distribution matrix; According to the true classification labels of the training samples, the probability distribution matrix is aligned and enhanced.
6. The method according to claim 1, characterized in that Also includes: Determine text nodes and word nodes according to the text to be classified; Based on the sliding window, determining the number of occurrences of a word and the number of co-occurrences of two words in the text to be classified; Calculate the connection weight between the word nodes and the edge weight between the text node and the word node according to the word occurrence count and the two word co-occurrence count; The text semantic graph is constructed based on the connection weights and the edge weights.
7. A large model text classification system based on alignment strategy, characterized in that: include: A prompt building module, used to build language prompts based on the text to be classified; A model output module, used to obtain an output vector based on the pre-trained language model and the language prompt; A probability distribution calculation module, used to determine the probability distribution matrix of all candidate words of the pre-trained language model according to the output vector; The probability distribution matrix includes predicted classification labels; The text classification reasoning module is used to perform label propagation based on the pre-built text semantic graph and the probability distribution matrix after alignment enhancement to obtain the text classification result of the text to be classified; wherein, the following steps are included: Based on the connection weights and edge weights in the text semantic graph, label propagation is performed in combination with the probability distribution matrix after alignment enhancement; wherein, the probability distribution matrix after alignment enhancement processing is constructed according to the adjusted predicted classification labels in the probability distribution matrix and the true classification labels of the training samples; According to the updated predicted classification labels of the text nodes in the text semantic graph, the text classification results of the text to be classified are determined; the following is the text classification result A way of expressing: ; Wherein, A represents the text semantic graph; k represents the number of hops of label propagation; H represents the probability distribution matrix after alignment enhancement; k-hop label propagation is run in the text semantic graph, starting from the text node, and according to the pre-defined label propagation rules, the probability distribution is propagated along the edge of the graph to the adjacent word nodes; for each word node, its own probability distribution is updated according to the probability distribution of the text node connected to it; the updated word node probability distribution is propagated back to the text node, or propagated to other text nodes that are not directly connected; after k-hop propagation, for each text node, the updated probability distribution is aggregated to obtain the final corresponding predicted classification label; according to the probability of the aggregated predicted classification label, the category with the highest probability is selected as the text classification result.
8. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the large model text classification method based on alignment strategy according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the large model text classification method based on alignment strategy described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Mutual learning text classification method and system based on graph enhancement
CN115599918A
Prompt learning small sample classification method, system and equipment based on pre-training language model and medium
CN116415170A
Conversational data generation method and natural language reasoning data classification method and system
CN118312589A