Text classification method and device, equipment and storage medium
By optimizing the pre-trained model and building the label set, the problems of low efficiency and poor accuracy of text classification are solved, efficient and accurate text classification results are achieved, and computing resource requirements are reduced.
Patent Information
- Application Number
- CN202510354227.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, text classification methods have low efficiency and poor accuracy, especially in the case of high computational complexity, limited long text processing capabilities and scarce data in specific fields, and the randomness of large language models leads to poor interpretability of the results.
By obtaining labeled text data, inputting it to the first pretrained model output training data set, and inputting it to the second pretrained model for optimization, using preset prompt words and reordering framework to build a set of positive and negative example tags, combining pretrained word embedding models and model acceleration components, optimizing parameters of large language models to improve text classification efficiency and accuracy.
It improves the efficiency and accuracy of text classification, reduces the demand for computing resources, and enhances the generalization ability of the model in specific fields and the interpretability of the results.
Smart Images

Figure CN120296170A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular, to a text classification method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of the Internet, text data has shown explosive growth. How to effectively classify a large amount of text has become an important topic in the field of natural language processing.
[0003] Traditional text classification methods mainly rely on manual feature engineering and shallow models, and have problems such as difficult feature extraction and insufficient generalization ability. In recent years, with the rise of deep learning technology, especially the development of large language models (LLMs), the performance of text classification has been significantly improved. These models can capture rich semantic information through pre-training on a large-scale corpus, and thus achieve excellent performance in various natural language processing tasks.
[0004] However, although large language models show powerful capabilities in text classification, they still face challenges in practical applications. For example, the computational complexity of the model is high and it is difficult to meet the real-time requirements; the processing ability for long texts is limited, which may lead to information loss; in the case of scarce data in a specific field, the generalization ability of the model may be restricted. In addition, the randomness of the results generated by large language models makes the interpretability of the results poor, affecting their applications in specific fields. Summary of the Invention
[0005] This application provides a text classification method, apparatus, device, and storage medium to solve the defects of low text classification efficiency and poor accuracy in the prior art.
[0006] In a first aspect, this application provides a text classification method, which includes:
[0007] Obtain labeled text data, where the labeled text data includes: text data and label data;
[0008] Input the labeled text data into a first pre-trained model, and output a training data set through the first pre-trained model;
[0009] Input the training data set into a second pre-trained model, and optimize the second pre-trained model according to the training data set;
[0010] Analyze and process the text data to be processed through the optimized second pre-trained model to obtain a text classification result, where the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
[0011] Optionally, inputting the labeled text data into the first pre-trained model and outputting a training data set through the first pre-trained model includes:
[0012] Determine the text labels corresponding to the text data in the label data according to a preset prompt word to obtain a plurality of first text labels, construct a first positive example label set and a first negative example label set based on the plurality of first text labels, and perform sorting processing on the first negative example label set, where the prompt word is constructed based on the chain of thought prompt engineering;
[0013] Input the text data, the first positive example label set, and the sorted first negative example label set into the first pre-trained model, adjust the parameters of the first pre-trained model to obtain an adjusted first pre-trained model;
[0014] Determine the text labels corresponding to the text data in the label data according to the preset prompt word through the adjusted first pre-trained model to obtain a plurality of second text labels, and construct a second positive example label set and a second negative example label set based on the plurality of second text labels;
[0015] Re-sort the negative example labels in the second negative example label set based on a re-ranking framework, and construct a third negative example label set according to the re-ranking result;
[0016] Construct the training data set based on the text data, the third negative example label set, and the second positive example label set.
[0017] Optionally, the determining the text labels corresponding to the text data in the label data according to a preset prompt word to obtain a plurality of first text labels, constructing a first positive example label set and a first negative example label set based on the plurality of first text labels, and performing sorting processing on the first negative example label set includes:
[0018] Determine the cosine similarity between the text data and multiple text labels in the label data, and use the text labels with a cosine similarity higher than a preset cosine similarity threshold as the first text labels;
[0019] Use the labels manually labeled among the first text labels as positive example labels, and construct the first positive example label set according to the plurality of positive example labels;
[0020] Use the labels not manually labeled among the first text labels as negative example labels, construct the first negative example label set according to the plurality of negative example labels, and sort the plurality of negative example labels in the first negative example label set according to the cosine similarity between the negative example labels and the text data.
[0021] Optionally, inputting the text data, the first positive example label set, and the sorted first negative example label set into the first pre-trained model to adjust the parameters of the first pre-trained model, obtaining an adjusted first pre-trained model, includes:
[0022] Respectively extracting features of the text data, the positive example labels in the first positive example label set, and the negative example labels in the first negative example label set through the first pre-trained model to obtain corresponding text data vectors, first positive example label vectors, and first negative example label vectors;
[0023] Constructing a cross-entropy loss according to the first positive example label vector and the first negative example label vector, and adjusting the parameters of the first pre-trained model based on the cross-entropy loss to obtain an adjusted first pre-trained model.
[0024] Optionally, re-ranking the negative example labels in the second negative example label set based on a re-ranking framework, and constructing a third negative example label set according to the re-ranking result, includes:
[0025] Determining the relevance between the negative example labels in the second negative example label set and the text data based on the re-ranking framework, sorting the negative example labels according to the relevance to obtain a sorting result;
[0026] Based on the sorting result, obtaining a third negative example label set composed of the negative example labels in the second negative example label set that meet the preset conditions through hard negative sample mining.
[0027] Optionally, inputting the training data set into a second pre-trained model, and optimizing the second pre-trained model according to the training data set, includes:
[0028] Using a pre-trained word embedding model to respectively extract features of the text data, the negative example labels in the third negative example label set, and the positive example labels in the second positive example label set to obtain corresponding text data vectors, second negative example label vectors, and second positive example label vectors;
[0029] Inputting the text data vector, the second negative example label vector, and the second positive example label vector into the second pre-trained model;
[0030] Taking the output of a preset loss function as an optimization target, and adjusting the weight parameters of the second pre-trained model through backpropagation algorithm until the output value of the preset loss function is less than a preset output threshold.
[0031] Optionally, performing a recall process on the text data to be processed through the optimized second pre-trained model to obtain a text classification result, includes:
[0032] The optimized second pre-trained model is accelerated by a model acceleration component;
[0033] The text data to be processed is input into the optimized and accelerated second pre-trained model, and the text labels in the label file are retrieved according to the preset prompt words to determine the text labels corresponding to the text data to be processed, and a text classification result is obtained.
[0034] In a second aspect, the present application provides a text classification device, and the device includes:
[0035] An acquisition module, configured to acquire labeled text data, where the labeled text data includes: text data and label data;
[0036] A processing module, configured to input the labeled text data into a first pre-trained model, and output a training data set through the first pre-trained model;
[0037] The processing module is further configured to input the training data set into a second pre-trained model, and optimize the second pre-trained model according to the training data set;
[0038] The processing module is further configured to analyze and process the text data to be processed through the optimized second pre-trained model to obtain a text classification result, where the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
[0039] In a third aspect, the present application provides a text classification device, including:
[0040] A memory;
[0041] A processor;
[0042] Wherein, the memory stores computer execution instructions;
[0043] The processor executes the computer execution instructions stored in the memory to implement the text classification method as described in the first aspect and various possible implementation manners of the first aspect.
[0044] In a fourth aspect, the present application provides a computer storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the text classification method as described in the first aspect and various possible implementation manners of the first aspect.
[0045] The present application provides a text classification method, apparatus, device, and storage medium. The method includes obtaining labeled text data, where the labeled text data includes text data and label data; inputting the labeled text data into a first pre-trained model, and outputting a training data set through the first pre-trained model; inputting the training data set into a second pre-trained model, and optimizing the second pre-trained model according to the training data set; analyzing and processing the text data to be processed through the optimized second pre-trained model to obtain a text classification result, improving the efficiency and accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0047] Figure 1 Flow diagram of the text classification method provided by the present application Figure 1 ;
[0048] Figure 2 Flow diagram of the text classification method provided by the present application Figure 2 ;
[0049] Figure 3 Structural diagram of the text classification apparatus provided by the present application;
[0050] Figure 4 Structural diagram of the text classification device provided by the present application.
[0051] Through the above accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0053] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.
[0054] In the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0055] With the rapid advancement of Internet technology, the amount of text data has shown a sharp growth trend. How to efficiently classify these massive texts has become a key issue that needs to be urgently solved in the field of natural language processing.
[0056] Traditional text classification mainly relies on artificial feature construction and shallow models, which are often accompanied by limitations such as difficulty in feature extraction and poor model generalization performance.
[0057] The computational cost of existing large language models is high, making it difficult to meet the needs of real-time processing; their capabilities are limited when processing long texts, which may lead to the omission of key information; and when data in specific fields is scarce, the generalization performance of the model may be affected.
[0058] In response to the above problems, the present application proposes a text classification method, which obtains annotated text data, inputs the annotated text data into a first pre-trained model, outputs a training data set through the first pre-trained model, inputs the training data set into a second pre-trained model, optimizes the second pre-trained model according to the training data set, and analyzes and processes the text data to be processed through the optimized second pre-trained model to obtain a text classification result, thereby improving the efficiency and accuracy of text classification.
[0059] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0060] Figure 1 A flowchart of a text classification method provided in an embodiment of the present application Figure 1 .like Figure 1As shown in the figure, the text classification method provided in this embodiment includes:
[0061] S101: Obtain labeled text data, where the labeled text data includes: text data and label data.
[0062] Among them, the labeled text data is a processed and marked data form, consisting of two parts: text data and label data. The text data refers to various forms of text content; the label data is a mark for specific attributes, categories, or features of the text data.
[0063] In the fields of machine learning and natural language processing, the labeled text data can provide clear learning objectives and reference standards for the model, helping the model better understand the semantics and features of the text, thereby improving the accuracy and generalization ability of the model.
[0064] S102: Input the labeled text data into the first pre-trained model, and output a training data set through the first pre-trained model.
[0065] Among them, the first pre-trained model is a model pre-trained on a large-scale data set. The pre-training process enables the model to learn general language knowledge and patterns, and has certain language understanding and processing capabilities. The training data set is a data set used to train machine learning or deep learning models, containing the labeled text data processed by the first pre-trained model and the corresponding label data.
[0066] It can be understood that a suitable first pre-trained model is selected according to specific task requirements and the characteristics of the labeled text data, the labeled text data is pre-processed, the pre-processed labeled text data is input into the first pre-trained model, and the output result of the first pre-trained model is received, and a training data set is formed according to the output result.
[0067] By inputting the labeled text data into the first pre-trained model for training, the quality of the training data set and the performance of the model can be improved.
[0068] S103: Input the training data set into the second pre-trained model, and optimize the second pre-trained model according to the training data set.
[0069] Among them, the second pre-trained model is a model used to complete the final text classification target task. The second pre-trained model is a large language model of the same type as the first pre-trained model, and the number of parameters of the second pre-trained model is less than that of the first pre-trained model, which can save computing resources and improve processing speed when deployed to the user server.
[0070] It can be understood that optimizing the second pre-trained model refers to adjusting the parameters of the second pre-trained model so that the model can better complete the task on the given training data set. The optimization goal can be, for example, to minimize the error between the model's prediction result and the true label.
[0071] Specifically, select an appropriate loss function according to the type of task, and select an appropriate optimizer according to the characteristics of the model and the requirements of the task. Divide the training data set into a training set and a validation set. The training set is used to update the model's parameters, and the validation set is used to evaluate the performance of the model on unseen data to prevent overfitting of the model. Divide the training data set into several small batches and input them into the model for training in turn. After each batch training is completed, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and use the optimizer to update the model's parameters. During the training process, regularly evaluate the performance of the model on the validation set, and adjust the hyperparameters according to the evaluation results to optimize the training effect of the model.
[0072] By optimizing the second pre-trained model using a specific training data set, the model can learn specific features and patterns related to the task, thereby improving the performance of the model on this task.
[0073] S104: Analyze and process the text data to be processed through the optimized second pre-trained model to obtain a text classification result.
[0074] Among them, the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
[0075] It can be understood that preprocess the text data to be processed, and input the preprocessed text data to be processed into the optimized second pre-trained model. The second pre-trained model will analyze and calculate the input text according to the features and patterns it has learned, obtain the label information corresponding to the text, and use the label information corresponding to the text as the result of this text classification.
[0076] Using the optimized second pre-trained model can achieve automated text classification, improving the efficiency and accuracy of classification.
[0077] A text classification method provided by an embodiment of the present application. This method obtains annotated text data, where the annotated text data includes: text data and label data; inputs the annotated text data into a first pre-trained model, and outputs a training data set through the first pre-trained model; inputs the training data set into a second pre-trained model, and optimizes the second pre-trained model according to the training data set; analyzes and processes the text data to be processed through the optimized second pre-trained model to obtain a text classification result, improving the efficiency and accuracy of text classification.
[0078] Figure 2 A text classification method process provided by an embodiment of the present application Figure 2 . This embodiment is based on the Figure 1 embodiment, and a possible implementation manner of the text classification method is described in detail. As Figure 2 shown, the method includes:
[0079] S201: Obtain labeled text data, where the labeled text data includes: text data and label data.
[0080] Among them, step S201 is similar to the above step S101 and will not be elaborated here.
[0081] S202: Determine the text label corresponding to the text data in the label data according to a preset prompt word, obtain a plurality of first text labels, construct a first positive example label set and a first negative example label set based on the plurality of first text labels, and perform a sorting process on the first negative example label set.
[0082] Among them, the preset prompt word is used to screen out keywords or phrases related to the text data in the label data. The prompt word is constructed based on the chain-of-thought prompt engineering. By constructing the prompt word, the model can generate a thinking process and optimize the output effect of the model. The first text label is the text label corresponding to the text data screened out in the label data by the preset prompt word. The first positive example label set is a set composed of labels that are positively correlated with the text data among the plurality of first text labels. The first negative example label set contains labels that are not relevant or weakly relevant to the text data.
[0083] In this embodiment, determining the text label corresponding to the text data in the label data according to a preset prompt word, obtaining a plurality of first text labels, constructing a first positive example label set and a first negative example label set based on the plurality of first text labels, and performing a sorting process on the first negative example label set includes: determining the cosine similarity between the text data and multiple text labels in the label data, and using the text labels with a cosine similarity higher than the preset cosine similarity threshold as the first text labels; using the labels manually annotated among the first text labels as positive example labels, and constructing a first positive example label set according to the multiple positive example labels; using the labels not manually annotated among the first text labels as negative example labels, constructing a first negative example label set according to the multiple negative example labels, and sorting the multiple negative example labels in the first negative example label set according to the cosine similarity between the negative example labels and the text data.
[0084] It is understandable that the set of positive example labels can provide accurate learning objectives for the model, enabling the model to identify the core features of the text. The set of negative example labels can help the model recognize information irrelevant to the text, improving the model's discriminative ability and generalization ability. Sorting the first set of negative example labels can further optimize the subsequent text analysis process and enhance the model's ability to handle different types of negative examples.
[0085] In some alternative embodiments, the steps of constructing a prompt using the chain of thought technique include:
[0086] S301: Define the required request or question. Specifically, clearly state the specific request or task to be solved through natural language, ensuring that the task or question has clear context information for subsequent parsing and reasoning using the chain of thought technique. The purpose of this step is to obtain a specific goal or request as the basis for constructing the prompt. Specifically, the specific goal is to perform labeled classification on the input text data after retrieving the given labels in the retrieval label file.
[0087] S302: Analyze the input request or question to identify each key component included in the task. The key components can be the background information of the request, the specific content required, possible constraints, etc. The analysis results help construct the prompt for the chain of thought in the subsequent step to support the reasoning process. Specifically, the model is required to carefully read the input text data and understand its key information.
[0088] S303: Based on the analysis results of the request or question, construct a chain of thought consisting of multiple intermediate reasoning steps. The reasoning steps will gradually drive the reasoning process from the input of the question to the final answer, and each reasoning step is associated with the goal of the request, ensuring the coherence and integrity of the reasoning. Specifically, provide all the classifications in different fields in the label file to the model and let it first classify the current text into one specific field, then let the model retrieve in all the label sets of that field, and finally let the model check the relevance between the returned label and the text and output the results in descending order.
[0089] S304: Provide a specific example to the model. The example is a complete project data, and the complete project data is text data with the process of information extraction from the text data in the relevant field by professionals in the relevant field and the labeled tags after extraction. This step can show the model the specific implementation form of the task process, making the model's returned results more in line with the format required by the project.
[0090] In some alternative embodiments, the construction of the chain of thought can be designed through preset logic or domain knowledge. By concatenating a series of thought steps into a chain structure, the generated prompt can prompt the model to clarify the specific meaning of each intermediate step during the reasoning process and build the final answer based on this. This process can not only help the model understand and organize complex thought structures, but also verify the correctness of each link in a step-by-step manner, thereby avoiding the generation and reasoning of errors.
[0091] S203: Input the text data, the first set of positive example labels, and the sorted first set of negative example labels into the first pre-trained model, and adjust the parameters of the first pre-trained model to obtain an adjusted first pre-trained model.
[0092] Specifically, the first pre-trained model extracts features from the text data, the positive example labels in the first set of positive example labels, and the negative example labels in the first set of negative example labels respectively to obtain corresponding text data vectors, first positive example label vectors, and first negative example label vectors; construct a cross-entropy loss based on the first positive example label vectors and the first negative example label vectors, and adjust the parameters of the first pre-trained model based on the cross-entropy loss to obtain an adjusted first pre-trained model.
[0093] It can be understood that the adjusted first pre-trained model can better adapt to a specific task and has a more accurate analysis and processing ability for the input text.
[0094] S204: Use the adjusted first pre-trained model to determine the text label corresponding to the text data in the label data according to the preset prompt to obtain multiple second text labels, and construct a second set of positive example labels and a second set of negative example labels based on the multiple second text labels.
[0095] Among them, the second text label is the corresponding text label determined by the adjusted first pre-trained model for the text data in the label data according to the preset prompt, which is the result of the model's re-screening and matching, and may contain multiple labels. The second set of positive example labels is a set composed of the labels that are positively related to the text data selected from the multiple second text labels. The second set of negative example labels is a set containing the second text labels that are not related or weakly related to the text data.
[0096] It can be understood that by using the adjusted first pre-trained model to determine the label corresponding to the text, the label related to the text can be determined more accurately, improving the accuracy and reliability of the label; by constructing the second set of positive example labels and the second set of negative example labels, it helps to further optimize the learning and classification effects of the model and provides more accurate input for subsequent text classification.
[0097] S205: Reorder the negative example labels in the second negative example label set based on the reordering framework, and construct a third negative example label set according to the reordering result.
[0098] Specifically, determine the relevance between the negative example labels in the second negative example label set and the text data based on the reordering framework, sort the negative example labels according to the relevance to obtain a sorting result; based on the sorting result, obtain a third negative example label set composed of the negative example labels in the second negative example label set that meet the preset conditions through hard negative sample mining.
[0099] In this embodiment, a reordering model based on the Bidirectional Generative Encoder (BGE) architecture is adopted to perform higher-precision reordering on the training data set, so as to ensure that the labels in the negative example label set are strictly sorted in descending order according to the text relevance degree, and a pre-trained large language model is used to obtain the most challenging negative example labels, that is, multiple labels with the highest text relevance degree closest to the positive example, through hard negative sample mining; among them, hard negative sample mining is a method that preferentially selects the most difficult-to-distinguish negative samples for the current model during the model training process, aiming to improve the model's discrimination ability for complex negative samples, thereby improving the overall performance.
[0100] S206: Construct the training data set based on the text data, the third negative example label set, and the second positive example label set.
[0101] S207: Use a pre-trained word embedding model to extract features from the text data, the negative example labels in the third negative example label set, and the positive example labels in the second positive example label set respectively, to obtain corresponding text data vectors, second negative example label vectors, and second positive example label vectors.
[0102] In this embodiment, a pre-trained word embedding model (such as Word2Vec, GloVe, BERT, etc.) is used to convert the training text data into a vector representation of a fixed dimension. The text data can be converted into a vector through sentence-level, paragraph-level, or document-level embedding. This vector can contain the semantic information of the text, enabling the model to perform calculations and comparisons in a high-dimensional space.
[0103] S208: Input the text data vector, the second negative example label vector, and the second positive example label vector into the second pre-trained model.
[0104] It can be understood that the generated vector is used as the input and input into a pre-selected large model for fine-tuning. During the fine-tuning process, the model gradually adapts to the specific task requirements by continuously adjusting the weight parameters, thereby improving the effect and accuracy of processing tasks.
[0105] S209: Use the output of the preset loss function as the optimization objective, and adjust the weight parameters of the second pre-trained model through the backpropagation algorithm until the output value of the preset loss function is less than the preset output threshold.
[0106] In this embodiment, a preset loss function (such as cross-entropy loss) is used to calculate the error between the generated result of the model and the manually annotated result, and the output of the loss function is used as the optimization objective. The weight parameters of the large model are adjusted through the backpropagation algorithm. The optimization process uses algorithms such as gradient descent to minimize the loss function until the model reaches better performance.
[0107] It can be understood that after the adjustment of the weight parameters of the second pre-trained model is completed, the adjusted model is evaluated through the validation set to check its effect in actual applications. If the evaluation results show that the performance does not meet expectations, further optimization can be carried out based on the evaluation feedback. The optimization includes adjusting the learning rate, modifying the network structure, or using other regularization methods, etc.
[0108] S210: Accelerate the optimized second pre-trained model through the model acceleration component.
[0109] In this embodiment, a large language model (Very Large Language Model, abbreviated as vLLM) component is loaded in the text data processing system, and the virtualization parameters of vLLM are configured. Among them, the virtualization parameters include but are not limited to video memory allocation strategy, computing task priority, micro-batch processing size, and caching mechanism. Through optimized configuration, it is ensured that while reducing the computing resource occupancy, the model inference performance is maintained.
[0110] The vLLM hierarchical technology can effectively improve the inference efficiency and adaptability of large-scale language models in different hardware platforms and application scenarios. Especially on devices with limited computing resources, it ensures that the model can maintain high inference performance.
[0111] S211: Input the text data to be processed into the optimized and accelerated second pre-trained model, and retrieve the text labels in the label file according to the preset prompt words to determine the text labels corresponding to the text data to be processed, and obtain the text classification result.
[0112] In this embodiment, the task to be solved is input and passed to the vLLM component. Its virtualization scheduling module allocates resources and splits tasks for the inference task, and dynamically allocates resources to the split tasks to reduce resource occupancy and speed up the task running rate. Among them, the task to be solved is the text data to be input into the model component. After the current task completes the inference, it is integrated with all results to generate a complete task output.
[0113] A text classification method provided by an embodiment of the present application determines a text label corresponding to text data in label data through a preset prompt word, obtains a plurality of first text labels, constructs a first positive example label set and a first negative example label set based on the plurality of first text labels, and performs a sorting process on the first negative example label set. The text data, the first positive example label set, and the sorted first negative example label set are input into a first pre-trained model, and the parameters of the first pre-trained model are adjusted to obtain an adjusted first pre-trained model. Through the adjusted first pre-trained model, a text label corresponding to the text data is determined in the label data according to the preset prompt word, obtaining a plurality of second text labels, constructing a second positive example label set and a second negative example label set based on the plurality of second text labels, re-ranking the negative example labels in the second negative example label set based on a re-ranking framework, constructing a third negative example label set according to the re-ranking result, constructing a training data set based on the text data, the third negative example label set, and the second positive example label set, using a pre-trained word embedding model to extract features from the text data, the negative example labels in the third negative example label set, and the positive example labels in the second positive example label set respectively, obtaining corresponding text data vectors, second negative example label vectors, and second positive example label vectors, inputting the text data vectors, the second negative example label vectors, and the second positive example label vectors into a second pre-trained model, using the output of a preset loss function as an optimization target, adjusting the weight parameters of the second pre-trained model through a backpropagation algorithm until the output value of the preset loss function is less than a preset output threshold, accelerating the optimized second pre-trained model through a model acceleration component, inputting the text data to be processed into the optimized and accelerated second pre-trained model, and retrieving the text labels in the label file according to the preset prompt word to determine the text label corresponding to the text data to be processed, obtaining a text classification result. This method processes the text data through the first pre-trained model to obtain a training data set, and adjusts the second pre-trained model through this training data set, improving the accuracy of the second pre-trained model and making the output result more accurate.
[0114] Figure 3 It is a schematic structural diagram of a text classification device provided by the present application. As Figure 3 shown, the text classification device 300 provided in this embodiment includes:
[0115] An acquisition module 301, configured to acquire annotated text data, where the annotated text data includes: text data and label data;
[0116] A processing module 302, configured to input the annotated text data into a first pre-trained model, and output a training data set through the first pre-trained model;
[0117] The processing module 302 is further configured to input the training data set into the second pre-trained model, and optimize the second pre-trained model according to the training data set.
[0118] The processing module 302 is further configured to analyze and process the text data to be processed through the optimized second pre-trained model, and obtain a text classification result, where the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
[0119] A text classification device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0120] Figure 4 It is a schematic structural diagram of a text classification device provided in the present application. As Figure 4 shown, the text classification device provided in the present application, the text classification device 400 includes: a receiver 401, a transmitter 402, a processor 403, and a memory 404.
[0121] The receiver 401 is configured to receive instructions and data;
[0122] The transmitter 402 is configured to send instructions and data;
[0123] The memory 404 is configured to store computer execution instructions;
[0124] The processor 403 is configured to execute the computer execution instructions stored in the memory 404 to implement each step executed by the text classification method in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing text classification method embodiment.
[0125] Optionally, the above memory 404 can be either independent or integrated with the processor 403.
[0126] When the memory 404 is independently provided, the electronic device further includes a bus for connecting the memory 404 and the processor 403.
[0127] The present application further provides a computer storage medium, in which computer execution instructions are stored, and when the processor executes the computer execution instructions, the energy storage power station capacity evaluation method executed by the above text classification device is implemented.
[0128] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0129] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0130] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A text classification method, characterized in that, The method includes: Obtaining labeled text data, where the labeled text data includes: text data and label data; Inputting the labeled text data into a first pre-trained model, and outputting a training data set through the first pre-trained model; Inputting the training data set into a second pre-trained model, and optimizing the second pre-trained model according to the training data set; Analyzing and processing the text data to be processed through the optimized second pre-trained model to obtain a text classification result, where the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
2. The method according to claim 1, characterized in that, The step of inputting the labeled text data into a first pre-trained model and outputting a training data set through the first pre-trained model includes: Determining the text labels corresponding to the text data in the label data according to a preset prompt word to obtain a plurality of first text labels, constructing a first positive example label set and a first negative example label set based on the plurality of first text labels, and sorting the first negative example label set, where the prompt word is constructed based on the chain-of-thought prompt engineering; Inputting the text data, the first positive example label set, and the sorted first negative example label set into the first pre-trained model, adjusting the parameters of the first pre-trained model to obtain an adjusted first pre-trained model; Determining the text labels corresponding to the text data in the label data according to the preset prompt word through the adjusted first pre-trained model to obtain a plurality of second text labels, and constructing a second positive example label set and a second negative example label set based on the plurality of second text labels; Re-ranking the negative example labels in the second negative example label set based on a re-ranking framework, and constructing a third negative example label set according to the re-ranking result; Constructing the training data set based on the text data, the third negative example label set, and the second positive example label set.
3. The method according to claim 2, wherein The step of determining the text labels corresponding to the text data in the label data according to a preset prompt word to obtain a plurality of first text labels, constructing a first positive example label set and a first negative example label set based on the plurality of first text labels, and sorting the first negative example label set includes: Determining the cosine similarity between the text data and a plurality of text labels in the label data, and using the text labels with a cosine similarity higher than a preset cosine similarity threshold as the first text labels; Using the labels that have been manually labeled among the first text labels as positive example labels, and constructing the first positive example label set according to the plurality of positive example labels; Using the labels that have not been manually labeled among the first text labels as negative example labels, constructing the first negative example label set according to the plurality of negative example labels, and sorting the plurality of negative example labels in the first negative example label set according to the cosine similarity between the negative example labels and the text data.
4. The method according to claim 3, characterized in that, Inputting the text data, the first positive example label set, and the sorted first negative example label set into the first pre-trained model to adjust the parameters of the first pre-trained model, obtaining an adjusted first pre-trained model, includes: Respectively extracting features from the text data, the positive example labels in the first positive example label set, and the negative example labels in the first negative example label set through the first pre-trained model, obtaining corresponding text data vectors, first positive example label vectors, and first negative example label vectors; Constructing a cross-entropy loss according to the first positive example label vector and the first negative example label vector, and adjusting the parameters of the first pre-trained model based on the cross-entropy loss, obtaining an adjusted first pre-trained model.
5. The method according to claim 2, wherein Re-ranking the negative example labels in the second negative example label set based on the re-ranking framework, and constructing a third negative example label set according to the re-ranking result, includes: Determining the relevance between the negative example labels in the second negative example label set and the text data based on the re-ranking framework, sorting the negative example labels according to the relevance, obtaining a sorting result; Based on the sorting result, obtaining a third negative example label set composed of the negative example labels in the second negative example label set that meet the preset conditions through hard negative sample mining.
6. The method according to claim 1, characterized in that Inputting the training data set into the second pre-trained model, and optimizing the second pre-trained model according to the training data set, includes: Respectively extracting features from the text data, the negative example labels in the third negative example label set, and the positive example labels in the second positive example label set using a pre-trained word embedding model, obtaining corresponding text data vectors, second negative example label vectors, and second positive example label vectors; Inputting the text data vector, the second negative example label vector, and the second positive example label vector into the second pre-trained model; Taking the output of a preset loss function as an optimization target, adjusting the weight parameters of the second pre-trained model through the backpropagation algorithm until the output value of the preset loss function is less than a preset output threshold.
7. The method according to claim 2, characterized in that, Performing a recall process on the text data to be processed through the optimized second pre-trained model, obtaining a text classification result, includes: Performing an acceleration process on the optimized second pre-trained model through a model acceleration component; Inputting the text data to be processed into the optimized and accelerated second pre-trained model, and retrieving the text labels in the label file according to the preset prompt words, determining the text labels corresponding to the text data to be processed, obtaining a text classification result.
8. A text classification device, characterized in that, The device includes: An acquisition module, configured to acquire annotated text data, where the annotated text data includes: text data and label data; A processing module, configured to input the annotated text data into the first pre-trained model, and output a training data set through the first pre-trained model; The processing module is further configured to input the training data set into the second pre-trained model, and optimize the second pre-trained model according to the training data set; The processing module is further configured to analyze and process the text data to be processed through the optimized second pre-trained model, and obtain a text classification result, where the text classification result is used to indicate the correspondence between the text data to be processed and the text label.
9. A text classification device, characterized in that, The device includes: a memory; a processor; wherein, the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement a text classification method according to any one of claims 1-7.
10. A computer storage medium, characterized in that, Computer-executable instructions are stored in the computer storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement a text classification method according to any one of claims 1-7.