A training method of a generative model and related apparatuses

CN122838633APending Publication Date: 2026-09-29BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510387333.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]但是,上述处理方式针对具有复杂标签体系的多级文本分类任务,分类准确率不高

Benefits of technology

[0052]由上述技术方案可以看出,该技术方案首先获取标签体系数据以及训练样本,其中,标签体系数据描述了多个样本标签之间的层级关系,该层级关系用于标识相邻两个层级之间具有联系的样本标签,包括第i层级的样本标签以及第i+1层级中该样本标签的一个或多个子标签,训练样本的样本标签可标识训练样本的真实类别,然后将标签体系数据,任务提示词以及训练样本输入到初始生成式模型,初始生成式模型通过学习训练样本的语义特征,可以确定其对应的预测类别,从而通过初始生成式模型得到了训练样本对应的分类预测结果,最后通过分类预测结果与样本标签之间的差异,对初始生成式模型进行训练,得到分类生成式模型,由此该分类生成式模型可以实现对文本的多级文本分类,且根据相邻层级的样本标签之间的内在联系关系确定类别标签,有效提高了分类的准确性和效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838633A_ABST
    Figure CN122838633A_ABST
Patent Text Reader

Abstract

This application discloses a training method and related apparatus for a generative model. First, it acquires label system data and training samples. The label system data identifies the hierarchical relationship between multiple sample labels, and the sample labels of the training samples identify their true categories. Then, it inputs task prompts, label system data, and training samples into an initial generative model. The initial generative model obtains the classification prediction results corresponding to the training samples. By analyzing the difference between the classification prediction results and the sample labels, the initial generative model is trained to obtain a classification generative model. This classification generative model can achieve multi-level text classification and determine category labels based on the inherent relationships between sample labels at adjacent levels, effectively improving the accuracy and efficiency of classification in multi-level text classification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, specifically to a training method and related apparatus for a generative model. Background Technology

[0002] Multilevel text classification tasks involve analyzing and understanding text content and classifying it according to the hierarchical relationship of multiple categories. This is of great significance in scenarios such as text risk identification and intention recognition of in-vehicle voice assistant commands.

[0003] In related technologies, a label system tree corresponding to a multi-level text classification task can be expanded according to the leaf nodes, thereby transforming the multi-level text classification task into a single-level text classification task, which can then be processed by a multi-classification model.

[0004] However, the above processing method is not very accurate for multi-level text classification tasks with complex label systems. Summary of the Invention

[0005] In view of this, this application provides a method and related apparatus for training a generative model, which can perform multi-level text classification tasks.

[0006] To solve the above problems, the technical solution provided in this application is as follows:

[0007] On the one hand, this application provides a method for training a generative model, the method comprising:

[0008] Acquire label system data and training samples. The label system data is used to identify the hierarchical relationship of multiple sample labels. The training samples have sample labels used to identify the true category of the training samples.

[0009] The label system data, task prompts, and training samples are input into the initial generative model to obtain the classification prediction results of the training samples;

[0010] Based on the difference between the classification prediction results and the sample labels, the initial generative model is trained to obtain a classification generative model.

[0011] In one possible implementation, the initial generative model is obtained as follows:

[0012] The text content of the training samples is segmented into words to obtain multiple segmented fragments;

[0013] For the first segment of the plurality of segmented segments, a character is selected in the first segmented segment as a reference character, and at least one character after the reference character is used as a target character for masking processing to obtain a segmented mask segment.

[0014] The word segmentation mask fragment is input into the original generative model, and the original generative model predicts the masked characters to generate predicted characters.

[0015] Based on the difference between the predicted character and the target character, the original generative model is pre-trained to obtain the initial generative model.

[0016] In one possible implementation, selecting a character as a reference character in the first segmented word fragment and performing masking processing on at least one character following the reference character as a target character to obtain a segmented word mask fragment includes:

[0017] Determine a string with complete semantics from the first segmented word fragment, the string comprising N characters;

[0018] Based on the character order of the first segmented word, select a character from the first N-1 characters of the string as a reference character, and determine the target character from the characters in the string after the reference character;

[0019] The target characters are masked to obtain the word segmentation mask fragment.

[0020] In one possible implementation, the step of inputting the label system data, task prompt words, and training samples into an initial generative model to obtain the classification prediction results of the training samples includes:

[0021] Model input data is generated based on the tag system data, the task prompt words, and the training samples. In the model input data, the task prompt words are configured with corresponding first prompts, which are used to identify the corresponding content as the task prompt words. The training samples are configured with corresponding second prompts, which are used to identify the corresponding content as the training samples.

[0022] The model input data is input into the initial generative model to obtain the classification prediction results of the training samples.

[0023] In one possible implementation, the classification prediction result includes at least one predicted category, which has a hierarchy identifier to indicate the label hierarchy of the predicted category in the label system data.

[0024] In one possible implementation, the sample label includes labels at least two label levels.

[0025] In another aspect, this application provides a classification method, the method comprising:

[0026] The label system data, task prompts, and text to be classified are input into the classification generative model to obtain the classification prediction result of the text to be classified. The label system data is used to identify the hierarchical relationship of multiple sample labels.

[0027] In another aspect, this application provides a training apparatus for a generative model, the apparatus comprising an acquisition unit, a prediction unit, and a training unit:

[0028] The acquisition unit is used to acquire label system data and training samples. The label system data is used to identify the hierarchical relationship of multiple sample labels. The training samples have sample labels, which are used to identify the true category of the training samples.

[0029] The prediction unit is used to input the label system data, task prompt words and training samples into the initial generative model to obtain the classification prediction result of the training samples;

[0030] The training unit is used to train the initial generative model based on the difference between the classification prediction result and the sample label to obtain a classification generative model.

[0031] In one possible implementation, the apparatus further includes a pre-training unit, the pre-training unit being used for:

[0032] The text content of the training samples is segmented into words to obtain multiple segmented fragments;

[0033] For the first segment of the plurality of segmented segments, a character is selected in the first segmented segment as a reference character, and at least one character after the reference character is used as a target character for masking processing to obtain a segmented mask segment.

[0034] The word segmentation mask fragment is input into the original generative model, and the original generative model predicts the masked characters to generate predicted characters.

[0035] Based on the difference between the predicted character and the target character, the original generative model is pre-trained to obtain the initial generative model.

[0036] In one possible implementation, the pre-trained unit is specifically used for:

[0037] Determine a string with complete semantics from the first segmented word fragment, the string comprising N characters;

[0038] Based on the character order of the first segmented word, select a character from the first N-1 characters of the string as a reference character, and determine the target character from the characters in the string after the reference character;

[0039] The target characters are masked to obtain the word segmentation mask fragment.

[0040] In one possible implementation, the prediction unit is specifically used for:

[0041] Model input data is generated based on the tag system data, the task prompt words, and the training samples. In the model input data, the task prompt words are configured with corresponding first prompts, which are used to identify the corresponding content as the task prompt words. The training samples are configured with corresponding second prompts, which are used to identify the corresponding content as the training samples.

[0042] The model input data is input into the initial generative model to obtain the classification prediction results of the training samples.

[0043] In one possible implementation, the classification prediction result includes at least one predicted category, which has a hierarchy identifier to indicate the label hierarchy of the predicted category in the label system data.

[0044] In one possible implementation, the sample label includes labels at least two label levels.

[0045] In another aspect, this application provides a classification device, the device comprising a classification subunit:

[0046] The classification subunit is used to input the label system data, task prompt words and the text to be classified into the classification generative model to obtain the classification prediction result of the text to be classified. The label system data is used to identify the hierarchical relationship of multiple sample labels.

[0047] In another aspect, this application provides a computer device, which includes a processor and a memory:

[0048] The memory is used to store computer programs;

[0049] The processor is configured to execute the method described in any of the above-described embodiments according to the computer program.

[0050] In another aspect, this application provides a computer-readable storage medium for storing a computer program that, when executed by a computer device, implements the method described in any of the above-mentioned embodiments.

[0051] In another aspect, this application provides a computer program product including a computer program, which, when run on a computer device, causes the computer device to perform any of the methods described above.

[0052] As can be seen from the above technical solution, this solution first acquires label system data and training samples. The label system data describes the hierarchical relationship between multiple sample labels. This hierarchical relationship is used to identify sample labels that are related between two adjacent levels, including the sample label of the i-th level and one or more sub-labels of the sample label in the (i+1)-th level. The sample labels of the training samples can identify the true category of the training samples. Then, the label system data, task prompts, and training samples are input into the initial generative model. The initial generative model can determine its corresponding predicted category by learning the semantic features of the training samples. Thus, the classification prediction result corresponding to the training samples is obtained through the initial generative model. Finally, the initial generative model is trained by the difference between the classification prediction result and the sample labels to obtain the classification generative model. Therefore, the classification generative model can realize multi-level text classification and determine the category label according to the inherent relationship between sample labels of adjacent levels, effectively improving the accuracy and efficiency of classification. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A flowchart illustrating a training method for a generative model provided in an embodiment of this application;

[0055] Figure 2 This is a schematic diagram of a labeling system data provided in an embodiment of this application;

[0056] Figure 3 A model architecture diagram of a generative model provided in an embodiment of this application;

[0057] Figure 4 A flowchart illustrating the training process of a generative model provided in an embodiment of this application;

[0058] Figure 5 This is a schematic diagram of a training device for a generative model provided in an embodiment of this application. Detailed Implementation

[0059] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0060] As described in the background section, related technologies typically expand the label tree corresponding to the multi-level labels into multiple single-level labels by leaf nodes for multi-level text classification tasks, thus transforming the multi-level text classification task into a single-level text classification task, which is then processed by a multi-classification model. However, in multi-level text classification tasks with complex label systems, classifying hundreds or thousands of single-level labels results in too many classification categories for the multi-classification model, leading to low classification accuracy.

[0061] This application provides a training method for a generative model. First, it acquires label system data and training samples. The label system data describes the intrinsic relationships between multiple sample labels at adjacent levels, and the sample labels of the training samples can identify the true category of the training samples. Then, the label system data, task prompts, and training samples are input into an initial generative model to predict the true category of the training samples. By comparing the difference between the classification prediction results and the sample labels, the initial generative model is trained to obtain a classification generative model. This classification generative model can achieve multi-level text classification and determine the category labels based on the intrinsic relationships between sample labels at adjacent levels, effectively improving the accuracy and efficiency of classification.

[0062] The solutions provided in this application relate to the field of natural language processing technology, and are specifically illustrated through the following embodiments.

[0063] See Figure 1 The diagram shown illustrates a flowchart of a generative model training method provided in this embodiment. In this embodiment, the method can be executed using a computer device and includes the following steps:

[0064] S101: Obtain the label system data and training samples.

[0065] In this system, the label hierarchy data is used to identify the hierarchical relationship between multiple sample labels. In multi-level text classification tasks, the hierarchical relationship is used to identify the sample label of the i-th level and the sample label of the (i+1)-th level as a sub-label. The label hierarchy data can be set according to the multi-level text classification task. The purpose of multi-level text classification tasks is to match the text categories to a label hierarchy with hierarchical relationships. Different multi-level text classification tasks can correspond to different label hierarchy data, so that models trained based on different label hierarchy data can be applied to different multi-level text classification tasks. Of course, different multi-level text classification tasks can also share label hierarchy data, so that the trained model can be applied to various multi-level text classification tasks.

[0066] The sample labels in the label system data can be expressed in natural language. That is, the sample labels are the true categories obtained by understanding and analyzing the text content and have a certain semantic relationship with the text content. Of course, the sample labels in the label system data can also be expressed by other identifiers. Other identifiers can be indicators corresponding to the true categories of text such as numbers. The model can learn the correspondence between other identifiers and true categories, thereby realizing the learning of the hierarchical relationship between the sample labels.

[0067] This application does not specify the form of the label system data. Taking the application scenario of library management as an example (multi-level text classification task), the text can be classified according to two levels to obtain multiple sample labels, where the sample labels between adjacent levels have an inherent relationship. The following example illustrates the label system data in this application scenario.

[0068] As an example, see reference Figure 2 This is a schematic diagram of a tag system data provided in an embodiment of this application. The tag system data is represented by an image and is shown as multiple tree structures. Each node in the tree is a sample tag, and sample tags at adjacent levels have a hierarchical subordinate structure. A sample tag at the second level is a sub-tag of a sample tag at the first level. For example, in the figure, mathematics, biology, and chemistry at the second level are all sub-tags of natural sciences at the first level.

[0069] As another example, the labeling system data can be represented in textual form: "Level 1 is divided according to subject areas, with sample labels including natural sciences, social sciences, engineering and technology, and humanities. Level 2 is divided according to specific disciplines, with sample labels including mathematics, biology, philosophy, education, computer science, electrical engineering, literature, history, etc. Among them, the sub-labels of natural sciences include mathematics and biology, the sub-labels of social sciences include philosophy and education, the sub-labels of engineering and technology include computer science and electrical engineering, and the sub-labels of humanities include literature and history."

[0070] By using a label system to represent the logical hierarchical relationship between all sample labels involved in a multi-level text classification task, complex texts can be classified according to different levels, and the sample labels corresponding to the text content can be determined, thereby achieving structured management of texts.

[0071] Training samples have corresponding sample labels, which refer to the true categories obtained through understanding and analyzing the training samples. They have a certain semantic relationship with the text content and can be single-level or multi-level sample labels. There is a hierarchical relationship between multi-level sample labels.

[0072] The sample labels can be in the form of text expressed in natural language, such as natural sciences, social sciences, or mathematics, or they can be identifiers such as numbers or symbols to refer to the real category, such as 1, 2, 3, etc., where 1 refers to natural sciences, 2 refers to social sciences, and 3 refers to mathematics. Therefore, in the example of the above label system data, the sample labels can be in the above form expressed in natural language, or they can be any other form of content, as long as they can establish a referential relationship with the real category of the label. The embodiments of this application will not be described in detail here.

[0073] S102: Input the label system data, task prompts, and training samples into the initial generative model to obtain the classification prediction results of the training samples.

[0074] Since the sample labels corresponding to the training samples have a certain semantic relationship with the training samples, the initial generative model can understand the semantic information of the training samples, so that its output content is related or consistent with the sample labels, thus realizing the task of performing text classification.

[0075] The task prompts can be used to instruct the initial generative model to classify training samples hierarchically, based on the hierarchical relationships identified by the label system data. Alternatively, the task prompts may not instruct the model to classify training samples hierarchically; in this case, the model can learn the hierarchical classification method during subsequent training.

[0076] In one possible implementation, the sample labels include labels at least two label levels. In this case, the initial generative model is used to perform a multi-level text classification task. In the process of predicting the text category, based on the category at the i-th level, one or more sample labels at the i+1-th level that have a hierarchical relationship with the category can be determined through this hierarchical relationship. This allows the initial generative model to further determine the category of the training samples within the range of these one or more sample labels, narrowing the prediction range of the category at the i+1-th level, thereby obtaining classification prediction results more efficiently and accurately.

[0077] The classification prediction result is obtained by predicting the true category of the training samples through an initial generative model. Its classification accuracy cannot be guaranteed, therefore, the predicted category may be inconsistent with the corresponding sample label of the training sample. This application embodiment illustrates the potential inconsistency between the predicted category and the sample label in a label system where multiple sample labels have multiple label levels.

[0078] As an example, a text content in the training samples has three levels of labels: "Natural Sciences, Mathematics, and Basic Mathematics". However, the corresponding classification prediction result is "Natural Sciences and Mathematics", which only includes the first two label levels. This may be due to the model's insufficient capability, resulting in the failure to refine the classification to the third level.

[0079] As another example, in the training samples, the sample label corresponding to a text content is "natural science, mathematics" with two levels, while the model obtains the classification result "natural science, mathematics, basic mathematics" based on the data with a three-level label system. This may be due to the model being over-refined.

[0080] S103: Based on the difference between the classification prediction results and the sample labels, train the initial generative model to obtain a classification generative model.

[0081] Since the classification prediction results obtained from the initial generative model cannot guarantee their accuracy, the sample labels of the training samples can guide the initial generative model and improve the classification accuracy.

[0082] Since both the training sample labels and the classification prediction results include at least two label levels of true categories, it is necessary to compare the differences between the true categories and the predicted categories at each level. Then, the parameters of the initial generative model are adjusted so that its output classification prediction results are closer to the sample labels. In other words, through continuous training, the classification accuracy of the initial generative model is continuously improved, resulting in a classification generative model with high classification accuracy. When the training sample labels correspond to at least two label levels, training the initial generative model can guide it towards outputting prediction categories corresponding to multiple label levels, thus achieving a prediction category output corresponding to multiple label levels.

[0083] In one possible implementation, the generative model has N sequentially connected decoding layers, where N is greater than 1 and the value of N is determined based on the classification accuracy of the multi-level text classification task.

[0084] The core of this generative classification model can be a Transformer Decoder, comprising multiple TransformerDecoderLayers. The output of each TransformerDecoderLayer serves as the input to the next layer. As the number of decoding layers increases, the model's classification accuracy initially rises and then plateaus or even declines. The number of decoding layers can be adjusted based on the classification accuracy required for the task. For example, by analyzing the relationship between the number of decoding layers and classification accuracy, it was found that 12 TransformerDecoderLayers resulted in the highest classification accuracy. Therefore, this generative classification model is determined to have 12 sequentially connected TransformerDecoderLayers. Thus, by understanding the relationship between classification accuracy and the number of model layers, the number of layers can be adjusted based on the model's classification accuracy, thereby ensuring the classification accuracy for multi-level text classification tasks.

[0085] Therefore, by inputting the label system data, task prompts, and training samples into the initial generative model, the generative model can learn the semantic features of the training samples and obtain the classification prediction results corresponding to the training samples. Finally, by using the difference between the classification prediction results and the sample labels of the corresponding training samples, the initial generative model can be trained to obtain a classification generative model. Since the label system data describes the inherent relationship between sample labels at adjacent levels, it can guide the classification generative model to classify text level by level, effectively improving the accuracy and efficiency of multi-level text classification.

[0086] The initial generative model has a certain text understanding ability. In order to make it have a stronger text understanding ability in multi-level text classification tasks, it can be pre-trained based on the training samples of multi-level text classification tasks.

[0087] In one possible implementation, the initial generative model is obtained as follows:

[0088] B1: Perform word segmentation on the text content of the training samples to obtain multiple word segments.

[0089] This application does not impose any restrictions on the specific method of word segmentation. For example, it can be rule-based word segmentation, statistical word segmentation, deep learning-based word segmentation, or any other method that can break down text content into words.

[0090] B2: For the first segment of multiple segmented segments, select a character in the first segmented segment as the base character, and perform masking on at least one character after the base character as the target character to obtain the segmented masked segment.

[0091] Characters can be words, letters, or symbols. Since generative models predict subsequent information based on preceding information, to improve text understanding, the first character of each segment is generally not masked. Instead, for each segment, a character is selected as the reference character, and at least one character after the reference character is masked to avoid masking the first character of the segment. This results in multiple masked segment segments.

[0092] B3: Input the segmented mask fragment into the original generative model, and use the original generative model to predict the masked characters and generate predicted characters.

[0093] A rudimentary generative model refers to a basic model architecture that has not undergone any training, possesses the potential to generate text, but has not learned any domain-specific knowledge, i.e., any knowledge specific to multi-level text classification tasks. This application does not impose any restrictions on the model architecture of rudimentary generative models; for example, it can be an autoregressive model architecture, an autoencoder model architecture, a Transformer architecture, or a hybrid architecture, etc.

[0094] By inputting multiple segmented mask fragments into the original generative model, the masked characters in each segmented fragment can be predicted, thereby generating corresponding predicted characters for each segmented fragment.

[0095] B4: Based on the difference between the predicted character and the target character, the original generative model is pre-trained to obtain the initial generative model.

[0096] The predicted character and its corresponding masked character are compared. Based on the difference between the two, the parameters of the original generative model are adjusted so that the predicted character output is closer to the corresponding masked character, thus obtaining the pre-trained initial generative model.

[0097] This maximizes the probability of the masked characters appearing in the output, enabling the initial generative model to have a certain text understanding ability for multi-level text classification tasks, thus allowing it to make preliminary classification predictions of the text content.

[0098] In one possible implementation, step B2 can be performed through the following steps:

[0099] C1: Determine a string with complete semantics from the first segmented word, the string consisting of N characters.

[0100] A string with complete semantics refers to a string that can express a complete meaning without relying on context. For example, it can be a character, word, or sentence. Part-of-speech analysis can identify strings with complete semantics in word segmentation fragments. These strings consist of N characters.

[0101] C2: Based on the character order of the first segment, select a character from the first N-1 characters of the string as the reference character, and determine the target character from the characters in the string after the reference character.

[0102] Based on the character order of the segmented fragments, select a character from the first N-1 characters of the string as the reference character, and then determine the target character from the characters following the reference character in the string, so that the target character can be determined in a string with complete semantics.

[0103] C3: Mask the target characters to obtain the word segmentation mask fragment.

[0104] Therefore, in word segmentation, selecting characters for masking from strings with complete semantics can further optimize the initial generative model's text understanding ability in multi-level text classification tasks.

[0105] In one possible implementation, step S102 includes:

[0106] D1: Generate model input data based on the label system data, task prompts, and training samples.

[0107] By setting the input format of the initial generative model, the label system data, task prompts, and training samples can be combined into the model input data, and thus input as a whole.

[0108] In the model input data, task prompts are configured with corresponding first prompts, which identify the corresponding content as task prompts. Training samples are configured with corresponding second prompts, which identify the corresponding content as training samples. For example, the first prompt is... <prompt>It can be configured before the task prompt word, and the second prompt is... <user>It can be configured in front of the training samples.

[0109] D2: Input the model input data into the initial generative model to obtain the classification prediction results of the training samples.

[0110] Therefore, the label system data, task prompts, and training samples can be input into the initial generative model as a whole at once. This strengthens the relationship between the input parts and allows for accurate identification of the content corresponding to each input part through identifiers. This helps the model better understand the task intent, thereby reducing the difficulty of model understanding and improving the training efficiency of the model.

[0111] In one possible implementation, the classification prediction result includes at least one predicted category, that is, it may include one or more predicted categories. The predicted category has a hierarchy identifier to indicate its label hierarchy within the label system data.

[0112] As an example, the classification prediction result is "Natural Sciences, Mathematics", which has two label levels. Natural Sciences is the sample label for level 1, and Mathematics is the sample label for level 2. The level identifier for level 1 is... <lv1>The level identifier for level 2 is <lv2>Then the classification prediction result can be " <lv1>Natural Sciences <lv2>math".

[0113] Therefore, the hierarchical identifier clarifies the hierarchy of each predicted category, serving as a structured cue to help improve the interpretability of the model, thereby enabling the model to learn and generate classification prediction results better.

[0114] To more clearly describe the training method of this generative model, the following explanation will be provided in conjunction with a specific implementation scenario. See [link / reference] Figure 3 The diagram shown is a model architecture diagram of a generative model provided in this application. The main body of the generative model is a Transformer Decoder, which consists of 12 TransformerDecoderLayer layers.

[0115] The model's input consists of two parts: task prompts and user input, organized into a text segment with the task prompts preceding the user input. To distinguish between these two parts, two special words are added to the input. <prompt>"and" <user>The phrases "task prompt" and "user input" need to be added before the task prompt and user input, respectively. The task prompt refers to a set of instructions abstracted by the developer based on a specific task. Taking a library management application scenario as an example, it could be designed as "You are a multi-level text classification model; please provide the correct answer based on the user input." In this case, the model's input is "...". <prompt>You have a multi-level text classification model. Please provide the correct answer based on the user input. <user>Text to be classified.

[0116] The model's output is the predicted Chinese representation of the sample labels, expressed through several special words. <lvx>The model is divided into levels. "x" represents 1, 2, 3…, indicating that the sample label belongs to level 1, level 2, level 3… of the label system data. If the first-level label is "Natural Sciences" and the second-level label is "Mathematics", then the model output is… <lv1>Natural Sciences <lv2>mathematics".

[0117] In order to maximize the accuracy of the generative model on classification tasks, the training of the generative model consists of two stages. See Figure 4 , which is a flow chart of training a generative model provided by the embodiments of the present application.

[0118] First stage: pre-training stage.

[0119] Developers first need to collect a large amount of corpora related to the multi-level text classification task and some general corpora. Subsequently, according to the actual situation of the multi-level text classification task, these texts are organized into training samples according to sentences or paragraphs.

[0120] During the pre-training process, a sequence can be received as input, and a sequence can be generated as output. For each sample, a character in the sample is randomly selected, the text segment after the character is discarded, and the text segment before the character is used as input to maximize the probability that this character appears in the output. By repeatedly selecting training samples and performing training operations, the probability of occurrence of the selected characters in all samples is maximized. For text content, the input sequence and output sequence can be divided according to characters. Taking "good morning" as an example, the input sequence is "good", "morning", and the expected output sequence is "ing".

[0121] The purpose of this step is to enable the initial generative model to learn the ability to understand and generate text content on large-scale text, and lay a foundation for the next stage of training.

[0122] Second stage: fine-tuning stage with a small number of labeled samples.

[0123] In this stage, developers need to combine multiple sample labels involved in the multi-level text classification task, collect a certain amount of text for each leaf node of the label system tree corresponding to multiple sample labels to construct training samples. For each training sample, the input is " <prompt>Task prompt <user>The "sample text" indicates that the task prompt instructs the model to classify the text hierarchically, label by label, based on the hierarchical relationship between multiple sample labels in a multi-level text classification task, and outputs the expanded result of its predicted sample labels. <lv1>Level 1 tags <lv2>Secondary labels…”. During training, for each sample, the training objective is to maximize the overall probability of the output occurring given the input.

[0124] Thus, this generative model can classify text content with complex tag systems with less development effort, enabling it to perform multi-level text classification tasks. Furthermore, through two-stage training, it can achieve or even surpass the combined effect of multiple classification models in terms of classification accuracy.

[0125] Based on the above embodiments, this application also provides a classification method, the execution steps of which are as follows:

[0126] By inputting the tag system data, task prompts, and the text to be classified into the classification generative model, the classification prediction results of the text to be classified are obtained.

[0127] Among them, classification generative models refer to generative models that can be used for multi-level text classification tasks, such as the classification generative models obtained by the above-mentioned generative model training methods.

[0128] The label system data is used to identify the hierarchical relationship of multiple sample labels. The text to be classified refers to the text content to be classified for the text classification task. This text content may have sample labels, including one or more levels.

[0129] Task prompts are abstract instructions based on label system data and specific tasks. They are used to instruct the classification generative model to classify the samples to be classified according to the hierarchical relationship identified by the label system data, thereby obtaining the classification result of the samples to be classified.

[0130] Therefore, the inherent relationship between sample labels at adjacent levels described by the label system data can instruct the generative model to predict the category of the sample to be classified, thereby achieving higher classification accuracy.

[0131] Based on the above embodiments, this application provides a training device for generative models, with reference to... Figure 5 This is a schematic diagram of a generative model training device provided in an embodiment of this application. The device 500 includes an acquisition unit 501, a prediction unit 502, and a training unit 503.

[0132] The acquisition unit is used to acquire label system data and training samples. The label system data is used to identify the hierarchical relationship of multiple sample labels. The training samples have sample labels, which are used to identify the true category of the training samples.

[0133] The prediction unit is used to input the label system data, task prompt words and training samples into the initial generative model to obtain the classification prediction result of the training samples;

[0134] The training unit is used to train the initial generative model based on the difference between the classification prediction result and the sample label to obtain a classification generative model.

[0135] Therefore, the resulting generative model can be used for multi-level text classification tasks. Based on the inherent relationship between sample labels in adjacent levels, the generative model can achieve a hierarchical classification method, which effectively improves the accuracy and efficiency of classification in multi-level text classification tasks with complex label systems.

[0136] In one possible implementation, the apparatus further includes a pre-training unit, the pre-training unit being used for:

[0137] The text content of the training samples is segmented into words to obtain multiple segmented fragments;

[0138] For the first segment of the plurality of segmented segments, a character is selected in the first segmented segment as a reference character, and at least one character after the reference character is used as a target character for masking processing to obtain a segmented mask segment.

[0139] The word segmentation mask fragment is input into the original generative model, and the original generative model predicts the masked characters to generate predicted characters.

[0140] Based on the difference between the predicted character and the target character, the original generative model is pre-trained to obtain the initial generative model.

[0141] This maximizes the probability of the masked characters appearing in the output, enabling the initial generative model to have a certain text understanding ability for multi-level text classification tasks, thus allowing it to make preliminary classification predictions of the text content.

[0142] In one possible implementation, the pre-trained unit is specifically used for:

[0143] Determine a string with complete semantics from the first segmented word fragment, the string comprising N characters;

[0144] Based on the character order of the first segmented word, select a character from the first N-1 characters of the string as a reference character, and determine the target character from the characters in the string after the reference character;

[0145] The target characters are masked to obtain the word segmentation mask fragment.

[0146] Therefore, masking within a string with complete semantics in word segmentation can further optimize the initial generative model's text understanding capabilities for multi-level text classification tasks.

[0147] In one possible implementation, the prediction unit is specifically used for:

[0148] Model input data is generated based on the tag system data, the task prompt words, and the training samples. In the model input data, the task prompt words are configured with corresponding first prompts, which are used to identify the corresponding content as the task prompt words. The training samples are configured with corresponding second prompts, which are used to identify the corresponding content as the training samples.

[0149] The model input data is input into the initial generative model to obtain the classification prediction results of the training samples.

[0150] Therefore, by simultaneously inputting the label system data, task prompts, and training samples as a whole into the initial generative model, the relationship between the various parts of the input can be strengthened. Furthermore, the identifiers can accurately identify the content corresponding to each part of the input, helping the model to better understand the task intent, thereby reducing the difficulty of model understanding and improving the training efficiency of the model.

[0151] In one possible implementation, the classification prediction result includes at least one predicted category, which has a hierarchy identifier to indicate the label hierarchy of the predicted category in the label system data.

[0152] Therefore, the hierarchical identifier clarifies the hierarchy of each predicted category, serving as a structured cue to help improve the interpretability of the model, thereby enabling the model to learn and generate classification prediction results better.

[0153] In one possible implementation, the sample label includes labels at least two label levels.

[0154] Therefore, this generative classification model can achieve hierarchical classification based on the hierarchical relationship of sample labels identified by the label system data, thereby improving the classification accuracy of text classification tasks.

[0155] Based on the above embodiments, this application provides a classification device, which includes a classification subunit:

[0156] The classification subunit is used to input the label system data, task prompt words and the text to be classified into the classification generative model to obtain the classification prediction result of the text to be classified. The label system data is used to identify the hierarchical relationship of multiple sample labels.

[0157] Based on the above embodiments, this application provides a computer device, which includes a processor and a memory:

[0158] The memory is used to store computer programs;

[0159] The processor is used to execute the training method of the generative model described above according to the computer program.

[0160] Based on the above embodiments, this application provides a computer-readable storage medium for storing a computer program that, when executed by a computer device, implements the training method for the above-described generative model.

[0161] Based on the above embodiments, this application provides a computer program product including a computer program, which, when run on a computer device, causes the computer device to execute the above-described generative model training method.

[0162] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

[0163] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. < / user> < / prompt> < / lvx> < / user> < / prompt> < / user> < / prompt> < / user> < / prompt>

Claims

1. A method for training a generative model, characterized in that, The method includes: Acquire label system data and training samples. The label system data is used to identify the hierarchical relationship of multiple sample labels. The training samples have sample labels used to identify the true category of the training samples. The label system data, task prompts, and training samples are input into the initial generative model to obtain the classification prediction results of the training samples; Based on the difference between the classification prediction results and the sample labels, the initial generative model is trained to obtain a classification generative model.

2. The method according to claim 1, characterized in that, The initial generative model is obtained in the following way: The text content of the training samples is segmented into words to obtain multiple segmented fragments; For the first segment of the plurality of segmented segments, a character is selected in the first segmented segment as a reference character, and at least one character after the reference character is used as a target character for masking processing to obtain a segmented mask segment. The word segmentation mask fragment is input into the original generative model, and the original generative model predicts the masked characters to generate predicted characters. Based on the difference between the predicted character and the target character, the original generative model is pre-trained to obtain the initial generative model.

3. The method according to claim 2, characterized in that, The step of selecting a character as a reference character in the first segmented word segment and performing masking processing on at least one character following the reference character as a target character to obtain a segmented word mask segment includes: Determine a string with complete semantics from the first segmented word fragment, the string comprising N characters; Based on the character order of the first segmented word, select a character from the first N-1 characters of the string as a reference character, and determine the target character from the characters in the string after the reference character; The target characters are masked to obtain the word segmentation mask fragment.

4. The method according to claim 1, characterized in that, The step of inputting the label system data, task prompt words, and training samples into the initial generative model to obtain the classification prediction results of the training samples includes: Model input data is generated based on the tag system data, the task prompt words, and the training samples. In the model input data, the task prompt words are configured with corresponding first prompts, which are used to identify the corresponding content as the task prompt words. The training samples are configured with corresponding second prompts, which are used to identify the corresponding content as the training samples. The model input data is input into the initial generative model to obtain the classification prediction results of the training samples.

5. The method according to claim 1, characterized in that, The classification prediction result includes at least one predicted category, and the predicted category has a hierarchy identifier to indicate the label hierarchy of the predicted category in the label system data.

6. The method according to any one of claims 1-5, characterized in that, The sample labels include labels at least two label levels.

7. A classification method, characterized in that, The method includes: The label system data, task prompts, and text to be classified are input into the classification generative model to obtain the classification prediction result of the text to be classified. The label system data is used to identify the hierarchical relationship of multiple sample labels.

8. A training device for a generative model, characterized in that, The device includes an acquisition unit, a prediction unit, and a training unit: The acquisition unit is used to acquire label system data and training samples, wherein the label system data is used to identify the hierarchical relationship of multiple sample labels; The training samples have sample labels used to identify the true category of the training samples; The prediction unit is used to input the label system data, task prompt words and training samples into the initial generative model to obtain the classification prediction result of the training samples; The training unit is used to train the initial generative model based on the difference between the classification prediction result and the sample label to obtain a classification generative model.

9. A sorting device, characterized in that, The device includes a classification subunit: The classification subunit is used to input the label system data, task prompt words and the text to be classified into the classification generative model to obtain the classification prediction result of the text to be classified.

10. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs; The processor is configured to perform the method according to any one of claims 1-7 according to the computer program.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when executed by a computer device, performs the method described in any one of claims 1-7.

12. A computer program product comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the method described in any one of claims 1-7.