Text risk classification system training method and device

Through the architecture of the base model and independent classification head, the problem of risk category changes in text risk control is solved, and a text risk classification effect with rapid adaptation and resource saving is achieved.

CN120781079APending Publication Date: 2025-10-14ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510896848.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing text risk control solutions are unable to flexibly respond to the continuous changes in risk categories, which affects the model effectiveness and causes serious waste of computing resources.

Method used

Adopting the architecture of base model + independent classification head, we respond to changes in risk categories by adding, deleting and modifying classification heads, train K independent classification heads, and optimize the model configuration using neural network architecture search and hyperparameter optimization techniques.

Benefits of technology

It can quickly adapt to changes in risk categories, ensuring classification effectiveness while saving computing resources, and improving the flexibility and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120781079A_ABST
    Figure CN120781079A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a training method and device for a text risk classification system and a training method and device for a business object classification system. The text risk classification system comprises a base model and K classification heads corresponding to K risk categories, and the training method of the text risk classification system comprises the steps that firstly, a training text set is acquired, and any first training text is associated with a first text label indicating the risk category to which the first training text belongs; then, inputting the first training text into the base model to obtain a first output; then, respectively inputting the first output into the K classification heads to obtain K classification probabilities corresponding to the output; and then, based on the K classification probabilities and the first text label, training the K classification heads. A text risk classification system trained according to the method can quickly adapt to changes of risk category labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of risk control technology, and in particular, to the application of artificial intelligence in the field of text risk control. Background Art

[0002] With the rapid development of internet technology, content recommendation platforms have become a core vehicle for information distribution, widely used in social media, news, e-commerce, and other fields. Text recommendation services, a key technology in this field, analyze user interests and behaviors to push personalized text content (such as articles, comments, and advertisements) to users, enhancing the user experience.

[0003] However, during the text recommendation process, platforms must process massive amounts of user-generated and third-party content, which may contain high-risk information (such as sensitive topics, false news, and illegal advertising). Without effective risk control, recommending such content to users could negatively impact user experience and raise legal compliance issues. To ensure the security of the content ecosystem, text risk control (or risk management) is crucial.

[0004] Therefore, an improved text risk control solution is needed that can better meet actual application needs, such as flexible deployment and the ability to adapt to the diverse and rapidly changing scenarios in the risk control field. Summary of the Invention

[0005] The embodiments of this specification describe a training method and device for a text risk classification system, which can quickly respond to changes such as the addition, deletion, and modification of labels in text risk control scenarios, thereby better meeting actual application needs.

[0006] According to a first aspect, a training method for a text risk classification system is provided. The text risk classification system includes a base model and K classification heads corresponding to K risk categories, and the method includes:

[0007] A training text set is obtained, wherein any first training text is associated with a first text label indicating the risk category to which it belongs; the first training text is input into the base model to obtain a first output; the first output is respectively input into the K classification heads to obtain K classification probabilities of the corresponding outputs; and the K classification heads are trained based on the K classification probabilities and the first text label.

[0008] In some embodiments, the first text label indicates that the first training text belongs to one or more risk categories.

[0009] In some embodiments, while training the K classification heads, the base model is also fine-tuned.

[0010] In some embodiments, after training the K classification heads, the method further includes: deleting a first classification head from the text risk classification system, which corresponds to a first risk category among the K risk categories.

[0011] In some embodiments, after training the K classification heads, the method further includes: adding a second classification head to the text risk classification system, which corresponds to a second risk category newly added on the basis of the K risk categories; obtaining a second training text whose associated second text label indicates that it belongs to the second risk category; and using the second training sample to train the second classification head.

[0012] In some embodiments, the K risk categories include a third risk category; wherein, after training the K classification heads, the method further includes: obtaining a third training text, whose associated third text label indicates that it belongs to the third risk category, and the meaning of the category has changed compared to when training the K classification heads; using the third training sample, retraining the third classification head corresponding to the third risk category.

[0013] In some embodiments, the types and structures of the K classification heads are the same, or there are two classification heads with different types or structures among the K classification heads.

[0014] In some embodiments, during the training process of the text risk classification system, at least one of the following configurations of the text risk classification system is determined by adopting a neural network architecture search technique: the type of the base model, the type of each of the K classification heads, and the structure of each classification head.

[0015] In some embodiments, the text risk classification system is trained using hyperparameter optimization techniques.

[0016] According to a second aspect, a method for training a business object classification system is provided. The classification system includes a base model and K classification heads corresponding to K object categories. The method includes:

[0017] A training sample set is obtained, wherein any first training sample includes a first object feature and a first object label of a first business object, where the label indicates the object category to which the first business object belongs; the first object feature is input into the base model to obtain a first output; the first output is respectively input into the K classification heads to obtain K classification probabilities of the corresponding outputs; and the K classification heads are trained based on the K classification probabilities and the first object label.

[0018] According to a third aspect, a training device for a text risk classification system is provided, wherein the text risk classification system includes a base model and K classification heads corresponding to K risk categories. The device includes:

[0019] A text set acquisition unit is configured to acquire a training text set, wherein any first training text is associated with a first text label indicating the risk category to which it belongs; a text representation unit is configured to input the first training text into the base model to obtain a first output; a classification prediction unit is configured to input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; and a training unit is configured to train the K classification heads based on the K classification probabilities and the first text label.

[0020] According to a fourth aspect, a training device for a business object classification system is provided, wherein the business object classification system includes a base model and K classification heads corresponding to K object categories. The device includes:

[0021] A sample set acquisition unit is configured to acquire a training sample set, wherein any first training sample includes a first object feature and a first object label of a first business object, and the label indicates the object category to which the first business object belongs; a sample characterization unit is configured to input the first object feature into the base model to obtain a first output; a classification prediction unit is configured to input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; and a training unit is configured to train the K classification heads based on the K classification probabilities and the first object label.

[0022] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method provided in the first aspect.

[0023] According to a sixth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method provided in the first aspect is implemented.

[0024] In summary, by adopting the above-mentioned method and device disclosed in the embodiments of this specification, on the basis of two-stage training using the base model + classification head, independent classification heads are trained for K risk categories. When the risk category changes, the changes can be quickly responded to by adding, deleting and modifying the classification heads, thereby effectively saving computing resources while ensuring the risk classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 Schematic diagram of the structure and data processing of the text risk classification system with a single classification head;

[0027] Figure 2 This is a schematic diagram of the structure and data processing of the text risk classification system with multiple classification heads disclosed in the embodiments of this specification;

[0028] Figure 3 This is a schematic diagram of the training process of the text risk classification system disclosed in the embodiments of this specification;

[0029] Figure 4 A schematic diagram of the training process of the business object classification system disclosed in the embodiments of this specification;

[0030] Figure 5 This is a schematic diagram of the structure of a training device for a text risk classification system disclosed in an embodiment of this specification;

[0031] Figure 6 This is a structural diagram of a training device for a business object classification system disclosed in an embodiment of this specification. DETAILED DESCRIPTION

[0032] Next, the solution provided in this specification is described with reference to the accompanying drawings.

[0033] As mentioned previously, an improved text risk control solution is needed that can better meet practical application needs. Specifically, for text risk classification in risk control scenarios, there are generally two types of solutions: training a small model from scratch, and fine-tuning a pre-trained encoder-based language model (such as BERT). For scenarios where the risk category remains fixed, both solutions generally achieve good classification results.

[0034] However, in actual text risk control scenarios, the number, types, and semantics of risk categories are subject to continuous change. In this case, the above two solutions cannot respond flexibly, and there is a possibility that the model effect will be affected after introducing new data and retraining.

[0035] Based on the above observations and analysis, this specification embodies an improved text risk control solution. This improved solution proposes a classification architecture using a base model + classification head, and trains a separate classification head for each risk category. This means that the multiple classification heads corresponding to different risk categories are independent of each other. Therefore, changes in risk categories can be quickly responded to by simply adding, deleting, or modifying the classification heads.

[0036] It should be understood that the above-mentioned base model, also known as the backbone network or feature extractor, is the core feature extraction module, which is responsible for converting the input text into a high-dimensional semantic representation (feature vector), and the classification head performs the final classification or prediction based on these features.

[0037] To help understand, the following Figure 1 and Figure 2 , briefly introduces the classification architecture when there is one or more classification heads, as well as the data flow when using each architecture to process text.

[0038] Figure 1 and Figure 2 The K in refers to the number of risk categories, and K is an integer greater than 1. Figure 1 The figure illustrates the case of a single classification head. The single classification head processes the text x according to the base model to obtain the feature representation, and can obtain the corresponding classification result. This classification result is a K-dimensional probability vector, where the i-th dimension element yi represents the probability that the text x belongs to the i-th risk category. Figure 2 As shown in the figure, in the case of multiple classification heads, the feature representation output by the basic model for the text x is fed into K classification heads respectively, where the i-th classification head only outputs the probability yi that the text x belongs to the i-th risk category.

[0039] The following introduces Figure 2 The training process of the text risk classification system (base model + multiple classification heads) is shown. It should be understood that the execution subject of the training process can be any device, platform, server or equipment cluster with computing and processing capabilities. The training process may include Figure 3 The following steps are shown in:

[0040] Step S310, obtaining a training text set, wherein any first training text is associated with a first text label indicating the risk category to which it belongs; step S320, inputting the first training text into the base model to obtain a first output; step S330, inputting the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; step S340, training the K classification heads based on the K classification probabilities and the first text label.

[0041] The above steps are expanded as follows:

[0042] Firstly, in step S310, a training text set is obtained, wherein any first training text is associated with a first text label indicating a risk category to which the first training text belongs.

[0043] The training text set involves K risk categories, such as sensitive topics, fake news, illegal advertisements, fraudulent information, etc.

[0044] The first training text refers to any one of the training texts in the training text set. It should be noted that the "first" in "first training text" and "first text label" and the "second", "third" and the like in the text are used to distinguish the same things, and do not have other limiting functions such as sorting.

[0045] The above first text label indicates that the first training text belongs to one or more risk categories, for example, both sensitive topics and fake news.

[0046] The above describes the obtained training text and its associated text label.

[0047] Next, in step S320, the first training text is input into the base model to obtain a first output.

[0048] It should be noted that the base model is generally a pre-trained large language model, and exemplarily can be any pre-trained language model shown in Table 1.

[0049] Table 1

[0050]

[0051] In addition, the base model can also be a pre-trained embedding model, such as a BGE (BAAI General Embedding) model, a GTE (General Text Embedding) model, etc.

[0052] In this step, the first training text is processed by the base model to obtain the first output, that is, the feature vector of the first training text, which can also be referred to as text representation, embedding representation, text representation vector, etc.

[0053] Then, in step S330, the first output is input into the K classification heads respectively to obtain K classification probabilities corresponding to the outputs.

[0054] It should be noted that the classification head can be a small neural network with a customized structure, and the types and structures of the classification heads can be the same or different. In some actual business scenarios, configuring classification heads with different types or structures for different risk categories can improve the risk classification accuracy.

[0055] Exemplarily, the type of the classification head can be: fully connected network, recurrent neural network, convolutional neural network, transformer network, etc.; the structure of the classification head includes: the number of neural network layers, the width of each layer (such as the dimension of the weight vector), the size of the convolution kernel, etc.

[0056] In addition, the K classification heads may also be non-neural network machine learning models, such as support vector machines, random forests, gradient boosting trees, etc. It should be understood that the K classification heads may have 1, 2 to K different types (or structures).

[0057] The decision-making processes of the K classification heads are independent of each other. After processing the first output, any i-th classifier outputs the probability that the first document belongs to the i-th risk category. It can be understood that the classification probability ranges from [0, 1].

[0058] From the above, we can obtain the K classification probabilities output by the K classification heads corresponding to the first training text.

[0059] In step S340 , the K classification heads are trained based on the K classification probabilities and the first text label.

[0060] Specifically, the training loss can be calculated based on the K classification probabilities and the first text label, and then the model parameters in the K classification heads can be adjusted based on the training loss.

[0061] It can be understood that the loss function used to calculate the training loss can be selected as needed.

[0062] In a multi-classification scenario (or single-label classification scenario), each training text belongs to only one risk category. For example, the first text label is (0, 1, 0, 0), indicating that the first training text belongs to the second risk category. In this scenario, the loss function can adopt a cross-entropy loss function or label smoothing.

[0063] For example, in a multi-label classification scenario, a training text may belong to more than one risk category. For example, the first text label is (0, 0, 1, 1), indicating that the first training text belongs to both the second and third risk categories. In this scenario, the loss function can be binary cross entropy or FocalLoss.

[0064] It should be noted that when training the K classification heads, the parameters of the base model can be frozen, or the parameters of the base model can be fine-tuned. In addition, when training the above-mentioned text risk classification system, neural architecture search technology and / or hyper-parameter optimization technology can be used to automatically find a configuration that can produce a better classification effect. This configuration may include:

[0065] Types of base models: such as BERT, RoBERTa, DeBERTa, GPT, BGE, GTE, etc.

[0066] Types of classification heads: such as fully connected networks, recurrent neural networks, convolutional neural networks, transformer networks, etc.

[0067] The structure of the classification head: the number of network layers, the width of each layer, the size of the convolution kernel, etc.

[0068] Training hyperparameters: learning rate, batch size, etc.

[0069] Feature-related hyperparameters: the number of hidden states, processing methods, or merging methods.

[0070] For example, after all candidate combinations of base model + K classification heads (such as Bert + K fully connected networks, GPT + K convolutional networks, etc.) are trained, the base model + K classification head combination with excellent performance can be selected for online reasoning by observing the performance indicators (such as accuracy or F1 score).

[0071] As described above, training the text risk classification system using the training text set yields a trained text risk classification system, which can be used for text risk control in real-world business scenarios. It should be noted that when performing online inference, the inference process follows the same training flow as above. The main difference is that online inference only requires obtaining prediction results, without requiring parameter adjustments to the text risk classification system.

[0072] Based on the above-trained text risk classification system, when K risk categories need to be updated due to business needs, such as adding new risk categories, deleting risk categories, or changing the semantics of risk categories, the text risk system can be flexibly adjusted to enable it to quickly adapt to the changed risk categories.

[0073] In a possible case, the first risk category among the K risk categories is deleted. In this case, the first classification head corresponding to the first risk category can be directly deleted from the text risk classification system.

[0074] In one possible scenario, a second risk category is added to the K risk categories. In this case, a second classification head corresponding to this second risk category can be added to the text risk classification system. Furthermore, this second classification head can be trained using training texts that do and do not belong to the second risk category. During this training process, model parameters in the text risk classification system other than the second classification head can be frozen, or other model parameters can be selectively fine-tuned.

[0075] It should be noted that compared with Figure 1 A single classification head in Figure 2 The K classification heads shown in the figure can be implemented as a more lightweight model (such as fewer model parameters, fewer network layers, etc.), which can achieve the same or even better classification results. Therefore, compared with using a large number of samples (it is necessary to consider the global sample distribution and the classification results of all categories), adjusting Figure 1 The improved scheme can achieve the goal of using only a small number of samples to adjust the single classification head Figure 2 A more lightweight classification head (such as the second classification head mentioned above) is used to achieve fast model tuning.

[0076] In one possible scenario, the semantics of the third risk category among the K risk categories change. It should be understood that a semantic change in a risk category means that the definition of the category has been adjusted due to changes in the external environment, rule updates, business needs, etc., causing the original classification system's recognition logic for the category to fail or become biased.

[0077] At this point, some training samples can be obtained, including third training text belonging to the third risk category after the semantic change. These training samples can then be used to retrain the third classification head corresponding to the third risk category. During this training process, the model parameters of the text risk classification system except for the third classification head can be frozen, or other model parameters can be selectively fine-tuned.

[0078] In summary, the training method of the text risk classification system disclosed in the embodiment of this specification is used to train independent classification heads for K risk categories. When the risk category changes, the changes can be quickly responded to by adding, deleting, and modifying the classification heads, thereby effectively saving computing resources while ensuring the risk classification effect.

[0079] The above describes the improved solution disclosed in the embodiments of this specification, and the improved concept of this solution is applied to text risk classification. The applicant proposes that the improved concept can actually be extended to other business objects besides text, such as pictures, videos, audio, payment events, etc., and the model tasks can also be expanded to other classification tasks besides risk classification, such as image recognition for pictures (e.g., identifying pictures containing cars and trees) and thematic classification of videos (such as entertainment or education, etc.).

[0080] Based on this, the embodiment of this specification also discloses a training method for a business object classification system. The business object classification system includes a base model and K classification heads corresponding to K object categories. The execution subject of the training method can be any device, server, platform or device cluster with computing and processing capabilities. The training method may include Figure 4 The following process steps are shown in the figure:

[0081] Step S410, obtaining a training sample set, wherein any first training sample includes a first object feature and a first object label of a first business object, and the label indicates the object category to which the first business object belongs; step S420, inputting the first object feature into the base model to obtain a first output; step S430, inputting the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; step S440, training the K classification heads based on the K classification probabilities and the first object label.

[0082] Regarding the above process steps, it should be noted that the content of the training sample set, the base model, and the selection of the classification head vary in different business scenarios. For example, in the video theme classification scenario, any first training video is associated with a first video tag indicating its theme (such as preschool education or baby games). The base model can be a large pre-trained video-based model such as CLIP-ViT or Florence. The classification head can use a multi-layer perceptron + Dropout, or GAP (Global Average Pooling) + linear layer, etc.

[0083] Need to explain, for Figure 4 For the relevant description of the method steps, including the light adjustment of the business object classification system to adapt to the changes in the object category labels, please refer to the relevant description in the aforementioned embodiment and will not be repeated here.

[0084] In summary, the training method of the business object classification system disclosed in the embodiment of this specification is used to train independent classification heads for K object categories. In business scenarios where object categories continue to change, changes can be quickly responded to by adding, deleting, and modifying classification heads, thereby effectively saving computing resources while ensuring the model classification effect.

[0085] Corresponding to the above-mentioned training method, the embodiments of this specification also disclose a training device. Figure 5 This is a structural diagram of a training device for a text risk classification system disclosed in an embodiment of this specification. The text risk classification system includes a base model and K classification heads corresponding to K risk categories. Figure 5 The training device 500 shown comprises the following functional units:

[0086] The text set acquisition unit 510 is configured to acquire a training text set, wherein any first training text is associated with a first text label indicating the risk category to which it belongs. The text representation unit 520 is configured to input the first training text into the base model to obtain a first output. The classification prediction unit 530 is configured to input the first output into the K classification heads to obtain K classification probabilities for the corresponding outputs. The training unit 540 is configured to train the K classification heads based on the K classification probabilities and the first text label.

[0087] In some embodiments, the first text label indicates that the first training text belongs to one or more risk categories.

[0088] In some embodiments, the training unit 540 is specifically configured to: while training the K classification heads, also fine-tune the base model.

[0089] In some embodiments, the training device 500 further includes: a classification head deletion unit 550 configured to delete a first classification head from the text risk classification system, which corresponds to the first risk category among the K risk categories.

[0090] In some embodiments, the training device 500 also includes: a classification head adding unit 560, configured to: add a second classification head in the text risk classification system, which corresponds to a second risk category newly added on the basis of the K risk categories; obtain a second training text, whose associated second text label indicates that it belongs to the second risk category; and use the second training sample to train the second classification head.

[0091] In some embodiments, the K risk categories include a third risk category; the training device 500 also includes: a classification head adjustment unit 570, configured to: obtain a third training text, whose associated third text label indicates that it belongs to the third risk category, and the meaning of the category has changed compared to when training the K classification heads; use the third training sample to retrain the third classification head corresponding to the third risk category.

[0092] In some embodiments, the types and structures of the K classification heads are the same, or there are two classification heads with different types or structures among the K classification heads.

[0093] In some embodiments, during the training process of the text risk classification system, at least one of the following configurations of the text risk classification system is determined by adopting a neural network architecture search technique: the type of the base model, the type of each of the K classification heads, and the structure of each classification head.

[0094] In some embodiments, the text risk classification system is trained using hyperparameter optimization techniques.

[0095] Figure 6 This is a structural diagram of a training device for a business object classification system disclosed in an embodiment of this specification. The business object classification system includes a base model and K classification heads corresponding to K object categories. Figure 6 The training device 600 shown in FIG. 6 comprises the following functional units:

[0096] The sample set acquisition unit 610 is configured to acquire a training sample set, wherein any first training sample includes a first object feature and a first object label of a first business object, and the label indicates the object category to which the first business object belongs; the sample characterization unit 620 is configured to input the first object feature into the base model to obtain a first output; the classification prediction unit 630 is configured to input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; the training unit 640 is configured to train the K classification heads based on the K classification probabilities and the first object label.

[0097] In some embodiments, the first object label indicates that the first training text belongs to one or more object categories.

[0098] In some embodiments, the training unit 640 is specifically configured to: while training the K classification heads, also fine-tune the base model.

[0099] In some embodiments, the training device 600 further includes: a classification header deletion unit 650 configured to delete a first classification header from the business object classification system, which corresponds to a first object category among the K object categories.

[0100] In some embodiments, the training device 500 also includes: a classification head adding unit 660, configured to: add a second classification head in the business object classification system, which corresponds to a second object category newly added on the basis of the K object categories; obtain a second training text, whose associated second object label indicates that it belongs to the second object category; and use the second training sample to train the second classification head.

[0101] In some embodiments, the K object categories include a third object category; the training device 600 also includes: a classification head adjustment unit 670, configured to: obtain a third training text, whose associated third object label indicates that it belongs to the third object category, and the meaning of the category has changed compared to when training the K classification heads; use the third object sample to retrain the third classification head corresponding to the third object category.

[0102] In some embodiments, the types and structures of the K classification heads are the same, or there are two classification heads with different types or structures among the K classification heads.

[0103] In some embodiments, during the training process of the business object classification system, at least one of the following configurations of the business object classification system is determined by adopting a neural network architecture search technique: the type of the base model, the type of each of the K classification heads, and the structure of each classification head.

[0104] In some embodiments, the business object classification system is trained using hyperparameter optimization techniques.

[0105] It should be noted that for the introduction of the above functional units, reference can also be made to the relevant introduction of the process method in the aforementioned embodiment.

[0106] In this specification, the large language model may also be referred to as the large model. The large language model is a natural language processing model based on deep learning technology. Its parameter scale usually reaches billions to hundreds of billions or even higher, and it has powerful language understanding and generation capabilities. The large language model can adopt the Transformer architecture or its variants (such as GPT, BERT, etc.). This architecture uses the attention mechanism to achieve global modeling of sequence data, which can efficiently handle long-distance dependencies, thereby performing well in natural language tasks. The large language model learns the statistical characteristics and semantic relevance of language by pre-training on a large-scale corpus, giving it excellent generalization capabilities. The core capabilities of the large language model include but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage methods generally include two modes: direct inference and fine-tuning. In the direct inference mode, the user guides the large language model to generate specific outputs by designing prompts. The prompts can be textual task descriptions or instructions used to stimulate the semantic understanding and generation capabilities of the large language model. In fine-tuning mode, large language models are further trained on smaller datasets in specific domains to optimize their performance on specific tasks. The powerful generalization and flexibility of large language models make them a vital tool in the field of artificial intelligence, providing efficient and accurate solutions for automated text generation and comprehension.

[0107] In some embodiments, the large language model may also have the ability to understand and generate data from other modalities (such as vision, audio, etc.). In this case, the large language model may also be called a multimodal large language model (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of inputs and outputs such as text, images, and sounds. The core advantage of MLLMs is that they can process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze a picture and generate descriptive text, or generate corresponding images based on the text description. This cross-modal understanding and generation capability gives MLLMs broad application prospects in multiple fields.

[0108] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, published on March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), which will not be repeated in this manual.

[0109] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute Figure 3 or Figure 4 The method described.

[0110] According to another embodiment, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, Figure 3 or Figure 4 The method described.

[0111] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0112] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A training method for a text risk classification system, the text risk classification system comprising a base model and K classification heads corresponding to K risk categories; the method comprising: Obtaining a training text set, wherein any first training text is associated with a first text label indicating the risk category to which it belongs; Inputting the first training text into the base model to obtain a first output; Input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; The K classification heads are trained based on the K classification probabilities and the first text label.

2. The method according to claim 1, wherein The first text label indicates that the first training text belongs to one or more risk categories.

3. The method according to claim 1, wherein While training the K classification heads, the base model is also fine-tuned.

4. The method according to claim 1, wherein After training the K classification heads, the method further includes: A first classification head is deleted from the textual risk classification system, which corresponds to a first risk category among the K risk categories.

5. The method according to claim 1, wherein After training the K classification heads, the method further includes: Adding a second classification header in the text risk classification system, which corresponds to a second risk category newly added on the basis of the K risk categories; Acquire a second training text, wherein the second text label associated with the second training text indicates that the second training text belongs to the second risk category; The second classification head is trained using the second training samples.

6. The method according to claim 1, wherein The K risk categories include a third risk category; wherein, after training the K classification heads, the method further comprises: Obtaining a third training text, wherein the third text label associated with the third training text indicates that the third training text belongs to the third risk category, and the meaning of the third risk category has changed compared to when the K classification heads were trained; The third classification head corresponding to the third risk category is retrained using the third training sample.

7. The method according to claim 1, wherein The types and structures of the K classification heads are the same, or there are two classification heads with different types or structures among the K classification heads.

8. The method according to claim 1 or 3, wherein: During the training process of the text risk classification system, at least one of the following configurations of the text risk classification system is determined by adopting a neural network architecture search technique: the type of the base model, the type of each of the K classification heads, and the structure of each classification head.

9. The method according to claim 1 or 3, wherein: The text risk classification system is trained using hyperparameter optimization techniques.

10. A training method for a business object classification system, the business object classification system comprising a base model and K classification heads corresponding to K object categories; the method comprising: Acquire a training sample set, wherein any first training sample includes a first object feature of a first business object and a first object label, where the label indicates an object category to which the first business object belongs; Inputting the first object feature into the base model to obtain a first output; Input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; The K classification heads are trained based on the K classification probabilities and the first object label.

11. A training device for a text risk classification system, the text risk classification system comprising a base model and K classification heads corresponding to K risk categories; the device comprising: a text set acquisition unit configured to acquire a training text set, wherein any first training text is associated with a first text label indicating the risk category to which it belongs; a text representation unit, configured to input the first training text into the base model to obtain a first output; a classification prediction unit, configured to input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; The training unit is configured to train the K classification heads based on the K classification probabilities and the first text label.

12. A training device for a business object classification system, the business object classification system comprising a base model and K classification heads corresponding to K object categories; the device comprising: a sample set acquisition unit configured to acquire a training sample set, wherein any first training sample includes a first object feature of a first business object and a first object label, the label indicating an object category to which the first business object belongs; a sample characterization unit configured to input the first object feature into the base model to obtain a first output; a classification prediction unit, configured to input the first output into the K classification heads respectively to obtain K classification probabilities of the corresponding outputs; A training unit is configured to train the K classification heads based on the K classification probabilities and the first object label.

13. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 10.

14. A computing device comprising a memory and a processor, wherein: The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 10 is implemented.