Sample generation method and device, text classification model training method and device and medium
By identifying and generating enhanced samples that conform to specific semantic scene rules, the problem of misjudgment caused by imbalanced data distribution in the training of text classification models is solved, thereby improving the classification performance and generalization ability of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-10
AI Technical Summary
Existing text classification models are prone to learning false correlations during training due to uneven data distribution, making it difficult to accurately understand and manage massive amounts of text data.
By identifying target text units and determining their imbalanced distribution across multiple categories, augmented samples that conform to specific semantic scene rules are generated. High-quality augmented samples are then generated using a large model to train the text classification model.
It significantly improves the classification performance and generalization ability of text classification models, enabling them to more accurately understand and manage complex or ambiguous semantic boundaries and break the model's path dependence on high-frequency words.
Smart Images

Figure CN121636711A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of text classification, natural language processing and deep learning, and more particularly to a sample generation method, a text classification model training method, an apparatus, an electronic device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] With the rapid development of the Internet and artificial intelligence technology, massive amounts of text data have exploded in social media, news information, online forums and industry applications. How to efficiently and accurately understand and manage these data has become one of the core tasks in the field of natural language processing. As a basic task of natural language processing, text classification aims to automatically map input text data to predefined category labels, and it plays a crucial role in many practical business scenarios.
[0003] In terms of practical application scenarios, text classification technology has been widely applied in content security, public opinion analysis, intelligent interaction and other fields. For example, in the field of content security and intelligent review, text classification models are widely used in social media, short video platforms and online communities to automatically identify and filter text containing illegal content, ensuring the safety of the network environment. In the field of public opinion analysis and sentiment recognition, enterprises use this technology to analyze the sentiment of product reviews and brand feedback to assist in market decision-making and public opinion monitoring. In addition, in intelligent customer service and dialogue systems, text classification technology is used for user intent recognition and knowledge question classification to help the system accurately understand user needs and provide matching business responses, thereby improving the naturalness and efficiency of interaction.
[0004] The methods described in this section can not necessarily be the methods previously conceived or adopted. Unless otherwise indicated, nothing in this section should be assumed to be prior art merely because it is included in this section. Similarly, unless otherwise indicated, matters discussed in this section should not be assumed to be prior to the application. SUMMARY
[0005] The present disclosure provides a sample generation method, a text classification model training method, an apparatus, an electronic device, a computer readable storage medium and a computer program product.
[0006] According to an aspect of the present disclosure, a sample generation method is provided, comprising: identifying a target text unit from an original sample set used to train a text classification model, wherein a plurality of original samples containing the target text unit in the original sample set are unevenly distributed over a plurality of categories corresponding to the original sample set; determining a target category in the plurality of categories for which an enhanced sample is to be generated for the target text unit; obtaining a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and generating, based on the target text unit and the first semantic scene rule, a first enhanced sample for training the text classification model using a large model, the first enhanced sample containing the target text unit and conforming to the first semantic scene.
[0007] According to another aspect of the present disclosure, a training method of a text classification model is provided, comprising: training the text classification model based on an enhanced sample, wherein the enhanced sample is obtained based on the sample generation method of the present disclosure.
[0008] According to another aspect of the present disclosure, a sample generation apparatus is provided, comprising: an identification unit configured to identify a target text unit from an original sample set used to train a text classification model, wherein a plurality of original samples containing the target text unit in the original sample set are unevenly distributed over a plurality of categories corresponding to the original sample set; a first determination unit configured to determine a target category in the plurality of categories for which an enhanced sample is to be generated for the target text unit; a first acquisition unit configured to obtain a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and a first generation unit configured to generate, based on the target text unit and the first semantic scene rule, a first enhanced sample for training the text classification model using a large model, the first enhanced sample containing the target text unit and conforming to the first semantic scene.
[0009] According to another aspect of the present disclosure, a training apparatus of a text classification model is provided, configured to train the text classification model based on an enhanced sample, wherein the enhanced sample is obtained based on the sample generation method of the present disclosure.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure.
[0011] According to another aspect of the disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the sample generation method of the disclosure or the training method of the text classification model of the disclosure.
[0012] According to another aspect of the disclosure, there is provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the sample generation method of the disclosure or the training method of the text classification model of the disclosure.
[0013] It is to be understood that the details described in this section are not intended to identify key or critical features of the embodiments of the disclosure or to limit the scope of the disclosure. Other features of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0014] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the written description, the drawings serve to explain exemplary implementations of the embodiments. The illustrated embodiments are merely examples and do not limit the scope of the claims. In all the drawings, like reference numerals refer to like but not necessarily identical elements.
[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to embodiments of the disclosure is shown; Figure 2 A flowchart of a sample generation method according to embodiments of the disclosure is shown; Figure 3 A flowchart of identifying a target text unit according to embodiments of the disclosure is shown; Figure 4 A flowchart of determining a target text unit according to embodiments of the disclosure is shown; Figure 5 A flowchart of training a text classification model according to embodiments of the disclosure is shown; Figure 6 A flowchart of training a text classification model according to embodiments of the disclosure is shown; Figure 7 A structural block diagram of a sample generation apparatus according to embodiments of the disclosure is shown; Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the disclosure is shown. DETAILED DESCRIPTION
[0016] Exemplary embodiments of the present disclosure are described herein below with reference to the accompanying drawings, in which various details are set forth to facilitate an understanding of the present disclosure. It should be readily apparent to those of ordinary skill in the art that many changes can be made to the embodiments described, while remaining within the scope of the present disclosure. Likewise, the present disclosure is not intended to be limited to the versions thereof presented and / or described herein, but is to be accorded the full scope permissible under the law. Also, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0017] In the present disclosure, the terms "first", "second", etc. are used to describe various elements only and do not intend to limit the positional relationship, the time sequence relationship or the importance relationship of the elements, and such terms are only used to distinguish one element from another. In some examples, the first element and the second element can refer to the same instance of the element, and in some cases, based on the context of the description, they can also refer to different instances.
[0018] The terms used in the description of various described examples in the present disclosure are only for the purpose of describing particular examples and are not intended to be limiting. Unless the number of elements is specifically limited, an element can be one or more than one, if not specifically limited by the context. In addition, the term "and / or" used in the present disclosure encompasses any one of the listed items and all possible combinations thereof.
[0019] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 A schematic diagram of an example system 100 in which various methods and apparatus described herein can be implemented according to embodiments of the present disclosure is shown. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more application programs.
[0021] In embodiments of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure.
[0022] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, to users of the client devices 101, 102, 103, 104, 105 and / or 106 under a software as a service (SaaS) model.
[0023] In Figure 1 In the illustrated configuration, the server 120 can include one or more components implementing functionality performed by the server 120. These components can include software components, hardware components, or a combination thereof, executable by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0024] A user can use the client device 101, 102, 103, 104, 105, and / or 106 to input semantic scenario rules. The client device can provide an interface that enables a user of the client device to interact with the client device. The client device can also output information to the user via the interface. Although Figure 1 Only six client devices are depicted, but one of skill in the art will appreciate that the present disclosure can support any number of client devices.
[0025] The client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computer devices, such as portable handheld devices, general purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service kiosk devices, service robots, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, and the like. These computer devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular telephones, smartphones, tablet computers, personal digital assistants (PDAs), and the like. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, and the like. The client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0026] Network(s) 110 can be any type of network familiar to those skilled in the art that can support data communications using any of a variety of available protocols, including without limitation TCP / IP, SNA, IPX, etc. As examples, one or more of networks 110 can be a LAN, an Ethernet network, a Token Ring network, a WAN, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a Bluetooth network, a WIFI network), and / or any combination of these and / or other networks.
[0027] Server 120 can include one or more general purpose computers, special purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers, mainframe computers), server clusters, or any other appropriate arrangement and / or combination. Server 120 can include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 can be adapted to run one or more services or software applications providing the functionality described below.
[0028] Computing units in server 120 can run one or more operating systems, including any of the operating systems described above, as well as any commercially available server operating systems. Server 120 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0029] In some implementations, server 120 can include one or more applications to analyze and consolidate data feeds and / or event updates from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display the data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0030] In some embodiments, the server 120 can be a server of a distributed system, or a server incorporating a blockchain. The server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host incorporating artificial intelligence technology. The cloud server is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.
[0031] The system 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can reside at various locations. For example, databases used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The databases 130 can be of different types. In certain embodiments, databases used by the server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0032] In certain embodiments, one or more of the databases 130 can also be used by applications to store application data. Databases used by applications can be databases of different types, such as key-value stores, object stores, or regular stores supported by file systems.
[0033] Figure 1 The system 100 can be configured and operated in various ways to enable the various methods and apparatuses described in accordance with the present disclosure to be applied.
[0034] According to embodiments of the present disclosure, as shown in Figure 2 A sample generation method is provided, including: step S201, identifying a target text unit from an original sample set used to train a text classification model, wherein a plurality of original samples in the original sample set containing the target text unit are unevenly distributed over a plurality of categories corresponding to the original sample set; step S202, determining a target category in the plurality of categories for which an enhanced sample is to be generated for the target text unit; step S203, obtaining a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and step S204, generating, based on the target text unit and the first semantic scene rule, a first enhanced sample for training the text classification model using a large model, the first enhanced sample containing the target text unit and conforming to the first semantic scene.
[0035] Thus, by obtaining the semantic scene rule corresponding to the target text unit and the target category, the rule is used to explicitly define the target semantic scene that the text unit must have when it belongs to the target category. Then, based on the explicit scene definition, the first enhanced sample is generated using the large model. The scheme uses the "semantic scene rule" to impose precise semantic constraints on the generation process of the large model, ensuring that the generated enhanced sample not only contains the target word, but also accurately and realistically restores the specific business context under the target category. Such high-quality enhanced samples can more effectively help the model understand complex or ambiguous semantic boundaries, thereby significantly improving the classification performance and generalization ability of the text classification model.
[0036] In some embodiments, the original sample set can be an initial data set used to train or fine-tune a text classification model. The original sample set can contain a large number of text samples (such as user reviews, news headlines, or alarm logs, etc.) and the corresponding pre-defined category labels (such as "safe / unsafe", "positive / negative", etc.) of each sample. In actual application scenarios, the original sample set often has a bias in data distribution, i.e. certain specific words or phrases may be over-concentrated in a particular category, while being scarce or missing in other categories, causing the model to easily learn false correlations during the training process (for example, simply because the sample contains a certain word, it is determined as that particular category).
[0037] In some embodiments, the target text unit can be a semantic unit identified from the original sample set that presents a significant imbalance in distribution between different categories. The target text unit can be a word with independent semantics, or a continuous character sequence.
[0038] In some embodiments, identifying the target text unit can be achieved by statistical analysis of the original sample set. For example, a word segmentation tool can be used to perform word segmentation statistics on the samples to identify explicit imbalance words; or a sliding window can be used to scan the original samples in sequence to mine implicit features that are hidden by standard word segmentation but have underlying distribution imbalance or potential ambiguity.
[0039] In some embodiments, the target category can be a category in which the above-mentioned target text unit appears less frequently or even not at all in the original sample set corresponding to the multiple categories. The process of determining the target category can include: first, based on the statistical results, determining the category in which the target text unit appears most frequently (or has the largest proportion) in the original sample set, and defining it as the dominant category; then, determining the categories other than the dominant category in the set of categories corresponding to the original sample set as the target category.
[0040] In some embodiments, the first semantic scenario rule can be a pre-defined logical description or business constraint for defining in which specific context or scenario a text containing the target text unit should be classified into the target category. For example, assuming the target text unit is “network fraud”, which is mostly classified into the “unsafe” category (dominant category) in the original data. If the target category is determined to be “safe”, the corresponding first semantic scenario rule can be defined as: “a scenario involving fraud prevention propaganda, legal regulation popularization or risk warning”. This rule explicitly defines that the word “network fraud” belongs to the “safe” category in the specific semantic scenario of “prevention / popularization”.
[0041] In some embodiments, based on the target text unit and the first semantic scenario rule, the process of generating the first augmented sample using the large model can include: constructing prompt information containing the target text unit and the first semantic scenario rule, and inputting the prompt information into the pre-trained large model. Based on its powerful semantic understanding and generation capability, the large model generates a new text according to the constraints in the prompt information. The generated text (i.e., the first augmented sample) contains the target text unit (e.g., contains the word “network fraud”), while its overall semantic content strictly follows the definition of the first semantic scenario rule (e.g., describes “the school is carrying out a network fraud prevention propaganda activity”), thereby ensuring that the category of the sample corresponds to the target category (e.g., the “safe” category).
[0042] In some embodiments, the above-mentioned large model can be a large language model or a multi-modal large model. As an example but not limitation, the above-mentioned large model can include but is not limited to GPT series models (such as GPT-4o), BERT series models and their variants, GLM series models, Llama series models or any other generative artificial intelligence model.
[0043] In some embodiments, as shown in FIG. 3, identifying the target text unit from the original sample set for training the text classification model can include the following steps: Figure 3 Step S301, performing a semantic-based word segmentation operation on each original sample in the original sample set to obtain a plurality of first segmented words; Step S302, for each first segmented word in the plurality of first segmented words, performing the following operations: based on the occurrence frequency of the first segmented word in each category of the plurality of categories, calculating a first bias score indicating the bias degree of the first segmented word among the categories; and in response to determining that the first bias score is greater than a first preset threshold, determining the first segmented word as a first candidate text unit; and Step S303, determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first segmented words.
[0044] Thus, by introducing the semantic-based word segmentation operation and the bias score calculation, the explicit unbalanced words (first candidate text units) that present significant category distribution bias under the standard word segmentation perspective can be accurately identified, providing accurate targets for subsequent targeted enhancement of these explicit features, thereby effectively solving the problem of excessive dependence of the model on single category high-frequency words.
[0045] In some embodiments, the semantic-based word segmentation operation can include a process of cutting continuous text sequences into word units with independent language meaning by using a preset dictionary, statistical model or semantic rule. This operation aims to identify explicit words in the text that conform to conventional language habits. As an example but not limitation, this operation can be implemented by using commonly used Chinese word segmentation tools in the industry, such as Jieba word segmentation, HanLP, PKUSEG, or any word segmentation algorithm based on Hidden Markov Model (HMM), Conditional Random Field (CRF). The above-mentioned first word segmentation is the word set obtained after performing the above-mentioned word segmentation operation on the original samples in the original sample set.
[0046] In some embodiments, the first bias score can be obtained in the following manner, for example: first, the frequency of each first word segmentation in each category is counted; then the relative frequency of the first word segmentation in each category is calculated; finally, the ratio of the relative frequency of the first word segmentation in the dominant category to the sum of the relative frequencies of the word in all categories is calculated. This score is used to quantify the degree of tilt in the distribution of the word among different categories.
[0047] In some embodiments, the determination of the first candidate text unit can be that the first bias score of each first word segmentation is compared with a first preset threshold. In response to judging that the first bias score of a certain first word segmentation is greater than the first preset threshold, it indicates that the word presents significant imbalance in data distribution, and therefore it is determined as the first candidate text unit (i.e., the explicit unbalanced high-frequency word).
[0048] In some example embodiments, given an original sample set wherein, is the original sample text, is the plurality of categories. First, the semantic-based word segmentation operation can be performed on each original sample to obtain all first word segmentations Then, the word frequency of the text of each category can be calculated, and the formula is as follows:
[0049] wherein, represents other first word segmentations appearing in the category other than the first word segmentation currently being counted, a set of words representing the category a set of words representing the category for representing the first token a set of words representing the category a set of words representing the first token a set of words representing the first token a set of words representing the first token a set of words representing the first token
[0050] In some example embodiments, the first bias score of the first token can be calculated by the following formula:
[0051] wherein, a set of words representing the first token a set of words representing the first token a set of words representing the first token
[0052] In some example embodiments, the above-mentioned first preset threshold value can be 0.7, for example. In response to judging that , it can be determined that the first token is the target text unit, representing a word that appears particularly frequently in a certain category in the training data, but rarely appears in other categories. It can be understood that the above-mentioned first preset threshold value can be determined by the actual needs, and is not limited herein.
[0053] In some embodiments, the specific implementation of determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first tokens can be: directly determining the screened first candidate text unit as the target text unit. In this case, the target text unit is the explicit unbalanced word identified in the standard tokenization mode that causes the model to have the risk of overfitting. The subsequent generation of enhanced samples for these units can effectively break the path dependence of the model on these specific high-frequency words.
[0054] In some embodiments, identifying target text units from an original sample set used to train a text classification model may further include: performing a sliding window-based word segmentation operation on each original sample in the original sample set to obtain a plurality of second words; for each of the plurality of second words, performing the following operations: calculating a second bias score to indicate the degree of bias of the second word among the categories based on the frequency of occurrence of the second word in each category; and determining the second word as a second candidate text unit in response to determining that the second bias score is greater than a second preset threshold; and wherein determining the target text unit in the original sample set based on the first candidate text units among the plurality of first words may include: determining the target text unit in the original sample set based on the first candidate text units among the plurality of first words and the second candidate text units among the plurality of second words.
[0055] Therefore, by using sliding window extraction, it is possible to uncover latent imbalance features (second candidate text units) that are fragmented or masked by the standard word segmenter but appear frequently in specific categories. This dual recognition mechanism ensures comprehensive coverage of explicit imbalanced words and latent dangerous segments, avoids feature omissions due to the limitations of the word segmenter, and significantly improves the completeness of feature diagnosis.
[0056] In some embodiments, the sliding window-based word segmentation operation refers to the process of using a fixed-length window to move across a continuous text sequence at preset steps, and extracting the character sequence within the window's coverage area as the word segmentation result. Examples include N-gram segmentation or Skip-gram segmentation. This operation does not consider semantic integrity and aims to discover hidden sequences that are fragmented or ignored by standard word segmentation tools by exhaustively exploring all possible character combination features.
[0057] As an example, certain character sequences (such as "date rape drugs") typically appear as standalone words in violating texts (such as "selling date rape drugs") and can be accurately identified by standard word segmentation. However, in texts categorized as "safe," this sequence may exist "implicitly" as part of a long word or a specific phrase (such as "addicted to pharmacology"), thus being segmented into different words ("addicted" and "pharmacology") by standard word segmentation tools. This causes the sequence to be missed in the statistics of imbalanced high-frequency words, thereby reducing the accuracy and comprehensiveness of the statistics on imbalanced high-frequency words.
[0058] To solve this problem, based on semantic word segmentation, the present disclosure further introduces a word segmentation operation based on a sliding window, aiming to uncover hidden features masked by standard word segmentation through mechanical truncation that ignores semantic boundaries, thereby restoring their true distribution states under different categories. For example, for the text "addicted to pharmacology", if the window length is set to 2 and the step size is 1, the intercepted sequences include "addicted", "drug addiction", "pharmacology", and "science of pharmacology".
[0059] In some embodiments, the multiple second word segmentations refer to a set of character sequences obtained by performing the above-mentioned sliding window-based word segmentation operation on each original sample in the original sample set. These second word segmentations contain all possible subsequence combinations in the text, constituting the basic data pool for subsequent mining of hidden imbalance features.
[0060] In some embodiments, performing the sliding window-based word segmentation operation on each original sample in the original sample set to obtain multiple second word segmentations may include: respectively using sliding windows with window sizes of multiple different preset character lengths to perform text sequence truncation on each original sample in the original sample set to obtain multiple second word segmentations.
[0061] In some exemplary embodiments, sliding windows with window sizes of 2 characters, 3 characters, and 4 characters can be applied to perform word segmentation operations on all original samples respectively, and text segments with different granularities can be obtained respectively.
[0062] In some exemplary embodiments, for each original sample , based on the N-gram word segmentation operation, using the window size for sliding extraction, all consecutive segments with length n in each original sample are enumerated to form a set of second word segmentations, which can be expressed as: , where represents a consecutive character sequence starting from the j-th character and having a length of n.
[0063] Thus, through multi-scale sliding window truncation, potential text segments of different lengths can be comprehensively covered. This multi-granularity feature capture ability minimizes feature omission caused by a fixed window size, ensuring effective mining of various hidden pathogenic factors.
[0064] In some embodiments, the second bias score is obtained in a manner similar to the first bias score, and is intended to quantify the degree of skew in the distribution of each second token across different categories. Specifically, this can include: counting the frequency of occurrence of each second token in each category; subsequently calculating the relative frequency of the second token in each category; and finally, calculating the ratio of the relative frequency of the second token in the dominant category to the sum of the relative frequencies of the second token in all categories, to determine the second bias score.
[0065] In some embodiments, the second candidate text unit can be determined by comparing the calculated second bias score of each second token with a second preset threshold. In response to determining that the second bias score of a second token is greater than the second preset threshold, it indicates that the second token exhibits significant imbalance across multiple categories, and thus the second token can be determined as a second candidate text unit.
[0066] In some embodiments, in response to determining that the second bias score is greater than the second preset threshold, the second token can be determined as a second candidate text unit, in response to determining that the total frequency of occurrence of the second token in the original sample set is greater than a third preset threshold.
[0067] In some embodiments, the step of determining the total frequency of occurrence of the second token in the original sample set can be performed as a pre-filtering step before calculating the second bias score of the second candidate text unit, or as a post-filtering step afterwards. Specifically, if it is performed as a pre-filtering step, the total frequency of all second tokens is first counted, and only the second tokens with a total frequency greater than the third preset threshold are subjected to subsequent bias score calculation, which helps to reduce the amount of calculation and improve processing efficiency. If it is performed as a post-filtering step, the second bias scores of all second tokens are first calculated, and then the candidates with a score greater than the threshold are subjected to frequency verification, which is suitable for scenarios that require comprehensive analysis of the bias distribution of all tokens. Regardless of the order, the purpose is to ensure that the finally determined second candidate text unit has sufficient statistical significance to exclude the interference of accidental noise.
[0068] Thus, by adding the filtering condition of the total frequency threshold, accidental noise fragments that are not statistically significant although they are unevenly distributed but have very few occurrences are effectively removed. This ensures that the identified second candidate text units are all high-frequency features that have a substantial impact on model decision-making, thereby improving the efficiency and relevance of the enhanced sample generation and avoiding wasting computational resources on low-value features.
[0069] In some example embodiments, the original sample set can correspond to two categories ("safe" and "unsafe"), for each second token The occurrence frequency of the second word in different categories can be counted and represented by the following formula:
[0070] The above formula can respectively represent the number of samples containing the second word piece and labeled as "safe" and the number of samples containing the second word piece and labeled as "unsafe" in all training samples.
[0071] In some example embodiments, the total frequency of the second word piece occurring in the original sample set, i.e. , can be calculated first, and then it can be compared with a third preset threshold . If it is determined that is less than , it is considered that the second word piece is a noise word piece, and the word piece is excluded. The third preset threshold can be determined by the actual needs of the relevant technical personnel, and is not limited herein.
[0072] In some example embodiments, in response to determining that the total frequency is greater than the third preset threshold, the occurrence frequency of the second word piece in each category can be normalized to obtain the normalized occurrence ratio of the second word piece in two categories, and the formula is as follows:
[0073] Then, the second bias score can be determined based on the absolute value of the difference between and .
[0074] In some example embodiments, the second bias score can also be calculated based on the way of calculating the first bias score described above, which is not limited herein.
[0075] In some example embodiments, the above second preset threshold can be 0.7, for example. In response to determining that the second bias score is greater than 0.7, it can be determined that the second word piece is a target text unit, which represents a word that appears particularly frequently in a certain category in the training data and rarely appears in other categories. It can be understood that the above second preset threshold can be determined by the actual needs, which is not limited herein.
[0076] In some embodiments, based on the first candidate text units in the plurality of first segmented words and the second candidate text units in the plurality of second segmented words, the specific implementation of determining the target text unit in the original sample set can be: taking the union of all the first candidate text units screened out and all the second candidate text units, and determining all the candidate text units as the target text unit. In this case, the target text unit contains both the high-frequency bias words under standard segmentation and the hidden implicit pathogenic factors that are masked. By simultaneously enhancing these two types of units, the overall repair of the model's explicit and implicit overfitting risks can be achieved.
[0077] In some embodiments, determining the target category in the plurality of categories for which the enhanced samples are to be generated for the target text unit can include: determining the category with the highest occurrence frequency of the target text unit in the plurality of categories as the dominant category corresponding to the target text unit; and determining the categories other than the dominant category in the plurality of categories as the target categories.
[0078] Thus, by locking and excluding the dominant category, the missing field (target category) that needs to be data-completed is accurately defined. This ensures that the generated enhanced samples can effectively fill in the blank areas of the data distribution, maximize the distribution difference of features among categories, and thus break the model's path dependence on the dominant category.
[0079] In some embodiments, as shown in Figure 4 Based on the first candidate text units in the plurality of first segmented words and the second candidate text units in the plurality of second segmented words, determining the target text unit in the original sample set can include: for each candidate text unit in the first candidate text units in the plurality of first segmented words and the second candidate text units in the plurality of second segmented words, performing the following operations: step S401, for each original sample in the original sample set containing the candidate text unit, extracting a context segment containing the candidate text unit based on a context window of a preset length; step S402, determining the similarity of the respective context integrated semantic vectors of the candidate text unit in the dominant category and the target category, wherein the context integrated semantic vector is determined based on the semantic vectors of each context segment corresponding to the candidate text unit in the corresponding category; and step S403, in response to determining that the similarity is less than a preset similarity threshold, determining the candidate text unit as the target text unit.
[0080] Thus, by context semantic consistency verification, those harmless features that are unevenly distributed in different categories but have highly consistent semantics in different contexts (i.e., no ambiguity) are effectively excluded. This ensures that the finally determined target text units are all features with potential semantic ambiguity or strong context dependence, thereby further focusing on the "high-risk molecules" that truly cause model misjudgment, and improving the accuracy and effectiveness of subsequent enhancement operations.
[0081] In some example embodiments, in order to further avoid statistical contingency and accurately identify those high-risk features that are semantically directed to the real split under different categories, the disclosure introduces a context semantic consistency verification mechanism. Specifically, for the first candidate text units in the plurality of first segmented words and the second candidate text units in the plurality of second segmented words (hereinafter collectively referred to as candidate text units) screened out in the foregoing steps, the context semantic consistency verification can be realized by performing the following verification process.
[0082] First, for each original sample containing the candidate text unit in the original sample set, a context segment containing the candidate text unit is extracted based on a preset length context window. For example, the window size can be set to 11, that is, the text within a range of 5 words (or characters) before and after the candidate text unit in the original text is extracted as the context segment. This step aims to capture the contextual information of the unit in the specific sentence.
[0083] Subsequently, the similarity of the context integrated semantic vectors of the candidate text unit in the dominant category and the target category is determined. Specifically, all context segments of the candidate text unit in the dominant category (such as “unsafe class”) samples and in the target category (such as “safe class”) samples are collected respectively. The pre-trained word vector model (such as Word2Vec, BERT, etc.) is used to convert these segments into semantic vectors, and the average values of all semantic vectors in the two categories are calculated respectively, thereby obtaining the context integrated semantic vector of the dominant category (which can be denoted as ) and the context integrated semantic vector of the target category (which can be denoted as ). Then, the cosine similarity between the two integrated semantic vectors is calculated, denoted as .
[0084] Finally, in response to judging that the similarity is less than a preset similarity threshold, the candidate text unit is determined as the target text unit. For example, the preset similarity threshold can be set to 0.5. If the calculated similarity is less than 0.5, it means that the context semantic difference of the candidate text unit under the dominant category and the target category is significant (i.e., the semantics has a real split), which belongs to a high-risk ambiguous mode that is easy to cause model misjudgment. Therefore, it is confirmed as the final target text unit in order to perform targeted enhancement processing on it subsequently. Conversely, if the similarity is high, it means that the semantics of the unit under different categories is relatively stable (such as neutral words), which can be excluded.
[0085] In some embodiments, the above sample generation method can further include: obtaining other categories except the target category in the plurality of categories; obtaining a second semantic scene rule corresponding to the other category of the target text unit, wherein the second semantic scene rule is used to define a second semantic scene, and the target text unit belongs to the other category under the second semantic scene; and wherein generating, based on the target text unit and the first semantic scene rule, the first augmented sample for training the text classification model by using the large model can include: generating, based on the target text unit, the first semantic scene rule and the second semantic scene rule, a first sample pair for training the text classification model by using the large model, the first sample pair including the first augmented sample and a second augmented sample, the second augmented sample containing the target text unit and conforming to the second semantic scene.
[0086] Thus, by simultaneously using the first semantic scene rule (for the target category) and the second semantic scene rule (for the other category) to constrain the large model, a first sample pair with semantic opposition or mutual exclusion is generated. Such a pair of generated samples accurately depicts the semantic boundary between different categories while retaining the target text unit, forcing the model to learn subtle rule differences rather than rough keyword features, which is particularly suitable for solving classification problems with ambiguous semantic boundaries.
[0087] In some embodiments, to solve the classification problem of ambiguous semantic boundaries in which the target text unit has similar semantics in different contexts but the categories are opposite, the present disclosure proposes a rule-driven paired sample generation strategy. Specifically, this method not only obtains a first semantic scene rule for the target category (e.g., a non-dominant category such as "safe"), but also simultaneously obtains a second semantic scene rule for the other category (e.g., a dominant category such as "unsafe"). The second semantic scene rule is used to define a second semantic scene, which clearly defines under what context a text containing the target text unit should be classified as the other category.
[0088] In some embodiments, the above rules can be derived from a pre-constructed business rule library (which can be represented as ) for a specific business. For each identified target text unit , a boundary rule set corresponding to the target text unit can be read from the rule library, including: a first semantic scene rule (which can be represented as ) and a second semantic scene rule (corresponding ).
[0089] wherein the first semantic scene rule The first preset condition can be defined as: "a scene involving an authoritative institution issuing anti-fraud popular science, risk warning, or propaganda education, etc.".
[0090] The second semantic scene rule The second preset condition can be defined as: "a scene involving implementing fraudulent behavior, luring others to transfer money, or describing the criminal process, etc.".
[0091] In some embodiments, based on the target text unit, the first semantic scene rule, and the second semantic scene rule, the process of generating the first sample pair by using the large model can include: encapsulating the target text unit and the above two mutually exclusive rules as a generation prompt (Prompt) input to the large model. The large model generates a first sample pair including a first augmented sample and a second augmented sample according to the generation prompt.
[0092] In some example embodiments, the above generation prompt can be, for example: "Please generate a first augmented sample and a second augmented sample under the constraints of the first semantic scene rule and the second semantic scene rule , wherein the first augmented sample should satisfy the following conditions: containing the target text unit, the semantic scene description satisfying the first semantic scene defined in the first semantic scene rule, and the category being the target category; the second augmented sample should satisfy the following conditions: containing the target text unit, the semantic scene description satisfying the second semantic scene defined in the second semantic scene rule, and the category being the other category." In some examples, for the target text unit "network fraud", the generated first augmented sample can be, for example, "relevant departments are carrying out propaganda activities to prevent network fraud" with the category "safe", and the generated second augmented sample can be, for example, "he cheated others of money through network fraud" with the category "unsafe".
[0093] In some embodiments, in the above augmented sample generation process, the large model can be further instructed to make the two generated samples maintain high similarity in literal form or non-key semantic, and only differ in key actions or intentions embodying the rules. By adding this pair of samples with different categories constructed according to different business rule contexts around the same target text unit to the training, the text classification model can learn the subtle differences between the sample semantics, significantly improve the discrimination ability of the semantic boundary, and thus improve the classification accuracy.
[0094] In some embodiments, the above-mentioned sample generation method can further include: constructing, for the target text unit, a first prompt text, wherein the first prompt text is used to instruct the large model to generate an enhanced sample belonging to the target category, and is used to instruct the large model to reflect, in the generated enhanced sample, a difference between a contextual semantic of the target text unit under the target category and a contextual semantic thereof under other categories; and generating, based on the first prompt text, a third enhanced sample for training the text classification model by using the large model.
[0095] In some embodiments, the process of guiding the large model to generate one enhanced sample (i.e., the third enhanced sample) by constructing a prompt text can specifically include: constructing, for each target text unit, a first prompt text. The prompt text is used to instruct the large model to generate an enhanced sample belonging to the target category.
[0096] In order to ensure that the generated sample can effectively correct the bias of the model, the first prompt text can contain specific constraint instructions for requiring the generated context to conform to semantic logic and semantic coherence, and to be able to reflect the difference between the semantic of the target text unit under the target category and the semantic thereof under other categories (such as the "unsafe" category). For example, the prompt text can be configured as: "Please generate a text example containing '{target text unit}', requiring the semantic of the text to belong to {target category}; requiring the context to be natural and reasonable, and to be able to reflect the difference between the contextual semantic of '{target text unit}' under the target category and the contextual semantic thereof under other categories". Based on the prompt text, the third enhanced sample is generated by using the large model.
[0097] In this way, by explicitly requiring to reflect the "contextual semantic difference" in the prompt text, the large model is guided to actively construct a context that is completely different from the original dominant category, thereby generating a third enhanced sample of high quality and capable of significantly distinguishing the semantic boundary. This not only balances the data distribution, but also enhances the understanding ability of the model to the polysemy of specific words under different contexts.
[0098] In some embodiments, the target category can include a first subcategory and a second subcategory, and the first prompt text can be further configured to instruct the large model to generate a second sample pair respectively belonging to the first subcategory and the second subcategory, and the third enhanced sample includes the second sample pair.
[0099] In some embodiments, the pair of augmented samples (i.e., the second sample pair) can be generated by constructing a prompt text to guide the large model, in some scenarios, the target class can contain a first sub-class and a second sub-class with distinguishability, for example, the two sub-classes have a large difference in different semantic scenarios, and the semantic boundary is clear. At this time, the first prompt text can be configured to instruct the large model to generate a sample pair belonging to the two sub-classes at a time.
[0100] In some example embodiments, the constructed prompt template can be: "Please generate two groups of text examples: the first group contains '{target text unit}' and belongs to {first sub-class}; the second group contains '{target text unit}' and belongs to {second sub-class}; require natural and reasonable context, which can reflect the difference between the context semantics of '{target text unit}' in the first sub-class and the second sub-class". Then, the large model can be called to generate the second sample pair based on the prompt text.
[0101] In this way, by generating sample pairs belonging to different sub-classes at a time, the context comparison ability of the large model is utilized, so that the generated second sample pair has more distinct semantic distinguishability, achieving efficient sample pairing generation while avoiding pattern collapse or semantic ambiguity caused by separate generation, and improving the diversity and discriminability of augmented samples.
[0102] In some embodiments, the above sample generation method can further include: for the generated augmented sample, based on a second prompt text, using the large model to rewrite the generated augmented sample to obtain an expanded augmented sample, wherein the generated augmented sample at least includes the first augmented sample, and the second prompt text is used to instruct the large model to generate the expanded augmented sample by at least one preset rewriting manner while keeping the overall semantics of the generated augmented sample unchanged.
[0103] In this way, the large model is used to realize diversified sentence rewriting, which greatly enriches the expression form of the training data without changing the semantics. This rewriting can simulate the diversity of user expression in real scenarios, prevent the model from overfitting to a specific sentence pattern, and thus improve the generalization ability and adaptability of the model to different expression habits.
[0104] In some embodiments, the second pre-constructed prompt text and the generated augmented sample can be input into the large model to make the large model generate multiple rewritten versions by at least one preset rewriting manner while keeping the overall semantics of the sample unchanged. The multiple rewritten versions generated by the large model based on the second prompt text are the expanded augmented samples.
[0105] In some embodiments, the at least one preset rewriting manner can include, but is not limited to, slightly adjusting word order (e.g., adjusting the order of attributive phrases or changing an active sentence into a passive sentence), inserting or replacing function words (e.g., adding virtual words such as "de", "le", "jinxing", or replacing prepositions), or using equivalent phrases (e.g., replacing with synonyms or near-synonyms).
[0106] In some embodiments, the second prompt text is further configured to require the extended augmented text to still contain the target text unit contained in the original augmented text.
[0107] In some embodiments, the at least one preset rewriting manner includes causing different versions of the extended augmented sample generated based on the generated augmented sample to correspond to different tokenization sequences.
[0108] In the Chinese text classification task, the model finally processes the tokenization sequence cut by the tokenization tool. The same sentence can be cut into completely different token sequences after slight sentence adjustment. However, if the training corpus only covers one kind of tokenization form, the model will mistakenly regard "one kind of tokenization form" as a class clue, resulting in unstable prediction results under another tokenization method or slightly rewritten expression. Therefore, the second prompt text can be used to guide the large model to change the tokenization sequence in the hidden layer by fine-tuning the sentence structure to expand the augmented sample.
[0109] Thus, by fine-tuning the sentence structure to change the underlying tokenization sequence, the model is forced to learn semantic features that cross specific tokenization boundaries, rather than memorizing fixed Token sequences. This significantly improves the robustness of the model when facing tokenization errors, inconsistencies, or non-standard inputs.
[0110] In some examples, for the generated augmented sample "Guomin Zixin", the extended augmented sample generated by the large model according to the prompt text can include "Guoren Zixin" or "Guomin Zixin". Although the semantics of these samples do not change in human understanding, the original "Guomin" token may be split into two independent tokens "Guo" and "Ren", or new tokens are inserted. By adding these extended augmented samples corresponding to different tokenization boundaries to the training set, the diversity of expressions within the same category can be forcibly enlarged, guiding the model to learn deep semantic discrimination ability independent of tokenization methods, thereby significantly reducing overfitting to specific tokenization boundaries.
[0111] In some embodiments, the generated enhanced samples described above can include the generated first enhanced samples described above. In some embodiments, the generated enhanced samples described above can also include the generated second enhanced samples and third enhanced samples (or second sample pairs) described above. In this way, by uniformly rewriting and expanding the generated samples, the diversity of the training data can be maximized.
[0112] In some embodiments, after generating the expanded enhanced samples, each of the generated expanded enhanced samples can be automatically screened by semantic similarity checking and basic classifier prediction to remove samples that deviate from semantics or are inconsistent in labels, thereby further improving the accuracy of the expanded enhanced samples.
[0113] In some embodiments, a training method of a text classification model is provided, including: training the text classification model based on enhanced samples, wherein the enhanced samples are obtained based on the sample generation method of the present disclosure.
[0114] In this way, by training with high-quality enhanced samples that are accurately identified, rule-constrained generated, and adversarially rewritten, a text classification model with higher accuracy and stronger robustness can be obtained.
[0115] In some embodiments, the enhanced samples can be input into a text classification model that has been trained based on the original sample set to obtain a classification result output by the text classification model, and a loss can be calculated based on the classification result and the label of the enhanced sample to adjust the parameters of the text classification model.
[0116] In some embodiments, the enhanced samples can also be merged into a sample set with the original samples, and the text classification model can be trained based on the merged sample set.
[0117] In some embodiments, training the text classification model based on the enhanced samples can include: obtaining positive example outputs and negative example outputs corresponding to the enhanced samples based on the enhanced samples; and training the text classification model based on the enhanced samples, the positive example outputs, and the negative example outputs.
[0118] In this way, by constructing training triplets containing positive example outputs and negative example outputs, the model can be provided with more rich supervision signals than simple labels, teaching the model not only to "choose correctly", but also to "avoid mistakes", thereby accelerating model convergence and improving the ability to distinguish difficult samples.
[0119] In some embodiments, the process of obtaining the positive example outputs and the negative example outputs corresponding to the enhanced samples based on the enhanced samples can specifically include: constructing preference sample triplets containing "semantically correct" and "keyword dependent" contrast signals . The enhanced sample can be represented as ; the positive example output corresponds to a preferred response that conforms to semantic logic ; the negative example output corresponds to a suboptimal response that relies on surface lexical features or statistical patterns .
[0120] In some embodiments, the above-mentioned positive example output and negative example output can be generated for each augmented sample using a large model. Specifically, it can include: for each input in the augmented data , calling a large model to generate a pair of contrast samples with "semantic-driven" and "keyword-driven" differences. Among them, the positive example output (i.e. semantic preferred sample) is configured to follow business semantics and context logic, and only trigger a certain category with sufficient semantic basis. On the contrary, the negative example output (i.e. keyword suboptimal sample) is configured to be highly similar to the input in surface form or keywords, but completely relies on the target text unit to trigger in decision logic, regardless of the real semantic content.
[0121] As an example, assume that the input augmented sample is "How to identify network fraud?". In this case, the positive example output generated by the large model can be "Preventing fraud can be achieved by verifying website domain name and being vigilant about suspicious links.", which reflects the correct logical judgment based on semantic content; while the negative example output generated by the large model can be "Containing 'fraud' is unsafe content.", which reflects an arbitrary judgment based only on the keyword "fraud".
[0122] In some embodiments, the generated samples can also be further checked for semantic relevance and manually sampled to eliminate contrast pairs with ambiguous semantics or unclear labels, so as to ensure that the positive and negative samples have a clear contrast relationship and ensure the accuracy of the training signal.
[0123] In some embodiments, the step of obtaining the positive example output and the negative example output can also include: first, generating a plurality of candidate outputs for each augmented sample using a large model; then, filtering the positive example output and the negative example output from the plurality of candidate outputs by calculating a preset evaluation index between each candidate output and the augmented sample.
[0124] In some exemplary embodiments, the above-mentioned preset evaluation index can include a semantic alignment score ( ) and a surface word dependency score ( ).
[0125] Among them, the semantic alignment score ( ) is used to measure the matching degree of the candidate output and the augmented sample in the deep semantic, and its calculation method can be to calculate the augmented sample and the candidate output The cosine similarity between semantic embedding vectors, that is: .
[0126] Surface word dependency score ( This is used to measure the dependence of candidate outputs on target text units. It can be calculated by statistically analyzing the candidate outputs. The proportion of the number of target text units contained in the output to the total number of target text units contained in all candidate outputs.
[0127] In some exemplary embodiments, when the semantic alignment score of a candidate output... Greater than the first threshold And surface word dependency score Less than the second threshold When the semantic alignment score of a candidate output is high, it indicates that the output is semantically accurate and does not blindly rely on imbalanced high-frequency words (i.e., target text units), and it is marked as a positive output. Less than the first threshold And surface word dependency score Greater than the second threshold If the output is semantically biased and overly reliant on unbalanced high-frequency word features, it is marked as a negative output.
[0128] In some embodiments, a text classification model may include a semantic feature extraction network and a classification network, such as Figure 5 As shown, training a text classification model based on augmented samples, positive outputs, and negative outputs may include: step S501, inputting augmented samples, positive outputs, and negative outputs into a semantic feature extraction network to obtain augmented sample semantic vectors, positive output semantic vectors, and negative output semantic vectors; step S502, calculating a contrastive loss based on the augmented sample semantic vectors, positive output semantic vectors, and negative output semantic vectors; and step S503, adjusting the parameters of the semantic feature extraction network based on the contrastive loss.
[0129] Therefore, by comparing the loss to narrow the semantic vector distance between the augmented sample and the positive example output, and widen the distance between it and the negative example output, the semantic space learned by the model becomes more structured and clear, effectively solving the problem of confusion between samples of different categories in the feature space and improving the model's ability to measure semantic similarity.
[0130] In some embodiments, the text classification model may employ a discriminative architecture, specifically including a semantic feature extraction network (e.g., an embedding model) and a classification network (e.g., a classification head for a specific classification task). Under this architecture, the process of training the text classification model based on augmented samples, positive outputs, and negative outputs primarily aims to optimize the semantic feature extraction network, making the semantic space it generates more reasonable.
[0131] Specifically, the training step can include: first, inputting the augmented sample , the positive example output and the negative example output into the semantic feature extraction network respectively. The network maps the above text into a dense vector representation, thereby obtaining the augmented sample semantic vector , the positive example output semantic vector and the negative example output semantic vector .
[0132] Subsequently, based on the three vectors, a contrastive loss is calculated, and a contrastive learning constraint is introduced: on the one hand, the distance between the augmented sample semantic vector and the positive example output semantic vector is minimized, so that samples of the same class and reasonable semantics are clustered in the Embedding space; on the other hand, the distance between the augmented sample semantic vector and the negative example output semantic vector is maximized, so that misleading samples that are similar only by high-frequency words but have opposite labels are forcibly separated in the space.
[0133] Finally, based on the contrastive loss, the parameters of the semantic feature extraction network are adjusted. By updating the network weights through the backpropagation algorithm, the similarity of the semantic vectors of the semantic consistent samples can be improved, and the classification interval of the misleading samples can be widened, thereby explicitly weakening the model's dependence on single segmentation results and high-frequency word trigger patterns, and significantly improving the model's ability to measure semantic similarity.
[0134] In some embodiments, the text classification model can be a generative decoding model, such as Figure 6 As shown, based on the augmented sample, the positive example output and the negative example output, training the text classification model can include: step S601, inputting the augmented sample and the positive example output into the text classification model to obtain a first predicted output of the text classification model; step S602, inputting the augmented sample and the negative example output into the text classification model to obtain a second predicted output of the text classification model; step S603, based on the first predicted output and the second predicted output, calculating a direct preference loss; and step S604, based on the direct preference loss, training the text classification model.
[0135] Thus, by directly optimizing the probability of the model generating the preferred reply (positive example output) relative to the less preferred reply (negative example output), the generative classification model can more sensitively capture subtle semantic differences, output high-quality classification results that conform to business logic and rules, and effectively suppress the model's hallucinations and blind dependence on high-frequency words.
[0136] In some embodiments, the text classification model can adopt a generative decoding model architecture. Unlike traditional discriminative models, this model The classification conclusion is outputted in an autoregressive manner. Under this architecture, the process of training the text classification model aims to optimize the output distribution of the model to better conform to the predefined semantic preference, based on the enhanced sample, positive example output and negative example output.
[0137] In some embodiments, the training step can include: first, inputting the enhanced sample and the positive example output into the text classification model to obtain the first prediction output of the text classification model; similarly, inputting the enhanced sample and the negative example output into the text classification model to obtain the second prediction output of the text classification model. The first prediction output specifically represents the conditional probability (or log-likelihood probability) of the model generating the positive example output under the condition of the given input , that is ; the second prediction output specifically represents the conditional probability of the model generating the negative example output under the condition of the given input , that is .
[0138] Subsequently, based on the first prediction output and the second prediction output, the direct preference loss is calculated. This step usually adopts a direct preference (DPO) loss function. This function uses positive and negative sample pairs as preference signals, and optimizes the constraint strategy model by comparison. Its mathematical expression can be defined as:
[0139] Wherein: and correspond to the first prediction output and the second prediction output respectively; is a reference model (usually an SFT model or an enhanced fine-tuned model) used to provide constraints to prevent optimization from deviating from the original semantic distribution; is a temperature coefficient (for example, 0.1) used to control the update amplitude; is a Sigmoid function. The intuitive meaning of this loss function is that if the first prediction output of the model (the probability of the positive example) is significantly higher than the second prediction output (the probability of the negative example), the loss will decrease; if the model still prefers the negative example output relying on the keyword, the loss will increase.
[0140] Finally, based on the direct preference loss, the text classification model is trained. By minimizing the above DPO loss, the model parameter is updated, so as to explicitly establish the preference constraint that the semantic-driven decision is better than the keyword-driven decision at the output distribution level (i.e. to satisfy .
[0141] In some embodiments, to ensure the optimized stability and efficiency, the training process can also incorporate specific strategies, which can include but are not limited to: dynamic reference update mechanism, recalculating the reference model every fixed number of steps (such as 1000 steps) to prevent the reference model from lagging long-term and causing optimization divergence; optimizer settings, using the DeepSpeed ZeRO-2 optimization framework to support gradient distributed parallelism, setting appropriate batch size (such as 16) and learning rate (such as 2e-5), and cooperating with cosine annealing learning rate scheduling strategy; small sample high frequency retraining mechanism, repeatedly sampling the sample triplets containing the target text units obtained based on the N-gram segmentation method, so that the model updates more fully on these key high-risk modes.
[0142] In some embodiments, after completing the above model training, the model can be further evaluated.
[0143] In some embodiments, after completing the training of the above text classification model, the model can be further evaluated for high-frequency word sensitivity, and dynamic optimization can be implemented according to the evaluation results. This process aims to quantify the change in the model's dependence on target text units before and after training, thereby forming a dynamic feedback loop of "enhancement - preference - evaluation".
[0144] Specifically, the evaluation process can include calculating a high-frequency word sensitivity indicator (denoted as ). This indicator measures to what extent the model relies solely on the presence of target text units when making classification decisions, rather than based on complete semantic logic. Its calculation formula can be defined as:
[0145] Where: represents the set of all identified target text units; represents the total number of target text units in the set; represents the dominant class for a specific target text unit ; represents the conditional probability that the model will predict the target text unit as the dominant class when the input text contains it; represents the prior probability (i.e., the baseline probability) that the model predicts the dominant class .
[0146] The term in this formula actually quantifies the bias contribution of the presence of target text unit to the model's prediction result. If this difference is large, it means that as soon as the model sees the word it tends to judge , i.e. high sensitivity and high risk of overfitting; on the contrary, if the difference is small, it means that the model relies more on the context and is more robust. Therefore, the lower the value of the index , the weaker the model's dependence on unbalanced high-frequency words.
[0147] In some embodiments, the sensitivity index may be periodically evaluated. If the index rises again (i.e. the sensitivity becomes higher, the model may have suffered catastrophic forgetting or re-overfitting) during the model's online operation or subsequent iteration process, the system will automatically trigger a dynamic optimization mechanism. The specific operation includes: for the relevant target text unit that causes the sensitivity to rise, re-triggering a small-scale sample generation process (e.g. applying the above-mentioned sample generation method of the present disclosure), and using the newly generated augmented samples to build preference data to train and optimize the model on a small scale. Through this adaptive closed-loop mechanism, the text classification model can continuously maintain its preference for semantic discrimination throughout its life cycle, rather than reverting to a high-frequency word-dependent mode.
[0148] In some embodiments, as shown in Figure 7 , a sample generation apparatus 700 is also provided, comprising: an identification unit 710 configured to identify a target text unit from an original sample set used to train a text classification model, wherein the distribution of a plurality of original samples containing the target text unit in the original sample set over a plurality of categories corresponding to the original sample set is unbalanced; a first determination unit 720 configured to determine a target category in the plurality of categories for which an augmented sample is to be generated for the target text unit; a first acquisition unit 730 configured to acquire a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and a first generation unit 740 configured to generate, based on the target text unit and the first semantic scene rule, a first augmented sample for training the text classification model using a large model, the first augmented sample containing the target text unit and conforming to the first semantic scene.
[0149] The operations performed by the units 710-740 in the above-mentioned sample generation apparatus 700 and the effects that can be achieved are similar to steps S201-S204 in the above-mentioned sample generation method, and will not be repeated here.
[0150] In some embodiments, the identifying unit can include: a first obtaining subunit configured to perform a semantic-based word segmentation operation on each original sample in the original sample set to obtain a plurality of first segmented words; a first performing subunit configured to, for each first segmented word in the plurality of first segmented words, perform the following operations: based on the occurrence frequency of the first segmented word in each category in the plurality of categories, calculate a first bias score indicating the bias degree of the first segmented word between the categories; and in response to judging that the first bias score is greater than a first preset threshold, determine the first segmented word as a first candidate text unit; and a first determining subunit configured to determine a target text unit in the original sample set based on the first candidate text unit in the plurality of first segmented words.
[0151] In some embodiments, the identifying unit can further include: a second obtaining subunit configured to perform a sliding window-based word segmentation operation on each original sample in the original sample set to obtain a plurality of second segmented words; a second performing subunit configured to, for each second segmented word in the plurality of second segmented words, perform the following operations: based on the occurrence frequency of the second segmented word in each category in the plurality of categories, calculate a second bias score indicating the bias degree of the second segmented word between the categories; and in response to judging that the second bias score is greater than a second preset threshold, determine the second segmented word as a second candidate text unit; and wherein the first determining subunit is further configured to determine the target text unit in the original sample set based on the first candidate text unit in the plurality of first segmented words and the second candidate text unit in the plurality of second segmented words.
[0152] In some embodiments, in response to judging that the second bias score is greater than the second preset threshold, determining the second segmented word as the second candidate text unit can include: in response to judging that the second bias score is greater than the second preset threshold, and in response to judging that the total frequency of occurrence of the second segmented word in the original sample set is greater than a third preset threshold, determining the second segmented word as the second candidate text unit.
[0153] In some embodiments, the first determining unit can include: a second determining subunit configured to determine the category with the highest occurrence frequency of the target text unit in the plurality of categories as a dominant category corresponding to the target text unit; and a third determining subunit configured to determine the categories other than the dominant category in the plurality of categories as target categories.
[0154] In some embodiments, determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first segmented words and the second candidate text unit in the plurality of second segmented words can include: for each of the first candidate text unit in the plurality of first segmented words and the second candidate text unit in the plurality of second segmented words, performing the following operations: for each original sample in the original sample set containing the candidate text unit, extracting a context segment containing the candidate text unit based on a preset length context window; determining a similarity of respective context comprehensive semantic vectors of the candidate text unit in the dominant category and the target category, wherein the context comprehensive semantic vector is determined based on semantic vectors of respective context segments corresponding to the candidate text unit in the corresponding category; and in response to judging that the similarity is less than a preset similarity threshold, determining the candidate text unit as the target text unit.
[0155] In some embodiments, the second obtaining subunit can be further configured to: respectively utilize sliding windows with a window size of a plurality of different preset character lengths to perform text sequence cutting on each original sample in the original sample set to obtain the plurality of second segmented words.
[0156] In some embodiments, the sample generation apparatus described above can further include: a second obtaining unit configured to obtain other categories in the plurality of categories except for the target category; a third obtaining unit configured to obtain a second semantic scene rule of the other categories corresponding to the target text unit, wherein the second semantic scene rule is used to define a second semantic scene in which the target text unit belongs to the other categories; and wherein the first generation unit can be further configured to: based on the target text unit, the first semantic scene rule and the second semantic scene rule, utilize the large model to generate a first sample pair for training the text classification model, the first sample pair including a first augmented sample and a second augmented sample, the second augmented sample containing the target text unit and conforming to the second semantic scene.
[0157] In some embodiments, the sample generation apparatus described above can further include: a construction unit configured to, for the target text unit, construct a first prompt text, wherein the first prompt text is used to instruct the large model to generate an augmented sample belonging to the target category, and is used to instruct the large model to reflect, in the generated augmented sample, a difference between a context semantic of the target text unit in the target category and a context semantic of the target text unit in the other categories; and a second generation unit configured to, based on the first prompt text, utilize the large model to generate a third augmented sample for training the text classification model.
[0158] In some embodiments, the target category can include a first subcategory and a second subcategory, and the first prompt text can be further configured to instruct the large model to generate a second sample pair respectively belonging to the first subcategory and the second subcategory, and the third augmented sample includes the second sample pair.
[0159] In some embodiments, the sample generation apparatus described above can further include a third generation unit configured to, for the generated enhanced samples, rewrite the generated enhanced samples based on a second prompt text by using the large model to obtain extended enhanced samples, wherein the generated enhanced samples at least include the first enhanced sample, and the second prompt text is used to instruct the large model to generate the extended enhanced samples by at least one preset rewriting manner while keeping the overall semantics of the generated enhanced samples unchanged.
[0160] In some embodiments, the at least one preset rewriting manner includes causing different versions of the extended enhanced samples generated based on the generated enhanced samples to correspond to different word piece segmentation sequences.
[0161] In some embodiments, a training apparatus of a text classification model is also provided, which is configured to train the text classification model based on the enhanced samples, wherein the enhanced samples are obtained based on the sample generation method of the present disclosure.
[0162] The operations performed by the training apparatus of the text classification model described above and the effects that can be achieved are similar to those of the training method of the text classification model described above, and will not be repeated here.
[0163] In some embodiments, the training apparatus can include an acquisition unit configured to acquire positive example outputs and negative example outputs corresponding to the enhanced samples based on the enhanced samples, and a training unit configured to train the text classification model based on the enhanced samples, the positive example outputs, and the negative example outputs.
[0164] In some embodiments, the text classification model can include a semantic feature extraction network and a classification network, and the training unit can be further configured to input the enhanced samples, the positive example outputs, and the negative example outputs into the semantic feature extraction network respectively to obtain enhanced sample semantic vectors, positive example output semantic vectors, and negative example output semantic vectors, calculate a contrastive loss based on the enhanced sample semantic vectors, the positive example output semantic vectors, and the negative example output semantic vectors, and adjust parameters of the semantic feature extraction network based on the contrastive loss.
[0165] In some embodiments, the text classification model can be a generative decoding model, and the training unit can be further configured to input the enhanced samples and the positive example outputs into the text classification model to obtain first prediction outputs of the text classification model, input the enhanced samples and the negative example outputs into the text classification model to obtain second prediction outputs of the text classification model, calculate a direct preference loss based on the first prediction outputs and the second prediction outputs, and train the text classification model based on the direct preference loss.
[0166] According to embodiments of the present disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0167] Reference Figure 8 A block diagram of the structure of an electronic device 800, which is an example of a hardware device that can be applied to aspects of the present disclosure, will now be described, which can serve as a server or a client of the present disclosure. The electronic device is intended to represent a wide variety of digital electronic computer devices, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent a wide variety of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0168] As Figure 8 shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0169] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, an output unit 807, the storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information to the electronic device 800, can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0170] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 801 performs various methods and processes described above, such as the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure. For example, in some embodiments, the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the sample generation method of the present disclosure or the training method of the text classification model of the present disclosure by any other appropriate means, such as by means of firmware.
[0171] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0172] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0173] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0174] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0175] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0176] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0177] It should be understood that various forms of flow shown above can be used with orders of steps reordered, added to, or deleted from. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure are achieved, which are not limited herein.
[0178] While embodiments or examples of the present disclosure have been described with reference to the figures, it is understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but only by the claims and equivalents thereof. Various elements in the embodiments or examples can be omitted or replaced by equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples can be combined in various ways. It is important that many of the elements described herein can be replaced by equivalent elements that appear after the present disclosure as technology evolves.
Claims
1. A sample generation method, comprising: identifying a target text unit from an original sample set used for training a text classification model, wherein a plurality of original samples in the original sample set containing the target text unit are unevenly distributed over a plurality of categories corresponding to the original sample set; determining a target category in the plurality of categories for which an augmented sample is to be generated for the target text unit; obtaining a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and generating, based on the target text unit and the first semantic scene rule, a first augmented sample for training the text classification model using a large model, the first augmented sample containing the target text unit and conforming to the first semantic scene.
2. The method of claim 1, wherein, The identifying a target text unit from an original sample set used for training a text classification model comprises: performing a semantic-based segmentation operation on each original sample in the original sample set to obtain a plurality of first segments; for each first segment in the plurality of first segments, performing the following operations: calculating, based on the occurrence frequency of the first segment in each category in the plurality of categories, a first bias score indicating the bias degree of the first segment between categories; and in response to determining that the first bias score is greater than a first preset threshold, determining the first segment as a first candidate text unit; and determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first segments.
3. The method of claim 2, wherein, The identifying a target text unit from an original sample set used for training a text classification model further comprises: performing a sliding window-based segmentation operation on each original sample in the original sample set to obtain a plurality of second segments; for each second segment in the plurality of second segments, performing the following operations: calculating, based on the occurrence frequency of the second segment in each category in the plurality of categories, a second bias score indicating the bias degree of the second segment between categories; and in response to determining that the second bias score is greater than a second preset threshold, determining the second segment as a second candidate text unit; and wherein The determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first segments comprises: determining the target text unit in the original sample set based on the first candidate text unit in the plurality of first segments and the second candidate text unit in the plurality of second segments.
4. The method of claim 3, wherein, The determining the second segment as a second candidate text unit in response to determining that the second bias score is greater than a second preset threshold comprises: in response to determining that the second bias score is greater than a second preset threshold, and in response to determining that the total frequency of occurrence of the second segment in the original sample set is greater than a third preset threshold, determining the second segment as a second candidate text unit.
5. The method of claim 3 or 4, wherein, The determining a target category in the plurality of categories for which an augmented sample is to be generated for the target text unit comprises: determine a category with a highest occurrence frequency of the target text unit in the plurality of categories as a dominant category corresponding to the target text unit; and determine a category other than the dominant category in the plurality of categories as the target category.
6. The method of claim 5, wherein, determining the target text unit in the original sample set based on a first candidate text unit in the plurality of first segmented words and a second candidate text unit in the plurality of second segmented words includes: for each candidate text unit in the first candidate text unit in the plurality of first segmented words and the second candidate text unit in the plurality of second segmented words, the following operations are performed: for each original sample containing the candidate text unit in the original sample set, a context segment containing the candidate text unit is extracted based on a context window of a preset length; determine the similarity of the context comprehensive semantic vectors of the candidate text unit in the dominant category and the target category, respectively, wherein the context comprehensive semantic vector is determined based on the semantic vectors of each context segment corresponding to the candidate text unit in the corresponding category; and in response to determining that the similarity is less than a preset similarity threshold, the candidate text unit is determined as the target text unit.
7. The method of any one of claims 3-6, wherein, the sliding window-based segmentation operation is performed on each original sample in the original sample set to obtain a plurality of second segmented words includes: text sequence truncation is performed on each original sample in the original sample set by using sliding windows with a plurality of different preset character lengths, respectively, to obtain the plurality of second segmented words.
8. The method of any one of claims 1-7, further comprising: obtaining other categories in the plurality of categories other than the target category; obtaining a second semantic scene rule corresponding to the other categories of the target text unit, wherein the second semantic scene rule is used to define a second semantic scene in which the target text unit belongs to the other categories; and wherein the first enhanced sample for training the text classification model is generated by using the large model based on the target text unit and the first semantic scene rule includes: based on the target text unit, the first semantic scene rule and the second semantic scene rule, a first sample pair for training the text classification model is generated by using the large model, the first sample pair includes the first enhanced sample and a second enhanced sample, and the second enhanced sample contains the target text unit and conforms to the second semantic scene.
9. The method of any one of claims 1-8, further comprising: for the target text unit, a first prompt text is constructed, wherein the first prompt text is used to instruct the large model to generate an enhanced sample belonging to the target category, and is used to instruct the large model to reflect the difference between the context semantics of the target text unit in the target category and the context semantics of the target text unit in other categories in the generated enhanced sample; and based on the first prompt text, a third enhanced sample for training the text classification model is generated by using the large model.
10. The method of claim 9, wherein, The target category includes a first subcategory and a second subcategory, the first prompt text is further configured to instruct the large model to generate a second sample pair respectively belonging to the first subcategory and the second subcategory, and the third enhanced sample includes the second sample pair.
11. The method of any one of claims 1-10, further comprising: for the generated enhanced sample, rewriting the generated enhanced sample by the large model based on a second prompt text to obtain an extended enhanced sample, wherein the generated enhanced sample at least includes the first enhanced sample, and the second prompt text is used to instruct the large model to generate the extended enhanced sample by at least one preset rewriting manner while keeping the overall semantics of the generated enhanced sample unchanged.
12. The method of claim 11, wherein, The at least one preset rewriting manner includes making different versions of the extended enhanced sample generated based on the generated enhanced sample correspond to different word piece segmentation sequences.
13. A method for training a text classification model, the method comprising: training the text classification model based on an enhanced sample, wherein the enhanced sample is obtained based on the sample generation method of any one of claims 1-12.
14. The method of claim 13, wherein, The training of the text classification model based on the enhanced sample includes: obtaining positive example output and negative example output corresponding to the enhanced sample based on the enhanced sample; and training the text classification model based on the enhanced sample, the positive example output and the negative example output.
15. The method of claim 14, wherein, The text classification model includes a semantic feature extraction network and a classification network, and the training of the text classification model based on the enhanced sample, the positive example output and the negative example output includes: inputting the enhanced sample, the positive example output and the negative example output into the semantic feature extraction network respectively to obtain an enhanced sample semantic vector, a positive example output semantic vector and a negative example output semantic vector; calculating a contrastive loss based on the enhanced sample semantic vector, the positive example output semantic vector and the negative example output semantic vector; and adjusting parameters of the semantic feature extraction network based on the contrastive loss.
16. The method of claim 14, wherein, The text classification model is a generative decoding model, and the training of the text classification model based on the enhanced sample, the positive example output and the negative example output includes: inputting the enhanced sample and the positive example output into the text classification model to obtain a first prediction output of the text classification model; inputting the enhanced sample and the negative example output into the text classification model to obtain a second prediction output of the text classification model; calculating a direct preference loss based on the first prediction output and the second prediction output; and training the text classification model based on the direct preference loss.
17. A sample generation apparatus, comprising: an identification unit configured to identify a target text unit from an original sample set used to train a text classification model, wherein a plurality of original samples in the original sample set containing the target text unit are unevenly distributed over a plurality of categories corresponding to the original sample set; a first determining unit configured to determine a target category to be used to generate an augmented sample for the target text unit from the plurality of categories; a first obtaining unit configured to obtain a first semantic scene rule corresponding to the target category of the target text unit, wherein the first semantic scene rule is used to define a first semantic scene in which the target text unit belongs to the target category; and a first generating unit configured to generate, based on the target text unit and the first semantic scene rule, a first augmented sample for training the text classification model using a large model, the first augmented sample containing the target text unit and conforming to the first semantic scene.
18. A device for training a text classification model, the device being configured to: training the text classification model based on the augmented samples, wherein, the augmented sample is obtained based on the sample generation method of any one of claims 1-12.
19. An electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-16.
20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer execute the method according to any one of claims 1-16.
21. A computer program product comprising a computer program, wherein, The computer program, when executed by the processor, implements the method of any one of claims 1-16. The computer program, when executed by the processor, implements the method of any one of claims 1-16.
Citation Information
Cited By
Psychological state text classification method based on large model generative data enhancement
CN121935378A