Large language model generated text detection method, system and equipment and storage medium
By constructing a detector that includes a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module, the poor generalization problem of existing detectors in unseen LLM and diverse scenarios is solved, achieving better text detection results.
Patent Information
- Application Number
- CN202510793789.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing large language model generated text detectors have poor generalization when dealing with unseen LLMs and diverse scenarios, and have difficulty effectively distinguishing between human-written and LLM-generated texts.
An LLM-generated text detector is constructed, which includes a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module. By learning domain-invariant features and enhancing feature extraction with dynamic perturbations, the robustness and generalization ability of the detector are improved.
The detection performance of the detector in unseen LLMs and diverse scenarios is significantly improved, the adaptability to text noise is enhanced, and the generalization ability across scenarios and LLMs is improved.
Smart Images

Figure CN120705320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security, and in particular to a method, system, device, and storage medium for detecting text generated by a large language model. Background Art
[0002] The task of detecting text generated by large language models (LLMs) aims to identify and distinguish content generated by these models from human-written content. This task is generally classified as a binary classification problem. LLM-generated text falls into the broader category of machine-generated text, which encompasses all text produced by automated systems, including simple language models and rule-based systems. However, with the rapid development of LLM technology, LLM-generated text has become increasingly important in the field of machine-generated text.
[0003] With the widespread application of LLMs in fields such as creative writing, programming, education, and office work, the highly realistic texts they generate are rapidly flooding the internet, raising a series of important societal issues, including false and misleading content on social media, spam and phishing attacks, and potential threats to academic integrity posed by generated scientific papers. Traditionally, evaluating the quality of machine-generated text relies on human judgment. However, Dugan et al., in their article "Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text. AAAI 2023," found that distinguishing between LLM-generated text and human-written text is a significant challenge for humans. Research shows that untrained human reviewers can typically identify text generated by state-of-the-art LLMs with an accuracy close to random guessing. Even experienced researchers and faculty members only achieve a success rate of approximately 50% when identifying LLM-generated scholarly work. Furthermore, as technology continues to advance, it is becoming increasingly difficult for human readers to distinguish between human-authored and LLM-generated content. This trend highlights the need to develop automated systems to detect LLM-generated text.
[0004] Mainstream large language model-generated text detectors can be divided into zero-shot based methods and training-based methods.
[0005] Zero-shot based methods are usually based on statistics and can identify LLM-generated text without additional training. Researchers use various statistical measurement methods for detection, including entropy, perplexity, average log probability score, fluency, etc., as well as n-grams, unified information density (UID), log-rank information and various language features, such as part-of-speech determiners, conjunctions, auxiliary relations, vocabulary, emotional tone, etc. Various zero-shot detection methods provide different strategies to improve detection accuracy and efficiency. Galle et al. introduced an unsupervised method in the article "Unsupervised and distributional detection of machine-generated text.arXiv 2021" to identify the excessive occurrence of repeated high-order n-grams in machine text, thereby distinguishing them from human-generated text. Mitchell et al. found in the article "Detectgpt: Zero-shot machine-generated text detection using probabilitycurvature.ICML 2023" (DetectGPT) that machine-generated text is often associated with areas of negative curvature in the LLM log probability function. Based on this insight, the authors proposed a text perturbation method to measure the log probability difference between the original text and the perturbed text. If a consistent positive difference occurs, it indicates that the text is generated by LLM. Bao et al. improved the above method in the article "Fast-detectgpt: Efficient zero-shot detection of machine-generated text via conditional probability curvature.ICLR 2024" (Fast-DetectGPT), which improves efficiency by eliminating the need for perturbation analysis by using conditional probability curvature. In addition, Krishna et al. proposed a detector in the article "Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. NeurIPS2023" that uses information retrieval to store LLM outputs in a database and search for semantically similar content to identify LLM-generated text, but this approach raises privacy concerns about storing user conversations.
[0006] Zero-shot methods typically require access to model weights to calculate the probability value of text, which limits their ability to cope with closed-source LLMs. To address this problem, researchers often use proxy language models as an alternative. However, this proxy method has a significant disadvantage: it is highly sensitive to the selected proxy model, resulting in large differences in performance on different proxy models. In addition, compared with training-based methods, zero-shot methods generally perform less than ideally in multiple benchmarks. Therefore, in practical applications, how to overcome these limitations and improve the robustness and effectiveness of zero-shot methods remains an urgent problem to be solved.
[0007] Training-based methods primarily use supervised learning to train data generated by a specific LLM. Early work, such as Ippolito et al.'s paper "Automatic detection of generated text is easiest when humans are fooled.arXiv 2019," demonstrated the effectiveness of detectors based on BERT (BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.NAACL 2019) in distinguishing between human-written text and text generated by GPT-2 (Language Models are Unsupervised MultitaskLearners.OpenAI Blog2019). To address the potential social risks of GPT-2-generated text, the OpenAI team fine-tuned RoBERTa (RoBERTa: A Robustly Optimized BERTPretraining Approach.arXiv 2019) to construct a detector. Zhan et al. also fine-tuned RoBERTa-large to detect text generated by ChatGPT in their paper "G3detector: General GPT-generated text detector.arXiv 2023." In the article "Gpt-sentinel: Distinguishing human and chatgpt generated content.arXiv 2023", Chen et al. also used their own collected datasets to train models such as RoBERTa to detect text generated by ChatGPT.In addition, many works dedicated to building benchmarks, such as Guo et al.'s "How close is chatgpt to human experts? comparison corpus, evaluation, and detection.arXiv2023" (HC3), Yu et al.'s "Cheat: A largescale dataset for detecting chatgpt-written abstracts.TBD 2025" (CHEAT), Li et al.'s "Mage: Machine-generated text detection in the wild.ACL 2024" (MAGE), He et al.'s "Mgtbench: Benchmarking machine-generated text detection.CCS2024" (MGTBench), and Wang et al.'s "M4: Multi-generator, multi-domain, and multi-lingual black-box machine-generated text detection.EACL 2024" (M4), have also explored the effectiveness of supervised training detectors. However, these detectors are generally not designed for generalization and therefore suffer from poor generalization capabilities.
[0008] With the emergence of new LLMs and their improved capabilities, it is becoming increasingly difficult to adapt LLM-generated text detectors to new models and scenarios. Bhattacharjee et al. provide a domain adaptation framework for LLM-generated text detection in the article "Conda: Contrastive domain adaptation for AI-generated text detection. IJCNLP-AACL 2023" (CONDA), but it still requires unlabeled data from the target domain. Their follow-up work "Eagle: A domain generalization framework for AI-generated text detection. arXiv2024" (EAGLE) proposes a domain generalization method applicable to multiple LLMs, but ignores the importance of cross-scenario generalization.
[0009] In view of this, the present invention is proposed. Summary of the Invention
[0010] The purpose of the present invention is to provide a method, system, device and storage medium for detecting text generated by a large language model, aiming to enhance generalizable features, thereby improving the ability of LLM generated text detectors in processing unseen LLMs and diverse scenarios.
[0011] The purpose of the present invention is achieved through the following technical solutions:
[0012] A method for detecting text generated by a large language model, comprising:
[0013] Build an LLM-generated text detector consisting of a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module; where LLM is a large language model;
[0014] The input text is extracted with text features through a text encoder and a prediction result is output through a classification head, and a binary classification loss is calculated by combining the prediction results; the text features of the LLM-generated text are screened out from the text features extracted by the text encoder through an LLM conditional feature alignment module, and the similarity loss is calculated by narrowing the distance between the text features of different LLM-generated texts; the input text is enhanced through a dynamic contrastive learning module, and the enhanced features are extracted through the text encoder, and the contrast loss is calculated by combining the enhanced features with the corresponding text features, and the prediction result corresponding to the enhanced features is output through the classification head, and the enhanced binary classification loss is calculated; all the calculated losses are combined to construct a total training loss, and the LLM-generated text detector is trained;
[0015] After training, the text to be detected is input into the LLM to generate a text detector, and the prediction results are output through the text encoder and classification head.
[0016] A large language model generated text detection system, used to implement the aforementioned method, includes:
[0017] A detector construction unit, which is used to construct an LLM-generated text detector including a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module; where LLM is a large language model;
[0018] The training unit is configured to input text, extract text features through a text encoder, output prediction results through a classification head, and calculate a binary classification loss based on the prediction results; screen out text features of LLM-generated text from the text features extracted by the text encoder through an LLM conditional feature alignment module, and calculate a similarity loss by narrowing the distance between text features of different LLM-generated texts; enhance the input text through a dynamic contrast learning module, extract enhanced features through the text encoder, calculate a contrast loss based on the enhanced features and corresponding text features, and output a prediction result corresponding to the enhanced features through a classification head, and calculate an enhanced binary classification loss; construct a total training loss based on all the calculated losses, and train the LLM-generated text detector;
[0019] The detection unit is used to input the text to be detected into the LLM to generate a text detector after training, and output the prediction result through the text encoder and classification head.
[0020] A processing device comprising: one or more processors; a memory for storing one or more programs;
[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0022] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.
[0023] It can be seen from the technical solution provided by the present invention that an LLM conditional feature alignment module is designed to guide the generative text detector to learn domain-invariant features for LLM-generated text. By ensuring that the learned features remain consistent in the text content of different scenarios, this method significantly enhances the generalization of the LLM-generated text detector. In addition, in order to further improve the adaptability of the generative text detector to text noise, a dynamic contrast learning module is introduced, that is, the data is dynamically perturbed during the training process, prompting the generative text detector to pay more attention to robust features, so as to better cope with changes and complexities in the real world. In general, the present invention can solve the problem of poor generalization of the LLM-generated text detector and enhance the detector's ability to handle unseen LLMs and diverse scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 A flowchart of a method for generating text detection using a large language model provided by an embodiment of the present invention;
[0026] Figure 2 A schematic diagram of a large model generating text detection task provided by an embodiment of the present invention;
[0027] Figure 3 The overall architecture diagram of the LLM-generated text detector provided by an embodiment of the present invention;
[0028] Figure 4 Schematic diagrams of RoBERTa (left) and XLNet (right) provided for embodiments of the present invention;
[0029] Figure 5 This is a diagram of the LLM conditional feature alignment module architecture provided by an embodiment of the present invention;
[0030] Figure 6 This is a diagram of the architecture of the dynamic comparative learning module provided by an embodiment of the present invention;
[0031] Figure 7 A schematic diagram showing the comparison results between the MLS dataset provided by an embodiment of the present invention and the existing LLM-generated text detection dataset;
[0032] Figure 8 Schematic diagram of an ablation experiment with different loss function coefficients provided by an embodiment of the present invention;
[0033] Figure 9 t-SNE visualization diagram of the representations learned by the pre-trained model, baseline model, and detector of the present invention provided in an embodiment of the present invention;
[0034] Figure 10 A schematic diagram of a large language model-generated text detection system provided by an embodiment of the present invention;
[0035] Figure 11 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0037] First, the following terms may be used in this article:
[0038] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.
[0039] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.
[0040] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," and "fixed" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this document based on specific circumstances.
[0041] The following describes in detail the method, system, device, and storage medium for detecting text generated by a large language model provided by the present invention. Any information not described in detail in the embodiments of the present invention is prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, the procedures are performed in accordance with conventional conditions in the art or the conditions recommended by the manufacturer. Instruments used in the embodiments of the present invention, where the manufacturer is not specified, are all commercially available conventional products.
[0042] Example 1
[0043] The embodiment of the present invention provides a method for detecting text generated by a large language model. Figure 1 As shown, it mainly includes the following steps:
[0044] Step 1: Build an LLM to generate text detector.
[0045] In an embodiment of the present invention, the LLM generated text detector includes: a text encoder, a classification head, an LLM conditional feature alignment module and a dynamic contrast learning module.
[0046] Step 2: Train the LLM to generate a text detector.
[0047] In this embodiment of the present invention, a training set containing a large amount of input text (including LLM-generated text and manually written text) is collected in advance, and the loss is calculated based on the output of each module in the LLM-generated text detector, and the LLM-generated text detector is trained accordingly. Specifically:
[0048] (1) Input text is passed through the text encoder to extract text features and output prediction results through the classification head, and the binary classification loss is calculated based on the prediction results.
[0049] Use the pre-trained language model PLM as the text encoder, extract the text feature f from the input text x, output the prediction result through the classification head H, and calculate the binary classification loss in combination with the classification label, which is expressed as:
[0050] f=PLM(x)
[0051]
[0052] in, represents the cross entropy loss function, y bcls is the classification label, L bcls is the binary classification loss.
[0053] (2) The text features of LLM-generated text are screened out from the text features extracted by the text encoder through the LLM conditional feature alignment module, and the similarity loss is calculated by shortening the distance between the text features of different LLM-generated texts.
[0054] Combined with the classification labels corresponding to the input text, the text features of the LLM generated text are screened out: f llmc =f[y bcls =1], where f represents the text feature, y bcls =1 means the classification label is 1, that is, the corresponding input text is LLM generated text;
[0055] Define the number of input text pairs of different scene categories as N, and the sum of the cosine similarities between all input text pairs of different scene categories as S:
[0056]
[0057] Among them, cos is cosine similarity; It is an indicator function. If the condition in the brackets is met, the output is 1, otherwise the output is 0. Represents the labels of n different scene categories, i, j are the indexes of the input text, b is the size of a batch, These are text features.
[0058] The similarity loss is calculated as follows:
[0059]
[0060] in, is the similarity loss.
[0061] (3) The input text is enhanced through the dynamic contrastive learning module, and the enhanced features are extracted through the text encoder. The contrast loss is calculated by combining the enhanced features with the corresponding text features. In addition, the prediction results corresponding to the enhanced features are output through the classification head, and the enhanced binary classification loss is calculated.
[0062] For input text x i After word segmentation and adding the tag [CLS], randomly select the words that need to be replaced, and obtain the corresponding synonyms based on the dictionary, and use the synonyms to replace the corresponding words to obtain enhanced text
[0063] For input text x i , select Enhanced Text As a positive sample, all other input texts in the same batch are regarded as negative samples; each input text and enhanced sample obtains the corresponding text features and enhanced features through the text encoder respectively, and calculates the cosine similarity s after dimensionality reduction through the projection layer P. ij With s ii , and the exponential similarity z ij With z ii , expressed as:
[0064] s ij =cos(P(f i ),P(f j ))
[0065]
[0066] Among them, cos is the cosine similarity function, exp is the natural exponential function, and f i ,f j Represents the input text x i ,x j Text features, i, j are the indexes of the input text, To enhance the text Enhanced features, s ij is the text feature f i ,f j The cosine similarity of s ii is the text feature f i and enhanced features The cosine similarity of , t is the temperature parameter.
[0067] The contrast loss is calculated as follows:
[0068]
[0069] Where b is the size of a batch, is the contrast loss.
[0070] (4) Combine all the calculated losses to construct the total training loss and train the LLM-generated text detector.
[0071] The total training loss is expressed as:
[0072]
[0073] in, is the total training loss, is the binary classification loss, is the similarity loss, is the contrast loss, α and β are the weight coefficients of the corresponding loss; To enhance the binary classification loss, it is expressed as:
[0074] Considering that the relevant training process can be implemented by referring to conventional technology, it will not be described in detail.
[0075] Step 3: After training, the text to be detected is input into the LLM to generate a text detector, and the prediction results are output through the text encoder and classification head.
[0076] After completing the training based on the above step 2, the text features of the text to be detected are extracted by the text encoder, and then the classification head outputs the prediction results, that is, the probability of belonging to the LLM generated text category and the probability of belonging to the manually written text category.
[0077] In view of the obvious deficiencies of existing datasets, the present invention constructs a dataset MLS for evaluating the trained LLM-generated text detector; wherein, data is obtained from NLP datasets, including: a story generation dataset XSum, a news writing dataset Wp, a scientific writing dataset SciGen, a financial question-answering dataset Finance, and a medical question-answering dataset Medicine; and a portion of the data in the story generation dataset XSum, the news writing dataset Wp, and the scientific writing dataset SciGen is translated into Chinese, and the above datasets and the text data obtained by translation are used as manually written text data; then, the LLM is guided by prompts to generate the LLM-generated text corresponding to each manually written text; all the manually written text data and the LLM-generated text are combined to form the dataset MLS.
[0078] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.
[0079] 1. Overall overview of the plan.
[0080] The core objective of this invention is to improve the generalization capability of LLM-generated text detectors in unseen LLMs and scenarios by designing a detector based on feature alignment and contrastive learning. The detector model designed in this invention is versatile and can adapt to different input texts, addressing the poor generalization problem of previous detectors. Specifically, this invention designs a module that aligns the features of LLM-generated texts in different scenarios, guiding the detector to learn domain-invariant features specific to LLM-generated texts, encouraging the increase in similarity among these text features, thereby achieving more generalizable feature extraction. This approach significantly enhances the detector's generalization. To further improve the detector's adaptability to textual noise, contrastive learning based on dynamic text perturbation is introduced. This method dynamically perturbs the data during training, reducing the distance between each text item and its perturbed text and increasing the distance between each text item and other text items. This results in more robust feature extraction and better adaptability to the changes and complexity of the real world. Based on the above ideas, these innovations are integrated into two key modules: the LLM-conditional feature alignment module and the dynamic contrastive learning module. Together with the text encoder and classification head, they form a novel LLM-generated text detector.
[0081] In addition, considering the rapid development of LLM, it is particularly important to use data that reflects real-world scenarios to effectively evaluate detection methods. Existing datasets have obvious shortcomings: they either fail to fully cover multiple LLMs, multiple languages, and diverse application scenarios, or rely on outdated LLM-generated data, which is difficult to meet current needs. In order to solve these problems, the present invention constructs a new dataset MLS. The MLS dataset covers Chinese and English data generated by ten cutting-edge LLMs from five different scenarios, including closed-source models with larger parameters and open-source models with smaller parameters, thereby achieving more comprehensive coverage. Finally, experiments on the MLS dataset demonstrate the effectiveness of the method of the present invention.
[0082] 2. Detailed introduction of the plan.
[0083] 1. An overall introduction to the LLM generated text detector.
[0084] like Figure 2Figure 1 shows a diagram of the text detection task using a large language model. The generating task obtains text of the corresponding category through human writing and large language model generation. The detector inputs text from unknown sources and outputs the text category.
[0085] like Figure 3 As shown in the figure, the overall framework of the LLM-generated text detector provided by the present invention mainly includes: a text encoder (Encoder), a classification head (Classifier), an LLM conditional feature alignment module (LCFA) and a dynamic contrastive learning module (DCL). Given a batch of input text x, it is first converted into word element embeddings by a pre-trained word segmenter, and then sent to a pre-trained text encoder to extract semantic features. Finally, the [CLS] embedding output by the text encoder is used as the text feature f. The text feature f is passed to the classification head to predict whether the input text is generated by LLM.
[0086] The LLM conditional feature alignment module and dynamic contrastive learning module introduced in the framework can help further optimize feature representation. The following briefly describes the functions of each module.
[0087] (1) Text encoder and classification head.
[0088] The text encoder is responsible for extracting high-level semantic features from the input text, while the classification head maps these features to binary classification output to determine whether the text is generated by LLM.
[0089] (2)LLM conditional feature alignment module.
[0090] This module receives text features f and filters them out from the LLM. Since these texts may come from different scenes, the module narrows the distance between text pairs from different scenes and learns domain-invariant features, thereby solving the cross-scene generalization problem and mitigating the degradation of detection performance in unseen scenes.
[0091] (3) Dynamic comparative learning module.
[0092] For the input text x, some words in each text are replaced with synonyms to generate enhanced text x with similar semantics but different expressions. aug The dynamic contrastive learning module reuses the same text encoder and converts x aug Convert to enhanced feature f aug Subsequently, the feature representation is optimized based on the contrastive learning method in SimCLR. In addition, the enhanced feature f augIt is also fed into the classification head to predict whether it is generated by LLM.
[0093] 2. LLM generates the total training loss of the text detector.
[0094] The total training loss mainly consists of several parts:
[0095] Binary classification loss: used to optimize the predictive power of the text encoder and classification head.
[0096] Similarity loss: ensures that the LLM conditional feature alignment module can narrow the distance between cross-scene text features.
[0097] Contrastive Loss: Enhancing feature robustness through a dynamic contrastive learning module.
[0098] Enhanced binary classification loss: Also used to optimize the predictive power of the text encoder and classification head.
[0099] These losses are weightedly combined to achieve multi-task joint optimization.
[0100] Overall, the detector's synergy between the LLM conditional feature alignment module and the dynamic contrastive learning module enhances its adaptability to LLM-generated text and its generalization performance in complex application scenarios. This design significantly improves the detector's performance in unseen scenarios.
[0101] The following introduces the specific calculation process of each module and related losses.
[0102] (1) Principles of text encoder and classification head and related loss calculation process.
[0103] The text encoder is a core module responsible for converting input text into high-dimensional semantic features, providing the foundation for subsequent classification and feature alignment. Pretrained language models (PLMs) perform well on downstream tasks after fine-tuning, ensuring that the detector captures rich linguistic structure. This paper experiments with two advanced pretrained language models as text encoders: RoBERTa and XLNet (XLNet: Generalized Autoregressive Pretraining for Language Understanding, NeurIPS 2019).
[0104] RoBERTa is an improved pre-trained language model based on BERT. It significantly improves the detector's language understanding capabilities through longer training time, larger batch sizes, and a dynamic masking strategy. It uses a bidirectional Transformer architecture that captures global semantic information in context, making it suitable for complex text classification tasks.
[0105] XLNet is a pre-trained language model based on autoregressive mechanisms. It overcomes the limitation of traditional autoregressive models in capturing bidirectional context through permutation language modeling. XLNet has strong context modeling capabilities when processing text, making it particularly well-suited for capturing complex patterns in text generated by LLM. Figure 4 The masked language modeling principles of RoBERTa and the permutation language modeling principles of XLNet are demonstrated respectively.
[0106] For a batch of input text x, each text is first split into multiple tokens and a special tag [CLS] is added. Subsequently, the text encoder encodes the token sequence and uses the [CLS] embedding output by the text encoder as the text feature f∈R b×h , where b represents the batch size and h represents the dimension of the hidden layer. These features not only contain the semantic information of the input text, but also provide a high-quality representation basis for subsequent modules. The text features are input to the classification head H, which outputs the prediction and calculates the binary classification loss L bcls :
[0107] f=PLM(x)
[0108]
[0109] Among them, the classification label of each sample (input text) 1 means the text comes from LLM, 0 means the text comes from human, represents the cross entropy loss.
[0110] (2) The principle of LLM conditional feature alignment module and the related loss calculation process.
[0111] Previous work (MAGE) showed that generalization across scenes is more challenging for detecting LLM-generated text than across LLMs. This is because text generated by different LLMs for the same prompt often exhibits some correlation, while text generated by the same LLM in different scenes is usually uncorrelated. As a result, detectors trained on a limited set of scenes may experience significant performance degradation when applied to unseen scenes.
[0112] To solve this cross-scenario generalization problem, this paper proposes the LLM Conditional Feature Alignment Module (LCFA), which can learn domain-invariant features. The specific architecture is as follows: Figure 5 To prevent the alignment of human-written text features from different scenarios, which would make it difficult to distinguish between human text and LLM text, the present invention filters out the human-generated text features and only retains the LLM-generated text features. These features are called LLM conditional features f llmc :
[0113] f llmc =f[y bcls =1]
[0114] Define the number of sample pairs of different scene categories as N, and the sum of the cosine similarities between all sample pairs of different scene categories as S:
[0115]
[0116] Similarity loss is calculated as follows:
[0117]
[0118] in, Represents labels for n different scenes.
[0119] By minimizing the similarity loss, LLM-generated text features from different scenes can be aligned, improving the domain invariance of the features and thus enhancing the generalization ability of the detector in unknown scenes.
[0120] (3) The principle of dynamic contrastive learning module and the related loss calculation process.
[0121] Inspired by the previous work CONDA, this paper combines contrastive learning to enhance the robustness of the detector to interference. Contrastive learning improves the generalization ability of representation by comparing similar samples with different samples. In CONDA, synonym replacement is used to perturb data and generate enhanced samples, which are then reused in training. However, this static contrastive learning limits the diversity of enhancement. To solve this problem, this paper dynamically perturbs text data in each training step, allowing the detector to learn more robust text representations. The specific architecture is as follows: Figure 6 shown.
[0122] For each original input text x i First, we use the Jieba library to segment words and process special tags, then randomly select words that need to be replaced, use WordNet to obtain synonyms, and finally perform synonym replacement to obtain enhanced samples.
[0123] The present invention adopts the contrast loss proposed in SimCLR. i ,choose as positive samples, while all other samples in the batch are considered negative samples. i and Get f through the text encoder respectively i and Then they are reduced in dimension by the projection layer P. The cosine similarity s between samples is defined as ij , s iiand exponential similarity z ij , z ii :
[0124] s ij =cos(P(f i ),P(f j ))
[0125]
[0126] Contrastive loss The calculation is as follows:
[0127]
[0128] Where t is the temperature parameter.
[0129] In addition, a batch of enhanced text features f aug Will be used to calculate the enhanced binary classification loss
[0130]
[0131] Combining the above losses, we get the total training loss:
[0132]
[0133] in, is the total training loss, and α and β are the weight coefficients of the corresponding loss.
[0134] 3. Construct MLS dataset.
[0135] With the rapid development of LLMs, it is crucial to evaluate the performance of LLM-generated text detectors on state-of-the-art LLMs. To meet real-world evaluation needs, we constructed the MLS dataset, which covers ten cutting-edge LLMs in five scenarios in English and Chinese. The included models cover closed-source systems (GPT-3.5, GPT-4o, GLM-4, Kimi) and open-source alternatives (LLaMA3-8b, Mistral-7b, Gemma-7b, Qwen2-7b, InternLM2-7b, Deepseek-v2-21b). Figure 7 The number of LLMs used in MLS and the LLM sophistication are compared with a series of previous datasets such as HC3, CHEAT, MAGE, MGTBench, and M4.
[0136] For human-written text, we obtained data from various traditional NLP datasets, including the story generation dataset XSum, the news writing dataset Wp, the scientific writing dataset SciGen, the financial question-answering dataset Finance, and the medical question-answering dataset Medicine. We used an automatic translation tool to translate portions of the XSum, Wp, and SciGen datasets into Chinese.
[0137] In order to generate the corresponding LLM-generated text for each manually written text, the present invention provides multiple prompts for each LLM. For XSum, Wp and SciGen, we follow the method in the previous work MAGE. For English data, the first 30 words are input as prompts, and the LLM is requested to generate the corresponding story, news or scientific text. For Chinese data, it is adjusted to input the first 20 Chinese characters as prompts. For the Finance and Medicine datasets consisting of question-answer pairs, the corresponding questions are used as prompts to let the LLM generate answers. For each closed-source system, queries are performed through its official API (application programming interface), and for open source models, the model is run using the weights in its official Hugging Face. All inference hyperparameters such as temperature, topp and topk follow the default values in the official documents of each company and are not modified.
[0138] The constructed MLS dataset contains 22,000 texts, 11,000 each in English and Chinese. Five domains—story generation, news writing, scientific writing, financial Q&A, and medical Q&A—each contain 4,400 texts. For each human-generated text, 10 corresponding LLM-generated texts are generated, resulting in a 1:10 ratio of human-generated texts to LLM-generated texts.
[0139] 4. Performance verification.
[0140] In order to illustrate the performance of the present invention, performance verification was carried out through experiments.
[0141] The text encoders used in the experiments, RoBERTa and XLNet, were pre-trained on English and Chinese corpora, with weights from the Hugging Face libraries hfl / chinese-roberta-wwm-ext and hfl / chinese-xlnet-base. The loss function coefficients were set to α = 1 and β = 0.1. We followed the setup of the previous work EAGLE, using the AdamW optimizer proposed in Loshchilov et al.'s paper "Decoupled weight decay regularization.arXiv 2017" with a weight decay of 0.02. All models were trained on the training set for three epochs with a learning rate of 2e-5 and a batch size of 32. 10% of the text words were randomly replaced to implement text perturbations.
[0142] 1. Dataset and Evaluation Metrics
[0143] In the experiment, the detector of the present invention was evaluated on the MLS dataset. The dataset was divided into training, validation, and test sets in a ratio of 8:1:1. The generalization ability was evaluated using text generated by GPT-4o, Qwen2, and Deepseek-v2, as well as randomly selected scientific writing texts. Following the settings of previous work such as MAGE and CONDA, the area under the receiver operating characteristic curve (AUC) was used as the evaluation metric, with the higher the AUC, the better. In the experiment, the detector was evaluated in four settings:
[0144] (1) In-Distribution Evaluation: Evaluate text from visible LLMs and scenes.
[0145] (2) Cross-LLM evaluation: evaluate text from seen scenes but without LLM.
[0146] (3) Cross-Scenario Evaluation: Evaluate text from scenes where LLMs have been seen but not seen.
[0147] (4) Cross-LLM & Scenario Evaluation: Evaluate texts without LLM and scenarios.
[0148] 2. Quantitative results.
[0149] Table 1 shows the performance comparison of previous state-of-the-art methods and our method on the MLS dataset.
[0150] Table 1: Performance comparison on the MLS dataset; the best result is in bold, the second best result is underlined
[0151]
[0152] As shown in Table 1, RoBERTa achieves state-of-the-art results on datasets such as HC3 and MGTbench. XLNet is also a popular pre-trained model. EAGLE uses adversarial training to align features between different LLMs, but ignores the importance of cross-scenario generalization.
[0153] It can be seen that the method of the present invention consistently outperforms other methods in all four settings. It is worth noting that the method achieves more than 8% improvement in cross-scene and cross-LLM and scene settings, highlighting its ability to learn generalized features. In addition, it is occasionally observed that the in-distribution results are slightly worse than the cross-LLM results. We believe that this can be explained from two perspectives. First, as mentioned by MAGE, generalization between different LLMs is relatively easy, and the experimental results also show that the AUC of the in-distribution evaluation and the AUC of the cross-LLM evaluation are indeed very close. Second, each model is evaluated after full training, without selecting the model with the best verification effect on the validation set, ensuring a fair performance comparison. Therefore, some models may overfit, resulting in a lower AUC for the in-distribution evaluation.
[0154] 3. Ablation research.
[0155] Ablation studies are conducted using XLNet to verify the effectiveness of the proposed components. Starting from the baseline, additional components are gradually introduced, and each variant is trained and evaluated. The results are shown in Table 2.
[0156] Table 2: Ablation experiments for static contrastive learning (SCL), dynamic contrastive learning (DCL) and LLM conditional features
[0157]
[0158] Adding static contrastive learning (SCL) used in CONDA to the baseline improves performance in both in-distribution and cross-LLM evaluations, but performance remains limited in cross-scene evaluation. Introducing dynamic contrastive learning (DCL) significantly improves performance across scenes and across LLM and scene settings, highlighting the importance of DCL in capturing robust features. Finally, adding LLM conditional feature alignment (LCFA) further improves performance across scenes and across LLM and scene settings, highlighting its role in learning domain-invariant features.
[0159] Table 3 shows the ablation experiment results for different text lengths. It can be observed that the performance reaches saturation when the length is 128. Therefore, following the settings of MAGE, the text length is set to 128 in the experiment.
[0160] Table 3: Ablation experiments for different text lengths
[0161]
[0162] Figure 8 Ablation results for different loss function coefficients are shown. It is observed that the best performance is achieved when α = 1 and β = 0.1. Increasing or decreasing α alone, β alone, and both α and β simultaneously lead to performance degradation compared to the optimal setting.
[0163] 4. Visualization.
[0164] In order to evaluate whether our method effectively captures domain-invariant features, we perform t-SNE visualization on the hidden layer embeddings of text from five different scenarios. The embeddings of the pre-trained model, the baseline model, and our detector are visualized, as shown in the following figure: Figure 9 shown. Figure 9 The visualization results shown show that: for the pre-trained model, features from the same scene are clustered together, while features from different scenes are more scattered. Compared with the baseline model, the detector of the present invention appears closer in the visualization to features from different scenes, which indicates that the method of the present invention effectively learns domain-invariant features, thereby improving the performance of the model in a cross-domain setting.
[0165] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.
[0166] Example 2
[0167] The present invention also provides a large language model generated text detection system, which is mainly used to implement the method provided in the above embodiment, such as Figure 10 As shown, the system mainly includes:
[0168] A detector construction unit, which is used to construct an LLM-generated text detector including a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module; where LLM is a large language model;
[0169] The training unit is configured to input text, extract text features through a text encoder, output prediction results through a classification head, and calculate a binary classification loss based on the prediction results; screen out text features of LLM-generated text from the text features extracted by the text encoder through an LLM conditional feature alignment module, and calculate a similarity loss by narrowing the distance between text features of different LLM-generated texts; enhance the input text through a dynamic contrast learning module, extract enhanced features through the text encoder, calculate a contrast loss based on the enhanced features and corresponding text features, and output a prediction result corresponding to the enhanced features through a classification head, and calculate an enhanced binary classification loss; construct a total training loss based on all the calculated losses, and train the LLM-generated text detector;
[0170] The detection unit is used to input the text to be detected into the LLM to generate a text detector after training, and output the prediction result through the text encoder and classification head.
[0171] Considering that the main technical details involved in the system have been introduced in the previous embodiments, they will not be repeated here.
[0172] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0173] Example 3
[0174] The present invention also provides a processing device, such as Figure 11 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.
[0175] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0176] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:
[0177] The input device can be a touch screen, image acquisition device, physical button or mouse;
[0178] The output device may be a display terminal;
[0179] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0180] Example 4
[0181] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0182] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0183] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.
Claims
1. A method for detecting text generated by a large language model, characterized in that: include: Build an LLM-generated text detector consisting of a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module; where LLM is a large language model; The input text is extracted with text features through a text encoder and a prediction result is output through a classification head, and a binary classification loss is calculated by combining the prediction results; the text features of the LLM-generated text are screened out from the text features extracted by the text encoder through an LLM conditional feature alignment module, and the similarity loss is calculated by narrowing the distance between the text features of different LLM-generated texts; the input text is enhanced through a dynamic contrastive learning module, and the enhanced features are extracted through the text encoder, and the contrast loss is calculated by combining the enhanced features with the corresponding text features, and the prediction result corresponding to the enhanced features is output through the classification head, and the enhanced binary classification loss is calculated; all the calculated losses are combined to construct a total training loss, and the LLM-generated text detector is trained; After training, the text to be detected is input into the LLM to generate a text detector, and the prediction results are output through the text encoder and classification head.
2. The method for detecting text generated by a large language model according to claim 1, characterized in that: The input text is extracted through the text encoder and the prediction result is output through the classification head. The binary classification loss is constructed by combining the prediction results. Use the pre-trained language model PLM as the text encoder, extract the text feature f from the input text x, output the prediction result through the classification head H, and calculate the binary classification loss in combination with the classification label, which is expressed as: f=PLM(x) in, represents the cross entropy loss function, y bcls is the classification label, is the binary classification loss.
3. The method for detecting text generated by a large language model according to claim 1, wherein: The LLM conditional feature alignment module filters out the text features of the LLM-generated text from the text features extracted by the text encoder, and calculates the similarity loss by shortening the distance between the text features of the texts generated by different LLMs. The method includes: Combined with the classification labels corresponding to the input text, the text features of the LLM generated text are screened out: f llmc =f[y bcls =1], where f represents the text feature, y bcls =1 means the classification label is 1, that is, the corresponding input text is LLM generated text; Define the number of input text pairs of different scene categories as N, and the sum of cosine similarities between all input text pairs of different scene categories as S: Among them, cos is cosine similarity; It is an indicator function. If the condition in the brackets is met, the output is 1, otherwise the output is 0. Represents the labels of n different scene categories, i, j are the indexes of the input text, b is the size of a batch, All are text features; The similarity loss is calculated as follows: in, is the similarity loss.
4. The method for detecting text generated by a large language model according to claim 1, wherein: The step of enhancing the input text by the dynamic contrast learning module includes: For input text x i After word segmentation and adding the tag [CLS], randomly select the words that need to be replaced, and obtain the corresponding synonyms based on the dictionary, and use the synonyms to replace the corresponding words to obtain enhanced text 5. The method for detecting text generated by a large language model according to claim 4, characterized in that: The step of extracting enhanced features through a text encoder and calculating contrast loss by combining the enhanced features with corresponding text features includes: For input text x i , select Enhanced Text As a positive sample, all other input texts in the same batch are regarded as negative samples; each input text and enhanced sample obtains the corresponding text features and enhanced features through the text encoder respectively, and calculates the cosine similarity s after dimensionality reduction through the projection layer P. ij With s ii , and the exponential similarity z ij With z ii , expressed as: s ij =cos(P(f i ),P(f j )) Among them, cos is cosine similarity, exp is natural exponential function; f i ,f j Represents the input text x i ,x j Text features, i, j are the indexes of the input text, To enhance the text Enhanced features, s ij is the text feature f i ,f j The cosine similarity of s ii is the text feature f i and enhanced features The cosine similarity of , t is the temperature parameter; The contrast loss is calculated as follows: Where b is the size of a batch, is the contrast loss.
6. The method for detecting text generated by a large language model according to claim 1, characterized in that: The total training loss constructed by combining all the calculated losses is expressed as: in, is the total training loss, is the binary classification loss, is the similarity loss, is the contrast loss, α and β are the weight coefficients of the corresponding loss; To enhance the binary classification loss, it is expressed as: represents the cross entropy loss function, y bcls is the classification label, and H is the classification head.
7. The method for detecting text generated by a large language model according to claim 1, wherein: Also includes: Construct the MLS dataset to evaluate the trained LLM-generated text detector; Data was obtained from NLP datasets, including the story generation dataset XSum, the news writing dataset Wp, the scientific writing dataset SciGen, the financial question-answering dataset Finance, and the medical question-answering dataset Medicine. Part of the data from the story generation dataset XSum, the news writing dataset Wp, and the scientific writing dataset SciGen was translated into Chinese. The above datasets and the translated text data were used as manually written text data. Afterwards, the LLM is guided by prompts to generate LLM-generated text corresponding to each manually written text; All manually written text data and LLM generated text are combined to form the dataset MLS.
8. A large language model generated text detection system, characterized in that The method for implementing any one of claims 1 to 7 comprises: A detector construction unit, which is used to construct an LLM-generated text detector including a text encoder, a classification head, an LLM conditional feature alignment module, and a dynamic contrastive learning module; where LLM is a large language model; The training unit is configured to input text, extract text features through a text encoder, output prediction results through a classification head, and calculate a binary classification loss based on the prediction results; screen out text features of LLM-generated text from the text features extracted by the text encoder through an LLM conditional feature alignment module, and calculate a similarity loss by narrowing the distance between text features of different LLM-generated texts; enhance the input text through a dynamic contrast learning module, extract enhanced features through the text encoder, calculate a contrast loss based on the enhanced features and corresponding text features, and output a prediction result corresponding to the enhanced features through a classification head, and calculate an enhanced binary classification loss; construct a total training loss based on all the calculated losses, and train the LLM-generated text detector; The detection unit is used to input the text to be detected into the LLM to generate a text detector after training, and output the prediction result through the text encoder and classification head.
9. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.