Machine-generated natural language detection method based on multiple features
Through a machine-generated natural language detection method based on multi-features, supervised learning is used to generate detectors, which solves the problems of weak generalization ability and low detection accuracy of existing detection methods, and achieves more efficient and accurate text source detection.
Patent Information
- Application Number
- CN202510226003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing machine-generated natural language detection methods have problems such as weak generalization ability, overfitting, and low detection accuracy, especially when the detection effect declines when facing unopen source language models and language replacement.
A machine-generated natural language detection method based on multi-features is proposed. By obtaining 11 features of the text to be detected, including log likelihood, rank number, log rank number, confusion degree, etc., and these features are synthesized into an 11-dimensional vector input machine learning classification algorithm for supervised learning, and a detector is generated.
It improves the accuracy and generalization ability of the detection method, enhances the robustness and resource-friendliness of the detection, and can more accurately distinguish machine-generated text from human-written text.
Smart Images

Figure CN120067333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically discloses a machine-generated natural language detection method based on multiple features. Background Art
[0002] With the rapid development of large language models (LLMs), their excellent performance has made it a huge challenge for ordinary users to distinguish between machine-generated texts (MGTs) and human-written texts (HWTs). Therefore, the application of large language models has also caused many problems. In actual application scenarios, criminals frequently use large language models to carry out phishing, spread false information, conduct academic fraud, and maliciously generate spam. For example, fraudulent product reviews are fabricated with the help of large language models, which seriously mislead consumers' purchasing decisions; in the academic field, plagiarism and cheating are frequent, which greatly undermines the academic integrity environment and disrupts the normal order of society. At the same time, when Internet users face texts of unknown origin, they often find it difficult to make accurate judgments due to the lack of effective means of discrimination, which provides an opportunity for malicious use of information and poses a potential threat to social stability.
[0003] The current machine-generated natural language detection methods mainly include: detection methods based on statistical features, detection methods based on language models, and detection methods based on watermark algorithms. There are the following defects: For detection methods based on statistical features, although they use pre-trained language models to extract statistical features and use machine learning classification models to carry out supervised learning, it is extremely difficult to obtain the internal weight parameters and network structure of the model when facing a non-open source language model, which seriously limits its promotion and use in practical applications; detection methods based on language models achieve text classification tasks by transforming the output layer of open source large language models. This method belongs to black box detection, which has weak generalization ability and is prone to overfitting. Once the text to be detected exceeds the scope of the training data or involves language change, its detection effect will drop significantly. Especially in fields such as education, this method often misjudges human-written text as machine-generated text, which seriously affects the accuracy of detection; detection methods based on watermark algorithms detect text by manipulating model generation behavior or embedding tags, but this process will have a negative impact on the quality of generated text. Moreover, this method has poor robustness when attacked. For example, when the watermark sampling algorithm is modified or the watermark information is cracked, the detection mechanism will fail. In addition, this method is not suitable for commercial and closed-source models.
[0004] In summary, existing detection methods are difficult to meet actual needs, and there is an urgent need for an efficient, accurate, machine-generated natural language detection method with strong generalization ability and robustness. Summary of the invention
[0005] To solve the problems of weak generalization ability, easy overfitting, and low detection accuracy in the existing detection of machine-generated natural language, the present invention proposes a method for detecting machine-generated natural language based on multiple features.
[0006] The present invention provides a method for detecting machine-generated natural language based on multiple features, comprising the following steps:
[0007] S1. Obtain the text to be detected, and perform preprocessing on the text to be detected to obtain a preprocessed text;
[0008] S2. Input the preprocessed text obtained in step S1 into the tokenizer of the pre-trained language model for text tokenization and convert the tokens into a sequence of tokens recognizable by the pre-trained language model. Input the sequence of tokens into the pre-trained language model for calculation to obtain the unnormalized prediction probability of each token;
[0009] S3. Calculate 11 features of the text according to the sequence of tokens and the unnormalized prediction probability of each token obtained in step S2. The 11 features include: log-likelihood, rank, log-rank, perplexity, ratio of rank to log-likelihood, variance of surprisal, mean of the difference in surprisal between adjacent tokens, information entropy, lexical density, variance of log-rank, and the result of supervised detection;
[0010] S4. Concatenate the 11 features obtained in step S3 into an 11-dimensional vector, and input the 11-dimensional vector into a machine learning classification algorithm for supervised learning to obtain a detector for machine-generated natural language;
[0011] S5. Input the text to be detected into the detector for machine-generated natural language obtained in step S4 to obtain a detection result.
[0012] According to a method for detecting machine-generated natural language based on multiple features according to some embodiments of the present application, the preprocessing in step S1 includes removing punctuation, spaces, URLs, special characters, and emojis from the text to be detected, converting the vocabulary of the text to be detected into lowercase vocabulary, and restoring the part of speech.
[0013] According to a method for detecting machine-generated natural language based on multiple features according to some embodiments of the present application, the pre-trained language model in step S2 is GPT2 XL, and the tokenizer parameters max_length are set to 512 and truncation is set to True.
[0014] A machine-generated natural language detection method based on multiple features according to some embodiments of the present application. In step S3, calculating the log-likelihood of the text includes: calculating the normalized value of the unnormalized prediction probability of each token in the text, that is, performing a Softmax function operation on the value of the unnormalized prediction probability to obtain the token probability, and taking the logarithm of the token probability. The negative mean of the log-likelihood of the text tokens is obtained through the mean function, that is, the log-likelihood of text S, as shown in formula (1):
[0015] LL(S) = -E(log(p(x))) (1)
[0016] Where LL(S) represents the log-likelihood of text S, S represents the text, x represents the token in text S, p(·) represents the probability function, and E(·) represents the mean function;
[0017] Calculating the rank number of the text includes: calculating the normalized value of the unnormalized prediction probability of each token, that is, performing a Softmax function operation on the value of the unnormalized prediction probability to obtain the token probability, and obtaining the rank number of the token probability. The mean value of the rank numbers of the text tokens is obtained through the mean function, abbreviated as the rank number, as shown in formula (2):
[0018] R(S) = E(rank(p(x))) (2)
[0019] Where R(S) represents the rank number of text S, and rank represents the function of obtaining the probability ranking of token x among the predicted candidate words in the pre-trained language model;
[0020] Calculating the log-rank number of the text can directly take the logarithm value of R(S), as shown in formula (3):
[0021] LR(S) = log(R(S)) (3)
[0022] Where LR(S) represents the log-rank number of text S;
[0023] Calculating the perplexity of the text, as shown in formula (4):
[0024]
[0025] Where PPL(S) represents the perplexity of text S, n represents the number of tokens in the text, i ∈ (1, 2, 3,..., n), and exp(·) represents the exponential function;
[0026] Calculating the ratio of the rank number to the log-likelihood of the text, as shown in formula (5):
[0027]
[0028] Among them, Ratio(S) represents the ratio of the rank number of text S to the log-likelihood;
[0029] Calculating the surprisal variance of the text includes: first calculating the surprisal value of the unnormalized prediction probability of each token, so as to obtain the surprisal variance of the text, as shown in formula (6):
[0030]
[0031] Among them, UID(S) represents the surprisal variance of text S, and s(x i ) represents the surprisal of the i-th token, and x i represents the i-th token, s(x i ) = -log(p(x i |x <i ))), and s(S) represents the surprisal of text S;
[0032] Calculating the mean of the surprisal differences between adjacent tokens of the text, as shown in formula (7):
[0033]
[0034] Among them, UID'(S) represents the mean of the surprisal differences between adjacent tokens of text S, and x i-1 represents the (i - 1)-th token;
[0035] Calculating the information entropy of the text, as shown in formula (8):
[0036]
[0037] Among them, H(S) represents the information entropy of text S;
[0038] Calculating the lexical density of the text, as shown in formula (9):
[0039]
[0040] Among them, LD(S) represents the lexical density of text S, and V represents the number of vocabulary used in the text;
[0041] Calculating the variance of the log-rank of the text, as shown in formula (10):
[0042]
[0043] Among them, VR(S) represents the variance of the log-rank of text S;
[0044] Calculating the result of the supervised detection of the text for the supervised detection model, as shown in formula (11):
[0045] SD(S) = F(S) (11)
[0046] Among them, SD(S) represents the result of the supervised detection of text S by the supervised detection model, and F represents the supervised detection model.
[0047] According to a multi-feature based machine-generated natural language detection method of some embodiments of the present application, the supervised detection model is AI generated text detection and roberta large openai detector, and the parameters of the supervised detection model tokenizer, max_length, are set to 512, and truncation is set to True.
[0048] According to a multi-feature based machine-generated natural language detection method of some embodiments of the present application, the machine learning classification algorithm in step S4 is the extreme gradient boosting tree machine learning algorithm.
[0049] According to a multi-feature based machine-generated natural language detection method of some embodiments of the present application, in the supervised learning in step S4, the ratio of the amount of data in the training set to the test set is 9:1, the learning rate is 0.01, the maximum depth of the weak classifier is set to 29, and the number of weak classifiers is set to 1800.
[0050] According to a multi-feature based machine-generated natural language detection method of some embodiments of the present application, the loss function of the supervised learning in step S4 is the binary logistic regression loss, as shown in formula (12):
[0051]
[0052] Among them, represents the binary logistic regression loss, y k represents the true probability of the k-th sample, k ∈ (1, 2, 3,..., N), represents the predicted probability of the k-th sample, and N represents the number of samples.
[0053] According to a multi-feature based machine-generated natural language detection method of some embodiments of the present application, the detection results in step S5 include that the text to be detected is machine-generated text and the text to be detected is human-written text.
[0054] A method for detecting machine-generated natural language based on multiple features proposed by the present invention uses eleven features during the detection process, making the prediction results have good interpretability, enhancing the credibility and transparency of the detection method. Moreover, this method uses a pre-trained language model with a relatively small number of parameters, having significant resource-friendly characteristics. Even in an environment with relatively scarce computing resources, it can still efficiently implement the detection function, greatly expanding its application scope and scenarios. The present invention can comprehensively consider the source problem of the text to be detected, thereby more accurately distinguishing machine-generated text from human-written text, significantly improving the accuracy of text source detection. Compared with existing baseline methods, the present invention can achieve more excellent performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic flowchart of a method for detecting machine-generated natural language based on multiple features of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following further describes in detail the embodiments of the present invention with reference to the drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0057] Embodiment 1. This embodiment provides a method for detecting machine-generated natural language based on multiple features, as Figure 1 shown, including the following steps:
[0058] S1. Obtain the text to be detected, and perform preprocessing on the text to be detected to obtain the preprocessed text;
[0059] S2. Input the preprocessed text obtained in step S1 into the tokenizer of the pre-trained language model for text tokenization and convert the tokens into a sequence of tokens recognizable by the pre-trained language model. Input the sequence of tokens into the pre-trained language model for calculation to obtain the unnormalized prediction probability of each token;
[0060] S3. Calculate 11 features of the text according to the sequence of tokens and the unnormalized prediction probability of each token obtained in step S2. The 11 features include: log-likelihood, rank, log-rank, perplexity, ratio of rank to log-likelihood, variance of surprisal, mean of the difference in surprisal between adjacent tokens, information entropy, lexical density, variance of log-rank, and the result of supervised detection;
[0061] S4. Concatenate the 11 features obtained in step S3 into an 11-dimensional vector, and input the 11-dimensional vector into a machine learning classification algorithm for supervised learning to obtain a machine-generated natural language detector;
[0062] S5. Input the text to be detected into the machine-generated natural language detector obtained in step S4 to obtain the detection result.
[0063] Embodiment 2. This embodiment provides a method for detecting machine-generated natural language based on multiple features, including the following steps:
[0064] S1. Obtain the text to be detected, and preprocess the text to be detected to obtain the preprocessed text;
[0065] Preferably, specifically, the preprocessing includes removing punctuation, spaces, URLs, special characters, and emojis from the text to be detected, converting the vocabulary of the text to be detected into lowercase vocabulary and restoring the part of speech, so as to improve the data quality, reduce noise, and provide a good data basis for subsequent steps;
[0066] S2. Input the preprocessed text obtained in step S1 into the tokenizer of the pre-trained language model for text tokenization and convert the tokens into a sequence of tokens recognizable by the pre-trained language model, input the sequence of tokens into the pre-trained language model for calculation, and obtain the unnormalized prediction probability of each token;
[0067] Preferably, specifically, the pre-trained language model is GPT2 XL, and the tokenizer is the corresponding tokenizer of GPT2 XL to ensure the adaptation of text processing and the model, and improve the accuracy and effectiveness of feature extraction.
[0068] The parameter max_length of the tokenizer is set to 512, and truncation is set to True. That is, the length of the imported word vector is at most 1024, and if it exceeds, the vector is truncated to ensure that the data length is within the range acceptable by the model, and at the same time, the data dimension sizes within the same batch are consistent. At the same time, check whether the length of the word vector is less than 50. If the length is less than 50, then remove the text. Because it is more difficult to detect shorter texts and the detection significance is not great.
[0069] S3. Calculate 11 features of the text based on the token sequence obtained in step S2 and the unnormalized prediction probability of each token. The 11 features include: log-likelihood, rank, log-rank, perplexity, ratio of rank to log-likelihood, variance of surprisal, mean of difference in surprisal between adjacent tokens, information entropy, lexical density, variance of log-rank, and the result of supervised detection. Among them, log-likelihood, rank, log-rank, perplexity, and ratio of rank to log-likelihood belong to the features of the decoding strategy. Variance of surprisal and mean of difference in surprisal between adjacent tokens belong to the features based on the principle of uniform information density. Information entropy belongs to the features based on information entropy. Lexical density and variance of log-rank belong to the features based on the study of the dataset. In natural language generation tasks, in order to generate a text sequence, a pre-trained language model usually needs to let the model predict each token one by one until a termination condition is reached. During the generation process, the pre-trained language model will generate the probability scores of all the words in the vocabulary according to the current state. The way in which the pre-trained language model selects the output token according to the probability is the decoding strategy of the language model. And statistical features that can be used for detection can be discovered through the study of the decoding strategy of the pre-trained language model. Usually, the pre-trained language model will select the words with higher probability scores for output. Therefore, if the probability scores represented by these feature values in the text to be detected are larger, the text is more likely to be machine-generated; on the contrary, it may be written by humans.
[0070] As a preference of this embodiment, specifically, calculating the log-likelihood of the text includes: calculating the normalized value of the unnormalized prediction probability of each token in the text, that is, performing the Softmax function operation on the value of the unnormalized prediction probability to obtain the token probability, taking the logarithm of the token probability, and obtaining the negative mean of the log-likelihood of the text tokens through the mean function, that is, the log-likelihood of text S, as shown in formula (1):
[0071] LL(S) = -E(log(p(x))) (1)
[0072] Where LL(S) represents the log-likelihood of text S, S represents the text, x represents the token in text S, p(·) represents the probability function, and E(·) represents the mean function;
[0073] Calculating the negative mean of the log-probability of tokens is to magnify the difference between the probabilities of different tokens. Therefore, the logarithm is introduced. The smaller probability value is higher, the larger probability value is smaller, and the final result is taken as a positive number. That is, the linear numerical change is converted into an exponential change to magnify the numerical difference. If the log-likelihood value is higher, the text may be machine-generated; if the log-likelihood value is lower, it may be a text written by humans;
[0074] Calculating the rank of a text includes: calculating the normalized value of the unnormalized prediction probability of each token, that is, performing the Softmax function operation on the value of the unnormalized prediction probability to obtain the token probability, and obtaining the rank of the token probability. The average rank of the text tokens is obtained through the mean function, which is simply referred to as the rank, as shown in formula (2), that is, it is used to calculate the average ranking of the token in the probability distribution in the vocabulary:
[0075] R(S) = E(rank(p(x))) (2)
[0076] Among them, R(S) represents the rank of text S, rank represents the function of obtaining the ranking of token x in terms of probability in the candidate words predicted by the pre-trained language model, calculating the average ranking of the token in the probability distribution in the vocabulary, which is the same as the idea of log-likelihood. However, when the probability is higher, the corresponding ranking of the token in the vocabulary is higher, that is, the ranking is smaller. If the rank value is smaller, then this text may be from a machine-generated text. If the rank value is larger, then it may be a human-written text;
[0077] Calculating the log-rank of a text can directly take the logarithm of R(S), as shown in formula (3), that is, it is used to calculate the logarithm of the rank:
[0078] LR(S) = log(R(S)) (3)
[0079] Among them, LR(S) represents the log-rank of text S. In the operation, the logarithm of R(S) can be directly taken. Since the size of the language model vocabulary is large, the logarithmic form of the rank value can be used, which has a better detection effect. If the log-rank value is smaller, then this text may be from a machine-generated text, while if the log-rank value is larger, then it may be a human-written text;
[0080] For the probability of a text, as shown in formula (4):
[0081]
[0082] Among them, w i represents the word in the text, n represents the length of the text, i ∈ (1, 2, 3,..., n),
[0083] Calculating the perplexity of a text, as shown in formula (5):
[0084]
[0085] Among them, PPL(S) represents the perplexity of text S, and exp(·) represents the exponential function. During the training of a language model, perplexity is usually an indicator used to evaluate the performance of the language model. According to the above formula for calculating perplexity, perplexity is based on the model's estimation of the probability distribution of data and reflects the degree of adaptation of the model to the real data. Since a pre-trained language model is being detected, given its superior performance, the smaller the perplexity eigenvalue of the text to be detected, the more likely it is to be from machine-generated text. On the contrary, the larger the perplexity eigenvalue, the more likely it is to be from human-written text;
[0086] Calculate the ratio of the rank number to the log-likelihood of the text, as shown in formula (6), for highlighting the difference to assist in classification:
[0087]
[0088] Among them, Ratio(S) represents the ratio of the rank number to the log-likelihood of text S; the ratio of the rank number to the log-likelihood combines the log-likelihood and the rank number. From the descriptions of the above two, for the same output token of the model, when the probability is higher, the rank value is lower; when the probability is lower, the rank value is higher. Therefore, by calculating the ratio of the log-rank number to the log-likelihood, this difference can be amplified, enabling the machine-generated natural language detector to better distinguish the differences between humans and machines;
[0089] The Uniform Information Density principle, abbreviated as the UID principle, holds that in order for humans to communicate efficiently, the information in language should be distributed as evenly as possible. At the same time, according to the content of information theory, language can be regarded as a communication system, and each language unit carries different amounts of information. The surprisal in information theory can be used to quantify the amount of information. For the i-th language unit u i in terms of which, surprisal(s(·)) is the negative logarithm of the conditional probability of u i for the model, as shown in formula (7):
[0090]
[0091] Such a definition means that when the probability of a word appearing is smaller, its information content is larger. According to the characteristic of the uniform distribution of information in human-written text in the UID principle, the degree of dispersion of the information content of tokens can be calculated, and then its distribution characteristics can be obtained, which are used as the detection criteria;
[0092] Calculating the variance of the surprisal of the text includes: calculating the surprisal value of the unnormalized predicted probability of each token, so as to obtain the variance of the surprisal of the text, as shown in formula (8):
[0093]
[0094] Among them, UID(S) represents the variance of surprise of text S, s(x i ) represents the surprise of the i-th word, x i represents the i-th word, s(x i )=-log(p(x i |x <i )), s(S) represents the surprise of text S, and checks the degree of dispersion of the surprise of the text to the model. According to the UID principle, the information content of human-written text is more evenly distributed. If the variance of the surprise is small, that is, the degree of dispersion is low, the information content is more evenly distributed, and the text is more likely to be from human-written text; on the contrary, if the variance of the surprise is large, that is, the degree of dispersion is high, the information content is more chaotic, and the text is more likely to be from machine-generated text;
[0095] Calculate the mean of the difference in surprise between adjacent words in the text, as shown in formula (9):
[0096]
[0097] Among them, UID'(S) represents the mean difference of the surprisingness of adjacent words in text S, x i-1 Represents the i-1th word. Another explanation of the UID principle is that UID is a resistance that prevents too fast transfer from the information-dense part to the information-sparse part. The signal should have a smooth transition between the information-sparse component and the dense component. Check the mean of the difference in the surprise of the model for adjacent words. According to the UID principle, the amount of information in human-written text changes more smoothly. If the mean of the difference in the surprise of adjacent words is small, the amount of information in the text changes more smoothly, and the text is more likely to come from human-written text; conversely, if the mean of the difference in the surprise of adjacent words is large, the amount of information changes dramatically, and the text is more likely to come from machine-generated text.
[0098] Calculate the information entropy of the text, as shown in formula (10):
[0099]
[0100] Among them, H(S) represents the information entropy of text S; Information Entropy is a basic concept in information theory. It describes the uncertainty of various possible events occurring in the information source, and the average amount of information after removing redundancy in the information is called "information entropy". Information entropy measures the uncertainty or unpredictability contained in the information source. The greater the information entropy, the greater the uncertainty in the information source and the more information there is; conversely, the smaller the information entropy, the smaller the uncertainty in the information source and the less information there is. In a language model, information entropy represents the amount of information in the text and the uncertainty of the model about the text. The calculation method is the same as that of information entropy in information theory, where the probability uses the predicted probability of the model for the token. Since information entropy describes the uncertainty of the text, and the probability corresponding to the text generated by the model should be relatively high, then the uncertainty is relatively low, and the information entropy is relatively small. Another explanation is that for the text generated by the model, it is the data within the "internal distribution" of the model, so its probability is relatively high, the uncertainty is relatively low, and the information entropy is relatively small. Calculate the mean value of the information entropy of the tokens on the text. If the information entropy eigenvalue is relatively small, then this text may be from machine-generated text; if the information entropy eigenvalue is relatively large, then it may be human-written text;
[0101] In an artificial study of datasets of machine-generated text and human-written text, it is found that the vocabulary used in machine-generated text is less than that in human-written text, that is, humans use a more diverse vocabulary. Therefore, calculating the vocabulary used in the text can also be used as a measure for detection. Calculate the vocabulary density of the text, as shown in formula (11):
[0102]
[0103] Among them, LD(S) represents the vocabulary density of text S, V represents the number of vocabulary used in the text, and calculate the vocabulary density of the text. The text with a large vocabulary density may be human-written text, and conversely, it may be machine-generated text;
[0104] Calculate the variance of the log-rank of the text, as shown in formula (12):
[0105]
[0106] Among them, VR(S) represents the variance of the log-rank of text S, which combines the language model decoding strategy and the measurement method of vocabulary research. Therefore, calculate the degree of dispersion of the rank values of the tokens. Since the choices during machine generation are more concentrated, the variance value of the log-rank is relatively low; while the choices during human writing are more dispersed, so the variance value of the log-rank is relatively high;
[0107] Calculate the result of the supervised detection of the text for the supervised detection model, as shown in formula (13):
[0108] SD(S) = F(S) (13)
[0109] Among them, SD(S) represents the result of the supervised detection of text S by the supervised detection model, and F represents the supervised detection model. In the detection of machine-generated natural language text, in addition to feature-based detection, there are also methods based on pre-trained language models. Usually, it is based on a pre-trained language model, the network structure is modified, the network parameters except for the output layer are frozen, and the output layer is changed to a classification layer. And it is supervised and trained with a large amount of data to enable it to perform classification and detection. After the pre-trained language model for machine-generated natural language detection inputs the text to be detected, the result obtained is usually the probability that the text is machine-generated text or human-written text;
[0110] The supervised detection models are AI generated text detection and roberta large openaidetector, and the parameters of the supervised detection model tokenizer max_length are set to 512 and truncation is set to True;
[0111] S4. Combine the 11 features obtained in step S3 into an 11-dimensional vector, and input the 11-dimensional vector into a machine learning classification algorithm for supervised learning to obtain a machine-generated natural language detector;
[0112] As a preference of this embodiment, specifically, the machine learning classification algorithm is the extreme gradient boosting tree machine learning algorithm. In the supervised learning, the ratio of the training set data volume to the test set data volume is 9:1, the learning rate is 0.01, the maximum depth of the weak classifier is set to 29, and the number of weak classifiers is set to 1800; at the same time, it is stipulated that the label of machine-generated text is 1; the label of human-written text is 0;
[0113] The loss function of the supervised learning is the binary logistic regression loss, as shown in formula (14):
[0114]
[0115] Among them, represents the binary logistic regression loss, y k represents the true probability of the k-th sample, k ∈ (1, 2, 3,..., N), represents the predicted probability of the k-th sample, and N represents the number of samples.
[0116] More preferably, before performing the supervised learning, it is necessary to first perform a standardization operation on the 11 features. After the data is standardized, the model can be created. In the parameter tuning stage, grid search can be used for parameter screening. After multiple iterative searches and tests, the model with the best performance index is selected as the classification model of the detection algorithm.
[0117] S5. Input the text to be detected into the machine-generated natural language detector obtained in step S4 to obtain a detection result;
[0118] As an optimization of this embodiment, specifically, the detection result includes that the text to be detected is machine-generated text and the text to be detected is human-written text.
[0119] To verify the detection effectiveness of the natural language detector in this embodiment, the following experiment was conducted:
[0120] To test the performance of the natural language detector in this embodiment, this embodiment collected 9 human-written text datasets and a human ChatGPT comparison corpus (HC3) dataset. Based on the human-written text datasets, corresponding machine-generated texts were generated through multiple language models (T5-3B, Davinci, GPT series), and then the 9 human-written text datasets, the generated corresponding machine-generated texts, and HC3 were combined as the dataset for this experiment.
[0121] To cover different fields and generation methods as much as possible, the 9 human-written text datasets were used as source data, and multiple large language models were used to construct machine-generated texts. Among them, the methods for constructing machine-generated texts in the following datasets are as follows: The first construction method is to select the first 30 tokens of the source text and input them into the large language model to continue generating text. The source text is human-written text, and the text generated by the language model is machine-generated text. The second is to require the large language model to generate machine-generated text according to a certain theme, such as an argument, a news headline, a story theme, an article theme, etc.
[0122] The overall situation of the training and test datasets is introduced as follows:
[0123] 1. Statements and restatements of opinions from the r / ChangeMyView (CMV) Reddit sub-community.
[0124] 2. Opinions from the Yelp review dataset.
[0125] 3. News articles from XSum and TLDR_news (TLDR).
[0126] 4. Q&A texts from the ELI5 dataset.
[0127] 5. Story texts according to the RedditWritingPrompts (WP) dataset.
[0128] 6. Story texts from the ROCStories Corpora (ROC).
[0129] 7. Commonsense reasoning texts from HellaSwag (HS).
[0130] 8. A reading comprehension dataset from SQuAD.
[0131] 9. Abstracts of scientific articles from SciGen (SGen).
[0132] 10. The Human ChatGPT Comparison Corpus (HC3) dataset.
[0133] Statistical information on human-written text data and machine-generated text data from each source is shown in Table 1.
[0134] Table 1 Dataset details
[0135]
[0136] Among them, HWT represents human-written text, with a total of 178,000 pieces; MGT represents machine-generated text, with a total of 318,000 pieces, and all data totals 497,000 pieces.
[0137] In this embodiment, multiple metrics for evaluating the performance of the classification model are adopted, including: Accuracy (Acc for short), Precision (Pre for short), Recall (Rec for short), F1-Score (F1 for short), Confusion Matrix, AUC value, False Positive Rate (FPR), and False Negative Rate (FNR). The definitions of the metrics in the classification problem are as follows:
[0138] TN (True Negative): The number of samples correctly predicted as negative by the model.
[0139] FN (False Negative): The number of actual positive-class samples wrongly predicted as negative by the model.
[0140] TP (True Positive): The number of samples correctly predicted as positive by the model.
[0141] FP (False Positive): The number of actual negative-class samples wrongly predicted as positive by the model.
[0142] The definitions of each evaluation metric are as follows:
[0143] Accuracy is as shown in formula (15):
[0144]
[0145] Precision is as shown in formula (16):
[0146]
[0147] The recall rate is shown in formula (17):
[0148]
[0149] The F1 score is shown in formula (18):
[0150]
[0151] The false detection rate is shown in formula (19):
[0152]
[0153] The missed detection rate is shown in formula (20):
[0154]
[0155] In the experiments of this embodiment, mainly 8 algorithms are used to compare with the natural language detection method of this embodiment.
[0156] 1. Log Likelihood: This method evaluates the average log probability of the text. The principle is that paragraphs with higher average log probability are more likely to be generated by a machine.
[0157] 2. Rank: This method evaluates the average rank number of each token in the text. The principle is that paragraphs with lower average rank number are more likely to be generated by a machine.
[0158] 3. Log Rank: This method does not directly use the rank number, but evaluates the average log rank of each token in the text. Paragraphs with lower average observed log rank are more likely to be generated by a machine.
[0159] 4. Entropy: This method is inspired by the assumption that text generated by a machine is more likely to have an overconfident (and thus lower entropy) prediction distribution. This method takes a higher average entropy as a signal of text generated by a machine.
[0160] 5. DetectGPT: DetectGPT is a perturbation-based detection method. Specifically, this method proves that text sampled from a large language model tends to be in the negative curvature region of the model's log probability function. Therefore, DetectGPT checks the average probability change of the text after being perturbed.
[0161] 6. Supervised Detect based on large language models: This is a classifier based on large language models. It can directly input text, and the output result is the probability that the text belongs to machine-generated text. Hereinafter referred to as supervised detection.
[0162] 7. DetectLLM-LRR: DetectLLM-LRR is a method based on a decoding strategy. This method uses the log-likelihood log-rank ratio as a feature. It utilizes the numerical characteristics of these two features to amplify the feature information, which is more conducive to detection.
[0163] 8. DetectLLM-NPR: DetectLLM-NPR is also a detection method based on perturbation. However, DetectLLM-NPR checks the average rank change of the text after being perturbed.
[0164] The experimental results of this embodiment are shown in Table 2:
[0165] Table 2 Experimental results of six comparison algorithms and the algorithm of the present invention on the entire dataset
[0166]
[0167]
[0168] In this experiment, first, 400,000 pieces of data in the entire dataset were used for training and testing, and the experimental results of the multi-feature-based machine-generated natural language detection method (hereinafter referred to as MFE-Detector) of this embodiment and four baseline methods were obtained respectively. It can be seen from Table 2 that the accuracy rate of the MFE-Detector algorithm is 93.95%, the F1 score is 0.9383, and the AUC value is 0.9396. Compared with the comparison algorithms, the accuracy rate of the MFE-Detector algorithm has increased by about 60% on average. It is experimentally obtained that for human texts, the probability of misjudging them as machine-generated is 3.790%, that is, the false detection rate. For machine-generated texts, the probability of misjudging them as human-written is 8.295%, that is, the missed detection rate. It can be seen that the false detection rate is slightly higher, that is, the MFE-Detector algorithm is more likely to classify human-written texts as machine-generated texts.
[0169] Compared with single-feature detection methods such as the entropy method, log-likelihood, and rank method, the MFE-Detector algorithm introduces multiple features, including features related to the probability concept based on the decoding strategy, features based on information entropy, features based on the difference between human and machine vocabulary, features based on the principle of uniform information density, and the results of the supervised detector. It covers a more comprehensive range of aspects, uses multiple features to comprehensively consider the source of the text to be detected, determines the text source from multiple perspectives, and makes the results more accurate. This reflects the necessity of introducing multi-feature classification. Therefore, the detection accuracy has been improved.
[0170] Then, in this embodiment, a comparison is also made with the perturbation-based comparison algorithm on the HC3 dataset, and the results are shown in Table 3:
[0171] Table 3 Experimental results of four comparison algorithms and the algorithm of the present invention on the HC3 dataset
[0172]
[0173] Judging from the experimental results, DetectLLM-LRR is better than other baseline methods, while the AUC value of DetectLLM-NPR is close to 0.5, indicating that the task has completely failed. Therefore, the extremely high recall rate observed in DetectLLM-NPR is abnormal. In contrast, the method MFE-Detector of this embodiment achieves the best detection performance on various metrics of the HC3 dataset. The excellent performance of the method MFE-Detector of this embodiment on the HC3 dataset may be because the HC3 dataset is constructed using ChatGPT, and MFE-Detector uses GPT2 XL as the feature extraction model. These two homologous models can extract features more effectively, thus improving the performance of the HC3 dataset. This finding shows that selecting a series of models as feature extractors can better extract text features from different sources.
[0174] This embodiment also conducts an experiment on the influence of the choice of language model on the detection accuracy. This embodiment studies the influence of different pre-trained language models on the detection performance while keeping other conditions unchanged. The purpose is to evaluate whether the choice of model will affect the detection results. The tested models include RoBERTa-large, GPT2 Medium, Mamba, and GPT2 XL, and the experimental results are shown in Table 4:
[0175] Table 4 Influence of different language models as embedding calculation models on performance
[0176]
[0177] As shown in Table 4, for the GPT model, as the model size (number of parameters) decreases, the performance of the MFE-Detector algorithm of the method in this embodiment slightly decreases in all metrics. This indicates that as the model capacity decreases, the detection performance also decreases. In contrast, although RoBERTa requires less computing time, its accuracy is lower. This indicates that RoBERTa may be more suitable for resource-constrained environments. Mamba is a new model architecture with potential across different fields, and its detection performance is not that strong, perhaps due to lower pre-training strength or sub-optimal model structure. In summary, the MFE-Detector algorithm uses GPT2 XL as the computational model for text embedding, which can not only meet the requirements of model detection performance but also has low requirements for computing resources, meeting the requirements of resource conservation.
[0178] In the design of this embodiment, it is crucial to select an effective classification model. The present invention evaluated several classification algorithms in the experiment, including the K-Nearest Neighbor algorithm (KNN), Support Vector Machine (SVM), Fully Connected Neural Network (FCNN), and Extreme Gradient Boosting Tree algorithm (XGBoost). The experimental results are shown in Table 5. The XGBoost algorithm demonstrated better performance, so XGBoost was selected as the classification model in the design of the present invention.
[0179] Table 5 Influence of Different Algorithms as Classification Models on Performance
[0180]
[0181] Tables 6 and 7 are the implementation cases of this embodiment
[0182] Table 6 Implementation Case 1 of this embodiment
[0183]
[0184] Table 7 Implementation Case 2 of this embodiment
[0185]
[0186] Tables 6 and 7 show the predicted labels of the present invention and the comparative methods. As can be seen from the table, the multi-feature-based machine-generated natural language detection method of this embodiment can correctly predict the source of the text, while other comparative algorithms may not necessarily correctly detect the source.
[0187] The method of the present invention is a multi-feature-based method for detecting machine-generated natural language. It is based on the statistical analysis of a series of feature data distributions, including log-likelihood, rank, log-rank, perplexity, the ratio of rank to log-likelihood, variance of surprisal, mean of the difference in surprisal between adjacent tokens, information entropy, lexical density, variance of log-rank, and the output result of supervised detection. In the specific implementation process, machine-generated text usually exhibits the following characteristics: a smaller log-likelihood value, a smaller rank, a smaller log-rank, a smaller ratio of rank to log-likelihood, a smaller perplexity, a larger variance of surprisal, a larger mean of the difference in surprisal between adjacent tokens, a smaller information entropy, a smaller variance of log-rank, a smaller lexical density, and a larger output result of the supervised detector. On the contrary, human-written text presents opposite characteristics. By performing multi-dimensional analysis on these features, the present invention can comprehensively consider the source problem of the text to be detected, thereby more accurately distinguishing machine-generated text from human-written text, significantly improving the accuracy of text source detection. Compared with the existing baseline methods, the present invention can achieve more superior performance.
[0188] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A multi-feature based machine-generated natural language detection method, characterized in that: The steps include: S1. Obtaining a text to be detected, and preprocessing the text to be detected to obtain a preprocessed text; S2. Input the preprocessed text obtained in step S1 into the word segmenter of the pretrained language model for text segmentation and convert the segmented words into a word unit sequence recognizable by the pretrained language model, input the word unit sequence into the pretrained language model for operation, and obtain the unnormalized prediction probability of each word unit; S3. Calculate 11 features of the text according to the word-unit sequence obtained in step S2 and the unnormalized predicted probability of each word-unit, the 11 features including: log likelihood, rank, log-rank, perplexity, ratio of rank to log likelihood, variance of surprise, mean of difference in surprise between adjacent word-units, information entropy, vocabulary density, variance of log-rank and result of supervised detection; S4. Combining the 11 features obtained in step S3 into an 11-dimensional vector, and inputting the 11-dimensional vector into a machine learning classification algorithm for supervised learning to obtain a machine-generated natural language detector; S5. Input the text to be detected into the machine-generated natural language detector obtained in step S4 to obtain a detection result.
2. The method for detecting machine-generated natural language based on multiple features according to claim 1, characterized in that: The preprocessing in step S1 includes removing punctuation marks, spaces, URLs, special characters and emoticons in the text to be detected, converting the words in the text to be detected into lowercase words and restoring the parts of speech.
3. The method for detecting a multi-feature machine-generated natural language according to claim 1, characterized in that: The pre-trained language model in step S2 is GPT2 XL, the word segmenter parameter max_length is set to 512, and truncation is set to True.
4. The method for detecting a machine-generated natural language based on multiple features according to claim 1, characterized in that: In step S3, the log-likelihood of the text is calculated as shown in formula (1): LL(S)=-E(log(p(x))) (1) Where LL(S) represents the log-likelihood of the text S, S represents the text, x represents the word in the text S, p(·) represents the probability function, and E(·) represents the mean function; Calculate the rank of the text, as shown in formula (2): R(S)=E(rank(p(x))) (2) Where R(S) represents the rank of the text S, and rank represents the function of finding the probability ranking of the word x in the candidate words predicted by the pre-trained language model; Calculate the logarithmic rank of the text, as shown in formula (3): LR(S)=log(R(S)) (3) Among them, LR(S) represents the logarithmic rank of text S; Calculate the perplexity of the text, as shown in formula (4): Where PPL(S) represents the perplexity of the text S, n represents the number of tokens in the text, i∈(1,2,3,…,n), and exp(·) represents the exponential function; The ratio of the rank number to the log likelihood of the text is calculated as shown in formula (5): Among them, Ratio(S) represents the ratio of the rank of text S to the log likelihood; The surprising variance of the text is calculated as shown in formula (6): Among them, UID(S) represents the variance of surprise of text S, s(x i ) represents the surprise of the i-th word, x i represents the i-th word, s(x i )=-log(p(x i |x <i )), s(S) represents the surprise of text S; The mean of the difference in the surprisingness of the adjacent words in the text is calculated, as shown in formula (7): Among them, UID'(S) represents the mean difference of the surprisingness of adjacent words in text S, x i-1 represents the i-1th word; The information entropy of the text is calculated as shown in formula (8): Among them, H(S) represents the information entropy of text S; The vocabulary density of the text is calculated as shown in formula (9): Among them, LD(S) represents the lexical density of text S, and V represents the number of words used in the text; The variance of the logarithmic rank of the text is calculated as shown in formula (10): Among them, VR(S) represents the variance of the log rank of text S; The result of the supervised detection of the text for the supervised detection model is calculated as shown in formula (11): SD(S)=F(S) (11) Among them, SD(S) represents the result of supervised detection of text S for the supervised detection model, and F represents the supervised detection model.
5. The method for detecting machine-generated natural language based on multiple features according to claim 4, characterized in that: The supervised detection models are AI generated text detection and roberta large openai detector, and the supervised detection model word segmenter parameter max_length is set to 512 and truncation is set to True.
6. The method for detecting machine-generated natural language based on multiple features according to claim 1, characterized in that: The machine learning classification algorithm described in step S4 is an extreme gradient boosting tree machine learning algorithm.
7. The method for detecting machine-generated natural language based on multiple features according to claim 1, characterized in that: In the supervised learning described in step S4, the ratio of the data volume of the training set to that of the test set is 9:1, the learning rate is 0.01, the maximum depth of the weak classifier is set to 29, and the number of weak classifiers is set to 1800.
8. The method for detecting machine-generated natural language based on multiple features according to claim 1, characterized in that: The loss function of the supervised learning in step S4 is a binary logistic regression loss, as shown in formula (12): in, represents the binary logistic regression loss, y k represents the true probability of the kth sample, k∈(1,2,3,…,N), represents the predicted probability of the kth sample, and N represents the number of samples.
9. The method for detecting machine-generated natural language based on multiple features according to claim 1, characterized in that: The detection result in step S5 includes that the text to be detected is a machine-generated text and the text to be detected is a human-written text.
Citation Information
Patent Citations
Method and system for detecting machine to generate Chinese text, terminal and medium
CN116468022A
Generated text detection method based on statistical feature fusion of multiple large language models
CN117291175A
Method and device for training text generation model based on reinforcement learning
CN118551824A
Large language model generation text detection method based on semantic decoupling
CN119167114A
Deepfake detection
US20240355322A1
Cited By
AI generation text detection method based on sentence length distribution and text predictability characteristics
CN122433705A
Ai-generated text detection method based on sentence length distribution and text predictability features
CN122433705B