Paper detection method and device, storage medium, terminal
By segmenting and word segmentation of papers, using the self-coded binary classification model and the token table of multiple natural language generation models to generate training corpus, the problem of low accuracy in AI-generated paper detection in the existing technology is solved, and more efficient paper detection results are achieved.
Patent Information
- Application Number
- CN202311035650.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-08-16
AI Technical Summary
The detection accuracy of papers generated based on AI in the prior art is low, mainly because the relying on training samples is too single and the model generalization ability is insufficient, so it is impossible to effectively identify papers generated by multiple generative models.
By segmenting and word segmentation of sample papers, using self-coded binary classification model to predict paper fragments, combining the token tables of multiple natural language generation models to generate token sets, perform token replacement, build training corpus, improve the diversity and comprehensiveness of training samples, and use classification statistical parameters to determine whether the paper is generated by a natural language generation model.
It improves the accuracy of paper detection, enhances the generalization ability of the model, and can more effectively identify papers generated by different generative models, improving the comprehensiveness and accuracy of the detection.
Smart Images

Figure CN117033555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text processing technology, and in particular to a paper detection method and device, a storage medium, and a terminal. Background Art
[0002] With the rapid development of generative pre-trained language model technology, the scope of content generated by artificial intelligence (AI) is becoming increasingly broad. It can generate intelligent question and answer, write program code, and even write papers based on pre-trained language models. However, in most application scenarios, AI-based paper writing is not allowed, such as student graduation theses. Therefore, in addition to checking for plagiarism, in order to ensure the authenticity of the paper, it is also necessary to test the writing method of the paper to confirm whether the paper is AI-generated.
[0003] The existing detection of AI-generated papers mainly uses the paper content generated by the AI model as training samples to train the model, and then identifies whether the paper is generated based on AI based on the trained model. However, this method relies too much on paper training samples and can only learn the generation logic of the model that generates the training samples. The model's generalization ability is weak, and there are many types of models that can be used to generate papers, resulting in low accuracy in detecting whether a paper is generated based on an AI model. Summary of the Invention
[0004] In view of this, the present invention provides a paper detection method and device, a storage medium, and a terminal, the main purpose of which is to solve the problem of low detection accuracy of papers generated based on AI models.
[0005] According to one aspect of the present invention, a method for detecting papers is provided, comprising:
[0006] Obtaining a paper to be tested, and segmenting the paper to be tested to obtain multiple paper segments;
[0007] Using the trained classification model to perform prediction processing on the paper fragments, obtaining classification results for each of the paper fragments, the classification results including the classification probability of each token;
[0008] Calculating a classification statistical parameter based on the classification probability of the token, and determining a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model;
[0009] Among them, the construction process of the training corpus of the classification model that has completed training includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample paper based on the token set to obtain training corpus generation.
[0010] Furthermore, the replacing of tokens in the sample paper based on the token set to generate training corpus includes:
[0011] Obtaining a sample paper, and segmenting the sample paper according to a preset character length to obtain a plurality of sample paper segments;
[0012] Segment the sample paper fragments, and match each original token obtained by segmentation with the tokens in the global token set to obtain the number of token hits for each sample paper fragment;
[0013] Determining the number of token replacements for each of the sample paper fragments based on the number of token hits;
[0014] For each of the sample paper fragments, replacing the tokens in the sample paper fragments according to the token replacement strategy and the corresponding token replacement quantity to obtain a training corpus;
[0015] The token replacement strategy includes at least one of a first replacement strategy for performing replacement based on the token cluster list set, a second replacement strategy for performing replacement based on the global token set, and a third replacement strategy for performing no replacement.
[0016] Furthermore, each replacement strategy in the token replacement strategy is respectively configured with a corresponding execution probability. For each of the sample paper fragments, the tokens in the sample paper fragment are replaced according to the token replacement strategy and the corresponding token replacement quantity, and the training corpus obtained includes:
[0017] For each of the sample paper fragments, randomly determine a position to be replaced that satisfies the number of token replacements;
[0018] For each of the positions to be replaced, if the token of the position to be replaced exists in the global token set, a target replacement strategy is determined from the replacement strategies by a replacement strategy probability generator;
[0019] Replace the position to be replaced according to the target replacement strategy to obtain a replaced sample paper fragment;
[0020] Based on the comparison results of the replaced sample paper fragment and the sample paper fragment at each position, the token corresponding to each position is marked to obtain the training corpus.
[0021] Furthermore, generating a token set based on a token table of multiple natural language generation models includes:
[0022] Obtaining token tables of multiple natural language generation models and token vector embedding representations corresponding to tokens in the token tables;
[0023] Merge and remove duplicates from the token tables to obtain a global token set;
[0024] The tokens in each of the token tables are clustered based on the token vector embedding representation, and a token clustering list set is generated based on each clustering result.
[0025] Furthermore, clustering the tokens in each of the token tables based on the token vector embedding representation, and generating a token cluster list set based on each clustering result includes:
[0026] Calculating the number of cluster categories of tokens corresponding to each of the natural language generation models based on the number of tokens in each of the token tables;
[0027] Clustering the tokens in each of the token tables based on the corresponding token vector embedding representation and the number of cluster categories to obtain a clustering result for each of the natural language generation models, the clustering result including a plurality of token clusters;
[0028] Token clusters corresponding to the same token in each of the clustering results are merged into a token cluster list to obtain a token cluster list set, where there is a token intersection between the token cluster lists.
[0029] Furthermore, before performing prediction processing on the paper fragments using the trained classification model to obtain classification results for each of the paper fragments, the method further includes:
[0030] Obtain training corpus and token set, and merge the token set with the original token table of the Bert model to obtain the expanded Berttoken table;
[0031] An initial classification model is constructed based on the Bert model and the expanded Berttoken table;
[0032] The initial classification model is trained using the training corpus to obtain a classification model that has completed training.
[0033] Furthermore, the classification statistical parameters include a global token classification probability mean, a first-class token classification probability mean, and a second-class token quantity ratio; the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio;
[0034] The step of determining the detection result of the paper to be detected based on the comparison result of the classification statistical parameter and the preset parameter threshold comprises:
[0035] If any one of the global token classification probability mean, the first-category token classification probability mean, and the second-category token quantity ratio is greater than the corresponding parameter threshold, it is determined that the paper to be detected is a paper generated based on a natural language generation model.
[0036] According to another aspect of the present invention, a paper detection device is provided, comprising:
[0037] An acquisition module is used to acquire the paper to be detected and divide the paper to be detected into segments to obtain multiple paper segments;
[0038] A prediction processing module is used to use the trained classification model to perform prediction processing on the paper fragments to obtain classification results for each of the paper fragments;
[0039] A determination module, configured to calculate a classification statistical parameter based on the classification probability of the token, and determine a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model;
[0040] Among them, the construction process of the training corpus of the classification model that has completed training includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample paper based on the token set to obtain training corpus generation.
[0041] Furthermore, the device further comprises:
[0042] A segmentation module is used to obtain a sample paper and segment the sample paper according to a preset character length to obtain a plurality of sample paper segments;
[0043] A matching module is used to segment the sample paper fragments and match each original token obtained by segmentation with the tokens in the global token set to obtain the number of token hits for each sample paper fragment;
[0044] A calculation module, configured to determine the number of token replacements for each of the sample paper fragments based on the number of token hits;
[0045] A replacement module is used to replace the tokens in each of the sample paper fragments according to the token replacement strategy and the corresponding token replacement quantity to obtain a training corpus;
[0046] The token replacement strategy includes at least one of a first replacement strategy for performing replacement based on the token cluster list set, a second replacement strategy for performing replacement based on the global token set, and a third replacement strategy for performing no replacement.
[0047] Furthermore, the replacement module includes:
[0048] A first determining unit is configured to randomly determine, for each of the sample paper fragments, a position to be replaced that satisfies the number of token replacements;
[0049] a second determining unit, configured to determine, for each of the positions to be replaced, a target replacement strategy from the replacement strategies using a replacement strategy probability generator if the token of the position to be replaced exists in the global token set;
[0050] A replacement unit, configured to replace the position to be replaced according to the target replacement strategy to obtain a replaced sample paper fragment;
[0051] The marking unit is used to mark the token corresponding to each position based on the comparison results of the replaced sample paper fragment and the sample paper fragment at each position to obtain a training corpus.
[0052] Furthermore, the device further comprises:
[0053] The acquisition module is further configured to acquire token tables of multiple natural language generation models and token vector embedding representations corresponding to tokens in the token tables;
[0054] A processing module, configured to merge and remove duplicates from the token tables to obtain a global token set;
[0055] A clustering module is used to cluster the tokens in each of the token tables based on the token vector embedding representation, and generate a token cluster list set based on each clustering result.
[0056] Furthermore, the clustering module includes:
[0057] A calculation unit, configured to calculate the number of cluster categories of tokens corresponding to each of the natural language generation models based on the number of tokens in each of the token tables;
[0058] Clustering the tokens in each of the token tables based on the corresponding token vector embedding representation and the number of cluster categories to obtain a clustering result for each of the natural language generation models, the clustering result including a plurality of token clusters;
[0059] A merging unit is used to merge the token clusters corresponding to the same token in each of the clustering results into a token cluster list to obtain a token cluster list set, where there is a token intersection between the token cluster lists.
[0060] Furthermore, the device further comprises:
[0061] The acquisition module is further configured to acquire training corpus and a token set, and merge the token set with the original token table of the Bert model to obtain an expanded Berttoken table; and construct an initial classification model based on the Bert model and the expanded Berttoken table;
[0062] The training module is used to train the initial classification model using the training corpus to obtain a classification model that has completed training.
[0063] Furthermore, in a specific application scenario, the classification statistical parameters include a global token classification probability mean, a first-class token classification probability mean, and a second-class token quantity ratio; the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio;
[0064] The determination module is also used to determine that the paper to be detected is a paper generated based on a natural language generation model if any one of the global token classification probability mean, the first-category token classification probability mean, and the second-category token quantity ratio is greater than the corresponding parameter threshold.
[0065] According to another aspect of the present invention, a storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction enables a processor to perform operations corresponding to the above-mentioned paper detection method.
[0066] According to another aspect of the present invention, there is provided a terminal, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0067] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned paper detection method.
[0068] By means of the above technical solution, the technical solution provided by the embodiment of the present invention has at least the following advantages:
[0069] The present invention provides a paper detection method and device, a storage medium, and a terminal. The embodiment of the present invention obtains a paper to be detected and divides the paper to be detected into segments to obtain multiple paper fragments; uses a trained classification model to predict the paper fragments to obtain classification results of each paper fragment, and the classification results include the classification probability of each token; calculates classification statistical parameters based on the classification probability of the token, and determines the detection result of the paper to be detected based on the comparison result of the classification statistical parameters with a preset parameter threshold. The detection result is used to indicate whether the paper to be detected is based on natural language generation. Model generation; wherein, the construction process of the training corpus of the classification model that has completed training includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample papers based on the token set to obtain training corpus generation, and obtaining training corpus by selecting Chinese characters, words, numbers, and punctuation marks to replace the original paper content, thereby improving the diversity and comprehensiveness of the training samples, and training the classification model based on the training corpus, increasing the learning difficulty of the model, realizing self-supervised learning of the model, and greatly improving the classification prediction ability of the classification model, thereby improving the accuracy of paper detection.
[0070] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0072] Figure 1 A flow chart of a paper detection method provided by an embodiment of the present invention is shown;
[0073] Figure 2 A flow chart of another paper detection method provided by an embodiment of the present invention is shown;
[0074] Figure 3 The following is a block diagram showing the composition of a paper detection device provided by an embodiment of the present invention;
[0075] Figure 4 A schematic structural diagram of a terminal provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0076] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0077] The existing detection of papers generated by AI mainly uses the paper content generated by the AI model as training samples to complete the model training, and then identifies whether the paper is generated based on AI based on the trained model. However, this method is too dependent on the paper training samples and can only learn the generation logic of the model that generates the training samples. The generalization ability of the model is weak, and there are many types of models that can be used to generate papers, resulting in low accuracy in detecting whether the paper is generated based on the AI model. The embodiment of the present invention provides a paper detection method, such as Figure 1 As shown, the method includes:
[0078] 101. Obtain a paper to be tested, and segment the paper to be tested to obtain multiple paper segments.
[0079] In an embodiment of the present invention, the paper to be detected is a Chinese academic paper whose content needs to be detected to see if it is generated based on an AI model. Among them, the paper to be detected can be a doctoral dissertation, a bachelor's thesis, or a professional title thesis, and the subjects involved can be natural sciences, literature, sociology, etc. The embodiment of the present invention does not limit the application scenarios, subjects involved, and professional categories of the paper. Due to the restrictions of the model on the input text, after obtaining the paper to be detected, the content of the paper is divided into paper fragments with a token number of less than 512 tokens. During the segmentation process, the paper content can be directly divided into paper fragments with 512 tokens per segment according to 512 tokens, or it can be segmented according to complete sentences in the paper. While ensuring the integrity of the sentences, it ensures that the number of tokens contained in the paper fragments is less than 512 tokens. The embodiment of the present invention does not make specific restrictions. This is to facilitate the subsequent prediction of each paper fragment based on the classification model that has been trained.
[0080] 102. Use the trained classification model to perform prediction processing on the paper fragments to obtain classification results for each of the paper fragments.
[0081] In an embodiment of the present invention, a paper fragment is input into a classification model that has been trained, and the probability of each token in the paper fragment being generated by the AI model is predicted based on the model, so as to obtain the probability that each token in each paper fragment is generated by the AI model, that is, the classification result. Among them, the classification model that has been trained is an autoencoder binary classification model. Its loss function is the cross entropy function of the predicted token classification probability and the corresponding token classification label. Among them, the classification model that has been trained is obtained based on a pre-constructed training corpus. The construction process of the training corpus includes: obtaining a sample paper and a token set, and replacing the tokens of the sample paper based on the token set to obtain the training corpus. Among them, the sample paper is a paper generated by a non-AI model, and can be a paper in the same or similar academic field as the paper to be tested, or a paper corresponding to each academic field. The selection of paper samples can be customized according to the actual application scenario, and the embodiment of the present invention does not make specific limitations. After obtaining the sample paper, the tokens with the same or similar semantics in the sample paper are replaced based on the tokens in the token set to obtain the training corpus. Among them, the token set is generated based on the token tables of multiple currently mainstream pre-trained language generation models. The specific generation method can be a simple merge, for example, the token tables of various pre-trained language generation models are combined together, and the set obtained by deduplicating the same tokens can be obtained; it can also be a cluster merge, for example, clustering is performed according to the token semantics, and the clustered tokens are further merged according to the semantics. The embodiment of the present invention does not make specific limitations.
[0082] It should be noted that since the token set includes token tables of multiple currently mainstream pre-trained semantic models, the tokens in the sample papers are replaced based on this token set, and the resulting training corpus contains different word usages of multiple pre-trained language generation models for the same semantics, and the word usage of each replaced token is not limited to the same model. For example, the Atoken and Btoken in a sentence will be replaced, and the A'token used to replace the Atoken comes from Model A, and the B'token used to replace the Btoken comes from Model B. Therefore, using this training corpus to train the model can break the limitations of model training on the corpus generated by one or several language generation models, so that the classification model can learn the word usage characteristics of different language generation models during the training process, greatly improving the generalization ability of the model, and thus improving the accuracy of the model classification.
[0083] 103. Calculate the classification statistical parameters based on the classification probability of the token, and determine the detection result of the paper to be detected based on the comparison result of the classification statistical parameters and the preset parameter threshold.
[0084] In the embodiment of the present invention, a token is the smallest unit of the word segmentation result and the smallest unit of the natural language generation model vocabulary. It can be a single word, a phrase, a number, a punctuation mark, a symbol, etc. The detection result is used to indicate whether the paper to be detected is generated based on the natural language generation model. The classification result only includes the classification prediction results of each token in the paper to be detected, and can only characterize the probability that a single token in the paper to be detected is generated by the natural language generation model. Since a paper includes tens of thousands of tokens, it is necessary to perform statistical calculations on the prediction results of each token to obtain the classification statistical parameters used to determine whether the entire paper to be detected is generated based on the natural language generation model. And based on the comparison results of the classification statistical parameters and the preset parameter threshold, it is identified whether the paper to be detected is generated based on the natural language generation model. Among them, the classification statistical parameters are not limited to one, and the corresponding preset parameter threshold is also not limited to one. The detection result can be determined based on the comparison results of a single classification statistical parameter and the corresponding preset parameter threshold, or it can be determined based on the comparison results of multiple classification statistical parameters and the corresponding preset parameter threshold. The embodiment of the present invention does not make specific limitations. The classification statistical parameters may include the mean of the classification probabilities of all tokens in the paper to be tested, the mean of the classification probabilities of some tokens, for example, the mean of the top 10% of tokens ranked from largest to smallest by token classification probability, the number of tokens with a probability greater than a preset probability threshold, for example, the number of tokens with a classification probability greater than 50%, etc., which are not specifically limited in the embodiments of the present invention. Statistics are performed based on the classification prediction results of the global tokens in the paper to be tested, and the detection results are identified based on the statistical results, which fully considers the global nature of paper detection and further ensures the accuracy of paper detection.
[0085] In one embodiment of the present invention, for further explanation and limitation, as Figure 2 As shown in Figure 2, the construction process of the training corpus includes:
[0086] 201. Obtain a sample paper, and segment the sample paper according to a preset character length to obtain a plurality of sample paper segments.
[0087] 202. Segment the sample paper fragments, and match each original token obtained by segmentation with the tokens in the global token set to obtain the number of token hits of each sample paper fragment.
[0088] 203. Determine the number of token replacements for each of the sample paper fragments based on the number of token hits.
[0089] 204. For each of the sample paper fragments, replace the tokens in the sample paper fragment according to the token replacement strategy and the corresponding token replacement quantity to obtain a training corpus.
[0090] In an embodiment of the present invention, the sample paper can be a paper extracted from a paper library, for example, 1 million Chinese academic papers are extracted from the paper library of the academic library of the A network. After obtaining the sample paper, the sample paper is cut into sample paper segments less than the preset character length according to punctuation marks. During the segmentation process, the last sentence of the previous sample paper segment is used as the starting sentence of the next segment for iteration. In order to avoid the token length of the sample paper segment changing due to token replacement, and generating a segment with a length exceeding 512 tokens, it is necessary to set the preset character length to a length less than 512 tokens, for example, 480 tokens, so as to reserve the token length for the length change of token replacement. After the segmentation is completed, each sample paper segment is segmented, and each word or token in the segmentation result is matched with the character or word in the global token set. In the segmentation process, the forward maximum matching segmentation method, the reverse maximum matching segmentation method or other segmentation methods can be used, and the embodiment of the present invention does not make specific limitations. The tokens in the sample paper fragment that match the content in the global token set are identified as hit tokens, and the number of hit tokens is counted. The token hit number is set to SegN, and based on the formula WRN = SegN / 10(1); the token replacement number WRN is calculated, and then the tokens in the sample paper fragment are replaced according to the token replacement number.
[0091] It should be noted that the token set includes a token cluster list set and a global token set. Among them, the token replacement strategy includes at least one of a first replacement strategy based on the token cluster list set, a second replacement strategy based on the global token set, and a third replacement strategy without replacement. That is, in the process of replacing tokens in the paper sample fragment based on the token replacement strategy, the replacement operation performed on any position to be replaced may be to take a token from the token cluster list set for replacement, or to take a token from the global token set for replacement, or to not replace. By replacing the tokens in the paper sample fragment with the token table of the natural language generation model to obtain the training corpus, it is possible to convert the words in the paper sample fragment into the words of the natural language generation model, and avoid the limitations of the overall paper generation based on the natural language generation model, thereby improving the diversity and comprehensiveness of the training corpus.
[0092] In an embodiment of the present invention, for further illustration and limitation, the steps are as follows: for each of the sample paper fragments, tokens in the sample paper fragments are replaced according to the token replacement strategy and the corresponding token replacement quantity, and the training corpus obtained includes:
[0093] For each of the sample paper fragments, randomly determine the positions to be replaced that meet the token replacement quantity;
[0094] For each of the positions to be replaced, if the token at the position to be replaced exists in the global token set, determine the target replacement strategy from the first replacement strategy, the second replacement strategy, and the third replacement strategy through the replacement strategy probability generator;
[0095] Replace the position to be replaced according to the target replacement strategy to obtain the replaced sample paper fragment;
[0096] Based on the comparison results of the replaced sample paper fragment and the sample paper fragment at each position, mark the tokens corresponding to each position to obtain the training corpus.
[0097] In an embodiment of the present invention, taking a paper sample fragment as an example, the token replacement process is described. Generate a set of non-repeating random numbers RN{R1, R2,..., Rn} according to the token replacement quantity, where Rn represents the Rnth token position in the paper sample fragment, 0 <= Rn < segN. Judge the word segmentation results of each position corresponding to the random numbers in the RN set. If the word segmentation result of the current position is not in the global token set, no operation is performed on this position; if the word segmentation result of the current position is in the global token set, replace it according to the token replacement strategy. Among them, each replacement strategy in the token replacement strategy is configured with a corresponding execution probability. This execution probability is determined based on the replacement strategy probability generator. For example, the probability of configuring the replacement strategy probability generator is that the occurrence probability of the first replacement strategy is 80%, the occurrence probability of the second replacement strategy is 10%, and the occurrence probability of the third replacement strategy is 10%. Then, the replacement strategy to be executed currently, that is, the target replacement strategy, is assigned by the replacement strategy probability generator. After completing the token replacement, record the positions where the replacement occurs, and compare them with the paper sample fragment before replacement. Configure the classification label of the position where the change occurs as 1, and configure the token classification label of the position where no change occurs as 0 to obtain the training corpus. For example, the paper sample fragment is:
[0098] "However, limited by the development environment and its own reasons, small and medium-sized enterprises have encountered development bottlenecks and are facing serious challenges."
[0099] Word segmentation results:
[0100] However, due to the limitations of the development environment and their own reasons, small and medium-sized enterprises have encountered development bottlenecks and are facing serious challenges.
[0101] Original labels = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
[0102] Among them, "limited by" and "encounter" exist in the W vocabulary list respectively. A word is randomly selected from the cluster list of each word and replaced. The selected words are "limit" and "encounter", and the position labels of these two words are changed to 1 respectively.
[0103] Replacement result:
[0104] "However, due to the restrictive development environment and their own reasons, small and medium-sized enterprises have encountered development bottlenecks and face serious challenges."
[0105] Modify label = [0,0,0,1,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0].
[0106] During the classification model training phase, self-supervised learning is achieved by selecting relevant Chinese characters, words, numbers, and punctuation marks to replace the original content, thereby increasing the learning difficulty of the model. It also has a higher accuracy rate for vocabulary classification with diverse content, thereby improving the model's discrimination ability.
[0107] In one embodiment of the present invention, for further explanation and limitation, before matching each original token obtained by word segmentation with the tokens in the global token set to obtain the number of token hits for each sample paper fragment, the method further includes:
[0108] Obtaining token tables of multiple natural language generation models and token vector embedding representations corresponding to tokens in the token tables;
[0109] Merge and remove duplicates from the token tables to obtain a global token set;
[0110] The tokens in each of the token tables are clustered based on the token vector embedding representation, and a token clustering list set is generated based on each clustering result.
[0111] In an embodiment of the present invention, in order to ensure the comprehensiveness of the natural language generation vocabulary covered by the training corpus, token tables of existing open source natural language generation models, such as the Baichuan model, the Harbin Institute of Technology iFlytek Language Cloud, and the Tsinghua Natural Language Open Source Model, are obtained. Since the sample paper is in Chinese, if the token table contains tokens other than Chinese characters, words, numbers, punctuation marks, and symbols, they are deleted. The tokens in the token tables corresponding to each model are merged into a token set, and the repeated tokens therein are deduplicated to obtain a global token set. In order to further improve the semantic information of the token, the tokens corresponding to each model are clustered based on the token vector embedding representation corresponding to each model, and a token clustering list set is generated based on the clustering results of the tokens corresponding to each model. By summarizing and clustering the token tables corresponding to each natural language generation model, a complete and comprehensive token can be obtained, and the language association between the tokens can be obtained, thereby providing comprehensive and accurate token information of the natural language generation model for the generation of subsequent training corpus.
[0112] In one embodiment of the present invention, for further explanation and limitation, clustering the tokens in each of the token tables based on the token vector embedding representation, and generating a token cluster list set based on each clustering result includes:
[0113] Calculating the number of cluster categories of tokens corresponding to each of the natural language generation models based on the number of tokens in each of the token tables;
[0114] For each token in the token table, clustering is performed based on the corresponding token vector embedding representation and the number of cluster categories to obtain a clustering result of each natural language generation model;
[0115] The token clusters corresponding to the same token in each of the clustering results are merged into a token cluster list to obtain a token cluster list set.
[0116] In the embodiment of the present invention, the number of cluster categories is determined based on the number of tokens in each token table, for example,
[0117] Model A: {token1[0.01,0.002,...],token2[...],token3[...],...}; the number of tokens in Model A is TNa;
[0118] Model B: {token1[0.01,0.002,...],token2[...],token3[...],...}; the number of tokens in Model B is TNb;
[0119] Model C: {token1[0.01, 0.002, ...], token2[...], token3[...], ...}; the number of tokens in Model C is TNc;
[0120] Assuming the number of cluster categories of each model is CNa, CNb, and CNc, the number of cluster categories of each model is calculated as:
[0121] CNa=TNa / M(2); CNb=TNb / M(3); CNc=TNc / M(4);
[0122] Wherein, M is an adjustment coefficient, and M can be customized within a numerical range of 1 to 5 according to the requirements of a specific scenario, and is not specifically limited in the embodiment of the present invention.
[0123] After determining the number of cluster categories, the tokens are clustered based on the token vector embedding representation until the number of cluster categories is equal to or less than the number of cluster categories, and the clustering results of each model are obtained. Specifically, the similarity between each vector can be calculated based on the vector dot method, and the tokens can be merged and clustered from high to low similarity. Clustering can also be performed based on other clustering methods, which are not specifically limited in the embodiments of the present invention. Among them, the clustering results include multiple token clusters, and the token clusters include multiple tokens. For example, ACL{Ac1, Ac2, Ac3, ...} is the token clustering result of model A, wherein Ac1, Ac2, and Ac3 each represent a token cluster, and Ac1 includes multiple tokens such as token19 and token25. After completing the clustering of the tokens corresponding to each model, the clustering results of each model are merged. Specifically, the token clusters corresponding to the same token in each clustering result are merged into a token cluster list. For example, for token 1, the token clusters containing token 1 in the clustering results of each model are merged into a token cluster list, token 1 list = [token 3, token 89...]. For token 2, the token clusters containing token 2 in the clustering results of each model are merged into a token cluster list, token 2 list = [token 118, token 8...]. The token cluster lists obtained after the merger have token intersections. For example, token 3 list = [token 1, token 280...]. Since token 1 and token 3 may exist in the same token cluster, the token 1 list and the token 3 list contain the same token cluster, and therefore, the same token exists. By clustering the token tables of each model and merging the clustering results, the token replacement process generated by the training corpus can be replaced based on the semantic relevance between tokens, so that the semantics of the replaced tokens can better meet the context of the sentence, thereby obtaining more accurate training corpus.
[0124] In one embodiment of the present invention, for further explanation and limitation, before the step of performing prediction processing on the paper fragments using the trained classification model to obtain the classification results of each of the paper fragments, the method further includes:
[0125] Obtain training corpus and token set, and merge the token set with the original token table of the Bert model to obtain the expanded Berttoken table;
[0126] An initial classification model is constructed based on the Bert model and the expanded Berttoken table;
[0127] The initial classification model is trained using the training corpus to obtain a classification model that has completed training.
[0128] In an embodiment of the present invention, an autoencoding binary classification model, i.e., an initial classification model, is constructed based on the Bert model. In order to make the vocabulary of the classification model match the training corpus, the global token set in the token set is merged with the Berttoken table and deduplicated to obtain an expanded Berttoken table. The token classification probability in the training corpus is predicted by the initial classification model, and the cross entropy of the predicted token classification probability and the corresponding token classification label is used as the loss function. The initial classification model is continuously trained so that the loss function converges to a preset cross entropy threshold, and the training of the model is completed to obtain a classification model that has been trained. Classifying tokens based on the autoencoding binary classification model enables the classification process to take into account all preceding and subsequent tokens, which is in line with the rules of human paper writing and is significantly different from the autoregressive model commonly used in AI content generation, which relies entirely on the generation method of the preceding sequence.
[0129] In one embodiment of the present invention, for further explanation and limitation, the classification statistical parameters include a global token classification probability mean, a first-class token classification probability mean, and a second-class token quantity ratio, and the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio;
[0130] The step of determining the detection result of the paper to be detected based on the comparison result of the classification statistical parameter and the preset parameter threshold comprises:
[0131] If any one of the global token classification probability mean, the first-category token classification probability mean, and the second-category token quantity ratio is greater than the corresponding parameter threshold, it is determined that the paper to be detected is a paper generated based on a natural language generation model.
[0132] In an embodiment of the present invention, the classification probability is the probability that a token is classified into a category generated by a natural language generation model. The global token classification probability mean is the mean of the classification probabilities of all tokens in the paper to be detected. The first-category token classification probability mean is the probability mean of a preset percentage of tokens in the paper to be detected before the classification probabilities are arranged from large to small, wherein the preset percentage can be 10%, 5%, 15%, or can be customized according to specific application requirements, which is not specifically limited in the embodiment of the present invention. The second-category token quantity ratio is the ratio of the number of tokens whose classification probability is higher than the preset probability threshold to the total number of tokens. Among them, the preset probability threshold can be 50%, or can be customized according to specific application requirements, which is not specifically limited in the embodiment of the present invention. Only when the global token classification probability mean is greater than the first parameter threshold, the first-category token classification probability mean is greater than the second parameter threshold, and the second-category token quantity ratio is greater than the third parameter threshold, is the paper to be detected determined to be a paper generated based on a natural language generation model, otherwise the paper to be detected is determined to be a paper generated by a non-natural language generation model. The first parameter threshold may be 0.35, the second parameter threshold may be 0.45, and the point parameter threshold may be 0.5. These values may also be customized according to application requirements, and are not specifically limited in the embodiment of the present invention.
[0133] The present invention provides a paper detection method. In an embodiment of the present invention, a paper to be detected is obtained and segmented to obtain multiple paper fragments; the paper fragments are predicted and processed using a classification model that has been trained to obtain classification results of each paper fragment, and the classification results include classification probabilities of each token; classification statistical parameters are calculated based on the classification probabilities of the tokens, and the detection results of the paper to be detected are determined based on the comparison results of the classification statistical parameters and preset parameter thresholds, and the detection results are used to indicate whether the paper to be detected is generated based on a natural language generation model; wherein, the construction process of the training corpus of the classification model that has been trained includes: generating a token set based on token tables of multiple natural language generation models, and replacing tokens in sample papers based on the token set to obtain training corpus generation, and obtaining training corpus by selecting Chinese characters, words, numbers, and punctuation marks to replace the original paper content to obtain the training corpus, thereby improving the diversity and comprehensiveness of the training samples, and training the classification model based on the training corpus to increase the learning difficulty of the model, realize self-supervised learning of the model, greatly improve the classification prediction ability of the classification model, and thus improve the accuracy of paper detection.
[0134] Furthermore, as a response to the above Figure 1 The embodiment of the present invention provides a paper detection device, such as Figure 3 As shown, the device includes:
[0135] An acquisition module 31 is used to acquire a paper to be detected and segment the paper to be detected to obtain multiple paper segments;
[0136] The prediction processing module 32 is used to perform prediction processing on the paper fragments using the trained classification model to obtain classification results for each of the paper fragments;
[0137] A determination module 33 is configured to calculate a classification statistical parameter based on the classification probability of the token, and determine a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model;
[0138] Among them, the construction process of the training corpus of the classification model that has completed training includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample paper based on the token set to obtain training corpus generation.
[0139] Furthermore, the device further comprises:
[0140] A segmentation module is used to obtain a sample paper and segment the sample paper according to a preset character length to obtain a plurality of sample paper segments;
[0141] A matching module is used to segment the sample paper fragments and match each original token obtained by segmentation with the tokens in the global token set to obtain the number of token hits for each sample paper fragment;
[0142] A calculation module, configured to determine the number of token replacements for each of the sample paper fragments based on the number of token hits;
[0143] A replacement module is used to replace the tokens in each of the sample paper fragments according to the token replacement strategy and the corresponding token replacement quantity to obtain a training corpus;
[0144] The token replacement strategy includes at least one of a first replacement strategy for performing replacement based on the token cluster list set, a second replacement strategy for performing replacement based on the global token set, and a third replacement strategy for performing no replacement.
[0145] Furthermore, the replacement module includes:
[0146] A first determining unit is configured to randomly determine, for each of the sample paper fragments, a position to be replaced that satisfies the number of token replacements;
[0147] a second determining unit, configured to determine, for each of the positions to be replaced, a target replacement strategy from the replacement strategies using a replacement strategy probability generator if the token of the position to be replaced exists in the global token set;
[0148] A replacement unit, configured to replace the position to be replaced according to the target replacement strategy to obtain a replaced sample paper fragment;
[0149] The marking unit is used to mark the token corresponding to each position based on the comparison results of the replaced sample paper fragment and the sample paper fragment at each position to obtain a training corpus.
[0150] Furthermore, the device further comprises:
[0151] The acquisition module 31 is further configured to acquire token tables of multiple natural language generation models and token vector embedding representations corresponding to tokens in the token tables;
[0152] A processing module, configured to merge and remove duplicates from the token tables to obtain a global token set;
[0153] A clustering module is used to cluster the tokens in each of the token tables based on the token vector embedding representation, and generate a token cluster list set based on each clustering result.
[0154] Furthermore, the clustering module includes:
[0155] A calculation unit, configured to calculate the number of cluster categories of tokens corresponding to each of the natural language generation models based on the number of tokens in each of the token tables;
[0156] Clustering the tokens in each of the token tables based on the corresponding token vector embedding representation and the number of cluster categories to obtain a clustering result for each of the natural language generation models, the clustering result including a plurality of token clusters;
[0157] A merging unit is used to merge the token clusters corresponding to the same token in each of the clustering results into a token cluster list to obtain a token cluster list set, where there is a token intersection between the token cluster lists.
[0158] Furthermore, the device further comprises:
[0159] The acquisition module 31 is further used to acquire training corpus and token set, and merge the token set with the original token table of the Bert model to obtain an expanded Berttoken table;
[0160] An initial classification model is constructed based on the Bert model and the expanded Berttoken table;
[0161] The training module is used to train the initial classification model using the training corpus to obtain a classification model that has completed training.
[0162] Furthermore, in a specific application scenario, the classification statistical parameters include a global token classification probability mean, a first-class token classification probability mean, and a second-class token quantity ratio; the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio;
[0163] The determination module 33 is further configured to determine that the paper to be detected is a paper generated based on a natural language generation model if any one of the global token classification probability mean, the first-category token classification probability mean, and the second-category token quantity ratio is greater than the corresponding parameter threshold.
[0164] The present invention provides a paper detection device. In an embodiment of the present invention, a paper to be detected is obtained and segmented to obtain multiple paper fragments; the paper fragments are predicted and processed using a classification model that has been trained to obtain classification results of each paper fragment, and the classification results include classification probabilities of each token; classification statistical parameters are calculated based on the classification probabilities of the tokens, and the detection results of the paper to be detected are determined based on the comparison results of the classification statistical parameters and preset parameter thresholds, and the detection results are used to indicate whether the paper to be detected is generated based on a natural language generation model; wherein, the construction process of the training corpus of the classification model that has been trained includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample paper based on the token set to obtain training corpus generation, and obtaining training corpus by selecting Chinese characters, words, numbers, and punctuation marks to replace the original paper content to obtain the training corpus, thereby improving the diversity and comprehensiveness of the training samples, and training the classification model based on the training corpus to increase the learning difficulty of the model, realize self-supervised learning of the model, greatly improve the classification prediction ability of the classification model, and thus improve the accuracy of paper detection.
[0165] According to one embodiment of the present invention, a storage medium is provided, wherein the storage medium stores at least one executable instruction, and the computer executable instruction can execute the paper detection method in any of the above method embodiments.
[0166] Figure 4 A schematic structural diagram of a terminal provided according to an embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the terminal.
[0167] like Figure 4 As shown, the terminal may include: a processor (processor) 402 , a communications interface (Communications Interface) 404 , a memory (memory) 406 , and a communication bus 408 .
[0168] The processor 402 , the communication interface 404 , and the memory 406 communicate with each other via a communication bus 408 .
[0169] The communication interface 404 is used to communicate with other devices such as clients or other servers.
[0170] The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the above-mentioned paper detection method embodiment.
[0171] Specifically, the program 410 may include program codes, which include computer operation instructions.
[0172] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The one or more processors included in the terminal may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.
[0173] The memory 406 is used to store the program 410. The memory 406 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0174] The program 410 may be specifically configured to cause the processor 402 to perform the following operations:
[0175] Obtaining a paper to be tested, and segmenting the paper to be tested to obtain multiple paper segments;
[0176] Using the trained classification model to perform prediction processing on the paper fragments, obtaining classification results for each of the paper fragments, the classification results including the classification probability of each token;
[0177] Calculating a classification statistical parameter based on the classification probability of the token, and determining a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model;
[0178] Among them, the construction process of the training corpus of the classification model that has completed training includes: generating a token set based on the token tables of multiple natural language generation models, and replacing the tokens in the sample paper based on the token set to obtain training corpus generation.
[0179] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0180] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A paper detection method, characterized in that: include: Obtaining a paper to be tested, and segmenting the paper to be tested to obtain multiple paper segments; Using the trained classification model to perform prediction processing on the paper fragments, obtaining classification results for each of the paper fragments, the classification results including the classification probability of each token; Calculating a classification statistical parameter based on the classification probability of the token, and determining a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model; The process of constructing the training corpus of the classification model that has been trained includes: generating a token set based on token tables of multiple natural language generation models, and replacing tokens in sample papers based on the token set to obtain training corpus; Among them, the classification statistical parameters include the global token classification probability mean, the first-class token classification probability mean, the first-class token classification probability mean is the probability mean of a preset percentage of tokens in the paper to be tested before the classification probability is arranged from large to small, the second-class token quantity ratio, the second-class token quantity ratio is the ratio of the number of tokens with a classification probability higher than a preset probability threshold to the total number of tokens, and the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio; The detection result of the paper to be detected is determined based on the comparison result of the classification statistical parameter and the preset parameter threshold, including: if any one of the global token classification probability mean, the first-class token classification probability mean, and the second-class token quantity ratio is greater than the corresponding parameter threshold, then it is determined that the paper to be detected is a paper generated based on the natural language generation model.
2. The method according to claim 1, characterized in that The token set includes a token cluster list set and a global token set. The tokens in the sample paper are replaced based on the token set to generate the training corpus, including: Obtaining a sample paper, and segmenting the sample paper according to a preset character length to obtain a plurality of sample paper segments; Segment the sample paper fragments, and match each original token obtained by segmentation with the tokens in the global token set to obtain the number of token hits for each sample paper fragment; Determining the number of token replacements for each of the sample paper fragments based on the number of token hits; For each of the sample paper fragments, replacing the tokens in the sample paper fragments according to the token replacement strategy and the corresponding token replacement quantity to obtain a training corpus; The token replacement strategy includes at least one of a first replacement strategy for performing replacement based on the token cluster list set, a second replacement strategy for performing replacement based on the global token set, and a third replacement strategy for performing no replacement.
3. The method according to claim 2, characterized in that Each replacement strategy in the token replacement strategy is respectively configured with a corresponding execution probability. For each sample paper fragment, the tokens in the sample paper fragment are replaced according to the token replacement strategy and the corresponding token replacement quantity, and the training corpus obtained includes: For each of the sample paper fragments, randomly determine a position to be replaced that satisfies the number of token replacements; For each of the positions to be replaced, if the token of the position to be replaced exists in the global token set, a target replacement strategy is determined from the replacement strategies by a replacement strategy probability generator; Replace the position to be replaced according to the target replacement strategy to obtain a replaced sample paper fragment; Based on the comparison results of the replaced sample paper fragment and the sample paper fragment at each position, the token corresponding to each position is marked to obtain the training corpus.
4. The method according to claim 1, wherein The token table based on multiple natural language generation models generates a token set including: Obtaining token tables of multiple natural language generation models and token vector embedding representations corresponding to tokens in the token tables; Merge and remove duplicates from the token tables to obtain a global token set; The tokens in each of the token tables are clustered based on the token vector embedding representation, and a token clustering list set is generated based on each clustering result.
5. The method according to claim 4, characterized in that The step of clustering the tokens in each of the token tables based on the token vector embedding representation and generating a token cluster list set based on each clustering result includes: Calculating the number of cluster categories of tokens corresponding to each of the natural language generation models based on the number of tokens in each of the token tables; Clustering the tokens in each of the token tables based on the corresponding token vector embedding representation and the number of cluster categories to obtain a clustering result for each of the natural language generation models, the clustering result including a plurality of token clusters; Token clusters corresponding to the same token in each of the clustering results are merged into a token cluster list to obtain a token cluster list set, where there is a token intersection between the token cluster lists.
6. The method according to claim 1, characterized in that Before performing prediction processing on the paper fragments using the trained classification model to obtain classification results for each of the paper fragments, the method further includes: Obtain training corpus and token set, and merge the token set with the original token table of the Bert model to obtain the expanded Berttoken table; An initial classification model is constructed based on the Bert model and the expanded Berttoken table; The initial classification model is trained using the training corpus to obtain a classification model that has completed training.
7. A paper detection device, characterized in that: include: An acquisition module is used to acquire the paper to be detected and divide the paper to be detected into segments to obtain multiple paper segments; A prediction processing module is used to use the trained classification model to perform prediction processing on the paper fragments to obtain classification results for each paper fragment, wherein the classification results include the classification probability of each token; A determination module, configured to calculate a classification statistical parameter based on the classification probability of the token, and determine a detection result of the paper to be detected based on a comparison result of the classification statistical parameter and a preset parameter threshold, wherein the detection result is used to indicate whether the paper to be detected is generated based on a natural language generation model; The process of constructing the training corpus of the classification model that has been trained includes: generating a token set based on token tables of multiple natural language generation models, and replacing tokens in sample papers based on the token set to obtain training corpus; Among them, the classification statistical parameters include the global token classification probability mean, the first-class token classification probability mean, the first-class token classification probability mean is the probability mean of a preset percentage of tokens in the paper to be tested before the classification probability is arranged from large to small, the second-class token quantity ratio, the second-class token quantity ratio is the ratio of the number of tokens with a classification probability higher than a preset probability threshold to the total number of tokens, and the preset parameter thresholds include a first parameter threshold corresponding to the global token classification probability mean, a second parameter threshold corresponding to the first-class token classification probability mean, and a third parameter threshold corresponding to the second-class token quantity ratio; The detection result of the paper to be detected is determined based on the comparison result of the classification statistical parameter and the preset parameter threshold, including: if any one of the global token classification probability mean, the first-class token classification probability mean, and the second-class token quantity ratio is greater than the corresponding parameter threshold, then it is determined that the paper to be detected is a paper generated based on the natural language generation model.
8. A storage medium storing at least one executable instruction, wherein the executable instruction enables a processor to perform operations corresponding to the paper detection method according to any one of claims 1 to 6.
9. A terminal comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the paper detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text classification method, text classification device, computer equipment and storage medium
CN115640394A
Classification model training method and related device
CN116401552A