Detection method for paper generated by artificial intelligence and related device

The three feature channels of the large language model were detected by pre-trained papers to extract features and fusion, which solved the problem of low accuracy in detecting artificial intelligence generated papers in the prior art, and achieved higher classification accuracy.

CN120218085AActive Publication Date: 2025-06-27INST OF MEDICAL INFORMATION CHINESE ACAD OF MEDICAL SCI

Patent Information

Application Number
CN202510704458.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-06-27
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

It is difficult for the existing technology to accurately detect the differences between artificial intelligence-generated papers and human-written papers, resulting in low accuracy of detection results.

Method used

The pre-trained paper detection large language model is used, and word-sentence semantic modeling, text feature statistics and quality evaluation are carried out through three feature channels. The deep semantic features, text statistical features and quality evaluation features of the paper to be processed are extracted, and the feature decoding module is fused to determine whether the paper is generated by artificial intelligence.

Benefits of technology

It improves the classification accuracy of detection results and can more accurately distinguish between artificial intelligence-generated papers and human-written papers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218085A_ABST
    Figure CN120218085A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence generated paper detection method and a related device, and relates to the technical field of artificial intelligence, and the method comprises the steps: respectively extracting a deep semantic feature, a text statistical feature and a quality evaluation feature of a to-be-processed paper through three feature channels of a paper detection large language model, the deep semantic features represent word-level and sentence-level semantic information, part-of-speech semantic information and inter-sentence semantic consistency information of the to-be-processed paper, the text statistical features represent language style features of the to-be-processed paper, and the quality evaluation features represent overall quality features of the to-be-processed paper. And a feature decoding module in the thesis detection large language model is utilized to obtain a target detection result representing whether the to-be-processed thesis is an artificial intelligence generated thesis according to the extracted features. According to the method, the feature information of the to-be-processed papers is extracted from multiple angles through the three feature channels, the to-be-processed papers can be accurately classified, and the classification accuracy of the target detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for detecting papers generated by artificial intelligence and related devices. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the emergence of large language models (LLMs), scholars can use artificial intelligence technology to generate papers.

[0003] On the one hand, the act of using artificial intelligence technology to generate papers may lead to academic misconduct; on the other hand, artificial intelligence technology has some limitations in academic writing, especially in terms of factual accuracy. Therefore, the need to detect whether a paper is generated by artificial intelligence is urgent. Summary of the Invention

[0004] In view of the above problems, this application provides a method for detecting papers generated by artificial intelligence and related devices to achieve the purpose of accurately detecting whether a paper to be processed is a paper generated by artificial intelligence. The specific solutions are as follows:

[0005] The first aspect of this application provides a method for detecting papers generated by artificial intelligence, including:

[0006] Performing word and sentence-level semantic modeling processing on the paper to be processed using the first feature channel in the pre-trained large language model for paper detection to obtain the deep semantic features of the paper to be processed, where the deep semantic features represent the word-level, sentence-level semantic information, part-of-speech semantic information, and inter-sentence semantic consistency information of the paper to be processed;

[0007] Performing text feature statistics processing on the paper to be processed using the second feature channel in the large language model for paper detection to obtain the text statistical features of the paper to be processed, where the text statistical features represent the language style characteristics of the paper to be processed;

[0008] Performing quality assessment processing on the paper to be processed using the third feature channel in the large language model for paper detection to obtain the quality assessment features of the paper to be processed, where the quality assessment features represent the overall quality characteristics of the paper to be processed;

[0009] Using the feature decoding module in the large language model for paper detection to obtain a target detection result representing whether the paper to be processed is a paper generated by artificial intelligence according to the deep semantic features, the text statistical features, and the quality assessment features;

[0010] Among them, the training data used by the large language model for paper detection in the training stage includes: positive sample data labeled with the tags of AI-generated papers and negative sample data labeled with the tags of normal papers. The positive sample data are papers marked as AI-generated content in the retraction watch database, and the negative sample data are papers in the literature database that meet the preset impact factor requirements. The topic similarity between the positive sample data and the negative sample data meets the preset similarity requirements.

[0011] In a possible implementation, before using the first feature channel in the pre-trained large language model for paper detection to perform word and sentence level semantic modeling on the paper to be processed to obtain the deep semantic features of the paper to be processed, using the second feature channel in the large language model for paper detection to perform text feature statistics on the paper to be processed to obtain the text statistical features of the paper to be processed, and using the third feature channel in the large language model for paper detection to perform quality assessment on the paper to be processed to obtain the quality assessment features of the paper to be processed, it further includes:

[0012] Performing structured parsing on the paper to be processed to obtain the corresponding structured paper of the paper to be processed;

[0013] Taking the structured paper as the paper to be processed.

[0014] In a possible implementation, the first feature channel includes a text input module, a feature encoding module, an inter-sentence semantic consistency modeling module, and a part-of-speech semantic distribution modeling module;

[0015] The using the first feature channel in the pre-trained large language model for paper detection to perform word and sentence level semantic modeling on the paper to be processed to obtain the deep semantic features of the paper to be processed includes:

[0016] Performing clause splitting, word segmentation, and part-of-speech tagging on the paper to be processed through the text input module to obtain the word segmentation sequence and the word segmentation part-of-speech tagging sequence corresponding to each clause in the paper to be processed;

[0017] Performing word-level context encoding on each word in the word segmentation sequence corresponding to each clause through the feature encoding module to obtain the word-level encoding vector sequence corresponding to each clause, performing weighted aggregation processing on the word segmentation within the sentence according to the word-level encoding vector sequence corresponding to each clause to obtain the sentence vector corresponding to each clause, performing sentence-level context encoding on the sentence vectors corresponding to all clauses in the paper to be processed to obtain the sentence-level encoding vectors corresponding to all clauses respectively, and performing weighted aggregation processing on the sentence-level encoding vectors corresponding to all clauses respectively to obtain a document vector representing the overall semantic information of the paper to be processed;

[0018] The inter-sentence semantic consistency modeling module constructs an inter-sentence similarity matrix based on the sentence-level encoded vector corresponding to each clause in the to-be-processed paper, and performs index feature extraction on the inter-sentence similarity matrix based on a preset statistical index to obtain a matrix vector representing the internal coherence and semantic consistency of the to-be-processed paper;

[0019] The part-of-speech semantic distribution modeling module performs statistical analysis processing based on the occurrence frequency on all parts of speech in the to-be-processed paper according to the word segmentation part-of-speech tagging sequence and word-level encoded vector sequence corresponding to each clause in the to-be-processed paper, and obtains a part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper.

[0020] In a possible implementation, the part-of-speech semantic distribution modeling module performs statistical analysis processing based on the occurrence frequency on all parts of speech in the to-be-processed paper according to the word segmentation part-of-speech tagging sequence and word-level encoded vector sequence corresponding to each clause in the to-be-processed paper, and obtains a part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper, including:

[0021] The part-of-speech semantic distribution modeling module statistically analyzes the occurrence frequency of each word segment under each part of speech according to the word-level encoded vector sequence corresponding to each clause for each part of speech among all parts of speech, obtains the occurrence frequency of each word segment under each part of speech, performs average pooling processing on the occurrence frequencies of all word segments under each part of speech to obtain the occurrence frequency of each part of speech, so as to obtain the occurrence frequencies of all parts of speech respectively, and splices the occurrence frequencies of all parts of speech respectively to obtain the part-of-speech vector.

[0022] In a possible implementation, the utilization of the second feature channel in the paper detection large language model to perform text feature statistical processing on the to-be-processed paper to obtain the text statistical features of the to-be-processed paper includes:

[0023] The second feature channel performs text feature statistical processing at the lexical level and text feature statistical processing at the sentence level on the to-be-processed paper respectively, and obtains the statistical features at the lexical level and the statistical features at the sentence level of the to-be-processed paper;

[0024] Among them, the lexical level includes one or more of the following levels: lexical richness, the number of unique words, and the usage conditions of quantitative words, stop words, conjunctions, target words, and template words respectively, and the target word refers to a word whose occurrence frequency in the to-be-processed paper is higher than a preset frequency threshold;

[0025] The sentence level includes one or more of the following levels: sentence structure complexity, sentence length change distribution, proportion of synonymous expressions used, and usage conditions of template sentences.

[0026] In a possible implementation, using the third feature channel in the paper detection large language model to perform quality evaluation processing on the to-be-processed paper to obtain the quality evaluation features of the to-be-processed paper includes:

[0027] Performing text quality evaluation on the to-be-processed paper using a preset quality evaluation index to obtain the quality evaluation features of the to-be-processed paper, where the quality evaluation index includes one or more of the following indexes: text readability index, text diversity index, and text predictability index.

[0028] In a possible implementation, using the feature decoding module in the paper detection large language model to obtain a target detection result indicating whether the to-be-processed paper is an AI-generated paper according to the deep semantic features, the text statistical features, and the quality evaluation features includes:

[0029] Using the feature fusion module in the feature decoding module to perform feature fusion based on attention weights on the deep semantic features, the text statistical features, and the quality evaluation features to obtain fused features;

[0030] Using the classification output module in the feature decoding module to obtain the target detection result according to the fused features.

[0031] The second aspect of this application provides a detection device for AI-generated papers, including:

[0032] A deep feature extraction module, configured to use the first feature channel in a pre-trained paper detection large language model to perform multi-level semantic modeling processing on a to-be-processed paper to obtain the deep semantic features of the to-be-processed paper, where the deep semantic features represent the semantic information at the word level and sentence level of the to-be-processed paper;

[0033] A shallow feature extraction module, configured to use the second feature channel in the paper detection large language model to perform text feature statistics processing on the to-be-processed paper to obtain the text statistical features of the to-be-processed paper, where the text statistical features represent the language style information of the to-be-processed paper;

[0034] A macro feature extraction module, configured to use the third feature channel in the paper detection large language model to perform quality evaluation processing on the to-be-processed paper to obtain the quality evaluation features of the to-be-processed paper, where the quality evaluation features represent the overall quality characteristics of the to-be-processed paper;

[0035] A feature fusion decoding module, configured to use the feature decoding module in the paper detection large language model to obtain a target detection result indicating whether the paper to be processed is an AI-generated paper according to the deep semantic feature, the text statistical feature, and the quality evaluation feature;

[0036] Wherein, the training data used by the paper detection large language model in the training stage includes: positive sample data labeled with the label of AI-generated papers and negative sample data labeled with the label of normal papers. The positive sample data is the papers marked as AI-generated content in the retraction watch database, and the negative sample data is the papers in the literature database that meet the preset impact factor requirements. The topic similarity between the positive sample data and the negative sample data meets the preset similarity requirements.

[0037] The third aspect of this application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the method for detecting AI-generated papers in the first aspect or any implementation manner of the first aspect.

[0038] The fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0039] The memory is used to store a computer program;

[0040] The processor is used to execute the computer program so that the electronic device can implement the method for detecting AI-generated papers in the first aspect or any implementation manner of the first aspect.

[0041] The fifth aspect of this application provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement the method for detecting AI-generated papers in the first aspect or any implementation manner of the first aspect.

[0042] With the above technical solutions, the method for detecting AI-generated papers provided by this application takes into account the obvious differences between AI-generated papers and normal papers written by humans in terms of both the deep semantic expression and the shallow language style expression of vocabulary and sentences. For example, compared with AI-generated papers, normal papers have higher semantic richness and more flexible language styles, and are not restricted to certain fixed template words, template sentences, etc. Based on this, this application uses the first feature channel in the pre-trained large language model for paper detection to perform word and sentence-level semantic modeling on the paper to be processed, obtaining the deep semantic features of the paper to be processed, and uses the second feature channel in the large language model for paper detection to perform text feature statistics on the paper to be processed, obtaining the text statistical features of the paper to be processed. Further, compared with AI-generated papers, normal papers are the result of in-depth thinking and research, and their overall quality is often significantly higher than that of AI-generated papers. Therefore, this application also performs quality assessment on the paper to be processed through the third feature channel of the large language model for paper detection, obtaining the quality assessment features of the paper to be processed. Since this application can extract the feature information of the paper to be processed from three perspectives: deep, shallow, and macro, through the three feature channels of the large language model for paper detection, the feature decoding module in the large language model for paper detection can accurately classify the paper to be processed based on the deep semantic features, text statistical features, and quality assessment features, thereby improving the classification accuracy of the target detection results obtained by the feature decoding module. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the elements and elements are not necessarily drawn to scale.

[0044] Figure 1 It is a schematic flowchart of a method for detecting AI-generated papers provided by this application;

[0045] Figure 2 It is a schematic structural diagram of a large language model for paper detection provided by this application;

[0046] Figure 3 It is a schematic structural diagram of another large language model for paper detection provided by this application;

[0047] Figure 4 It is a schematic structural diagram of a device for detecting AI-generated papers provided by this application;

[0048] Figure 5 It is a schematic structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.

[0050] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0051] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0052] Current text classification methods, such as the text classification model based on the XLNet-CNN and XLNet-BiGRU dual channels, can quickly classify text information on website platforms such as forums and blogs. However, the text information on website platforms such as forums and blogs is almost all human-written documents. Using human-written documents as training data to train the above dual-channel text classification model enables the dual-channel text classification model to fully learn the differences between human-written documents and accurately classify human-written documents.

[0053] However, with the rapid rise and development of artificial intelligence technology, it has become possible to generate papers using artificial intelligence technology. And with the increasing research pressure on researchers or for reasons such as the need for some scholars to publish papers, etc., artificial intelligence-generated papers have repeatedly appeared in various literature databases. After in-depth research, the present application has found that the differences between multiple human-written documents are mainly reflected in the differences in the expression habits of different authors. However, the differences between artificial intelligence-generated papers and human-written papers are mainly reflected in aspects such as semantic expression, language style expression, and paper quality. These two types of differences are significantly different. Therefore, based on the dual-channel text classification model to detect whether a paper is an artificial intelligence-generated paper, the accuracy of the detection result is relatively low.

[0054] To accurately detect whether a paper (referred to as the paper to be processed in this application) is an AI-generated paper, this application provides a method for detecting AI-generated papers, which can be applicable to scenarios for detecting whether a paper is an AI-generated paper.

[0055] Optionally, the method for detecting AI-generated papers provided by this application can be applied to a terminal with data processing capabilities, a server with data processing capabilities, or a system composed of a terminal and a server.

[0056] Referring to Figure 1 , which is a schematic flowchart of a method for detecting AI-generated papers provided by an embodiment of this application. As Figure 1 , the method for detecting AI-generated papers can specifically include the following steps:

[0057] Step S101: Use the first feature channel in the pre-trained large language model for paper detection to perform word and sentence-level semantic modeling on the paper to be processed, and obtain the deep semantic features of the paper to be processed.

[0058] This embodiment is used to detect whether the paper to be processed is generated by AI technology. For example, when scholars publish papers in various academic journals, they can first use the scholars' papers as the papers to be processed, and only publish the papers to be processed in academic journals when it is detected that the papers to be processed are not AI-generated papers.

[0059] To improve the detection accuracy, this application pre-trains a large language model for paper detection, which includes three feature channels to extract feature information of the paper to be detected from three perspectives that can significantly distinguish AI-generated papers and normal papers (i.e., papers written by humans).

[0060] Here, the large language model for paper detection is obtained by performing supervised fine-tuning on a pre-constructed large language model using training papers labeled with paper type tags as training data.

[0061] Optionally, the classification cross-entropy is used as the loss function in the fine-tuning stage, and the AdamW optimizer is used for gradient update. Here, AdamW is a variant of the Adam optimizer (Adaptive Moment Estimation), mainly used for the training of deep learning models.

[0062] Optionally, the training papers labeled with paper type tags include: positive sample data labeled with AI-generated paper tags and negative sample data labeled with normal paper tags.

[0063] Optionally, the positive sample data are papers marked as AI Generated Content (AIGC) in the Retraction Watch database, and the negative sample data are papers in a literature database that meet the preset impact factor requirements. The topic similarity between the positive sample data and the negative sample data meets the preset similarity requirements.

[0064] Optionally, the literature database that meets the preset impact factor requirements (such as the PubMed literature database) may refer to the top n literature databases sorted by impact factor, where n is a positive integer. Generally, literature databases with high impact factors often have more rigorous review systems, and the quality of papers in them is generally high, and the probability of being generated by AIGC is relatively low. Therefore, using papers in a literature database with a high impact factor as negative sample data to train the paper detection large language model can effectively improve the classification accuracy of the paper detection large language model.

[0065] Of course, the "literature database that meets the preset impact factor requirements" can also be others, such as a literature database with an impact factor greater than or equal to a preset impact factor threshold, which is not specifically limited in this application.

[0066] Optionally, in order to enable the paper detection large language model to learn more difference information between AI-generated papers and normal papers, the above training data can be used to perform more detailed hyperparameter tuning on the paper detection large language model. Among them, the hyperparameters to be tuned include, but are not limited to, the following hyperparameters: learning rate, batch size, regularization strength (such as ridge regression (L2) regularization, regularization (Dropout) rate, etc.), optimizer parameters (such as beta value).

[0067] Optionally, when performing hyperparameter tuning, a preset hyperparameter tuning algorithm can be used, such as grid search, random search, or Bayesian optimization algorithm.

[0068] To prevent the paper detection large language model from overfitting during training, optionally, the early stopping method can be used to monitor the performance of the validation set in the training data.

[0069] Optionally, K-fold cross-validation can be used in the training and tuning stages of the paper detection large language model to obtain more robust performance estimates and model selection, and the generalization ability and robustness of the model can also be tested on training data from different sources, different fields, and even with minor adversarial perturbations.

[0070] See Figure 2, which is a schematic structural diagram of a large language model for paper detection provided by this application. As Figure 2 The large language model for paper detection consists of a first feature channel 21, a second feature channel 22, a third feature channel 23 and a feature decoding module 24. The first feature channel 21, the second feature channel 22 and the third feature channel 23 are respectively connected to the feature decoding module 24.

[0071] Considering that there are obvious differences in the deep semantic expressions of words and sentences between papers generated by artificial intelligence and normal papers written by humans. For example, compared with papers generated by artificial intelligence, normal papers are superior in multiple dimensions such as semantic richness, semantic logic, semantic emotion, semantic professionalism, semantic innovation, and semantic readability. Based on this, in the large language model for paper detection in this application embodiment, a first feature channel 21 is designed to perform word and sentence level semantic modeling processing on the paper to be processed through the first feature channel, so as to obtain the deep semantic features of the paper to be processed.

[0072] Here, the deep semantic features represent the word-level and sentence-level semantic information, part-of-speech semantic information, and inter-sentence semantic consistency information of the paper to be processed. Here, the inter-sentence semantic consistency information refers to the coherence and unity maintained in logic, context, and information transmission between multiple sentences or paragraphs, avoiding contradictions, concept jumps, or information conflicts.

[0073] Among them, the above-mentioned word set and sentence-level semantic information include one or more of the semantic information in the following dimensions: semantic richness, semantic logic, semantic emotion, semantic professionalism, semantic innovation, and semantic readability.

[0074] Step S102: Use the second feature channel in the large language model for paper detection to perform text feature statistical processing on the paper to be processed, so as to obtain the text statistical features of the paper to be processed.

[0075] After careful research on papers generated by artificial intelligence and normal papers written by humans, it is found that there are obvious differences in the shallow language style expressions of words and sentences between papers generated by artificial intelligence and normal papers. For example, in the use of words, sentence structures, and logical connections in various parts of papers generated by artificial intelligence, there are some recognizable templatized features (such as template sentences and template words), resulting in the overall paper being relatively rigid and mechanical.

[0076] For example, the common template words in papers generated by artificial intelligence are listed as follows.

[0077] First, the introduction section. Template words indicating development include: with the development of..., in recent years, with the progress of technology, currently, widely used in recent years; template words indicating introducing a topic include: attracting wide attention, becoming a research hotspot, receiving extensive attention, being highly regarded; template words indicating introducing challenges include: facing numerous challenges, having certain problems, not yet effectively solved, difficult to meet the requirements; template words emphasizing necessity include: it is necessary, urgently needed, urgently required, worthy of in-depth research; template words indicating research value include: theoretical significance, practical significance, promoting the development of the field, enriching the theoretical system.

[0078] Second, the related work section. Template words indicating summarizing research include: research shows, a large number of studies, existing research, existing methods; template words indicating method examples include: proposed a kind of, adopted the method of..., used, method based on...; template words indicating comparison differences include: different from, compared with, the difference lies in, compared with..., this study is more...; template words indicating pointing out deficiencies include: there are deficiencies, there are limitations, not considered, ignored, lacking the modeling of....

[0079] Third, the method section. Template words indicating method description include: this paper proposes, we design, constructs a, introduces, proposes a kind of; template words indicating structural expression include: the framework includes, mainly contains, the module includes, consists of..., includes the following parts; template words indicating process include: first, then, then, finally, ultimately, the overall process is shown in the figure.

[0080] Fourth, the experiment section. Template words indicating data introduction include: on the... dataset, the experiment adopts, the data contains, selects, the experimental settings are as follows; template words indicating result evaluation include: surpassing existing methods, having better performance, significantly improved, performing superiorly, verifying the effectiveness of...; template words indicating visual analysis include: it can be seen from the figure, as shown in the figure, indicates that, by comparison, the trends are consistent.

[0081] Fifth, the conclusion section. Template words indicating summary and restatement include: this paper proposes, this paper studies, the experimental results verify, the method is effective, achieving good results; template words indicating looking ahead to the future include: subsequent research, future work, further expansion, considering more complex situations, applying to other tasks; template words indicating limitation description include: there are still certain limitations, the method depends on, not considered, need to be further improved.

[0082] Sixth, the logical connection part. The template words indicating parallelism are: not only that, in addition, at the same time, moreover, and; the template words indicating a turn are: however, but, nevertheless, it is worth noting that; the template words indicating cause and effect are: therefore, thus, resulting in, due to, the result is; the template words indicating order or time are: first, then, second, finally, currently, in recent years; the template words indicating summary or induction are: in summary, it can be seen from this, generally speaking, it can be seen that, generally speaking.

[0083] Common template sentences in papers generated by artificial intelligence are listed as follows.

[0084] First, the introduction part. The template sentences indicating background statements are: "With the rapid development of artificial intelligence, the problem of XXX has received extensive attention", "XXX has become one of the hot issues in current research"; the template sentences indicating problem introduction are: "Despite a large number of studies, the problem of XXX has not been fully solved", "Existing studies mostly focus on XXX, however, there are still deficiencies in XXX"; the template sentences indicating research purposes are: "This paper aims to propose a new method to solve the problem of XXX", "To solve the above problems, this paper proposes a new method based on XXX"; the template sentences indicating research significance are: "This research helps to further promote the development of the field of XXX", "The research results have important value for understanding XXX".

[0085] Second, the related work part. The template sentences indicating overall generalization are: "A large number of literatures have conducted in-depth research on XXX"; the template sentences indicating method introduction are: "XXX et al. improved the performance of XXX by introducing the XXX mechanism"; the template sentences indicating comparison with others' methods are: "Different from the above methods, the method in this paper has obvious advantages in XXX"; the template sentences indicating discovering deficiencies are: "There are still certain limitations in existing research in XXX", "Most of the above studies are based on the assumption of XXX and have limitations in practical applications".

[0086] Third, the method part. The template sentences indicating model description are: "We designed a deep learning framework composed of XXX", "The model proposed in this paper mainly consists of three parts: XXX, XXX and XXX"; the template sentences indicating step description are: "First,... Subsequently,... Finally,..."; the template sentences indicating parameter definition are: "To simplify the representation, we use the following symbols to define...".

[0087] Fourth, the experimental part. The template sentences for expressing experimental settings are: "We evaluated the proposed method on the XXX dataset"; the template sentences for expressing performance descriptions are: "The results show that the method outperforms the existing methods in terms of XXX metrics"; the template sentences for expressing visualization descriptions are: "Figure X shows the comparison of different methods on XXX" and "Table X shows the comparison results of different methods on the XXX task."

[0088] Fifth, the conclusion part. The template sentences for summarizing contributions are: "This paper proposes an effective XXX method to solve the XXX problem"; the template sentences for looking ahead to the future are: "Future work can further explore the XXX direction" and "We plan to extend this method to more complex scenarios"; the template sentences for emphasizing advantages are: "The experimental results fully verify the effectiveness and robustness of the proposed method" and "This study provides a new solution idea for the XXX problem."

[0089] It should be noted that the above template words and template sentences are only examples and are not intended to limit this application.

[0090] Using the above template words and template sentences extensively makes the overall paper generated by artificial intelligence rather mechanical and lacks flexibility. In addition, the paper generated by artificial intelligence is not a paper based on real research, and the quantitative vocabulary it uses is less. On the contrary, when writing a normal paper, humans tend to pay attention to using flexible and diverse vocabulary and sentence patterns, and their language style expressions will be more rich and diverse, and a large amount of research result data will be presented in a normal paper.

[0091] Based on the above differences, in the large language model for paper detection in this application embodiment, a second feature channel 22 is designed to perform text feature statistical processing on the paper to be processed through the second feature channel, so as to obtain the text statistical features of the paper to be processed. Here, the text statistical features characterize the language style characteristics of the paper to be processed.

[0092] Step S103: Use the third feature channel in the large language model for paper detection to perform quality evaluation processing on the paper to be processed, so as to obtain the quality evaluation features of the paper to be processed.

[0093] It is understandable that compared with papers generated by artificial intelligence, normal papers are obtained through in-depth thinking and research. Therefore, the overall quality of normal papers is usually significantly higher than that of papers generated by artificial intelligence. For example, in terms of text readability, normal papers are written after in-depth human thinking and are carefully considered in terms of word and sentence usage, sentence connection, etc. Therefore, compared with papers generated by artificial intelligence, normal papers have higher text readability. Another example is in terms of text diversity. When humans write normal papers and encounter words and sentences with the same or similar meanings, they tend to change the expression methods. Therefore, compared with papers generated by artificial intelligence, normal papers have better text diversity.

[0094] Based on this, in the embodiment of the present application, a third feature channel 23 is designed in the above-mentioned paper detection large language model, so as to perform quality evaluation processing on the paper to be processed through the third feature channel and obtain the quality evaluation features of the paper to be processed. Here, the quality evaluation features represent the overall quality characteristics of the paper to be processed.

[0095] Step S104: Use the feature decoding module in the paper detection large language model to obtain the target detection result indicating whether the paper to be processed is a paper generated by artificial intelligence according to the deep semantic features, text statistical features, and quality evaluation features.

[0096] In this embodiment, the feature decoding module 24 in the paper detection large language model can perform feature decoding according to the deep semantic features, text statistical features, and quality evaluation features, and can accurately classify the paper to be processed according to the decoding result to obtain the target detection result indicating whether the paper to be processed is a paper generated by artificial intelligence.

[0097] The detection method for AI-generated papers provided in this application takes into account the obvious differences between AI-generated papers and normal papers written by humans in terms of both the deep semantic expression and the shallow language style expression of vocabulary and sentences. For example, compared with AI-generated papers, normal papers have a higher semantic richness and a more flexible and changeable language style, and are not restricted to certain fixed template words, template sentences, etc. Based on this, this application uses the first feature channel in the pre-trained large language model for paper detection to perform word and sentence level semantic modeling on the paper to be processed, obtaining the deep semantic features of the paper to be processed, and uses the second feature channel in the large language model for paper detection to perform text feature statistics on the paper to be processed, obtaining the text statistical features of the paper to be processed. Further, compared with AI-generated papers, normal papers are the result of in-depth thinking and research, and their overall quality is often significantly higher than that of AI-generated papers. Therefore, this application also performs quality assessment on the paper to be processed through the third feature channel of the large language model for paper detection, obtaining the quality assessment features of the paper to be processed. Since this application can extract the feature information of the paper to be processed from three perspectives: deep, shallow, and macroscopic through the three feature channels of the large language model for paper detection, the feature decoding module in the large language model for paper detection can accurately classify the paper to be processed based on the deep semantic features, text statistical features, and quality assessment features, thereby improving the classification accuracy of the target detection results obtained by the feature decoding module.

[0098] In some embodiments of this application, considering that the structure of the text to be processed may be complex and changeable, it is difficult to extract features or the extracted features are inaccurate and incomplete in the above steps S101 to S103.

[0099] To improve the accuracy and comprehensiveness of the features, the embodiments of this application can, before the steps S101 to S103 of "using the first feature channel in the pre-trained large language model for paper detection to perform word and sentence level semantic modeling on the paper to be processed, obtaining the deep semantic features of the paper to be processed, using the second feature channel in the large language model for paper detection to perform text feature statistics on the paper to be processed, obtaining the text statistical features of the paper to be processed, and using the third feature channel in the large language model for paper detection to perform quality assessment on the paper to be processed, obtaining the quality assessment features of the paper to be processed", first perform structured parsing on the paper to be processed to obtain the structured paper corresponding to the paper to be processed, and then use the structured paper as the paper to be processed in steps S101 to S103.

[0100] Optionally, the structured paper can be a multi-level structure of "title - abstract - chapter title - section title - body content". For example, the structured parsing method of the portable document format file provided in CN117473980A can be used to perform structured parsing on the paper to be processed.

[0101] The above structural analysis method is only an example and does not limit this application.

[0102] In this embodiment, by performing structural analysis on the paper to be processed, the subsequent process can more accurately extract the context features of the paper to be processed, thereby improving the classification accuracy of the paper to be processed.

[0103] To make those skilled in the art better understand this application, the process of using the paper detection large language model to detect papers in this application will be introduced in detail through the following embodiments.

[0104] See Figure 3 As shown, it is a schematic structural diagram of another paper detection large language model provided by this application. As Figure 3 shown, the first feature channel 21 includes a text input module 211, a feature encoding module 212, an inter-sentence semantic consistency modeling module 213, and a part-of-speech semantic distribution modeling module 214. The text input module 211 is connected to the feature encoding module 212, and the feature encoding module 212 is also connected to the inter-sentence semantic consistency modeling module 213 and the part-of-speech semantic distribution modeling module 214.

[0105] Here, the text input module 211 is used to perform sentence splitting, word segmentation, and part-of-speech tagging on the paper to be processed, and obtain the word segmentation sequence and word segmentation part-of-speech tagging sequence corresponding to each sentence in the paper to be processed.

[0106] Optionally, after receiving the paper to be processed, the text input module 211 can first determine whether the paper to be processed is a Chinese paper or a foreign language paper. If it is a Chinese paper, it uses regular expressions combined with Chinese punctuation to split the paper to be processed into sentences. If it is a foreign language paper, it uses the spaCy sentence splitter to split the paper to be processed into sentences; further, it uses the BERT (Bidirectional Encoder Representations from Transformers) tokenizer and the WordPiece tokenization algorithm to tokenize each sentence in the paper to be processed, and obtain the word segmentation sequence corresponding to each sentence.

[0107] In addition, the text input module 211 can also perform part-of-speech tagging on each token in the token sequence corresponding to each sentence, and obtain the token part-of-speech tagging sequence corresponding to each sentence.

[0108] For example, a clause in the paper to be processed is "Xiaoming likes reading books in the library", then the corresponding word segmentation sequence for this clause is ["Xiaoming", "likes", "in", "the library", "reading books"], and the corresponding part-of-speech tagging sequence for the word segmentation of this clause is ["PER", "v", "p", "LOC", "v"], where "PER" represents a person's name, "v" represents a verb, "p" represents a preposition, and "LOC" represents a place name.

[0109] Optionally, for Chinese papers, the text input module 211 can call HanLP for part-of-speech tagging. For corresponding foreign papers, it can call spaCy for part-of-speech tagging.

[0110] The feature encoding module 212 is used to perform word-level context encoding on each word in the word segmentation sequence corresponding to each clause, obtaining a word-level encoded vector sequence corresponding to each clause. According to the word-level encoded vector sequence corresponding to each clause, weighted aggregation processing of the word segmentation within the sentence is performed to obtain a sentence vector corresponding to each clause. Sentence-level context encoding is performed on the sentence vectors corresponding to all clauses in the paper to be processed, obtaining sentence-level encoded vectors corresponding to all clauses respectively. Weighted aggregation processing between sentences is performed on the sentence-level encoded vectors corresponding to all clauses respectively, obtaining a document vector representing the overall semantic information of the paper to be processed.

[0111] For example, assuming the clause results of the paper to be processed are [clause 1, clause 2,..., clause N], then the word segmentation results are: clause 1 = [w11, w12,..., w1T], clause 2 = [w21, w22,..., w2T], and so on, where N and T are positive integers.

[0112] Then, the feature encoding module 212 performs word-level context encoding to obtain: the word-level encoded vector sequence [h11, h12,..., h1T] corresponding to clause 1, the word-level encoded vector sequence [h21, h22,..., h2T] corresponding to clause 2, and so on, where h11 is obtained by encoding w11 in combination with the following word segmentation of w11 (e.g., w12), h12 is obtained by encoding w12 in combination with the context word segmentation of w12 (e.g., w11 and w13), and so on.

[0113] The feature encoding module 212 also calculates the word-level attention weights of each word in each word segmentation sequence, and performs weighted aggregation processing of the word-level encoded vectors in the corresponding word-level encoded vector sequence according to the calculated word-level attention weights to obtain a sentence vector corresponding to each clause. For example, the sentence vector S1 obtained for the word-level encoded vector sequence [h11, h12,..., h1T] is: , where, The word-level attention weights representing h1j, where h1j represents the j-th word-level encoding vector corresponding to the first clause, that is, the word-level encoding vector corresponding to the j-th word segment in the first word segmentation. j is a positive integer between 1 and T. Similarly, the sentence vectors S2 to SN corresponding to other clauses can be obtained.

[0114] Next, the feature encoding module 212 performs sentence-level context encoding to obtain: the sentence-level encoding vector hS1 corresponding to the first clause, the sentence-level encoding vector hS2 corresponding to the second clause, and so on. Among them, hS1 is obtained by encoding S1 by combining the sentence vector (such as S2) corresponding to the subsequent clause of S1, and hS2 is obtained by encoding S2 by combining the sentence vectors (such as S1 and S3) corresponding to the context clauses of S2, and so on.

[0115] Finally, the feature encoding module 212 calculates the sentence-level attention weights of each clause in the paper to be processed, and performs weighted aggregation processing between sentences on the sentence-level encoding vectors corresponding to each clause according to the calculated sentence-level attention weights to obtain the document vector. For example, the document vector SD obtained for the sentence-level encoding vectors hS1 to hSN is: , where represents the sentence-level attention weight of hSk, and hSk represents the k-th sentence-level encoding vector, where k is a positive integer between 1 and N.

[0116] Optionally, the above word-level encoding vector can be a 768-dimensional vector, and the sentence-level encoding vector can be a 256-dimensional vector.

[0117] Optionally, when calculating the word-level attention weights and sentence-level attention weights, the above feature encoding module 212 can be implemented based on an MLP (Multi-Layer Perceptron).

[0118] Optionally, the above feature encoding module 212 can be a BERT-HAN model. This BERT-HAN model can replace the static word vectors used in the traditional HAN (Hierarchical Attention Network) with context-sensitive dynamic word vectors generated by BERT. This improvement can effectively strengthen local dependence relationships, reduce semantic information loss, thereby enhancing the overall feature expression ability and the final classification and detection performance.

[0119] In this embodiment, the feature encoding module 212 obtains word-level encoded vectors through word-level context encoding. The word-level encoded vectors contain not only the semantic information of the current word segmentation, but also the semantic information of the context word segmentations, improving the semantic accuracy of word segmentation encoding. Similarly, through sentence-level context encoding, the obtained sentence-level encoded vectors contain not only the semantic information of the current sub-sentence, but also the semantic information of the context sub-sentences, improving the semantic accuracy of sub-sentence encoding, and thus improving the semantic accuracy of the document vector.

[0120] The inter-sentence semantic consistency modeling module 213 is used to construct an inter-sentence similarity matrix based on the sentence-level encoded vectors corresponding to all sub-sentences in the paper to be processed, and perform index feature extraction on the inter-sentence similarity matrix based on a preset statistical index to obtain a matrix vector representing the internal coherence and semantic consistency of the paper to be processed.

[0121] Specifically, the inter-sentence semantic consistency modeling module 213 can calculate the similarity of the sentence-level encoded vectors corresponding to each pair of sub-sentences to construct an inter-sentence similarity matrix. Optionally, the cosine similarity algorithm can be used to calculate the similarity of the sentence-level encoded vectors. Of course, other similarity algorithms can also be used, and the present application does not make specific limitations.

[0122] In this embodiment, multiple statistical indexes are preset to perform index feature extraction on the inter-sentence similarity matrix to obtain a matrix vector. For example, optionally, the statistical index can be the mean, variance, standard deviation, maximum value, density, range, quantile, norm, rank, trace, skewness, entropy, or singular value. Of course, the statistical index can also be other, and the present application does not make specific limitations.

[0123] In a possible implementation, 10 statistical indexes can be used to extract the index features of the inter-sentence similarity matrix from 10 aspects, and the 10 obtained index features can be concatenated into a vector to obtain a 10-dimensional matrix vector.

[0124] The part-of-speech semantic distribution modeling module 214 is used to perform statistical analysis processing based on the occurrence frequency on all parts of speech in the paper to be processed according to the part-of-speech tagging sequences and word-level encoded vector sequences corresponding to all sub-sentences in the paper to be processed, and obtain a part-of-speech vector representing the part-of-speech semantic information of the paper to be processed.

[0125] Optionally, the process of "performing statistical analysis processing based on the occurrence frequency on all parts of speech in the paper to be processed according to the part-of-speech tagging sequence and word-level encoding vector sequence corresponding to each clause in the paper to be processed, and obtaining a part-of-speech vector representing the part-of-speech semantic information of the paper to be processed" may include: for each part of speech among all parts of speech, statistically analyzing the occurrence frequency of each token under this part of speech according to the word-level encoding vector sequence corresponding to each clause, obtaining the occurrence frequency of each token under this part of speech, performing average pooling processing on the occurrence frequencies of all tokens under this part of speech to obtain the occurrence frequency of this part of speech, so as to obtain the occurrence frequencies of all parts of speech respectively, and concatenating the occurrence frequencies of all parts of speech respectively to obtain a part-of-speech vector.

[0126] Optionally, the calculation formula for the occurrence frequency of a token is: the number of occurrences of the token / the total number of tokens in the paper to be processed × 100%.

[0127] For example, in the paper to be processed, the token "Xiaoming" appears 2 times, "Xiaohong" appears 4 times, and the total number of tokens in the paper to be processed is 20. Then the occurrence frequency of "Xiaoming" is 10%, the occurrence frequency of "Xiaohong" is 20%, and the occurrence frequency of the PER part of speech is (20% + 10%) / 2 = 15%.

[0128] Assume that all parts of speech include nouns, verbs, adjectives, adverbs, prepositions, numeral-classifiers, conjunctions, and auxiliary words, and the occurrence frequencies are 10%, 14%, 0%, 0%, 6%, 15%, 3%, and 0% respectively. Then the part-of-speech vector is [10%, 14%, 0%, 0%, 6%, 15%, 3%, 0%].

[0129] It should be noted that when concatenating the occurrence frequencies of all parts of speech respectively, it is necessary to concatenate them in a fixed part-of-speech order to avoid affecting the final classification result due to inconsistent concatenation order.

[0130] In summary, the deep semantic features in this embodiment include a document vector, a matrix vector, and a part-of-speech vector.

[0131] Still referring to Figure 3 , this embodiment will also use the second feature channel 22 in the paper detection large language model to perform text feature statistical processing on the paper to be processed, and obtain the text statistical features of the paper to be processed. In one possible implementation, this process may include: performing text feature statistical processing on the paper to be processed at the lexical level and the sentence level respectively through the second feature channel, obtaining the lexical-level statistical features and sentence-level statistical features of the paper to be processed, and using the lexical-level statistical features and sentence-level statistical features as text statistical features.

[0132] After in-depth research, it is found that at the lexical level, there are the following differences between AI-generated papers and normal papers:

[0133] First, in terms of lexical richness, normal papers are the result of careful human thinking and polishing. To improve the quality of the papers, a variety of professional, advanced, and colloquial words are often used to package the papers. However, the training data of AI-generated papers mostly come from databases such as online texts and academic literature, resulting in a relatively single language style of AI-generated papers. For example, they usually mainly use formal written language and tend to repeat high-frequency words, with limited lexical diversity. Based on this, in terms of lexical richness, normal papers are often significantly higher than AI-generated papers.

[0134] Second, in terms of the number of unique words (unique words include domain terms, low-frequency words, personalized expression words, etc.), human scholars, through long-term learning and research, can master a large number of domain terms and cutting-edge concepts and can choose the most appropriate words according to the context during the writing process. In addition, there may be innovative expressions during the writing process. However, AI-generated papers tend to repeat high-frequency words due to the lack of creative thinking, resulting in a significantly higher number of unique words in normal papers than in AI-generated papers.

[0135] Third, in terms of the use of quantitative words, humans tend to display a large amount of their own research result data during the process of writing papers, while AI-generated papers are often content generated by referring to existing papers, so the number of their quantitative words is extremely small. Optionally, a regular detection method can be used to detect quantitative words.

[0136] Fourth, in terms of the use of stop words, AI-generated papers often frequently use stop words such as "of", "is", "in" to fill the sentence structure, resulting in redundant papers; while the use frequency of stop words in normal papers is moderate and stop words are often not used when not necessary. For example, a possible expression in an AI-generated paper is "Although this experiment is a failure", while a normal paper usually expresses it as "Although the experiment failed".

[0137] Fifth, in terms of the use of conjunctions, AI-generated papers often rely too much on conjunctions such as "firstly", "secondly", "finally", "although...but..." to show the logic of the papers; while normal papers present the logical relationship naturally through sentence structures. For example, a possible expression in an AI-generated paper is "Although X increases, but Y decreases", while a normal paper usually expresses it as "The research found that when X increases, Y instead decreases".

[0138] Fifth, in the use of high-frequency words, AI-generated papers will overuse high-frequency terms or keywords in the training data, resulting in abnormal frequency of certain words and a relatively mechanical distribution of high-frequency words. For example, the repeated use of transition words such as "in addition" and "however" that have no substantive semantics. The high-frequency words in normal papers are often closely related to the core arguments, and scholars will flexibly adjust the use of vocabulary according to the needs of the argument to avoid excessive repetition of a single word.

[0139] Fifth, in the use of template words, almost all AI-generated papers will use the template words listed above. However, normal papers do not necessarily use template words. Even if they use template words, the template words can be appropriately integrated into the context without being abrupt.

[0140] Based on the above distinction, optionally, the lexical level includes one or more of the following levels: the lexical level includes one or more of the following levels: lexical richness, the number of unique words, and the usage of quantitative words, stop words, conjunctions, target words and template words, where target words refer to words whose frequency of appearance in the papers to be processed is higher than a preset frequency threshold.

[0141] Similarly, at the vocabulary level, there are the following differences between AI-generated papers and normal papers:

[0142] First, in terms of sentence structure complexity, AI-generated papers are almost all combinations of simple sentences, while the ratio of simple and complex sentences in normal papers is moderate.

[0143] Second, in terms of the distribution of sentence length changes, the sentence length distribution of AI-generated papers presents a narrow peak, mostly concentrated in the 15-25 word range, lacking alternation between long and short sentences; while the sentence length distribution of normal papers presents a bimodal distribution, with both short sentences of 5-10 words to enhance the sense of rhythm and long sentences of more than 30 words to develop complex arguments, showing an overall effect of natural fluctuations in length. For example, the distribution of sentence length changes can be evaluated by the average sentence length, the ratio of the maximum sentence length to the minimum sentence length, and the sentence length variance.

[0144] Third, in terms of the proportion of synonyms used, AI-generated papers will repeatedly use the same core vocabulary, while normal papers will focus on changing synonyms, so AI-generated papers use a higher proportion of synonyms. Optionally, an algorithm based on word vector neighbors (such as the K-nearest neighbor algorithm KNN) can be used to determine the proportion of synonyms used in an article.

[0145] Fourth, in the use of template sentences, similar to template words, almost all AI-generated papers will use the template sentences listed above. However, normal papers will appropriately integrate the template sentences into the context without being abrupt.

[0146] Based on this, optionally, the sentence level includes one or more of the following levels: sentence structure complexity, sentence length variation distribution, proportion of synonymous expressions used, and usage of template sentences.

[0147] It should be noted that the above content at the lexical level and sentence level is only an example. In addition, it can be other things, such as paragraph structure, etc. The present application does not make specific limitations.

[0148] In a possible implementation, in order to better evaluate the characteristics of the paper to be processed in terms of shallow linguistic expressions, this embodiment can use multiple statistical features for evaluation. For example, a total of 40 features are selected from the lexical layer statistical features and sentence layer statistical features, and then they are concatenated into a vector form as the text statistical features.

[0149] Since there are many features, it may consume a large amount of storage resources and computing resources. In order to save resources and reduce the data volume, optionally, after obtaining the lexical layer statistical features and sentence layer statistical features in the foregoing, before concatenating them into a vector form, this embodiment can use the Z-score (standard score) algorithm to scale and process the lexical layer statistical features and sentence layer statistical features.

[0150] This embodiment can accurately capture the text usage rules of the paper to be processed at the shallow layer through the second feature channel. By using the differences in this text usage rule between the paper generated by artificial intelligence and the normal paper, it can accurately determine whether the paper to be processed is a paper generated by artificial intelligence. And the second feature channel can not rely on a deep learning model and only needs to perform simple statistical calculations, with higher efficiency.

[0151] Still referring to Figure 3 , this embodiment will also use the third feature channel 23 in the large language model for paper detection to perform quality evaluation processing on the paper to be processed, and obtain the quality evaluation features of the paper to be processed. Optionally, this process may include: using a preset quality evaluation index to perform text quality evaluation on the paper to be processed to obtain the quality evaluation features of the paper to be processed, where the quality evaluation index includes one or more of the following indexes: text readability index, text diversity index, and text predictability index.

[0152] That is to say, this embodiment can use the text readability index to perform text readability evaluation on the paper to be processed to obtain the text readability evaluation features of the paper to be processed; use the text diversity index to perform text diversity evaluation on the paper to be processed to obtain the text diversity evaluation features of the paper to be processed; use the text predictability index to perform text predictability evaluation on the paper to be processed to obtain the text predictability evaluation features of the paper to be processed; and obtain the quality evaluation features according to the text readability evaluation features, text diversity evaluation features, and text predictability evaluation features.

[0153] Text readability metrics are used to measure the comprehension difficulty and reading fluency of text. Optionally, the text readability metrics can be the Flesch Reading Ease (FRE) metric and / or the Gunning Fog Index (GFI) metric. Here, the FRE metric is a method for scoring text readability, with scores ranging from 0 to 100, where a higher score indicates that the article is easier to read; the GFI estimates the reading difficulty of text by examining the complexity of sentences and the difficulty of vocabulary.

[0154] Optionally, the text diversity metric can be the Self-BLEU metric, which is used to evaluate the semantic overlap degree between internal segments of the text.

[0155] The text predictability metric is used to reflect the modeling difficulty of the text for the language model. Optionally, the text predictability metric can be the text Perplexity metric and / or the inverse exponential operation metric of Log-Likelihood. Here, the inverse exponential operation of Log-Likelihood performs an exponential operation on the model's Log-Likelihood to restore it to the probability scale.

[0156] Of course, the text readability metric, the text diversity metric, and the text predictability metric can also be others, which are not specifically limited in this application.

[0157] Optionally, in this embodiment, the text readability evaluation feature, the text diversity evaluation feature, and the text predictability evaluation feature can be concatenated into a vector form as the quality evaluation feature.

[0158] This embodiment mines the deep semantic relationships of the paper to be processed through the first feature channel, captures the usage rules on the language surface through the second feature channel, and assists in modeling the generation characteristics of the paper to be processed through the third feature channel, realizing the complementarity of multi-dimensional and multi-granularity information.

[0159] See Figure 3 , optionally, the feature decoding module 24 may include: a feature fusion module 241 and a classification output module 242. Then, the process of "using the feature decoding module in the large language model for paper detection to obtain the target detection result indicating whether the paper to be processed is an AI-generated paper based on the deep semantic features, text statistical features, and quality evaluation features" may include: using the feature fusion module 241 to perform feature fusion based on attention weights on the deep semantic features, text statistical features, and quality evaluation features to obtain the fused features; using the classification output module 242 to obtain the target detection result based on the fused features.

[0160] As introduced above, the deep semantic features may include: document vectors, matrix vectors, and part-of-speech vectors. In a possible implementation, the feature fusion module 241 may first fuse the document vectors, matrix vectors, and part-of-speech vectors to obtain deep semantic features, and then perform attention-weighted feature fusion on the deep semantic features, text statistical features, and quality assessment features to obtain fused features.

[0161] In a possible implementation, the feature fusion module 241 may directly perform attention-weighted feature fusion on the document vectors, matrix vectors, part-of-speech vectors, text statistical features, and quality assessment features to obtain fused features.

[0162] However, no matter which of the above fusion methods is used, it is possible that the features to be fused (such as deep semantic features, text statistical features, and quality assessment features) are in different dimensions, which affects the accuracy of fusion. For this reason, in this embodiment, before fusion, each feature to be fused may be mapped to the same dimension. For example, a linear layer may be used to uniformly map the features to be fused in different dimensions to 256 dimensions.

[0163] Optionally, a small attention mechanism network may be used to dynamically adjust the contribution of each channel feature to the final representation by learning the importance weights of different channel features, obtain the importance weights of each channel, and then perform attention-weighted feature fusion based on the determined importance weights to obtain fused features.

[0164] It should be noted that the process of "performing attention-weighted feature fusion on deep semantic features, text statistical features, and quality assessment features" described above is only an example. In addition, there may be other fusion methods, such as vector concatenation, which are not specifically limited in this application.

[0165] Optionally, the classification output module 242 may be a feature classification model based on the XGBoost (eXtreme Gradient Boosting) classification algorithm. Then, the process of "obtaining the target detection result according to the fused features" may include: inputting the fused features into the feature classification model to obtain the target detection result output by the feature classification model.

[0166] In summary, this embodiment provides a large language model for paper detection based on the fusion of semantic-text-index features. In the semantic part, a feature encoding module (such as the BERT-HAN model) is introduced for multi-level semantic modeling, which can deeply enhance the semantic relationship between sentences and words of the paper to be processed. At the same time, an inter-sentence semantic consistency modeling module and a part-of-speech semantic distribution modeling module are introduced to further extract the semantics of the intermediate encoding results (word-level encoding vectors and sentence-level encoding vectors) of the feature encoding module, thereby enhancing the overall semantic feature expression ability of the paper to be processed; in the text part, by statistically analyzing the text features at the lexical and sentence levels, the significant difference features between AI-generated papers and normal papers at the shallow text feature level can be effectively extracted, and the obtained text statistical features can be used as an important basis for identifying AI-generated papers; in the index part, combined with existing multi-dimensional indexes such as readability, perplexity, and text diversity in the market, the subtle differences between AI-generated papers and normal papers written by humans in terms of naturalness, complexity, and fluency are further revealed, so as to effectively assist the paper discrimination task. Through the complementary advantages of the above three aspects, the difference features between AI-generated papers and normal papers can be more comprehensively captured, and the final classification accuracy is improved.

[0167] The process of model application is introduced above. Next, the process of model training will be introduced.

[0168] Step 1: Data collection

[0169] To prevent the large language model for paper detection from directly comparing the text similarity between the paper to be processed and the papers in a mature literature database (such as PubMed), thus falling into the misunderstanding of using the theme as the judgment basis, a labeled dataset is first constructed, including the following two types of sample data:

[0170] Positive sample data: The full text of the papers clearly marked as being generated by AIGC in the retraction watch database;

[0171] Negative sample data: The full text of normal papers with similar themes to the positive sample data. Specifically, papers with the same or similar themes are retrieved through literature databases such as PubMed that meet the preset impact factor requirements.

[0172] The above collected labeled dataset can be divided into a training set, a validation set, and a test set according to a preset ratio (such as 7∶1.5∶1.5).

[0173] Step 2: Paper structure parsing

[0174] Structural parsing is performed on each paper in the training set, validation set, and test set. For example, the structural parsing method of the portable document format file provided by CN117473980A is used to perform structural parsing on the paper to be processed.

[0175] Step 3: Data Annotation

[0176] Label the positive and negative sample data in the training set, validation set, and test set as 1 and 0 respectively for subsequent text classification research.

[0177] Step 4: Model Construction

[0178] For example, construct a model as shown in Figure 2 as the large language model for paper detection.

[0179] Step 5: Supervised Tuning

[0180] Use the training set and validation set labeled in Step 3 to perform supervised tuning on the constructed large language model for paper detection, so that the model performs better in the classification task.

[0181] The specific optimization method is to use categorical cross-entropy as the loss function and the AdamW optimizer for gradient update; use the labeled training set and validation set to perform hyperparameter tuning with a preset hyperparameter tuning algorithm (such as grid search, random search, or Bayesian optimization algorithm).

[0182] Step 6: Model Validation and Testing

[0183] Use the validation set labeled in Step 3 to validate the pre-trained large language model for paper detection, and use the test set labeled in Step 3 to test the pre-trained large language model for paper detection.

[0184] Optionally, the evaluation metrics used in the validation and testing phases can include: Accuracy, Precision, Recall, and F1-score.

[0185] For example, in a possible implementation, when the accuracy, precision, and recall all exceed 85%, and the F1-score reaches above 0.85, it is considered that the pre-trained large language model for paper detection has strong discriminative ability and can be applied to actual classification.

[0186] If the evaluation metrics do not meet the requirements, for example, at least one of the accuracy, precision, and recall does not exceed 85%, or the F1-score does not reach above 0.85, error analysis can be performed, for example, checking the samples misclassified by the large language model for paper detection, and further optimizing the model architecture or data cleaning.

[0187] Optionally, the processes in Step 5 and Step 6 above can use K-fold cross-validation to obtain more robust performance estimates and model selection, and early stopping can be used to prevent overfitting.

[0188] Optionally, in this embodiment, the generalization ability and robustness of the model can be tested on test sets from different sources, different fields, and even with slightly adversarial perturbations.

[0189] Next, a detection device for AI-generated papers provided by an embodiment of the present application will be described. The detection device for AI-generated papers described below can be correspondingly referred to the detection method for AI-generated papers described above.

[0190] See Figure 4 , Figure 4 which is a schematic structural diagram of a detection device for AI-generated papers provided by an embodiment of the present application.

[0191] As Figure 4 shown, the device may include:

[0192] A deep feature extraction module 401, configured to perform multi-level semantic modeling processing on a to-be-processed paper using a first feature channel in a pre-trained large language model for paper detection, to obtain deep semantic features of the to-be-processed paper, where the deep semantic features represent word-level and sentence-level semantic information of the to-be-processed paper;

[0193] A shallow feature extraction module 402, configured to perform text feature statistics processing on the to-be-processed paper using a second feature channel in the large language model for paper detection, to obtain text statistical features of the to-be-processed paper, where the text statistical features represent language style information of the to-be-processed paper;

[0194] A macro feature extraction module 403, configured to perform quality assessment processing on the to-be-processed paper using a third feature channel in the large language model for paper detection, to obtain quality assessment features of the to-be-processed paper, where the quality assessment features represent overall quality characteristics of the to-be-processed paper;

[0195] A feature fusion and decoding module 404, configured to use a feature decoding module in the large language model for paper detection to obtain a target detection result indicating whether the to-be-processed paper is an AI-generated paper according to the deep semantic features, the text statistical features, and the quality assessment features;

[0196] Wherein, the training data used in the training stage of the large language model for paper detection includes: positive sample data labeled with AI-generated paper labels and negative sample data labeled with normal paper labels. The positive sample data is papers marked as AI-generated content in a retraction watch database, and the negative sample data is papers in a literature database that meet preset impact factor requirements. The topic similarity between the positive sample data and the negative sample data meets a preset similarity requirement.

[0197] In one possible implementation, the detection device for AI-generated papers provided by the embodiments of the present application may further include: a structure analysis module.

[0198] The structure analysis module is configured to perform sentence-level semantic modeling processing on the paper to be processed by using the first feature channel in the pre-trained large language model for paper detection to obtain the deep semantic features of the paper to be processed, perform text feature statistical processing on the paper to be processed by using the second feature channel in the large language model for paper detection to obtain the text statistical features of the paper to be processed, and perform quality evaluation processing on the paper to be processed by using the third feature channel in the large language model for paper detection to obtain the quality evaluation features of the paper to be processed. Before that, perform structural parsing on the paper to be processed to obtain a structured paper corresponding to the paper to be processed, and use the structured paper as the paper to be processed.

[0199] In one possible implementation, the above-mentioned first feature channel includes a text input module, a feature encoding module, an inter-sentence semantic consistency modeling module, and a part-of-speech semantic distribution modeling module.

[0200] Based on this, when the above-mentioned deep feature extraction module performs sentence-level semantic modeling processing on the paper to be processed by using the first feature channel in the pre-trained large language model for paper detection to obtain the deep semantic features of the paper to be processed, it may specifically be used for:

[0201] Perform clause segmentation, word segmentation, and part-of-speech tagging processing on the paper to be processed through the text input module to obtain a word segmentation sequence and a word segmentation part-of-speech tagging sequence corresponding to each clause in the paper to be processed;

[0202] Perform word-level context encoding on each word in the word segmentation sequence corresponding to each clause through the feature encoding module to obtain a word-level encoding vector sequence corresponding to each clause, perform weighted aggregation processing on the intra-clause word segmentation according to the word-level encoding vector sequence corresponding to each clause to obtain a sentence vector corresponding to each clause, perform sentence-level context encoding on the sentence vectors corresponding to all clauses in the paper to be processed to obtain a sentence-level encoding vector corresponding to each of all clauses, and perform weighted aggregation processing on the sentence-level encoding vectors corresponding to each of all clauses to obtain a document vector representing the overall semantic information of the paper to be processed;

[0203] Construct an inter-sentence similarity matrix according to the sentence-level encoding vectors corresponding to all clauses in the paper to be processed through the inter-sentence semantic consistency modeling module, and perform index feature extraction on the inter-sentence similarity matrix based on a preset statistical index to obtain a matrix vector representing the internal coherence and semantic consistency of the paper to be processed;

[0204] The part-of-speech semantic distribution modeling module performs statistical analysis based on the occurrence frequency on all parts of speech in the to-be-processed paper according to the part-of-speech tagging sequences and word-level encoding vector sequences corresponding to all clauses in the to-be-processed paper, to obtain a part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper.

[0205] In a possible implementation, when the above-mentioned deep feature extraction module performs statistical analysis based on the occurrence frequency on all parts of speech in the to-be-processed paper according to the part-of-speech tagging sequences and word-level encoding vector sequences corresponding to all clauses in the to-be-processed paper, to obtain a part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper, it can specifically be used for:

[0206] The part-of-speech semantic distribution modeling module performs statistics on the occurrence frequency of each token under each part of speech according to the word-level encoding vector sequences corresponding to all clauses, to obtain the occurrence frequency of each token under each part of speech, performs average pooling on the occurrence frequencies of all tokens under each part of speech, to obtain the occurrence frequency of each part of speech, so as to obtain the occurrence frequencies of all parts of speech, and splices the occurrence frequencies of all parts of speech to obtain the part-of-speech vector.

[0207] In a possible implementation, when the above-mentioned shallow feature extraction module performs text feature statistics on the to-be-processed paper using the second feature channel in the paper detection large language model to obtain the text statistical features of the to-be-processed paper, it can specifically be used for:

[0208] Performing text feature statistics on the to-be-processed paper at the lexical level and at the sentence level respectively through the second feature channel, to obtain the lexical-level statistical features and sentence-level statistical features of the to-be-processed paper;

[0209] Among them, the lexical level includes one or more of the following levels: lexical richness, number of unique words, and the usage of quantitative words, stop words, conjunctions, target words, and template words respectively, and the target word refers to a word whose occurrence frequency in the to-be-processed paper is higher than a preset frequency threshold;

[0210] The sentence level includes one or more of the following levels: syntactic structure complexity, sentence length variation distribution, proportion of synonymous expressions used, and usage of template sentences.

[0211] In a possible implementation, when the above-mentioned macro feature extraction module performs quality assessment on the to-be-processed paper using the third feature channel in the paper detection large language model to obtain the quality assessment features of the to-be-processed paper, it can specifically be used for:

[0212] Perform text quality assessment on the to-be-processed paper using preset quality assessment metrics to obtain the quality assessment features of the to-be-processed paper. Among them, the quality assessment metrics include one or more of the following metrics: text readability metric, text diversity metric, and text predictability metric.

[0213] In a possible implementation, when the above-mentioned feature fusion decoding module uses the feature decoding module in the paper detection large language model to obtain the target detection result indicating whether the to-be-processed paper is an AI-generated paper based on the deep semantic features, the text statistical features, and the quality assessment features, it can specifically be used for:

[0214] Use the feature fusion module in the feature decoding module to perform feature fusion based on attention weights on the deep semantic features, the text statistical features, and the quality assessment features to obtain fused features;

[0215] Use the classification output module in the feature decoding module to obtain the target detection result according to the fused features.

[0216] In an embodiment of the present application, an electronic device is further provided, including at least one processor and a memory connected to the processor. Among them: the memory is used to store a computer program, and the processor is used to execute the computer program so that the electronic device can implement each step of the method for detecting AI-generated papers as described above.

[0217] For example, referring to Figure 5 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.

[0218] As Figure 5 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0219] Typically, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, a memory card, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0220] An embodiment of the present application also provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the artificial intelligence-generated paper detection methods provided by the embodiments of the present application.

[0221] An embodiment of the present application also provides a computer-readable storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, they can enable the electronic device to implement any one of the artificial intelligence-generated paper detection methods provided by the embodiments of the present application.

[0222] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0224] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0225] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A detection method for papers generated by artificial intelligence, characterized in that, Including: Performing word and sentence level semantic modeling processing on the paper to be processed using the first feature channel in the pre-trained large language model for paper detection, to obtain the deep semantic features of the paper to be processed, where the deep semantic features represent the word-level and sentence-level semantic information, part-of-speech semantic information, and inter-sentence semantic consistency information of the paper to be processed; Performing text feature statistical processing on the paper to be processed using the second feature channel in the large language model for paper detection, to obtain the text statistical features of the paper to be processed, where the text statistical features represent the language style characteristics of the paper to be processed; Performing quality assessment processing on the paper to be processed using the third feature channel in the large language model for paper detection, to obtain the quality assessment features of the paper to be processed, where the quality assessment features represent the overall quality characteristics of the paper to be processed; Using the feature decoding module in the large language model for paper detection to obtain a target detection result indicating whether the paper to be processed is an AI-generated paper based on the deep semantic features, the text statistical features, and the quality assessment features; Among them, the training data used by the large language model for paper detection in the training stage includes: positive sample data labeled with AI-generated paper labels and negative sample data labeled with normal paper labels. The positive sample data is papers marked as AI-generated content in the retraction watch database, and the negative sample data is papers in the literature database that meet the preset impact factor requirements. The topic similarity between the positive sample data and the negative sample data meets the preset similarity requirements.

2. The detection method for papers generated by artificial intelligence according to claim 1, characterized in that Before performing word and sentence level semantic modeling processing on the paper to be processed using the first feature channel in the pre-trained large language model for paper detection to obtain the deep semantic features of the paper to be processed, performing text feature statistical processing on the paper to be processed using the second feature channel in the large language model for paper detection to obtain the text statistical features of the paper to be processed, and performing quality assessment processing on the paper to be processed using the third feature channel in the large language model for paper detection to obtain the quality assessment features of the paper to be processed, it further includes: Performing structured parsing on the paper to be processed to obtain the corresponding structured paper of the paper to be processed; Taking the structured paper as the paper to be processed.

3. The detection method for papers generated by artificial intelligence according to claim 1, characterized in that The first feature channel includes a text input module, a feature encoding module, an inter-sentence semantic consistency modeling module, and a part-of-speech semantic distribution modeling module; The performing word and sentence level semantic modeling processing on the paper to be processed using the first feature channel in the pre-trained large language model for paper detection to obtain the deep semantic features of the paper to be processed includes: Performing clause splitting, word segmentation, and part-of-speech tagging processing on the paper to be processed through the text input module to obtain the word segmentation sequence and the word segmentation part-of-speech tagging sequence corresponding to each clause in the paper to be processed; The word-level context encoding is performed on each token in the token sequence corresponding to each clause through the feature encoding module, to obtain the word-level encoding vector sequence corresponding to each clause. The weighted aggregation processing of the intra-sentence tokens is performed according to the word-level encoding vector sequence corresponding to each clause, to obtain the sentence vector corresponding to each clause. The sentence-level context encoding is performed on the sentence vectors corresponding to all the clauses in the to-be-processed paper, to obtain the sentence-level encoding vectors corresponding to all the clauses. The weighted aggregation processing among the sentences is performed on the sentence-level encoding vectors corresponding to all the clauses, to obtain the document vector representing the overall semantic information of the to-be-processed paper; The inter-sentence semantic consistency modeling module constructs an inter-sentence similarity matrix according to the sentence-level encoding vectors corresponding to all the clauses in the to-be-processed paper, and performs index feature extraction on the inter-sentence similarity matrix based on a preset statistical index, to obtain the matrix vector representing the internal coherence and semantic consistency of the to-be-processed paper; The part-of-speech semantic distribution modeling module performs statistical analysis processing based on the occurrence frequency on all the part-of-speech in the to-be-processed paper according to the part-of-speech tagging sequence and the word-level encoding vector sequence corresponding to all the clauses in the to-be-processed paper, to obtain the part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper.

4. The detection method for an AI-generated paper according to claim 3, characterized in that, The part-of-speech semantic distribution modeling module performs statistical analysis processing based on the occurrence frequency on all the part-of-speech in the to-be-processed paper according to the part-of-speech tagging sequence and the word-level encoding vector sequence corresponding to all the clauses in the to-be-processed paper, to obtain the part-of-speech vector representing the part-of-speech semantic information of the to-be-processed paper, including: The part-of-speech semantic distribution modeling module, for each part-of-speech among all the part-of-speech, counts the occurrence frequency of each token under this part-of-speech according to the word-level encoding vector sequence corresponding to all the clauses, to obtain the occurrence frequency of each token under this part-of-speech, performs average pooling processing on the occurrence frequencies of all the tokens under this part-of-speech, to obtain the occurrence frequency of this part-of-speech, so as to obtain the occurrence frequencies of all the part-of-speech, concatenates the occurrence frequencies of all the part-of-speech, to obtain the part-of-speech vector.

5. The detection method for an AI-generated paper according to claim 1, wherein The text feature statistics processing of the to-be-processed paper is performed by using the second feature channel in the paper detection large language model, to obtain the text statistical features of the to-be-processed paper, including: The second feature channel performs text feature statistics processing on the to-be-processed paper at the vocabulary level and the sentence level respectively, to obtain the vocabulary-level statistical features and the sentence-level statistical features of the to-be-processed paper; Among them, the vocabulary level includes one or more of the following levels: vocabulary richness, the number of unique words, and the usage conditions of quantitative words, stop words, conjunctions, target words, and template words respectively, and the target word refers to the word whose occurrence frequency in the to-be-processed paper is higher than the preset frequency threshold; The sentence level includes one or more of the following levels: sentence structure complexity, sentence length change distribution, proportion of synonymous expressions used, and usage conditions of template sentences.

6. The detection method for an AI-generated thesis according to claim 1, characterized in that, Performing quality assessment processing on the to-be-processed paper by using the third feature channel in the paper detection large language model to obtain the quality assessment features of the to-be-processed paper, including: Performing text quality assessment on the to-be-processed paper by using a preset quality assessment index to obtain the quality assessment features of the to-be-processed paper, where the quality assessment index includes one or more of the following indexes: text readability index, text diversity index, and text predictability index.

7. The detection method for papers generated by artificial intelligence according to claim 1, characterized in that, Using the feature decoding module in the paper detection large language model to obtain a target detection result indicating whether the to-be-processed paper is an AI-generated paper according to the deep semantic features, the text statistical features, and the quality assessment features, including: Using the feature fusion module in the feature decoding module to perform feature fusion based on attention weights on the deep semantic features, the text statistical features, and the quality assessment features to obtain fused features; Using the classification output module in the feature decoding module to obtain the target detection result according to the fused features.

8. A detection device for detecting papers generated by artificial intelligence, characterized in that, Including: A deep feature extraction module, configured to perform multi-level semantic modeling processing on the to-be-processed paper by using the first feature channel in the pre-trained paper detection large language model to obtain the deep semantic features of the to-be-processed paper, where the deep semantic features represent the semantic information at the word level and sentence level of the to-be-processed paper; A shallow feature extraction module, configured to perform text feature statistics processing on the to-be-processed paper by using the second feature channel in the paper detection large language model to obtain the text statistical features of the to-be-processed paper, and the text statistical features represent the language style information of the to-be-processed paper; A macro feature extraction module, configured to perform quality assessment processing on the to-be-processed paper by using the third feature channel in the paper detection large language model to obtain the quality assessment features of the to-be-processed paper, and the quality assessment features represent the overall quality characteristics of the to-be-processed paper; A feature fusion and decoding module, configured to use the feature decoding module in the paper detection large language model to obtain a target detection result indicating whether the to-be-processed paper is an AI-generated paper according to the deep semantic features, the text statistical features, and the quality assessment features; Wherein, the training data used in the training stage of the paper detection large language model includes: positive sample data labeled with AI-generated paper labels and negative sample data labeled with normal paper labels, the positive sample data is the papers marked as AI-generated content in the retraction watch database, the negative sample data is the papers in the literature database that meet the preset impact factor requirements, and the topic similarity between the positive sample data and the negative sample data meets the preset similarity requirements.

9. An electronic device, characterized in that, Including at least one processor and a memory connected to the processor, where: The memory is used to store a computer program; The processor is configured to execute the computer program so that the electronic device can implement the method for detecting AI-generated papers as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The storage medium stores one or more computer programs which, when executed by an electronic device, enable the electronic device to implement the method for detecting an AI-generated paper as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Composition scoring method based on attention mechanism

    CN107133211A

  • Large language model generation text detection method based on semantic decoupling

    CN119167114A

  • Information retrieval method and device, electronic equipment and storage medium

    CN119513334A

  • Method for detecting Chinese paper module generated by large language model

    CN119886120A

  • System and method for training and operating large language models using codewords

    US12271696B1

Cited By

  • Recognition method, model training method and device and electronic equipment

    CN120804323A

  • Method and system for detecting AI generation text and medium

    CN120849593A

  • Large model generation text detection method, device and system

    CN121412746A

  • A large model generates a text detection method, device and system

    CN121412746B