Methods, devices, and related products for judging AI-generated academic texts

By classifying and rewriting academic texts, calculating the difference in information volume, and using pre-constructed difference distribution data to judge the source of text generated by AI, the problem of large resources and energy investment in the existing technology is solved, and efficient and accurate identification is achieved across disciplines.

CN119357388BActive Publication Date: 2025-08-19TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411300646.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-08-19
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

When identifying AI to generate academic texts, the existing technology requires training models for different disciplines, and the resources and energy are invested greatly, and the recognition effect of a single model in different disciplines is poor, making it difficult to achieve efficient and accurate judgments.

Method used

By classifying the academic text of judgment, rewriting it using the preset big model, calculating the information difference value, and calculating the number of votes generated by AI and written by humans based on the preconstructed difference value distribution data, and judging the text source with the preset threshold.

Benefits of technology

Without the need to train the model separately for each subject, the source of academic texts can be accurately determined using a difference distribution data, saving time and labor costs, and improving the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357388B_ABST
    Figure CN119357388B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and discloses a method, device and related products for judging AI-generated academic texts. The method classifies the academic text to be judged to obtain a target subject category; uses a large model to rewrite the academic text to be judged to obtain a rewritten text; calculates the target information volume of the academic text to be judged and the average information volume of the rewritten text, and subtracts them to obtain the target information volume difference; based on the first difference distribution data, uses the target subject category and the standardized target information volume difference to calculate the first AI-generated vote number and the first human-written vote number; calculates the possibility score of the academic text to be judged being an AI-generated text based on the two types of votes; and judges whether the academic text to be judged is generated by AI based on a preset threshold. The present disclosure does not need to train a model separately for each subject category, and can accurately determine the source of the academic text to be judged using a difference distribution data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more specifically, to a method, device, and related products for determining AI-generated academic texts. Background Art

[0002] In today's digital and intelligent age, academic research and exchange are unprecedentedly active. With the rapid development of artificial intelligence (AI) technology, its application in academic text generation is becoming increasingly widespread. However, in the academic field, ensuring the originality and reliability of research results is crucial. Failure to clearly identify the source of academic texts can undermine the balance and health of the academic ecosystem. Furthermore, for academic evaluation systems, the ability to accurately determine the generation method of texts is directly related to the fair assessment of the value of academic achievements. Therefore, exploring methods for judging AI-generated academic texts has become a pressing and critical task in current academic development.

[0003] At present, the common method of detecting whether academic texts are generated by AI mainly focuses on using the differences in text features between academic content generated by AI and academic content written by humans, such as differences in language style, logical structure, word usage habits and other features, to automatically classify them. However, academic texts in different disciplines are quite different. Therefore, when using existing academic content judgment methods to identify academic texts generated by AI, if you want to achieve high recognition accuracy in each subject area, you need to train different models or classification methods for different disciplines, which will require more resources and energy. The recognition effect of a single model on academic content in different disciplines also varies greatly. Therefore, there is an urgent need to solve this technical problem. Summary of the Invention

[0004] In response to the above situation, the embodiments of the present disclosure provide a method, device and related products for determining whether AI generates academic texts, aiming to solve the above problems or at least partially solve the above problems.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for determining whether an AI-generated academic text is generated, the method comprising:

[0006] Classify the acquired academic texts to be judged and obtain the target subject category;

[0007] Rewriting the academic text to be judged using a first preset large model and a preset rewriting instruction to obtain at least one rewritten text; the preset rewriting instruction includes the following rewriting requirements: maintaining the original meaning and retaining the original style;

[0008] Calculating target information volume data of the academic text to be judged and average information volume data of the at least one rewritten text; subtracting the target information volume data from the average information volume data to obtain a target information volume difference;

[0009] Based on the pre-constructed first difference distribution data, the first AI-generated vote count and the first human-written vote count are calculated using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference between the rewritten AI-generated sample texts and the rewritten human-written sample texts in different subject categories in different intervals;

[0010] Based on the first AI-generated vote number and the first human-written vote number, a possibility score of the academic text to be judged being an AI-generated text is calculated; and based on the possibility score and a preset threshold, it is judged whether the academic text to be judged is generated by AI to obtain a judgment result.

[0011] In a second aspect, the embodiments of the present disclosure further provide a device for determining whether AI-generated academic texts are used, the device comprising:

[0012] The classification module is used to classify the acquired academic texts to be judged and obtain the target subject category;

[0013] A rewriting module is configured to rewrite the academic text to be judged using a first preset large model and preset rewriting instructions to obtain at least one rewritten text; the preset rewriting instructions include the following rewriting requirements: maintaining the original meaning and retaining the original style;

[0014] a calculation module, configured to calculate target information amount data of the academic text to be judged and average information amount data of the at least one rewritten text; and subtract the target information amount data from the average information amount data to obtain a target information amount difference;

[0015] A voting module is configured to calculate a first AI-generated vote count and a first human-written vote count based on pre-constructed first difference distribution data, using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference between the rewritten AI-generated sample text and the rewritten human-written sample text in different intervals under different subject categories;

[0016] The judgment module is used to calculate the possibility score of the academic text to be judged being an AI-generated text based on the first AI-generated vote number and the first human-written vote number; and to judge whether the academic text to be judged is generated by AI based on the possibility score and a preset threshold, to obtain a judgment result.

[0017] In a third aspect, an embodiment of the present disclosure further provides an electronic device comprising: a processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to execute the steps of the above-mentioned method for determining whether AI generates academic text.

[0018] In a fourth aspect, an embodiment of the present disclosure also provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple applications, the electronic device executes the steps of the above-mentioned method for judging whether AI generates academic text.

[0019] By means of the above technical solution, the judgment method, device and related products for AI-generated academic texts provided by the embodiments of the present disclosure can use the pre-constructed first difference distribution data to identify the source of the academic text to be judged. The first difference distribution data is equivalent to a reference standard or characteristic pattern, which reflects the difference in information volume between AI-generated sample texts and human-written sample texts in a large number of different subject categories before and after rewriting. The distribution pattern in different intervals reflects the characteristic difference in information expression between AI-generated texts and human-written texts. During implementation, according to the target subject category corresponding to the academic text to be judged and the standardized target information volume difference, the total number of information volume differences after rewriting of AI-generated sample texts and the total number of information volume differences after rewriting of human-written sample texts can be obtained under the standard given by the first difference distribution data. These two total numbers are respectively used as the first AI-generated vote number and the first human-written vote number for the academic text to be judged. Then, based on these votes, the possibility score of the academic text to be judged being an AI-generated text is calculated, and then judged in combination with the preset threshold, thereby achieving the purpose of accurately distinguishing whether the academic text is generated by AI. Compared with the existing technology, the technical solution provided by the embodiment of the present disclosure does not require separate model training for each subject category. It uses a difference interval distribution data to accurately determine whether the academic text to be judged is generated by AI, which also greatly saves time and labor costs.

[0020] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0022] Figure 1 A schematic diagram of a flow chart of a method for determining AI-generated academic texts provided in an embodiment of the present disclosure is shown;

[0023] Figure 2 A schematic diagram of the structure of the device for judging AI-generated academic texts provided in an embodiment of the present disclosure is shown;

[0024] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present disclosure more clear, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present disclosure.

[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0027] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such usage is interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in sequences other than those illustrated or described herein. In addition, the term "including" and its variations are to be interpreted as open-ended terms meaning "including but not limited to."

[0028] As mentioned above, in the prior art, the common method for detecting whether academic texts are generated by AI mainly focuses on automatically classifying academic content by using the differences in text features between AI-generated academic content and human-written academic content. Specifically, there are the following detection methods: (1) Extracting features such as punctuation, word preference, and paragraph distribution in the text, and applying traditional machine learning methods to classify and analyze the statistical features of academic texts generated by AI and written by humans. Through statistical analysis of human texts and AI-generated texts, it can be found that human texts and AI-generated texts have obvious differences in word usage habits, named entity expressions, vocabulary richness, and sentence structure. Therefore, the statistical features of text fragments, such as high-frequency words, stop words, conjunctions, sentence length, paragraph distribution, and other fine-grained text statistical features, are used as feature sets to distinguish human texts from AI-generated texts. (2) Model verification: When a large model generates text content, it calculates the probability distribution of each possible word based on the generated text content. At the same time, in order to ensure the coherence and logic of the generated content, the model usually selects words with a higher sampling probability. Therefore, through the autoregressive model of the same architecture, the probability characteristics of all texts in the text generation process are calculated to see if they meet the predictions of the AI generation model. (3) Deep learning model classification: directly train the deep learning model using AI-generated academic content and human-written content as input. (4) Train the language model based solely on a large amount of human-written academic content, and then calculate the difference in the threshold performance of AI-generated academic content and human-written content to distinguish AI-generated academic content from human-written content.

[0029] However, academic texts in different disciplines are quite different. Therefore, when using existing academic content judgment methods to identify academic texts generated by AI, if you want to achieve high recognition accuracy in each subject area, you need to train different models or classification methods that are applicable to different disciplines, which will require more resources and energy. The recognition effect of a single model on academic content in different disciplines also varies greatly. Based on this, the present invention proposes a judgment method, device and related products for AI-generated academic texts. The following is a detailed description of the present disclosure through specific embodiments.

[0030] To facilitate understanding of this embodiment, a method for determining whether an AI-generated academic text is generated is first introduced in detail. The execution subject of the method for determining whether an AI-generated academic text is generated provided by the embodiment of this disclosure is generally a computer device with certain computing capabilities. The computer device includes, for example, a terminal device or a server or other processing device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a personal digital assistant (PDA), etc. In some possible implementations, the method for determining whether an AI-generated academic text is generated can be implemented by a processor calling computer-readable instructions stored in a memory.

[0031] Figure 1 A flow chart of the method for judging the AI-generated academic text provided by the embodiment of the present disclosure is shown. Figure 1 It can be seen that the embodiment of the present disclosure includes at least steps S101-S105:

[0032] S101: Classify the acquired academic texts to be judged to obtain target subject categories.

[0033] This embodiment does not limit the content, length, or source of the academic text to be judged. The academic text to be judged can be any academic-related text, such as all or part of an academic paper, research report, conference paper, case study, or the like. The academic text to be judged can be from the internet or user input.

[0034] During implementation, for example, the academic text to be judged can be input by the user through a front-end interface (such as a web page, a mobile application interface, a desktop application interface, etc.), and the user terminal directly or indirectly sends the academic text to be judged input by the user to the execution entity of this embodiment (such as a server) through wired or wireless communication.

[0035] After obtaining the academic text to be judged, the academic text to be judged is classified to obtain one or more target subject categories CID = {CID1, CID2, ..., CID n}, where n is a natural number. During implementation, any academic classification method can be used, such as the Chinese Library Classification System or the Tongfang Knowledge Network 168 Subject Classification. In specific implementation, existing large models and text classification tools can be used for classification.

[0036] S102: Rewrite the academic text to be judged using a first preset large model and preset rewriting instructions to obtain at least one rewritten text; the preset rewriting instructions include the following rewriting requirements: maintaining the original meaning and retaining the original style.

[0037] In this step, the preset rewriting instructions are limited to include two rewriting requirements: maintaining the original meaning and retaining the original style. Among them, the original style includes many unique features of the text, such as word usage habits, sentence structure, tone, etc. Texts written by humans and texts written by AI have their own unique styles. This embodiment uses preset rewriting instructions to enable the model to retain the original style when rewriting the text, avoiding the introduction of too many new changes and uncertainties, so that the subsequent information difference between the academic text to be judged and the rewritten text can accurately reflect the characteristics of the academic text to be judged, and then use the subsequent information difference distribution data to accurately identify the source of the academic text to be judged. The preset rewriting instructions can, for example, be "Rewrite text X, maintain the original meaning, and retain the original style. Text X:..." Here, the first preset large model can be any existing large language model, which is not limited in this embodiment.

[0038] S103: Calculate the target information volume data of the academic text to be judged and the average information volume data of the at least one rewritten text; subtract the target information volume data from the average information volume data to obtain a target information volume difference.

[0039] Here, information content data is used to measure the uncertainty or randomness of the text, which can reflect the amount of information contained in the text.

[0040] S104: Based on the pre-constructed first difference distribution data, the first AI-generated vote number and the first human-written vote number are calculated using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference of the AI-generated sample text after rewriting and the information difference of the human-written sample text after rewriting in different subject categories in different intervals.

[0041] For example, assume that the target subject categories include subject category 1 and subject category 2, and the standardized target information difference is a. Under subject category 1 in the pre-constructed first difference distribution data, a difference interval A1 including a can be found. Based on difference interval A1, the information difference quantity 1 of the AI-generated sample text after rewriting and the information difference quantity 1 of the human-written sample text after rewriting can be determined. Under subject category 2 in the first difference distribution data, a difference interval A2 including a can be found. Based on difference interval A2, the information difference quantity 2 of the AI-generated sample text after rewriting and the information difference quantity 2 of the human-written sample text after rewriting can be determined. The information difference quantities 1 and 2 of the AI-generated sample text after rewriting are added to obtain the first AI-generated voting number; the information difference quantities 1 and 2 of the human-written sample text after rewriting are added to obtain the first human-written voting number.

[0042] S105: Based on the first AI-generated vote count and the first human-written vote count, calculate the possibility score of the academic text to be judged to be an AI-generated text; and based on the possibility score and a preset threshold, judge whether the academic text to be judged is generated by AI to obtain a judgment result.

[0043] In specific implementation, for example, the probability score of whether the academic text to be judged is AI-generated text can be calculated according to the following formula (1):

[0044]

[0045] Among them, SCORE is the probability score of the academic text to be judged as AI-generated text, CAN is the number of votes generated by the first AI, and CBN is the number of votes written by the first human.

[0046] After obtaining the likelihood score, the relationship between the likelihood score and the preset threshold can be determined. If the likelihood score is greater than or equal to the preset threshold, the judgment result is that the academic text to be judged is generated by AI; if the likelihood score is less than the preset threshold, the judgment result is that the academic text to be judged is written by a human. Here, the preset threshold can be set according to actual needs and is not limited to this embodiment. For example, the preset threshold can be set to 0.85.

[0047] During implementation, the judgment result can be displayed on the front-end page, and can also be displayed on the front-end page together with the possibility score.

[0048] from Figure 1It can be seen from the method shown that the embodiment of the present disclosure uses the pre-constructed first difference distribution data to identify the source of the academic text to be judged. The first difference distribution data is equivalent to a reference standard or characteristic pattern, which reflects the difference in the amount of information of a large number of AI-generated sample texts and human-written sample texts before and after rewriting under different subject categories. The distribution pattern in different intervals reflects the characteristic difference in information expression between AI-generated texts and human-written texts. During implementation, according to the target subject category corresponding to the academic text to be judged and the standardized target information difference, the total number of information difference values of the AI-generated sample texts after rewriting and the total number of information difference values of the human-written sample texts after rewriting can be obtained under the standard given by the first difference distribution data. These two total numbers are respectively used as the first AI-generated vote number and the first human-written vote number for the academic text to be judged. Then, based on these votes, the possibility score of the academic text to be judged being an AI-generated text is calculated, and then judged in combination with the preset threshold, thereby achieving the purpose of accurately distinguishing whether the academic text is generated by AI. Compared with the existing technology, the technical solution provided by the embodiment of the present disclosure does not require separate model training for each subject category. It uses a difference interval distribution data to accurately determine whether the academic text to be judged is generated by AI, which also greatly saves time and labor costs.

[0049] Furthermore, in order to better illustrate the process of the above-mentioned method for determining AI-generated academic texts, as a refinement and extension of the above-mentioned embodiment, the embodiment of the present invention provides several optional embodiments, but is not limited thereto, as shown below:

[0050] In one possible implementation, the method further includes: obtaining academic texts of various subject categories written by humans to form a collection of human-written academic texts; for any target subject category, using a second preset macro model and a target generation instruction corresponding to the target subject category to generate multiple AI-generated academic texts; the AI-generated academic texts corresponding to each target subject category constitute a collection of AI-generated academic texts; using the first preset macro model and the preset rewriting instruction to rewrite any text in the collection of human-written academic texts and the collection of AI-generated academic texts to obtain a collection of human-rewritten academic texts and a collection of AI-rewritten academic texts; calculating the information volume data of any first target text in the collection of human-written academic texts, as well as the information volume data of the human-rewritten academic texts. The first information difference value set is obtained by calculating the difference between the average information amount data of at least one rewritten text corresponding to the first target text in the set; and the information amount data of any second target text in the AI-generated academic text set and the difference between the average information amount data of at least one rewritten text corresponding to the second target text in the AI-rewritten academic text set are calculated to obtain the second information difference value set; the elements in the first information difference value set and the second information difference value set are standardized respectively to obtain a standardized first information difference value set and a standardized second information difference value set; the first difference distribution data is generated according to the standardized first information difference value set and the standardized second information difference value set.

[0051] In the present embodiment, multiple subject categories include the aforementioned target subject categories. Multiple subject categories can be full subject categories. The subject classification method is the same as the classification method used in the aforementioned S101, and will not be described in detail here. During implementation, multiple academic content texts of a certain length can be randomly extracted from the academic content (academic papers, conference papers, dissertations, etc.) under each subject category to form a human-written academic text set. If the total number of subject categories is CN, and each subject category has Num texts, then the human-written academic text set includes Num*CN texts.

[0052] During implementation, a comparison table of text serial numbers and subject category IDs (Table 1) can also be generated simultaneously to facilitate subsequent generation of Table 2. The text serial number indicates the position of the text in the collection.

[0053] Table 1 Comparison table of text serial numbers and subject categories

[0054]

[0055]

[0056] Then, for any target subject category, a second preset large model and a target generation instruction corresponding to the target subject category can be used to generate multiple AI-generated academic texts, thereby obtaining an AI-generated academic text set. Here, the second preset large model can be any existing large language model, which is not limited in this embodiment. The first preset large model and the second preset large model can be the same or different. If the total number of subject categories is CN, and each subject category has Num texts, then the AI-generated academic text set includes Num*CN texts. During implementation, similarly, a text sequence number and subject category ID comparison table can also be generated for subsequent calculations.

[0057] A target generation instruction, for example, may be "Generate an academic text in the subject category X." In one possible implementation, the target generation instruction is generated based on the subject category, research topic, and academic content type, where the academic content type includes at least one of the following: introduction, research background, research methods, and research conclusions.

[0058] For example, the target generation instruction is "Generate a Z-length academic content in the subject area X with the research topic Y. Requirements: (1) the content length is MinLen; (2) the content complies with academic standards", where X represents the name of a specific subject category; Y represents the research topic under the subject; and Z represents different academic contents such as introduction, research background, research methods, and research conclusions.

[0059] The target generation instructions in this embodiment clearly define the subject categories, research topics, and academic content types, so that the generated AI-generated academic text collection is more in line with the specific needs of actual application scenarios, with greater pertinence, richness and practicality, thereby significantly improving the accuracy of the first difference distribution data obtained subsequently.

[0060] Then, the first preset large model can be used to generate multiple rewritten texts for each academic text in the set of human-written academic texts, thereby obtaining a set of human-rewritten academic texts. Let BM represent the set of human-rewritten academic texts. Suppose there are Num*CN texts in the set of human-written academic texts, then BM = {{B1M1, B1M2, ..., B1Mr}, {B2M1, B2M2, ..., B2Mr}, ...}, where there are Num*CN subsets in BM, and each subset has r texts. Using the first preset large model, multiple rewritten texts are generated for each academic text in the set of AI-generated academic texts, thereby obtaining a set of AI-rewritten academic texts. Let AM represent the AI rewritten academic text set. Suppose there are Num*CN texts in the AI rewritten academic text set, then AM = {{A1M1, A1M2, ..., A1Mr}, {A2M1, A2M2, ..., A2Mr}, ...}, where there are Num*CN subsets in AM and each subset has r texts.

[0061] In a possible implementation, the type of information volume data is: information entropy, vocabulary richness, or sentence complexity.

[0062] In specific implementation, if the type of information data is information entropy, the information data of a text can be calculated according to the following formula (2):

[0063]

[0064] Here, we can first count the frequency of occurrence of each word (or character) in the text. Assume that there are n different words (or characters) in the sentence, denoted as W1, W2, ..., Wn, and their frequencies of occurrence are f1, f2, ..., fn. The word frequencies are transformed into probability distributions. For each word Wi, its probability p(x i ) can be expressed as formula (3):

[0065]

[0066] The denominator in formula (3) is the total frequency of all words (or characters) in the text.

[0067] If the type of information data is lexical richness, the information data of a text can be calculated as follows: count the number of different words (characters) appearing in the text, denoted as V; count the total number of words (characters) in the text, denoted as T; and use V / T as the information data of the text. If the type of information data is sentence complexity, the information data of a text can be calculated as follows: measure the length of each sentence (which can be the number of words or characters). Assume that there are m sentences in the text, and the length of each sentence is L1, L2, ..., Lm; the information data = (L1 + L2 + ... + Lm) / m.

[0068] For example, if the type of information data is information entropy, the information entropy of any first target text in the set of human-written academic texts can be calculated using the above formula (2). Let B = {B1, B2, ...} represent the set of human-written academic texts, and then the information entropy set HB = (HB1, HB2, ...) corresponding to set B can be obtained, where HBi = H(bi). Let the set BM {{B1M1, B1M2 ...}, {B2M1, B2M2 ...} ...} represent the set of human-rewritten academic texts, where {BiM1, BiM2 ...} corresponds to Bi in set B. The above formula (2) can be used to calculate the average information entropy of at least one rewritten text corresponding to each first target text, and obtain the average information entropy set HBM = {HB1M, HB2M, ...}, HBiM = Mean(H({BiM1, BiM2, ...})) = Mean(H(BiM1), H(BiM2), ...). Finally, the corresponding elements of set HB and set HBM are subtracted to obtain the first information difference set. If BX represents the first information difference set, then BX = {x1, x2, ...} = HB-HBM = {HB1-HB1M, HB2-HB2M, ...}, and the number of elements in BX is the same as the number of elements in the set of human-written academic texts. Similarly, we can calculate the information entropy set HA corresponding to the AI-generated academic text set A, the average information entropy set HAM corresponding to the AI-rewritten academic text set AM, and the first information difference set AX = {x1, x2, ...} = HA-HAM = {HA1-HA1M, HA2-HA2M, ...}, and the number of elements in AX is the same as the number of elements in the AI-generated academic text set.

[0069] In order to ensure that each information difference is within a reasonable range, it is convenient for subsequent data processing and analysis. In this embodiment, the elements in the first information difference set and the second information difference set can be standardized to obtain a standardized first information difference set and a standardized second information difference set. For example, the elements in the set can be standardized using a zero-mean standardization method. In a possible implementation of the present disclosure, the elements in the first information difference set and the second information difference set are standardized to obtain a standardized first information difference set and a standardized second information difference set, including: obtaining the maximum difference and the minimum difference in the first information difference set and the second information difference set; for any target difference in the first information difference set and the second information difference set, using the maximum difference and the minimum difference, the target difference is standardized to obtain a standardized target difference; each standardized target difference constitutes the standardized first information difference set and the standardized second information difference set.

[0070] In specific implementation, the standardized target difference can be calculated according to the following formula (4):

[0071]

[0072] Where x′ represents the normalized target difference, x represents the target difference, minV represents the minimum difference, and MaxV represents the maximum difference.

[0073] During implementation, if x is greater than the maximum difference, then x=maximum difference; if x is less than the minimum difference, then x=minimum difference.

[0074] Finally, based on the two standardized sets, the distribution data of the information difference between the AI-generated text and the human-written text after rewriting in different subject categories is generated, that is, the first difference distribution data. In one possible implementation, the first difference distribution data is generated based on the standardized first information difference set and the standardized second information difference set, including: grouping the elements in the standardized first information difference set and the standardized second information difference set according to subject categories to obtain multiple groups of information differences; for any target information difference group, dividing the elements in the target information difference group into intervals according to a first preset difference interval set, and obtaining a first total number of differences belonging to the standardized first information difference set and a second total number of differences belonging to the standardized second information difference set under each difference interval; based on each of the subject categories, each of the difference intervals and the corresponding first total number and second total number, the first difference distribution data is obtained.

[0075] Assume that the standardized first information difference value set is BX' = {x'1, x'2, x'3, x'4, x'5, x'6, x'7, x'8, x'9, x'10}, and the standardized second information difference value set is AX' = {y'1, y'2, y'3, y'4, y'5, y'6, y'7, y'8, y'9, y'10}. Among them, the text corresponding to x'1, x'2, y'1, y'2 belongs to subject category 1; the text corresponding to x'3, x'4, y'3, y'4 belongs to subject category 2; the text corresponding to x'5, x'6, y'5, y'6 belongs to subject category 3; the text corresponding to x'7, x'8, y'7, y'8 belongs to subject category 4; and the text corresponding to x'9, x'10, y'9, y'10 belongs to subject category 5. The following is an exemplary description of this embodiment using the set BX' and the set AX'.

[0076] First, the elements in sets BX' and AX' are grouped according to subject categories to obtain multiple groups of information differences: information difference group 1 (x'1, x'2, y'1, y'2), information difference group 2 (x'3, x'4, y'3, y'4), information difference group 3 (x'5, x'6, y'5, y'6), information difference group 4 (x'7, x'8, y'7, y'8), and information difference group 5 (x'9, x'10, y'9, y'10).

[0077] Then, for any target information difference value group, the elements in the target information difference value group can be divided into intervals according to the first preset difference interval set, and the total number of differences belonging to the standardized first information difference value set (first total number) and the total number of differences belonging to the standardized second information difference value set (second total number) under each difference interval corresponding to the target information difference value group are obtained. Here, the first preset difference interval set can be set according to actual needs. For example, the first preset difference interval set is set to: [0-0.1], [0.1-0.2], [0.2-0.3], [0.3-0.4], [0.4-0.5], [0.5-0.6], [0.6-0.7], [0.7-0.8], [0.8-0.9], [0.9-1]. For example, assume that in information difference group 1, x'1 = 0.85, x'2 = 0.8, y'1 = 0.25, and y'2 = 0.24. Based on the first set of preset difference intervals, the elements in information difference group 1 are divided into intervals. The first and second totals for each difference interval can be statistically obtained: in the difference interval [0.2-0.3], the first total is 0, and the second total is 2; in the difference interval [0.8-0.9], the first total is 2, and the second total is 0; and in other difference intervals, both the first and second totals are 0. Similarly, the first and second totals for each difference interval corresponding to information difference group 2, information difference group 3, information difference group 4, and information difference group 5 can be obtained.

[0078] Finally, based on each subject category, difference interval, and the corresponding first total and second total, the first difference distribution data is obtained. For example, the following Table 2 shows the first difference distribution data.

[0079] Table 2 First difference distribution data

[0080]

[0081]

[0082] This example rewrites a large number of known human-written and AI-generated texts and calculates the difference in information between the texts before and after the rewriting to obtain difference distribution data. This distribution data can reflect the general patterns and characteristics of these two types of texts after specific processing. This difference distribution data can subsequently be used to accurately determine the source of new academic texts. Furthermore, the information difference calculation requires little effort, and the distribution data can be constructed quickly, efficiently, and at a low cost.

[0083] In actual application scenarios, if the academic text to be judged has problems such as not having a clear disciplinary tendency, spanning multiple disciplinary fields, or being too new to accurately determine the disciplinary category, the accuracy of the probability score calculated using the above formula (1) is relatively poor. Based on this, in a possible embodiment of the present disclosure, the method further includes:

[0084] Dividing the elements in the standardized first information difference value set and the standardized second information difference value set into intervals according to a second preset difference value interval set, to obtain a third total number of difference values belonging to the standardized first information difference value set and a fourth total number of difference values belonging to the standardized second information difference value set in each interval;

[0085] Based on the second set of preset difference intervals, the third totals and the fourth totals corresponding to each interval, second difference distribution data is obtained, where the second difference distribution data reflects the distribution of the difference in text information between the AI-generated text and the human-written text after rewriting in different intervals;

[0086] Based on the second difference distribution data, the second AI-generated vote count and the second human-written vote count are calculated using the standardized target information difference;

[0087] The method of calculating the possibility score of the academic text to be judged as an AI-generated text based on the first AI-generated vote number and the first human-written vote number includes: calculating a first possibility score based on the first AI-generated vote number, the first human-written vote number and a preset first weight; calculating a second possibility score based on the second AI-generated vote number, the second human-written vote number and a preset second weight; and adding the first possibility score and the second possibility score to obtain a possibility score of the academic text to be judged as being AI-generated.

[0088] In this embodiment, for any target difference interval in the second preset difference interval set, the total number of difference values in the standardized first information amount difference value set that fall within the target difference interval can be counted to obtain a third total number corresponding to the target difference interval. The total number of difference values in the standardized second information amount difference value set that fall within the target difference interval can be counted to obtain a fourth total number corresponding to the target difference interval. Similarly, the third total number and fourth total number corresponding to each difference interval can be obtained.

[0089] Then, based on the second preset difference interval set, the third total and the fourth total corresponding to each interval, the second difference distribution data is obtained. For example, the second difference distribution data is shown in Table 3 below.

[0090] Table 3 Second difference distribution data

[0091] Difference interval Fourth total The third total 0-0.1 50086 267 0.1-0.2 408960 5341 ... ... ... 0.9-1.0 178 50889

[0092] Next, based on the second difference distribution data, the second AI-generated vote count and the second human-written vote count are calculated using the standardized target information difference. For example, let the standardized target information difference be b. A difference interval B including b can be determined in the second difference distribution data. Based on the difference interval B, the number of information difference values after rewriting the AI-generated sample text (the second AI-generated vote count) and the number of information difference values after rewriting the human-written sample text (the second human-written vote count) can be determined within this interval.

[0093] Finally, the first likelihood score is calculated based on the first AI-generated vote count, the first human-written vote count, and the preset first weight; the second likelihood score is calculated based on the second AI-generated vote count, the second human-written vote count, and the preset second weight; the first likelihood score and the second likelihood score are added together to obtain the likelihood score of the academic text to be judged to be AI-generated. That is, the likelihood score of the academic text to be judged to be AI-generated is calculated according to the following formula (5):

[0094]

[0095] Wherein, W is a preset first weight, 1-W is a preset second weight, AN is a fourth total, and BNw is a third total. In implementation, W is greater than 0.5 and less than or equal to 1, and can be set according to actual needs and is not limited in this embodiment.

[0096] This embodiment constructs second difference distribution data, which is based on the first difference distribution data and ignores the factor of subject category. During implementation, based on the second difference distribution data, the second AI-generated vote count and the second human-written vote count can be calculated using the standardized target information difference; when calculating the probability score of the academic text to be judged to be AI-generated, the second AI-generated vote count and the second human-written vote count are also taken into account, so that when the academic text to be judged has problems such as no obvious subject tendency, spanning multiple subject fields, or the subject category is difficult to accurately judge, the accuracy of the final probability score can be guaranteed.

[0097] The present disclosure also provides a method for determining whether AI-generated academic texts are valid, including the following steps:

[0098] Step S1: Acquire academic texts of various subject categories written by humans to form a collection of academic texts written by humans.

[0099] Step S2: For any target subject category, use the second preset large model and the target generation instructions corresponding to the target subject category to generate an AI-generated academic text set.

[0100] During implementation, target generation instructions are generated based on subject categories, research topics, and academic content types, which include introduction, research background, research methods, and research conclusions.

[0101] Step S3: Using the first preset large model and the preset rewriting instructions, rewrite any text in the human-written academic text set and the AI-generated academic text set to obtain the human-rewritten academic text set and the AI-rewritten academic text set.

[0102] Step S4: Calculate the difference between the information entropy of any first target text in the set of academic texts written by humans and the average information entropy of at least one rewritten text corresponding to the first target text in the set of academic texts rewritten by humans, and obtain a first information difference set.

[0103] Step S5: Calculate the information entropy of any second target text in the AI-generated academic text set and the difference between the average information entropy of at least one rewritten text in the AI-rewritten academic text set corresponding to the second target text to obtain a second information difference set.

[0104] Step S6: Obtain the maximum difference and the minimum difference in the first information amount difference set and the second information amount difference set.

[0105] Step S7: For any target difference in the first information difference set and the second information difference set, the target difference is standardized using the maximum difference and the minimum difference to obtain a standardized target difference; each standardized target difference constitutes a standardized first information difference set and a standardized second information difference set.

[0106] Step S8: Group the elements in the standardized first information difference value set and the standardized second information difference value set according to subject categories to obtain multiple groups of information difference values.

[0107] Step S9: For any target information difference group, the elements in the target information difference group are divided into intervals according to the first preset difference interval set, and a first total number of differences belonging to the standardized first information difference set and a second total number of differences belonging to the standardized second information difference set are obtained under each difference interval.

[0108] Step S10: Based on each subject category, each difference interval and the corresponding first total and second total, first difference distribution data is obtained.

[0109] Step S11: According to the second preset difference interval set, the elements in the standardized first information difference set and the standardized second information difference set are divided into intervals to obtain the third total number of differences belonging to the standardized first information difference set and the fourth total number of differences belonging to the standardized second information difference set in each interval.

[0110] Step S12: obtaining second difference distribution data based on the second preset difference interval set, the third total number and the fourth total number corresponding to each interval.

[0111] Step S13: Classify the acquired academic text to be judged to obtain the target subject category.

[0112] Step S14: rewrite the academic text to be judged using the first preset large model and the preset rewriting instruction to obtain at least one rewritten text.

[0113] Step S15: Calculate the target information volume data of the academic text to be judged and the average information volume data of at least one rewritten text; subtract the target information volume data from the average information volume data to obtain the target information volume difference.

[0114] Step S16: Based on the pre-constructed first difference distribution data, using the target subject category and the standardized target information difference, calculate the first AI-generated vote number and the first human-written vote number.

[0115] Step S17: Based on the second difference distribution data, the second AI-generated vote number and the second human-written vote number are calculated using the standardized target information difference.

[0116] Step S18: Calculate a first likelihood score based on the number of votes generated by the first AI, the number of votes written by the first human, and the preset first weight.

[0117] Step S19: Calculate a second likelihood score based on the second AI-generated vote count, the second human-written vote count, and a preset second weight.

[0118] Step S20: Add the first possibility score and the second possibility score to obtain a possibility score for whether the academic text to be determined is generated by AI.

[0119] Step S21: Based on the possibility score and the preset threshold, determine whether the academic text to be judged is generated by AI and obtain a judgment result.

[0120] Those skilled in the art will understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order, but does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0121] It should be noted that, in practical applications, all possible implementation methods described above can be arbitrarily combined to form possible embodiments of the present disclosure, and will not be described one by one here.

[0122] Based on the same concept, the present disclosure also provides a device for judging whether AI generates academic texts. Figure 2 The schematic diagram of the structure of the judgment device for AI-generated academic text provided by the embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the AI-generated academic text judgment device 200 provided in the embodiment of the present disclosure includes:

[0123] The classification module 201 is used to classify the acquired academic text to be judged to obtain the target subject category;

[0124] The rewriting module 202 is configured to rewrite the academic text to be judged using a first preset large model and preset rewriting instructions to obtain at least one rewritten text; the preset rewriting instructions include the following rewriting requirements: maintaining the original meaning and retaining the original style;

[0125] The calculation module 203 is configured to calculate the target information amount data of the academic text to be judged and the average information amount data of the at least one rewritten text; and subtract the target information amount data from the average information amount data to obtain a target information amount difference;

[0126] Voting module 204 is configured to calculate a first AI-generated vote count and a first human-written vote count based on pre-constructed first difference distribution data, using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference between the rewritten AI-generated sample text and the rewritten human-written sample text in different subject categories in different intervals;

[0127] The judgment module 205 is used to calculate the possibility score of the academic text to be judged being an AI-generated text based on the first AI-generated vote number and the first human-written vote number; and to judge whether the academic text to be judged is generated by AI based on the possibility score and a preset threshold, to obtain a judgment result.

[0128] In one possible implementation, the device further includes a first construction module, which is used to: obtain academic texts of various subject categories written by humans to form a collection of human-written academic texts; for any target subject category, use a second preset large model and a target generation instruction corresponding to the target subject category to generate multiple AI-generated academic texts; the AI-generated academic texts corresponding to each target subject category constitute an AI-generated academic text collection; use the first preset large model and the preset rewriting instruction to rewrite any text in the human-written academic text collection and the AI-generated academic text collection to obtain a human-rewritten academic text collection and an AI-rewritten academic text collection; calculate the information volume data of any first target text in the human-written academic text collection, as well as the information volume data of the human-rewritten academic text collection. The first information difference value set is obtained by calculating the difference between the average information amount data of at least one rewritten text corresponding to the first target text in the academic text set; and the second information difference value set is obtained by calculating the information amount data of any second target text in the academic text set generated by the AI and the difference between the average information amount data of at least one rewritten text corresponding to the second target text in the academic text set rewritten by the AI; the elements in the first information difference value set and the second information difference value set are standardized respectively to obtain a standardized first information difference value set and a standardized second information difference value set; the first difference distribution data is generated according to the standardized first information difference value set and the standardized second information difference value set.

[0129] In one possible implementation, in the above-mentioned device, the target generation instruction is generated based on subject category, research topic, and academic content type, and the academic content type includes at least one of the following: introduction, research background, research method, and research conclusion.

[0130] In a possible implementation, in the above device, the type of information volume data is: information entropy, vocabulary richness, or sentence complexity.

[0131] In one possible embodiment, the device also includes a standardization processing module, which is used to: obtain the maximum difference and the minimum difference in the first information amount difference value set and the second information amount difference value set; for any target difference in the first information amount difference value set and the second information amount difference value set, use the maximum difference and the minimum difference to standardize the target difference to obtain a standardized target difference; each standardized target difference constitutes the standardized first information amount difference value set and the standardized second information amount difference value set.

[0132] In one possible implementation, in the above-mentioned device, the first construction module is used to: group the elements in the standardized first information difference value set and the standardized second information difference value set according to subject categories to obtain multiple groups of information difference values; for any target information difference value group, divide the elements in the target information difference value group into intervals according to a first preset difference interval set to obtain a first total number of differences belonging to the standardized first information difference value set and a second total number of differences belonging to the standardized second information difference value set under each difference interval; obtain the first difference distribution data based on each of the subject categories, each of the difference intervals and the corresponding first total number and second total number respectively.

[0133] In one possible implementation, the apparatus further includes a second construction module for: performing interval division on the elements in the standardized first information difference set and the standardized second information difference set according to a second preset difference interval set, to obtain a third total number of differences belonging to the standardized first information difference set and a fourth total number of differences belonging to the standardized second information difference set in each interval; obtaining second difference distribution data based on the second preset difference interval set, the third total number, and the fourth total number corresponding to each interval, wherein the second difference distribution data reflects the distribution of text information difference between the AI-generated text and the human-written text in different intervals after rewriting; and calculating a second AI-generated vote number and a second human-written vote number based on the second difference distribution data and the standardized target information difference.

[0134] The judgment module is used to: calculate a first possibility score based on the first AI-generated vote number, the first human-written vote number and a preset first weight; calculate a second possibility score based on the second AI-generated vote number, the second human-written vote number and a preset second weight; and add the first possibility score and the second possibility score to obtain a possibility score for the academic text to be judged to be AI-generated.

[0135] It should be noted that any of the above-mentioned AI-generated academic text judgment devices can implement the above-mentioned AI-generated academic text judgment method one by one, which will not be repeated here.

[0136] Figure 3 FIG. 1 shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 3As shown, at the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device. Of course, the electronic device may also include other hardware required for its services.

[0137] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0138] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0139] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a judgment device for AI-generated academic text at the logical level. The processor executes the program stored in the memory and is specifically used to perform the aforementioned method.

[0140] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0141] The electronic device can execute the method for judging whether AI generates academic texts provided by multiple embodiments of the present disclosure, and realize the judgment device for AI generating academic texts. Figure 2 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present disclosure.

[0142] The embodiments of the present disclosure also propose a computer-readable storage medium, which stores one or more programs, and the one or more programs include instructions. When the instructions are executed by an electronic device including multiple applications, the electronic device can execute the judgment method of AI-generated academic text provided by multiple embodiments of the present disclosure.

[0143] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0145] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0147] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0148] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0149] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0150] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the method for judging the AI-generated academic text described in the above method embodiment. For details, please refer to the above method embodiment, which will not be repeated here.

[0151] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0152] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0153] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0154] The above are merely examples of the present disclosure and are not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims of the present disclosure.

Claims

1. A method for judging AI-generated academic texts, characterized in that: The method comprises: Classify the acquired academic texts to be judged and obtain the target subject category; Rewriting the academic text to be judged using a first preset large model and a preset rewriting instruction to obtain at least one rewritten text; the preset rewriting instruction includes the following rewriting requirements: maintaining the original meaning and retaining the original style; Calculating target information volume data of the academic text to be judged and average information volume data of the at least one rewritten text; subtracting the target information volume data from the average information volume data to obtain a target information volume difference; Based on the pre-constructed first difference distribution data, the first AI-generated vote count and the first human-written vote count are calculated using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference between the rewritten AI-generated sample texts and the rewritten human-written sample texts in different subject categories in different intervals; Based on the first AI-generated vote count and the first human-written vote count, a probability score is calculated for the academic text to be determined to be an AI-generated text; and based on the probability score and a preset threshold, a determination is made as to whether the academic text to be determined is AI-generated, thereby obtaining a determination result. The method further comprises: Acquire academic texts in various subject categories written by humans to form a collection of academic texts written by humans; For any target subject category, a second preset macromodel and a target generation instruction corresponding to the target subject category are used to generate multiple AI-generated academic texts; the AI-generated academic texts corresponding to each target subject category constitute a set of AI-generated academic texts; Rewriting any text in the human-written academic text set and the AI-generated academic text set using the first preset large model and the preset rewriting instruction to obtain a human-rewritten academic text set and an AI-rewritten academic text set; Calculating the difference between the information content data of any first target text in the set of human-written academic texts and the average information content data of at least one rewritten text corresponding to the first target text in the set of human-rewritten academic texts to obtain a first information content difference value set; and Calculating the difference between the information content data of any second target text in the AI-generated academic text set and the average information content data of at least one rewritten text corresponding to the second target text in the AI-rewritten academic text set to obtain a second information content difference set; performing standardization processing on the elements in the first information difference value set and the second information difference value set respectively to obtain a standardized first information difference value set and a standardized second information difference value set; The first difference distribution data is generated according to the standardized first information amount difference value set and the standardized second information amount difference value set.

2. The method according to claim 1, characterized in that The target generation instruction is generated based on subject category, research topic, and academic content type, and the academic content type includes at least one of the following: introduction, research background, research method, and research conclusion.

3. The method according to claim 1, characterized in that The types of information data are: information entropy, vocabulary richness or sentence complexity.

4. The method according to claim 1, wherein The step of respectively standardizing the elements in the first information difference value set and the second information difference value set to obtain a standardized first information difference value set and a standardized second information difference value set includes: Obtaining a maximum difference and a minimum difference in the first information amount difference value set and the second information amount difference value set; For any target difference in the first information amount difference set and the second information amount difference set, the target difference is standardized using the maximum difference and the minimum difference to obtain a standardized target difference; each standardized target difference constitutes the standardized first information amount difference set and the standardized second information amount difference set.

5. The method according to claim 1, characterized in that The generating the first difference distribution data according to the standardized first information amount difference value set and the standardized second information amount difference value set includes: Grouping the elements in the standardized first information difference value set and the standardized second information difference value set according to subject categories to obtain multiple groups of information difference values; For any target information amount difference value group, according to a first preset difference value interval set, the elements in the target information amount difference value group are divided into intervals, and a first total number of differences belonging to the standardized first information amount difference value set and a second total number of differences belonging to the standardized second information amount difference value set are obtained in each difference interval; Based on each of the subject categories, each of the difference intervals and the corresponding first totals and second totals, the first difference distribution data is obtained.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Dividing the elements in the standardized first information difference value set and the standardized second information difference value set into intervals according to a second preset difference value interval set, to obtain a third total number of difference values belonging to the standardized first information difference value set and a fourth total number of difference values belonging to the standardized second information difference value set in each interval; Based on the second set of preset difference intervals, the third totals and the fourth totals corresponding to each interval, second difference distribution data is obtained, where the second difference distribution data reflects the distribution of the difference in text information between the AI-generated text and the human-written text after rewriting in different intervals; Based on the second difference distribution data, the second AI-generated vote count and the second human-written vote count are calculated using the standardized target information difference; The calculating, based on the first AI-generated vote count and the first human-written vote count, a likelihood score of the to-be-determined academic text being an AI-generated text includes: Calculate a first likelihood score based on the number of votes generated by the first AI, the number of votes written by the first human, and a preset first weight; Calculate a second likelihood score based on the number of votes generated by the second AI, the number of votes written by the second human, and a preset second weight; The first possibility score and the second possibility score are added together to obtain a possibility score of whether the academic text to be determined is generated by AI.

7. A judgment device for AI-generated academic texts, characterized in that: The device comprises: The classification module is used to classify the acquired academic texts to be judged and obtain the target subject category; A rewriting module is configured to rewrite the academic text to be judged using a first preset large model and preset rewriting instructions to obtain at least one rewritten text; the preset rewriting instructions include the following rewriting requirements: maintaining the original meaning and retaining the original style; a calculation module, configured to calculate target information amount data of the academic text to be judged and average information amount data of the at least one rewritten text; and subtract the target information amount data from the average information amount data to obtain a target information amount difference; A voting module is configured to calculate a first AI-generated vote count and a first human-written vote count based on pre-constructed first difference distribution data, using the target subject category and the standardized target information difference; the first difference distribution data reflects the distribution of the information difference between the rewritten AI-generated sample text and the rewritten human-written sample text in different intervals under different subject categories; A judgment module is configured to calculate a likelihood score of the academic text to be judged to be AI-generated based on the first AI-generated vote count and the first human-written vote count; and determine whether the academic text to be judged is AI-generated based on the likelihood score and a preset threshold, thereby obtaining a judgment result; The first construction module is used to: obtain academic texts of various subject categories written by humans to form a collection of human-written academic texts; for any target subject category, use the second preset large model and the target generation instruction corresponding to the target subject category to generate multiple AI-generated academic texts; the AI-generated academic texts corresponding to each target subject category constitute an AI-generated academic text collection; use the first preset large model and the preset rewriting instruction to rewrite any text in the human-written academic text collection and the AI-generated academic text collection to obtain a human-rewritten academic text collection and an AI-rewritten academic text collection; calculate the information volume data of any first target text in the human-written academic text collection, as well as the information volume data of any first target text in the human-rewritten academic text collection and the information volume data of any first target text in the human-rewritten academic text collection. The first information difference value set is obtained by calculating the difference between the average information amount data of at least one rewritten text corresponding to the first target text; and the second information difference value set is obtained by calculating the information amount data of any second target text in the AI-generated academic text set and the difference between the average information amount data of at least one rewritten text corresponding to the second target text in the AI-rewritten academic text set; the elements in the first information difference value set and the second information difference value set are standardized respectively to obtain a standardized first information difference value set and a standardized second information difference value set; the first difference distribution data is generated according to the standardized first information difference value set and the standardized second information difference value set.

8. An electronic device comprising: processor; as well as A memory arranged to store computer-executable instructions, wherein when the executable instructions are executed, the processor is caused to perform the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing one or more programs, characterized in that: When the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text recognition method and device, processor and electronic equipment

    CN117076670A