Text detection method, apparatus, device, medium, and program product

CN122735705APending Publication Date: 2026-09-11CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610795551.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]然而,现有技术方案所依赖的编码器模型通常基于掩码语言建模任务进行预训练,其核心能力在于根据上下文预测被掩码的词汇,导致其提取文本的语义特征时,容易过度关注文本中表面的词汇共现和搭配模式,难以充分捕捉由AI生成的具有复杂模式文本的深层特征

Benefits of technology

[0081] The beneficial effects of the second to sixth aspects mentioned above are described in the corresponding description of the first aspect and will not be repeated here.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735705A_ABST
    Figure CN122735705A_ABST
Patent Text Reader

Abstract

This application provides a text detection method, apparatus, device, medium, and program product, relating to the field of data processing technology, for improving the accuracy of AIGC text detection. The specific technical solution includes: obtaining the perplexity level corresponding to a first text; masking the first text according to a preset masking probability to obtain a first masked text, where the preset masking probability is the probability of masking each word in the first text; constructing text prompt words based on the perplexity level and the first masked text; determining a first probability and a second probability based on the text prompt words, where the first probability is the predicted probability that the first text is manually written, and the second probability is the predicted probability that the first text is AI-generated; and determining the text generation method of the first text based on the first probability and the second probability, where the text generation method includes either manual writing or AI generation. This application is applied to AIGC text detection scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a text detection method, apparatus, device, medium, and program product. Background Technology

[0002] In recent years, Artificial Intelligence Generated Content (AIGC) technology, centered on large language models, has developed rapidly. While improving the efficiency of content creation in various fields, it has also brought risks such as the spread of misinformation. To maintain the security and trustworthiness of the online information environment, developing detection technologies that can effectively distinguish between AI-generated content and human-created content has become an important research topic in this field.

[0003] Currently, the main approach used in AIGC text detection is to extract semantic representations of the text using a pre-trained language model encoder, then construct a dual-branch network structure. One branch integrates global and local features at the document level for overall judgment, while the other branch performs sentence-by-sentence analysis at the sentence level. Finally, the judgment results from both levels are fused to obtain the recognition result of the AI-generated text.

[0004] However, existing technologies typically rely on encoder models pre-trained based on masked language modeling tasks. Their core capability lies in predicting masked words based on context. This leads to an overemphasis on surface-level word co-occurrence and collocation patterns when extracting semantic features from text, making it difficult to fully capture the deep features of AI-generated text with complex patterns. Consequently, the accuracy of existing AIGC text detection is relatively low. Summary of the Invention

[0005] This application provides a text detection method, apparatus, device, medium, and program product to improve the accuracy of text detection.

[0006] In a first aspect, embodiments of this application provide a text detection method, the method comprising:

[0007] Get the perplexity level corresponding to the first text;

[0008] According to the preset masking probability, the first text is masked to obtain the first masked text. The preset masking probability is the probability of masking each word in the first text.

[0009] Based on the above perplexity levels and the first mask text, construct text prompt words;

[0010] Based on the above text prompts, a first probability and a second probability are determined. The first probability is the prediction probability that the first text is manually written, and the second probability is the prediction probability that the first text is generated by AI.

[0011] Based on the first probability and the second probability mentioned above, the text generation method of the first text is determined, including manual writing or AI generation.

[0012] The technical solution provided in this application brings at least the following beneficial effects: Since it is possible to obtain the perplexity level corresponding to the first text, and to mask the first text according to the preset masking probability to obtain the first mask text, and then to construct text prompt words based on the perplexity level and the first mask text, and to determine the first probability that the first text is manually written and the second probability that it is generated by AI based on the text prompt words, and finally to determine the text generation method of the first text based on the first probability and the second probability, compared with the prior art, by introducing the perplexity level as a global statistical feature prior, and performing low-probability masking on the text in the inference stage to destroy the surface word co-occurrence pattern, the complementary fusion of global statistical features and deep semantic features is achieved, breaking through the inherent limitation of the existing technology that it is difficult to capture the deep abnormal features of AI-generated text, thereby improving the accuracy of AIGC text detection.

[0013] One possible implementation method for obtaining the perplexity level corresponding to the first text includes:

[0014] Retrieve the text content of the first text with a preset character length;

[0015] Calculate the mean cross-entropy of the above text content to obtain the first perplexity;

[0016] Based on the first mapping relationship, the perplexity level corresponding to the first perplexity is determined. The first mapping relationship is a mapping relationship between at least one perplexity level and at least one perplexity value range.

[0017] Another possible implementation, based on the first probability and the second probability, determines the text generation method of the first text, including:

[0018] Normalize the first probability and the second probability mentioned above;

[0019] The ratio of the normalized first probability to the first sum is determined as the target probability, which is the sum of the normalized first probability and the normalized second probability.

[0020] If the target probability is greater than a preset probability threshold, the first text is determined to be generated by AI.

[0021] If the target probability is less than or equal to the preset probability threshold, the first text is determined to be manually written.

[0022] Another possible implementation involves masking the first text according to a preset mask probability to obtain the first masked text, including:

[0023] Based on k different random seeds, each word in the first text is replaced with a mask marker k times with the preset mask probability to obtain k different first mask texts, where k is a positive integer;

[0024] The above-mentioned ratio of the normalized first probability to the first sum is determined as the target probability, including:

[0025] For each first masked text, the ratio of the normalized first probability to the first sum is determined as the first candidate probability, thus obtaining k candidate probabilities corresponding to the k first masked texts; the first sum is the sum of the normalized first probability and the normalized second probability corresponding to the first masked text.

[0026] The target probability is determined based on the median and mean of the k second candidate probabilities mentioned above.

[0027] Another possible implementation, before masking the first text according to a preset mask probability to obtain the first masked text, includes the following:

[0028] If the character length of the first text is greater than a preset length threshold, the first text is segmented according to the delimiters in the first text to obtain N text blocks, where N is a positive integer;

[0029] Each word in the above text block is replaced with a mask marker with a preset mask probability to obtain the second mask text corresponding to each text block;

[0030] The first mask text mentioned above includes the second mask text corresponding to each of the above text blocks.

[0031] Another possible implementation is that the aforementioned perplexity levels include the first perplexity level corresponding to each text block; the method also includes:

[0032] Based on the first perplexity level and the second mask text corresponding to each text block, N first text prompt words are constructed, and each first text prompt word corresponds to a text block;

[0033] Based on the first text prompt word corresponding to each text block, the first probability and the second probability corresponding to each text block are obtained; the first probability corresponding to a text block is the prediction probability that the text block is manually written, and the second probability corresponding to a text block is the prediction probability that the text block is generated by AI.

[0034] Based on the first and second probabilities corresponding to each of the above text blocks, the text generation method corresponding to each of the above text blocks is determined.

[0035] Another possible implementation, after obtaining the first probability and second probability corresponding to each text block based on the first text prompt word corresponding to each text block, the method further includes:

[0036] Based on the N first probabilities and N second probabilities corresponding to the N text blocks mentioned above, determine N second candidate probabilities;

[0037] Based on the above N second candidate probabilities, the target probability is determined, and based on the above target probability and the above preset probability threshold, the above target generation method is determined.

[0038] Secondly, embodiments of this application provide a text detection device, including:

[0039] The acquisition module is used to obtain the perplexity level corresponding to the first text.

[0040] The processing module is used to perform masking processing on the first text according to a preset masking probability to obtain the first masked text. The preset masking probability is the probability of masking each word in the first text.

[0041] The first construction module is used to construct text prompt words based on the above-mentioned perplexity level and the above-mentioned first mask text;

[0042] The first determining module is used to determine a first probability and a second probability based on the aforementioned text prompt words, wherein the first probability is the predicted probability that the aforementioned first text is manually written, and the aforementioned second probability is the predicted probability that the aforementioned first text is generated by AI.

[0043] The second determining module is used to determine the text generation method of the first text based on the first probability and the second probability, wherein the text generation method includes manual writing or AI generation.

[0044] Another possible implementation, the second determining module mentioned above, is specifically used for:

[0045] Normalize the first probability and the second probability mentioned above;

[0046] The ratio of the normalized first probability to the first sum is determined as the target probability, which is the sum of the normalized first probability and the normalized second probability.

[0047] If the target probability is greater than a preset probability threshold, the first text is determined to be generated by AI.

[0048] If the target probability is less than or equal to the preset probability threshold, the first text is determined to be manually written.

[0049] Another possible implementation is that the above processing module is specifically used for:

[0050] Based on k different random seeds, each word in the first text is replaced with a mask marker k times with the preset mask probability to obtain k different first mask texts, where k is a positive integer;

[0051] The above-mentioned ratio of the normalized first probability to the first sum is determined as the target probability, including:

[0052] For each first masked text, the ratio of the normalized first probability to the first sum is determined as the first candidate probability, thus obtaining k candidate probabilities corresponding to the k first masked texts; the first sum is the sum of the normalized first probability and the normalized second probability corresponding to the first masked text.

[0053] The target probability is determined based on the median and mean of the k second candidate probabilities mentioned above.

[0054] One possible implementation is that the above-mentioned acquisition module is specifically used for:

[0055] Retrieve the text content of the first text with a preset character length;

[0056] Calculate the mean cross-entropy of the above text content to obtain the first perplexity;

[0057] Based on the first mapping relationship, the perplexity level corresponding to the first perplexity is determined. The first mapping relationship is a mapping relationship between at least one perplexity level and at least one perplexity value range.

[0058] Another possible implementation, in which the above processing module is specifically used for:

[0059] Based on k different random seeds, each word in the first text is replaced with a mask marker k times with the preset mask probability to obtain k different first mask texts, where k is a positive integer.

[0060] Another possible implementation is that the first determining module mentioned above is specifically used for:

[0061] Normalize the first probability and the second probability corresponding to any mask text among the k first mask texts respectively to obtain the normalized first probability and the normalized second probability corresponding to any of the mask texts.

[0062] The ratio of the normalized first probability to the first sum is determined as the first candidate probability. The first sum is the sum of the normalized first probability and the normalized second probability. Similarly, the first candidate probabilities corresponding to the remaining first mask texts in the k first mask texts are calculated to obtain the k candidate probabilities.

[0063] The target probability is determined based on the median and mean of the k second candidate probabilities.

[0064] If the target probability is greater than a preset probability threshold, the first text is determined to be generated by AI.

[0065] If the target probability is less than or equal to the preset probability threshold, the first text is determined to be manually written.

[0066] Another possible implementation, the text detection device provided in this application embodiment further includes:

[0067] The segmentation module is used to segment the first text according to the delimiters in the first text when the character length of the first text is greater than a preset length threshold, so as to obtain N text blocks, where N is a positive integer;

[0068] The replacement module is used to replace each word in the above text block with a mask marker with a preset mask probability to obtain the second mask text corresponding to each text block;

[0069] The first mask text mentioned above includes the second mask text corresponding to each of the above text blocks.

[0070] Another possible implementation is that the aforementioned perplexity level includes the first perplexity level corresponding to each text block; the text detection device provided in this application embodiment further includes:

[0071] The second construction module is used to construct N first text prompt words based on the first perplexity level and the second mask text corresponding to each of the above text blocks, with each first text prompt word corresponding to a text block;

[0072] The third determining module is used to obtain the first probability and the second probability corresponding to each text block based on the first text prompt word corresponding to each text block; the first probability corresponding to a text block is the predicted probability that the text block is manually written, and the second probability corresponding to a text block is the predicted probability that the text block is generated by AI.

[0073] The fourth determining module is used to determine the text generation method corresponding to each of the above text blocks based on the first probability and the second probability corresponding to each of the above text blocks.

[0074] Another possible implementation, the text detection device provided in this application embodiment further includes:

[0075] The fifth determining module is used to determine N second candidate probabilities based on the N first probabilities and N second probabilities corresponding to the N text blocks mentioned above;

[0076] The sixth determining module is used to determine the target probability based on the above N second candidate probabilities, and to determine the above target generation method based on the above target probability and the above preset probability threshold.

[0077] Thirdly, this application provides an electronic device comprising: a processor and a memory; the memory stores a program or instructions executable on the processor, wherein the program or instructions, when executed by the processor, implement the method of the first aspect described above.

[0078] Fourthly, this application provides a readable storage medium on which a program or instructions are stored, which, when executed by a computer, implement the method of the first aspect described above.

[0079] Fifthly, this application provides a computer program product stored in a storage medium, which, when executed by a computer, implements the method described in the first aspect.

[0080] In a sixth aspect, embodiments of this application provide a chip including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0081] The beneficial effects of the second to sixth aspects mentioned above are described in the corresponding description of the first aspect and will not be repeated here. Attached Figure Description

[0082] Figure 1 A schematic diagram of the network architecture for a text detection method application provided in this application embodiment;

[0083] Figure 2 A flowchart illustrating a text detection method provided in an embodiment of this application;

[0084] Figure 3 A flowchart illustrating another text detection method provided in an embodiment of this application;

[0085] Figure 4 A flowchart illustrating yet another text detection method provided in this application embodiment;

[0086] Figure 5 A flowchart illustrating yet another text detection method provided in this application embodiment;

[0087] Figure 6 A flowchart illustrating yet another text detection method provided in this application embodiment;

[0088] Figure 7 A flowchart illustrating yet another text detection method provided in this application embodiment;

[0089] Figure 8 A flowchart illustrating yet another text detection method provided in this application embodiment;

[0090] Figure 9 This application provides a schematic diagram of the overall architecture of a text detection method according to an embodiment of the present application.

[0091] Figure 10 This is a schematic diagram of the structure of a text detection device provided in an embodiment of this application;

[0092] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0093] The text detection method, apparatus, equipment, medium, and program products provided in this application will now be described in detail with reference to the accompanying drawings.

[0094] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0095] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0096] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0097] In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0098] The text detection method, apparatus, device, medium, and program product provided in this application embodiment can be applied to AIGC text detection scenarios.

[0099] In existing technologies, a pre-trained language model encoder is used to extract semantic representations of text. Then, a two-branch network structure is constructed. One branch integrates global and local features at the document level to make an overall judgment, while the other branch performs sentence-by-sentence analysis at the sentence level. Finally, the judgment results at the two levels are fused to obtain the recognition result of AI-generated text.

[0100] However, existing technologies typically rely on encoder models pre-trained based on masked language modeling tasks. Their core capability lies in predicting masked words based on context. This leads to an overemphasis on surface-level word co-occurrence and collocation patterns when extracting semantic features from text, making it difficult to fully capture the deep features of AI-generated text with complex patterns. Consequently, the accuracy of existing AIGC text detection is relatively low.

[0101] To address the aforementioned technical problems, this application provides a text detection method, apparatus, device, medium, and program product. By acquiring the perplexity level corresponding to the first text and masking it according to a preset masking probability to obtain a first masked text, and then constructing text prompts based on the perplexity level and the first masked text, determining a first probability that the first text was manually written and a second probability that it was AI-generated based on the text prompts, and finally determining the text generation method of the first text based on the first and second probabilities, compared with existing technologies, by introducing the perplexity level as a priori global statistical feature and performing low-probability masking on the text during the inference stage to disrupt the surface word co-occurrence pattern, complementary fusion of global statistical features and deep semantic features is achieved. This overcomes the inherent limitation of existing technologies in capturing deep abnormal features of AI-generated text, thereby improving the accuracy of AIGC text detection.

[0102] The text detection method, apparatus, device, medium, and program products provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0103] Figure 1 The network architecture of a text detection method provided in an embodiment of this application is illustrated. For example... Figure 1 As shown, the network architecture includes a text detection device 101 and a terminal device 102. The text detection device 101 and the terminal device 102 are interconnected.

[0104] In some embodiments, the text detection device 101 may be a server, a computer, or a processor or processing unit within a server or computer. The server may be a single server or a server cluster consisting of multiple servers. It should be noted that the specific device form of the text detection device 101 is not limited in the embodiments of this application. Figure 1 The text detection device 101 is used as an example of a single server.

[0105] In some embodiments, the terminal device may be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and the embodiments of this application do not specifically limit it. Figure 1 The example shown is a mobile phone, with terminal device 102 as an example.

[0106] In some embodiments, the text detection device 101 responds to a text detection instruction sent by the terminal device 102 to perform the following actions: obtaining the perplexity level corresponding to the first text; masking the first text according to a preset masking probability to obtain a first masked text, wherein the preset masking probability is the probability of masking each word in the first text; constructing text prompt words based on the perplexity level and the first masked text; determining a first probability and a second probability based on the text prompt words, wherein the first probability is the predicted probability that the first text is manually written and the second probability is the predicted probability that the first text is AI-generated; determining the text generation method of the first text based on the first probability and the second probability, wherein the text generation method includes manual writing or AI generation, and sending the text generation method to the terminal device 102; the terminal device 102 is used to send a text detection instruction to the text detection device 101 and receive the text generation method sent by the text detection device 101.

[0107] It should be noted that the network architecture described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As network architectures evolve, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0108] See Figure 2 This is a flowchart illustrating a text detection method provided in an embodiment of this application. Figure 2 As shown, the text detection method provided in this application embodiment can be implemented by the above-mentioned text detection device, specifically including the following steps 201 to 205.

[0109] Step 201: The text detection device obtains the perplexity level corresponding to the first text.

[0110] In some embodiments, the text detection device uses a general large language model to calculate the sequence cross-entropy of the first text and converts it into perplexity, then maps it to a preset discrete level range to obtain the corresponding perplexity level.

[0111] In some embodiments, perplexity is a quantitative index calculated based on the sequence cross-entropy of a language model, used to measure how difficult it is for the model to predict the text. The lower the value, the more the text conforms to the statistical laws of general language.

[0112] In some embodiments, the perplexity level is a classification label obtained by dividing continuous perplexity values ​​into multiple semantic discrete intervals, which is used to intuitively represent the overall predictability level of the text.

[0113] Step 202: The text detection device performs masking processing on the first text according to a preset masking probability to obtain the first masked text. The preset masking probability is the probability of masking each word in the first text.

[0114] In some embodiments, the text detection device performs a masking operation on each segmentation unit of the first text with a preset independent probability, replacing the segmentation units that meet the conditions with the [MASK] tag, and can generate multiple different masked texts through different random seeds.

[0115] In some embodiments, the preset mask probability is a probability threshold for independently masking each word segmentation unit in the text during the inference phase.

[0116] Step 203: The text detection device constructs text prompt words based on the above-mentioned confusion level and the above-mentioned first mask text.

[0117] In some embodiments, the text prompt is a string that integrates the perplexity level of the text to be detected and the mask text information in a preset format, and is used to guide the general large language model to perform the AIGC text detection task.

[0118] In some embodiments, the text detection device concatenates the perplexity level label of the first text with the first mask text according to a preset format to construct a unified detection prompt word, i.e., a text prompt word, which includes global statistical features and local semantic features.

[0119] Step 204: The text detection device determines a first probability and a second probability based on the aforementioned text prompts. The first probability is the predicted probability that the first text was written manually, and the second probability is the predicted probability that the first text was generated by AI.

[0120] In some embodiments, the text detection device inputs the constructed detection prompt words into a general large language model, and after model inference, outputs two normalized prediction probabilities: one for the first text being manually written and the other for AI-generated.

[0121] Step 205: The text detection device determines the text generation method of the first text based on the first probability and the second probability mentioned above. The text generation method includes manual writing or AI generation.

[0122] In some embodiments, the text detection device calculates the target probability based on the first probability and the second probability, and determines the final text generation method of the first text based on the relationship between the target probability and a preset probability threshold.

[0123] In some embodiments, the method of determining the text generation of the first text based on the first probability and the second probability may also be implemented in the following ways, but not limited to: normalizing the first probability and the second probability to obtain normalized first probability and normalized second probability; determining the ratio of the normalized first probability to the first sum as the target probability; the first sum as the sum of the normalized first probability and the normalized second probability; determining that the first text is AI-generated when the target probability is greater than a preset probability threshold; and determining that the first text is manually written when the target probability is less than or equal to the preset probability threshold.

[0124] In some embodiments, combined with Figure 2 ,like Figure 3 As shown, step 201 above can be implemented through steps 201a to 201c.

[0125] Step 201a: The text detection device acquires the text content of the first text with a preset character length.

[0126] In some embodiments, the preset character length is a pre-defined fixed token quantity threshold, which is used to unify the perplexity calculation standard for texts of different lengths and eliminate the interference of text length differences on the detection results.

[0127] In some embodiments, the text detection device extracts a continuous text content of a preset character length from the starting position of the first text. If the total length of the first text is less than the preset character length, the entire text is used directly for subsequent calculations.

[0128] Step 201b: The text detection device calculates the mean cross-entropy of the above text content to obtain the first perplexity.

[0129] In some embodiments, the mean cross-entropy is the arithmetic mean of the negative log probabilities of each token in the text content, reflecting the average prediction difficulty of a general large language model for the entire text sequence.

[0130] In some embodiments, the text detection device substitutes the calculated mean cross-entropy into the natural exponential function for transformation to obtain the first perplexity, thereby realizing the mapping from information content indicators to predictability indicators.

[0131] In some embodiments, the first perplexity is a quantitative index calculated based on the mean of cross-entropy, and the lower the value, the more the text conforms to the statistical distribution pattern of a common language.

[0132] Step 201c: The text detection device determines the perplexity level corresponding to the first perplexity according to the first mapping relationship, wherein the first mapping relationship is a mapping relationship between at least one perplexity level and at least one perplexity value range.

[0133] In some embodiments, the first mapping relationship is a pre-configured discretized mapping table, where each perplexity level corresponds to a continuous perplexity value range, and the range boundary can be dynamically adjusted according to the accuracy requirements of different detection scenarios.

[0134] In some embodiments, the text detection device matches the first perplexity with each numerical interval in the first mapping relationship one by one to determine the semantic perplexity level label corresponding to the interval to which it belongs.

[0135] In this way, by standardizing the text length calculation, quantifying the global language statistical characteristics, and converting them into semantic level labels, a standardized, interpretable, and reliable global statistical feature foundation that is not affected by differences in text length is provided for subsequent AIGC text detection.

[0136] The text detection method provided in this application can obtain the perplexity level corresponding to the first text, and perform masking processing on the first text according to a preset masking probability to obtain the first mask text. Then, based on the perplexity level and the first mask text, text prompt words are constructed. Based on the text prompt words, the first probability that the first text is manually written and the second probability that it is generated by AI are determined. Finally, based on the first probability and the second probability, the text generation method of the first text is determined. Compared with the prior art, by introducing the perplexity level as a global statistical feature prior and performing low-probability masking processing on the text during the inference stage to destroy the surface word co-occurrence pattern, the complementary fusion of global statistical features and deep semantic features is achieved. This overcomes the inherent limitation of the prior art in capturing the deep abnormal features of AI-generated text, thereby improving the accuracy of AIGC text detection.

[0137] In some embodiments, combined with Figure 2 ,like Figure 4 As shown, step 205 above can be implemented through steps 205a to 205d.

[0138] Step 205a: The text detection device normalizes the first probability and the second probability mentioned above.

[0139] In some embodiments, the normalization process uses a linear scaling method to map the original probability values ​​to the [0,1] interval, eliminating the problem of inconsistent model output scales between different mask samples and ensuring the fairness of subsequent probability fusion.

[0140] In some embodiments, the first probability is the probability that the corresponding first mask text is predicted to be the original output written by a human, and the second probability is the probability that it is predicted to be the original output generated by AI. After normalization, the sum of the two approaches 1.

[0141] Step 205b: The text detection device determines the ratio of the normalized first probability to the first sum as the target probability, where the first sum is the sum of the normalized first probability and the normalized second probability.

[0142] In some embodiments, a reference benchmark is constructed by summing two types of normalized probabilities, and the candidate probability in the form of relative confidence is calculated to unify the judgment calculation standard for different masked samples.

[0143] In some embodiments, the influence of model output bias is eliminated by relying on ratio calculation, so that the candidate probability can accurately reflect the credibility of the text bias towards the corresponding generated category.

[0144] Step 205c: If the target probability is greater than a preset probability threshold, the text detection device determines that the first text is generated by AI.

[0145] In some embodiments, the preset probability threshold is a pre-defined classification decision boundary, which can be dynamically adjusted according to the tolerance for false positives in different detection scenarios.

[0146] In some embodiments, when the target probability is greater than a preset probability threshold, it indicates that the comprehensive detection result of multiple independent mask enhancements is more likely to indicate that the first text was generated by AI, thereby making a corresponding classification decision.

[0147] In some embodiments, the preset probability threshold can be determined by performing grid search optimization on a labeled validation set to balance the precision and recall of detection.

[0148] Step 205d: If the target probability is less than or equal to the preset probability threshold, the text detection device determines that the first text was written manually.

[0149] In some embodiments, when the target probability is less than or equal to a preset probability threshold, it indicates that the comprehensive detection result of multiple independent mask enhancements is more likely to indicate that the first text was written by a human, thereby making a corresponding classification decision.

[0150] In some embodiments, for boundary samples where the difference between the target probability and the preset probability threshold is less than a preset value, a secondary detection process can be triggered to further verify the sample by increasing the number of masking operations or using a higher precision model.

[0151] In some embodiments, the secondary detection results of boundary samples will overwrite the initial detection results to ensure the detection accuracy of high-risk samples.

[0152] In this way, by normalizing the probability to unify the output standard, and then converting it to obtain the standardized candidate probability, the category determination is completed based on the fixed threshold. This effectively unifies the determination criteria, reduces the determination error caused by the difference in model output, and improves the accuracy and consistency of the text generation method recognition results.

[0153] In some embodiments, combined with Figure 4 ,like Figure 5 As shown, step 202 above can be implemented through step 202a as follows.

[0154] Step 202a: The text detection device, based on k different random seeds, replaces each word in the first text with a mask marker k times with the preset mask probability, to obtain k different first mask texts, where k is a positive integer.

[0155] In some embodiments, k is typically a positive integer between 5 and 10, and the value of k can be dynamically adjusted according to the detection accuracy requirements and computing resource limitations.

[0156] In some embodiments, each random seed corresponds to a completely independent masking process, and the mask position sequences generated by different seeds are independent of each other and have no statistical correlation, ensuring that the resulting k first mask texts have sufficient diversity.

[0157] In some embodiments, the text detection device independently generates a uniformly distributed random number between 0 and 1 for each word segmentation unit in the first text, and replaces the random number with a [MASK] tag when the random number is less than or equal to a preset mask probability.

[0158] In some embodiments, the k different random seeds are k distinct integers used to initialize the pseudo-random number generator to generate k sets of independent mask position sequences.

[0159] In some embodiments, the independent masking operation is an operation that makes a masking decision for each word segmentation unit in the text separately, and the masking result of a single word segmentation unit does not affect other word segmentation units.

[0160] In some embodiments, the first masked text is a text variant obtained by performing an independent masking operation on the original first text, wherein some word segmentation units are replaced with [MASK] tags, while the remaining word segmentation units remain unchanged.

[0161] In this way, by obtaining multiple different variations of the same original text through multiple independent masking enhancements, the random fluctuations of a single masking operation can be effectively offset, significantly improving the stability and robustness of subsequent AIGC text detection results.

[0162] The above step 205b can be implemented through the following steps 205b1 and 205b2.

[0163] Step 205b1: For each first masked text, the text detection device determines the ratio of the normalized first probability to the first sum as the first candidate probability, thereby obtaining k candidate probabilities corresponding to the k first masked texts; the first sum is the sum of the normalized first probability and the normalized second probability corresponding to the first masked text.

[0164] In some embodiments, the first sum is the sum of the normalized probabilities of the two classes corresponding to a single first mask text. The absolute probability is converted into relative classification confidence by calculating the ratio, thereby eliminating the influence of the overall offset of the model output on the detection result.

[0165] Step 205b2: The text detection device determines the target probability based on the median and mean of the above k second candidate probabilities.

[0166] In some embodiments, by combining the median and the mean, the median is used to resist the interference of a single anomaly detection result, while the mean reflects the overall trend of multiple detections, thus achieving a more robust probability fusion.

[0167] In some embodiments, a weighted fusion method is used to calculate the target probability, for example, assigning a weight of 0.6 to the median and a weight of 0.4 to the mean, and the weight allocation can be dynamically adjusted according to the actual detection effect.

[0168] In some embodiments, when k is even, the median is taken as the arithmetic mean of the two middle first candidate probabilities to ensure the accuracy and consistency of the median calculation.

[0169] Thus, by normalizing and calibrating the detection results of multiple independent mask enhancements, transforming relative confidence, and robustly fusing multiple statistics, the random fluctuations and outlier interference of a single mask operation are effectively offset, significantly improving the stability, reliability, and robustness of AIGC text detection results.

[0170] In some embodiments, combined with Figure 2 ,like Figure 6 As shown, before step 202 above, the text detection method provided in this application embodiment may further include the following steps 301 to 302.

[0171] Step 301: When the character length of the first text is greater than a preset length threshold, the text detection device performs text segmentation on the first text according to the delimiters in the first text to obtain N text blocks, where N is a positive integer.

[0172] In some embodiments, the preset length threshold is matched with the maximum context window length of the general large language model, typically set to 512 or 1024 tokens, to ensure that each text block can be processed completely and efficiently by the model.

[0173] In some embodiments, the separators include semantic delimiters such as periods, question marks, exclamation marks, semicolons, and paragraph newlines, which are preferentially used to divide at the boundaries of complete semantic units to avoid disrupting the internal logic of sentences or paragraphs.

[0174] In some embodiments, when the length of a single segmented text block is still greater than a preset length threshold, the text block is further segmented at the positions of secondary delimiters such as commas and pauses until the length of all text blocks meets the requirements.

[0175] In some embodiments, when the total length of the first text is less than or equal to a preset length threshold, no segmentation operation is performed, and the entire first text is directly entered into the subsequent processing flow as a text block.

[0176] Step 302: The text detection device replaces each word in the above text block with a mask marker with a preset mask probability to obtain the second mask text corresponding to each text block.

[0177] In some embodiments, the first mask text includes the second mask text corresponding to each of the above text blocks.

[0178] In some embodiments, a masking operation is performed independently for each text block, and each word segmentation unit in each text block is replaced with a [MASK] tag with a preset independent probability. The masking processes of different text blocks are independent of each other and do not interfere with each other.

[0179] In some embodiments, the preset mask probability adopts a low probability of up to 4%, which moderately disrupts the surface word co-occurrence pattern while preserving the core semantic information of each text block to the maximum extent.

[0180] In some embodiments, at least one second mask text is generated for each text block. All second mask texts are input into the model for inference, and the detection results of all text blocks are finally fused to obtain the overall generation method of the first text.

[0181] In some embodiments, when the multiple masking enhancement mechanism is enabled, the text detection device generates k corresponding second mask texts for each text block using k different random seeds, thereby ensuring the stability of the detection results for each text block.

[0182] In some embodiments, a text block is a plurality of independent semantic units obtained by dividing a long text according to semantic boundaries, and the length of each text block does not exceed a preset length threshold.

[0183] In this way, by segmenting long texts into multiple independent text blocks that adapt to the model context window at the semantic boundary and performing masking processing on each block, the problem of long texts not being able to be detected completely due to the limitation of the context length of large language models is solved, while maintaining the semantic integrity of each text block, which significantly improves the accuracy and reliability of long text AIGC detection.

[0184] In some embodiments, combined with Figure 6 ,like Figure 7 As shown, after step 302 above, the perplexity level includes the first perplexity level corresponding to each text block. The text detection method provided in this application embodiment may also include the following steps 401 to 403.

[0185] Step 401: The text detection device constructs N first text prompt words based on the first perplexity level and the second mask text corresponding to each text block, with each first text prompt word corresponding to a text block.

[0186] In some embodiments, each first text prompt word adopts a preset format that is completely consistent with short text detection. The text detection device concatenates the perplexity level label of the corresponding text block with the second mask text to ensure the consistency of the long text and short text detection processes.

[0187] In some embodiments, the text detection device independently constructs a unique first text cue word for each text block. The cue word construction process for different text blocks is independent of each other and does not interfere with each other, supporting subsequent parallel inference processing.

[0188] In some embodiments, when the multiple masking enhancement mechanism is enabled, each text block is constructed with k first text prompt words, each corresponding to k different second mask texts of the text block.

[0189] Step 402: Based on the first text prompt word corresponding to each text block, the text detection device obtains the first probability and the second probability corresponding to each text block; the first probability corresponding to a text block is the predicted probability that the text block is manually written, and the second probability corresponding to a text block is the predicted probability that the text block is generated by AI.

[0190] In some embodiments, the text detection device inputs each first text prompt word into a general large language model for independent reasoning, and the model outputs two normalized predicted probabilities for the corresponding text block: one manually written and one generated by AI.

[0191] In some embodiments, the reasoning process for all text blocks can be executed in parallel, making full use of computing resources and significantly shortening the overall detection time for long texts.

[0192] In some embodiments, when each text block corresponds to multiple first text prompt words, the text detection device first fuses the multiple inference results of the text block, and then obtains the final first probability and second probability of the text block.

[0193] In some embodiments, the calculation methods for the first and second probabilities corresponding to text blocks are exactly the same as those for short text detection, ensuring consistency in detection standards.

[0194] Step 403: The text detection device determines the text generation method corresponding to each text block based on the first probability and the second probability corresponding to each text block.

[0195] In some embodiments, the text detection device compares the magnitude of the first probability and the second probability corresponding to each text block, and determines the category corresponding to the text block as the text generation method of the text block with the higher probability value.

[0196] In some embodiments, a relative confidence calculation method consistent with short text detection is adopted. The text detection device first uses the ratio of the first probability to the first sum as the candidate probability, and then compares it with a preset probability threshold to determine the text generation method.

[0197] In some embodiments, the text detection device records the text generation method and corresponding confidence value of each text block, providing a basis for subsequently determining the overall generation method of the entire first text.

[0198] In some embodiments, text blocks with a confidence level below a preset threshold can be marked as "uncertain" and a secondary detection process can be triggered to improve the accuracy of single-block detection.

[0199] In this way, by independently constructing prompt words, making independent inferences and judgments for each text block, the block-based parallel detection of long texts is realized. This not only solves the problem of the context length limitation of large language models, but also maintains the same accuracy standard as short text detection. At the same time, it provides fine-grained block-level detection results for the comprehensive judgment of the overall generation method of long texts.

[0200] In some embodiments, combined with Figure 7 ,like Figure 8As shown, after step 402 above, the text detection method provided in this application embodiment may further include the following steps 501 to 502.

[0201] Step 501: The text detection device determines N second candidate probabilities based on the N first probabilities and N second probabilities corresponding to the N text blocks mentioned above.

[0202] In some embodiments, the text detection device independently calculates the second candidate probability for each text block, which is calculated as the ratio of the first probability corresponding to the text block to the first sum of the text block, wherein the first sum is the sum of the first probability and the second probability corresponding to the text block.

[0203] In some embodiments, the second candidate probability calculation process for all text blocks can be executed in parallel to make full use of computing resources and further shorten the overall detection time for long texts.

[0204] Step 502: The text detection device determines the target probability based on the above N second candidate probabilities, and determines the target generation method based on the above target probability and the above preset probability threshold.

[0205] In some embodiments, a strategy consistent with the fusion of multiple masking enhancement results is adopted, combining the median and mean of the N second candidate probabilities to calculate the target probability, which both resists interference from individual anomalous text blocks and reflects the overall detection trend.

[0206] In this way, by robustly fusing the detection results of all blocks of the long text with multiple statistics, the comprehensive detection confidence of the entire long text is obtained. This not only makes full use of the fine-grained detection information of each text block, but also effectively resists the interference of a single abnormal text block on the overall result. This achieves the unification of detection standards for long text and short text, and significantly improves the accuracy and reliability of AIGC detection for long text.

[0207] The text detection method of this application will be described below through specific embodiments.

[0208] like Figure 9 As shown, the overall architecture of the text detection method provided in this application includes:

[0209] A four-stage technical architecture is adopted, consisting of "perplexity feature enhancement + data masking enhancement + instruction fine-tuning (SFT) + probabilistic ensemble decision". In the first stage, the original text input is acquired, and the mean cross-entropy of the first 512 characters of the text is calculated using a base model as the perplexity representation. Perplexity is then divided into seven levels (very small, relatively small, small, medium, large, relatively large, very large) according to the training set quantiles, and perplexity level labels are output. In the second stage, multiple text blocks are constructed for the same sample. Long texts are segmented into blocks of up to 1500 characters using delimiters, and overlapping regions of three words (or characters) are set to maintain coherence. Random masking enhancement is applied to these text blocks, and the enhanced text blocks are output. After this process, each sample corresponds to multiple text blocks. The training stage is used for corpus enhancement, with each text block serving as a training sample. The inference stage is used to perform multiple predictions on the same sample to improve accuracy. In the third stage, based on the text blocks, perplexity labels, and sample labels obtained in the first two stages, prompt words are written and SFT training data in question-and-answer format is constructed. Using the LLama Factory large-scale model fine-tuning framework, LoRA is used to fine-tune the SFT large-scale model. The 0 / 1 label samples are rebalanced, and the model weights corresponding to the loss convergence are selected based on the loss curve. The trained detection model is then output. In the fourth stage, the inference stage, the text to be detected is obtained, the Logits probability of the key tokens representing the category labels in the output is calculated, and Softmax normalization is performed to obtain the AI-generated prediction probability value. Multiple corpus augmentations are performed on the same sample to obtain multiple samples. The above prediction is performed on these multiple samples to obtain multiple prediction probability values. The mean and median of the multiple prediction probability values ​​are calculated, and the average of the two is used to obtain the final AI-generated probability value. The optimal classification threshold is obtained through a calibration set to determine whether it is AI-generated. The AI-generated probability value and the final label are output, where label = 1 indicates that the text is machine-generated, and label = 0 indicates that the text is human-written.

[0210] In some embodiments, to demonstrate the complete processing flow from raw text input to perplexity level label output, this application provides an exemplary description, namely... Figure 9 The stages of confusion characteristics shown include:

[0211] The mean cross-entropy of the first 512 characters of the text to be detected is calculated using the Qwen2.5-14B-Instruct pedestal model, and used as a representation of perplexity. The specific calculation process is as follows:

[0212] Load the pre-trained Qwen2.5-14B-Instruct model and tokenizer, perform word segmentation on the text to be detected, extract the first 512 characters, input the segmented token sequence into the model, and calculate the cross-entropy loss value. The calculation formula is as follows:

[0213] Formula (1)

[0214] Where M is the total number of valid tokens involved in the calculation, and V is the size of the large model vocabulary. For the i-th valid token position, the predicted Logit value is the j-th word in the vocabulary. : The real label index of the i-th valid token. This represents the softmax probability of the i-th token.

[0215] Using the cross-entropy loss value as a representation of perplexity, all samples in the training set are divided into 7 level intervals according to the quantiles of perplexity, with the boundary values ​​of each interval shown in Table 1. Samples outside the training set are classified according to the boundaries of the perplexity intervals determined within the training set to ensure feature consistency.

[0216] Table 1 shows the range of mean cross-entropy values ​​corresponding to the perplexity labels.

[0217]

[0218] In some embodiments, Figure 9 The data masking enhancement stage shown includes: For long texts exceeding 1500 characters, segmentation is performed according to the following rules to obtain text blocks: Delimiter priority: Segmentation is performed sequentially according to the delimiters [[".", "\n"], [":"], and [" "]]. Maximum length limit: The maximum length of each text block is 1500 characters. Overlapping area setting: Three word (or character) overlap areas are set between adjacent text blocks to maintain semantic coherence. Segmentation strategy: Segmentation is prioritized at periods and newlines to avoid cutting from the middle of the sentence.

[0219] Each segmented text block is randomly masked according to the following masking rules to obtain the masked text block. The specific masking rules are as follows: Training phase: Each word is replaced with "[MASK]" with a 15% probability. Inference phase: Each word is replaced with "[MASK]" with a 4% probability (to reduce information loss). Masking interval control: Once a word is replaced with "[MASK]", the next 5 consecutive words are not masked to avoid continuously disrupting the semantic structure. Multi-view augmentation: Each sample is augmented 5 times using a different random number seed to form multi-view training data.

[0220] In some embodiments, Figure 9 The fine-tuning phase of the instructions includes: constructing a Prompt template that includes role settings, workflow, and output format by combining the perplexity level labels of the samples and the masked text blocks, ensuring that the large model clearly defines the detection task requirements. Data balancing and checkpoint selection during training involve rebalancing the training set samples labeled 0 (human-written) and 1 (AI-generated) before training to ensure that the number of samples in both classes is consistent. Based on the LLama Factory large model fine-tuning framework, the Qwen2.5-14B-Instruct model is trained using SFT using the LoRA method. During training, the loss curve is monitored, and the model weights at which the training loss converges are selected as the final detection model.

[0221] In some embodiments, Figure 9 The inference decision and result integration phase shown includes: calculating the predicted probability of samples; for each text block to be detected, calculating the probability of the key token (0 or 1) representing the category label in the model output, the specific process is as follows:

[0222] The prompt word template is combined with the text to be detected and then input into the fine-tuned model to obtain the model's output. <label>The predicted probability value for the next token after " is 0 or 1. Extract the predicted probability values ​​corresponding to token IDs "0" and "1", and perform Softmax normalization on the two predicted probability values ​​to obtain the AI-generated predicted probability value. The probability calculation formula is as follows:

[0223] Formula (2)

[0224] Where P represents the probability that the text was generated by AI. and These represent the predicted probability values ​​for output labels 0 and 1, respectively.

[0225] Multi-view probability fusion is employed, and the optimal probability threshold for classification is determined based on the training set. Multiple augmentation results for the same sample are integrated probabilities. Specifically, the masking ratio during the inference stage is adjusted from 15% in training to 4% to reduce information loss and maintain detection accuracy. Five masking augmentations with different random number seeds are performed on the same sample, resulting in five predicted probability values. The mean and median of these five predicted probability values ​​are calculated, and the average of the mean and median is taken as the final AI-generated probability prediction value.

[0226] Formula (3)

[0227] The final result is generated based on the threshold: the final AI-generated probability value. If the value is greater than 0.998688, it is considered AI-generated text (label 1); otherwise, it is considered human-written text (label 0). This threshold is optimized in the training set and the same threshold is used in the inference phase.

[0228] In some embodiments, in addition to "perplexity binning based on the mean of cross-entropy", one or more statistical features such as information entropy, log-rank, word frequency distribution skewness, syntactic complexity, and punctuation distribution can be used in combination to represent the data, and the continuous values ​​can be discretized into level labels to input the detection model. The binning method is not limited to a fixed seven levels; equal-frequency binning, equal-width binning, clustering binning, or adaptive binning based on a calibration set can also be used to obtain the same discriminative ability as the original scheme.

[0229] In some embodiments, masking enhancement can also be replaced by synonym rewriting, back-translation enhancement, local perturbation enhancement, or deletion-insertion enhancement. The enhancement strength during the training and inference phases can be fixed parameters or dynamically adjusted according to text complexity. All of the above alternatives can construct multi-view samples while maintaining the semantic subject unchanged, thereby improving robustness and generalization.

[0230] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there is no conflict, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0231] As can be seen, the above mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0232] This application embodiment can divide the text detection device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. Optionally, the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0233] In some embodiments, this application also provides a text detection apparatus. The text detection apparatus may include one or more functional modules for implementing the text detection method of the above method embodiments.

[0234] For example, Figure 10 This is a schematic diagram of the structure of a text detection device provided in an embodiment of this application. Figure 10 As shown, the text detection device 900 includes:

[0235] The module 901 is used to acquire the perplexity level corresponding to the first text. The processing module 902 is used to mask the first text according to a preset mask probability to obtain the first mask text. The preset mask probability is the probability of masking each word in the first text. The first construction module 903 is used to construct text prompt words based on the perplexity level and the first mask text. The first determination module 904 is used to determine a first probability and a second probability based on the text prompt words. The first probability is the predicted probability that the first text is manually written, and the second probability is the predicted probability that the first text is AI-generated. The second determination module 905 is used to determine the text generation method of the first text based on the first probability and the second probability. The text generation method includes manual writing or AI generation.

[0236] The text detection device provided in this application can obtain the perplexity level corresponding to the first text, and perform masking processing on the first text according to a preset masking probability to obtain the first masked text. Then, based on the perplexity level and the first masked text, text prompt words are constructed. Based on the text prompt words, a first probability that the first text is manually written and a second probability that it is generated by AI are determined. Finally, based on the first probability and the second probability, the text generation method of the first text is determined. Compared with the prior art, by introducing the perplexity level as a global statistical feature prior and performing low-probability masking processing on the text during the inference stage to destroy the surface word co-occurrence pattern, the complementary fusion of global statistical features and deep semantic features is achieved. This overcomes the inherent limitation of the prior art in capturing the deep abnormal features of AI-generated text, thereby improving the accuracy of AIGC text detection.

[0237] In some embodiments, the acquisition module 901 is specifically used for:

[0238] Obtain the text content of the first text with a preset character length, calculate the average cross-entropy of the text content to obtain the first perplexity, and determine the perplexity level corresponding to the first perplexity based on the first mapping relationship. The first mapping relationship is a mapping relationship between at least one perplexity level and at least one perplexity value range.

[0239] In other embodiments, the second determining module 905 is specifically used to: normalize the first probability and the second probability, determine the ratio of the normalized first probability to the first sum as the target probability, the target probability being the sum of the normalized first probability and the normalized second probability, determine that the first text is generated by AI when the target probability is greater than a preset probability threshold, and determine that the first text is written manually when the target probability is less than or equal to the preset probability threshold.

[0240] Another possible implementation is that the processing module 902 is specifically used to: based on k different random seeds, replace each word in the first text with a mask marker k times using the preset mask probability to obtain k different first mask texts, where k is a positive integer; the ratio of the normalized first probability to the first sum is determined as the target probability, which includes: for each first mask text, determining the ratio of the normalized first probability to the first sum as the first candidate probability, obtaining k candidate probabilities corresponding to the k first mask texts; the first sum is the sum of the normalized first probability and the normalized second probability corresponding to the first mask text; and the target probability is determined based on the median and mean of the k second candidate probabilities.

[0241] Another possible implementation, the text detection device provided in this application embodiment further includes:

[0242] The segmentation module is used to segment the first text according to the delimiters in the first text when the character length of the first text is greater than a preset length threshold, so as to obtain N text blocks, where N is a positive integer.

[0243] The replacement module is used to replace each word in the above text block with a mask marker with a preset mask probability to obtain the second mask text corresponding to each text block.

[0244] The first mask text mentioned above includes the second mask text corresponding to each of the above text blocks.

[0245] Another possible implementation is that the aforementioned perplexity level includes the first perplexity level corresponding to each text block; the text detection device provided in this application embodiment further includes:

[0246] The second construction module is used to construct N first text prompt words based on the first perplexity level and the second mask text corresponding to each of the above text blocks, with each first text prompt word corresponding to a text block.

[0247] The third determining module is used to obtain the first probability and the second probability corresponding to each text block based on the first text prompt word corresponding to each text block; the first probability corresponding to a text block is the prediction probability that the text block is manually written, and the second probability corresponding to a text block is the prediction probability that the text block is generated by AI.

[0248] The fourth determining module is used to determine the text generation method corresponding to each of the above text blocks based on the first probability and the second probability corresponding to each of the above text blocks.

[0249] Another possible implementation, the text detection device provided in this application embodiment further includes:

[0250] The fifth determining module is used to determine N second candidate probabilities based on the N first probabilities and N second probabilities corresponding to the N text blocks mentioned above.

[0251] The sixth determining module is used to determine the target probability based on the above N second candidate probabilities, and to determine the above target generation method based on the above target probability and the above preset probability threshold.

[0252] It should be noted that the text detection device can implement all the processes implemented in the above method embodiments and achieve the same beneficial effects. To avoid repetition, it will not be described again here.

[0253] In the case where the functions of the integrated modules described above are implemented in hardware, this application provides a possible structural schematic diagram of the electronic device involved in the above embodiments. For example... Figure 11 As shown, the electronic device 90 includes: a processor 92, a communication interface 93, and a bus 94. Optionally, the electronic device 90 may also include a memory 91.

[0254] Processor 92 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0255] Communication interface 93 is used to connect with other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0256] The memory 91 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0257] In one possible implementation, the memory 91 can exist independently of the processor 92. The memory 91 can be connected to the processor 92 via a bus 94 and is used to store instructions or program code. When the processor 92 calls and executes the instructions or program code stored in the memory 91, it can implement the text detection method provided in the embodiments of this application.

[0258] In another possible implementation, memory 91 can also be integrated with processor 92.

[0259] Bus 94 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 94 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0260] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the service calling device can be divided into different functional modules to complete all or part of the functions described above.

[0261] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described text detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0262] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0263] This application also provides a readable storage medium storing a program or instructions that, when executed by a computer, implement the text detection method provided in the above embodiments. It is understood that all or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware; the readable storage medium can be any of the foregoing embodiments or memory; the readable storage medium can also be an external storage device of the service invocation device, such as a pluggable hard drive, Smart MediaCard (SMC), Secure Digital (SD) card, flash card, etc., equipped on the service invocation device. Further, the readable storage medium can include both internal storage units of the service invocation device and external storage devices. The readable storage medium is used to store the computer program and other programs and data required by the service invocation device. The readable storage medium can also be used to temporarily store data that has been output or will be output.

[0264] This application also provides a computer program product, which is stored in a storage medium and implements the text detection method provided in the above embodiments when the computer program product is executed by a computer.

[0265] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0266] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0267] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.< / label>

Claims

1. A text detection method, characterized in that, include: Get the perplexity level corresponding to the first text; The first text is masked according to a preset masking probability to obtain the first masked text. The preset masking probability is the probability of masking each word in the first text. Based on the perplexity level and the first mask text, construct text prompt words; Based on the text prompt words, a first probability and a second probability are determined. The first probability is the predicted probability that the first text is manually written, and the second probability is the predicted probability that the first text is generated by AI. Based on the first probability and the second probability, the text generation method of the first text is determined, and the text generation method includes manual writing or AI generation.

2. The text detection method according to claim 1, characterized in that, The step of obtaining the perplexity level corresponding to the first text includes: Obtain the text content of the first text with a preset character length; Calculate the mean cross-entropy of the text content to obtain the first perplexity; Based on the first mapping relationship, the perplexity level corresponding to the first perplexity is determined, wherein the first mapping relationship is a mapping relationship between at least one perplexity level and at least one perplexity value range.

3. The text detection method according to claim 1, characterized in that, Determining the text generation method of the first text based on the first probability and the second probability includes: Normalize the first probability and the second probability; The ratio of the normalized first probability to the first sum is determined as the target probability, wherein the target probability is the sum of the normalized first probability and the normalized second probability; If the target probability is greater than a preset probability threshold, the first text is determined to be AI-generated; If the target probability is less than or equal to the preset probability threshold, the first text is determined to be manually written.

4. The text detection method according to claim 1, characterized in that, The step of masking the first text according to a preset mask probability to obtain the first masked text includes: Based on k different random seeds, each word in the first text is replaced with a mask marker k times with the preset mask probability to obtain k different first mask texts, where k is a positive integer; Determining the ratio of the normalized first probability to the first sum as the target probability includes: For each first masked text, the ratio of the normalized first probability to the first sum is determined as the first candidate probability, thus obtaining k candidate probabilities corresponding to the k first masked texts; the first sum is the sum of the normalized first probability and the normalized second probability corresponding to the first masked text. The target probability is determined based on the median and mean of the k second candidate probabilities.

5. The text detection method according to claim 1, characterized in that, Before performing masking processing on the first text according to a preset masking probability to obtain the first masked text, the method further includes: If the character length of the first text is greater than a preset length threshold, the first text is segmented according to the delimiters in the first text to obtain N text blocks, where N is a positive integer; Each word in the text block is replaced with a mask marker with a preset mask probability to obtain the second mask text corresponding to each text block; The first mask text includes the second mask text corresponding to each text block.

6. The text detection method according to claim 5, characterized in that, The perplexity level includes a first perplexity level corresponding to each text block; the method further includes: Based on the first perplexity level and the second mask text corresponding to each text block, N first text prompt words are constructed, and each first text prompt word corresponds to a text block; Based on the first text prompt word corresponding to each text block, a first probability and a second probability corresponding to each text block are obtained; the first probability corresponding to a text block is the predicted probability that the text block is manually written, and the second probability corresponding to a text block is the predicted probability that the text block is generated by AI. Based on the first probability and the second probability corresponding to each text block, the text generation method corresponding to each text block is determined.

7. The text detection method according to claim 6, characterized in that, After obtaining the first probability and the second probability corresponding to each text block based on the first text prompt word corresponding to each text block, the method further includes: Based on the N first probabilities and N second probabilities corresponding to the N text blocks, determine N second candidate probabilities; Based on the N second candidate probabilities, the target probability is determined, and the target generation method is determined based on the target probability and the preset probability threshold.

8. A text detection device, characterized in that, include: The acquisition module is used to obtain the perplexity level corresponding to the first text. The processing module is used to perform masking processing on the first text according to a preset masking probability to obtain the first masked text, wherein the preset masking probability is the probability of masking each word in the first text. The first construction module is used to construct text prompt words based on the confusion level and the first mask text; The first determining module is used to determine a first probability and a second probability based on the text prompt words, wherein the first probability is the predicted probability that the first text is manually written, and the second probability is the predicted probability that the first text is generated by AI. The second determining module is used to determine the text generation method of the first text based on the first probability and the second probability, wherein the text generation method includes manual writing or AI generation.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the text detection method as described in any one of claims 1-7.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a computer, implement the text detection method as described in any one of claims 1-7.

11. A computer program product, characterized in that, The computer program product is stored in a storage medium, and when executed by a computer, the computer program product implements the text detection method as described in any one of claims 1-7.