Method for detecting generated text based on a maximum mean discrepancy of depth

By constructing a deep composite kernel network and introducing a Wild Bootstrap mechanism, the problems of insufficient capture of high-dimensional semantic differences and sequence correlation in generated text detection are solved, and the stability and generalization ability are improved under various training data conditions, thereby improving the accuracy and robustness of generated text detection.

CN122287591APending Publication Date: 2026-06-26SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-03-20
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing generated text detection methods are insufficient in capturing high-dimensional semantic differences, handling sentence sequence correlations within paragraphs, and in terms of stability and generalization ability in scenarios with limited or imbalanced training data, resulting in poor detection performance.

Method used

Paragraph-level generated text detection is modeled as a nonparametric two-sample test problem. A deep composite kernel network is constructed to integrate shallow syntactic features and deep semantic features of the text. A Wild Bootstrap mechanism is introduced to generate random weight sequences with autoregressive structures. The deep kernel network is trained by maximizing the unbiased estimation of test power, and the observation statistics are calculated and the detection results are output.

Benefits of technology

It significantly improves the ability to capture high-dimensional semantic differences in generated text detection, solves the problem of decreased test power caused by sequence correlation, enhances stability under limited or imbalanced training data and generalization ability to unknown generative models, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287591A_ABST
    Figure CN122287591A_ABST
Patent Text Reader

Abstract

This invention discloses a paragraph-level generated text detection method based on deep maximum mean difference (MDD) metric. The method includes: modeling generated text detection as a nonparametric two-sample test problem; designing a deep composite kernel to fuse shallow syntactic features and deep semantic features; training a deep kernel network with the goal of maximizing unbiased estimation of test power; in the testing phase, introducing a Wild Bootstrap mechanism to generate random weight sequences with autoregressive structures; performing a weighted quadratic transformation on the MMD kernel difference matrix; preserving paragraph sentence sequence correlation; accurately estimating the null distribution of the test statistic; and completing the detection by comparing the observed statistic with the test threshold. This method improves sensitivity to high-dimensional semantic differences, alleviates the problem of decreased test power caused by sequence correlation, and can efficiently distinguish between human and AI-generated text. It is suitable for generated text detection in scenarios such as news and academic writing.
Need to check novelty before this filing date? Find Prior Art