Paper AI generation rate detection system with modification auxiliary prompt

By extracting text features and combining them with contextual analysis, a sequence of modification instructions is generated, which solves the problem of insufficient identification of logical connections between paragraphs in existing technologies, and improves the accuracy and readability of paper detection.

CN121637334APending Publication Date: 2026-03-10CHANGSHA HUAHUA NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing plagiarism detection tools cannot effectively identify the logical connections between paragraphs, resulting in insufficient detection accuracy and making it difficult to accurately determine whether a paper was generated by an intelligent tool.

Method used

The module extracts text features and calculates the initial probability of generated content by obtaining specific paragraph probabilities, combined with the distribution characteristics of cross-domain terms; the module determines the paragraph probability sequence and analyzes the contextual consistency, and generates a sequence of modification instructions based on the judgment set; the module determines the overall generation ratio index using a weighted average method and outputs a complete report.

Benefits of technology

It significantly improves the logical consistency and professionalism of academic papers, enhances the accuracy and readability of detection, and provides intelligent support for academic writing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121637334A_ABST
    Figure CN121637334A_ABST
Patent Text Reader

Abstract

The invention discloses a paper AI generation rate detection system with a modification auxiliary prompt, which is characterized in that text features are extracted, paragraph probabilities are calculated, a paragraph sequence is adjusted in combination with context analysis, cross-domain terms and term habits are accurately judged and optimally adjusted for high-probability target paragraphs, and a modification instruction sequence is generated. And finally outputting a complete report containing the optimized text unit. Cross-domain term distribution is taken as a core index, and a weighted average method is fused to determine an overall generation proportion, so that the specialty and consistency of texts are ensured, and a business closed loop is realized. According to the method, the quality and readability of the paper text are remarkably improved, and intelligent support is provided for academic writing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of paper detection technology, and in particular discloses a paper AI generation rate detection system with modification assistance prompts. Background Technology

[0002] In academic research and education, the detection of originality in academic papers has always been a crucial issue. With the development of technology, the phenomenon of papers being generated using intelligent tools is increasingly prevalent, posing a serious challenge to academic integrity. How to accurately identify content in papers that may have been generated by intelligent tools and provide effective guidance for improvement to students and researchers has become an urgent problem to be solved. Research in this area is not only related to maintaining academic norms but also directly affects educational equity and the authenticity of knowledge innovation.

[0003] Currently, although some detection tools exist on the market, they often have significant limitations. Many methods, when analyzing papers, lack a comprehensive consideration of the content context, relying solely on surface features for judgment, which easily leads to misjudgments or omissions. Furthermore, these tools typically fail to deeply analyze the connections between different parts of the paper, ignoring the impact of the overall text structure and logic on the judgment results, resulting in insufficient detection accuracy and failing to meet practical needs.

[0004] A deeper technical challenge lies in how to fully consider the logical connections between different parts of a paper during the detection process to ensure the comprehensiveness and accuracy of the judgment. In particular, the content of each paragraph in a paper is not isolated; its expression, word choice, and the scope of knowledge it covers are often closely related to the preceding and following paragraphs. Analyzing only individual paragraphs may lead to biased conclusions due to a lack of contextual reference. For example, a paragraph in a paper might be suspected of being non-human-written because it uses multiple technical terms from different fields, but if the content of the preceding and following paragraphs provides a reasonable background explanation, this suspicion may be unfounded. How to incorporate the logical relationships between paragraphs during detection and make more realistic judgments based on this has become a technical bottleneck that urgently needs to be overcome.

[0005] Therefore, how to comprehensively consider the contextual relationships between paragraphs and accurately determine the generation source of each paragraph and the entire paper based on specific characteristics when identifying whether the content of a paper was generated by intelligent tools has become a key problem that this study urgently needs to solve. Summary of the Invention

[0006] This invention provides a paper AI generation rate detection system with modification assistance prompts, aiming to solve at least one defect in the above-mentioned prior art.

[0007] This invention relates to a paper AI generation rate detection system with modification assistance prompts, comprising: The paragraph-specific probability acquisition module is used to obtain the initial paragraph sequence from the paper text through the text analysis model, extract text features for each initial paragraph sequence and calculate the initial generated content probability index, and incorporate cross-domain terminology distribution features as the core judgment index to obtain the paragraph-specific probability. The paragraph probability sequence determination module is used to combine the specific probability of a paragraph with the text of the preceding and following paragraphs based on the integration attributes of adjacent paragraphs, and to analyze the consistency using the context analysis method to determine the adjusted paragraph probability sequence. The judgment set generation module is used to obtain cross-domain terminology and usage habit judgment indicators for target paragraphs in the adjusted paragraph probability sequence that exceed the threshold condition, and generate a judgment set based on the judgment set. The text unit acquisition module is used to construct a sequence of modification instructions based on a judgment set if the number of cross-domain terms in the target paragraph exceeds the average of the associated paragraphs, and then replace and adjust the text units according to the usage habits to obtain the adjusted text units. The overall generation ratio determination module is used to aggregate the overall probability from all adjusted text units and the probability sequence of the remaining paragraphs, and to determine the overall generation ratio by using a weighted average method to integrate the influence of cross-domain terms. The complete report output module is used to output a complete report containing detected reduced attributes based on the overall generation ratio indicators and the sequence of modification instructions for each target paragraph, covering the adjusted text units to achieve business closure.

[0008] Furthermore, the paragraph-specific probability acquisition module includes: The preliminary probability value acquisition unit is used to extract the initial paragraph sequence from the full text of the paper using a text segmentation tool. For the initial paragraph sequence, a text feature extraction tool is used to obtain word frequency and sentence structure features, calculate the preliminary content generation probability index, and obtain the preliminary probability value of the paragraph. The intermediate value determination unit is used to obtain the distribution density and domain relevance characteristics of cross-domain terms in a paragraph based on the preliminary probability value and the cross-domain term distribution feature database. The relevance characteristics are incorporated as the core judgment criteria to determine the intermediate value of a specific probability of the paragraph. The final value judgment unit is used to perform secondary calibration on the cross-domain term distribution characteristics in the paragraph through the cross-domain term weight adjustment tool if the intermediate value is lower than the preset threshold, obtain the adjusted feature weight, and judge the final value of the paragraph's specific probability.

[0009] Furthermore, the paragraph probability sequence determination module includes: The paragraph relevance judgment result acquisition unit is used to obtain at least one adjacent paragraph sequence from the target text through a text segmentation tool, extract semantic coherence and content coherence features from the adjacent paragraph sequence, and obtain preliminary paragraph relevance judgment results. The attribute fusion degree determination unit is used to determine the attribute fusion degree of adjacent paragraph sequences by comparing the text consistency and contextual logic features between adjacent paragraphs based on the preliminary paragraph correlation judgment results and using semantic comparison tools. The probability distribution difference acquisition unit is used to calibrate paragraph boundary values ​​and association weight ratios through a weight adjustment tool if the attribute fusion degree is lower than a preset threshold, and then obtain the adjusted probability distribution difference. The paragraph probability sequence judgment unit is used to perform a secondary comparison between the specific probability of a paragraph and the preceding and following related text, based on the adjusted probability distribution difference, using content integration tools, to determine the final paragraph probability sequence.

[0010] Furthermore, the judgment set generation module includes: The cross-domain terminology distribution set acquisition unit is used to obtain target paragraphs with higher than a preset threshold from the adjusted paragraph probability sequence through semantic extraction tools, and to extract cross-domain terminology and usage habit features from the obtained target paragraphs with higher than the preset threshold to obtain a preliminary cross-domain terminology distribution set. The cross-domain terminology relevance judgment index determination unit is used to determine the cross-domain terminology relevance judgment index by analyzing the contextual consistency characteristics of cross-domain terms and usage habits in the target paragraph based on the preliminary cross-domain terminology distribution set and using text comparison tools. The indicator set acquisition unit is used to calibrate the correlation between cross-domain terms and usage habits through a weight adjustment tool if the cross-domain term relevance judgment indicator is lower than a preset threshold, and to obtain the adjusted indicator set. The judgment unit based on the judgment set is used to compare and integrate cross-domain terminology and usage habits with the probability sequence of the target paragraph using content integration tools to determine the final judgment set based on the judgment set.

[0011] Furthermore, the text unit acquisition module includes: The adjustment requirement identifier generation unit is used to obtain cross-domain term quantity data from the target paragraph, compare it with the average value of related paragraphs using a text comparison tool, and generate an adjustment requirement identifier if it exceeds a preset threshold. The matching degree index acquisition unit is used to obtain the language habit characteristics of the target paragraph based on the adjustment requirement identifier, and to obtain the matching degree index by analyzing the contextual consistency through semantic extraction tools. The modification instruction sequence forming unit is used to extract adaptation rules from a pre-established language habit library based on the matching degree index and the judgment set, and to determine the replacement content to form a modification instruction sequence. The text unit generation unit is used to adjust the language usage of the target paragraph according to the sequence of modification instructions using a text integration tool, and generate the adjusted text unit.

[0012] Furthermore, the overall generation ratio indicator determination module includes: The preliminary probability aggregation result determination unit is used to extract the distribution data of cross-domain terms from the adjusted text units, match them with the probability sequence of the remaining paragraphs using a text comparison tool, obtain the correlation value between the distribution data and the probability sequence of the remaining paragraphs, and determine the preliminary probability aggregation result. The overall probability distribution data acquisition unit is used to fuse the probability values ​​of text units based on the preliminary probability aggregation results and the degree of influence of cross-domain terms, using a weighted average calculation tool to obtain the overall probability distribution data. The correction coefficient judgment unit is used to redistribute the weights of cross-domain terms through probability adjustment tools and determine the appropriate correction coefficient if the overall probability distribution data exceeds the preset threshold range. The overall generation ratio indicator acquisition unit is used to perform secondary fusion of the adjusted text units and the probability sequence of the remaining paragraphs with the correction coefficient using a data integration tool to obtain the final overall generation ratio indicator.

[0013] Furthermore, the complete report output module includes: The paragraph adjustment scheme acquisition unit is used to obtain the sequence of modification instructions for the target paragraph from a pre-established database based on the overall generation ratio index, and to perform content feature analysis using a data comparison tool to obtain a preliminary paragraph adjustment scheme. The text unit determination unit is used to obtain matching data of detected attributes and reduced attributes for the initial paragraph adjustment plan. The attribute filtering tool is used to compare them item by item. If the data exceeds the preset threshold range, local optimization is performed to determine the optimized text unit. The text output result acquisition unit is used to perform secondary calibration of the adjustment unit based on the optimized text unit and the correspondence between the modification instruction sequence and the modification instruction, using a data integration tool to obtain the text output result that meets the requirements. The business closed-loop report acquisition unit is used to format the complete report based on the text output results using a report generation tool, covering relevant attribute data, to obtain the final business closed-loop report.

[0014] The beneficial effects achieved by this invention are as follows: This invention provides a paper generation rate detection system with modification assistance prompts, offering a complete solution to the problems of uneven distribution of cross-disciplinary terminology, insufficient paragraph consistency, and non-standard language usage in academic papers. The core issue lies in how to improve the logical consistency and professionalism of text through technical means. This invention extracts text features and calculates paragraph probabilities, adjusts paragraph sequences based on contextual analysis, and accurately judges and optimizes cross-disciplinary terminology and language usage for high-probability target paragraphs, generating a sequence of modification instructions and ultimately outputting a complete report containing optimized text units. Using cross-disciplinary terminology distribution as the core indicator, a weighted average method is used to determine the overall generation ratio, ensuring the professionalism and consistency of the text and achieving a closed-loop business process. The technical effect of this invention is to significantly improve the quality and readability of paper texts, providing intelligent support for academic writing. Attached Figure Description

[0015] Figure 1 This is a functional block diagram of an embodiment of the AI-generated paper generation rate detection system with modification assistance prompts of the present invention.

[0016] Explanation of icon numbers: 10. Paragraph-specific probability acquisition module; 20. Paragraph probability sequence determination module; 30. Judgment set generation module; 40. Text unit acquisition module; 50. Overall generation ratio index determination module; 60. Complete report output module. Detailed Implementation

[0017] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0018] like Figure 1As shown, the first embodiment of the present invention proposes a paper AI generation rate detection system with modification assistance prompts, including a paragraph-specific probability acquisition module 10, a paragraph probability sequence determination module 20, a judgment set generation module 30, a text unit acquisition module 40, an overall generation ratio index determination module 50, and a complete report output module 60. The paragraph-specific probability acquisition module 10 is used to obtain an initial paragraph sequence from the paper text through a text analysis model, extract text features from each initial paragraph sequence, calculate a preliminary generated content probability index, and incorporate cross-domain terminology distribution features as a core judgment index to obtain the paragraph-specific probability. The paragraph probability sequence determination module 20 is used to combine the paragraph-specific probability with the preceding and following related paragraph texts based on the adjacent paragraph integration attributes, and use context analysis methods to analyze consistency to determine the adjusted paragraph probability sequence. The judgment set generation module 30 is used to obtain cross-domain terminology and usage habit judgment indicators for target paragraphs in the adjusted paragraph probability sequence that exceed the threshold condition, and generate a judgment set based on these indicators. The text unit acquisition module 40 is used to construct a modification instruction sequence based on the judgment set if the number of cross-domain terms in the target paragraph exceeds the average value of the associated paragraphs, and then replace and adjust the text units according to the usage habits to obtain the adjusted text units. The overall generation ratio indicator determination module 50 is used to aggregate the overall probability from all adjusted text units and the remaining paragraph probability sequences, and use a weighted average method to integrate the influence of cross-domain terms to determine the overall generation ratio indicator. The complete report output module 60 is used to output a complete report containing the detected reduced attributes based on the overall generation ratio indicator and the modification instruction sequence of each target paragraph, covering the adjusted text units to achieve business closure.

[0019] The paragraph-specific probability acquisition module 10 takes the text of the paper to be detected as the analysis object. First, it uses a pre-trained text analysis model (such as BERT (Bidirectional Encoder Representations from Transformers, a bidirectional pre-trained language model based on Transformer encoders) and RoBERTa (a pre-trained language model based on BERT) adapted to academic text processing to naturally split the text according to the paragraph structure of the original paper, and obtain an ordered initial paragraph sequence (ensuring paragraph integrity and semantic independence). For each initial paragraph in the sequence, features are extracted and probabilities are calculated in two steps: First, basic text features (including sentence complexity, semantic coherence, repetition rate of common words, expression logic patterns, etc.) are extracted, and the features are converted into vectors and input into the probability calculation module to obtain a preliminary generated content probability index (value range 0-1, the higher the value, the greater the probability that the paragraph is generated by AI); Second, cross-domain terminology distribution features are extracted (covering the frequency of occurrence of professional terms, the adaptability to the field of the paper, the rationality of terminology collocation, the semantic contribution of terms in the paragraph, etc., to distinguish between "human-original professional expressions" and "AI-generated terminology piling / misuse"). This cross-domain terminology distribution feature is used as the core judgment indicator and is weighted and fused with the preliminary generated content probability index (the weight of cross-domain terminology features is higher than that of basic text features to avoid misjudging professional paragraphs). Finally, the "paragraph-specific probability" of each paragraph is output - that is, the precise probability value of the paragraph being generated by AI, providing core data support for subsequent context calibration and target paragraph selection.

[0020] The paragraph probability sequence determination module 20, based on the "paragraph-specific probability" (quantified probability value generated by AI for a single paragraph) output by the paragraph-specific probability acquisition module 10, first clarifies the "integration attributes" of adjacent paragraphs, which are key features reflecting the inherent connections between paragraphs. These include three core attributes: thematic relevance (continuity of core arguments and keywords), logical continuity (completeness of argumentation links such as causality, progression, and transition), and consistency of expression style (sentence complexity, density of terminology, and consistency of person / tense). Based on these integration attributes, the specific probability of a single paragraph is deeply linked with the text content of preceding and following paragraphs. A contextual analysis method of "semantic similarity calculation + logical link verification + style coherence matching" is used to conduct multi-dimensional consistency verification: algorithms such as cosine similarity and BERT semantic vector matching are used to determine the naturalness of content connection; the rationality of paragraph embedding is checked through argumentation logic graphs; and abrupt changes in writing style are investigated through style feature vector comparison.

[0021] Based on the "adjusted paragraph probability sequence" output by the paragraph probability sequence determination module 20 and the judgment set generation module 30, the AI ​​generation probability threshold is first preset (which can be customized according to the academic level of the paper and the strictness of detection, such as 0.7 for ordinary papers and 0.6 for core journals). Paragraphs with probability values ​​higher than this threshold are defined as "target paragraphs" (i.e., high-risk AI-generated paragraphs). For each target paragraph, the system extracts key judgment indicators in two categories: one is cross-domain terminology judgment indicators, including the suitability of terms to the field of the paper, terminology usage density (deviation from the average of human writing in the same field), logical coherence of terminology collocation, and semantic accuracy of terms (whether they conform to academic definitions); the other is usage habit judgment indicators, including sentence repetition rate, concentration of high-frequency vocabulary preferences, and completeness of argumentation logic (whether there is a gap between the argument, evidence, and conclusion). By integrating the specific quantitative data of the two types of indicators (such as "term matching degree 45%" and "sentence repetition rate 38%"), the anomaly judgment criteria (such as matching degree <60% is abnormal), and the problem location results (such as "term stacking" and "logical break"), a structured "basis judgment set" is formed. Each target paragraph corresponds to a dedicated set, clearly presenting the specific evidence and problem dimensions that it is judged as high-risk AI generation, providing a clear basis for subsequent targeted modifications.

[0022] The text unit acquisition module 40, based on the target paragraph (high-probability AI-generated paragraph) determined by the judgment set generation module 30 and the judgment set, first performs a cross-domain term quantity comparison—statistically comparing the total number of cross-domain terms in the target paragraph with the average number of terms in related paragraphs (adjacent paragraphs or paragraphs on the same topic). If the number of terms in the target paragraph significantly exceeds this average (e.g., exceeding 30% or more; a threshold can be customized based on domain characteristics), it is determined that there is a "term stuffing" problem (a typical characteristic of AI-generated text, manifested as illogical stacking of professional terms to cover up empty content).

[0023] The overall generation ratio determination module 50 uses the "adjusted text units" (high-risk AI-generated paragraphs that have been modified and optimized) and the "remaining paragraph probability sequence" (low-risk original paragraphs with probabilities below the threshold and not listed as target paragraphs in the paragraph probability sequence determination module 20) obtained by the text unit acquisition module 40 as core inputs, and completes the overall generation ratio calculation in two steps: Basic probability aggregation: First, the basic AI generation probability of the two types of paragraphs is defined - the basic probability of the adjusted text unit is its optimized correction probability (based on the effect of the modification instruction execution, such as the original paragraph specific probability of 0.8, which is reduced to 0.3 after terminology optimization and sentence structure adjustment); the basic probability of the remaining paragraphs directly adopts the paragraph specific probability adjusted by the paragraph probability sequence determination module 20 (such as low-risk values ​​such as 0.4, 0.2, etc.).

[0024] Weighted average fusion: The overall generation ratio is calculated using a weighted average method. The weight allocation focuses on the "influence of cross-domain terms" (key distinguishing features of AI-generated text), while also taking into account paragraph importance and optimization effect. The specific weighting rules are as follows: Cross-domain terminology influence weight (0.5): The weight is quantified based on the reasonableness of the use of cross-domain terms in the paragraph (such as whether there is still term stuffing and whether the suitability meets the standard). The more reasonable the terminology is used, the lower the weight (reducing the negative impact on the overall proportion).

[0025] Paragraph optimization effect weight (0.3): For the adjusted text unit, the weight is assigned according to the execution effect of the modification instructions (such as the reduction of sentence repetition rate and the improvement of logical integrity). The better the optimization effect, the lower the weight.

[0026] Paragraph length weight (0.2): Assigned according to the proportion of paragraph word count to the total word count of the paper. The higher the proportion of paragraph length, the higher the weight (to ensure that the core content has a more significant impact on the overall proportion).

[0027] By using the weighted average formula (overall generation ratio = Σ(basic probability of a single paragraph × corresponding weight) / Σ weight), the probability data of all paragraphs and the influencing factors of cross-domain terms are integrated to finally output a quantitative "overall generation ratio index" (value range 0-100%), which intuitively reflects the AI ​​generation ratio of the entire paper after optimization.

[0028] The complete report output module 60 takes the "overall generation ratio index" (the optimized AI generation quantitative ratio of the entire paper) determined by the overall generation ratio index determination module 50 and the "modification instruction sequence" (structured optimization scheme) for each target paragraph from the text unit acquisition module 40 as core inputs. It combines the "basis judgment set" (evidence for high-risk paragraph judgment) from the basis judgment set generation module 30 and the "adjusted text unit" (optimized paragraph text) from the text unit acquisition module 40 to form a complete report containing "detection results - modification scheme - optimization verification - follow-up suggestions".

[0029] The report should cover four key information categories, highlighting the "detection reduction attribute" (i.e., the magnitude of the decrease in AI generation rate before and after optimization, the reasons for the decrease, and the implementation effect): The quantitative detection results sub-unit clearly presents the overall generation ratio index (e.g., "AI generation rate after optimization is 28%), and compares it with the overall generation ratio before optimization (e.g., "45% before optimization"), and quantifies the reduction in detection (e.g., "reduced by 17 percentage points"); it also includes a detailed probability distribution of each paragraph (the corrected probability of the adjusted text unit and the original probability of the remaining paragraphs), and marks the optimization effect of high-risk paragraphs.

[0030] The target paragraph modification scheme sub-unit presents the "basis judgment set" (such as "term stuffing, sentence repetition rate of 38%), complete modification instruction sequence (such as term replacement suggestions, sentence adjustment scheme) and "adjusted text unit" (optimized text that can be directly copied and replaced) for each target paragraph according to the paragraph number, so that users can clearly know "where the problem is, how to modify it, and the effect after modification".

[0031] The sub-unit on attribute reduction in detection provides a detailed analysis of the core reasons for the overall decrease in the generation ratio, including the contribution of cross-domain terminology optimization (such as "solving the problem of terminology piling up and reducing the matching degree of AI-generated features"), the effectiveness of language habit adjustment (such as "sentence diversification and reducing traces of AI templated expression"), and the impact of improved logical integrity (such as "after supplementing arguments, the features of human writing are enhanced"), making the reduction effect of detection explainable and traceable.

[0032] Sub-unit for subsequent optimization suggestions: Based on the overall generation ratio index and paragraph optimization, provide targeted suggestions to further reduce the AI ​​generation rate (such as "a certain adjusted text unit still has a few logical gaps, it is recommended to supplement specific experimental data", "some sentences in the remaining low-risk paragraphs are too mechanical, the expression style can be appropriately optimized"); at the same time, clarify the implementation and use of the report (such as "the adjusted text unit can directly replace the corresponding paragraph of the original text, and it is recommended to re-test and verify after replacement").

[0033] Furthermore, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes a paragraph-specific probability acquisition module 10 comprising a preliminary probability value acquisition unit, an intermediate value determination unit, and a final value judgment unit. The preliminary probability value acquisition unit is used to extract an initial paragraph sequence from the full text of the paper using a text segmentation tool, and to obtain word frequency and sentence structure features for the initial paragraph sequence using a text feature extraction tool, calculate a preliminary content generation probability index, and obtain the preliminary probability value of the paragraph.

[0034] The following formula is used to determine the content generation probability of a paragraph by calculating the proportion of the weighted sum of the frequencies of each word in the paragraph to the total frequency of all words in the paragraph: (1) In formula (1), Indicates the first The initial probability value of each paragraph, Indicates the first The first paragraph Frequency of each word Indicates the first The weight coefficient of each word, Indicates the total number of paragraphs. Indicates the total number of words.

[0035] The following formula is used to extract lexical features of a paragraph through logarithmic transformation and inverse document frequency weighting: (2) In formula (2), This represents the word frequency feature value of a paragraph. Indicates the number of distinct words in a paragraph. Indicates the first The number of times each word appears in the paragraph. Indicates the first Inverse document frequency of each word.

[0036] When processing a paper on the application of artificial intelligence in medical diagnosis, the initial probability value acquisition unit first extracts an initial sequence of paragraphs from the full text using a text segmentation tool. This text segmentation tool is typically based on natural language processing techniques, such as sentence segmentation algorithms and paragraph boundary detection, to break down the paper content into independent paragraph units. Specifically, the tool scans the text, identifying paragraph start and end markers, such as line breaks or specific punctuation marks, to obtain an ordered list of paragraphs, such as a sequence from abstract to conclusion. This facilitates subsequent analysis, ensuring that each paragraph is processed as an independent entity. Assuming the full paper has 2000 words, the tool extracts 20 initial paragraphs, each averaging 100 words, forming sequences S1 to S20.

[0037] For the initial paragraph sequence, a text feature extraction tool was used to obtain word frequency and sentence structure features. This tool is a statistical model-based system. For example, it uses the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to calculate word frequency, which counts the number of times each word appears in the paragraph and divides it by the total number of words, while also considering inverse document frequency to highlight unique words. Sentence structure features are analyzed through parse tree analysis of sentence length, complexity, and grammatical patterns, such as calculating average sentence length or the proportion of subordinate clauses. For instance, in paragraph S5 of a medical paper, the tool extracted a word frequency of 0.05 for "diagnosis" and 0.03 for "neural network," and identified that the sentence structure was mainly complex sentences with an average sentence length of 25 words. These features reflect the language density and style of the paragraph, helping to quantify content characteristics.

[0038] Calculate a preliminary content generation probability index to obtain the initial probability value of the paragraph. Here, the content generation probability index is an evaluation based on a machine learning model, such as using a Bayesian network to combine word frequency and sentence structure features to calculate the probability value. The specific process includes inputting features into the model, which infers whether the paragraph is original content generated through training data. For example, the probability formula simplifies to P = (word frequency weight + sentence structure similarity) / 2. If the word frequency matching degree of S5 is high, then P = 0.8, indicating that the preliminary probability value is 80%, which provides a basic quantitative basis for subsequent steps.

[0039] The intermediate value determination unit is used to obtain the distribution density and domain relevance characteristics of cross-domain terms in a paragraph based on the preliminary probability value and the cross-domain term distribution feature database. The relevance characteristics are incorporated as the core judgment criteria to determine the intermediate value of a specific probability of the paragraph.

[0040] The distribution density of cross-disciplinary terms in a paragraph is obtained using the following formula: (3) In formula (3), This indicates the distribution density of cross-disciplinary terms in a paragraph. Indicates the total length of the paragraph. Indicates the total number of cross-domain terms. Indicates the first Weight coefficients for cross-domain terms, Indicates the first The frequency of cross-disciplinary terms in a paragraph Indicating cross-disciplinary terms The number of times it appears in cross-domain databases This indicates the total number of documents in the database.

[0041] The domain relevance characteristics of a paragraph are derived using the following formula: (4) In formula (4), This represents the domain relevance feature value of the paragraph. Indicates the total number of domain categories. Indicates the first The importance weight of each domain Indicates that the paragraph belongs to the first... The number of cross-disciplinary terms in each field Indicates the first The total number of cross-disciplinary terms in each field Represents the entropy decay coefficient. Indicates the first The cross-domain terminology distribution entropy of each field.

[0042] The median value of a paragraph's specific probability is obtained using the following formula: (5) In formula (5), The median value representing a specific probability of a paragraph. This represents the initial probability value. The coefficient representing the influence strength of the correlation feature. This represents the relevance feature score after fusion. The mean of the correlation feature is represented. The standard deviation represents the correlation characteristic. This represents the probability adjustment factor.

[0043] The intermediate value determination unit, based on the preliminary probability value and combined with a cross-domain terminology distribution feature database, obtains the distribution density and domain relevance features of cross-domain terms in a paragraph. This cross-domain terminology distribution feature database is a pre-built knowledge base containing lists and distribution statistics of cross-domain terms from multiple fields such as medicine and computer science. For example, density refers to the concentration of cross-domain terms in a paragraph, and relevance is calculated by matching the domain vector using cosine similarity. In a medical paper, for S5, the database query shows that the density of "neural network" is 2 times per 100 words, with a relevance of 0.9 to the AI ​​domain. This is incorporated as a core judgment criterion to determine the intermediate value of a specific probability for the paragraph; for example, after fusing the preliminary 0.8 with the relevance, an intermediate value of 0.75 is obtained.

[0044] The final value judgment unit is used to perform secondary calibration on the cross-domain term distribution characteristics in the paragraph through the cross-domain term weight adjustment tool if the intermediate value is lower than the preset threshold, obtain the adjusted feature weight, and judge the final value of the paragraph's specific probability.

[0045] The following formula is used to perform a secondary calibration adjustment of cross-domain term weights when the median value is below a threshold: (6) In formula (6), This represents the adjusted feature weights. Indicates the weight of the original cross-domain terms. Indicates the weight adjustment coefficient Indicates the preset threshold. This represents the current intermediate value.

[0046] The final probability value of the paragraph is calculated using the sigmoid function according to the following formula: (7) In formula (7), The final value representing a specific probability of a paragraph. Represents the base of the natural constant. Indicates the first Adjustment weights for cross-disciplinary terms, Indicates the first The distribution characteristic values ​​of cross-domain terms, This indicates the total number of cross-disciplinary terms in a paragraph.

[0047] The final value judgment unit is used to perform secondary calibration of the cross-domain terminology distribution characteristics in the paragraph using a cross-domain terminology weight adjustment tool if the median value is lower than a preset threshold, thereby obtaining the adjusted feature weights. This cross-domain terminology weight adjustment tool is an optimization algorithm, such as using gradient descent to adjust the weights. The process includes identifying low-relevance cross-domain terms, such as "diagnosis," and adjusting their weights from 0.5 to 0.7 to improve density balance. If the median value of 0.75 is lower than the threshold of 0.8, the calibrated weights are used to determine the final value of a specific probability for the paragraph, such as a final value of 0.85. This improves the accuracy of content analysis in business applications; for example, in academic integrity checks, it can effectively identify the probability of generated text, reduce misjudgments, and improve system robustness.

[0048] Preferably, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes a paragraph probability sequence determination module 20 comprising a paragraph correlation judgment result acquisition unit, an attribute fusion degree determination unit, a probability distribution difference acquisition unit, and a paragraph probability sequence judgment unit. The paragraph correlation judgment result acquisition unit is used to obtain at least one adjacent paragraph sequence from the target text through a text segmentation tool, extract semantic coherence and content coherence features from the adjacent paragraph sequences, and obtain preliminary paragraph correlation judgment results.

[0049] The preliminary results of the paragraph relevance assessment are obtained using the following formula: (8) In formula (8), Indicates the first The results of paragraph correlation judgment for adjacent paragraph sequences. This indicates the weighting of semantic coherence in the final judgment. This indicates the number of paragraph pairs included in the calculation. Indicates the first The semantic coherence score of each paragraph pair. Indicates the first Each paragraph is scored for content coherence.

[0050] When processing a paper on the application of artificial intelligence in education, the paragraph relevance assessment unit first extracts at least one sequence of adjacent paragraphs from the target text using a text segmentation tool. This text segmentation tool is a system based on natural language processing technology, typically using sentence segmentation algorithms and boundary recognition mechanisms to decompose the entire text into continuous paragraph units. Specifically, the text segmentation tool scans the text content, detecting line breaks, punctuation marks, or semantic breaks to extract adjacent paragraph sequences, such as three consecutive paragraphs from the introduction to the methodology section, forming sequences P1 to P3. This helps ensure that subsequent analysis focuses on context-related units, avoiding isolated processing. Assuming the full text of the paper is approximately 1500 words, the text segmentation tool extracts sequences P1 discussing the AI ​​overview, P2 describing educational applications, and P3 detailing case studies; these adjacent paragraphs total approximately 300 words, forming a compact analytical foundation. In this way, the system can initially capture the continuity of the text, laying the foundation for further feature extraction.

[0051] Semantic cohesion and content coherence features were extracted from adjacent paragraph sequences to obtain preliminary results on paragraph relevance. Specifically, semantic cohesion was assessed by calculating the keyword overlap rate and topic vector similarity between paragraphs. For example, word embedding models such as Word2Vec were used to generate vectors, and then cosine similarity was calculated to quantify the cohesion strength. Content coherence features involved analyzing the frequency of transition words and logical fluency in sentences, such as statistically analyzing the proportion of conjunctions like "therefore" or "in addition." In the sequence P1 to P3 of the educational paper, the extraction showed that the cohesion between P1 and P2 was 0.7, based on shared cross-domain terms such as "machine learning," while the coherence features indicated that transition words accounted for 15%, thus leading to a preliminary judgment of moderate relevance, reflecting the strength of the logical connection between paragraphs.

[0052] The attribute fusion degree determination unit is used to determine the attribute fusion degree of adjacent paragraph sequences by comparing the textual consistency and contextual logical features between adjacent paragraphs based on the preliminary paragraph correlation judgment results and using semantic comparison tools.

[0053] The following formula is used to calculate the attribute blending degree between adjacent paragraphs: (9) In formula (9), This indicates the paragraph attribute integration value. This represents the total number of paragraph attribute features. and Each represents a paragraph and paragraphs In the Feature values ​​on each attribute, This represents the attribute similarity weight coefficient. This represents the location distance weighting coefficient. This represents the distance attenuation parameter. Paragraph and paragraphs The positional distance between them.

[0054] The attribute fusion determination unit, based on the preliminary judgment results, uses a semantic comparison tool to compare the textual consistency and contextual logic features between adjacent paragraphs to determine the attribute fusion degree of the adjacent paragraph sequence. This semantic comparison tool is a system based on a deep learning framework, such as using the BERT model for semantic encoding, comparing the similarity of paragraph embeddings, and evaluating logical features such as the continuity of causal relationship chains. Specifically, it inputs P1 and P2 into the model, calculates a consistency score by comparing the consistency rate of entity mentions, such as whether the expression of "AI system" is consistent in the two paragraphs; and checks the smoothness of the reasoning chain through sequence modeling. If P2 continues the discussion of educational challenges in P1, the fusion degree is calculated as 0.65, representing the overall level of attribute integration. In business applications, this step can improve the accuracy of text quality assessment, for example, in academic publishing, helping to identify content fluency issues and thus optimizing the editing process.

[0055] The probability distribution difference acquisition unit is used to obtain the adjusted probability distribution difference by calibrating the paragraph boundary value and the associated weight ratio through a weight adjustment tool if the attribute fusion degree is lower than a preset threshold.

[0056] The following formula is used to calibrate paragraph boundary values: (10) In formula (10), This indicates the adjusted paragraph boundary value. Indicates the original paragraph boundary value. This represents the learning rate parameter of the weight adjustment tool. This indicates the change in the correlation weight ratio. Represents the target probability distribution. Indicates the current probability distribution; The following formula quantifies the degree of change in the probability distribution obtained after calibration using the weight adjustment tool: (11) In formula (11), This represents the difference in the adjusted probability distribution. The number of dimensions representing the probability distribution. Indicates the first The weight coefficients of the dimension, Indicates the weighted number after adjustment. The probability value of dimension, Indicates the number before weight adjustment The probability value of dimension.

[0057] The probability distribution difference acquisition unit is used to calibrate paragraph boundary values ​​and association weight ratios using a weight adjustment tool if the attribute fusion degree is lower than a preset threshold, thereby obtaining the adjusted probability distribution difference. The weight adjustment tool is an optimization algorithm system that dynamically adjusts parameters based on gradient methods. For example, it first identifies the weights of boundary values ​​such as paragraph start words, and then calibrates the association ratio by multiplying it by a correction factor. The process involves iteratively calculating the distribution difference; for example, an initial difference of 0.2 might be adjusted to 0.15 to ensure a more accurate probability representation.

[0058] The paragraph probability sequence judgment unit is used to perform a secondary comparison between the specific probability of a paragraph and the preceding and following related text, based on the adjusted probability distribution difference, using content integration tools, to determine the final paragraph probability sequence.

[0059] The final paragraph probability sequence is obtained using the following formula: (12) In formula (12), Paragraph The final probability sequence value, A parameter representing the balance between local probability and association probability. Paragraph The local specific probability, This indicates the total number of related paragraphs. Paragraph With paragraph The correlation strength coefficient between them.

[0060] The paragraph probability sequence judgment unit, based on the adjusted probability distribution difference, uses a content integration tool to perform a secondary comparison between the specific probability of the paragraph and the preceding and following related text to determine the final paragraph probability sequence. This content integration tool is an integrated framework that combines a rule engine and a machine learning module. For example, it matches the specific probability of P2 (0.8) with the text of P1 and P3, performing a secondary verification of logical consistency, and ultimately generating a sequence such as [0.75, 0.82, 0.78]. This improves the reliability of the generated text in content moderation operations.

[0061] Furthermore, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes a cross-domain terminology distribution set acquisition unit, a cross-domain terminology relevance judgment index determination unit, an index set acquisition unit, and a judgment set judgment unit based on the judgment set in the judgment set generation module 30. The cross-domain terminology distribution set acquisition unit is used to obtain target paragraphs with higher than a preset threshold from the adjusted paragraph probability sequence through semantic extraction tools, and to extract cross-domain terminology and usage habit features from the obtained target paragraphs with higher than the preset threshold to obtain a preliminary cross-domain terminology distribution set.

[0062] The following formula is used to filter out target paragraphs that are higher than a preset threshold from the adjusted paragraph probability sequence: (13) In formula (13), Represents the target set of paragraphs. Indicates the first The probability value of each paragraph. Indicates the preset threshold. Indicates the total number of paragraphs.

[0063] The process of extracting cross-domain terms from multiple target paragraphs is described by the following formula: (14) In formula (14), Represents a cross-domain terminology set. Indicates the number of target paragraphs. Indicates the first A cross-domain terminology extraction function for each paragraph. Indicates the first The semantic features of each paragraph Indicates the first The domain identifier for each paragraph.

[0064] The following formula is used to quantify the stylistic features of language usage in a target paragraph: (15) In formula (15), Indicates the characteristic value of usage habits. This indicates the number of custom feature dimensions. Indicates the first The weights of each feature Indicates the first The feature extraction function, Indicates the first The context information corresponding to each feature.

[0065] When processing a paper on the application of artificial intelligence in the medical field, the cross-domain terminology distribution set acquisition unit first uses a semantic extraction tool to extract target paragraphs with probabilities higher than a preset threshold from an adjusted paragraph probability sequence. This semantic extraction tool is a system based on natural language processing technology, typically utilizing attention mechanisms and sequence models to identify high-probability units. Specifically, the semantic extraction tool scans the probability sequence; for example, assuming the sequence value is [0.6, 0.85, 0.7] and the preset threshold is 0.8, it extracts the second paragraph as the target because its value is higher than the threshold. This helps focus on the core content and avoids low-relevance parts interfering with subsequent analysis. In one implementation example, the full paper is approximately 2000 words long. The adjusted paragraph probability sequence corresponds to the introduction, methods, and conclusion sections. The extracted target paragraph discusses the application of AI in diagnosis, totaling approximately 400 words, forming the basis for analysis. Through this extraction, the system can initially identify high-quality text units, providing support for feature extraction.

[0066] For target paragraphs exceeding a preset threshold, cross-domain terminology and usage characteristics are extracted to obtain a preliminary cross-domain terminology distribution set. Specifically, cross-domain terminology refers to terms like "neural network," borrowed from AI and applied to medical diagnosis. Usage characteristics include professional expressions, such as formal academic terminology or the use of specific abbreviations. The extraction process involves word frequency statistics and pattern matching, for example, using the TF-IDF algorithm to calculate cross-domain term weights, generating a distribution set such as {"deep learning": 0.4, "patient data": 0.3}, reflecting the mixed-domain language features in the paragraphs. In one implementation, for target paragraphs in medical papers, the extraction shows that cross-domain terminology accounts for 20%, and the usage is predominantly passive voice, thus obtaining a preliminary set to help assess the cross-domain integration degree of the text.

[0067] The cross-domain terminology relevance judgment index determination unit is used to determine the cross-domain terminology relevance judgment index by analyzing the contextual consistency characteristics of cross-domain terms and usage habits in the target paragraph based on the preliminary cross-domain terminology distribution set and using text comparison tools.

[0068] The following formula is used to analyze the contextual consistency characteristics of cross-domain terminology and usage in the target paragraph: (16) In formula (16), Indicating cross-disciplinary terms and cross-disciplinary terminology Contextual consistency coefficient between This represents the total number of context feature dimensions. Indicates the first The weights of each context feature Indicating cross-disciplinary terms and cross-disciplinary terminology In the A similarity function on each feature.

[0069] The relevance of cross-domain terms is assessed by comprehensively considering the frequency and distribution characteristics of cross-domain terms using the following formula: (17) In formula (17), Indicating cross-disciplinary terms The correlation index value, Represents the target set of paragraphs. Indicating cross-disciplinary terms In paragraph Frequency of occurrence in Indicating cross-disciplinary terms In paragraph Distribution density in Paragraph Total vocabulary.

[0070] The cross-domain terminology relevance assessment unit determines the relevance index based on a preliminary cross-domain terminology distribution set and employs a text comparison tool to analyze the contextual consistency features of cross-domain terms and usage habits in the target paragraph. This text comparison tool is an integrated comparison system based on a similarity calculation framework such as the Siamese network, comparing the semantic consistency of cross-domain terms in context. The specific process includes embedding cross-domain terms into vectors and then calculating Euclidean distance to assess consistency. For example, it checks whether "AI model" maintains the same meaning before and after a paragraph, resulting in an index such as 0.75, indicating a relevance level. In one implementation, when analyzing paragraphs from medical papers, high cross-domain terminology consistency was found, but occasional deviations in usage habits, such as colloquial expressions, were observed, resulting in an index of 0.68. This step can identify potential inconsistencies in academic review processes, improving content coherence.

[0071] The indicator set acquisition unit is used to calibrate the correlation between cross-domain terms and usage habits through a weight adjustment tool if the cross-domain terminology relevance judgment indicator is lower than a preset threshold, and to obtain an adjusted indicator set.

[0072] The process of generating a new set of indicators by adjusting the deviation ratio is described by the following formula: (18) In formula (18), This represents the adjusted set of indicators. Indicates the first One original indicator value, This indicates an adjustment to the strength coefficient. Indicates the first The deviation of each indicator from the target value Represents standardized parameters. This indicates the total number of indicators.

[0073] The indicator set acquisition unit is used to calibrate the correlation between cross-domain terms and usage habits using a weight adjustment tool if the cross-domain terminology relevance judgment indicator is lower than a preset threshold, thereby obtaining an adjusted indicator set. The weight adjustment tool is an optimization system that dynamically calibrates parameters using a gradient descent method. For example, if the initial relevance is 0.5, and it is lower than the threshold of 0.7, it is multiplied by a correction factor of 1.2, resulting in a new set [0.72, 0.65] after iteration. In one embodiment, for a medical paragraph with an indicator of 0.6, the calibrated set is increased to 0.71, ensuring a more accurate distribution, which optimizes text quality in the editing process.

[0074] The judgment unit based on the judgment set is used to compare and integrate cross-domain terminology and usage habits with the probability sequence of the target paragraph using content integration tools to determine the final judgment set based on the judgment set.

[0075] The optimal set of judgment criteria is selected using the following formula through multi-dimensional weighted scoring: (19) In formula (19), This indicates the index of the set used for final judgment. Indicates the first Content consistency score for each candidate set Indicates the first Reliability score of each candidate set Indicates the first The fusion score of each candidate set , , These represent the weight parameters for the three rating dimensions.

[0076] Based on the adjusted set of indicators, the judgment unit uses a content integration tool to fuse and compare cross-domain terminology and usage habits with the probability sequence of the target paragraph to determine the final judgment set. This content integration tool is a fusion framework that combines rule matching and embedding models for comparisons, such as fusing cross-domain terminology features with a probability of 0.85 to generate a set like {Consistency: 0.8, Fusion Degree: 0.75}. In one implementation, applied to medical papers, the final set supports the determination of text reliability, improving the verification efficiency of cross-domain content in publishing.

[0077] Preferably, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes a text unit acquisition module 40 comprising an adjustment requirement identifier generation unit, a matching degree index acquisition unit, a modification instruction sequence formation unit, and a text unit generation unit. The adjustment requirement identifier generation unit is used to obtain cross-domain term quantity data from the target paragraph, compare it with the average value of related paragraphs using a text comparison tool, and generate an adjustment requirement identifier if it exceeds a preset threshold.

[0078] The following formula is used to define the conditions for generating adjustment requirement identifiers: (20) In formula (20), This indicates the generation status of the adjustment requirement identifier. This represents the evaluation value obtained from the comparative analysis. This indicates the system's preset judgment threshold. When the evaluation value exceeds the threshold, the flag value is 1, indicating that adjustment is needed; otherwise, it is 0, indicating that no adjustment is needed.

[0079] When processing a paper on the application of artificial intelligence in the medical field, the adjustment requirement generation unit first obtains cross-domain term quantity data from the target paragraph. This cross-domain term quantity data refers to the total number of words introduced from the AI ​​field into medical text, such as "machine learning algorithm." Specifically, this process involves scanning the target paragraph using a word frequency analysis tool. For example, in a 400-word text discussing AI-assisted diagnosis, 15 cross-domain terms are identified, including "convolutional neural network" and "data training set." Then, a text comparison tool is used to compare the number of terms in the target paragraph with the average of related paragraphs. This text comparison tool is a system based on cosine similarity calculation. It compares the number of terms in the target paragraph with the average of other related paragraphs in the entire text. For example, if the average number of terms in related paragraphs is 10, and the target has 15 terms with a preset threshold of 12, then the threshold is exceeded, thus generating an adjustment requirement identifier. This identifier is a binary flag, such as "adjustment required = 1," used to trigger subsequent processing.

[0080] The matching degree index acquisition unit is used to obtain the language habit characteristics of the target paragraph based on the adjustment requirement identifier, and to obtain the matching degree index by analyzing the contextual consistency through semantic extraction tools.

[0081] The degree of matching between the requirement identifier and the target paragraph is calculated using the following formula in Gaussian function form: (twenty one) In formula (21), This represents the matching degree metric. The coefficient representing the maximum matching degree. The feature vector representing the demand identifier, The feature vector representing the target paragraph. This represents the matching sensitivity parameter.

[0082] The matching score acquisition unit, based on the adjustment requirement identifier, obtains the linguistic habits of the target paragraph. These linguistic habits include the frequency of passive voice use or the occurrence pattern of professional abbreviations such as "MRI". Specifically, contextual consistency is analyzed using a semantic extraction tool. This semantic extraction tool is a BERT-based system that evaluates consistency after feature vectorization. For example, in the target paragraph of a medical paper, it checks whether the semantics of "AI model" are semantically coherent in the surrounding sentences. By calculating the similarity between the embedded vectors, a matching score of 0.82 is obtained, indicating a high level of consistency. This process starts from the identifier trigger, ensuring that in-depth analysis is only performed on the parts that need adjustment.

[0083] The modification instruction sequence forming unit is used to extract adaptation rules from a pre-established language habit library based on the matching degree index and the judgment set, and to determine the replacement content to form a modification instruction sequence.

[0084] The following formula is used to extract matching rules from a pre-established database of idioms: (twenty two) In formula (22), This represents the adaptation rules extracted from the idiom database. This refers to a pre-established database of common expressions and habits. This represents the candidate rules in the library. Indicates the characteristics of the target content. Indicates historical usage frequency. The function represents the calculation of content similarity. The function represents the frequency matching degree calculation. and This represents the weighting parameter.

[0085] The following formula is used to generate the sequence of modification instructions: (twenty three) In formula (23), This indicates the first instruction in the sequence of modification instructions. One instruction, This indicates the total length of the instruction sequence to be modified. Indicates the operation type identifier. Indicates the replacement of positional parameters. This indicates the specific content value after the replacement.

[0086] The modification instruction sequence forming unit extracts adaptation rules from a pre-established language habit database using a content replacement tool, based on a matching degree index and a judgment set. This content replacement tool is a rule-driven framework that combines a machine learning classifier to match rules from the database. For example, the database stores formal expression rules in the medical field, such as replacing "our AI" with "this artificial intelligence system". After determining the replacement content, a modification instruction sequence is formed. This sequence is an ordered list, such as [Replacement position 1: Replace "our" with "system", position 2: Adjust to passive voice]. Based on the fusion of a matching degree of 0.82 and the judgment set, the specificity of the replacement is ensured.

[0087] The text unit generation unit is used to adjust the language usage of the target paragraph according to the sequence of modification instructions using a text integration tool, and generate the adjusted text unit.

[0088] The adjusted text unit is derived using the following formula: (twenty four) In formula (24), This represents the generated adjusted text unit. Indicates the first Performance factors of a text integration tool Indicates the first The amount of language habit adjustment processed by each tool Indicates the number of text integration tools used. This represents the overall adjustment coefficient.

[0089] The text unit generation unit adjusts the language usage of the target paragraph according to the sequence of modification instructions using a text integration tool. This text integration tool is a sequence generation model that uses a recurrent neural network to apply adjustments instruction by instruction. For example, for a medical paper paragraph, the instruction sequence guides the transformation of colloquial expressions into formal language, generating an adjusted text unit. For instance, the original sentence "AI is quite useful" becomes "The artificial intelligence system performs excellently." This text unit is approximately 350 characters long, maintaining the original meaning while enhancing professionalism, thus completing the entire process.

[0090] Furthermore, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes an overall generation ratio index determination module 50 comprising a preliminary probability aggregation result determination unit, an overall probability distribution data acquisition unit, a correction coefficient judgment unit, and an overall generation ratio index acquisition unit. The preliminary probability aggregation result determination unit is used to extract the distribution data of cross-domain terms from the adjusted text units, match it with the probability sequence of the remaining paragraphs using a text comparison tool, obtain the correlation value between the distribution data and the probability sequence of the remaining paragraphs, and determine the preliminary probability aggregation result.

[0091] The following formula is used to quantify and extract the weighted distribution features of terms from different domains in the adjusted text units: (25) In formula (25), Indicates the first Distribution data of cross-domain terms in each adjusted text unit. This indicates the total number of categories for cross-domain terms. Indicates the first Class terms in text units Frequency in Indicates the first The weight coefficient of the class term. The control logic of formula (25) is to combine the "occurrence density" and "domain importance" of cross-domain terms in the text unit with "frequency weighted summation + weight benchmark normalization" to quantify their distribution intensity in the text.

[0092] The Pearson correlation coefficient between the distribution data and the remaining paragraph probability sequence is calculated using the following formula to measure the degree of matching: (26) In formula (26), This represents the correlation value between the distributed data and the probability sequence of the remaining paragraphs. Indicates the total number of remaining paragraphs. Indicates the first The distribution data values ​​of each paragraph, Indicates the first The probability sequence values ​​of each paragraph. This represents the mean of the distributed data. The mean of the probability sequence is represented. The control logic of formula (26) is to quantify the linear correlation between the "distribution data" and the "remaining paragraph probability sequence" through "bias covariance + standardization", thereby measuring the matching consistency between the two.

[0093] The final probability aggregation output is determined using the following formula, which combines maximum value selection and weighted averaging: (27) In formula (27), This indicates the preliminary probability aggregation result. Indicates the aggregate weight parameter. This indicates the total number of paragraphs participating in the aggregation. Indicates the first The probability score of each paragraph. Indicates the first The confidence coefficient of each paragraph. The control logic of formula (27) combines the two strategies of "maximum value selection" and "confidence weighted average" to take into account both the "optimal value" of the probability score and the "overall confidence" and output a more reliable aggregate result.

[0094] The initial probability aggregation results determine the unit's performance when processing a paper on the application of artificial intelligence in the medical field. First, it extracts the distribution data of cross-disciplinary terms from the adjusted text units. This extraction process typically involves using natural language processing tools to scan and statistically analyze the text units, identifying cross-disciplinary terms like "deep learning model" that are introduced into medicine from the AI ​​field, and calculating their frequency of distribution within each unit. Specifically, this process includes breaking down the text units into word sequences and generating distribution maps using word frequency statistics algorithms. For example, in a unit discussing AI image diagnosis, cross-disciplinary terms such as "neural network training" appear 5 times and "big data analysis" appears 3 times, forming a distribution dataset of approximately 50 data points. This aids in subsequent matching analysis.

[0095] A text comparison tool is used to match the probability sequence of the remaining paragraphs to obtain the correlation value between the distributed data and the probability sequence of the remaining paragraphs. This text comparison tool is a similarity-based system that uses the Pearson correlation coefficient framework to assess association. For example, after vectorizing the extracted distributed data, it is compared with the probability sequence of the remaining paragraphs of the paper (such as the sequence of probabilities of cross-disciplinary terms). The specific process involves calculating the correlation coefficient between the two sequences. For example, if the distributed data sequence is [0.2, 0.3, 0.15] and the remaining sequence is [0.18, 0.28, 0.16], the correlation value is 0.95, indicating a high correlation, thus determining the preliminary probability aggregation result. This aggregation result is obtained through simple averaging or weighted summation, for example, merging the correlation value with the original probability into a preliminary value such as 0.85.

[0096] The overall probability distribution data acquisition unit is used to fuse the probability values ​​of text units based on the preliminary probability aggregation results and the degree of influence of cross-domain terms using a weighted average calculation tool to obtain the overall probability distribution data.

[0097] The final distribution is obtained by aggregating the probability values ​​of all text units using the following formula: (28) In formula (28), This represents the aggregated result of the overall probability distribution data. Indicates the total number of text units. Indicates the first The normalization coefficients of each text unit. Indicates the first The probability value of each text unit. Indicates the first The cross-domain influence correction factor of each unit. The control logic of formula (28) is to weight the probability value of each text unit with "normalization coefficient + cross-domain correction factor", and then sum the correction contributions of all units to obtain the overall probability distribution of information from all units.

[0098] The overall probability distribution data acquisition unit, based on the preliminary probability aggregation results, uses a weighted average calculation tool to fuse the probability values ​​of text units according to the influence degree of cross-domain terms, thus obtaining the overall probability distribution data. The weighted average calculation tool is a numerical processing system that works by weighting and summing the probability values ​​according to the influence weights of cross-domain terms (such as weights based on frequency of occurrence). For example, in a medical paper, a weight of 0.4 might be assigned to "convolutional neural network" and 0.3 to "patient data privacy," and then the fused distribution data is calculated. Specifically, this fusion process involves listing all probability values, applying formulas such as weighted sum divided by the total weights, and obtaining an overall distribution such as [0.25, 0.35, 0.4]. This connects the aforementioned aggregation results, ensuring the comprehensiveness of the distribution data.

[0099] The correction coefficient judgment unit is used to redistribute the weights of cross-domain terms through probability adjustment tools and determine the appropriate correction coefficient if the overall probability distribution data exceeds the preset threshold range.

[0100] The following formula is used to determine whether the overall probability distribution exceeds a preset threshold range: (29) In formula (29), This indicates the total number of cross-domain terms. Indicates the first The probability distribution values ​​of each term. Indicates an indicator function, This indicates the upper limit of the preset threshold. This represents the lower limit of the preset threshold. The control logic of formula (29) is to use an indicator function to filter out terms that exceed the threshold, and then calculate their proportion to quantify the degree of deviation of the overall probability distribution.

[0101] The following formula is used to redistribute the weights of cross-domain terms using the probability adjustment tool: (30) In formula (30), Indicates the first The weights of cross-domain terms after redistribution Represents the original weights. This indicates the domain relevance factor of term i. Representation of terms The frequency adjustment factor, This represents the total number of terms whose weights need to be adjusted. The control logic of formula (30) is... The final correction factor is calculated using the following formula: (31) In formula (31), This represents the adjustment factor for the fit. Indicates the basic correction parameter. This indicates the number of data points involved in the correction calculation. Indicates the first The deviation value of each data point This represents the mean of the deviation values. Indicates the stability adjustment parameter. The regularization parameter is represented. The control logic of formula (31) is to redistribute the weights through normalization based on the priority of "original weight × domain-related factor × frequency factor" so that the weights are more in line with the requirements of "domain association + frequency adaptation".

[0102] The correction coefficient determination unit is used to redistribute the weights of cross-domain terms using a probability adjustment tool if the overall probability distribution data exceeds a preset threshold range, thereby determining the appropriate correction coefficient. The probability adjustment tool is an optimization system that uses an iterative algorithm to adjust the weights. For example, if the distribution data exceeds the threshold of 0.5, the coefficient is recalculated using gradient descent. The specific process involves initializing the weights, calculating the deviation, and then iteratively updating. For instance, adjusting the original weight of 0.4 to 0.35 yields a correction coefficient of 1.2. This step continues the fusion analysis, correcting for cases exceeding the threshold.

[0103] The overall generation ratio indicator acquisition unit is used to perform secondary fusion of the adjusted text units and the probability sequence of the remaining paragraphs with the correction coefficient using a data integration tool to obtain the final overall generation ratio indicator.

[0104] The final overall generation ratio is derived using the following formula: (32) In formula (32), This represents the final overall generation ratio indicator. This represents the correction factor. This indicates the adjusted text unit proportions. This indicates the total number of remaining paragraphs. Indicates the first The probability value of each paragraph. Indicates the first The weighting coefficients for each paragraph. The control logic of formula (32) uses a correction coefficient. By balancing the "basic proportion of the adjusted text unit" and the "probability-weighted contribution of the remaining paragraphs", the final overall generation proportion index is obtained.

[0105] The overall generation ratio acquisition unit, based on the correction coefficient, uses a data integration tool to perform a secondary fusion of the adjusted text units and the remaining paragraph probability sequences to obtain the final overall generation ratio index. The data integration tool is a merging system that works by performing secondary data fusion through matrix operations, such as multiplying the corrected text unit probability with the remaining sequence and then normalizing it. Specifically, in the medical paper example, this process includes constructing a fusion matrix, applying a correction coefficient of 1.2, and calculating a ratio index such as 0.92, representing the final generation ratio. This forms a complete logical chain from extraction to fusion.

[0106] Preferably, the paper AI generation rate detection system with modification assistance prompts provided in this embodiment includes a complete report output module 60 comprising a paragraph adjustment scheme acquisition unit, a text unit determination unit, a text output result acquisition unit, and a business closed-loop report acquisition unit. The paragraph adjustment scheme acquisition unit is used to obtain the modification instruction sequence of the target paragraph from a pre-established database based on the overall generation ratio index, and to perform content feature analysis using a data comparison tool to obtain a preliminary paragraph adjustment scheme.

[0107] The following formula is used to quantify the effectiveness and scope of paragraph adjustment plans: (33) In formula (33), This indicates the evaluation value of the paragraph adjustment plan. This indicates the total number of content units in the paragraph that need to be adjusted. Indicates the first The original feature values ​​of each content unit, Indicates the first The adjusted feature values ​​of each content unit Indicates the first Adjustment weighting factor for each content unit.

[0108] When processing a paper on the application of artificial intelligence in education, the paragraph adjustment scheme acquisition unit retrieves a sequence of modification instructions for the target paragraph from a pre-established database based on an overall generation ratio index, such as 0.92. This database is a structured storage system containing modification rules for various cross-disciplinary terms. For example, for AI cross-disciplinary terms introduced in education papers, such as "machine learning algorithm," the database provides instruction sequences including specific operations such as replacement, insertion, or deletion. Specifically, this acquisition process involves querying the database index, first inputting the ratio index as a filtering condition, and then extracting matching sequences. For example, instruction 1 in the sequence is "replace 'traditional teaching' with 'AI-assisted teaching'," and instruction 2 is "insert privacy protection statement," thus forming an ordered list of instructions. This helps with subsequent content adjustments and ensures the coherence of the paper. In one implementation, a data comparison tool is used to perform content feature analysis to obtain a preliminary paragraph adjustment scheme. The data comparison tool is a feature extraction-based system that compares text similarity through semantic vectors, for example, using a cosine similarity framework to evaluate the matching degree between the target paragraph and a standard template. The specific process involves converting paragraphs of the paper into vector representations, such as converting "online learning platform" into a vector of [0.1, 0.5, 0.3], and then comparing it with the database template vector [0.12, 0.48, 0.32] to calculate a similarity score of 0.95. If the score is high, an adjustment plan is generated, such as "enhancing the explanation of AI cross-domain terms in the paragraphs." This connects the application of the aforementioned modification instruction sequence and forms the basis for the preliminary plan.

[0109] The text unit determination unit is used to obtain matching data of detected attributes and reduced attributes for the initial paragraph adjustment plan. The data is compared item by item using an attribute filtering tool. If the data exceeds the preset threshold range, local optimization is performed to determine the optimized text unit.

[0110] The following formula is used to calculate the degree of matching between the detection attribute and the reduction attribute: (34) In formula (34), This represents the matching metric between the detected attribute and the reduced attribute. Indicates the first The value of each detected attribute, Indicates the first A value that reduces the attribute. This indicates the preset maximum threshold range. This represents the weighting coefficient between attributes.

[0111] The text unit determines the matching data of the detected attributes and the attributes to be reduced for the initial paragraph adjustment plan. This matching data is then compared item by item using an attribute filtering tool. The attribute filtering tool is a filtering system that checks attributes one by one according to preset rules. For example, detected attributes include "accuracy of cross-domain terminology" and "logical fluency," while reduced attributes include "redundant descriptions." Specifically, this comparison process involves listing the attribute data in the plan. For example, if the detected attribute value is 0.8 and the reduced attribute value is 0.2, then a threshold of 0.5 is compared. If the detected attribute exceeds this threshold, it is marked as needing optimization, thus determining the necessity of local optimization. This step continues the feature analysis, ensuring the plan's relevance.

[0112] The text unit determination unit is used to perform local optimization if the text exceeds a preset threshold range, thus determining the optimized text unit. The specific process includes iterative adjustments. For example, for paragraphs exceeding the threshold, optimization algorithms such as the stepwise replacement method are applied. Problem areas, such as excessive repetition of cross-disciplinary terms, are first identified, and then replaced with concise expressions to obtain optimized units such as "The application of AI in education has improved the efficiency of personalized learning." This optimizes the redundancy problem of the original unit.

[0113] The text output result acquisition unit is used to perform secondary calibration of the adjustment unit based on the optimized text unit and the correspondence between the modification instruction sequence and the modification instruction, using a data integration tool to obtain the text output result that meets the requirements.

[0114] The required text output is obtained using the following formula: (35) In formula (35), This indicates the text output that meets the requirements. This indicates the total number of original text units. Indicates the first The weight coefficient of each text unit, Indicates the first The content value of each text unit Indicates the first The optimization adjustment factor for each unit. This indicates the impact coefficient of the modification instruction. Indicates the total number of modification commands. Indicates the first The calibration value of each modification instruction.

[0115] The text output acquisition unit, based on the optimized text units and the correspondence between the modification instruction sequence and the modification instructions, uses a data integration tool to perform a secondary calibration of the adjustment unit. The data integration tool is a fusion system that merges data through mapping relationships; for example, it maps "insert privacy statement" in the instruction sequence to a specific unit. The secondary calibration process involves checking consistency; if the matching degree between the adjusted unit and the instruction reaches 0.9, the output result is confirmed. This forms a logical chain from optimization to calibration.

[0116] The business closed-loop report acquisition unit is used to format the complete report based on the text output results using a report generation tool, covering relevant attribute data, to obtain the final business closed-loop report.

[0117] The following formula is used to evaluate the quality of a complete report processed by a report generation tool: (36) In formula (36), This indicates the quality score of the final business loop report. This indicates the total number of attribute data included in the report. Indicates the first Weight coefficients for each attribute data, Indicates the first Completeness score of each attribute data This represents the normalization factor for the formatting process.

[0118] The business loop report acquisition unit formats the text output results into a complete report using a report generation tool, covering relevant attribute data. This report generation tool is a template-based system. Specifically, this process includes importing the output results, adding attribute data such as "similarity before and after adjustment 0.85," and then formatting it into a standard report structure. This step achieves business loop closure, ensuring the completeness and readability of the final report, and effectively improving the quality of cross-domain integration in educational paper processing.

[0119] This embodiment provides a paper generation rate detection system with modification assistance prompts. Compared with existing technologies, it extracts text features and calculates paragraph probabilities, adjusts paragraph sequences based on contextual analysis, and accurately judges and optimizes cross-domain terminology and usage habits for high-probability target paragraphs, generating a sequence of modification instructions and finally outputting a complete report containing optimized text units. Using cross-domain terminology distribution as the core indicator, a weighted average method is integrated to determine the overall generation ratio, ensuring the professionalism and consistency of the text and achieving a closed-loop business process. This embodiment significantly improves the quality and readability of paper texts, providing intelligent support for academic writing.

[0120] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A thesis AI generation rate detection system with modification assistance prompts, characterized in that, The method comprises the following steps: A paragraph-specific probability acquisition module (10) is used to acquire an initial paragraph sequence from a paper text by a text analysis model, extract text features from each initial paragraph sequence, calculate a preliminary content generation probability indicator, incorporate cross-field term distribution features as a core judgment indicator, and obtain a paragraph-specific probability; A paragraph probability sequence determination module (20) is used to combine the paragraph-specific probability with the text of the adjacent paragraphs before and after according to the adjacent paragraph attribute, analyze the consistency by using a context analysis method, and determine an adjusted paragraph probability sequence; A basis judgment set generation module (30) is used to acquire the cross-field term and language habit judgment indicators of a target paragraph in the adjusted paragraph probability sequence that is higher than a threshold condition, and generate a basis judgment set; A text unit acquisition module (40) is used to construct a modification instruction sequence by the basis judgment set if the number of cross-field terms in the target paragraph exceeds the average value of the associated paragraphs, replace and adjust the language habits to obtain an adjusted text unit; An overall generation proportion indicator determination module (50) is used to aggregate the overall probability from all adjusted text units and the remaining paragraph probability sequence, determine the overall generation proportion indicator by using a weighted average method to incorporate the influence of cross-field terms; A complete report output module (60) is used to output a complete report containing a detection reduction attribute according to the overall generation proportion indicator and the modification instruction sequence of each target paragraph, and cover the adjusted text unit to achieve a business closed loop.

2. The paper AI generation rate detection system with modification assisted prompt of claim 1, wherein, The paragraph-specific probability acquisition module (10) comprises: A preliminary probability value acquisition unit is used to extract an initial paragraph sequence from a paper text by a text segmentation tool, acquire the word frequency and sentence structure features of the initial paragraph sequence by using a text feature extraction tool, calculate a preliminary content generation probability indicator, and obtain a preliminary probability value of the paragraph; An intermediate value determination unit is used to acquire the distribution density and field relevance features of the cross-field terms in the paragraph according to the preliminary probability value and in combination with a cross-field term distribution feature database, incorporate the relevance features as a core judgment basis, and determine the intermediate value of the paragraph-specific probability; A final value judgment unit is used to perform secondary calibration on the cross-field term distribution features in the paragraph by using a cross-field term weight adjustment tool if the intermediate value is lower than a preset threshold, acquire the adjusted feature weight, and judge the final value of the paragraph-specific probability.

3. The paper AI generation rate detection system with modification assisted hints of claim 1, wherein, The paragraph probability sequence determination module (20) comprises: A paragraph correlation judgment result acquisition unit is used to acquire at least one adjacent paragraph sequence from a target text by a text segmentation tool, extract the semantic cohesion degree and content coherence features of the adjacent paragraph sequence, and obtain a preliminary paragraph correlation judgment result; An attribute fusion degree determination unit is used to determine the attribute fusion degree of the adjacent paragraph sequence according to the preliminary paragraph correlation judgment result by using a semantic comparison tool to compare the text consistency and context logic features between adjacent paragraphs. The probability distribution difference acquisition unit is configured to, if the attribute fusion degree is lower than a preset threshold, calibrate the paragraph boundary value and the correlation weight ratio by using a weight adjustment tool, and acquire an adjusted probability distribution difference. The paragraph probability sequence judgment unit is configured to, for the adjusted probability distribution difference, perform secondary comparison between the paragraph-specific probability and the front and rear related text by using a content integration tool, and judge a final paragraph probability sequence.

4. The paper AI generation rate detection system with modification assisted hints of claim 1, wherein, The judgment set generation module (30) comprises: The cross-domain term distribution set acquisition unit is configured to acquire a target paragraph higher than a preset threshold from the adjusted paragraph probability sequence by using a semantic extraction tool, extract cross-domain terms and language habit features for the acquired target paragraph higher than the preset threshold, and obtain a preliminary cross-domain term distribution set. The cross-domain term relevance judgment index determination unit is configured to, according to the preliminary cross-domain term distribution set, analyze the context consistency features of the cross-domain terms and the language habits in the target paragraph by using a text comparison tool, and determine a cross-domain term relevance judgment index. The index set acquisition unit is configured to, if the cross-domain term relevance judgment index is lower than a preset threshold, calibrate the correlation degree of the cross-domain terms and the language habits by using a weight adjustment tool, and acquire an adjusted index set. The judgment set judgment unit is configured to, for the adjusted index set, perform fusion comparison between the cross-domain terms and the language habit features and the probability sequence of the target paragraph by using a content integration tool, and judge a final judgment set.

5. The paper AI generation rate detection system with modification assisted hints of claim 1, wherein, The text unit acquisition module (40) comprises: The adjustment demand identifier generation unit is configured to acquire cross-domain term quantity data from the target paragraph, perform comparison with an average value of related paragraphs by using a text comparison tool, and generate an adjustment demand identifier if a preset threshold is exceeded. The matching degree index acquisition unit is configured to acquire language habit features of the target paragraph according to the adjustment demand identifier, analyze context consistency by using a semantic extraction tool, and obtain a matching degree index. The modification instruction sequence formation unit is configured to, according to the matching degree index and the judgment set, extract an adaptation rule from a pre-established language habit library by using a content replacement tool, judge replacement content, and form a modification instruction sequence. The text unit generation unit is configured to, according to the modification instruction sequence, adjust the language habits of the target paragraph by using a text integration tool, and generate an adjusted text unit.

6. The paper AI generation rate detection system with modification assisted hints of claim 1, wherein, The overall generation proportion index determination module (50) comprises: The preliminary probability aggregation result determination unit is configured to extract distribution data of cross-domain terms from the adjusted text unit, perform matching with a remaining paragraph probability sequence by using a text comparison tool, acquire a relevance value of the distribution data and the remaining paragraph probability sequence, and determine a preliminary probability aggregation result. The following formula is used to quantitatively extract the weighted distribution features of different domain terms in the adjusted text unit: ; in, Indicates the first Distribution data of cross-domain terms in each adjusted text unit. This indicates the total number of categories for cross-domain terms. Indicates the first Class terms in text units Frequency in Indicates the first Weighting coefficients for class terms; The following formula is used to calculate the Pearson correlation coefficient between the distribution data and the remaining paragraph probability sequence to measure the matching degree: ; wherein, represents a correlation value of the distribution data and the probability sequence of the remaining passages, represents the total number of the remaining passages, represents a distribution data value of the passage, represents a probability sequence value of the passage, represents a mean value of the distribution data, represents a mean value of the probability sequence;​​ The following formula is used to determine the final probability aggregation output by combining maximum value selection and weighted average: ; wherein, denotes the preliminary probability aggregation result, denotes the aggregation weight parameter, denotes the total number of passages participating in the aggregation, denotes the probability score of the th passage, denotes the confidence coefficient of the th passage; The overall probability distribution data acquisition unit is configured to fuse the probability values of the text units by using a weighted average calculation tool according to the influence degree of the cross-field term based on the preliminary probability aggregation result, so as to obtain overall probability distribution data. The correction coefficient judgment unit is configured to re-distribute the weight of the cross-field term by using a probability adjustment tool and judge the adaptive correction coefficient if the overall probability distribution data exceeds a preset threshold range. The overall generation proportion index acquisition unit is configured to secondarily fuse the adjusted text unit and the remaining paragraph probability sequence by using a data integration tool based on the correction coefficient, so as to obtain a final overall generation proportion index.

7. The paper AI generation rate detection system with modification assisted prompt of claim 6, wherein, In the overall probability distribution data acquisition unit, the probability values of all text units are aggregated by using the following formula to obtain a final distribution: ; wherein, represents an aggregated result of the overall probability distribution data, represents a total number of text units, represents a normalization coefficient of the th text unit, represents a probability value of the th text unit, represents a cross-domain influence correction factor of the th unit.

8. The paper AI generation rate detection system with modification assisted prompt of claim 7, wherein, In the correction coefficient judgment unit, the following formula is used to judge whether the overall probability distribution exceeds the preset threshold range: ; wherein, represents the total number of cross-domain terms, represents a probability distribution value of the th term, represents an indicator function, represents an upper limit of the preset threshold value, represents a lower limit of the preset threshold value; The following formula is used to realize the re-distribution of the weight of the cross-field term by using the probability adjustment tool: ; in, Indicates the first The weights of cross-domain terms after redistribution Represents the original weights. This indicates the domain relevance factor of term i. Representation of terms The frequency adjustment factor, This indicates the total number of terms whose weights need to be adjusted. The following formula is used to calculate the final correction coefficient: ; wherein, denotes an adapted correction coefficient, denotes a base correction parameter, denotes the number of data points participating in the correction calculation, denotes the deviation value of the th data point, denotes the mean value of the deviation values, denotes a stability adjustment parameter, denotes a regularization parameter.

9. The paper AI generation rate detection system with modification assisted prompt of claim 8, wherein, In the overall generation proportion index acquisition unit, the final overall generation proportion index is obtained by using the following formula: ; wherein, represents the final overall generation proportion indicator, represents the correction coefficient, represents the adjusted text unit proportion, represents the total number of remaining paragraphs, represents the probability value of the th paragraph, represents the weight coefficient of the th paragraph.

10. The paper AI generation rate detection system with modification assisted hints of claim 1, wherein, The complete report output module (60) comprises: The paragraph adjustment scheme acquisition unit is configured to obtain a modification instruction sequence of the target paragraph from a pre-established database based on the overall generation proportion index, analyze the content features by using a data comparison tool, and obtain a preliminary paragraph adjustment scheme. The text unit determination unit is configured to obtain matching data of the detection attribute and the reduction attribute based on the preliminary paragraph adjustment scheme, compare the data item by item by using an attribute screening tool, and determine the optimized text unit by performing local optimization if the data exceeds a preset threshold range. The text output result acquisition unit is configured to secondarily calibrate the adjustment unit by using a data integration tool based on the correspondence between the modification instruction sequence and the modification instruction based on the optimized text unit, so as to obtain a text output result meeting the requirements. The business closed loop report acquisition unit is configured to format the complete report by using a report generation tool based on the text output result, cover the related attribute data, and obtain a final business closed loop report.