Text washing detection method based on deep learning

Through deep learning methods, the BERT model, multi-sentence feature fusion mechanism and loss function are used to optimize the neural network, and the accuracy problem of traditional text scheduling detection methods in complex scenarios is solved, and efficient and robust text scheduling detection is achieved.

CN120297260APending Publication Date: 2025-07-11UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510216363.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional text scheduling detection methods are not accurate when dealing with complex text scheduling scenarios, making it difficult to accurately judge the scheduling behavior, especially those scheduling techniques that change the surface shape of the text but maintain the original intention.

Method used

Using the text scheduling detection method based on deep learning, text features are extracted through the BERT model, combined with the multi-sentence feature fusion mechanism and triple loss and positive example enhancement loss function, the neural network model is optimized, and a robust text feature representation is constructed.

Benefits of technology

It significantly improves the accuracy and robustness of text draft scheduling detection, can effectively identify text draft scheduling to varying degrees, and improves the accuracy and efficiency of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297260A_ABST
    Figure CN120297260A_ABST
Patent Text Reader

Abstract

The invention discloses a text washing detection method based on deep learning. The method comprises the following steps: S1, carrying out feature coding on an input text through a BERT model; s2, grouping and fusing the features of the draft washing text extracted in the above step to obtain higher-quality feature representation of the draft washing text; s3, performing multi-statement feature fusion on the text features in the above step, integrating multi-view context information, and capturing deep text semantic features; s4, optimizing the neural network model through triple loss and a positive example enhancement loss function; and S5, designing and constructing a manuscript washing detection data set through a data set generation module based on a comparative learning thought. According to the method disclosed by the invention, the method has good robustness on retouching and multi-round translation manuscript washing, while the text features are efficiently extracted, the text manuscript washing behavior is accurately detected, and the method is suitable for the field of text manuscript washing detection and copyright protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing and deep learning, and particularly relates to a method for detecting text rewriting based on deep learning. Background Art

[0002] With the continuous development of content creation and dissemination methods, text rewriting detection has gradually become an important challenge in the field of intellectual property protection. Traditional text rewriting detection methods often have low accuracy when dealing with complex text rewriting scenarios. The reason is that these rewriting techniques can effectively change the surface form of the text while maintaining the original meaning. Rewriting behaviors may include, but are not limited to, translating or polishing the original text multiple times and then recombining it, making it difficult for rewriting detection methods that rely solely on traditional text similarity metrics to accurately judge. Summary of the Invention

[0003] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a method for detecting text rewriting based on deep learning. By fusing text rewriting features at different levels and strengthening the discriminative features in the rewritten text, higher-quality text rewriting features can be obtained. The present invention also proposes a multi-sentence feature fusion mechanism that combines multi-perspective context information to capture deep text semantic features. In addition, the present invention combines the advantages of triplet loss and positive example enhancement loss functions in semantic representation learning to adapt to complex text rewriting detection scenarios and guide the proposed neural network model to more accurately detect text rewriting. To achieve the above object and other advantages of the present invention, a method for detecting text rewriting based on deep learning is provided, including:

[0004] S1. Feature encoding the input text through a BERT model;

[0005] S2. Grouping and fusing the text rewriting features extracted in the above step to obtain a higher-quality representation of text rewriting features;

[0006] S3. Performing multi-sentence feature fusion on the text features in the above step, integrating multi-perspective context information, and capturing deep text semantic features;

[0007] S4. Optimizing the neural network model through triplet loss and positive example enhancement loss functions;

[0008] S5. Designing and constructing a text rewriting detection dataset based on the contrast learning idea through a dataset generation module.

[0009] Preferably, the step S2 specifically includes the following steps:

[0010] S21. Concatenate the feature matrices of a certain amount of paraphrased text and the feature matrix of the full-text paraphrased text to obtain combined features, and divide the feature maps of the combined features into multiple groups;

[0011] S22. Approximately represent the semantic information of each group by calculating the average value of each group of sub-features;

[0012] S23. Multiply the average vector back to each sub-feature in the form of a dot product to obtain the importance coefficient of each sub-feature in this group of text features;

[0013] S24. Perform normalization processing on the importance coefficient;

[0014] S25. Adjust the normalized output through the introduced scaling parameter g and translation parameter w, and generate a new normalized importance coefficient through the sigmoid function σ(·);

[0015] S26. The normalized importance coefficient Multiply by the original sub-feature q i , to obtain an enhanced sub-feature vector;

[0016] S27. Concatenate all the enhanced sub-vectors together to form an enhanced vector group. After the features of all groups are enhanced, re-concatenate to obtain an enhanced feature representation of the paraphrased text.

[0017] Preferably, the specific steps in step S3 include the following steps:

[0018] S31. Convolve the feature matrix of the original text through convolution kernels of multiple different sizes, and use all-zero matrices of 3×d, 2×d, and 1×d to pad the feature matrix respectively, so that the output vector dimensions are kept consistent;

[0019] S32. After convolving the feature matrix of the original text above to obtain multiple sentence feature matrices, perform max pooling on the sentence feature matrices to reduce the feature dimensions, and after multiple convolution and pooling operations, concatenate multiple feature vectors to form a comprehensive feature representation;

[0020] S33. Similarly obtain the final feature matrices of the manuscript text and the contradictory text through steps S31 - S32.

[0021] Preferably, the input text in step S1 includes the original text, a small amount of paraphrased text, full-text paraphrased text, and contradictory text. The original text, a small amount of paraphrased text, full-text paraphrased text, and contradictory text are respectively sent into the BERT model to extract features to obtain corresponding feature matrices.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: compared with the prior art, it overcomes the problems of poor robustness of the extracted text features and independent detection of plagiarism for each sentence.

[0023] Combining the grouped enhancement fusion mechanism and the multi-sentence feature fusion mechanism, strengthening the discriminative features in the plagiarized text, and combining multi-perspective context information to obtain a robust text feature representation; integrating the advantages of the triplet loss and the positive example enhancement loss in semantic representation learning to guide the model to more accurately detect text plagiarism.

[0024] In summary, the present invention realizes the perfect combination of high efficiency, strong robustness and high precision in the field of text plagiarism detection, can significantly improve the overall performance of text plagiarism detection, and provides an effective solution for copyright protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is the overall flowchart of the deep learning-based text plagiarism detection method according to the present invention;

[0026] Figure 2 It is the overall framework diagram of the deep learning model of the deep learning-based text plagiarism detection method according to the present invention;

[0027] Figure 3 It is the operation schematic diagram of the multi-sentence feature fusion module of the deep learning-based text plagiarism detection method according to the present invention;

[0028] Figure 4 It is the detection performance result diagram based on cosine similarity and Spearman correlation coefficient of the deep learning-based text plagiarism detection method according to the present invention;

[0029] Figure 5 It is the t-SNE dimensionality reduction and visualization of the text distribution of the deep learning-based text plagiarism detection method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Refer to Figure 1 , a deep learning-based text plagiarism detection method, including the following steps:

[0032] S1. Feature encoding the input text through the BERT model:

[0033] In this step, the original text data is first encoded and converted into a form that can be processed by a deep learning model, that is, the linguistic attributes such as word meanings, grammatical structures, and context relationships in the text data are mapped to an abstract feature space, and the input text is represented by a feature matrix.

[0034] The present invention uses a BERT model to perform feature encoding on the input text. Assuming that there are a total of m sentences in the original text o, the original text sentence set can be represented by For a given sentence, first, the special character " <pad>"The sentence length is expanded to 512, and then the words in the sentence are one-hot encoded and encoded into word vectors. Then the obtained encoding vector is input into the BERT model for feature extraction to obtain the semantic vector of each sentence. Finally, these sentence features are concatenated to obtain the feature matrix V of the original text. o The BERT model can capture long-range dependencies between text words through the self-attention mechanism, thereby achieving a global understanding of the entire input text sequence. This mechanism allows the model to dynamically adjust the degree of attention to each word in the sentence according to the context, thereby better understanding the ambiguity and complexity of language.

[0035] The inputs of the present invention are the original text o, a small amount of plagiarized text b, the full plagiarized text h, and a text completely unrelated to the original text, called contradictory text c. The purpose of this design is first to enable the method proposed by the present invention to effectively detect and identify different degrees of text plagiarism, while also not mistakenly detecting contradictory text that is unrelated to the original text as plagiarized text. Similar to the original text, a small amount of plagiarized text b, the full plagiarized text h, and the contradictory text c are also input into the BERT model to extract features, and the feature matrices V are obtained respectively. b 、V h and V c .

[0036] S2. By grouping and fusing the plagiarized text features extracted in the above steps, a higher quality plagiarized text feature representation is obtained:

[0037] This step aims to strengthen the discriminative features in the plagiarized text by fusing plagiarized text features of different degrees, so as to obtain higher quality plagiarized text features and improve the plagiarism detection capability of the method proposed in the present invention.

[0038] First, the feature matrix V of a small amount of plagiarized text b And the feature matrix V of the full text plagiarism text h After concatenation, we get the joint feature V, and divide the feature map into n groups, which are recorded as:

[0039] V=[p1,p2,…,p n ]∈R m×k ,

[0040] Since the processing methods for n groups are the same, one of them is selected and recorded as:

[0041] p j =[q1,q2,…,q m ]∈R m×t ,j=1,2,…,n,

[0042] in Its sub - features are represented as:

[0043] q i =[a1,a2,…,a t T ∈R t , i = 1, 2, …, m.

[0044] During the learning process of the deep neural network, each group of features contains specific semantic information. The semantic information of each group can be approximately represented by calculating the average value of each group of sub - features. The calculation formula is:

[0045]

[0046] Multiplying this average vector back to each sub - feature through the dot - product method, the importance coefficient of each sub - feature in this group of text features can be obtained. This coefficient can be expressed as:

[0047]

[0048] This coefficient also measures to a certain extent the average semantic vector i and the similarity between the sub - feature q.

[0049] To eliminate the measurement influence between different samples, the obtained importance coefficients are normalized. This operation accelerates the convergence of the deep - learning model while maintaining the correlation between features. The specific operation is as follows:

[0050]

[0051] where ε is a constant introduced in the normalization process, and E(ω) and D(ω) represent the mean and variance of each group of importance coefficients respectively. The formula is:

[0052]

[0053] The normalized output is adjusted by the introduced scaling parameter g and translation parameter w to ensure that the text changes can be effectively represented after the normalization operation, and a new normalized importance coefficient is generated through the sigmoid function σ(·):

[0054]

[0055] Finally, the normalized importance coefficient is multiplied by the original sub - feature q i , strengthening the important features and weakening the unimportant features to obtain an enhanced sub - feature vector:

[0056] ​

[0057] All enhancer vectors form an enhancer vector group After the features of all groups are enhanced, they are re - spliced to obtain an enhanced paraphrased text feature representation

[0058] S3. Through multi - sentence feature fusion of the text features described in the above steps, integrating multi - perspective context information to capture deep - level text semantic features:

[0059] After steps S1 and S2, the corresponding feature matrices V o and V c are obtained for the original text o and the contradictory text c, while the two different degrees of paraphrased texts are transformed into a higher - quality paraphrased text feature matrix through grouped enhancement fusion Next, this step will perform multi - sentence feature fusion on these three feature matrices respectively, breaking the limitation of independent comparison between existing sentences, connecting the context before and after sentences, and obtaining a more comprehensive text feature representation. Since the processing methods for these three feature matrices are the same, the following will take the original text feature matrix V o as an example for introduction

[0060] The feature matrix of the original text where m is the number of sentences in the original text and d is the feature length of 768 encoded by BERT. Convolution is performed on the feature matrix V o using convolutional kernels of sizes 4×256, 3×256, 2×256, and 1×256 respectively. There are two convolutional kernels of each size, for a total of 8 convolutional kernels. These four sizes of convolutional kernels respectively correspond to grouping every 4 sentences as a group, every 3 sentences as a group, every 2 sentences as a group, and only performing feature extraction and fusion on each sentence itself. In order to make the features convolved by each convolutional kernel have the same size, padding needs to be performed during convolution to keep the vector dimensions consistent. When performing convolution for the three different - sized convolutional kernels of 4×256, 3×256, and 2×256, zero - filled matrices of 3×d, 2×d, and 1×d are used to pad the feature matrix V o respectively. After padding, the dimensions of the feature matrix are (m + 3)×d, (m + 2)×d, and (m + 1)×d respectively. During convolution, the convolutional stride set in the present invention in the horizontal direction is 256, so the size of the feature matrix in the horizontal direction after convolution is 3. After passing through 8 convolutional kernels, the feature matrix will obtain 8 multi - sentence feature matrices The process of multi - sentence feature fusion convolution is as Figure 3 As shown in the figure, the convolution kernel and the feature matrix obtained by using the convolution kernel are represented by the same color. In each multi-sentence feature matrix O, each small square represents the multi-sentence feature obtained by fusion of each sentence and the surrounding sentences. Then these feature matrices are max-pooled to reduce the feature dimension, and after multiple convolution and pooling operations, the 8 feature vectors are cascaded to form a comprehensive feature representation. Similarly, we can also get the final feature matrix of plagiarized text and contradictory text and

[0061] S4. Optimize the neural network model through triple loss and positive example enhancement loss function:

[0062] The present invention combines triplet loss and positive example enhancement loss to guide the proposed neural network model to detect text plagiarism more accurately.

[0063] The triplet loss function improves the model's ability to distinguish between positive and negative samples by shortening the distance between semantically similar texts in the feature space and increasing the distance between semantically dissimilar texts. The formula for the triplet loss function is as follows:

[0064]

[0065] in, is the feature matrix of the original text, is the feature matrix of the plagiarized text, is the feature matrix of the contradictory text, D() represents the distance between the two feature matrices in the brackets in the feature space, and margin is a constant greater than 0. The optimization goal is to bring and The distance in feature space, push away and distance, and and The distance in feature space is much smaller than and The distance between the two texts is 1.5. As a result, the trained neural network model can effectively distinguish plagiarized texts and will not regard contradictory texts as plagiarized texts.

[0066] Positive example enhancement loss optimizes the model's ability to capture text semantics by comparing the similarity between positive sample pairs and a large number of negative sample pairs. Its formula can be expressed as:

[0067]

[0068] Among them, sim() represents the cosine similarity of the two feature matrices in the brackets, and τ is the temperature parameter used to adjust the data distribution.

[0069] These two loss functions promote the learning of more robust text feature representations by the deep learning model proposed in this solution from different perspectives. The triplet loss function helps enhance the sensitivity of the model to the subtle differences between the plagiarized text and the original text, while the positive example enhancement loss further improves the model's understanding and representation ability of text semantics through contrastive learning with a large number of negative samples. The present invention combines these two loss functions with a certain weight ratio to obtain the final loss function, which can be expressed as:

[0070] J = αL + βl

[0071] S5. Design and construct a plagiarized text detection dataset based on the idea of contrastive learning through the dataset generation module.

[0072] The most common intelligent plagiarized text methods in real life are plagiarized text based on polishing and translation. Polishing plagiarized text is to use a large language model such as ChatGPT to polish the original text, which can evade traditional plagiarized text detection algorithms while retaining the semantics of the original text. Translating plagiarized text is to perform back translation between multiple languages through translation software, for example, translating the Chinese original text into English and then back into Chinese to achieve the purpose of plagiarized text. In this step, datasets for polishing plagiarized text and translating plagiarized text are respectively collected and constructed.

[0073] The original text of the polishing plagiarized text detection dataset is selected from the CSL dataset, which is provided by the National Engineering Technology Research Center for the Sharing Service of Scientific and Technological Resources and contains the meta-information (title, abstract, and keywords) of journal papers published from 2010 to 2020. Then, it is screened according to the Chinese core journal directory and labeled with subject and category tags, divided into 13 categories and 67 subjects. After the original text is determined, it is required that ChatGPT polish e% of the sentences in the original text while keeping the original semantics unchanged. By changing the proportion of the polished sentences, the present invention can obtain locally polished texts b with different degrees. Similarly, it is required that ChatGPT perform full-text polishing on the original text to obtain the full-text polished text h, aiming to improve the text quality while keeping the original semantics. The present invention designs the dataset using the idea of contrastive learning. Each group of data contains positive sample pairs and negative sample pairs. The positive sample pairs are the original text and two texts polished to different degrees, and the negative sample pairs are the original text and irrelevant contradictory texts. These four text information are used as a group of inputs and sent to steps S1 to S4 for learning.

[0074] The original text of the translation and paraphrasing detection dataset is selected from the LCSTS news summary dataset. After the original text is determined, first, a machine translator is randomly selected from four machine translators: Baidu Translate, Google Translate, Youdao Translate, and Bing Translate. Then, a transition language is randomly selected from languages such as Chinese, English, Spanish, German, Japanese, and Tibetan for translation. Finally, the translated and paraphrased text b and the translated and paraphrased text h are obtained by randomly translating 2 to 4 rounds. Similar to the paraphrasing detection dataset, it is also designed using the idea of contrastive learning. The positive sample pairs are the original text and the two translated and paraphrased texts, and the negative sample pairs are the original text and unrelated contradictory texts.

[0075] Finally, the training data generated in step S5 is sent to the deep neural network composed of steps S1 to S4 for training. This network can generate robust text feature representations by combining context. By simply comparing the distances between two input texts in the feature space, it can be determined whether it is paraphrased text.

[0076] To verify the technical effects of the present invention, the following experimental analysis is carried out:

[0077] The deep neural network described in steps S1 to S4 is built using the Pytorch framework. The number of network layers is 12. Each time 32 groups of text pairs are input during training. The training environment used is NVIDIA GeForce RTX 2080Ti. The dataset used for training is the paraphrasing detection dataset described in step S5, abbreviated as LMP.

[0078] The experimental content mainly includes the detection performance analysis based on cosine similarity and Spearman correlation coefficient, t-SNE visualization to analyze the quality of text representation, optimal loss function weight analysis, and ablation experiments.

[0079] 1. Detection performance analysis based on cosine similarity and Spearman correlation coefficient

[0080] In the task of detecting rewritten manuscripts, it is usually necessary to calculate the similarity between different texts. Cosine similarity evaluates the similarity between two text features by calculating the cosine value of the angle between them in the semantic feature space. The closer the cosine value is to 1, the more similar the two text features are. In addition to cosine similarity, another commonly used similarity evaluation metric is the Spearman correlation coefficient. It is a non-parametric statistical method that does not depend on the specific distribution of the data, but evaluates the relationship between two variables by ranking the data. After calculating the cosine similarity of all text representation pairs, the Spearman correlation coefficient is used to compare the correlation between the cosine similarity generated by the model and the manually marked similarity. The value range of the Spearman correlation coefficient is from -1 to 1. The closer the correlation coefficient is to 1 or -1, the stronger the correlation between the two texts. The closer the correlation coefficient is to 0, the weaker the correlation. The formula for the Spearman correlation coefficient is:

[0081]

[0082] where d i represents the rank difference between a pair of texts, and n represents the number of text pairs.

[0083] By adjusting the similarity threshold and comparing the Spearman correlation coefficients under different similarity thresholds, the threshold standard of the model is determined. As shown in Figure 4 , when the similarity threshold is 0.64 on the LMP dataset, the Spearman correlation coefficient of the model proposed in the present invention reaches the maximum of 0.8659. Therefore, in the task of detecting rewritten manuscripts, a rewritten manuscript threshold of 0.64 is set. When the cosine similarity of two text features exceeds 0.64, they will be determined as rewritten manuscripts. After setting the threshold, the results are shown in the following table. The cosine similarity of the polished text pair is 0.9855, while the cosine similarity of the contradictory text pair is 0.5627.

[0084]

[0085]

[0086] In addition, since the deep learning model proposed in the present invention can learn comprehensive and rich text feature representations, the model trained on the LMP dataset for detecting rewritten manuscripts in step S5 is directly tested on the dataset for detecting translated rewritten manuscripts mentioned above. The highest Spearman correlation coefficient reaches 0.8536, demonstrating that the method proposed in the present invention has good robustness.

[0087] 2. Use t-SNE to visually analyze the quality of text representations

[0088] To analyze the quality of the text representations generated by the deep learning model proposed in this invention, the text vectors generated by the model are clustered to observe whether sentences with similar or related semantics are accurately clustered together.

[0089] First, 100 original texts belonging to the engineering category, their polished and plagiarized texts, and 100 contradictory texts belonging to the agricultural category are respectively extracted from the LMP dataset for cleaning and preprocessing, including removing stop words and punctuation marks, etc. The original texts, polished and plagiarized texts, and contradictory texts are processed through the model and converted into fixed-length feature vectors, which contain the comprehensive semantic features of the texts. Then, the K-Means clustering algorithm is used for clustering operations. Before the clustering operation, all text vectors are standardized to eliminate the differences in features of different dimensions. The initial center selection of the K-Means algorithm adopts the K-Means++ method to avoid the problem of local optimal solutions and improve the quality and stability of clustering. The clustering results are visually displayed through the dimensionality reduction technique t-SNE. The distribution of the original texts, polished and plagiarized texts, and contradictory texts in the two-dimensional space is as Figure 5 shown. Through this visualization result, it can be found that the polished texts are very close to the original texts in the feature space, which reflects that the polishing operation retains the core semantics of the original text while introducing subtle changes. In contrast, the contradictory texts are far from the original texts and polished texts in the feature space, showing a large semantic difference from the original text. This result proves that the model proposed in this invention can be effectively used to detect text polishing and plagiarism.

[0090] 3. Optimal Loss Function Weight Analysis and Ablation Experiment

[0091] This invention combines the triplet loss and the positive example enhancement loss with a certain weight ratio to obtain the final loss function. To better find a set of weighting factors to fuse the two loss functions in this invention, thereby improving the accuracy of text similarity calculation and plagiarism detection. The control variable method is used to adjust the weighting factors. Given a similarity threshold, the value combinations of the weighting factors α and β are adjusted, and by observing the results of the Spearman correlation coefficient in various combination cases, the best value combination of the parameters is determined. The following table shows the results of some combinations.

[0092]

[0093] It can be seen from this that by reducing the weight of the triplet loss, the Spearman correlation coefficient value will increase accordingly. This is because the triplet loss pays more attention to accurately measuring the distance relationship between samples, emphasizing that the distance between the original text and the positive sample should be less than that between the negative sample. Therefore, its proportion is considered relatively small. By increasing the weighting factor of the positive example enhancement loss, the Spearman correlation coefficient value has a significant improvement. The positive example enhancement loss learns an effective representation of the data by maximizing the similarity between positive samples and minimizing the similarity between negative samples, and it can better express the original meaning of words, extract local and global features of the text, and mine the meaning information of the text itself. Therefore, its comprehensive influence is greater, so its proportion is considered relatively the largest. It can be obtained from the above table that the Spearman correlation coefficient value of the model reaches the optimum under the combination of weights of 0.25 and 0.75.

[0094] In addition, ablation experiments were also carried out to verify the roles of the multi-sentence feature fusion mechanism and the loss function. A total of four datasets with different degrees of polishing were used for the experiments to observe whether adding the multi-sentence feature fusion mechanism and the loss function could improve the model performance. The experimental results are shown in the following table.

[0095]

[0096] It can be seen from this that the multi-sentence feature fusion mechanism and the loss function proposed by the present invention can significantly improve the performance of the plagiarism detection task.

[0097] The number of devices and the processing scale described here are used to simplify the description of the present invention. The application, modification, and variation of the present invention are obvious to those skilled in the art. Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.< / pad>

Claims

1. A text rewriting detection method based on deep learning, characterized in that, It includes the following steps: S1. Perform feature encoding on the input text through the BERT model; S2. Obtain a higher-quality paraphrased text feature representation by grouping and fusing the paraphrased text features extracted in the above step; S3. Integrate multi-perspective context information and capture deep text semantic features by performing multi-sentence feature fusion on the text features in the above step; S4. Optimize the neural network model through the triplet loss and the positive example enhancement loss function; S5. Design and construct a paraphrasing detection dataset based on the contrastive learning idea through the dataset generation module.

2. The method for detecting text rewriting based on deep learning according to claim 1, characterized in that, The specific steps in step S2 include the following steps: S21. Concatenate the feature matrices of a certain amount of paraphrased text and the feature matrix of the full-text paraphrased text to obtain joint features, and divide the feature maps of the joint features into multiple groups; S22. Approximately represent the semantic information of each group by calculating the average value of each group of sub-features; S23. Multiply the average vector back to each sub-feature in the form of a dot product to obtain the importance coefficient of each sub-feature in this group of text features; S24. Perform normalization processing on the importance coefficient; S25. The output after normalization is adjusted by the introduced scaling parameter g and translation parameter w, and a new normalized importance coefficient is generated through the sigmoid function σ(·); S26. Multiply the normalized importance coefficient by the original sub - feature q i to obtain the enhanced sub - feature vector; S27. All enhanced sub-vectors are concatenated together to form an enhanced vector group. After the features of all groups are enhanced, they are re-concatenated to obtain an enhanced paraphrased text feature representation.

3. The method for detecting text rewriting based on deep learning according to claim 1, characterized in that, The specific steps in step S3 include the following steps: S31. Convolve the feature matrix of the original text through multiple convolution kernels of different sizes, and pad the feature matrix with all-zero matrices of 3×d, 2×d, and 1×d respectively to keep the output vector dimension consistent; S32. After convolving the feature matrix of the above original text, obtain multiple sentence feature matrices, perform max pooling on the sentence feature matrices to reduce the feature dimension, and after multiple convolution and pooling operations, concatenate multiple feature vectors to form a comprehensive feature representation; S33. Similarly obtain the final feature matrices of the manuscript text and the contradictory text through steps S31 - S32.

4. The method for detecting text rewriting based on deep learning according to claim 1, characterized in that, The input text in step S1 includes the original text, a small amount of paraphrased text, the full-text paraphrased text, and the contradictory text. The original text, a small amount of paraphrased text, the full-text paraphrased text, and the contradictory text are respectively sent into the BERT model to extract features and obtain the corresponding feature matrices.

5. The method for detecting text rewriting based on deep learning according to claim 4, wherein The specific steps in step S1 include the following steps: S11. Use special characters for the original text sentence set composed of m sentences of the input text <pad>"Expand the sentence length to 512;< / pad> S12. One-hot encode the words in the sentence to encode them into word vectors; S13. Input the obtained encoded vectors into the BERT model for feature extraction; S14. Concatenate all sentence features to obtain the feature matrix of the original text.

6. The method for detecting text rewriting based on deep learning according to claim 1, characterized in that, In step S5, datasets for polished paraphrasing and translated paraphrasing are respectively collected and constructed. The original texts of the polished paraphrasing detection dataset are selected from the CSL dataset. After the original text is determined, ChatGPT is required to polish e% of the sentences in the original text while keeping the original semantics unchanged; By changing the proportion of polished sentences, the present invention can obtain locally polished texts with different degrees; Furthermore, the idea of contrastive learning is adopted to design the dataset. The positive sample pairs are the original text and two texts polished to different degrees, and the negative sample pairs are the original text and irrelevant contradictory texts; After the original text is determined, through different translators, a random language is selected as the intermediate language for translation, and the translated paraphrasing texts are obtained by randomly translating 2 to 4 rounds respectively; Designed by adopting the idea of contrastive learning, the positive sample pairs are the original text and two translated paraphrasing texts, and the negative sample pairs are the original text and irrelevant contradictory texts.

Citation Information

Cited By

  • Visual arrangement system and method for large language model workflow

    CN121615664A

  • A large language model workflow visual arrangement system and method

    CN121615664B