Text data processing method and apparatus, text data detection method and apparatus
By employing deep learning and contrastive learning methods, an encoder model is trained using multi-domain text data. Combined with the MD-LoRA auxiliary network, this solves the problem of low efficiency in multi-domain text AIGC detection, achieving more efficient and accurate text detection.
Patent Information
- Application Number
- CN202411739848.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing technologies struggle to effectively detect AI-generated text content (AIGC), especially when the semantic distribution of texts across multiple domains is uneven, where traditional methods are inefficient and have limited effectiveness.
We employ a deep learning and contrastive learning approach, acquiring multi-domain artificial text and AIGC text, using an encoder model for feature extraction and contrastive learning loss adjustment, and combining it with an MD-LoRA auxiliary network for training to enhance the robustness and efficiency of the detection model.
It improves the robustness and efficiency of AIGC text detection, effectively distinguishing between artificial text and AIGC text in different fields, and reducing the detection difficulty and resource load.
Smart Images

Figure CN119621984B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and contrastive learning, and particularly to a text data processing method and apparatus, and a text data detection method and apparatus. Background Technology
[0002] Currently, large-scale model tools and robots such as Keling AI and GPT-4V have the ability to process multi-modal data, including text, images, and videos, and can automatically output Artificial Intelligence Generated Content (AIGC). Among these, the continuous refinement of technology in the text modality best demonstrates the rapid development of large-scale model technology. Training strategies for other modalities are also basically based on text-domain technologies with minor modifications. Through training strategies such as collecting large-scale noisy data for text pre-training, fine-tuning instructions with carefully designed template data, and aligning with human-labeled data based on human preferences, large-scale models are gradually approaching the writing habits of real humans, and the difficulty of detecting text AIGC continues to improve.
[0003] Manual detection methods have become extremely difficult and inefficient. Furthermore, as large-scale model technology advances over time, AIGC (Artificial Intelligence Generated Text) has emerged in multiple domains within text, making AIGC identification by experts in those domains highly manpower-intensive. Additionally, the varying degrees of understanding of semantic distribution across different domains by large models alter the detection difficulty. Traditional AIGC detection methods, aside from white-box solutions that are inaccessible to ordinary users or developers, often employ black-box methods—end-to-end training using "artificial text - AIGC text" samples—which typically involve joint training with data from multiple domains. This ignores the differences in text distribution across domains and the varying degrees of understanding of semantics by large models within those domains, resulting in limited effectiveness for current text AIGC detection. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a text data processing method and apparatus, and a text data detection method and apparatus, which enhance the effectiveness of existing AIGC detection models by employing a multi-domain text AIGC detection enhancement method based on deep learning and contrastive learning techniques.
[0005] In a first aspect, the present invention provides a text data processing method, the method comprising:
[0006] Acquire multi-domain artificial text, which includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data.
[0007] Multi-domain artificial texts are input into a large model, and multi-domain AIGC texts are output; the multi-domain AIGC texts are generated at the document granularity and sentence granularity respectively according to heuristic rules for each artificial text.
[0008] The encoder model to be trained is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text, so as to obtain the artificial text encoding vector and the AIGC text encoding vector at each level. The encoder model includes a multi-level encoder network and a corresponding multi-level auxiliary encoding network. The encoder network includes several sub-network layers at multiple levels, and the auxiliary encoding network includes several auxiliary encoding layers at multiple levels. Each sub-network layer is connected to the corresponding level of auxiliary encoding layer.
[0009] A contrastive learning loss is constructed based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and the model parameters of the auxiliary encoding network are adjusted based on the contrastive learning loss to obtain the trained encoder model.
[0010] In a second aspect, the present invention also proposes a text data detection method, the method comprising:
[0011] Obtain the text data to be tested;
[0012] The text data to be tested is input into the trained encoder model, and the text encoding vector of the text data to be tested is output; the encoder model is the trained encoder model described in the text data processing method of the first aspect of the present invention;
[0013] The text encoding vector of the text data to be tested is detected to obtain the text detection result of the text data to be tested.
[0014] The text detection result is either artificial text or AIGC text, which indicates the source of the text content of the text data to be tested.
[0015] In a third aspect, the present invention also provides a text data processing apparatus, the apparatus comprising:
[0016] The first acquisition unit is used to acquire multi-domain artificial text, which includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data.
[0017] An enhanced processing unit is used to input multi-domain artificial text into a large model and output multi-domain AIGC text; the multi-domain AIGC text is generated for each artificial text according to heuristic rules at both document granularity and sentence granularity.
[0018] The first extraction unit is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text using the encoder model to be trained, so as to obtain the artificial text encoding vector and the AIGC text encoding vector at each level; the encoder model includes a multi-level encoder network and corresponding multi-level auxiliary encoding networks.
[0019] The first adjustment unit is used to construct a loss based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and to adjust the model parameters of the auxiliary encoding network based on the loss to obtain the trained encoder model.
[0020] In a fourth aspect, the present invention also provides a text data detection device, the device comprising:
[0021] The second acquisition unit is used to acquire the text data to be tested;
[0022] The second extraction unit is used to extract features from the text data to be tested using an encoder model to obtain the text encoding vector of the text data to be tested. The encoder model is the encoder model trained in the text data processing method described in the first aspect of the present invention.
[0023] The detection unit is used to detect the text encoding vector of each text data to be tested, and obtain the text detection result corresponding to each text data to be tested; the text detection result is artificial text or AIGC text, which is used to indicate the source of the text content of the text data to be tested.
[0024] The beneficial effects of this invention are:
[0025] 1. The heuristic rules for generating multi-domain AIGC text using large model tools in this invention can ensure coverage of various AIGC scenarios in various domains. In addition to the commonly used text rewriting, correction and polishing using large models, there are also cases of using large models for continuation writing, or translation followed by polishing and continuation writing, or only modifying some sentences in the text. This fully guarantees the comparative learning effect of the model and improves the robustness of the model in the field of text AIGC detection.
[0026] 2. This invention employs a multi-domain joint low-rank adaptive augmentation network (MD-LoRA), also known as an auxiliary encoding network, added to each layer of the encoder model. During training, the encoder model is frozen, and only the corresponding MD-LoRA auxiliary network is trained. This accelerates the model's training speed and reduces resource load, while providing a pluggable text AIGC detection enhancement scheme. The original model can be used for detection when appropriate, and the MD-LoRA auxiliary network can be combined for processing when facing challenging samples. Simultaneously, when text vectors pass through the MD-LoRA auxiliary network, the auxiliary encoding dimensionality reduction module, composed of unbiased linear layers, maintains scaling invariance to compress vectors into an abstract semantic low-rank space, preserving the uniformity of the compressed feature space. Meanwhile, the encoding dimensionality enhancement module, composed of biased linear layers, ensures that different translational distributions are used to enhance the dimensionality of text from different domains to the corresponding semantically separable space, dynamically adjusting the differences in domain distributions and preventing semantic distributions from interfering with each other.
[0027] 3. This invention employs a multi-level knowledge-protected contrastive learning loss function to simultaneously record multiple hidden layer result vectors. It dynamically samples hidden layer vectors based on the network layer number to calculate the loss, deepening the contrast process between artificial text and AIGC text to multiple layer nodes within the encoder model. Simultaneously, it merges the text vector encoded by the MD-LoRA auxiliary network (which acts as an update network) with the encoder text vector, bringing the semantic space distribution closer to the encoder text vector and distancing it from the AIGC text. This ensures that the semantic distribution interval between artificial text and AIGC text is increased without excessive loss of prior information in the encoder model. By combining the contrastive loss obtained within a batch, the contrastive learning effect is enhanced, reducing the AIGC detection difficulty of existing black-box models. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating the text data processing method according to an embodiment of the present invention;
[0030] Figure 2 This is a diagram illustrating the heuristic rule principle for generating multi-domain AIGC text using a large model tool in an embodiment of the present invention.
[0031] Figure 3 This is a schematic diagram of the auxiliary coding network structure according to an embodiment of the present invention;
[0032] Figure 4This is a flowchart illustrating the text data detection method according to an embodiment of the present invention;
[0033] Figure 5 This is a schematic diagram of the structure of the text data processing device according to an embodiment of the present invention;
[0034] Figure 6 This is a schematic diagram of the structure of the text data detection device according to an embodiment of the present invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] This invention provides a text data processing method and apparatus, and a text data detection method and apparatus. The text data processing method can be used in the text data processing apparatus, and the text data detection method can be used in the text data detection apparatus. The text data processing apparatus can be integrated into a computer device, and the text data detection apparatus can also be integrated into a computer device, which can be a terminal or a server. The terminal can be a mobile phone, tablet computer, laptop computer, smart TV, wearable smart device, personal computer (PC), or vehicle terminal, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The server can also be a node in a blockchain.
[0037] Please see Figure 1 ,like Figure 1 The diagram shown is a flowchart of the text data processing method provided in this application. The method includes:
[0038] 101. Obtain multi-domain artificial text, wherein the multi-domain artificial text includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data;
[0039] Artificial text is text directly created and written by humans. It has wide applications in various fields, such as education, scientific research, media, and advertising. In education, textbooks and lesson plans written by teachers are examples of artificial text; in scientific research, papers and reports written by researchers are also important components of artificial text. Artificial text can take the form of words, phrases, sentences, etc., and can include textual information such as characters, numbers, symbols, etc. Artificial text can be in any language, such as Chinese or English, or it can be a mixture of different languages, such as a mixture of Chinese and English. This application does not limit the form of artificial text.
[0040] In the text data processing method provided in this application embodiment, the large amount of artificial text obtained does not need to be labeled, which can save a lot of time in adding labels and thus greatly improve the efficiency of training the model.
[0041] 102. Input multi-domain artificial texts into a large model and output multi-domain AIGC texts; the multi-domain AIGC texts are generated for each artificial text according to heuristic rules at the document granularity and sentence granularity respectively.
[0042] In this embodiment of the invention, AIGC text is text automatically generated using artificial intelligence techniques such as large model tools. This embodiment does not limit the specific large language model used.
[0043] In embodiments of the present invention, such as Figure 2 As shown, generating AIGC text for the multi-domain applications according to heuristic rules at both document and sentence granularity levels includes: generating AIGC translated text, polished text, and continuation text using a combination of one or more methods among translation, polishing, and continuation based on prompts; and generating semantically granular AIGC text by polishing sentence data. The document data consists of man-made text from each domain; the sentence data is obtained by sampling sentence-granular text clusters according to a Gaussian distribution; and the sentence-granular text clusters are obtained by segmenting man-made text from each domain into sentences based on punctuation marks.
[0044] In some embodiments, the statement data is polished, and the polishing of the statement data to generate semantic granular AIGC text includes polishing the statement data to generate sampled polished text; inserting the sampled polished text into the semantic granular text cluster to generate AIGC sentence clusters; and merging the AIGC sentence clusters according to punctuation marks to generate semantic granular AIGC text.
[0045] In some embodiments, exemplarily, the step of generating multi-domain AIGC text using a large model tool according to heuristic rules includes:
[0046] Step 201: Input multi-domain artificial text into the large model tool and polish the text using prompts to obtain multi-domain AIGC polished text;
[0047] Step 202: After segmenting the multi-domain artificial text into sentences according to punctuation marks, merge the sentences in order from front to back until the length of the merged text reaches half the length of the original text to obtain truncated text. Input the truncated text into the large model tool and use the prompt to continue writing the text to obtain the multi-domain AIGC continuation text.
[0048] Step 203: Input the Chinese text from the multi-domain artificial text into the large model tool and translate it into English text using the prompt. Then, follow the steps 201 and 202 to obtain the English multi-domain AIGC polished text and the multi-domain AIGC continuation text, respectively.
[0049] Step 204: Input the English text from the multi-domain artificial text into the large model tool and translate it into Chinese text using the prompt. Then, follow the steps 201 and 202 to obtain the Chinese multi-domain AIGC polished text and multi-domain AIGC continuation text respectively.
[0050] Step 205: Segment the multi-domain artificial text according to punctuation marks to obtain sentence-level text clusters. Sample the sentences in the text clusters according to Gaussian distribution, and input the sampled sentences into the large model tool for polishing with prompts. Insert the polished text back into the sentence-level text clusters in the original order to obtain AIGC sentence clusters. Finally, merge the sentence clusters according to the original punctuation marks to obtain sentence-level AIGC text.
[0051] The heuristic rules for generating multi-domain AIGC text using large model tools in this invention can ensure coverage of various AIGC scenarios in different domains. In addition to the commonly used methods of rewriting, correcting, and polishing text using large models, there are also cases where large model tools are used for continuation writing, or translation followed by polishing and continuation writing, or only modifying parts of the text. This fully guarantees the comparative learning effect of the model and improves the robustness of the model in the field of text AIGC detection.
[0052] 103. The encoder model to be trained is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text, so as to obtain the artificial text encoding vector and AIGC text encoding vector at each level.
[0053] The encoder model includes a multi-level encoder network and corresponding multi-level auxiliary coding networks; the encoder network includes several sub-network layers of multiple levels, and the auxiliary coding network includes several auxiliary coding layers of multiple levels, with each sub-network layer connected to an auxiliary coding layer of the corresponding level.
[0054] In this embodiment of the invention, the encoder model can be a Transformers network. In each network layer of the encoder network (such as the encoder network of the i-th layer), there is a block. The block consists of attention blocks and multilayer perceptron blocks. The attention block contains multiple linear layers and a normalization network. The multilayer perceptron block also contains multiple linear layers. Each such linear layer is a sub-network layer. Each sub-network layer corresponds to an auxiliary coding layer. All the auxiliary coding layers together constitute an auxiliary coding network called the MD-LoRA auxiliary network.
[0055] Therefore, the encoder model in this embodiment adds an auxiliary encoder network after each network layer of the traditional encoder network. During training, the encoder model is frozen and only the corresponding MD-LoRA auxiliary network is trained. This speeds up the training of the model and reduces resource load. It also provides a pluggable text AIGC detection enhancement scheme for the model. The original encoder network can be used for detection when appropriate, and the MD-LoRA auxiliary network can be combined for relevant processing when facing difficult samples. This can improve the robustness of the model.
[0056] In embodiments of the present invention, such as Figure 3 As shown, the auxiliary coding network includes an auxiliary coding dimensionality reduction module and an auxiliary coding dimensionality increase module corresponding to the number of domains. The auxiliary coding dimensionality reduction module includes a linear layer with no bias term, and the auxiliary coding dimensionality increase module includes a linear layer with a bias term. When text vectors pass through the MD-LoRA auxiliary network, the auxiliary coding dimensionality reduction module, composed of a linear layer with no bias term, can maintain scaling invariance and help compress the vectors to an abstract semantic low-rank space, keeping the distribution compression feature space uniform. Meanwhile, the coding dimensionality increase module, composed of a linear layer with bias term, ensures that different translation distributions are used to increase the dimensionality of text from different domains to the corresponding semantically separable space, dynamically adjusting the differences in domain distributions and preventing semantic distributions from interfering with each other.
[0057] In some embodiments, exemplarily, the encoder network may include multiple Transformer encoders, wherein each encoder may contain multiple hidden layers and multiple self-attention heads. Preferably, the encoder network may include 48 Transformer encoders, each encoder containing 1024 hidden layers and 16 self-attention heads, etc.
[0058] In some embodiments, exemplarily, the process of extracting the artificial text encoding vector and the AIGC text encoding vector at each level includes:
[0059] 301. Pass the artificial text encoding vector generated by the previous encoder network layer and the AIGC text encoding vector generated by the previous auxiliary encoding network layer through each sub-network layer of the current encoder network layer to obtain the artificial text intermediate encoding vector and AIGC text intermediate encoding vector generated by each sub-network layer of the current encoder network layer.
[0060] 302. Use the intermediate encoding vector of the artificial text generated by the last layer of the current encoder network layer as the encoding vector of the artificial text generated by the current encoder network layer.
[0061] 303. Pass the artificial text intermediate coding vector and AIGC text intermediate coding vector generated by each sub-network layer of the current encoder network layer through the auxiliary coding dimensionality reduction module of each corresponding auxiliary coding layer of the current auxiliary coding network layer to obtain the artificial text low-rank coding intermediate vector and AIGC text low-rank coding intermediate vector generated by each auxiliary coding layer of the current auxiliary coding network layer.
[0062] 304. Pass the low-rank encoded intermediate vectors of the artificial text and the low-rank encoded intermediate vectors of the AIGC text generated by each auxiliary coding layer of the current auxiliary coding network layer through the auxiliary coding dimension-upgrading module of each corresponding auxiliary coding layer of the current auxiliary coding network layer to obtain the high-dimensional encoded intermediate vectors of the artificial text and the high-dimensional encoded intermediate vectors of the AIGC text generated by each auxiliary coding layer of the current auxiliary coding network layer.
[0063] 305. Add the artificial text intermediate encoding vector and AIGC text intermediate encoding vector generated by each sub-network layer of the current encoder network layer to the artificial text high-dimensional encoding intermediate vector and AIGC text high-dimensional encoding intermediate vector generated by each corresponding auxiliary encoding layer of the current auxiliary encoding network layer, and obtain the artificial text encoding vector and AIGC text encoding vector generated by the current auxiliary encoding network layer.
[0064] In this embodiment of the invention, in the encoder model, each level of the encoder network is configured with a corresponding auxiliary encoding network; initially, each artificial text T is... person and corresponding AIGC text T aigc The initial layer of the artificial text encoding vector is obtained by passing it through the embedding layer of the encoder network. With AIGC text encoding vector Combining the two vectors above yields the combined vector. vector Combined vectors Input the subsequent hidden layers of the encoder network and the corresponding auxiliary encoder network for each layer to generate the intermediate encoded vector of the artificial text generated by the initial layer encoder network. With AIGC text intermediate encoding vector The artificial text encoding vector generated by the previous auxiliary coding network is compared with the AIGC text encoding vector. The low-rank encoded intermediate vector generated by the initial-level auxiliary coding network is obtained through the auxiliary coding dimensionality reduction module of the initial-level auxiliary coding network. The low-rank encoded intermediate vector generated by the initial level auxiliary coding network The high-dimensional intermediate encoded vector generated by the current-level auxiliary encoding network is obtained through the auxiliary encoding upscaling module of the current-level auxiliary encoding network. The artificial text encoding vector generated by the current level encoder network With AIGC text encoding vector and the high-dimensional encoding intermediate vector of the current hierarchical auxiliary encoding network By concatenating the vectors, we obtain the artificial text encoding vector generated by the current level auxiliary encoding network and the AIGC text encoding vector. Similarly, the encoding vectors for the remaining levels can also be obtained in the same way.
[0065] In this embodiment of the invention, the artificial text encoding vectors at each level and the AIGC text encoding vectors can be constructed into a multi-level, multi-domain encoding vector group G. v ={l1, l2, ..., l n} can be represented as:
[0066]
[0067] In the formula, l i This represents the result vector set of the encoder network at layer i and the corresponding auxiliary encoder network, where n represents the number of hidden layers in the encoder model. i Let A represent the encoder network of layer i. i and These represent the auxiliary coding dimensionality reduction module of the auxiliary coding network corresponding to the i-th layer encoder network and the j-th type of auxiliary coding dimensionality increase module corresponding to the data domain, respectively, where m represents the number of data domains.
[0068] 104. Construct a contrastive learning loss based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and adjust the model parameters of the auxiliary encoding network based on the contrastive learning loss to obtain the trained encoder model.
[0069] In this embodiment of the invention, the loss constructed based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level includes:
[0070] 401. Select a portion of artificial text encoding vectors and AIGC text encoding vectors at the same level based on a preset hierarchical interval;
[0071] In this embodiment of the invention, the step of selecting a portion of the artificial text encoding vector and AIGC text encoding vector of the same level based on a preset level interval includes selecting a portion of the level according to a first interval when the level is within a first threshold range; selecting a portion of the level according to a second interval when the level is within a second threshold range; selecting a portion of the level according to a third interval when the level is within a third threshold range; and selecting a portion of the level according to a fourth interval when the level is within a fourth threshold range.
[0072] This embodiment can encode vector groups G from multiple levels and multiple domains. v The result vector group l corresponding to the index layer is extracted according to the interval value I. i This reduces some computational load while widening the potential semantic space distribution variation space within the intermediate layers, increasing the semantic receptive field of each intermediate layer, and enhancing the alignment effect brought about by contrastive learning.
[0073] For example, the formula for calculating the interval value is shown below:
[0074]
[0075] In the formula, n represents the number of hidden layers in the encoder model. This represents the floor symbol.
[0076] 402. The first loss is obtained based on the similarity between the classification head of the artificial text encoding vector generated by the encoder network at the same level and the classification head of the artificial text encoding vector generated by the auxiliary encoding network.
[0077] For example, compute the result vector group l of the i-th hidden layer (not the last layer). i The contrast loss is calculated first by taking the encoder network result vector. The "[CLS]" vector Take the result vector of the MD-LoRA auxiliary network The "[CLS]" vector Calculate vectors with vector The cosine similarity is used to calculate the cross-entropy loss with a label of 1. Obtain the semantic contrast loss of the i-th hidden layer. i .
[0078] For example, the result vector group l of the nth hidden layer (the last layer) is calculated. i The contrast loss is calculated first by taking the encoder network result vector. The "[CLS]" vector Take the result vector of the MD-LoRA auxiliary network The "[CLS]" vector Obtain the semantic contrast loss of the last hidden layer. n .
[0079] The first loss can be obtained by combining the semantic losses of each non-last layer and the semantic loss of the last layer. The first loss can bring the semantic space distribution of artificial text vectors closer together and increase the semantic distribution gap between artificial text and AIGC text.
[0080] 403. The second loss is obtained based on the similarity between the classification head of the artificial text encoding vector generated by the auxiliary coding network at the same level and the classification head of the AIGC text encoding vector generated by the auxiliary coding network.
[0081] For example, compute the result vector group l of the i-th hidden layer (not the last layer). i The contrast loss is calculated first by taking the encoder network result vector. The "[CLS]" vector Take the result vector of the MD-LoRA auxiliary network The "[CLS]" vector Calculate vectors with vector The cosine similarity is used to calculate the cross-entropy loss with the label set to 0. Obtain the semantic contrast loss of the i-th hidden layer. i .
[0082] For example, the result vector group l of the nth hidden layer (the last layer) is calculated. i The contrast loss is calculated first by taking the encoder network result vector. The "[CLS]" vector Take the result vector of the MD-LoRA auxiliary network The "[CLS]" vector Obtain the semantic contrast loss of the last hidden layer. n .
[0083] The second loss can be obtained by combining the semantic loss of each non-last layer and the semantic loss of the last layer. The second loss can widen the semantic space distribution of AIGC text vectors and increase the semantic distribution interval between artificial text and AIGC text.
[0084] 404. The first matrix is obtained by multiplying the classification head vector group of the artificial text encoding vector generated by the multilayer auxiliary coding network with the classification head vector group of the artificial text encoding vector generated by the multilayer encoder network.
[0085] For example, combining the vectors of all text within a batch. Get the vector group cls person_v cls person_r cls aigc Take the vector group cls person_v With vector group cls person_r Multiplying them together yields the similarity matrix M. pv_pr As the first matrix.
[0086] 405. The second matrix is obtained by multiplying the classification head vector group of the artificial text encoding vector generated by the multilayer auxiliary coding network with the classification head vector group of the AIGC text encoding vector generated by the multilayer auxiliary coding network.
[0087] For example, combining the vectors of all text within a batch. Get the vector group cls person_v cls person_r cls aigc Take the vector group cls person_v With vector group cls aigc Multiplying them together yields the similarity matrix M. pv_aigc As the second matrix.
[0088] 406. The third loss is obtained by multiplying the diagonal text pair vectors of the first matrix with the classification head vector group of the artificial text encoding vectors generated by the multilayer encoder network.
[0089] For example, for the first matrix M pv_pr The InfoNCE loss is calculated using diagonal text pair vectors as positive samples. As the third loss.
[0090] 407. Based on the cross-entropy relationship between the second matrix and the labels, the fourth loss is obtained.
[0091] For example, for the fourth matrix M pv_aigc Calculate cross-entropy loss with label 0. As the fourth loss.
[0092] In this embodiment of the invention, the above losses can be combined to obtain the final contrastive learning loss, and the calculation formula for the contrastive learning loss can be expressed as:
[0093]
[0094]
[0095] In the formula, Cross_Entropy(·) represents the cross-entropy loss function, Cosine_Similarity(·) represents the cosine similarity calculation function, and bs represents the batch size during training.
[0096] Therefore, the text data processing method provided in this application employs a large number of AIGC texts generated according to heuristic rules for sample augmentation and expansion. Then, a comparative learning method is used to train the encoder model based on multi-domain artificial texts and the expanded AIGC texts, enabling the encoder model to learn accurate text attribute feature extraction capabilities. During training, the encoder model is frozen and only the corresponding MD-LoRA auxiliary network is trained. This accelerates the model's training speed and reduces resource load, while providing a plug-in text AIGC detection augmentation scheme for the model. The original model can be used for detection when appropriate, and the MD-LoRA auxiliary network can be combined for relevant processing when facing difficult samples. At the same time, when the text vector passes through the MD-LoRA auxiliary network, the auxiliary encoding dimensionality reduction module composed of unbiased linear layers can maintain scaling invariance to help compress the vector to the abstract semantic low-rank space, maintaining the uniformity of the distributed compressed feature space. Meanwhile, the encoding dimensionality enhancement module composed of biased linear layers ensures that different translation distributions are used to enhance the dimensionality of texts from different domains to the corresponding semantically separable space, dynamically adjusting the differences in domain distributions and preventing semantic distributions from interfering with each other; thus greatly improving the training efficiency of the model.
[0097] This application also provides a text data detection method, which can be specifically applied to a text data detection device, which can be installed in a terminal. For example... Figure 4 The diagram shown is a flowchart of the text data detection method provided in this application. The method includes:
[0098] 111. Obtain the text data to be tested;
[0099] After training the encoder model, the trained encoder model can be deployed on the terminal, and corresponding text data detection tasks can be carried out based on the deployed model.
[0100] Specifically, the terminal can obtain the text data to be tested, for example, it can obtain the text data to be tested from a public database.
[0101] 112. Input the text data to be tested into the trained encoder model, and output the text encoding vector of the text data to be tested; the encoder model is the trained encoder model described in the text data processing method.
[0102] In this embodiment of the invention, the text data to be tested can be input into the trained encoder model for feature extraction to obtain the text encoding vector of the text data to be tested. Specifically, the encoder model here can be the encoder model trained in the above embodiment.
[0103] 113. Detect the text encoding vector of the text data to be tested to obtain the text detection result of the text data to be tested;
[0104] The text detection result is either artificial text or AIGC text, which indicates the source of the text content of the text data to be tested.
[0105] In this embodiment of the invention, the text detection result of the text data to be tested can be further determined by detecting the text encoding vector of the text data to be tested through network structures such as fully connected layers, classifiers, and decoder models.
[0106] For example, when using a decoder model, the main task of the decoder model is to convert the text encoding vector of the text data to be tested output by the encoder model into the target sequence. The decoder model can use any existing pre-trained decoder model, such as the decoder model in the Transformer model. Through the processing of the decoder model, the text detection result of the text data to be tested can be obtained.
[0107] As described above, the text data detection method provided in this application acquires text data to be tested; inputs the text data to be tested into a trained encoder model, and outputs the text encoding vector of the text data to be tested; the encoder model is the trained encoder model described in the text data processing method of the first aspect of this invention; and detects the text encoding vector of the text data to be tested to obtain the text detection result of the text data to be tested. Thus, by improving the training efficiency of the encoder model, the efficiency of text data detection can be improved. Moreover, using the encoder model provided in this application to perform text detection on the text data to be tested can also greatly improve the accuracy of text detection.
[0108] To better implement the above text data processing methods, embodiments of this application also provide a text data processing apparatus, which can be integrated into a terminal or server. For example, Figure 5 The diagram shown is a structural schematic of a text data processing apparatus provided in an embodiment of this application. The apparatus includes:
[0109] The first acquisition unit 121 is used to acquire multi-domain artificial text, which includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data.
[0110] The enhancement processing unit 122 is used to input multi-domain artificial text into a large model and output multi-domain AIGC text; the multi-domain AIGC text is generated by each artificial text according to heuristic rules at the document granularity and sentence granularity respectively.
[0111] The first extraction unit 123 is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text using the encoder model to be trained, so as to obtain the artificial text encoding vector and the AIGC text encoding vector at each level; the encoder model includes a multi-level encoder network and a corresponding multi-level auxiliary encoding network.
[0112] The first adjustment unit 124 is used to construct a loss based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and to adjust the model parameters of the auxiliary encoding network based on the loss to obtain the trained encoder model.
[0113] As described above, the text data processing apparatus provided in this application embodiment acquires multi-domain artificial text through a first acquisition unit. The multi-domain artificial text includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data; the multi-domain artificial text is input into a large model through an enhancement processing unit, and multi-domain AIGC text is output; the multi-domain AIGC text is generated at the document granularity and sentence granularity respectively according to heuristic rules for each artificial text; the first extraction unit uses an encoder model to be trained to extract features from the enhanced text data composed of each artificial text and the corresponding AIGC text, obtaining the artificial text encoding vector and AIGC text encoding vector at each level; the encoder model includes a multi-level encoder network and corresponding multi-level auxiliary encoding networks; the first adjustment unit constructs a loss based on at least two artificial text encoding vectors and AIGC text encoding vectors at the same level, and adjusts the model parameters of the auxiliary encoding network based on the loss, to obtain the trained encoder model.
[0114] To better implement the above text data detection method, embodiments of this application also provide a text data detection device, which can be integrated into a terminal or server. For example, Figure 6 The diagram shown is a structural schematic of the text data detection device provided in an embodiment of this application. The device may include:
[0115] The second acquisition unit 131 is used to acquire the text data to be tested;
[0116] The second extraction unit 132 is used to extract features from the text data to be tested using an encoder model to obtain the text encoding vector of the text data to be tested. The encoder model is the encoder model trained by the text data processing method described above.
[0117] The detection unit 133 is used to detect the text encoding vector of each text data to be tested and obtain the text detection result corresponding to each text data to be tested; the text detection result is artificial text or AIGC text, which is used to indicate the source of the text content of the text data to be tested.
[0118] As described above, the text data detection device provided in this application acquires the text data to be tested through a second acquisition unit. A second extraction unit uses an encoder model to extract features from the text data to be tested, obtaining a text encoding vector for the text data. A detection unit then detects the text encoding vector of each text data to be tested, obtaining a text detection result for each text data. The text detection result is either artificial text or AIGC text, used to represent the source of the text content of the text data to be tested. Therefore, by improving the training efficiency of the encoder model, the efficiency of text data detection can be improved. Furthermore, using the text data detection device provided in this application to detect the text data to be tested can also significantly improve the accuracy of text detection.
[0119] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.
[0120] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A text data processing method, characterized in that, The method includes: Acquire multi-domain artificial text, which includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data. Multi-domain artificial texts are input into a large model, and multi-domain AIGC texts are output; the multi-domain AIGC texts are generated at the document granularity and sentence granularity respectively according to heuristic rules for each artificial text. Generating AIGC texts for the multi-domain applications according to heuristic rules at both document and sentence granularity levels includes: generating AIGC translated text, polished text, and continuation text using a combination of one or more methods (translation, polishing, and continuation) based on prompts; and generating semantically granular AIGC text by polishing sentence data. The document data consists of man-made texts from each domain; the sentence data is obtained by sampling sentence-granular text clusters according to a Gaussian distribution; and the sentence-granular text clusters are obtained by segmenting man-made texts from each domain into sentences based on punctuation marks. The encoder model to be trained is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text, so as to obtain the artificial text encoding vector and the AIGC text encoding vector at each level. The encoder model includes a multi-level encoder network and a corresponding multi-level auxiliary encoding network. The encoder network includes several sub-network layers at multiple levels, and the auxiliary encoding network includes several auxiliary encoding layers at multiple levels. Each sub-network layer is connected to the corresponding level of auxiliary encoding layer. A contrastive learning loss is constructed based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and the model parameters of the auxiliary encoding network are adjusted based on the contrastive learning loss to obtain the trained encoder model. The auxiliary coding network includes an auxiliary coding dimensionality reduction module and an auxiliary coding dimensionality increase module corresponding to the number of domains; the auxiliary coding dimensionality reduction module includes a linear layer without bias terms, and the auxiliary coding dimensionality increase module includes a linear layer with bias terms. The comparative learning loss constructed based on at least two artificial text encoding vectors of the same level and AIGC text encoding vectors includes: Select a portion of the same level of artificial text encoding vectors and AIGC text encoding vectors based on a preset hierarchical interval; The first loss is obtained based on the similarity between the classifier head of the artificial text encoding vector generated by the encoder network at the same level and the classifier head of the artificial text encoding vector generated by the auxiliary encoding network. The second loss is obtained based on the similarity between the classifier head of the artificial text encoding vector generated by the auxiliary coding network at the same level and the classifier head of the AIGC text encoding vector generated by the auxiliary coding network. The first matrix is obtained by multiplying the classification head vector group of the artificial text encoded vector generated by the multilayer auxiliary coding network with the classification head vector group of the artificial text encoded vector generated by the multilayer encoder network. The second matrix is obtained by multiplying the classification head vector group of the artificial text encoding vector generated by the multilayer auxiliary coding network with the classification head vector group of the AIGC text encoding vector generated by the multilayer auxiliary coding network. The third loss is obtained by multiplying the diagonal text pair vectors of the first matrix with the classification head vector group of the artificial text encoding vectors generated by the multilayer encoder network; The fourth loss is obtained based on the cross-entropy relationship between the second matrix and the labels.
2. The text data processing method according to claim 1, characterized in that, The process of refining sentence data to generate semantic granular AIGC text includes refining the sentence data to generate sampled refined text; inserting the sampled refined text into a semantic granular text cluster to generate an AIGC sentence cluster; and merging the AIGC sentence clusters according to punctuation marks to generate semantic granular AIGC text.
3. The text data processing method according to claim 1, characterized in that, The process of obtaining the artificial text encoding vector and AIGC text encoding vector at each level includes: By passing the artificial text encoding vector generated by the previous encoder network layer and the AIGC text encoding vector generated by the previous auxiliary encoder network layer through each sub-network layer of the current encoder network layer, we obtain the artificial text intermediate encoding vector and the AIGC text intermediate encoding vector generated by each sub-network layer of the current encoder network layer. Use the intermediate encoding vector of the artificial text generated by the last layer of the current encoder network layer as the encoding vector of the artificial text generated by the current encoder network layer. The artificial text intermediate coding vectors and AIGC text intermediate coding vectors generated by each sub-network layer of the current encoder network layer are passed through the auxiliary coding dimensionality reduction module of each auxiliary coding layer of the current auxiliary coding network layer to obtain the artificial text low-rank coding intermediate vectors and AIGC text low-rank coding intermediate vectors generated by each auxiliary coding layer of the current auxiliary coding network layer. The low-rank encoded intermediate vectors of the artificial text and the low-rank encoded intermediate vectors of the AIGC text generated by each auxiliary coding layer of the current auxiliary coding network layer are passed through the auxiliary coding dimension-up module of each auxiliary coding layer of the current auxiliary coding network layer to obtain the high-dimensional encoded intermediate vectors of the artificial text and the AIGC text generated by each auxiliary coding layer of the current auxiliary coding network layer. The artificial text intermediate encoding vector and AIGC text intermediate encoding vector generated by each sub-network layer of the current encoder network layer are added to the artificial text high-dimensional encoding intermediate vector and AIGC text high-dimensional encoding intermediate vector generated by each corresponding auxiliary encoding layer of the current auxiliary encoding network layer, in order to obtain the artificial text encoding vector and AIGC text encoding vector generated by the current auxiliary encoding network layer.
4. The text data processing method according to claim 1, characterized in that, The selection of artificial text encoding vectors and AIGC text encoding vectors of the same level based on a preset level interval includes selecting a portion of the level according to a first interval when the level is within a first threshold range; and selecting a portion of the level according to a second interval when the level is within a second threshold range. When the level is within the third threshold range, select some levels according to the third interval; when the level is within the fourth threshold range, select some levels according to the fourth interval.
5. A text data detection method, characterized in that, The method includes: Obtain the text data to be tested; The text data to be tested is input into the trained encoder model, and the text encoding vector of the text data to be tested is output; the encoder model is the trained encoder model as described in any one of claims 1 to 4; The text encoding vector of the text data to be tested is detected to obtain the text detection result of the text data to be tested. The text detection result is either artificial text or AIGC text, which indicates the source of the text content of the text data to be tested.
6. A text data processing device, characterized in that, The apparatus is applied to a text data processing method according to any one of claims 1 to 4, comprising: The first acquisition unit is used to acquire multi-domain artificial text, which includes text data of multiple different language types; each language type of text data corresponds to multiple different domain types of text data. An enhanced processing unit is used to input multi-domain artificial text into a large model and output multi-domain AIGC text; the multi-domain AIGC text is generated for each artificial text according to heuristic rules at both document granularity and sentence granularity. The first extraction unit is used to extract features from the augmented text data composed of each artificial text and the corresponding AIGC text using the encoder model to be trained, so as to obtain the artificial text encoding vector and the AIGC text encoding vector at each level; the encoder model includes a multi-level encoder network and a corresponding multi-level auxiliary encoding network. The encoder network includes several sub-network layers at multiple levels, and the auxiliary encoding network includes several auxiliary encoding layers at multiple levels. Each sub-network layer is connected to the corresponding level of the auxiliary encoding layer. The first adjustment unit is used to construct a contrastive learning loss based on at least two artificial text encoding vectors and AIGC text encoding vectors of the same level, and to adjust the model parameters of the auxiliary encoding network based on the contrastive learning loss to obtain the trained encoder model.
7. A text data detection device, characterized in that, The device includes: The second acquisition unit is used to acquire the text data to be tested; The second extraction unit is used to extract features from the text data to be tested using an encoder model to obtain the text encoding vector of the text data to be tested. The encoder model is the trained encoder model described in any one of the text data processing methods of claims 1 to 4. The detection unit is used to detect the text encoding vector of each text data to be tested, and obtain the text detection result corresponding to each text data to be tested; the text detection result is artificial text or AIGC text, which is used to indicate the source of the text content of the text data to be tested.
Citation Information
Patent Citations
Text abstract auxiliary generation method based on comparative learning
CN116910233A
Text pre-training model generation method, text prediction method and related equipment
CN117493866A