Content discrimination model training method, content discrimination method and device
By training the content discriminant model, extracting and fusing the semantic and stylistic features of sample text and perturbing text, the problem of indistinguishability between artificial intelligence generated content and artificially generated content is solved, and higher discrimination accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510331782.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to effectively distinguish the content generated by artificial intelligence from the content generated by artificial intelligence, which leads to difficulty in identification and affects the accuracy of information extraction and discrimination.
By training the content discriminant model, the semantic and style features of the sample text and sample perturbation text are extracted, the feature fusion algorithm is used to generate discriminant results, and the model parameters are iteratively optimized to improve the robustness and accuracy of the model.
It improves the accuracy of discrimination of AI-generated content, reduces the probability of misjudgment, optimizes the stickiness of content creators and users, and enhances the stability and defensiveness of content judgment.
Smart Images

Figure CN120372005A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular to artificial intelligence technical fields such as natural language processing and computer vision. Background Art
[0002] With the development of technology, artificial intelligence has become increasingly important in people's work and life. In the daily use of artificial intelligence, people can generate the required content through artificial intelligence. In this scenario, there may be a certain degree of similarity between the content constructed by people themselves and the content generated by artificial intelligence, thus having a certain degree of impact on the recognition of AI-generated content. Summary of the Invention
[0003] The present disclosure provides a training method, a content discrimination method, and a device for a content discrimination model.
[0004] According to a first aspect of the present disclosure, a training method for a content discrimination model is provided, including: obtaining a candidate content discrimination model to be trained, as well as sample texts of the candidate content discrimination model and sample perturbation texts of the sample texts; extracting respective sample style semantic fusion features of the sample texts and the sample perturbation texts to obtain a first AI generation discrimination result of the candidate content discrimination model for the sample texts and a second AI generation discrimination result for the sample perturbation texts; obtaining a target training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result; and iterating the candidate content discrimination model according to the target training loss to obtain a trained target content discrimination model.
[0005] According to a second aspect of the present disclosure, a content discrimination method is provided, including: obtaining a text object to be discriminated and a trained target content discrimination model, where the target content discrimination model is obtained according to the method proposed in the first aspect; obtaining a target detector in the target content discrimination model to obtain a target semantic feature and a target style feature of the text object, where the target style feature is obtained based on a target triple feature of the text object; performing feature fusion on the target semantic feature and the target style feature through the target detector to obtain a target semantic style fusion feature of the text object, and outputting a target AI generation discrimination result of the text object through the target detector based on the target semantic style fusion feature; and determining a target text category to which the text object belongs based on the target AI generation discrimination result.
[0006] According to a third aspect of the present disclosure, a training device for a content discrimination model is provided, including: a first acquisition module configured to acquire a candidate content discrimination model to be trained, as well as sample texts of the candidate content discrimination model and sample perturbation texts of the sample texts; a first discrimination module configured to extract respective sample style semantic fusion features of the sample texts and the sample perturbation texts, so as to obtain a first AI generation discrimination result of the candidate content discrimination model for the sample texts and a second AI generation discrimination result of the candidate content discrimination model for the sample perturbation texts; a second acquisition module configured to obtain an objective training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result; and an iteration module configured to iterate the candidate content discrimination model according to the objective training loss to obtain a trained objective content discrimination model.
[0007] According to a fourth aspect of the present disclosure, a content discrimination device is provided, including: a third discrimination module configured to acquire a text object to be discriminated and a trained objective content discrimination model, where the objective content discrimination model is obtained based on the device proposed in the above third aspect; a feature extraction module configured to acquire an objective detector in the objective content discrimination model to obtain an objective semantic feature and an objective style feature of the text object, where the objective style feature is obtained based on an objective triple feature of the text object; a second discrimination module configured to perform feature fusion on the objective semantic feature and the objective style feature through the objective detector to obtain an objective semantic style fusion feature of the text object, and output an objective AI generation discrimination result of the text object through the objective detector based on the objective semantic style fusion feature; and a determination module configured to determine an objective text category to which the text object belongs based on the objective AI generation discrimination result.
[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the content discrimination model training method proposed in the above first aspect and / or the content discrimination method proposed in the above second aspect.
[0009] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause the computer to execute the content discrimination model training method proposed in the above first aspect and / or the content discrimination method proposed in the above second aspect.
[0010] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the training method of the content discrimination model proposed in the first aspect above and / or the content discrimination method proposed in the second aspect above.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 is a schematic flowchart of a method for training a content discrimination model according to an embodiment of the present disclosure;
[0014] Figure 2 is a schematic flowchart of a method for training a content discrimination model according to another embodiment of the present disclosure;
[0015] Figure 3 is a schematic flowchart of a method for training a content discrimination model according to another embodiment of the present disclosure;
[0016] Figure 4 is a schematic flowchart of a method for training a content discrimination model according to another embodiment of the present disclosure;
[0017] Figure 5 is a schematic flowchart of a method for training a content discrimination model according to another embodiment of the present disclosure;
[0018] Figure 6 is a schematic flowchart of a content discrimination method according to an embodiment of the present disclosure;
[0019] Figure 7 is a schematic flowchart of a content discrimination method according to another embodiment of the present disclosure;
[0020] Figure 8 is a schematic structural diagram of a device for training a content discrimination model according to an embodiment of the present disclosure;
[0021] Figure 9 is a schematic structural diagram of a content discrimination device according to an embodiment of the present disclosure;
[0022] Figure 10 is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0024] Data processing is a basic link in systems engineering and automatic control. Data is a form of expression of facts, concepts, or instructions and can be processed by manual or automated devices. After data is interpreted and given a certain meaning, it becomes information. Data processing is the collection, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be chaotic and difficult to understand.
[0025] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Natural language processing is mainly applied to machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR, and other aspects.
[0026] Artificial Intelligence, abbreviated as AI in English, is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new technical science that studies, develops, and applies theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is an important part of the intelligent discipline. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is a very broad science, including robotics, speech recognition, image recognition, natural language processing, expert systems, machine learning, computer vision, etc.
[0027] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant regulations of national laws and regulations and do not violate public order and good customs.
[0028] Figure 1 It is a schematic flowchart of the training method of the content discrimination model according to an embodiment of the present disclosure. As Figure 1 shown, the method includes:
[0029] S101. Obtain a candidate content discrimination model to be trained, as well as sample texts of the candidate content discrimination model and sample perturbation texts of the sample texts.
[0030] In daily work and life, people need to distinguish between content generated by artificial intelligence (AI) and content generated manually for purposes such as extracting and identifying real information.
[0031] In the embodiments of the present disclosure, a trained model can be used to discriminate whether the text content is AI-generated content. Among them, the model to be trained can be determined as a candidate content discrimination model.
[0032] In this scenario, based on the sample construction method in the related art, sample texts for use in training the candidate content discrimination model can be constructed. Among them, the sample texts can be constructed based on text content generated by AI or based on text content generated manually, and no specific limitation is made here.
[0033] Optionally, when training the candidate content discrimination model, the model can also be trained using texts obtained by perturbing the sample texts. Among them, the sample texts can be processed by an algorithm based on the perturbation algorithm in the related art, and then the perturbed texts of the sample texts can be obtained according to the result of the algorithm processing as sample perturbation texts.
[0034] S102. Extract the sample style semantic fusion features of the sample texts and the sample perturbation texts respectively to obtain a first AI generation discrimination result of the candidate content discrimination model for the sample texts and a second AI generation discrimination result for the sample perturbation texts.
[0035] In the embodiments of the present disclosure, the sample texts can be input into the candidate content discrimination model, and the sample texts can be feature-extracted through the feature extraction layer in the candidate content discrimination model. Among them, the sample texts can be feature-extracted in the semantic dimension through a preset semantic feature extraction algorithm in the model, and the sample texts can be feature-extracted in the style dimension through a preset style feature extraction algorithm in the model.
[0036] Among them, the feature extraction of the sample texts in the style dimension can be understood as extracting relevant features such as the preference and habit of the diction and sentence structure of the text content included in the sample texts, so as to obtain the features of the sample texts in the style dimension.
[0037] In this scenario, the feature fusion layer in the candidate content discrimination model can be obtained, and the features in the semantic dimension and the features in the style dimension extracted by the feature extraction layer are input into the feature fusion layer. The two are processed by the feature fusion algorithm deployed in the feature fusion layer, and then the fused features after the fusion of the two are obtained. This fused feature is the sample style semantic fusion feature of the sample text.
[0038] It should be noted that for the process of obtaining the sample style semantic fusion feature of the sample perturbed text, reference can be made to the process of obtaining the sample style semantic fusion feature of the above sample text for understanding, and no specific elaboration will be made here.
[0039] When AI generates text content, there is a set text construction style. It can be understood that when AI generates text content, its word choice and sentence structure have set writing habits, which can include the complexity of words used, the distribution of sentence lengths, the word frequency distribution, and the syntactic structure and other related writing habits. Based on this part of the habits, the text construction style when AI generates text can be characterized. Correspondingly, when people write text independently, their word choice and sentence structure have their own writing habits, which can be characterized based on relevant information such as the complexity of the vocabulary used, the distribution of sentence lengths, the word frequency distribution, and the syntactic structure.
[0040] Optionally, the text semantics of the sample text and the sample perturbed text can be obtained through the sample style semantic fusion features of the sample text and the sample perturbed text respectively, and the text construction styles of the sample text and the sample perturbed text can be obtained. Optionally, during the construction process of the candidate content discrimination model, relevant information such as the text construction style when AI generates text can be assigned to the model, so that the attention of the assigned candidate content discrimination model to the text construction style can be improved.
[0041] In this scenario, the candidate content discrimination model can perform a comparative analysis based on the feature representation information of the sample style semantic fusion feature and the relevant information assigned when the model is constructed for AI-generated text, and obtain the probability that the text content corresponding to this sample style semantic fusion feature is AI-generated text content based on the result of the comparative analysis, and then obtain the discrimination result of this text content. Among them, when this text content is the sample text, this discrimination result is the first AI generation discrimination result of the sample text, and when this text content is the sample perturbed text, this discrimination result is the second AI generation discrimination result of the sample perturbed text.
[0042] S103. Obtain the target training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result.
[0043] Optionally, according to the loss value acquisition algorithm in the related technology, algorithm processing can be performed on the discriminant result generated by the first AI and the label information of the sample text, and then based on the result of the algorithm processing, the loss value of the discriminant result generated by the first AI based on the label information of the sample text can be obtained.
[0044] And, based on the same processing process, the loss value of the discriminant result generated by the second AI based on the label information of the sample perturbation text can be obtained.
[0045] In this scenario, the preset training strategy of the candidate content discriminant model can be obtained, and based on the setting algorithm for the model training loss in the training strategy, algorithm processing is performed on the above two loss values. Then, based on the result of the algorithm processing, the loss value of the candidate content discriminant model in the current training round is obtained, and this loss value is determined as the target training loss of the candidate content discriminant model in the current training round.
[0046] S104, Iterate the candidate content discriminant model according to the target training loss to obtain the trained target content discriminant model.
[0047] Optionally, based on the iterative optimization method of model parameters in the related technology, the model parameters of the candidate content discriminant model in the current round can be adjusted based on the target training loss, so as to achieve the parameter iteration of the candidate content discriminant model in the current training round.
[0048] Furthermore, return to obtain the next sample text and the corresponding next sample perturbation text of the next sample text, and continue to train the candidate content discriminant model with adjusted parameters until the preset model training end condition is met, end the training of the candidate content discriminant model, and determine the model obtained at the end of the training as the trained target content discriminant model.
[0049] Optionally, the training end condition of the candidate content discriminant model can be set based on the training round or based on the output result of the model, and no specific limitation is made here.
[0050] The training method of the content discrimination model proposed by the present disclosure obtains a candidate content discrimination model, corresponding sample texts, and sample perturbation texts, and extracts the sample style semantic fusion features of the sample texts and the sample perturbation texts respectively, so as to obtain the first AI-generated discrimination result output by the candidate content discrimination model based on the sample text and the second AI-generated discrimination result output based on the sample perturbation text. Optionally, based on the first AI-generated discrimination result and the second AI-generated discrimination result, the target training loss of the candidate content discrimination model is obtained, and the candidate content discrimination model is iteratively modeled based on the target training loss to obtain the trained target content discrimination model. In the present disclosure, the candidate content discrimination model is trained based on the sample text and the sample perturbation text, which improves the robustness of the trained target content discrimination model, extracts the features of the text content input into the model in the semantic dimension and the style dimension, enables the candidate content discrimination model to learn the differences between AI-generated content and human-generated content in the semantic dimension and the style dimension respectively, improves the accuracy of the trained target content discrimination model for discriminating AI-generated content, reduces the possibility of misjudging AI-generated content in the scenario of discriminating AI-generated content based on the target content discrimination model, thereby reducing the occurrence probability of abnormal events caused by content misjudgment, improving the stickiness of content creators and users, and optimizing the discrimination method and discrimination effect of AI-generated content.
[0051] In the above embodiment, regarding the acquisition of the first AI-generated discrimination result and the second AI-generated discrimination result, it can also be combined with Figure 2 For further understanding, Figure 2 is a schematic flowchart of the training method of the content discrimination model according to another embodiment of the present disclosure. As Figure 2 shown, the method includes:
[0052] S201, obtain the candidate perturbator in the candidate content discrimination model.
[0053] In the embodiment of the present disclosure, a perturbator is set in the candidate content discrimination model, which can be determined as the candidate perturbator. The sample text input into the model can be perturbed by the candidate perturbator to obtain the sample perturbation text corresponding to the sample text.
[0054] As an example, as Figure 3 shown, the candidate content discrimination model includes Figure 3 the candidate perturbator shown. In the Figure 3 shown scenario, the sample text can be input into Figure 3 the candidate perturbator shown. Through the preset perturbation strategy in the candidate perturbator, the sample text input into it is perturbed to obtain Figure 3 the sample perturbation text shown.
[0055] S202. Input the sample text into the candidate perturbator to obtain the candidate perturbed text of the sample text.
[0056] In the embodiments of the present disclosure, the candidate perturbator has a preset text perturbation strategy. Among them, the text perturbation strategy may include operations such as synonym replacement, sentence pattern conversion, active-passive rewriting, and noise character addition.
[0057] In this scenario, the sample text can be input into the candidate perturbator, and the sample text is rewritten through the preset perturbation strategy in the candidate perturbator to realize the perturbation of the sample text, and the text obtained after perturbation is determined as the candidate perturbed text corresponding to the sample text.
[0058] S203. Perform semantic similarity evaluation and classification consistency evaluation on the candidate perturbed text and the sample text to obtain the semantic similarity evaluation parameter and the classification consistency evaluation parameter.
[0059] In the embodiments of the present disclosure, the semantic similarity between the sample text and the sample perturbed text needs to meet the preset conditions. In this scenario, it is necessary to perform similarity analysis in the semantic dimension on the candidate perturbed text obtained by the candidate perturbator and the sample text.
[0060] Optionally, the evaluation items used for the preset semantic similarity analysis can be obtained, and the sample text and the candidate perturbed text are processed by an algorithm based on the analysis algorithm corresponding to the evaluation item. Then, based on the result of the algorithm processing, the evaluation parameter under the semantic similarity evaluation item between the sample text and the candidate perturbed text is obtained and marked as the semantic similarity evaluation parameter between the two.
[0061] As an example, the evaluation items used for semantic similarity analysis can be obtained based on the similarity score (Bidirectional Encoder Representations from Transformers Score, BERTScore) of the pre-trained language model (BERT) in related technologies to obtain the semantic similarity evaluation parameter between the sample text and the candidate perturbed text. It is also possible to obtain the evaluation items used for semantic similarity analysis based on the text similarity evaluation method (Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence, ROUGE-L) based on the Longest Common Subsequence (LCS) in related technologies to obtain the semantic similarity evaluation parameter between the sample text and the candidate perturbed text. No specific limitation is made here.
[0062] In the embodiments of the present disclosure, when a candidate disturber perturbs a sample text, the category to which the text content obtained by the perturbation belongs needs to be the same as the category to which the sample text belongs. For example, if the category to which the sample text belongs is set to AI-generated, then for the candidate perturbed text obtained by the candidate disturber after perturbing it, the category to which it belongs is AI-generated. That is to say, when the candidate disturber perturbs the sample text, it will not change the original text writing style of the sample text.
[0063] In this scenario, for the candidate perturbed text obtained by the candidate disturber, it is necessary to evaluate whether the category to which the candidate perturbed text belongs is the same as the category to which the sample text belongs. Among them, this evaluation can be determined as a classification consistency evaluation for the two.
[0064] Optionally, the sample text and the candidate perturbed text can be evaluated and analyzed based on the classification consistency evaluation method in related technologies, and then according to the results of the evaluation and analysis, the evaluation parameters obtained from the classification consistency evaluation of the sample text and the candidate perturbed text are used as the classification consistency evaluation parameters for the two.
[0065] S204, in response to the semantic similarity evaluation parameter being greater than or equal to a preset semantic similarity threshold, and the classification consistency evaluation parameter indicating that the candidate perturbed text and the sample text are classified consistently, determine that the candidate perturbed text is the sample perturbed text of the sample text.
[0066] In the embodiments of the present disclosure, when the semantic similarity evaluation parameter is greater than or equal to a preset semantic similarity threshold, it can be determined that the candidate perturbed text and the sample text are semantically similar to each other, and further, it can be determined that the perturbation of the candidate disturber to the sample text does not modify the original semantics of the sample text.
[0067] And when the classification consistency evaluation parameter indicates that the category to which the candidate perturbed text belongs is the same as the category to which the sample text belongs, it can be determined that the candidate perturbed text obtained by the candidate disturber perturbing the sample text has not changed its belonging category.
[0068] In this scenario, for any candidate perturbed text, when the semantic similarity evaluation parameter and the classification consistency evaluation parameter corresponding to the candidate perturbed text meet the above conditions, it can be determined that the similarity between the candidate perturbed text and the sample perturbed text meets the training requirements of the downstream candidate classifier. Further, the candidate perturbed text can be determined as the sample perturbed text of the sample text.
[0069] S205, obtain the candidate semantic encoder and the candidate style encoder in the candidate content discrimination model.
[0070] In the embodiments of the present disclosure, an encoder for the semantic dimension and an encoder for the style dimension are provided in the candidate content discrimination model, which are used to extract the features of the semantic dimension and the features of the style dimension of the text content input into the model.
[0071] Among them, the encoder for extracting the features of the semantic dimension can be determined as the candidate semantic encoder in the candidate content discrimination model, and the encoder for extracting the features of the style dimension can be determined as the candidate style encoder in the candidate content discrimination model.
[0072] As an example, as Figure 3 shown, the candidate semantic encoder and the candidate style encoder in the candidate content discrimination model can be set in Figure 3 the candidate detector shown.
[0073] S206, extract the first semantic feature of the sample text through the candidate semantic encoder, and extract the second semantic feature of the sample perturbed text.
[0074] As an example, as Figure 4 shown, assuming Figure 4 the text shown is the sample text, then the sample text can be input into Figure 4 the text preprocessing layer shown, preprocess the sample text through the preset processing strategy in the text preprocessing layer, and input the preprocessed text content into Figure 4 the candidate semantic encoder shown, extract the features in the semantic dimension through the candidate semantic encoder, and determine the extracted features as the first semantic feature of the sample text.
[0075] And, when assuming Figure 4 the text shown is the sample perturbed text, the sample perturbed text can be processed accordingly through the above process of obtaining the first semantic feature, and then the semantic feature of the sample perturbed text extracted by the candidate semantic encoder can be obtained, which is marked as the second semantic feature.
[0076] S207, extract the first style feature of the sample text through the candidate style encoder, and extract the second style feature of the sample perturbed text.
[0077] In the embodiments of the present disclosure, a candidate style encoder is provided in the candidate content discrimination model. Through the candidate style encoder, the writing style features of the sample text or the sample perturbed text can be extracted respectively. Among them, the features corresponding to the writing style of the sample text extracted can be determined as the first style feature of the sample text, and the features corresponding to the writing style of the sample perturbed text extracted can be determined as the second style feature of the sample perturbed text.
[0078] It should be noted that in the process of constructing the candidate content discrimination model, the corresponding model parameters can be generated based on the method of Style Preference Optimization (SPO) in related technologies, and the model can be assigned based on these model parameters, so that the assigned candidate content discrimination model can pay more attention to the unique words and sentence patterns used in AI-generated content during the training process. Among them, the above effect can be achieved by increasing the penalty or reward for the candidate content discrimination model to recognize features such as AI-unique words and sentence patterns during the training process.
[0079] Optionally, when constructing the text content, the sentences included in the text content can be composed of a subject, a predicate, and an object. In this scenario, the grammatical structure features of the text sentences can be extracted based on the composition of the subject, predicate, and object when writing the sentences, and the features in the dimension of the grammatical structure extracted are determined as the triple (Subject-Predicate-Object, SPO) features of the text content.
[0080] Optionally, the first triple features of the sample text and the second triple features of the sample perturbed text are extracted.
[0081] As an example, as Figure 4 shown, assuming Figure 4 the text shown is the sample text, then in Figure 4 the scenario shown, the sample text can be input into Figure 4 the shown preprocessing layer. Through the preprocessing layer, based on the grammar of the subject, predicate, and object, the grammatical structure features of each sentence in the sample text are extracted, so as to obtain the triple features of the sample text in the dimension of the grammatical structure, which are determined as the first triple features of the sample text.
[0082] Correspondingly, when Figure 4 the text shown is the sample perturbed text, the triple features of the sample perturbed text can be extracted based on the same processing process and used as the second triple features.
[0083] Optionally, based on the sample text and the first triple features, the input of the candidate style encoder is obtained to extract the first style features of the sample text.
[0084] In the embodiments of the present disclosure, the first triple features and the sample text can be jointly used as the input of the candidate style encoder. Among them, the input requirements of the candidate style encoder can be obtained, and the first triple features and the sample text can be integrated based on this requirement, so as to obtain the input data of the candidate style encoder.
[0085] Further, input the input data into the candidate style encoder, extract the features of the writing style dimension from the text content input therein through the candidate style encoder, and determine the extracted style features as the first style features of the sample text in the input data.
[0086] Optionally, based on the sample perturbed text and the second triple feature, obtain the input of the candidate style encoder to extract the second style feature of the sample perturbed text.
[0087] As an example, as Figure 4 shown, as Figure 4 shown, the sample text of the text shown extracts the first triple feature of the sample text through Figure 4 the shown preprocessing layer, and inputs the first triple feature and the sample text into Figure 4 the shown candidate style encoder to extract the first style feature of the sample text.
[0088] In the embodiments of the present disclosure, the process of obtaining the second style feature can be understood in combination with the process of obtaining the first style feature described above, and will not be elaborated here.
[0089] Optionally, the preprocessing layer can be set in Figure 3 the shown candidate detector.
[0090] S208, perform feature fusion on the first semantic feature and the first style feature to obtain the first fusion feature of the sample text, and perform feature fusion on the second semantic feature and the second style feature to obtain the second fusion feature of the sample perturbed text, so as to obtain the sample style semantic fusion features of the sample text and the sample perturbed text respectively.
[0091] As an example, as Figure 4 shown, for the sample text, the first semantic feature of the sample text extracted by the candidate semantic encoder and the first style feature of the sample text extracted by the candidate style encoder can be input into Figure 4 the shown feature fusion layer.
[0092] In this example, there is a preset feature fusion algorithm in the feature fusion layer, and the first semantic feature and the first style feature input into the feature fusion layer can be processed based on this algorithm, and then the fusion feature of the two can be obtained based on the result of the algorithm processing. This fusion feature is the first fusion feature of the sample text, that is, the sample style semantic fusion feature of the sample text proposed in the above embodiment.
[0093] It should be noted that the acquisition of the second fusion feature of the sample perturbation text can be understood in combination with the acquisition process of the first fusion feature of the above-mentioned sample text, which will not be elaborated here. Among them, the second fusion feature of the second semantic feature and the second style feature is the sample style semantic fusion feature of the sample perturbation text proposed in the above embodiment.
[0094] S209, obtain the candidate classifier in the candidate content discrimination model.
[0095] In the embodiment of the present disclosure, a classifier is set in the candidate content discrimination model, which can be determined as the candidate classifier in the candidate content discrimination model. Through the candidate classifier, it can be discriminated whether the text content is AI-generated text content.
[0096] Among them, the candidate classifier can be set in Figure 3 the candidate detector shown.
[0097] S210, based on the first fusion feature through the candidate classifier, obtain the first AI generation discrimination result of the sample text, and based on the second fusion feature through the candidate classifier, obtain the second AI generation discrimination result of the sample perturbation text.
[0098] In the embodiment of the present disclosure, for the sample text, the first fusion feature of the sample text can be input into the candidate classifier. Through the candidate classifier, it is classified whether the category of the sample text is AI-generated text, and the classification result and the confidence corresponding to the classification result are output. Among them, based on the classification result and the confidence of the classification result, the first AI generation discrimination result of the sample text can be obtained.
[0099] Correspondingly, for the sample perturbation text, its corresponding second fusion feature can be input into the candidate classifier. Through the candidate classifier, it is classified whether the category of the sample perturbation text is AI-generated text, and the classification result and the confidence corresponding to the classification result are output. Among them, based on the classification result and the confidence of the classification result, the second AI generation discrimination result of the sample perturbation text can be obtained.
[0100] As an example, as Figure 4 shown, for the sample text, the first fusion feature, which is the sample style semantic fusion feature of the sample text output by Figure 4 the shown feature fusion layer, can be input into Figure 4 the shown candidate classifier. Through the candidate classifier, it is classified whether the category of the sample text is AI-generated text based on the first fusion feature, so as to obtain Figure 4 the discrimination result shown as the first AI generation discrimination result of the sample text.
[0101] In this example, for the sample perturbed text, the second AI-generated discrimination result of the sample perturbed text can be obtained based on the method for obtaining the first AI-generated discrimination result of the above sample text, which will not be elaborated here.
[0102] As another example, as Figure 3 shown, the sample text and the sample perturbed text can be respectively input into the Figure 3 candidate detectors shown. Through the text preprocessing layer, candidate semantic encoder, candidate style encoder, and candidate classifier included in the candidate detector, the determination of the categories to which the sample text and the sample perturbed text belong is respectively performed, so as to obtain the Figure 3 first AI-generated discrimination result of the sample text shown, and the second AI-generated discrimination result of the sample perturbed text.
[0103] The training method of the content discrimination model proposed by the present disclosure perturbs the sample text through a candidate perturbator, so that the sample perturbed text and the sample text belong to the same category and are semantically similar, improving the similarity between the sample perturbed text and the sample text in the semantic dimension and the style dimension. And based on the sample text and the sample perturbed text, the candidate content discrimination model is trained, improving the defensive ability of the candidate content discrimination model against content rewriting attacks, improving the robustness and stability of the trained target content discrimination model. Based on the sample style semantic features fused by semantic features and style features, the corresponding AI-generated discrimination result is obtained, enabling the candidate content discrimination model to learn the differences between AI-generated content and human-generated content in the semantic dimension and the style dimension respectively, and improving the accuracy of the trained target content discrimination model in discriminating AI-generated content.
[0104] In the above embodiment, regarding the training of the candidate content discrimination model, it can also be combined with Figure 5 to further understand that Figure 5 is a schematic flowchart of the training method of the content discrimination model according to another embodiment of the present disclosure. As Figure 5 shown, the method includes:
[0105] S501, according to the first AI-generated discrimination result and the sample text, obtain the first classification loss of the candidate content discrimination model, and according to the second AI-generated discrimination result and the sample perturbed text, obtain the second classification loss of the candidate content discrimination model.
[0106] In the embodiment of the present disclosure, there are multiple loss values during the training process of the candidate content discrimination model, which may include the loss value between the classification discrimination result output by the candidate classifier and the category to which the sample belongs.
[0107] Among them, for the sample text and the first AI-generated discrimination result corresponding to the sample text, an algorithm processing can be performed on the first AI-generated discrimination result and the sample label carried by the sample text based on the loss value acquisition algorithm in the related art. Then, based on the result of the algorithm processing, the loss value of the first AI-generated discrimination result based on the sample label of the sample text can be obtained, and this loss value can be determined as the first classification loss of the candidate content discrimination model.
[0108] Correspondingly, for the sample perturbation text and the second AI-generated discrimination result of the sample perturbation text, the loss value of the second AI-generated discrimination result based on the sample label carried by the sample perturbation text can be obtained based on the same loss value acquisition algorithm, and used as the second classification loss of the candidate content discrimination model.
[0109] As an example, an algorithm processing can be performed on the first AI-generated discrimination result and the sample label of the sample text based on the cross-entropy loss value algorithm in the related art, and the cross-entropy loss value obtained from the algorithm processing can be determined as the first classification loss of the first AI-generated discrimination result based on the sample label of the sample text. It is also possible to calculate the first classification loss based on other loss value algorithms, which will not be specifically limited here.
[0110] S502. Obtain the output consistency loss of the candidate content discrimination model according to the first AI-generated discrimination result and the second AI-generated discrimination result.
[0111] In the embodiments of the present disclosure, the first AI-generated discrimination result is the discrimination result corresponding to the sample text obtained by the candidate classifier, and the second AI-generated discrimination result is the discrimination result corresponding to the sample perturbation text obtained by the candidate classifier.
[0112] In the scenario where the sample text and the sample perturbation text are semantically similar and belong to the same category, the classification discrimination result of the candidate content discrimination model for the sample text and the classification discrimination result for the sample perturbation text need to be consistent. That is to say, the error value between the first AI-generated discrimination result and the second AI-generated discrimination result needs to be within a preset error range.
[0113] It can be understood that when the class determination results of the first AI-generated discrimination result and the second AI-generated discrimination result are different, the error value between the first AI-generated discrimination result and the second AI-generated discrimination result is outside the preset error range.
[0114] And when the class determination results of the first AI-generated discrimination result and the second AI-generated discrimination result are the same and the confidence levels of the class determination results are different, the error value between the first AI-generated discrimination result and the second AI-generated discrimination result is outside the preset error range.
[0115] In this scenario, the discrimination results generated by the first AI and the second AI can be processed algorithmically based on the loss value algorithm in the related technology. Then, based on the algorithm results, the loss value of the discrimination results generated by the first AI and the second AI in the consistency dimension can be obtained, and this loss value is the output consistency loss of the candidate content discrimination model.
[0116] As an example, the discrimination results generated by the first AI and the second AI can be processed algorithmically based on the algorithm of the constraint loss based on the Kullback-Leibler Divergence in the related technology, and the loss value obtained by the algorithm can be determined as the output consistency loss of the candidate content discrimination model.
[0117] Among them, in the algorithm of the constraint loss based on the KL divergence, the acquisition of the KL divergence can be as follows:
[0118]
[0119] In the above formula, P represents the confidence distribution in the discrimination result generated by the first AI of the sample text, Q represents the confidence distribution in the discrimination result generated by the second AI of the sample perturbed text, and D KL represents the KL divergence.
[0120] Optionally, the constraint loss based on the KL divergence can be as follows:
[0121]
[0122] In the above formula, L KL represents the constraint loss based on the KL divergence.
[0123] It should be noted that based on the loss value shown in the above formula, the candidate content discrimination model can increase the degree of attention to the features in related dimensions such as the writing style and writing tone in the text content.
[0124] S503. Based on the first classification loss, the second classification loss, and the output consistency loss, obtain the target training loss of the candidate content discrimination model.
[0125] In the embodiments of the present disclosure, the weights of the first classification loss, the second classification loss, and the output consistency loss can be obtained, and the first classification loss, the second classification loss, and the output consistency loss are weighted based on this weight, so as to obtain the training loss of the candidate content discrimination model in the current round, which is determined as the target training loss of the candidate content discrimination model in the current round.
[0126] As an example, the acquisition algorithm of the target training loss can be understood in combination with the following formula:
[0127] Ltotal = L cls L(x) + L cls L(x') + λL KL
[0128] In the above expressions, x represents the sample text, x' represents the sample perturbation text, and L total represents the target training loss, and L cls L(x) represents the first classification loss corresponding to the sample text, and L cls L(x') represents the second classification loss corresponding to the sample perturbation text, and L KL represents the output consistency loss corresponding to the first AI-generated discrimination result and the second AI-generated discrimination result, and λ represents the weighting weight.
[0129] As an example, as Figure 3 shown, the first classification loss and the second classification loss included in the classification loss shown by Figure 3 , and Figure 3 the output consistency loss included in the consistency shown by Figure 3 are weighted and calculated to obtain the target training loss shown by
[0130] It should be noted that when adjusting the parameters of the candidate perturbator and the candidate detector, the model parameters of the candidate perturbator can be adjusted and iterated by methods such as gradient ascent and reinforcement learning in related technologies, or other methods can also be used, which are not specifically limited here.
[0131] S504. Iterate the candidate content discrimination model according to the target training loss to obtain a trained target content discrimination model.
[0132] Optionally, obtain the candidate perturbator and the candidate detector in the candidate content discrimination model, where the candidate detector includes a candidate semantic encoder, a candidate style encoder, and a candidate classifier.
[0133] In the embodiments of the present disclosure, the candidate content discrimination model includes a candidate perturbator for sample perturbation and a candidate detector for classifying AI-generated content. In this scenario, the candidate content discrimination model can be iteratively optimized by iterating the model parameters of the candidate perturbator and the candidate detector.
[0134] It should be noted that the candidate detector includes a candidate semantic encoder for semantic feature extraction, a candidate style encoder for style feature extraction, and a candidate classifier for classifying AI-generated content based on the style-semantic fusion feature.
[0135] Optionally, according to the target training loss, the parameters of the candidate disturber and the candidate detector are adjusted to adjust the parameters of the candidate content discrimination model, and the next sample text and the next sample perturbation text of the next sample text are obtained and returned. The candidate content discrimination model after parameter adjustment is continuously trained to obtain a trained target content discrimination model.
[0136] In the embodiments of the present disclosure, according to the method of iterative adjustment of model parameters in related technologies, the model parameters of the candidate content discrimination model in the current round can be adjusted based on the target training loss. For example, Figure 3 As shown, iterative adjustment of the parameters of the candidate disturber and the candidate detector included in the candidate content discrimination model is performed according to the target training loss, so as to realize the parameter iteration of the candidate content discrimination model.
[0137] Among them, for the candidate semantic encoder, candidate style encoder, and candidate classifier included in the candidate detector, the Figure 4 shown backpropagation method can be used to perform iterative adjustment of the parameters of the Figure 4 shown candidate semantic encoder, candidate style encoder, and candidate classifier respectively based on the target training loss obtained by the loss value calculation module, thereby realizing the parameter adjustment of the candidate detector.
[0138] Furthermore, the next sample text can be obtained and returned, and the next sample perturbation text corresponding to the next sample text can be obtained. Then, based on the next sample text and the next sample perturbation text, the candidate content discrimination model after parameter adjustment is continuously trained until the end condition of the model training is met, and thus a trained target content discrimination model can be obtained.
[0139] Among them, the end condition of the model training can be set based on the number of rounds of model training or the output result of the model, which is not specifically limited here.
[0140] It should be noted that training the candidate content discrimination model based on the sample text and the sample perturbation text proposed in the above embodiments can improve the accuracy of the trained target content discrimination model in determining the category of text content with similar semantic modifications, thereby improving the adversarial ability of the target content discrimination model against similar semantic modification attacks.
[0141] The training method of the content discrimination model proposed in the present disclosure obtains the target training loss of the candidate content discrimination model based on the first classification loss, the second classification loss, and the output consistency loss, improves the generalization ability of the candidate content discrimination model, reduces the probability of the candidate content discrimination model overfitting, and optimizes the training effect of the candidate content discrimination model.
[0142] The present disclosure also proposes a content discrimination method, which can be combined withFigure 6 Understand Figure 6 The following is a schematic flowchart of a content discrimination method according to an embodiment of the present disclosure. As Figure 6 shown, the method includes:
[0143] S601, Obtain the text object to be discriminated and the trained target content discrimination model.
[0144] In the embodiment of the present disclosure, the text content that needs to be determined whether it is AI-generated content can be determined as the text object to be discriminated.
[0145] Optionally, it is possible to discriminate whether the text object belongs to AI-generated content based on the trained target content discrimination model, where the target content discrimination model is obtained based on the training method of the content discrimination model proposed in the above Figures 1 to 5 embodiment.
[0146] S602, Obtain the target detector in the target content discrimination model to obtain the target semantic feature and target style feature of the text object, where the target style feature is obtained based on the target triple feature of the text object.
[0147] In the embodiment of the present disclosure, the target content discrimination model includes a trained target detector. In this scenario, it is possible to discriminate whether the text object belongs to AI-generated content through the target detector.
[0148] As an example, as Figure 7 shown, in the Figure 7 shown target detector, the text object can be preprocessed through the target preprocessing layer. Among them, the target preprocessing layer can extract features in the dimension of the grammatical structure of the text object based on relevant grammatical structures such as the subject, predicate, and object, so as to obtain the target triple feature of the text object.
[0149] Furthermore, the target triple feature and the text object are input into the Figure 7 shown target style encoder. The target style encoder extracts features of the text object in the dimension of the writing style, and determines the extracted style feature as the target style feature of the text object.
[0150] And, the text object is input into the Figure 7 shown target semantic encoder. The target semantic encoder extracts features of the text object in the semantic dimension, and determines the extracted semantic feature as the target semantic feature of the text object.
[0151] S603. Feature fusion of the target semantic features and the target style features is performed by a target detector to obtain the target semantic style fusion features of the text object, and the target detector outputs the target AI-generated discrimination result of the text object based on the target semantic style fusion features.
[0152] In the embodiments of the present disclosure, a target classifier included in the target detector can be used to discriminate and classify whether the text object belongs to AI-generated content. Among them, the target detector can perform feature fusion on the target semantic features and the target style features, and determine the fused features as the target semantic style fusion features of the text object.
[0153] As an example, as Figure 7 shown, the target semantic features output by the target semantic encoder and the target style features output by the target style encoder can be feature-fused through the target feature fusion layer shown in Figure 7 to obtain the fused target semantic style fusion features.
[0154] As Figure 7 shown, the target semantic style fusion features can be input into the target classifier shown in Figure 7 . The target classifier determines the category to which the text object belongs based on the target semantic style fusion features, and determines the output category result as the target AI-generated discrimination result of the text object obtained by the target detector in the target content discrimination model.
[0155] S604. Based on the target AI-generated discrimination result, determine the target text category to which the text object belongs.
[0156] In the embodiments of the present disclosure, the target AI-generated discrimination result includes the category to which the text object belongs and the confidence level of the category. In this scenario, the category to which the text object belongs can be determined according to the confidence level included in the target AI-generated discrimination result, and this category is the target text category to which the text object belongs.
[0157] As an example, as Figure 7 shown, for the confidence level of any category included in the target AI-generated discrimination result, the confidence level can be compared with the confidence level threshold shown in Figure 7 . When the confidence level is greater than or equal to the preset confidence level threshold, it can be determined that this category is the target text category to which the text object belongs.
[0158] Correspondingly, when the confidence level is less than the preset confidence level threshold, it can be determined that this category is not the target text category to which the text object belongs.
[0159] Among them, the target text category can include the AI-generated category and the manually written category.
[0160] It should be noted that among the various categories included in the target AI-generated discrimination result and the confidence levels of each category, there may be a situation where all the confidence levels are less than the confidence level threshold. In this scenario, it is impossible to accurately determine the target text category to which the text object belongs based on the target AI-generated discrimination result.
[0161] Furthermore, when the above situation occurs, a secondary category determination can be performed on the text object based on a preset category determination review strategy, so as to obtain the target text category of the text object. Among them, the category determination review strategy can be a manual review strategy or a secondary discrimination strategy implemented based on the target content discrimination model, and no specific limitation is made here.
[0162] The content discrimination method proposed in this disclosure obtains the target semantic feature and target style feature of the text object through the target detector in the target content discrimination model, so as to obtain the target semantic style fusion feature of the two. Furthermore, based on the target semantic style fusion feature, the target AI-generated discrimination result of the text object output by the target detector is obtained, and the target text category to which the text object belongs is obtained based on the target AI-generated discrimination result. In this disclosure, the target detector in the trained target content discrimination model is used to perform category discrimination on whether the text object belongs to AI-generated content, which improves the accuracy and efficiency of the category determination of the text object, reduces the possibility of misjudging the category to which the text object belongs, reduces the occurrence probability of abnormal events caused by the confusion between manually written and AI-generated content, and optimizes the discrimination method and discrimination effect of AI-generated content.
[0163] An embodiment of this disclosure also proposes a training device for a content discrimination model. Since the training device for the content discrimination model proposed in the embodiment of this disclosure corresponds to the training method for the content discrimination model proposed in the above several embodiments, the implementation manners of the above training method for the content discrimination model are also applicable to the training device for the content discrimination model proposed in the embodiment of this disclosure, and will not be described in detail in the following embodiments.
[0164] Figure 8 is a schematic structural diagram of a training device for a content discrimination model according to an embodiment of this disclosure. As Figure 8 shown, the training device 800 for the content discrimination model includes a first acquisition module 81, a first discrimination module 82, a second acquisition module 83, and an iteration module 84, where:
[0165] The first acquisition module 81 is configured to acquire a candidate content discrimination model to be trained, as well as a sample text of the candidate content discrimination model and a sample perturbation text of the sample text;
[0166] The first discrimination module 82 is configured to extract the sample style semantic fusion features of the sample text and the sample perturbation text respectively, so as to obtain the first AI generation discrimination result of the candidate content discrimination model for the sample text and the second AI generation discrimination result for the sample perturbation text;
[0167] The second acquisition module 83 is configured to obtain the target training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result;
[0168] The iteration module 84 is configured to iterate the candidate content discrimination model according to the target training loss to obtain the trained target content discrimination model.
[0169] In the embodiment of the present disclosure, the first discrimination module 82 is further configured to: obtain the candidate semantic encoder and the candidate style encoder in the candidate content discrimination model; extract the first semantic feature of the sample text and the second semantic feature of the sample perturbation text through the candidate semantic encoder; extract the first style feature of the sample text and the second style feature of the sample perturbation text through the candidate style encoder; perform feature fusion on the first semantic feature and the first style feature to obtain the first fusion feature of the sample text, and perform feature fusion on the second semantic feature and the second style feature to obtain the second fusion feature of the sample perturbation text, so as to obtain the sample style semantic fusion features of the sample text and the sample perturbation text respectively.
[0170] In the embodiment of the present disclosure, the first discrimination module 82 is further configured to: obtain the candidate classifier in the candidate content discrimination model; obtain the first AI generation discrimination result of the sample text through the candidate classifier based on the first fusion feature, and obtain the second AI generation discrimination result of the sample perturbation text through the candidate classifier based on the second fusion feature.
[0171] In the embodiment of the present disclosure, the first discrimination module 82 is further configured to: extract the first triple feature of the sample text and the second triple feature of the sample perturbation text; based on the sample text and the first triple feature, obtain the input of the candidate style encoder to extract the first style feature of the sample text; based on the sample perturbation text and the second triple feature, obtain the input of the candidate style encoder to extract the second style feature of the sample perturbation text.
[0172] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: acquire a candidate disturber in the candidate content discrimination model; input the sample text into the candidate disturber to obtain a candidate perturbed text of the sample text; perform semantic similarity evaluation and classification consistency evaluation on the candidate perturbed text and the sample text to obtain a semantic similarity evaluation parameter and a classification consistency evaluation parameter; and determine the candidate perturbed text as the sample perturbed text of the sample text in response to the semantic similarity evaluation parameter being greater than or equal to a preset semantic similarity threshold and the classification consistency evaluation parameter indicating that the candidate perturbed text and the sample text are classified consistently.
[0173] In an embodiment of the present disclosure, the second acquisition module 83 is further configured to: obtain a first classification loss of the candidate content discrimination model according to the first AI generation discrimination result and the sample text, and obtain a second classification loss of the candidate content discrimination model according to the second AI generation discrimination result and the sample perturbed text; obtain an output consistency loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result; and obtain a target training loss of the candidate content discrimination model based on the first classification loss, the second classification loss, and the output consistency loss.
[0174] In an embodiment of the present disclosure, the iteration module 84 is further configured to: acquire a candidate disturber and a candidate detector in the candidate content discrimination model, where the candidate detector includes a candidate semantic encoder, a candidate style encoder, and a candidate classifier; adjust the parameters of the candidate disturber and the candidate detector according to the target training loss to adjust the parameters of the candidate content discrimination model, and return to acquire the next sample text and the next sample perturbed text of the next sample text, and continue to train the candidate content discrimination model with adjusted parameters to obtain a trained target content discrimination model.
[0175] The training device of the content discrimination model proposed by the present disclosure obtains a candidate content discrimination model, corresponding sample texts and sample perturbation texts, and extracts the sample style semantic fusion features of the sample texts and sample perturbation texts respectively, so as to obtain the first AI-generated discrimination result output by the candidate content discrimination model based on the sample text and the second AI-generated discrimination result output based on the sample perturbation text. Optionally, the target training loss of the candidate content discrimination model is obtained based on the first AI-generated discrimination result and the second AI-generated discrimination result, and the candidate content discrimination model is iteratively modeled based on the target training loss to obtain the trained target content discrimination model. In the present disclosure, the candidate content discrimination model is trained based on the sample text and the sample perturbation text, which improves the robustness of the trained target content discrimination model, extracts features in the semantic dimension and style dimension of the text content input into the model, so that the candidate content discrimination model learns the differences between AI-generated content and human-generated content in the semantic dimension and style dimension respectively, improves the accuracy of the trained target content discrimination model in discriminating AI-generated content, reduces the possibility of misjudging AI-generated content in the scenario of discriminating AI-generated content based on the target content discrimination model, thereby reducing the occurrence probability of abnormal events caused by content misjudgment, improving the stickiness of content creators and users, and optimizing the discrimination method and discrimination effect of AI-generated content.
[0176] An embodiment of the present disclosure also proposes a content discrimination device. Since the content discrimination device proposed in the embodiment of the present disclosure corresponds to the content discrimination methods proposed in the above several embodiments, the implementation manners of the above content discrimination methods are also applicable to the content discrimination device proposed in the embodiment of the present disclosure and will not be described in detail in the following embodiments.
[0177] Figure 9 It is a schematic structural diagram of the content discrimination device according to an embodiment of the present disclosure, as Figure 9 shown, the content discrimination device 900 includes a third discrimination module 91, a feature extraction module 92, a second discrimination module 93 and a determination module 94, wherein:
[0178] The third discrimination module 91 is configured to obtain a text object to be discriminated and a trained target content discrimination model, wherein the target content discrimination model is obtained based on the above Figure 8 device proposed in the embodiment.
[0179] The feature extraction module 92 is configured to obtain a target detector in the target content discrimination model to obtain the target semantic feature and target style feature of the text object, wherein the target style feature is obtained based on the target triple feature of the text object;
[0180] The second discrimination module 93 is configured to perform feature fusion on the target semantic feature and the target style feature through a target detector to obtain the target semantic style fusion feature of the text object, and output the target AI generation discrimination result of the text object based on the target semantic style fusion feature through the target detector;
[0181] The determination module 94 is configured to determine the target text category to which the text object belongs based on the target AI generation discrimination result.
[0182] The content discrimination device proposed by the present disclosure obtains the target semantic feature and the target style feature of the text object through the target detector in the target content discrimination model, so as to obtain the target semantic style fusion feature of the two, and then obtains the target AI generation discrimination result of the text object output by the target detector based on the target semantic style fusion feature, and obtains the target text category to which the text object belongs based on the target AI generation discrimination result. In the present disclosure, the target detector in the trained target content discrimination model is used to perform category discrimination on whether the text object belongs to AI-generated content, which improves the accuracy and efficiency of the category determination of the text object, reduces the possibility of misjudgment of the category to which the text object belongs, reduces the occurrence probability of abnormal events caused by the confusion between manually written and AI-generated content, and optimizes the discrimination method and discrimination effect of AI-generated content.
[0183] According to an embodiment of the present disclosure, the present disclosure also proposes an electronic device, a readable storage medium, and a computer program product.
[0184] Figure 10 A schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0185] As Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1009 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0186] Multiple components in device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1006, such as various types of displays, speakers, etc.; a storage unit 1009, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0187] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the training method and / or the content discrimination method of the content discrimination model. For example, in some embodiments, the training method and / or the content discrimination method of the content discrimination model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1009. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the training method and / or the content discrimination method of the content discrimination model described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the training method and / or the content discrimination method of the content discrimination model by any other appropriate means (e.g., by means of firmware).
[0188] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0189] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0190] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0191] To initiate an interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user account can provide input to the computer. Other types of devices can also be used to initiate an interaction with the user account; for example, the feedback provided to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including acoustic input, voice input, or tactile input).
[0192] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user account computer having a graphical user account interface or a web browser through which the user account can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area networks (LANs), wide area networks (WANs), and the Internet.
[0193] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0194] It should be understood that the various forms of the processes shown above can be reordered, added to, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0195] The above specific embodiments do not limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A training method for a content discrimination model, characterized in that, The method includes: Obtaining a candidate content discrimination model to be trained, as well as the sample text of the candidate content discrimination model and the sample perturbation text of the sample text; Extracting the sample style semantic fusion features of the sample text and the sample perturbation text respectively to obtain a first AI generation discrimination result of the candidate content discrimination model for the sample text and a second AI generation discrimination result for the sample perturbation text; Obtaining the target training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result; Iterating the candidate content discrimination model according to the target training loss to obtain a trained target content discrimination model.
2. The method according to claim 1, wherein The extracting the sample style semantic fusion features of the sample text and the sample perturbation text respectively includes: Obtaining a candidate semantic encoder and a candidate style encoder in the candidate content discrimination model; Extracting a first semantic feature of the sample text and a second semantic feature of the sample perturbation text through the candidate semantic encoder; Extracting a first style feature of the sample text and a second style feature of the sample perturbation text through the candidate style encoder; Performing feature fusion on the first semantic feature and the first style feature to obtain a first fusion feature of the sample text, and performing feature fusion on the second semantic feature and the second style feature to obtain a second fusion feature of the sample perturbation text, so as to obtain the sample style semantic fusion features of the sample text and the sample perturbation text respectively.
3. The method according to claim 2, wherein The obtaining the first AI generation discrimination result of the candidate content discrimination model for the sample text and the second AI generation discrimination result for the sample perturbation text includes: Obtaining a candidate classifier in the candidate content discrimination model; Obtaining the first AI generation discrimination result of the sample text through the candidate classifier based on the first fusion feature, and obtaining the second AI generation discrimination result of the sample perturbation text through the candidate classifier based on the second fusion feature.
4. The method according to claim 3, wherein Before the extracting the first style feature of the sample text and the second style feature of the sample perturbation text through the candidate style encoder, it includes: Extracting a first triple feature of the sample text and a second triple feature of the sample perturbation text; Based on the sample text and the first triple feature, obtaining the input of the candidate style encoder to extract the first style feature of the sample text; Based on the sample perturbation text and the second triple feature, obtaining the input of the candidate style encoder to extract the second style feature of the sample perturbation text.
5. The method according to claim 1, wherein The obtaining the candidate content discrimination model to be trained, as well as the sample text of the candidate content discrimination model and the sample perturbation text of the sample text includes: Obtaining a candidate perturbator in the candidate content discrimination model; Inputting the sample text into the candidate perturbator to obtain the candidate perturbation text of the sample text; Perform semantic similarity evaluation and classification consistency evaluation on the candidate perturbation text and the sample text to obtain a semantic similarity evaluation parameter and a classification consistency evaluation parameter; In response to the semantic similarity evaluation parameter being greater than or equal to a preset semantic similarity threshold, and the classification consistency evaluation parameter indicating that the candidate perturbation text and the sample text are classified consistently, determine that the candidate perturbation text is the sample perturbation text of the sample text.
6. The method according to claim 1, wherein, The obtaining the target training loss of the candidate content discrimination model according to the first AI-generated discrimination result and the second AI-generated discrimination result includes: Obtain a first classification loss of the candidate content discrimination model according to the first AI-generated discrimination result and the sample text, and obtain a second classification loss of the candidate content discrimination model according to the second AI-generated discrimination result and the sample perturbation text; Obtain an output consistency loss of the candidate content discrimination model according to the first AI-generated discrimination result and the second AI-generated discrimination result; Based on the first classification loss, the second classification loss, and the output consistency loss, obtain the target training loss of the candidate content discrimination model.
7. The method according to claim 6, wherein The iterating the candidate content discrimination model according to the target training loss to obtain a trained target content discrimination model includes: Obtain a candidate perturbator and a candidate detector in the candidate content discrimination model, where the candidate detector includes a candidate semantic encoder, a candidate style encoder, and a candidate classifier; According to the target training loss, adjust the parameters of the candidate perturbator and the candidate detector to adjust the parameters of the candidate content discrimination model, and return to obtain the next sample text and the next sample perturbation text of the next sample text, and continue to train the candidate content discrimination model with adjusted parameters to obtain the trained target content discrimination model.
8. A content discrimination method, characterized in that, The method includes: Obtain a text object to be discriminated and a trained target content discrimination model, where the target content discrimination model is obtained based on the method described in any one of the above claims 1-7; Obtain a target detector in the target content discrimination model to obtain a target semantic feature and a target style feature of the text object, where the target style feature is obtained based on a target triple feature of the text object; Perform feature fusion on the target semantic feature and the target style feature through the target detector to obtain a target semantic style fusion feature of the text object, and output a target AI-generated discrimination result of the text object through the target detector based on the target semantic style fusion feature; Based on the target AI-generated discrimination result, determine the target text category to which the text object belongs.
9. A training device for a content discrimination model, characterized in that, The apparatus includes: A first acquisition module, configured to acquire a candidate content discrimination model to be trained, a sample text of the candidate content discrimination model, and a sample perturbation text of the sample text; The first discrimination module is configured to extract the sample style semantic fusion features of the sample text and the sample perturbation text respectively, so as to obtain the first AI generation discrimination result of the candidate content discrimination model for the sample text and the second AI generation discrimination result for the sample perturbation text; The second acquisition module is configured to obtain the target training loss of the candidate content discrimination model according to the first AI generation discrimination result and the second AI generation discrimination result; The iteration module is configured to iterate the candidate content discrimination model according to the target training loss to obtain a trained target content discrimination model.
10. A content discrimination device, characterized in that, The device includes: The third discrimination module is configured to obtain a text object to be discriminated and a trained target content discrimination model, where the target content discrimination model is obtained based on the device according to any one of claims 1-7 above; The feature extraction module is configured to obtain a target detector in the target content discrimination model to obtain the target semantic feature and the target style feature of the text object, where the target style feature is obtained based on the target triple feature of the text object; The second discrimination module is configured to perform feature fusion on the target semantic feature and the target style feature through the target detector to obtain the target semantic style fusion feature of the text object, and output the target AI generation discrimination result of the text object based on the target semantic style fusion feature through the target detector; The determination module is configured to determine the target text category to which the text object belongs based on the target AI generation discrimination result.
11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7 and / or 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7 and / or 8.
13. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7 and / or 8.