An anti-adhesion film for a semiconductor device

By introducing an adversarial topic awareness method into cross-topic automatic essay scoring, combined with topic sharing and specific topic prompts, the problem of neglecting topic-related features in existing methods is solved, achieving higher scoring accuracy and robustness.

CN119849477BActive Publication Date: 2025-11-28SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510024090.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-11-28
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing cross-topic automatic essay scoring methods mostly focus on learning shared features of the topic while neglecting the learning of topic-related features, resulting in insufficient scoring accuracy across different topics.

Method used

We adopt an adversarial topic awareness approach, which introduces topic-sharing and topic-specific cues, combines a joint modeling framework of regression and classification, utilizes a pre-trained model for feature representation, and learns topic-sharing and topic-specific features through adversarial training and pseudo-label generation techniques, thereby improving the robustness and prediction accuracy of the model.

Benefits of technology

It improves the accuracy and robustness of cross-topic automatic essay scoring, better adapts to the characteristics of different topics, and enhances the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849477B_ABST
    Figure CN119849477B_ABST
Patent Text Reader

Abstract

The application discloses an automatic essay scoring method and system based on an adversarial topic perception, and mainly relates to the technical field of automatic essay scoring across topics, and aims to solve the problem that existing methods mostly focus on learning topic shared features while ignoring topic related feature learning. The method comprises the following steps: obtaining a scored essay set of a source topic and an un-scored essay set of a target topic; setting a topic shared prompt, setting a specific topic prompt of each topic, splicing the essays to obtain a prompt enhanced embedding x, and obtaining a feature representation; calculating a scoring loss of a current regression task through predicted scores of each attribute; learning the topic shared prompt and the specific topic prompt based on adversarial training and pseudo-label generation technology of a classification task, and obtaining a total classification loss and an adversarial loss; and determining a final loss according to the regression loss, the total classification loss and the adversarial loss.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-topic automatic essay scoring, and in particular to a cross-topic automatic essay scoring method and system based on adversarial topic awareness. BACKGROUND

[0002] Automatic essay scoring aims to automatically evaluate the quality of essays written by students according to various writing topics. Automatic essay scoring systems can reduce the manual scoring burden of educators while providing timely and comprehensive feedback to students. Automatic essay scoring is usually formulated as a regression task, and early methods mainly rely on handcrafted features. Recently, deep learning has greatly promoted the development of the field of automatic essay scoring, and some works have achieved good results by utilizing the powerful feature learning ability of deep neural networks.

[0003] Generally, automatic essay scoring methods based on deep learning can be divided into specific topic-based automatic essay scoring and cross-topic automatic essay scoring according to whether the training data and test data come from the same topic. Specific topic-based automatic essay scoring methods aim to train an automatic essay scoring model using training and test data from the same topic. Most existing methods are specific topic-based automatic essay scoring methods, but it is difficult to collect a large number of specific prompt scored essays in actual scenarios. Therefore, cross-topic automatic essay scoring has emerged. Cross-topic automatic essay scoring methods aim to train a model with good transferability through data from a source topic, so that it can make accurate scoring on a target topic.

[0004] However, existing cross-topic automatic essay scoring methods do not apply prompt fine-tuning techniques to cross-topic automatic essay scoring. In addition, most existing methods focus on learning topic-shared features while ignoring the learning of topic-related features. SUMMARY

[0005] In view of the above deficiencies of the prior art, the present application provides a cross-topic automatic essay scoring method and system based on adversarial topic awareness to solve the problem that most existing methods focus on learning topic-shared features while ignoring the learning of topic-related features.

[0006] In a first aspect, the present application provides a cross-topic automatic essay scoring method based on adversarial topic awareness, the method comprising:

[0007] obtaining a set of scored essays of a source topic and a set of un-scored essays of a target topic; setting topic-shared prompts and specific topic prompts for each topic; and ​The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features ;Establish a joint modeling framework for regression and classification by A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. ;

[0008] In a joint regression and classification modeling framework, a classifier C is built consisting of an MLP layer with a softmax activation function. The essays in the graded essay set of each source topic are input into the classifier, and then the classifier is calculated to obtain... Classification loss of individual source topics Adversarial training and pseudo-label generation techniques based on classification tasks are used to learn topic-shared cues and topic-specific cues, respectively, to obtain adversarial loss. Classification loss of target topic , and then combine Classification loss of individual source topics Obtain total classification loss and combat losses Based on the regression loss, total classification loss, and adversarial loss, the final loss of the joint modeling framework is determined, and the trained joint modeling framework is obtained.

[0009] In one implementation of this application, obtaining The source topic includes a set of scored essays, and the target topic includes a set of unscored essays, specifically:

[0010] Get A collection of graded essays on a single topic. ;in, Indicates the first A collection of graded essays on various topics. Let represent the s-th rated essay in the i-th source topic, where i ∈ [1, 2, 3]. ], s∈[1, ], Number of essays under a set This represents the set of essays for the i-th source topic. Indicates and The corresponding essay score, This represents the score of the s-th graded essay. Let K represent the score value of the Kth attribute of the s-th rated essay in the i-th source topic, where K represents the number of scoring attributes in the essay. Collection of essays The corresponding score category This represents the category of the s-th graded essay in the i-th source topic. C represents the total number of scoring categories; retrieve the set of unscored essays on the target topic. ;in, Let t represent the t-th unscored essay, t∈[1, ..., ... ] This indicates the number of essays in the unrated essay collection.

[0011] In one implementation of this application, topic sharing prompts and specific topic prompts for each topic are set; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features

[0012] Setting up topic sharing prompts ;in, This represents the nth learnable vector; for each topic, a topic-specific hint is set. ;in, Represent the m-th learnable vector under the i-th topic; obtain the essay. ,in, Indicates the length of the essay This represents the Lth word in the essay;

[0013] Through the formula:

[0014] The essay's shared topic hints and corresponding topic-specific hints are combined to obtain the hint enhancement embedding x;

[0015] in, and These represent the start and end identifiers, respectively. Represents the embedding operation of a pre-trained model; The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features: ;in, , This represents the dimension of the output features of the pre-trained model. This represents a pre-trained model.

[0016] In one implementation of this application, a system is established by A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. Specifically, it includes:

[0017] Establish Each feature attribute generator;

[0018] Feature representation Labeled features of the input By splicing them together, we can obtain ;Will Enter the number The nonlinear feature transformation layer in the scorer obtains the... Feature representation of each attribute scoring task :

[0019] ;in and This represents the weight matrix and bias terms of the nonlinear layer. Represents a non-linear activation function;

[0020] Introducing the attention mechanism formula:

[0021]

[0022]

[0023]

[0024]

[0025] Calculate the first The final attribute feature representation of each attribute ;in, It is a concatenation of the feature representations for each attribute scoring task. Indicates the first The attribute feature for the first Attention weights for each attribute; Indicates the first Attention features of each attribute, Indicates the first The final attribute feature representation of each attribute;

[0026] For each attribute scoring task, use with The linear layer of the activation function will Mapped to an attribute score:

[0027] in, Indicates the first Predicted scores for each attribute, and These represent the corresponding weight matrix and bias term, respectively. This represents the sigmoid activation function;

[0028] Mean squared error is used as the loss function for each attribute scoring task, combined with A collection of graded essays on a single topic and The final rating loss is calculated based on the following attributes:

[0029]

[0030]

[0031] in, This represents the final score loss. Let i represent the set of rated essays for the i-th source topic, where i∈[1,N].

[0032] In one implementation of this application, a classifier C is constructed consisting of an MLP layer with a softmax activation function, which... The essays in the graded essay set of each source topic are input into the classifier, and then the classifier is calculated to obtain... Classification loss of individual source topics Specifically, it includes:

[0033] The overall score is divided into four levels: excellent, good, average, and poor. Essays from a set of graded essays on a given topic are input into a classifier. The classifier C consists of an MLP layer with a softmax activation function.

[0034] For each essay in the current source topic's set of scored essays, the cross-entropy loss between the predicted result and the true label is calculated as the model's classification loss:

[0035]

[0036] wherein, represents the prediction result of the i-th source topic in the scored essay set output by the classifier, represents the true label in the scored essay set of the i-th source topic;

[0037] For source topic datasets, the overall scoring classification loss calculation process is as follows:

[0038]

[0039] wherein, represents the classification loss.

[0040] In an implementation manner of the present application, the topic sharing prompt and the specific topic prompt are learned based on the adversarial training and the pseudo label generation technology of the classification task, respectively, to obtain the adversarial loss and the classification loss of the target topic , and then the classification loss of source topics is combined to obtain the total classification loss and the adversarial loss , specifically including:

[0041] The adversarial training module for the topic sharing prompt is implemented as follows:

[0042] A topic discriminator is designed for each source topic and the target topic to determine whether the input essay is from the target topic; the total topic discriminators of the source topics are represented as ; for the essays from the source topics or the target topic , they are input into the corresponding discriminators;

[0043] The probability that belongs to the target topic is calculated by the formula:

[0044]

[0045] ;

[0046] The adversarial loss is calculated by the formula:

[0047]

[0048] ;

[0049] ​The pseudo-tag generation module for specific topic suggestions is implemented as follows:

[0050] For each essay on the target topic , obtain feature representation and category probability prediction results ;

[0051] Sharpening the category probability prediction:

[0052] ,

[0053] in, , Indicate category The category prediction probability; the feature and category probability prediction results of all target topic essays are stored in a memory bank. In this process, an exponential moving average strategy is used for updating at each iteration:

[0054] ,

[0055] in, It is a smoothing hyperparameter. and This represents the feature representation and class probability prediction results of the current sample in the memory.

[0056] Find the cosine similarity between the feature representation of the current sample and the feature representations of all essays stored in the memory bank. of One nearest neighbor sample;

[0057] Through the formula:

[0058]

[0059] will come from Probabilistic prediction aggregation is performed on the categories of the nearest neighbor samples to provide... Generate category probability prediction and pseudo-tags ;

[0060] in, Indicates the current sample of The set of nearest neighbor indexes in the memory;

[0061] Calculate the cross-entropy loss of the target topic essay based on the generated pseudo-labels:

[0062]

[0063] Therefore, the total classification loss is calculated as follows:

[0064] .

[0065] In one implementation of this application, the final loss is determined based on regression loss, classification loss, and adversarial loss, specifically including:

[0066] Through the formula: Determine the final loss ;in, and This indicates the preset trade-off parameters.

[0067] Secondly, this application provides a cross-topic automatic essay scoring system based on adversarial topic awareness, the system comprising:

[0068] The theme-based essay acquisition module is used to acquire... A collection of graded essays on a single topic. A collection of unrated essays on the target topic

[0069] The suggestion enhancement module for the essay encoder is used to set shared suggestion for each topic and specific suggestion for each topic; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features ;

[0070] The joint modeling module is used to build a joint modeling framework for regression and classification. A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. In a joint modeling framework of regression and classification, a classifier C is built consisting of an MLP layer with a softmax activation function. The essays in the graded essay set of each source topic are input into the classifier, and then the classifier is calculated to obtain... Classification loss of individual source topics Adversarial training and pseudo-label generation techniques based on classification tasks are used to learn topic-shared cues and topic-specific cues, respectively, to obtain adversarial loss. Classification loss of target topic , and then combine Classification loss of individual source topics Obtain total classification loss and combat losses Based on the regression loss, total classification loss, and adversarial loss, the final loss of the joint modeling framework is determined, and the trained joint modeling framework is obtained.

[0071] In one implementation of this application, the topic-based essay acquisition module includes a source topic acquisition unit and a target topic acquisition unit.

[0072] The source topic acquisition unit is used to acquire... A collection of graded essays on a single topic. ;in, Indicates the first A collection of graded essays on various topics. Let represent the s-th rated essay in the i-th source topic, where i ∈ [1, 2, 3]. ], s∈[1, ], Number of essays under a set This represents the set of essays for the i-th source topic. Indicates and The corresponding essay score, This represents the score of the s-th graded essay. Let K represent the score value of the Kth attribute of the s-th rated essay in the i-th source topic, where K represents the number of scoring attributes in the essay. Collection of essays The corresponding score category This represents the category of the s-th graded essay in the i-th source topic. C represents the total number of scoring categories; the target topic acquisition unit is used to acquire the set of unscored essays on the target topic. ;in, Let t represent the t-th unscored essay, t∈[1, ..., ... ] This indicates the number of essays in the unrated essay collection.

[0073] In one implementation of this application, the prompt-enhanced essay encoder module includes a feature representation unit.

[0074] Used to set up topic sharing prompts ;in, This represents the nth learnable vector; for each topic, a topic-specific hint is set. ;in, Represent the m-th learnable vector under the i-th topic; obtain the essay. wherein, denotes the length of the composition denotes the Lth word in the composition;

[0075] By formula:

[0076] Splicing the composition theme shared prompt and the specific theme prompt corresponding to the theme to obtain a prompt enhanced embedding x;

[0077] wherein, and respectively represent the start and end identifiers, represents an embedding operation of the pre-trained model; inputting into the pre-trained model, obtaining a feature representation at the position of the pre-trained model output: ; wherein, , denotes the dimension of the pre-trained model output feature, denotes the pre-trained model.

[0078] Those skilled in the art can understand that the present application has at least the following beneficial effects:

[0079] The application discloses a cross-theme automatic composition scoring method and system based on adversarial theme perception, introduces theme shared prompts and specific theme prompts for each theme, guides the corresponding knowledge of the pre-trained model, and comprehensively improves the performance by using the pre-trained model. In the joint modeling framework of regression and classification, the theme shared prompt is learned by introducing adversarial training, thereby enhancing the robustness of the model to feature scale changes. At the same time, the specific theme prompt is learned by using the pseudo label generation technology, and the prediction accuracy of the model is further improved. The problem that existing methods mostly focus on learning theme shared features while ignoring theme related feature learning is solved. BRIEF DESCRIPTION OF DRAWINGS

[0080] In order to more clearly illustrate the technical solutions of the present application, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0081] Figure 1 is a flow chart of a cross-theme automatic composition scoring method based on adversarial theme perception provided by an embodiment of the present application.

[0082] Figure 2 is an internal structure schematic diagram of a cross-theme automatic composition scoring system based on adversarial theme perception provided by an embodiment of the present application. DETAILED DESCRIPTION

[0083] It is to be understood that the embodiments described hereinbelow are merely preferred embodiments of the present disclosure and do not represent the only embodiment that can be implemented in the present disclosure. The preferred embodiments are merely used to explain the technical principles of the present disclosure and should not be used to limit the scope of protection of the present disclosure. Based on the preferred embodiments provided in the present disclosure, any other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present disclosure.

[0084] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices that comprise a list of elements not only include those elements, but also include other elements that are not expressly listed, or other elements inherent in such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.

[0085] The technical solutions of the embodiments of the present application will be described in detail below with reference to the drawings.

[0086] The embodiments provide a cross-topic automatic essay scoring method based on adversarial topic perception, as shown in Figure 1 The method provided by the embodiments of the present application mainly includes the following steps:

[0087] Step 110, obtaining a scored essay set of a source topic and an un-scored essay set of a target topic. In some embodiments, the step can be specifically:

[0088] Obtaining a scored essay set of a source topic

[0089] ; wherein, represents the scored essay set of the i th source topic, represents the s th scored essay of the i th source topic, wherein i∈[1, ], s∈[1, ], represents the number of essays under the set, represents the essay set of the i th source topic, represents the essay score corresponding to , and represents the essay score of the s th scored essay, ​​​Let K represent the score value of the Kth attribute of the s-th rated essay in the i-th source topic, where K represents the number of scoring attributes in the essay. Collection of essays The corresponding score category This represents the category of the s-th graded essay in the i-th source topic. C represents the total number of scoring categories; retrieve the set of unscored essays on the target topic. ;in, Let t represent the t-th unscored essay, t∈[1, ..., ... ] This indicates the number of essays in the unrated essay collection.

[0090] Step 120: Set up topic sharing tips and specific topic tips for each topic; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features .

[0091] This step can be specifically described as follows:

[0092] Setting up topic sharing prompts ;in, This represents the nth learnable vector; for each topic, a topic-specific hint is set. ;in, Represent the m-th learnable vector under the i-th topic; obtain the essay. ,in, Indicates the length of the essay This represents the Lth word in the essay;

[0093] Through the formula:

[0094] The essay's shared topic hints and corresponding topic-specific hints are combined to obtain the hint enhancement embedding x;

[0095] in, and These represent the start and end identifiers, respectively. Represents the embedding operation of a pre-trained model; The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features: ;in, , This indicates the dimension of the output features of the pre-trained model, which will be used subsequently. As a characteristic feature of composition This represents a pre-trained model.

[0096] Step 130: Establish a joint modeling framework for regression and classification by... A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted score of each attribute; based on the predicted score of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. .

[0097] It should be noted that j∈[1,K].

[0098] This step can be specifically described as follows:

[0099] Establish Each feature attribute generator;

[0100] Feature representation Labeled features of the input By splicing them together, we can obtain ;Will Enter the number The nonlinear feature transformation layer in the scorer obtains the... Feature representation of each attribute scoring task :

[0101] ;in and This represents the weight matrix and bias terms of the nonlinear layer. Represents a non-linear activation function;

[0102] Introducing the attention mechanism formula:

[0103]

[0104]

[0105]

[0106]

[0107] Calculate the first The final attribute feature representation of each attribute ;in, It is a concatenation of the feature representations for each attribute scoring task. Indicates the first attribute feature of the i-th attribute; the attention weight of the i-th attribute; the attention feature of the i-th attribute, the final attribute feature representation of the i-th attribute;

[0108] For each attribute score task, a linear layer with an activation function is used to map to an attribute score:

[0109] wherein, the predicted score of the i-th attribute, and respectively represent the corresponding weight matrix and bias term, represents the sigmoid activation function. The mean square error is used as the loss function of each attribute score task, combined with the scored essay set of the i-th source topic

[0110] attributes, and the final scoring loss is calculated as follows:

[0111]

[0112]

[0113] wherein, the final scoring loss, represents the scored essay set of the i-th source topic, i∈[1,N].

[0114] Step 140, in the joint modeling framework of regression and classification, a classifier C composed of an MLP layer with a softmax activation function is established, and the essays in the scored essay set of the i-th source topic are input into the classifier, and then the classification loss of the i-th source topic is obtained.

[0115] In some embodiments, this step can be specifically: dividing the overall score into four levels of excellent, good, medium and poor; inputting the essays in the scored essay set of the i-th source topic into the classifier C; wherein the classifier C is composed of an MLP layer with a softmax activation function:

[0116] ​​​​​​​​For the composition in the scored composition set of the current source topic, the cross-entropy loss between the prediction result and the true label is calculated as the classification loss of the model:

[0117]

[0118] wherein, represents the prediction result of the i-th source topic in the scored composition set output by the classifier, represents the true label in the scored composition set of the i-th source topic;

[0119] For the source topic data set, the classification loss calculation process of the overall score is as follows:

[0120]

[0121] wherein, represents the classification loss.

[0122] Step 150, based on the classification task, the adversarial training and the pseudo label generation technology learn the topic sharing prompt and the specific topic prompt respectively, obtain the adversarial loss and the classification loss of the target topic , and then combine the classification loss of the source topic to obtain the total classification loss and the adversarial loss .

[0123] This step can be specifically: dividing the overall score into four grades of excellent, good, and poor; the composition in the scored composition set of the source topic ; wherein the classifier C is composed of an MLP layer with a softmax activation function:

[0124] For the composition in the scored composition set of the current source topic, the cross-entropy loss between the prediction result and the true label is calculated as the classification loss of the model:

[0125]

[0126] wherein, represents the prediction result of the i-th source topic in the scored composition set output by the classifier, represents the true label in the scored composition set of the i-th source topic;

[0127] For the source topic data set, the classification loss calculation process of the overall score is as follows:

[0128]

[0129] wherein, represents a classification loss.

[0130] Step 160, determining the final loss of the joint modeling framework according to the regression loss, the total classification loss and the adversarial loss, obtaining the trained joint modeling framework.

[0131] The final loss is determined by the formula: ; wherein, and represents a preset trade-off parameter.

[0132] Based on the above description, in order to better learn the theme shared features and theme related features, and apply the prompt fine-tuning technology to the field of cross-theme automatic essay scoring in the present application. In the present application, an adversarial theme-aware prompt fine-tuning method for cross-theme automatic essay scoring is proposed. First, the definition of cross-theme automatic essay scoring is described. Then, the theme-aware prompt enhancement coding module is described in detail, which contains theme-shared prompts and theme-specific prompts. At the same time, the joint modeling module is introduced, which models the automatic essay scoring as a classification and regression task at the same time. Finally, the theme-shared prompt adversarial training module and the theme-specific prompt pseudo-label generation module are introduced, which are used to learn the two prompts in the coding module, respectively, to promote the learning of theme-shared features and theme-related features.

[0133] In addition, the present application Figure 2 provides an adversarial theme-aware cross-theme automatic essay scoring system. As Figure 2 shown, the system provided by the present application mainly includes:

[0134] The theme essay acquisition module 210 is configured to acquire a set of scored essays of a source theme and a set of un-scored essays of a target theme.

[0135] The theme essay acquisition module 210 includes a source theme acquisition unit and a target theme acquisition unit,

[0136] The source theme acquisition unit is configured to acquire a set of scored essays of a source theme ; wherein, represents a set of scored essays of the i-th source theme, represents the s-th scored essay of the i-th source theme, wherein i∈[1, ], s∈[1, ], the number of essays under the set​​ This represents the set of essays for the i-th source topic. Indicates and The corresponding essay score, This represents the score of the s-th graded essay. Let K represent the score value of the Kth attribute of the s-th rated essay in the i-th source topic, where K represents the number of scoring attributes in the essay. Collection of essays The corresponding score category This represents the category of the s-th graded essay in the i-th source topic. C represents the total number of scoring categories;

[0137] The target topic acquisition unit is used to retrieve a collection of unrated essays on a target topic. ;in, Let t represent the t-th unscored essay, t∈[1, ..., ... ] This indicates the number of essays in the unrated essay collection.

[0138] The suggestion enhancement essay encoder module 220 is used to set topic-shared suggestions and topic-specific suggestions for each topic; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features .

[0139] The prompt-enhanced essay encoder module 220 includes feature representation units.

[0140] Used to set up topic sharing prompts ;in, This represents the nth learnable vector; for each topic, a topic-specific hint is set. ;in, Represent the m-th learnable vector under the i-th topic; obtain the essay. ,in, Indicates the length of the essay This represents the Lth word in the essay;

[0141] Through the formula:

[0142] The essay's shared topic hints and corresponding topic-specific hints are combined to obtain the hint enhancement embedding x;

[0143] in, and These represent the start and end identifiers, respectively. Represents the embedding operation of a pre-trained model; The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features: ;in, , This represents the dimension of the output features of the pre-trained model. This represents a pre-trained model.

[0144] Joint modeling module 230 is used to build a joint modeling framework for regression and classification. A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. In a joint modeling framework of regression and classification, a classifier C is built consisting of an MLP layer with a softmax activation function. The essays in the graded essay set of each source topic are input into the classifier, and then the classifier is calculated to obtain... Classification loss of individual source topics Adversarial training and pseudo-label generation techniques based on classification tasks are used to learn topic-shared cues and topic-specific cues, respectively, to obtain adversarial loss. Classification loss of target topic , and then combine Classification loss of individual source topics Obtain total classification loss and combat losses Based on the regression loss, total classification loss, and adversarial loss, the final loss of the joint modeling framework is determined, and the trained joint modeling framework is obtained.

[0145] The technical solutions of this disclosure have been described in conjunction with the preceding embodiments. However, it will be readily understood by those skilled in the art that the scope of protection of this disclosure is not limited to these specific embodiments. Without departing from the technical principles of this disclosure, those skilled in the art can disassemble and combine the technical solutions in the above embodiments, and can also make equivalent changes or substitutions to the relevant technical features. Any changes, equivalent substitutions, improvements, etc., made within the technical concept and / or technical principles of this disclosure will fall within the scope of protection of this disclosure.

Claims

1. An adversarially theme-aware based cross-theme automated essay scoring method, characterized in that, The method comprises: acquiring a scored set of essays of a source topic and an un-scored set of essays of a target topic; Set up topic sharing tips and specific topic tips for each topic; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features ; Establishing a joint modeling framework for regression and classification by A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. ; In a joint modeling framework of regression and classification, a classifier C consisting of an MLP layer with a softmax activation function is built, which takes as input an essay from a scored set of essays of one source topic, and computes the probability distribution over the set of classes for the source topic The classification loss for the source topic is computed as ; The classification task-based adversarial training and the pseudo-label generation technology respectively learn the topic sharing prompt and the specific topic prompt to obtain an adversarial loss and the classification loss of the target topic , and then combine the classification loss of the source topic to obtain a total classification loss and the adversarial loss ; According to the regression loss, the total classification loss and the adversarial loss, the final loss of the joint modeling framework is determined, and a trained joint modeling framework is obtained.

2. The cross-topic automated essay scoring method based on adversarial theme awareness according to claim 1, wherein, acquiring a scored essay set of a source topic and an un-scored essay set of a target topic, specifically comprising: Obtaining a scored essay set of the i-th source topic ; wherein, represents the i-th source topic, a scored essay set of the i-th source topic, represents the s-th scored essay of the i-th source topic, wherein i∈[1, ], s∈[1, ], represents the number of essays under the set, represents the essay set of the i-th source topic, represents the essay score corresponding to , represents the essay score of the s-th scored essay, represents the score value of the K-th attribute of the s-th scored essay of the i-th source topic, K represents the number of scoring attributes of the essay, represents the classification corresponding to the essay set , represents the classification of the s-th scored essay of the i-th source topic, , C represents the total number of classification categories; obtaining a non-scored essay set of a target topic ; wherein, represents the t-th non-scored essay, t∈[1, ], represents the number of essays in the non-scored essay set.​ 3.The cross-topic automated essay scoring method based on adversarial topic-awareness according to claim 1, wherein, Setting a theme sharing prompt, a specific theme prompt of each theme;Will Each composition in the scored composition set of the source theme and the ungraded composition set of the target theme is spliced with the theme sharing prompt and the specific theme prompt of the composition corresponding theme to obtain a prompt enhanced embedding x, and Input into the pre-training model, and obtain a feature representation Position output by the pre-training model , Specifically includes: Setting topic sharing prompt ; wherein, represents the nth learnable vector; for each topic, setting a specific topic prompt ; wherein, represents the mth learnable vector under the ith topic; obtaining an essay , wherein, represents the length of the essay, represents the Lth word in the essay; Through the formula: The theme sharing prompt of the composition and the specific theme prompt corresponding to the theme are spliced to obtain a prompt enhancement embedding x; wherein, and respectively represent a start and end identifier, represents an embedding operation of the pre-trained model; the input into the pre-trained model, and a feature representation is obtained at a position of the pre-trained model output: ; wherein, , represents a dimension of the pre-trained model output feature, represents the pre-trained model.

4. The cross-topic automated essay scoring method based on adversarial theme perception according to claim 1, wherein, Establish by A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. Specifically, it includes: establishing a feature attribute generator; The feature representation is inputted into the nonlinear feature conversion layer in the first score r, and a feature representation corresponding to the first attribute score task is obtained . The feature representation is inputted into the nonlinear feature conversion layer in the first score r, and a feature representation corresponding to the first attribute score task is obtained . : ; wherein and denote the weight matrix and bias term of the nonlinear layer, denotes a nonlinear activation function; The attention mechanism formula is introduced: Calculate the first The final attribute feature representation of each attribute ;in, It is a concatenation of the feature representations for each attribute scoring task. Indicates the first The attribute feature for the first Attention weights for each attribute; Indicates the first Attention features of each attribute, Indicates the first The final attribute feature representation of each attribute; For each attribute rating task, a linear layer with activation function maps to an attribute score: wherein, represents the predicted score of the th attribute, and respectively represent the corresponding weight matrix and bias term, represents the sigmoid activation function; Using mean square error as the loss function for each attribute scoring task, combining a set of scored essays for the source topic and attributes, the final scoring loss is computed as follows: wherein, represents the final score loss, represents the scored essay set of the i-th source topic, i ∈ [1, N].

5. The cross-theme auto-essay scoring method based on adversarial theme-awareness according to claim 1, wherein, A classifier C consisting of a MLP layer with a softmax activation function is established, and the composition of the classifier C is as follows: The composition of the classifier C is as follows: The composition of the classifier C is as follows: The composition of the classifier C is as follows: The overall score is divided into four levels of excellent, good, medium and poor; the The compositions in the scored composition set of the source topic are input to a classifier ; wherein the classifier C is composed of an MLP layer with a softmax activation function: For the composition in the scored composition set of the current source topic, the cross-entropy loss between the prediction result and the real label is calculated as the classification loss of the model: wherein, represents the predicted outcome in the scored essay set for the i-th source topic output by the classifier, represents the true label in the scored essay set for the i-th source topic. For each source topic dataset, the classification loss computation for the overall score is as follows: wherein, denotes the classification loss.

6. The cross-theme auto-essay scoring method based on adversarial theme perception according to claim 1, wherein, The classification task-based adversarial training and the pseudo-label generation technology learn the topic sharing prompt and the specific topic prompt respectively, and obtain an adversarial loss and the classification loss of the target topic , and then combine the classification loss of the source topic to obtain a total classification loss and the adversarial loss , specifically comprising: The implementation of the adversarial training module for the topic sharing prompt is as follows: For each source topic and target topic Design a topic discriminator. Determine whether the input essay originates from the target topic; [and then analyze the total number of source essays]. The topic discriminator is represented as follows: For essays from the source topic or the target topic The input is then fed into the corresponding discriminator; Through the formula: Computing the probability of belonging to the target topic ; Through the formula: Computing the adversarial loss ; The implementation of the pseudo-label generation module for the specific topic prompt is as follows: For each target topic, the essay , obtains a feature representation and a class probability prediction result ; The class probability prediction is sharpened: , wherein, , representing the class prediction probability of the class ; storing the features of all target topic compositions and the class probability prediction results in the memory bank and updating at each iteration using an exponential moving average strategy: , wherein, is a smoothing hyperparameter, and denote the feature representation and the class probability prediction result of the current sample in the memory bank; finding the k nearest neighbors of the current sample by computing the cosine similarity between the feature representation of the current sample and the feature representations of all the essays stored in the memory bank of nearest neighbor samples Through the formula: Aggregating the class probability predictions for the k nearest neighbor samples to generate a class probability prediction for the input sample Generating a class probability prediction And pseudo labels ;​ in, Indicates the current sample of The set of nearest neighbor indexes in the memory; The cross-entropy loss of the composition of the target topic is calculated based on the generated pseudo-label: Further, the total classification loss is calculated as follows: 。 7. The cross-topic automatic composition scoring method based on adversarial topic awareness according to claim 1, characterized in that, According to the regression loss, the classification loss and the adversarial loss, the final loss of the joint modeling framework is determined, and a trained joint modeling framework is obtained, specifically comprising: The final loss is determined by the formula: ; wherein, and represents a preset trade-off parameter.​ 8. An adversarially theme-aware based cross-theme automated essay scoring system, characterized by, The system comprises: a subject essay acquisition module for acquiring a scored essay set of the source subject an un-scored essay set of the target subject The suggestion enhancement module for the essay encoder is used to set shared suggestion for each topic and specific suggestion for each topic; The set of scored essays from the source topic and the set of unscored essays from the target topic are concatenated with the shared prompts for each essay and the specific topic prompts for each essay's corresponding topic, to obtain the prompt enhancement embedding x. The input is fed into the pre-trained model, and the output of the pre-trained model is... Location is represented by features ; The joint modeling module is used to build a joint modeling framework for regression and classification. A regression task consisting of feature attributers, and then using essay feature representation Input labeled features Attention mechanism, calculating the current essay in the [number]th position. The predicted scores of each attribute; where j∈[1,K]; through the predicted scores of each attribute, A collection of graded essays on a single topic and 1. Calculate the score loss for the current regression task based on 1. Attributes. In a joint modeling framework of regression and classification, a classifier C is built consisting of an MLP layer with a softmax activation function. The essays in the graded essay set of each source topic are input into the classifier, and then the classifier is calculated to obtain... Classification loss of individual source topics Adversarial training and pseudo-label generation techniques based on classification tasks are used to learn topic-shared cues and topic-specific cues, respectively, to obtain adversarial loss. Classification loss of target topic , and then combine Classification loss of individual source topics Obtain total classification loss and combat losses Based on the regression loss, total classification loss, and adversarial loss, the final loss of the joint modeling framework is determined, and the trained joint modeling framework is obtained.

9. The cross-topic automated essay scoring system based on adversarial theme awareness according to claim 8, wherein, The topic composition acquisition module comprises a source topic acquisition unit and a target topic acquisition unit, The source topic acquisition unit is used to acquire... A collection of graded essays on a single topic. ;in, Indicates the first A collection of graded essays on various topics. Let represent the s-th rated essay in the i-th source topic, where i ∈ [1, 2, 3]. ], s∈[1, ], Number of essays under a set This represents the set of essays for the i-th source topic. Indicates and The corresponding essay score, This represents the score of the s-th graded essay. Let K represent the score value of the Kth attribute of the s-th rated essay in the i-th source topic, where K represents the number of scoring attributes in the essay. Collection of essays The corresponding score category This represents the category of the s-th graded essay in the i-th source topic. C represents the total number of scoring categories; The target subject obtaining unit is configured to obtain an ungraded composition set of a target subject ; wherein, denotes the tth ungraded composition, t∈[1, ] denotes the number of compositions in the ungraded composition set.

10. The cross-theme automated essay scoring system based on adversarial theme awareness according to claim 8, wherein, The prompt enhanced composition encoder module comprises a feature representation unit, For setting a theme sharing prompt ; wherein, represents the nth learnable vector; for each theme, set a specific theme prompt ; wherein, represents the mth learnable vector under the ith theme; obtain an essay , wherein, represents the length of the essay represents the Lth word in the essay Through the formula: The theme sharing prompt of the composition and the specific theme prompt corresponding to the theme are spliced to obtain a prompt enhanced embedding x; wherein, and respectively represent a start and end identifier, represents an embedding operation of a pre-trained model; the input into the pre-trained model, and a feature representation is obtained at a position of the pre-trained model output: ; wherein, , represents a dimension of the pre-trained model output feature, represents a pre-trained model.

Citation Information

Patent Citations

  • Cross-theme composition automatic scoring method and system and medium

    CN115455178A

  • Cross-theme composition automatic evaluation method and system based on paired double-layer adversarial alignment

    CN117648921A