Generation method and device for common sense type after-class exercises in low-resource scene
Patent Information
- Application Number
- CN202380078034.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-06-27
AI Technical Summary
Existing technology is difficult to generate common sense after-school exercises in low-resource scenarios, especially due to the lack of a large-scale common sense database and high labeling costs. As a result, the machine question generation model has insufficient performance in practical applications and cannot effectively support intelligent education and natural science. needs in the field of language processing.
Pre-training language models are used for post-training, combined with encyclopedia concepts and causal relationships in external knowledge graphs. Through decoupled learning and adversarial frameworks, the generator design contains encoders and decoders, using variational reasoning and hint learning techniques. Generate diverse common sense reasoning exercises and reduce dependence on annotated data.
Generating high-quality common sense after-school exercises in low-resource scenarios improves the robustness and generalization capabilities of the model, can effectively alleviate the problem of resource scarcity, supports intelligent education and other applications, and is significantly better than traditional methods.
Smart Images

Figure CN120226003A_ABST
Abstract
Description
A method and device for generating common sense after-class exercises in low-resource scenarios Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for generating common sense after-class exercises in a low-resource scenario. Background Art
[0002] Common sense reasoning exercises differ from simple word matching questions in that they require not only coherent expression but also common sense reasoning. This requires the ability to reason across multiple lines of evidence, potentially involving common sense entities and relationships. Common sense entities and relationships refer to knowledge that is not explicitly presented in the input text but can be inferred from the common knowledge shared by humans in daily life. These common sense exercises are more challenging and motivate students, finding widespread application in industries such as smart education. Traditional literal-based mapping models lack a common sense understanding module, making them difficult to generate common sense exercises. Common sense understanding is a recognized fundamental challenge in artificial intelligence. Furthermore, applying these models requires extensive training resources, including large-scale common sense libraries and annotated training datasets. However, the cost of manually constructing these common sense libraries and annotations is high, making widespread application of this technology difficult.
[0003] Generating exercises tailored to the course content a user is studying is a labor-intensive and resource-intensive task. Different users learn different content, further increasing the cost of manual question generation. Because machine-generated questions can significantly reduce labor costs, this demand has led to the rapid development of automated question generation. Consequently, machine-generated questions have gradually become a hot topic in the fields of intelligent education and natural language processing. This technology requires the ability to generate coherent and relevant exercises based on a given text. This requires an understanding of the text's complex semantics and logical structure. Machine-generated questions can quickly generate adaptive questions relevant to the textbook content to assess student learning outcomes. Furthermore, this technology can support many applications, such as serving as a data augmentation strategy to alleviate the problem of scarce annotations in question-answering systems by generating a large number of samples.
[0004] Question generation is a cognitively demanding process, requiring varying degrees of comprehension. Simple questions typically only address the superficial meaning of the text and can be addressed through literal matching. However, matching falls far short of truly understanding complex semantics and fails to fully assess learners' understanding of the key points, including their ability to deeply reflect and reason across multiple text entities and relationships, as well as their ability to infer implicit common sense. This common sense includes encyclopedic concepts and knowledge of causal relationships, implicit in the text and reflecting human consensus and understanding of the objective world. In practice, literal memory questions are too simplistic to capture user interest. Common sense reasoning questions, on the other hand, are more likely to stimulate learning interest and more comprehensively assess a user's mastery of various knowledge areas. However, automatically generating these common sense questions is challenging. For example, as shown in Figure 1, the exercise asks about a type of clothing in a multiple-choice format. Unlike simple literal matching questions, the literal meaning of this question is not directly related to the answer. However, questions and answers can be logically linked by linking multiple lines of evidence within a given text, including "traditional local clothing," "mountain," "Mount Fuji," and additional commonsense relations, namely (Mount Fuji, Japan, part), (Japan, related to, Japanese), (Japanese, related to, kimono), (kimono, a type of, robe). Such multi-hop reasoning chains are crucial for both question direction and answering. However, research on these commonsense questions is quite weak. These questions not only tap into the discrete and combinatorial nature of language, but also require multi-hop reasoning capabilities based on complex context and commonsense.
[0005] Early generative approaches primarily generated questions by syntactically transforming input text using rules or templates. These rules or templates, being manually customized, were not only expensive to build but also lacked scalability. Currently, mainstream approaches employ data-driven neural network models. This model approaches the generative task as a sequence-to-sequence mapping problem similar to machine translation. Specifically, the output question is considered a sequence, and the input text is also considered a sequence. By learning the mapping pattern between input and output sequences from a large amount of training data, the input text is mapped into the question in a translation-like manner. However, this approach is generally only suitable for generating simple literal matching questions and struggles with commonsense reasoning questions that require comprehensive understanding. This is because commonsense reasoning questions are not about transforming or mapping the given content into a semantically equivalent form, but rather about generating questions subject to various grammatical and semantic constraints. Generated questions must not only be fluent but also answerable and reasonable. In other words, questions must be answerable based on the logical context of the text, and the answering process involves more than simple literal matching, but also commonsense reasoning. This solvability feedback and implicit commonsense knowledge provide an indispensable bridge for guiding the questioning process. Traditional methods lack the ability to model underlying common sense, ultimately resulting in only simple, superficial questions. Furthermore, learning such models requires significant resources. A large-scale common sense database is essential. However, input content may contain a large amount of non-standard textual expressions, such as internet jargon and colloquialisms. These expressions are difficult to capture with manually constructed, standardized knowledge graphs. For example, some informal online terms may not literally match synonymous concepts in the graph. Furthermore, the scale of currently constructed graphs is far smaller than the vast common sense knowledge in the real world, making it difficult to manually enumerate all common sense. Furthermore, as complexity increases, neural models require more training data. Their performance depends significantly on the availability of annotated data. This data-hungry nature makes them unsuitable for real-world applications. Furthermore, due to the high cost of annotation, real-world businesses may not always be able to provide sufficient common sense and annotation resources. Therefore, constructing common sense-based homework exercises in low-resource settings is a cutting-edge technical challenge with significant academic and commercial value. According to research, no patents currently address this topic.
[0006] Traditional approaches to address the low-resource, low-sample problem generally employ two approaches. The first is data transfer, which involves transferring abundant knowledge from a source domain to a scarce target domain. Recent research has shown that pre-trained language models (PLMs) can represent extensive, encyclopedic conceptual and causal knowledge from large corpora. This powerful representational capability helps models effectively encode the semantics and context of input text without requiring additional labeled data. Therefore, they can be viewed as a large-scale common sense repository. However, due to a lack of task-specific tuning, generalized pre-trained semantic models struggle to adapt to specialized domains. For example, they struggle to infer relevant, low-frequency relationships to support deductive reasoning in common sense tasks. Another technique that can address the low-sample scenario is generalization models. These models capture key features of questions from a small amount of labeled data and use a method of learning to predict unknown test cases. However, these features are often complex and difficult to discern in a sparse data space. Furthermore, these models are susceptible to misleading effects from unstable correlations and sporadic variance caused by insufficient labeled data, resulting in reduced robustness. Furthermore, exercises assessing the same knowledge point can be asked in a variety of ways, and the content of the questions must align with the course content and answers. This diversity of expression and logical consistency exceeds the representational capabilities of shallow language features. Because current neural network methods mostly learn a one-to-one mapping function from input to output based on sequence models, this type of inference-based questioning is difficult to achieve.
[0007] The current mainstream approaches in the academic field for machine-generated exams can be categorized into two types. The first uses a grammar or syntax analyzer to convert text into an intermediate form, such as a grammar or syntax tree. Templates or rules are then used to extract questions and answers from this intermediate form. Because templates and rules are manually designed and expensive to build and update, the models' scalability and coverage are limited. To address these issues, another approach uses sequence models to directly convert text into exam questions. This conversion relies on alignment between the text and the question learned from training data. This sequence model approach is detailed in the paper "X.Du and C.Cardie.2016.Neural Machine Translation by Jointly Learning to Align and Translate.In Journal of Computer Science." This model is completely data-driven and does not require the manual definition of numerous rules or templates. For example, a method for generating geography exam questions based on knowledge guidance proposed in an existing patent includes the following steps: S1: obtaining unstructured geography knowledge text corpus to construct a geography text corpus; S2: setting a syntactic template to identify corresponding logical sentences from the geography text corpus; S3: extracting geographical events from the logical sentences; S4: generalizing geographical events and constructing a structured geography knowledge graph based on the generalized geographical events; S5: constructing a graph knowledge-guided sequence model based on the structured geography knowledge graph; S6: generating geography exam questions based on the graph knowledge-guided sequence model. The present invention also provides a device for generating geography exam questions based on knowledge guidance, which is used to implement the method for generating geography exam questions based on knowledge guidance. This patent uses a sequence model, and the conversion process relies on the alignment relationship between the text and the question learned from the training data, but it is difficult to achieve this kind of analogical questioning.
[0008] Generating commonsense reasoning exercises first requires a large-scale commonsense database. However, manually constructed knowledge graphs are generally limited in size and struggle to cover the vast commonsense knowledge in the real world. Furthermore, question-generating models are typically based on deep neural networks, requiring large amounts of training data. However, this data is often manually annotated, which is very costly and makes large-scale training datasets difficult to obtain in real-world applications. However, the performance of neural models is strongly correlated with the size of the training dataset; smaller datasets are insufficient to train high-performance models. Traditional approaches to address this problem in low-resource scenarios can be summarized into two categories. The first is data transfer, which leverages external auxiliary resources through data augmentation or knowledge enhancement to better represent the semantics of the input. Unlike manually constructed graphs, recent research has found that large-scale pre-trained language models contain a richer and more extensive commonsense database. These models, such as BERT, GPT-2, and GPT-3, are widely used in academia. Another direction is generalization learning, which involves capturing key features to predict unknown examples and learn more robust models. By sampling the data distribution of key features, one-to-many generation can be supported. Typical current frameworks include variational autoencoders (VAEs) and generative adversarial learning (GANs). However, due to the non-differentiability of discrete samples, GANs are monotonic in the original data space, while VAEs can be uncontrollable without solvability feedback. Differently, this approach utilizes disentangled learning to build a robust generalization model that is adept at discovering latent question-asking factors in the data. Furthermore, an adversarial framework is employed to enhance commonsense answerability.
[0009] Existing technologies mainly focus on generating shallow questions, while neglecting the research on deep common sense questions that are of great value in practical applications. Therefore, there is a broad space for improvement in generating common sense reasoning exercises.
[0010] Summary of the Invention
[0011] The purpose of the present invention is to provide a method and device that can generate common sense reasoning type after-class exercises based on the textbook content studied by users in a low-resource scenario.
[0012] To achieve the above objectives, the present invention provides a method for generating common sense homework exercises in a low-resource scenario, comprising the following steps:
[0013] S1: Post-train the pre-trained language model using encyclopedic concepts and causal relationships in an external knowledge graph. The post-trained pre-trained language model contains a vocabulary, and when a word is input, a corresponding distributed vector representation is obtained;
[0014] S2: Input the sample into the pre-trained language model after post-training, and convert the input sample into a distributed vector representation through the pre-trained language model. The sample includes paragraph c, answer a and question y, and obtains the corresponding distributed vector e c 、e a , and e y , the distributed vector e of the problem y Input into the transformer model to obtain the encoding vector h of the problem y =Transformer(e y ), the distributed vector e of the paragraph c and the distributed vector e of the answer a Concatenate and input into the transformer model to obtain the paragraph-answer encoding vector h s =Transformer([e a ;e c ]), where [·;·] represents the vector concatenation operator;
[0015] S3: Encode the problem vector h y =Transformer(e y ) and the paragraph-answer encoding vector h s =Transformer([e a ;e c ]) is mapped into the latent space, from which the key question-asking factors are decoupled;
[0016] S4: By sampling the key question-asking factors and inputting the sampling results into a generator, the generator is a sequence model, which can generate exercises The generator consists of an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder uses the vector to generate the output sequence.
[0017] As a preferred solution, in step S1, when post-training the pre-trained language model, the template method is first used to convert the data in the form of triples in the graph into plain text sentences. Based on these sentences, the parameters in the pre-trained language model can be updated by minimizing the negative likelihood loss function, thereby injecting domain-related common sense knowledge into the model. The corresponding loss function is:
[0018] Among them, r t is the tth word in sentence r, and there are T words in the sentence.
[0019] As a preferred solution, in step S3, a latent variable z is introduced y To express the problem y, according to the information bottleneck principle, maximize (z y,y) mutual information, encouraging z y Focus on the key expression patterns in y; similarly, use z s To quantify the content of paragraph c and answer a that can reflect the user's query intention, and add divergence-based regularization to constrain variables, using the maximum mean difference (MMD) as the constraint indicator, and the decoupled variables are obtained by reverse learning training through reconstruction error, where the reconstruction error refers to the encoding (z y ,z s ) to what extent they can decode the actual exercise y.
[0020] As a preferred solution, when the encoder represents the input text as a vector, the prefix tuning technology of prompt learning is adopted to freeze the distributed representation parameters of the pre-trained model and only learn a few prompt parameters, reducing the scale of parameter adjustment while less affecting the performance; and in the decoder, the continuous prompt is designed to be M θ [i,:]=MLP θ ([M′θ[i,:];z y ;z s ]), where M′ θ is a learnable matrix, MLP(·) is a multi-layer feedforward neural network. Based on this prompt, exercises are generated word by word using the marginal probability of the following formula
[0021] Based on this probability, the softmax function is used to calculate the word with the highest probability in the predefined vocabulary, where Indicates the 1st to (t-1)th words of the generated results, Represents the tth word generated; Softmax is used as the last layer of the neural network, and each time the word with the highest probability in the vocabulary is output based on the output vector;
[0022] The objective function of the generator is:
[0023] The objective function of the generator is rewritten by factorization:
[0024] As a preferred solution, the classical variational inference mathematical solution method is used to solve the marginal probability objective function by maximizing the lower bound of evidence. Specifically,
[0025] Introducing variational posterior q φ (·) to approximate the true prior distribution p ψ (·), mathematically deriving the lower bound of the maximum evidence of the marginal probability objective function is:
[0026] By calculating the minimum reconstruction loss function on a small amount of labeled data At the same time, KL divergence is used to regularize the potential distribution to approximate the prior distribution p ψ (z), the maximum evidence lower bound ELBO can be decomposed into the following equation, where represents the labeled dataset used for training, and Represent the regularization loss functions related to the exercises and course content respectively;
[0027] To optimize the loss function related to the problem Assume that the prior distribution p ψ (·) and the posterior distribution q φ (·) obeys Gaussian distribution; under this assumption, the prior distribution p ψ (z y |y) can be distributed To calculate, where Represents the mean value. Different from the above distribution, its mean value can be compared with the language characteristics of question y by Φ(y) associated, where W y is the shared projection matrix, Φ(y) is the encoding feature function of question y; similarly, the posterior distribution By distribution To calculate, where By using the reparameterization technique, z y Calculated as μ y +σ y ⊙∈ y , where ∈ y It is from The resulting Gaussian noise, ⊙ is the element-wise product; the loss function based on KL divergence is calculated as:
[0028] Another loss function related to the question content such as input paragraph and answer The calculation method of is as follows: considering that the clues in these question contents may contain multiple inquiry topics, the input paragraph c and answer c are classified into k clusters in the encoding stage, where each centroid of the cluster corresponds to the prototype of a certain exercise inquiry topic. This can be achieved by the prior distribution p ψ (z s |s) regularization constraint is Gaussian mixture distribution of, and the posterior distribution q φ (z s |s) is regularized to Gaussian distribution is realized, where M k It is an indicator variable used to characterize the cluster corresponding to the kth Gaussian component. Average value is constrained to correspond to a cluster of the query prototype, that is, the mean can be calculated as W s s k , where s k is the centroid of the kth cluster; during training, the centroid remains unchanged and only the corresponding projection matrix W needs to be updated s ; By sampling all possible Gaussian components from the mixed distribution, the output is the diverse question results for the same knowledge point; based on the use of reparameterization techniques, the latent variable z s It can be calculated as μ s +σ s ⊙∈ s ; From this, through further mathematical deduction, the loss The solution is obtained by the soft EM algorithm based on formula (8), where Representation sample The probability of belonging to the kth question prototype can be calculated as τ is a temperature parameter that is usually set to 1, and dist(·) is the average value between the Gaussian distributions of the components and the latent variable z s The Euclidean distance between them is:
[0029] As a preferred solution, when establishing the generator of step S4, a discriminator based on adversarial learning is constructed as feedback to optimize the training of the generator; the discriminator is used to calculate the continuous probability value of common sense solvability. The discriminator is constructed with the help of a common sense question-answering model. The common sense question-answering model outputs a probability based on the input question, text, and answer, which is used to measure whether the input answer can answer the question. In addition, the question-answering model can also output the entities and relationships involved in the answering process. By using literal matching to determine whether these entities and relationships appear in the input text, it can be determined whether the answering process involves common sense outside the text. In this way, the probability value of common sense solvability can be obtained;
[0030] The loss function for each evaluation sample can be calculated as When the predicted result of the exercise does not match the input answer, or the solution process does not involve the implicit common sense of multi-hop deduction, the loss value is large. Based on such a discriminator, the generator is optimized according to the overall loss function, which is:
[0031] where γ g is the harmonic parameter.
[0032] As a preferred solution, when building a discriminator, given a small labeled dataset {y, a, c}, a supervised loss function is used. To train the discriminator, are the parameters of the discriminator; and the samples produced by the generator are used to expand the training data;
[0033] By combining manually annotated and machine-generated data, the discriminator can be trained based on the following loss function: Where H(·) is the result of generating Follow the estimated distribution q v The empirical Shannon entropy of β is β, and β is the parameter learned on the validation data set. This entropy term can regularize the discriminator to filter out noise and obtain better predictions. By using manually labeled and machine-generated samples, the discriminator can be optimized by the following formula, where γ v is the weighting factor.
[0034] As a preferred solution, based on the trained generator, we first select the prior distribution p ψ (z y |y) and p ψ (z s |s), and then these sample vectors are jointly fed into the encoder, through the formula The probability p in θ (·) to decode the questions word by word by As input, a transformer can be used to decode similar questions in different forms. The similarity measure is determined by the prior distribution p ψ (z y |y) and selects the top k similar predictions as the result.
[0035] The present invention also provides a device for generating common sense after-class exercises in a low-resource scenario, comprising:
[0036] Common sense enhanced representation unit: includes a post-trained pre-trained language model and a converter model; the pre-trained language model is post-trained using encyclopedia concepts and causal relationships in the external knowledge graph to obtain a post-trained pre-trained language model. The post-trained pre-trained language model converts the input sample into a distributed vector representation. The sample includes paragraph c, answer a, and question y, and obtains the corresponding distributed vector e c 、e a , and e y; The transformer model transforms the distributed vector e of the problem y Converted into the encoding vector h of the problem y =Transformer(e y ), and the distributed vector e of the spliced paragraphs c and the distributed vector e of the answer a Converted into paragraph-answer encoding vector h s =Transformer([e a ;e c ]);
[0037] Question generation unit: includes decoupler, sampler and generator; decoupler is used to encode the question vector h y =Transformer(e y ) and the paragraph-answer encoding vector h s =Transformer([e a ;e c ]) is mapped into the latent space, from which the key questioning factors are decoupled; the sampler is used to sample the decoupled key questioning factors and input the sampling results into the generator; the generator is used to generate exercises The generator is a sequence model. The generator includes an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder uses the vector to generate the output sequence.
[0038] As a preferred solution, the device further comprises:
[0039] Function optimizer, which uses the classic variational inference mathematical solution method to solve the marginal probability objective function of the generator by maximizing the lower bound of evidence based on the objective function of the generator and the marginal probability objective function;
[0040] The discriminator is used to calculate common sense-parseable continuous probability values to optimize the generator;
[0041] The question generation unit also includes a prediction module, which is used to decode the questions generated by the generator word by word based on the encoding vector of the question and paragraph-answer, and decode similar questions with different expressions. The similarity measure is determined by the prior distribution p ψ (z y |y), and the top k similar predictions are selected as the results.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The present invention provides a method for generating common sense after-class exercises in low-labeled resource scenarios, enabling machine-generated question generation. This method leverages the extensive knowledge contained in pre-trained language models to fully represent the context and implicit common sense of the input text. To avoid the problem of general pre-trained models forgetting specialized domain knowledge, the present invention uses encyclopedic concepts and causal relationships in an external knowledge graph to post-train the pre-trained language model to enhance its ability to express task-related domain common sense. Considering that the questioning process involves multiple factors and features, including the macroscopic test point intent and the microscopic verbal expression, the present invention projects samples into a dense latent space to facilitate learning these key features that control the questioning process. These features are interrelated but do not directly correspond to sample points in the original data space. Therefore, simply assuming they are independent oversimplifies the latent manifold, making it easy for the model to erroneously retain mixed noise, resulting in unsatisfactory results. To improve the model's combinatorial generalization ability, the present invention decouples key questioning factors. These factors can reflect the content of the test point intent and the diverse expression methods. Introducing these factors can enhance the robustness of the model and reduce the impact of misleading variance caused by insufficient training data. This invention can generate sufficient and diverse common sense exercises. It transfers the rich semantic knowledge from external pre-trained semantic models and learns key question patterns in a dense latent space with decoupled priors, effectively alleviating the problem of resource scarcity. This invention also provides a common sense after-school exercise generation device for low-resource scenarios, which has the aforementioned beneficial effects and will not be elaborated on here. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG1 is an example of a test question requiring common sense reasoning in background technology.
[0045] FIG2 is a flowchart of a method for generating common sense after-class exercises in a low-resource scenario according to an embodiment of the present invention.
[0046] FIG3 is a schematic structural diagram of a device for generating common sense after-class exercises in a low-resource scenario according to an embodiment of the present invention.
[0047] In the figure, 101 is the common sense enhanced representation unit; 102 is the question generation unit; 103 is the discriminator; 201 is the pre-trained language model; 202 is the transformer model; 203 is the decoupler; 204 is the generator; 205 is the function optimizer; and 206 is the prediction module. DETAILED DESCRIPTION
[0048] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0049] Example 1
[0050] As shown in FIG2 , a method for generating common sense homework exercises in a low-resource scenario according to a preferred embodiment of the present invention includes the following steps:
[0051] S1: Post-train the pre-trained language model using encyclopedic concepts and causal relationships in an external knowledge graph. The post-trained pre-trained language model (PLM) contains a vocabulary. When a word is input, a corresponding distributed vector representation is obtained.
[0052] S2: Input the sample into the pre-trained language model after post-training, and convert the input sample into a distributed vector representation through the pre-trained language model. The sample includes paragraph c, answer a and question y, and obtains the corresponding distributed vector e c 、e a , and e y , the distributed vector e of the problem y Input into the transformer model to obtain the encoding vector h of the problem y =Transformer(e y ), the distributed vector e of the paragraph c and the distributed vector e of the answer a Concatenate and input into the transformer model to obtain the paragraph-answer encoding vector h s =Transformer([e a ;e c ]), where [·;·] represents the vector concatenation operator;
[0053] S3: Encode the problem vector h y =Transformer(e y ) and the paragraph-answer encoding vector h s =Transformer([e a ;e c ]) is mapped into the latent space, from which the key question-asking factors are decoupled;
[0054] S4: By sampling the key question-asking factors and inputting the sampling results into a generator, the generator is a sequence model, which can generate exercises The generator consists of an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder uses the vector to generate the output sequence.
[0055] This embodiment leverages the extensive knowledge embedded in the pre-trained language model to fully represent the context and implicit common sense of the input text. To avoid the problem of general pre-trained models forgetting specialized domain knowledge, this embodiment post-trains the pre-trained language model using encyclopedic concepts and causal relationships from an external knowledge graph to enhance its ability to express task-related domain common sense. Considering that the questioning process involves multiple factors and features, including macro-level test point intent and micro-level verbal expression, this embodiment projects samples into a dense latent space to facilitate learning these key features that control the questioning process. These features are interrelated but do not directly correspond to sample points in the original data space. Therefore, simply assuming they are independent oversimplifies the latent manifold, making the model prone to incorrectly retaining mixed noise, leading to unsatisfactory results. To improve the model's combinatorial generalization ability, this embodiment decouples key questioning factors. These factors can reflect the content of the test point intent and the diverse expression methods. Introducing these factors enhances the robustness of the model and reduces the impact of misleading variance caused by insufficient training data. This embodiment can generate sufficient and diverse common sense exercises. It will transfer the rich semantic knowledge from the external pre-trained semantic model and learn key question patterns in the dense latent space of the decoupled prior, thereby effectively alleviating the problem of resource scarcity.
[0056] Example 2
[0057] The difference between this embodiment and the first embodiment is that, based on the first embodiment, each step is further explained.
[0058] This embodiment provides a method for generating common sense homework exercises in a low-resource scenario, including the following steps:
[0059] S1: Use encyclopedia concepts and causal relationships in the external knowledge graph to post-train the pre-trained language model. The post-trained PLM model contains a vocabulary, and when a word is input, a corresponding distributed vector representation can be obtained.
[0060] To generate common sense exercises, it is first necessary to fully represent the semantics of the input text and the common sense it implies. A simple source of common sense is the knowledge graph. For example, the ConceptNet graph contains a large number of encyclopedic concepts and various parent-child relationships between concepts, and the ATOMIC graph contains a large number of causal relationships. Given that the concept and relationship representations in the graph are discrete and non-normalized, it is difficult to align them with the normalized entity words in the graph. Entities at a single node lack sufficient disambiguation context to accurately calculate relevance for alignment, and due to the large scale of the graph, the computational complexity of integrating all contextually relevant nodes is very high. Furthermore, graphs are generally constructed manually and it is difficult to cover the vast common sense knowledge in the objective world. To address this problem, this embodiment utilizes a large language model (PLM) pre-trained from a large corpus. It can encode various useful contexts in the corpus into a continuous rather than discrete parameter space. In this space, the relevance between entity words can be easily calculated by calculating the cosine similarity of the vectors, thereby efficiently aligning the entities. Because the pre-trained model has a large representation space, it can fully cover all potential grammatical and semantic features. Commonly used pre-trained language models include BERT and GPT-2. In this embodiment, GPT-2 is used as the pre-trained language model. To enhance the common sense representation capabilities of these models in specialized domains, this embodiment performs post-training on the pre-trained models using task-specific knowledge graphs. These specialized graphs contain rich common sense concepts and relationships, facilitating deductive reasoning.
[0061] Considering that relationships in graphs are usually represented by special notations, such as the UsedFor symbol in the ConceptNet graph representing "used for" and the xIntent symbol in the ATOMIC graph representing "the intent of x is," these notations lack spaces and are difficult to match with ordinary language expressions.
[0062] Therefore, in step S1, when post-training the pre-trained language model, the template method is first used to convert the data in the form of triples in the graph into plain text sentences. Based on these sentences, the parameters in the pre-trained language model PLM can be updated by minimizing the negative likelihood loss function, thereby injecting domain-related common sense knowledge into the model. The corresponding loss function is formula (1),
[0063] Among them, r t is the tth word in sentence r, and there are T words in the sentence.
[0064] S2: Input the sample into the pre-trained language model after post-training, and convert the input sample into a distributed vector representation through the pre-trained language model. The sample includes paragraph c, answer a and question y, and obtains the corresponding distributed vector e c 、e a , and e y , the distributed vector e of the problem y Input into the transformer model to obtain the encoding vector h of the problem y =Transformer(e y ), the distributed vector e of the paragraph c and the distributed vector e of the answer a Concatenate and input into the transformer model to obtain the paragraph-answer encoding vector h s =Transformer([e a ;e c ]), where [·;·] represents the operator of vector concatenation.
[0065] The post-trained PLM model contains a vocabulary. When a word is input, a corresponding distributed vector representation can be obtained. Step S2 uses the model to convert the input sample into a distributed vector representation. For each sample including paragraph c, answer a and question y, this embodiment first uses the word recognizer of the pre-trained model GPT-2 to recognize each word, and then finds the corresponding distributed vectors of these words in GPT-2, i.e., e c , e a , and e y . GPT-2 is an open source pre-trained model library. Through this method, the rich common sense knowledge stored in the pre-trained language model can be transferred to this task, so as to better represent the semantics of the input content. In order to capture the rich context in the text, this method inputs these vectors into the transformer model (Transformer), which is good at learning context through the attention mechanism. After passing through the transformer model, the encoding h of the question can be obtained y =Transformer(e y ); Considering that the questioning process generally requires understanding the paragraph and the answer at the same time, this embodiment concatenates the two types of vectors c and a and inputs them into the converter model to obtain the corresponding encoding vector h s =Transformer([e a ;e c ]), where [·;·] represents the operator of vector concatenation.
[0066] S3: Encode the problem vector h y =Transformer(e y ) and the paragraph-answer encoding vector hs =Transformer([e a ;e c ]) is mapped into the latent space, from which the key question-asking factors are decoupled.
[0067] In order to find the key generating factors, this embodiment first maps the input samples into the latent space. Compared with the data space, the latent space allows the invariance of distributed transformations, so that samples with similar query content are adjacent. It is easier to find key question features in a group of similar samples than in isolated samples. Considering that these features are usually correlated, using these features directly can easily mislead the model into being irrelevant noise. In order to improve the robustness of the model, this embodiment will decouple the key question factors from it. Each factor only focuses on one type of feature, avoiding entanglement with other features, especially those that are not explicitly modeled. Specifically, in step S3, a latent variable z is introduced. y To express the problem y, according to the information bottleneck principle, maximize (z y ,y), which is equivalent to encouraging z y Focus on the key expression patterns in y; similarly, use z s To quantify the content in paragraph c and answer a that reflects the user's query intent. In order to enhance the controllability of the model, this embodiment constrains the variables by adding divergence-based regularization to ensure their independence. This embodiment uses the maximum mean difference (MMD) as the constraint indicator, and the decoupled variables are obtained by reverse learning training through reconstruction error, where the reconstruction error refers to the encoding (z y ,z s ) can decode the actual exercise y. For the specific learning process, see the generator's marginal probability objective function solution.
[0068] In summary, to leverage these key generative factors to achieve one-to-many generation through inference, this embodiment also introduces non-overlapping conditional prior distributions to regularize these factors. This includes using an isotropic Gaussian distribution to constrain the expression factor and a conditional mixture Gaussian distribution to act on the question content factor. In this conditional mixture Gaussian, each component Gaussian represents a set of question prototypes with clues. By sampling the distribution, this embodiment can integrate the necessary clues to form a deducible result.
[0069] S4: By sampling the key question-asking factors and inputting the sampling results into a generator, the generator is a sequence model, which can generate exercises The generator consists of an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder uses the vector to generate the output sequence.
[0070] In order to further reduce the model's dependence on labeled data, when the encoder represents the input text as a vector, the prefix tuning technology of prompt learning is adopted to freeze the distributed representation parameters of the pre-trained model and only learn a few prompt parameters, thereby reducing the scale of parameter adjustment while minimizing the impact on performance. θ [i,:]=MLP θ ([M′ θ [i,:]; z y ;z s ]), where M′ θ is a learnable matrix, MLP(·) is a multi-layer feedforward neural network, based on the prompt, the marginal probability of formula (2) is used to generate exercises word by word Formula (2) is the marginal probability function,
[0071] Based on this probability, the softmax function is used to calculate the word with the highest probability in the predefined vocabulary, where Indicates the 1st to (t-1)th words of the generated results, represents the tth word generated. The softmax function, also known as the logistic function, is a classic activation function in mathematical statistics. It normalizes a numerical vector into a probability distribution vector, where the sum of all probabilities is 1. Softmax is used as the final layer of a neural network. Each time, based on the output vector, it outputs the word with the highest probability in the vocabulary.
[0072] This embodiment adopts the prefix tuning technology in hint learning, which can reduce the scale of parameter adjustment while less affecting performance.
[0073] To avoid the problem of mode collapse, which means that the model degenerates into monotonicity and produces very simple results, this embodiment uses the prior distribution p ψ (·) to regularize the potential z y and z s This helps in the prediction phase by sampling p ψ (·) produces diverse results. In summary, the objective function of the generator in this embodiment is:
[0074] The objective function of the generator is rewritten by factorization:
[0075] It is difficult to directly calculate the marginal probability objective function in formula (4). This embodiment uses the classic variational reasoning mathematical solution method to solve the marginal probability objective function by maximizing the evidence lower bound (ELBO). Specifically,
[0076] Introducing variational posterior qψ (·) to approximate the true prior distribution p ψ (·), mathematically deriving the lower bound of the maximum evidence of the marginal probability objective function is:
[0077] By calculating the minimum reconstruction loss function on a small amount of labeled data At the same time, KL divergence is used to regularize the potential distribution to approximate the prior distribution p ψ (z), the maximum evidence lower bound ELBO can be decomposed into the following equations, which are formula (6), where represents the labeled dataset used for training, and Represent the regularization loss functions related to the exercises and course content respectively;
[0078] To optimize the loss function related to the problem Assume that the prior distribution p ψ (·) and the posterior distribution q φ (·) obeys Gaussian distribution; under this assumption, the prior distribution p ψ (z y |y) can be distributed To calculate, where Represents the mean value. Different from the above distribution, its mean value can be compared with the language characteristics of question y by Φ(y) associated, where W y is the shared projection matrix, Φ(y) is the encoding feature function of question y; similarly, the posterior distribution By distribution To calculate, where By using the reparameterization technique, z y Calculated as μ y +σ y ⊙∈ y , where ∈ y It is from The resulting Gaussian noise, ⊙, is the element-wise product. This technique is a classic method for finding optimization functions, used to separate the uncertainty of random variables, making it possible to differentiate intermediate nodes that were previously incapable of gradient propagation. Based on this technique, the loss function based on KL divergence can be calculated as:
[0079] Another loss function related to the question content such as input paragraph and answer The calculation method of is as follows: considering that the clues in these question contents may contain multiple inquiry topics, this embodiment classifies the input paragraph c and answer c into k clusters in the encoding stage, where each centroid of the cluster corresponds to the prototype of a certain exercise inquiry topic. This can be achieved by the prior distribution p ψ (z s |s) regularization constraint is Gaussian mixture distribution of, and the posterior distribution q φ (z s |s) is regularized to Gaussian distribution is realized, where M k It is an indicator variable used to characterize the cluster corresponding to the kth Gaussian component. Average value is constrained to correspond to a cluster of the query prototype, that is, the mean can be calculated as W s s k , where s k is the centroid of the kth cluster; during training, the centroid remains unchanged and only the corresponding projection matrix W needs to be updated s By sampling all possible Gaussian components from the mixture distribution, we can output diverse question results for the same knowledge point; based on the reparameterization technique, the latent variable z s It can be calculated as μ s +σ s ⊙∈ s ; From this, through further mathematical deduction, the loss The solution is obtained by the soft EM algorithm based on formula (8), where Representation sample The probability of belonging to the kth question prototype can be calculated as τ is a temperature parameter that is usually set to 1, and dist(·) is the average value between the Gaussian distributions of the components and the latent variable z s The Euclidean distance between them is:
[0080] A reasonable homework exercise should be able to deduce the answer from the context of the input content. In other words, the question should be paired with the input answer. Without the guidance of this additional feedback, it is difficult to produce reasonable results. Therefore, this embodiment constructs a discriminator as an indicator to provide solvability feedback. Considering that traditional discrete indicators cannot propagate gradients, it makes it difficult to optimize and train the model as a whole. Therefore, this embodiment designs a differentiable discriminator to calculate the continuous probability value that can be resolved by common sense. This discriminator is constructed by using the current most advanced common sense question answering model. This aspect refers to the performance ranking of the evaluation dataset and finds the best performing model UNICORN to implement the discriminator.
[0081] Specifically, when establishing the generator of step S4, a discriminator based on adversarial learning is constructed as feedback to optimize the training of the generator; the discriminator is used to calculate the continuous probability value of common sense solvability. The discriminator is constructed with the help of a common sense question-answering model. The common sense question-answering model outputs a probability based on the input question, text, and answer, which is used to measure whether the input answer can answer the question. In addition, the question-answering model can also output the entities and relationships involved in the answering process. By using literal matching to determine whether these entities and relationships appear in the input text, it can be determined whether the answering process involves common sense outside the text. In this way, the probability value of common sense solvability can be obtained;
[0082] The loss function for each evaluation sample can be calculated as When the predicted result of the exercise does not match the input answer, or the solution process does not involve the implicit common sense of multi-hop deduction, the loss value is large. Based on such a discriminator, the generator is optimized according to the overall loss function of formula (9), and the overall loss function is:
[0083] where γ g is the harmonic parameter.
[0084] In addition, when building the discriminator, given a small labeled dataset {y, a, c}, a supervised loss function is used to To train the discriminator, is the parameter of the discriminator; considering that the scale of manually labeled data may be small, resulting in the discriminator being insufficient to be fully trained, this method proposes to use the samples produced by the generator to expand the training data.
[0085] By combining manually annotated and machine-generated data, the discriminator can be trained based on the following loss function: Where H(·) is the result of generating Follow the estimated distribution q vThe empirical Shannon entropy of β is β, and β is the parameter learned on the validation data set. This entropy term can regularize the discriminator to filter out noise and obtain better predictions. By using manually labeled and machine-generated samples, the discriminator can be optimized by the following formula, where γ v is the weighting factor.
[0086] In summary, to prevent logical conflicts between generated results and answers, this example develops a differentiable discriminator to provide feedback on commonsense solvability. Through adversarial training, the generator reconstructs real problems to produce credible new results, while the discriminator collaboratively drives the generator to produce logically consistent results by reflectively providing feedback. This mutually reinforcing joint learning approach and the introduction of explicit decoupling constraints achieves this goal.
[0087] In order to solve the problem of scarce test case resources, this method generates multiple questions for each test case in different expressions. For a given article, these questions may correspond to the same answer. Based on the trained generator, first, we first extract the test case from the prior distribution p. ψ (z y |y) and p ψ (z s |s) and the encoding vectors of the sampled questions and contents, and then these sampled vectors are jointly fed to the encoder, and the encoding vectors are obtained by formula (2) The probability p in θ (·) to decode the questions word by word by As input, a transformer can be used to decode similar questions in different forms. The similarity measure is determined by the prior distribution p ψ (z y |y) and selects the top k similar predictions as the result.
[0088] To measure the performance of the model, this example conducted experiments using two currently popular datasets: the ConmosQA dataset and the MCScript dataset. These datasets contain 35,600 and 20,000 examples, respectively. These examples are multiple-choice questions that require common sense reasoning. This example uses three traditional metrics to measure the quality of the generated questions: BLEU-4, METEOR, and ROUGE-L. The experimental results show that this method significantly outperforms traditional methods.
[0089] Example 3
[0090] As shown in Figure 3, this embodiment provides a common sense after-class exercise generation device in a low-resource scenario based on the generation method of Example 2, including: a common sense enhanced representation unit 101 and a question generation unit 102. The question generation unit 102 is a generalized model for low-resource question generation. The common sense enhanced representation unit 101 includes a pre-trained language model 201 and a converter model 202 that have been post-trained. The question generation unit 102 includes a decoupler 203, a sampler, and a generator 204. The generator 204 includes an encoder and a decoder. The pre-trained language model 201, the converter model 202, the decoupler 203, the sampler, the encoder, and the decoder are connected in sequence.
[0091] Specifically, the common sense enhanced representation unit 101 includes a post-trained pre-trained language model 201 and a converter model 202; the pre-trained language model 201 is post-trained using the encyclopedia concepts and causal relationships in the external knowledge graph to obtain the post-trained pre-trained language model 201, and the post-trained pre-trained language model 201 converts the input sample into a distributed vector representation, where the sample includes paragraph c, answer a, and question y, and obtains the corresponding distributed vector e c 、e a , and e y ; The converter model 202 converts the distributed vector e of the problem y Converted into the encoding vector h of the problem y =Transformer(e y ), and the distributed vector e of the spliced paragraphs c and the distributed vector e of the answer a Converted into paragraph-answer encoding vector h s =Transformer([e a ;e c ]);
[0092] Question generation unit 102: includes a decoupler 203, a sampler and a generator 204; the decoupler 203 is used to convert the encoding vector h of the question y =Transformer(e y ) and the paragraph-answer encoding vector h s =Transformer([e a ;e c ]) is mapped into the latent space, from which the key questioning factors are decoupled; the sampler is used to sample the decoupled key questioning factors and input the sampling results into the generator 204; the generator 204 is used to generate exercises The generator 204 is a sequence model, which includes an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder generates an output sequence based on the vector.
[0093] In addition, the device also includes:
[0094] Function optimizer 205, configured to solve the marginal probability objective function of generator 204 by maximizing the lower bound of evidence based on the objective function and marginal probability objective function of generator 204 using a classical variational inference mathematical solution method;
[0095] Discriminator 103, used to calculate common sense-parseable continuous probability values to optimize generator 204;
[0096] The question generation unit 102 further includes a prediction module 206, which is used to decode the questions generated by the generator 204 word by word based on the encoding vectors of the question and the paragraph-answer, and decode similar questions in different expression forms. The similarity measure is determined by the prior distribution p ψ (z y |y), and the top k similar predictions are selected as the results.
[0097] In summary, embodiments of the present invention provide a method for generating commonsense homework exercises in low-resource scenarios. This method leverages the extensive knowledge contained in a pre-trained language model to fully represent the context and implicit commonsense of the input text content. To avoid the problem of general pre-trained models forgetting specialized domain knowledge, the present invention uses encyclopedic concepts and causal relationships in an external knowledge graph to post-train the pre-trained language model to enhance its ability to express task-related domain commonsense. Considering that the questioning process involves multiple factors and features, including the macro-level test point intent and the micro-level verbal expression, the present invention projects samples into a dense latent space to facilitate learning these key features that control the questioning process. These features are interrelated but do not directly correspond to sample points in the original data space. Therefore, simply assuming they are independent will oversimplify the latent manifold, making the model prone to erroneously retaining mixed noise, resulting in unsatisfactory results. To improve the model's combinatorial generalization ability, the present invention decouples key questioning factors. These factors can reflect the content of the test point intent and the diverse expression methods. Introducing these factors can enhance the robustness of the model and reduce the misleading variance caused by insufficient training data. The present invention can generate sufficient and diverse common sense exercises, which will transfer the rich semantic knowledge in the external pre-trained semantic model and learn key question patterns in the dense latent space of the decoupled prior, thereby effectively alleviating the problem of resource scarcity. In order to use these factors to achieve one-to-many generation of inferences, this aspect also introduces non-overlapping conditional prior distributions to regularize these factors, including using isotropic Gaussian distributions to constrain the expression method factors and using conditional mixed Gaussian distributions to act on the question content factors. In the conditional mixed Gaussian, each component Gaussian represents a set of question prototypes with clues. By sampling the distribution, this method can integrate the necessary clues to form a reasonable result. In order to reduce the dependence on labeled data, this method adopts the prefix tuning technique in hint learning, which can reduce the scale of parameter adjustment while less affecting performance. In addition, in order to prevent logical conflicts between the generated results and the answers, this method develops a differentiable verifier to provide feedback on common sense solvability. Through adversarial training, the generator reconstructs real problems to produce credible new results, while the verifier collaboratively drives the generator to produce logically consistent results by reflectively providing feedback. Through this mutually reinforcing joint learning method and the introduction of explicit decoupling constraints, this method can generate sufficient common sense homework exercises in low-resource scenarios. A large number of experimental results on two typical data sets have demonstrated the effectiveness of this method. The present invention can also support applications such as test question generation and intelligent teaching assistants, such as letting machines read textbooks to automatically generate exercises for students. These applications have huge commercial value. Moreover, this method also solves key bottleneck problems such as resource scarcity encountered in model deployment, and therefore has great practical value.Therefore, this method utilizes disentangled learning to build a robust generalization model that is good at discovering the underlying question factors in the data. In addition, an adversarial framework is used to improve the ability of common sense answerability.
[0098] In summary, this invention provides a large-scale knowledge base by transferring knowledge from pre-trained language models. By decoupling learning from a small amount of data, it captures key questioning factors to achieve generalized generation through analogy. Furthermore, it further reduces the scale of parameter adjustment through methods such as prompt learning. Furthermore, this invention constructs a verifier based on an adversarial framework to provide feedback on common sense solvability, thereby generating grammatically correct and logically consistent results. The advantages of the method proposed in this invention include:
[0099] (1) This method proposes a low-resource generation model with good generalization to generate sufficient and diverse common sense exercises. It can transfer the rich semantic knowledge from the external pre-trained semantic model and learn key question patterns in the dense latent space of the disentangled prior, thus effectively alleviating the problem of resource scarcity.
[0100] (2) This method develops an adversarial framework with a differentiable verifier to guide the generator to produce results that are both commonsense and logically consistent. Experiments on mainstream datasets demonstrate the effectiveness of this method.
[0101] The embodiment of the present invention provides a generation device based on the above generation method, which has the same technical effect and will not be described in detail here.
[0102] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention. These improvements and substitutions should also be regarded as the scope of protection of the present invention.
Claims
1. A method for generating common sense after-class exercises in a low-resource scenario, characterized in that: The steps include: S1: Post-train the pre-trained language model using encyclopedia concepts and causal relationships in the external knowledge graph. The post-trained pre-trained language model contains a vocabulary, and when a word is input, a corresponding distributed vector representation can be obtained; S2: Input the sample into the post-trained pre-trained language model, and convert the input sample into a distributed vector representation through the pre-trained language model. The sample includes paragraph c, answer a and question y, and obtains the corresponding distributed vector e c 、e a , and e y , the distributed vector e of the problem y Input into the transformer model to obtain the encoding vector h of the problem y = Transformer(e y ), the distributed vector e of the paragraph c and the distributed vector e of the answer a Concatenate them and input them into the transformer model to obtain the paragraph-answer encoding vector h s = Transformer([e a ;e c ]), where [·;·] represents the operator of vector concatenation; S3: The encoding vector h of the problem y = Transformer(e y ) and the paragraph-answer encoding vector h s = Transformer([e a ;e c ]) into the latent space, from which the key question-asking factors are decoupled; S4: By sampling the decoupled key question-asking factors and inputting the sampling results into a generator, the generator is a sequence model, and exercises can be generated The generator consists of an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder generates an output sequence based on the vector.
2. According to the method for generating common sense after-class exercises in a low-resource scenario in claim 1, it is characterized in that: In step S1, when the pre-trained language model is post-trained, the data in the form of triples in the graph is first converted into plain text sentences using the template method. Based on these sentences, the parameters in the pre-trained language model can be updated by minimizing the negative likelihood loss function, thereby injecting domain-related common sense knowledge into the model. The corresponding loss function is: Among them, r t is the tth word in sentence r, and there are T words in the sentence.
3. The method for generating common sense after-class exercises in a low-resource scenario according to claim 2, characterized in that: In step S3, a latent variable z is introduced y To express the problem y, according to the information bottleneck principle, maximize (z y ,y), encouraging z y Focus on the key expression patterns in y; similarly, use z s To quantify the content in paragraph c and answer a that can reflect the user's query intention, and add divergence-based regularization to constrain variables, the maximum mean difference (MMD) is used as the constraint indicator, and the decoupled variables are obtained by reverse learning training through reconstruction error, where the reconstruction error refers to the encoding (z y ,z s ) to what extent they can decode the actual exercise y.
4. The method for generating common sense after-class exercises in a low-resource scenario according to claim 3 is characterized in that: When the encoder represents the input text as a vector, it uses the prefix tuning technology of prompt learning to freeze the distributed representation parameters of the pre-trained model and only learn a few prompt parameters, thus reducing the scale of parameter tuning with less impact on performance. And design continuous prompts in the decoder as M θ [i,:]=MLP θ ([M′θ[i,:];z y ;z s ]), where M′θ is a learnable matrix and MLP(·) is a multi-layer feedforward neural network. Based on this prompt, exercises are generated word by word using the marginal probability of the following formula Based on this probability, the word with the highest probability in the predefined vocabulary is calculated through the softmax function, where Indicates the 1st to (t-1)th words of the generated results, Indicates the tth word generated; Softmax is used as the last layer of the neural network, and each time the word with the highest probability in the vocabulary is output according to the output vector; The objective function of the generator is: The objective function of the generator is rewritten by factorization:
5. The method for generating common sense after-class exercises in a low-resource scenario according to claim 4, characterized in that: Using the classical variational inference mathematical solution method, the marginal probability objective function is solved by maximizing the lower bound of evidence. Specifically, Introducing variational posterior q φ (·) to approximate the true prior distribution p ψ (·), mathematically derive the maximum evidence lower bound of the marginal probability objective function as: By calculating the minimized reconstruction loss function on a small amount of labeled data At the same time, the KL divergence is used to regularize the potential distribution to approximate the prior distribution p ψ (z), the maximum evidence lower bound ELBO can be decomposed into the following equation, where represents the labeled dataset used for training, and Represent the regularized loss functions related to the exercises and course content respectively; In order to optimize the loss function related to the exercise Assume that the prior distribution p ψ (·) and the posterior distribution q φ (·) obeys a Gaussian distribution; under this assumption, the prior distribution p ψ (z y |y) can be distributed To calculate, Represents the mean value. Different from the above distribution, its mean can be related to the linguistic features of question y by is associated with y is the shared projection matrix, Φ(y) is the encoding feature function of question y; similarlyGround, posterior distribution By distribution To calculate, By using the reparameterization technique, z y Calculated as μ y +σ y ⊙∈ y , where ∈ y is from The resulting Gaussian noise, ⊙ is the element-wise product; the loss function based on KL divergence is calculated as: Another loss function related to the question content such as input paragraph and answer The calculation method of is as follows: considering that the clues in these questions may contain multiple inquiry topics, the input paragraph c and answer c are classified into k clusters in the encoding stage, where each centroid of the cluster corresponds to the prototype of a certain exercise inquiry topic. This can be achieved by ψ (z s |s) The regularization constraint is Gaussian mixture distribution of , and the posterior distribution q φ (z s |s) is regularized to Gaussian distribution is realized, where M k is an indicator variable used to characterize the cluster corresponding to the kth Gaussian component. The average is constrained to correspond to a cluster of the query prototype, that is, the mean can be calculated as W s s k , where s k is the centroid of the kth cluster; during training, the centroid remains unchanged and only the corresponding projection matrix W needs to be updated s ; By sampling from all possible Gaussian components in the mixed distribution, the output is a variety of question results for the same knowledge point; based on the use of reparameterization techniques, the latent variable z s It can be calculated as μ s +σ s ⊙∈ s ; From this, through further mathematical derivation, the loss The soft EM algorithm is used to solve the problem based on formula (8), where Representation sample The probability of belonging to the kth question prototype can be calculated as τ is a temperature parameter that is usually set to 1, and dist(·) is the average value between the Gaussian distributions of the individual components and the latent variable z s The Euclidean distance between them gives:
6. The method for generating common sense after-class exercises in a low-resource scenario according to claim 5, characterized in that: When establishing the generator of step S4, a discriminator based on adversarial learning is constructed as feedback to optimize the training of the generator; The discriminator is used to calculate the continuous probability value of common sense resolvability. The discriminator is constructed with the help of the common sense question answering model. The common sense question answering model outputs a probability based on the input question, text and answer to measure whether the input answer can answer the question. In addition, the question answering model can also output the entities and relationships involved in the answering process. By using literal matching to determine whether these entities and relationships appear in the input text, it can be determined whether the answering process involves common sense outside the text. In this way, the probability value of common sense resolvability can be obtained; The loss function for each evaluation sample can be calculated as When the predicted result of the exercise does not match the input answer, or the solution process does not involve the implicit common sense of multi-hop deduction, the loss value is large. Based on such a discriminator, the generator is optimized according to the overall loss function, which is: where γ g is the harmonic parameter.
7. The method for generating common sense after-class exercises in a low-resource scenario according to claim 6, characterized in that: When building the discriminator, given a small labeled dataset {y, a, c}, a supervised loss function is used To train the discriminator, are the parameters of the discriminator; and the samples produced by the generator are used to expand the training data; Combining manually annotated and machine-generated data, the discriminator can be trained based on the following loss function: Where H(·) is the result of generating The estimated distribution q v The empirical Shannon entropy of ; and β is the parameter learned on the validation data set; this entropy term can regularize the discriminator to filter out noise and obtain better predictions. By using manually annotated and machine-generated samples, the discriminator can be optimized by the following formula, where γ v is the weight factor.
8. The method for generating common sense after-class exercises in a low-resource scenario according to claim 6, characterized in that: Based on the trained generator, first, we first select the prior distribution p ψ (z y |y) and p ψ (z s |s), and then these sampled vectors are jointly fed to the encoder, through the formula The probability p in θ (·) to decode the question word by word by As input, a transformer can be used to decode similar questions in different forms. The similarity measure is given by the prior distribution p ψ (z y |y), and selects the top k similar predictions as the result.
9. A device for generating common sense after-class exercises in a low-resource scenario, characterized in that: include: Common sense enhanced representation unit: including pre-trained language model and transformer model after post-training; The pre-trained language model is post-trained using the encyclopedia concepts and causal relationships in the external knowledge graph to obtain a post-trained pre-trained language model. The post-trained pre-trained language model converts the input samples into distributed vector representations. The samples include paragraphs c, answers a, and questions y, and obtain the corresponding distributed vectors e. c 、e a , and e y ; The transformer model transforms the distributed vector e of the problem y Converted into the encoding vector h of the problem y = Transformer(e y ), and the distributed vector e of the concatenated paragraphs c and the distributed vector e of the answer a Converted into the paragraph-answer encoding vector h s = Transformer([e a ;e c ]); Question generation unit: includes decoupler, sampler and generator; decoupler is used to encode the question vector h y = Transformer(e y ) and the paragraph-answer encoding vector h s = Transformer([e a ;e c ]) is mapped into the latent space, from which the key questioning factors are decoupled; the sampler is used to sample the decoupled key questioning factors and input the sampling results into the generator; Generator is used to generate exercises The generator is a sequence model. The generator includes an encoder and a decoder. The encoder is used to represent the input text as a vector, and the decoder generates an output sequence based on the vector.
10. The device for generating common sense after-class exercises in a low-resource scenario according to claim 9, characterized in that: Also includes: Function optimizer, which uses the classical variational inference mathematical solution method to solve the marginal probability objective function of the generator by maximizing the lower bound of evidence according to the objective function of the generator and the marginal probability objective function; The discriminator is used to calculate common sense-parseable continuous probability values to optimize the generator; The question generation unit also includes a prediction module, which is used to decode the questions generated by the generator word by word based on the encoding vector of the question and paragraph-answer by using the marginal probability of the generator to decode similar questions in different expressions. The similarity measure is given by the prior distribution p ψ (z y |y), and the top k similar predictions are selected as the results.