A knowledge-guided method and device for generating geography exam questions

By extracting and generalizing geographical events from unstructured geographical knowledge text corpus, building a geographical knowledge graph, and using a sequence model guided by graph knowledge, the problem of difficulty in generating high-quality geography test questions in the existing technology is solved, and a higher quality and diverse test question generation is achieved.

CN115455167BActive Publication Date: 2025-06-17SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211175334.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2025-06-17
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

The existing machine automatic question setting technology is difficult to generate high-quality geography test questions, especially the lack of effective methods when dealing with humanities and science knowledge, resulting in insufficient coverage and scalability of the generated test questions, which cannot meet the diversity needs of geography tests.

Method used

By obtaining unstructured geography knowledge text corpus, setting up syntax templates to identify factual sentences, extracting geographical events and generalizing them, building a structured geography knowledge graph, and generating geography test questions based on a sequence model guided by graph knowledge.

Benefits of technology

The model's knowledge reasoning ability and machine questioning ability can be improved, and high-quality geography test questions that are closer to reality can be generated to meet the diversity needs of geography tests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115455167B_ABST
    Figure CN115455167B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating geography test questions based on knowledge guidance, comprising the following steps: S1: obtaining unstructured geography knowledge text corpus to construct a geography text corpus; S2: setting a syntactic template, identifying corresponding factual sentences from the geography text corpus; S3: extracting geography events from factual sentences; S4: generalizing geography events, and constructing a structured geography knowledge graph based on the generalized geography events; S5: constructing a graph knowledge-guided sequence model based on the structured geography knowledge graph; S6: generating geography test questions based on the graph knowledge-guided sequence model. The present invention also provides a geography test question generation device based on knowledge guidance, which is used to implement the geography test question generation method based on knowledge guidance. The present invention provides a method and device for generating geography test questions based on knowledge guidance, which solves the problem that the current machine automatic question setting technology can only generate simple geography test questions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic examination question generation, and more specifically, to a method and device for generating geography examination questions based on knowledge guidance. Background Art

[0002] The essential knowledge for geography examinations is the basic knowledge system of the geography discipline, which consists of basic facts, basic concepts, basic logics, and basic qualities of the geography discipline. People generally measure the examinees' mastery of geography knowledge through examinations. High-quality examination questions generally not only assess the literal matching and memorization ability of knowledge points, but also can measure the comprehensive application of students' basic knowledge and basic principles by students. Examinations aim to enable students to gradually understand the characteristics of different natural and human things starting from several basic facts and form a good cognitive structure. However, proposition is a very laborious and resource-consuming task, and the subjectivity of manual proposition is relatively large. In examinations with a wide influence such as the college entrance examination, people need a low-cost and more objective proposition method. Therefore, it has promoted the rapid development of machine automatic proposition.

[0003] For machine automatic proposition, traditional atlas construction models are good at dealing with some factual knowledge, such as facts about the length of rivers and the order of celestial bodies in physical geography, but it is difficult to deal with more ambiguous human affairs knowledge. In geography college entrance examination questions, the proportion of examination questions related to human affairs is very large, and traditional methods are difficult to meet the generation requirements of such examination questions. In addition, due to the traditional template-based proposition method, the templates are manually designed, resulting in weak coverage and scalability of the model; moreover, the question types are relatively old-fashioned and cannot meet the requirements of the diversity of examination questions. On the other hand, neural models based on sequence models are prone to generating simple questions that do not require reasoning or semantically irrelevant examination questions due to the lack of understanding of high-order associated knowledge. This method is mainly used to generate simple literal comprehension questions, but simple questions are difficult to comprehensively evaluate students' knowledge structure, stimulate their self-learning ability, and are even less conducive to cultivating students' divergent thinking and logical reasoning abilities.

[0004] Therefore, the current machine automatic proposition technology can only generate simple geography examination questions and is difficult to meet the generation requirements of geography examination questions. Summary of the Invention

[0005] The present invention aims to overcome the technical defect that the current machine automatic proposition technology can only generate simple geography examination questions, and provides a method and device for generating geography examination questions based on knowledge guidance.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] A method for generating geography examination questions based on knowledge guidance includes the following steps:

[0008] S1: Obtain unstructured geographic knowledge text corpus to build a geographic text corpus;

[0009] S2: Setting a syntactic template and identifying corresponding factual sentences from the geographic text corpus according to the syntactic template;

[0010] S3: Extract geographical events from logical sentences based on dependency parsing and semantic role labeling;

[0011] S4: Generalize geographic events to obtain generalized geographic events, and construct a structured geographic knowledge graph based on the generalized geographic events;

[0012] S5: Construct a graph knowledge-guided sequence model based on the structured geographic knowledge graph;

[0013] S6: Generate geography test questions based on a sequential model guided by graph knowledge.

[0014] In the above scheme, geographical events are extracted from unstructured data and generalized, a geographical knowledge graph is constructed based on the generalized geographical events, and a sequence model guided by graph knowledge is further constructed, which shows the hierarchical relationship between knowledge points, improves the knowledge reasoning ability of the model, enhances the machine's questioning ability, and can generate high-quality geographical test questions that are closer to reality.

[0015] Preferably, the syntactic template comprises:

[0016] Cause-to-effect front-end syntax template:

[0017] <Conj...> {Cause},{Effect};

[0018] The centered syntax template from cause to effect:

[0019] {Cause} <verb>{Effect};

[0020] Cause-to-effect centered and supporting syntactic template:

[0021] <conj>{Cause} <conj verb>{Effect};

[0022] Cause-to-effect front-end matching syntactic template:

[0023] <conj>{Cause} <verb>,{Effect};

[0024] Causal-tracing middle-position syntactic template:

[0025] {Effect}<Conj...>{Cause};

[0026] Causal-tracing matching syntactic template:

[0027] <Conj...>{Effect}<Conj...>{Cause};

[0028] Cause-to-effect middle-position three-layer causal relationship syntactic template:

[0029] <Conj...>{Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0030] Syntactic template of a three-layer causal relationship guided by a cause-and-effect verb:

[0031] {Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0032] A three - layer causal relationship syntactic template led by a front - end supporting verb from cause to effect:

[0033] <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0034] Syntactic template for guiding the two-layer causal relationship formula in the middle from cause to effect:

[0035] <conj>{Cause1} <conj verb>,{Effect1 / cause2}, <verb>{effect2};

[0036] The syntactic template of the two-layer causal relationship formula led by the front-end supporting verb from cause to effect:

[0037] <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb>{effect2};

[0038] Syntactic template of two-layer causal relationship guided by a middle verb from cause to effect:

[0039] {Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause2}.

[0040] Preferably, the specific steps of step S3 are:

[0041] For any sentence, the trigger words of the geographical event in the sentence are first determined according to the syntactic template, and the participants of the geographical event in the sentence are identified through semantic role labeling. Then, the subject-predicate-object structure of the geographical event in the sentence is identified through dependency syntactic analysis, so as to extract the geographical event in the sentence.

[0042] Among them, it is determined whether there is an agent A0 of the action or a recipient A1 of the action in the result of semantic role labeling;

[0043] If A0 exists, the subject of the geographical event is represented by A0, otherwise the dependent child node of the dependency syntactic structure SBV is used as the subject; if the dependency syntactic structure SBV is also missing, the geographical event is represented as a verb-object structure;

[0044] If A1 exists, the object of the geographical event is represented by A1, otherwise the dependent child node of the dependency syntactic structure VOB is used as the object; if the dependency syntactic structure VOB is also missing, the geographical event is represented as a subject-predicate structure;

[0045] If no verb appears, a noun is used to represent the geographical event.

[0046] Preferably, the geographical events are generalized by the following steps:

[0047] S4.1: Use the most frequently occurring syntactic combinations in the geographic text corpus to abstract geographic events and obtain abstract geographic events;

[0048] S4.2: Calculate the cosine similarity of abstract geographic events;

[0049] S4.3: Generalize abstract geographic events based on cosine similarity:

[0050] If an abstract geographical event E has at least five similar abstract geographical events, the common components of the abstract geographical event E and its similar abstract geographical events are extracted as the generalized geographical event;

[0051] Otherwise, it is considered that the abstract geographical event E lacks generality and is not generalized;

[0052] If the similarity between two abstract geographic events is greater than a preset similarity threshold, they are similar abstract geographic events.

[0053] Preferably, the graph knowledge-guided sequence model includes a graph encoder, a K-BERT text encoder, and a decoder. First, the geographical knowledge graph obtained in step S4 is input into the graph encoder to grasp the structural context information in the geographical knowledge graph, and any geographical knowledge text corpus is input into the K-BERT text encoder to capture the context information of the text. Then, the decoder is used to generate corresponding geographical exam questions.

[0054] Preferably, the graph encoder includes a graph preprocessing unit and a graph transformation unit.

[0055] In the graph preprocessing unit, the TransH method is used to represent the high-dimensional predicates and entities of the geographical knowledge graph as low-dimensional matrices P and E. By training matrices P and E, the total distance of all facts (s, r, o) is minimized; where s represents the head entity, r represents the relationship, o represents the tail entity, e s represents the semantic vector of the object entity, p r represents the semantic vector of the predicate, and e o represents the semantic vector of the subject entity.

[0056] In the graph transformation unit, the following formula is used to calculate and capture semantic information:

[0057]

[0058] where represents the encoded information after splicing with attention, e i represents the original input encoded information, | represents the concatenation operation of N attn, j ∈ N, and attn j is the dot product calculation.

[0059]

[0060]

[0061] q i 、k i 、v i are the d k -dimensional vector representations after the i-th stacked block linearly transforms the input;

[0062]

[0063]

[0064] FF(·) is a two-layer feed-forward network, and LN is a normalization layer;

[0065] e1 = Concat(e s ; p r ; e o )

[0066] Perform layer normalization on the final output result e N :

[0067] e N = LN output (e N ).

[0068] Preferably, the K-BERT text encoder includes a knowledge layer, an embedding layer, a noise filtering layer, and a mask learning layer;

[0069] In the knowledge layer, given an input sentence s = [w0, w1, w2,..., w n and a knowledge graph,

[0070] Select all entity names involved in the sentence s through the following formula, and query their corresponding triples from the graph:

[0071] E = K-Query(s, KG)

[0072] E = [(w i , ri0, w i0 ),..., (w i , r ik , w ik )]

[0073] where E represents the set of triples, the K-Query() function is a formulaic representation of knowledge query, KG represents the knowledge graph, and (w i , r ik , w ik ) represents the corresponding triple queried;

[0074] Associate the triples in E to the appropriate positions in the entity relationship graph t through the K-Inject function:

[0075] t = K-Inject(s, E);

[0076] In the embedding layer, use the vocabulary provided by Google-BERT. Through a trainable lookup table, convert each token in the entity relationship graph into an embedding vector of dimension H. Then use [CLS] as the classification token and [MASK] as the masking token. Then, through soft position embedding, sort the word vectors and add the structural information lost by the input sentence. Finally, identify different sentences containing multiple sentences through segment embedding;

[0077] In the noise filtering layer, use a visibility matrix to limit the visible area of each vector:

[0078]

[0079] Among them, M ij represents the visualization matrix, represents the word w j and the word w i are on the same branch;

[0080] There are multiple stacks of Mask-self-attention in the mask learning layer, and the Mask-self-attention is:

[0081] Q i+ 1 , K i+1 , V i+1 = h i W q , h i W k , h i W v

[0082]

[0083] h i+1 = S i+1 V i+1

[0084] Among them, Q i+1 is the information to be queried, K i+1 is the vector to be queried, V i+1 is the value queried, h i is the hidden state of the i-th masked self-attention block, S i+1 is the attention score, M is the visible matrix calculated by the visual layer, d k is the scaling factor, W q , W k and W v are trainable model parameters, h i+1 is the hidden state of the (i + 1)-th masked self-attention block.

[0085] Preferably, the decoder uses the GPT-2 language model to decode and generate geography exam questions; the decoder is stacked by multiple decoding modules, and each decoding module includes a position encoding layer, a multi-head attention mechanism layer, and a batch normalization layer;

[0086] After obtaining the output vector h of the K-BERT text encoder and the output vector En of the atlas encoder, it further includes concatenating h and En to obtain the vector z = [h:En], and then inputting z = [h:En] into the multi-head attention mechanism layer as the output of the encoding layer;

[0087] Generate high-quality geography exam questions by sampling the conditional probability p(T|S):

[0088]

[0089] where T represents the output target sequence T = x m+1 ,…,x N , S represents the already generated sequence S = x1,…,x m , N represents the total length of the target output sequence, and p(x n |x1,...,x n-1 ) represents the conditional probability distribution of predicting the next word based on the already generated sequence.

[0090] Preferably, it further includes generating interference items by combining the negative sampling mechanism CTRL. The specific steps are as follows:

[0091] CTRL learns p(x i |x< i , c) by inputting a text sequence with condition c:

[0092]

[0093] Decompose it using the chain rule and inject it into the loss function for training. Use the answer as the condition to adjust the generated interference factor D:

[0094]

[0095] Represent the obtained n sequences with d-dimensional vectors, that is, obtain the matrix And inject it into the multi-head attention mechanism layer and the batch normalization layer, and finally select the best three interference items through scoring:

[0096] Scores(X0) = LayerNorm(x t )W vocab

[0097] where Scores(X0) is the score of the generated interference item, LayerNorm(X t ) is the normalization layer, and W vocab is the vocabulary weight matrix.

[0098] A geography exam question generation device based on knowledge guidance, used to implement the described geography exam question generation method based on knowledge guidance, including a data construction module, a geography knowledge graph construction module, a question generation module, and an interference item generation module;

[0099] The data construction module is used to obtain unstructured geography knowledge text corpora to construct a geography text corpus;

[0100] The geographical knowledge graph construction module is used to construct a structured geographical knowledge graph, which includes:

[0101] A sentence recognition module, which is used to recognize corresponding event sentences from the geographical text corpus according to the syntactic template;

[0102] An event extraction module, which is used to extract geographical events from the event sentences;

[0103] A knowledge generalization module, which is used to generalize geographical events;

[0104] The question generation module is used to generate geographical exam questions according to the graph knowledge-guided sequence model;

[0105] The distracter generation module is used to generate distracters by combining the negative sampling mechanism CTRL.

[0106] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0107] The present invention provides a method and device for generating geographical exam questions based on knowledge guidance, which extracts geographical events from unstructured data and generalizes them, constructs a geographical knowledge graph according to the generalized geographical events, and further constructs a graph knowledge-guided sequence model, showing the hierarchical relationship of knowledge points, improving the knowledge reasoning ability of the model, enhancing the machine question-asking ability, and being able to generate more practical and high-quality geographical exam questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] Figure 1 It is a flowchart of the implementation steps of the technical solution of the present invention;

[0109] Figure 2 It is a schematic structural diagram of the graph conversion unit in the present invention;

[0110] Figure 3 It is a schematic diagram of the data processing process of the K-BERT text encoder in the present invention;

[0111] Figure 4 It is a schematic diagram of the module connection in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0112] The drawings are only for illustrative purposes and cannot be construed as a limitation of this patent;

[0113] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;

[0114] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0115] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0116] Example 1

[0117] like Figure 1 As shown, a method for generating geography test questions based on knowledge guidance includes the following steps:

[0118] S1: Obtain unstructured geographic knowledge text corpus to build a geographic text corpus;

[0119] S2: Setting a syntactic template and identifying corresponding factual sentences from the geographic text corpus according to the syntactic template;

[0120] S3: Extract geographical events from logical sentences based on dependency parsing and semantic role labeling;

[0121] S4: Generalize geographic events to obtain generalized geographic events, and construct a structured geographic knowledge graph based on the generalized geographic events;

[0122] S5: Construct a graph knowledge-guided sequence model based on the structured geographic knowledge graph;

[0123] S6: Generate geography test questions based on a sequential model guided by graph knowledge.

[0124] During the specific implementation process, geographical events are extracted from unstructured data and generalized, a geographical knowledge graph is constructed based on the generalized geographical events, and a sequence model guided by graph knowledge is further constructed to show the hierarchical correlation of knowledge points, improve the model's knowledge reasoning ability, enhance the machine's questioning ability, and generate high-quality geographical test questions that are closer to reality.

[0125] Example 2

[0126] A method for generating geography test questions based on knowledge guidance comprises the following steps:

[0127] S1: Obtain unstructured geographic knowledge text corpus to build a geographic text corpus;

[0128] S2: Setting a syntactic template and identifying corresponding factual sentences from the geographic text corpus according to the syntactic template;

[0129] More specifically, the syntax template includes:

[0130] (1) Cause-to-effect front-end syntax template:

[0131] <Conj...> {Cause},{Effect};

[0132] The causal clue word is a conjunction that appears at the beginning of the causal sentence. The causal sentence it leads contains obvious causal relationships. The regular expression is used to construct rules to match and identify the cause part and the result part. The clue words are "因[為]、由于".

[0133] (2) Cause-to-effect centered syntax template:

[0134] {Cause} <verb>{Effect};

[0135] Generally, causal sentences are led by causal cue verbs. These verbs indicate the cause in front and the result in the back. Causal sentences can be directly identified by constructing regular expressions for matching. The clue words are verbs such as "lead, cause, trigger, cause, induce, cause". In addition, considering that the clue words can be two conjunctions, a conjunction and a verb, or two verbs. In the text of the geographical field, the latter leads to more relations between things and principles. Since such sentences have obvious relations between things and principles, regular expressions can be directly constructed to match and identify sentences related to things and principles.

[0136] (3) The syntactic template from cause to effect:

[0137] <conj>{Cause} <conj verb>{Effect};

[0138] The causal relationship is formed by the following causal clue word pairs. The first word of each word pair leads to the cause, and the second clue word leads to the result. The position of the causal clue word is generally in the middle of the clause, and a regular expression is constructed to identify the causal sentence. The clue word pairs are "<because, cause>, <because, cause>, <because, lead>, <because, and>, <due to, trigger>, <because, make>, <due to, so that>, <due to, thus>, <by, cause>, <due to, cause>" and other conjunction pairs.

[0139] (4) Cause-to-effect front-end matching syntax template:

[0140] <conj>{Cause} <verb>,{Effect};

[0141] The causal relationship is represented by a pair of supporting causal cue words, which appears at the beginning of a causal sentence and is separated by a comma. The part between the two causal cue words is the cause part, and the part after the causal cue word is the result part. Regular expressions are used to construct rules to match and identify the cause part and the result part. The cue word pair is "<subject, affected>".

[0142] (5) The middle-position syntactic template for inferring cause from effect:

[0143] {Effect}<Conj...>{Cause};

[0144] The causal cue word is located in the middle of the sentence, leading the result in front and the cause behind. Causal sentences can be identified by constructing regular expressions. The cue words are "because" and "due to".

[0145] (6) The supporting syntactic template for inferring cause from effect:

[0146] <Conj...>{Effect}<Conj...>{Cause};

[0147] The causal relationship is guided by a pair of causal cue words. The part between the two causal cue words represents the result, and the part after the latter causal cue word leads to the cause. Regular expression rules are constructed for the identification and extraction of the cause part and the result part. The cue words are gerund combinations such as "<caused by, reason>", "<led to, reason>", etc.

[0148] (7) The middle-position supporting three-layer causal relationship syntactic template for inferring effect from cause:

[0149] <Conj...>{Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0150] By constructing the following causal syntactic pattern, directly perform regular matching to obtain each cause part and result part; the clue words are a combination of conjunction-verb pairs and causal cue verbs.

[0151] (8) The three-layer causal relationship syntactic template guided by cause-to-effect verbs:

[0152] {Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0153] The three - layer causal relationship is guided by a ternary causal cue verb and can be regarded as a nested combination of the syntactic template (2). Each cause part and result part can be directly recognized by constructing a regular expression match; the cue word is the ternary causal cue verb.

[0154] (9) The syntactic template of the three - layer causal relationship formula guided by the front - end supporting verb from cause to effect:

[0155] <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb>{effect2 / cause3}, <verb>{effect3};

[0156] It is composed of a cause-and-effect cue pair and a cause-and-effect cue verb. The combination of syntactic template (4) and syntactic template (2) can identify causal sentences by matching regular expressions; the cue words are the cue words of syntactic template (4) and the combination of cause-and-effect cue verbs.

[0157] (10) The two-layer causal relationship formula syntactic template centered from cause to effect:

[0158] <conj>{Cause1} <conj verb>,{Effect1 / cause2}, <verb>{effect2};

[0159] Guided by the supporting conjunctions and causal cue verbs, a two-layer causal relationship is formed, and the causal sentences are recognized by directly constructing a regular expression through the syntactic template; the cue words are the pairs of supporting conjunctions plus all causal cue verbs.

[0160] (11) The syntactic template of the two-layer causal relationship formula guided by the front-end supporting verb from cause to effect:

[0161] <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb>{effect2};

[0162] Two-layer causal relationships are guided by supporting verbs and causal cue verbs, and causal sentences are recognized by directly constructing regular expressions through syntactic templates; the clue words are the affected plus all causal cue verbs.

[0163] (12) Syntactic template for two-layer causal relationship guided by the middle verb from cause to effect:

[0164] {Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause2}

[0165] The two-layer causal relationship is guided by two causal cue verbs, and the causal sentence clue words are directly identified and extracted as binary causal cue verbs by matching syntactic templates.

[0166] S3: Extract geographical events from logical sentences based on dependency parsing and semantic role labeling;

[0167] The dependency syntax analyzer is used to identify syntactic components such as "subject, predicate, object", "attributive, adverbial, complement" in relevant sentences, and obtain a suitable representation method for geographical events. Common types of dependency syntax analysis include subject-predicate relationship (SBV), verb-object relationship (VOB), attributive-mediate relationship (ATT), verb-complement structure (CMP), parallel relationship (COO), etc.; the semantic role tagger is used to identify the relationship between the predicate and each component in the sentence, including core semantic roles (such as agent A0, patient A1, etc.) and subsidiary semantic roles (such as time TMP, location LOC, etc.).

[0168] More specifically, the specific steps of step S3 are:

[0169] For any logical sentence, the trigger word of the geographical event in the logical sentence is first determined according to the syntactic template (in this embodiment, the verb closest to the clue word is used as the trigger word of the geographical event), and the participants of the geographical event in the logical sentence are identified through semantic role labeling, and then the subject-predicate-object structure of the geographical event in the logical sentence is identified through dependency syntactic analysis, so as to extract the geographical event of the logical sentence;

[0170] For example, in the sentence "The mountain torrents raged, the river water surged, causing serious disasters", "surge" is the trigger word of the geographical event, and "river water" is the agent of the action, which can also be called a participant;

[0171] Among them, it is determined whether there is an agent A0 of the action or a recipient A1 of the action in the result of semantic role labeling;

[0172] If A0 exists, the subject of the geographical event is represented by A0, otherwise the dependent child node of the dependency syntactic structure SBV is used as the subject; if the dependency syntactic structure SBV is also missing, the geographical event is represented as a verb-object structure;

[0173] If A1 exists, the object of the geographical event is represented by A1, otherwise the dependent child node of the dependency syntactic structure VOB is used as the object; if the dependency syntactic structure VOB is also missing, the geographical event is represented as a subject-predicate structure;

[0174] If no verb appears, a noun is used to represent the geographical event.

[0175] S4: Generalize the geographical event to obtain the generalized geographical event, and construct a structured geographical knowledge graph based on the generalized geographical event;

[0176] More specifically, the geographical event is generalized through the following steps:

[0177] S4.1: Abstract the geographical event using the syntactic combination with the highest frequency of occurrence in the geographical text corpus to obtain an abstract geographical event;

[0178] The syntactic combination is the subject-predicate combination, the predicate-object combination, or other combinations of verbs or nouns; in this embodiment, the verbs are abstractly replaced with the verb categories in the causal prompt dictionary, and the nouns are replaced with the high-frequency synonyms in the causal prompt dictionary. In this embodiment, deep learning word embedding technology is used to determine the synonyms of the seed trigger words and construct a causal prompt dictionary. The word embedding model Word2vec widely used in the field of natural language processing is adopted to expand the causal prompt word dictionary in the geographical domain. First, the manually annotated causal prompt words are used as seed words, and the causal prompt words are expanded through the Word2vec model, that is, the words with high semantic relevance to these seed words are found, and it is manually judged whether they are causal relationship prompt words and added to the causal prompt word dictionary. In this embodiment, 15 words most relevant to each seed word are calculated, and it is analyzed whether these words lead to causal sentences in the corpus to determine whether they are causal prompt words. Then the irrelevant words are deleted, and the qualified words are added to the causal prompt word dictionary.

[0179] S4.2: Calculate the cosine similarity of the abstract geographical event;

[0180] S4.3: Generalize the abstract geographical event according to the cosine similarity:

[0181] The cosine similarity measures the size of the angle between two vectors, and the result is represented by the cosine value of the angle. The value range of the cosine similarity is [-1, 1], and the larger the value, the more similar. In this embodiment, starting from the word frequency of the words appearing in the geographical event, the types of words appearing are used as the dimensions of the event vector, and the number of times a single type of word appears is used as the length of each dimension to form an event vector to calculate its similarity;

[0182] If there are at least 5 similar abstract geographical events for an abstract geographical event E, then extract the common components in the abstract geographical event E and its similar abstract geographical events as the generalized geographical event;

[0183] Otherwise, it is considered that the abstract geographical event E lacks generality and is not generalized;

[0184] Among them, if the similarity between two abstract geographical events is greater than the preset similarity threshold (the similarity threshold is set to 0.8 in this embodiment), they are mutually similar abstract geographical events.

[0185] In the specific implementation process, the extracted causal geographical events are abstracted to avoid the one-sidedness of the understanding of geographical knowledge caused by individual cases.

[0186] It also includes verifying causal relationships by combining existing literature research and consulting domain experts, and filtering out pairs of geographical events that are not causal relationships.

[0187] S5: Construct a graph knowledge-guided sequence model based on the structured geographical knowledge graph;

[0188] S6: Generate geographical exam questions based on the graph knowledge-guided sequence model

[0189] More specifically, the graph knowledge-guided sequence model includes a graph encoder, a K-BERT text encoder, and a decoder; first, the geographical knowledge graph obtained in step S4 is input into the graph encoder to grasp the structural context information in the geographical knowledge graph, and any geographical knowledge text corpus is input into the K-BERT text encoder to capture the context information of the text, and then the decoder is used to generate corresponding geographical exam questions.

[0190] In the specific implementation process, the Graph Transformer in the graph encoder solves the dilemma that the previous graph neural network (GNN) cannot be used for heterogeneous graphs, realizes explicit information interaction between each node, and uses the shortest path relationship representation between nodes as the basis for retaining graph structure information. In addition, compared with the GNN-based method that only considers the node information aggregation within a single-hop range, the Graph Transformer can realize multi-hop high-quality domain knowledge reasoning.

[0191] More specifically, the graph encoder includes a graph preprocessing unit and a graph transformation unit;

[0192] In the graph preprocessing unit, the TransH method is used to represent the high-dimensional predicates and entities of the geographical knowledge graph as low-dimensional matrices P and E; by training matrices P and E, the total distance of all facts (s, r, o) is minimized; where s represents the head entity, r represents the relationship, o represents the tail entity, e s represents the semantic vector of the object entity, p r represents the semantic vector of the predicate, and e o represents the semantic vector of the subject entity;

[0193] As Figure 2 shown, in the graph transformation unit, the following formula is used to calculate and capture semantic information:

[0194]

[0195] Among them, represents the encoded information after splicing with attention, and e i represents the encoded information of the original input. | represents the concatenation operation of N attn, j ∈ N, and attn j is dot product calculation;

[0196]

[0197]

[0198] q i , k i , v i are the d k -dimensional vector representations after the i-th stacked block linearly transforms the input;

[0199]

[0200]

[0201] FF(·) is a two-layer feed-forward network, and LN is a normalization layer;

[0202] e1 = Concat(e s ; p r ; e o )

[0203] Perform layer normalization on the final output result e N :

[0204] e N = LN output (e N ).

[0205] In the specific implementation process, the entity is vectorized through the transH technology based on the translation model, realizing the representation of the semantic information of the entities on the non-relationship path while considering the semantic information of the entities on the relationship path, avoiding the problem of information heterogeneity caused by polysemy of entity information, further improving the reasoning ability of the model, and at the same time establishing connections between more entities without relationships, complementing the knowledge graph relationships, and achieving the role of mining deep relationships.

[0206] More specifically, the K-BERT text encoder includes a knowledge layer, an embedding layer, a noise filtering layer, and a mask learning layer;

[0207] In the knowledge layer, given an input sentence s = [w0, w1, w2,..., w n and a knowledge graph,

[0208] Select all entity names involved in sentence s through the following formula, and query their corresponding triples from the knowledge graph:

[0209] E = K-Query(s, KG)

[0210] E = [(w i , ri0, w i0 ),...,(w i , r ik , w ik )]

[0211] Among them, E represents the set of triples, the K-Query() function is the formal representation of knowledge query, KG represents the knowledge graph, and (w i , r ik , w ik ) represents the corresponding triple retrieved;

[0212] Associate the triples in E to the appropriate positions in the entity relationship graph t through the K-Inject function:

[0213] t = K-Inject(s, E);

[0214] In the embedding layer, use the vocabulary provided by Google-BERT. Through a trainable lookup table, convert each token in the sentence tree into an embedding vector of dimension H. Then use [CLS] as the classification token and [MASK] as the masking token. Then, through soft position embedding, sort the word vectors and add the structural information lost by the input sentence; finally, identify different sentences containing multiple sentences through segment embedding;

[0215] In the noise filtering layer, use the visibility matrix to limit the visible area of each vector:

[0216]

[0217] Among them, M ij represents the visualization matrix, represents that word w j and word w i are on the same branch;

[0218] In the mask learning layer, there are multiple stacks of Mask-self-attention, and the Mask-self-attention is:

[0219] Q i+1 , K i+1 , V i+1 = h i W q , h i W k , h i W v

[0220]

[0221] h i+1 = S i+1 V i+1

[0222] Among them, Q i+1 is the information to be queried, K i+1 is the vector to be queried, V i+1 is the value queried, h i is the hidden state of the i-th masked self-attention block, S i+1 is the attention score, M is the visible matrix calculated by the visual layer, d k is the scaling factor, W q , W k and W v are trainable model parameters, h i+1 is the hidden state of the (i + 1)-th masked self-attention block. As Figure 3 shown, the sentence "Irrigation agriculture is a mode of ensuring agricultural production by means of large-scale irrigation during drought" is input into the K-BERT text encoder for processing.

[0223] More specifically, the decoder uses the GPT-2 language model to decode and generate geography exam questions; the decoder is stacked by multiple decoding modules, and each decoding module includes a position encoding layer, a multi-head attention mechanism layer, and a batch normalization layer;

[0224] After obtaining the output vector h of the K-BERT text encoder and the output vector En of the graph encoder, it further includes concatenating h and En to obtain the vector z = [h:En], and then inputting z = [h:En] into the multi-head attention mechanism layer as the output of the encoding layer;

[0225] By sampling the conditional probability p(T|S), high-quality geography exam questions are generated:

[0226]

[0227] Among them, T represents the output target sequence T = x m+1 ,..., x N , S represents the already generated sequence S = x1,..., x m , N represents the total length of the target output sequence, p(x n |x1,..., x n-1 ) represents the conditional probability distribution of predicting the next word based on the already generated sequence.

[0228] In the specific implementation process, to improve the feature representation of triples, graph and word-level representations are adopted, initialized by combining a knowledge graph and K-BERT respectively, then generalized by a neural network, and finally connected to the GPT-2 language model for generating distractors and questions, which has a certain improvement in the semantic understanding ability of question generation and greatly enhances the generalization ability of the question generation model.

[0229] More specifically, it also includes generating distractors by combining the negative sampling mechanism CTRL (A Conditional Transformer Language Model For Controllable Generation). The specific steps are as follows:

[0230] CTRL learns p(x i |x <i , c) by inputting a text sequence with condition c:

[0231]

[0232] Decompose it using the chain rule and inject it into the loss function for training. Use the answer as the condition to adjust the generated distractor D:

[0233]

[0234] Represent the obtained n sequences with d-dimensional vectors, that is, obtain the matrix And inject it into the multi-head attention mechanism layer and the batch normalization layer.

[0235]

[0236] MultiHead(X, k) = [h1;...; h k W o

[0237] where h j = Attention(XW j 1 , XW j 2 , XW j 3 )

[0238] Use a feed-forward neural network layer with a ReLU activation function to project the input to the internal dimension f, where

[0239] FF(X) = max(0, XU)V

[0240] Finally, select the best three distractors through scoring:

[0241] Scores(X0)=LayerNorm(X t )W vocab

[0242] Among them, Scores(X0) is the score of the generated interference item, LayerNorm(X t ) is the normalization layer, W vocab is the vocabulary weight matrix. The best three distractors are selected by scoring. When the model can sometimes only generate less than three distractors, multiple iterations can be performed until three distractors are generated.

[0243] In the specific implementation process, the condition c is controlled to penalize the generation of similar texts to generate grammatically different distributions, thereby avoiding the generated distractors being consistent with the answers. At the same time, the generated distractors can be made similar to the answers to the original test questions, and have the characteristics of being confusing but wrong.

[0244] Example 3

[0245] like Figure 4 As shown, a device for generating geography test questions based on knowledge guidance is used to implement the method for generating geography test questions based on knowledge guidance, including a data construction module, a geographic knowledge graph construction module, a question generation module and a distractor generation module;

[0246] The data construction module is used to obtain unstructured geographic knowledge text corpus to construct a geographic text corpus;

[0247] The geographic knowledge graph construction module is used to construct a structured geographic knowledge graph; wherein, it includes:

[0248] The sentence recognition module is used to identify corresponding factual sentences from the geographic text corpus according to the syntactic template;

[0249] Event extraction module, used to extract geographical events from logical sentences;

[0250] Knowledge generalization module, used to generalize geographic events;

[0251] The question generation module is used to generate geography test questions according to the sequence model guided by graph knowledge;

[0252] The interference term generation module is used to generate interference terms in combination with the negative sampling mechanism CTRL.

[0253] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.< / verb> < / verb> < / verb> < / conj> < / conj> < / verb> < / conj> < / conj> < / verb> < / verb> < / conj> < / conj> < / verb> < / verb> < / verb> < / verb> < / verb> < / verb> < / verb> < / conj> < / conj> < / conj> < / verb> < / verb> < / verb> < / verb> < / conj> < / conj> < / verb> < / conj> < / conj> < / verb> < / verb> < / conj> < / conj> < / verb> < / verb> < / verb> < / verb> < / verb> < / verb> < / verb> < / conj> < / conj> < / conj> < / verb>

Claims

1. A knowledge-guided geographical exam question generation method, characterized in that, The following steps are involved: S1: Obtain unstructured geographic knowledge text corpus to build a geographic text corpus; S2: Setting a syntactic template and identifying corresponding factual sentences from the geographic text corpus according to the syntactic template; S3: Extract geographical events from logical sentences based on dependency parsing and semantic role labeling; S4: Generalize geographic events to obtain generalized geographic events, and construct a structured geographic knowledge graph based on the generalized geographic events, including: S4.1: Use the most frequently occurring syntactic combinations in the geographic text corpus to abstract geographic events and obtain abstract geographic events; S4.2: Calculate the cosine similarity of abstract geographic events; S4.3: Generalize abstract geographic events based on cosine similarity: If an abstract geographical event E has at least five similar abstract geographical events, the common components of the abstract geographical event E and its similar abstract geographical events are extracted as the generalized geographical event; Otherwise, it is considered that the abstract geographical event E lacks generality and is not generalized; If the similarity between two abstract geographical events is greater than a preset similarity threshold, they are similar abstract geographical events. S5: Construct a graph knowledge-guided sequence model based on the structured geographic knowledge graph; S6: Generate geography test questions based on sequential model guided by graph knowledge; Among them, the graph knowledge-guided sequence model includes a graph encoder, a K-BERT text encoder and a decoder; the geographic knowledge graph obtained in step S4 is first input into the graph encoder to grasp the structural context information in the geographic knowledge graph, and any geographic knowledge text corpus is input into the K-BERT text encoder to capture the context information of the text, and then the decoder is used to generate corresponding geography test questions.

2. The knowledge-guided geographical exam question generation method according to claim 1, characterized in that, The syntax template includes: Cause-to-effect front-end syntax template: <Conj...> {Cause},{Effect}; The centered syntax template from cause to effect: {Cause} <verb> {Effect};< / verb> The syntactic template from cause to effect: <conj>{Cause} <conj verb> {Effect};< / conj> < / conj> Cause-to-effect front-end matching syntax template: <conj> {Cause} <Verb>,{Effect};< / conj> The syntactic template of tracing back to the cause and effect: {Effect}<Conj...> {Cause}; Syntax template for tracing cause from effect: < Conj... >{Effect}< Conj... > {Cause}; From cause to effect, the three-layer causal relationship syntax template is centered: <Conj...>{Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb> {effect3};< / verb> < / verb> < / verb> The three-layer causal relationship syntactic template introduced by the cause-to-effect verb: {Cause1} <verb>{effect1 / Cause2}, <verb>{effect2 / cause3}, <verb> {effect3};< / verb> < / verb> < / verb> From cause to effect, the front-end supporting verb guides the three-layer causal relationship syntax template: <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb>{effect2 / cause3}, <verb> {effect3};< / verb> < / verb> < / conj> < / conj> From cause to effect, the second-level causal relationship syntax template is centered: <conj>{Cause1} <conj verb>,{Effect1 / cause2}, <verb> {effect2};< / verb> < / conj> < / conj> From cause to effect, the front-end supporting verb guides the second-level causal relationship syntax template: <conj>{Cause1} <conj>,{Effect1 / cause2}, <verb> {effect2};< / verb> < / conj> < / conj> From cause to effect, the middle verb introduces the two-level causal relationship syntax template: {Cause1} <verb>{effect1 / Cause2}, <verb> {effect2 / cause2}.< / verb> < / verb> 3. A method for generating geography exam questions based on knowledge guidance according to claim 2, characterized in that The specific steps of step S3 are: For any sentence, the trigger words of the geographical event in the sentence are first determined according to the syntactic template, and the participants of the geographical event in the sentence are identified through semantic role labeling. Then, the subject-predicate-object structure of the geographical event in the sentence is identified through dependency syntactic analysis, so as to extract the geographical event in the sentence. Among them, it is determined whether there is an agent A0 of the action or a recipient A1 of the action in the result of semantic role labeling; If A0 exists, the subject of the geographical event is represented by A0, otherwise the dependent child node of the dependency syntactic structure SBV is used as the subject; if the dependency syntactic structure SBV is also missing, the geographical event is represented as a verb-object structure; If A1 exists, the object of the geographical event is represented by A1, otherwise the dependent child node of the dependency syntactic structure VOB is used as the object; if the dependency syntactic structure VOB is also missing, the geographical event is represented as a subject-predicate structure; If no verb appears, a noun is used to represent the geographical event.

4. A method for generating geography exam questions based on knowledge guidance according to claim 1, characterized in that The graph encoder includes a graph preprocessing unit and a graph conversion unit; In the map preprocessing unit, the TransH method is used to represent the high-dimensional predicates and entities of the geographical knowledge map as low-dimensional matrices P and E; by training the matrices P and E, the total distance of all facts ( s , r , o ) is minimized; Among them, s represents the head entity, r represents the relationship, o represents the tail entity, represents the semantic vector of the object entity, represents the semantic vector of the predicate, represents the semantic vector of the subject entity; In the graph conversion unit, the captured semantic information is calculated by the following formula: Among them, represents the encoded information after splicing with attention, represents the encoded information of the original input, and | represents N a attn concatenation operation of j ∈ N , is the dot product calculation; , , is the i -dimensional vector representation after the input is linearly transformed by the th stacked block; FF(·) is a two-layer feedforward network, and LN is a normalization layer; For the final output result Perform layer normalization: 。 5. A method for generating geography exam questions based on knowledge guidance according to claim 1, characterized in that The K-BERT text encoder includes a knowledge layer, an embedding layer, a noise filtering layer, and a mask learning layer; In the knowledge layer, given an input sentence s = [w0, w1, w2, …, w n ], and a knowledge graph, All entity names involved in sentence s are selected by the following formula, and their corresponding triples are queried from the graph: E = K-Query(s, KG) E = [(w i , ri0, w i0 ), ..., (w i , r ik , w ik )] Among them, E represents the set of triples, the K-Query() function is a formal representation of knowledge query, KG represents the knowledge graph, and (w i , r ik , w ik ) represents the corresponding triple retrieved; The triples in E are associated to the appropriate positions in the entity relationship graph t through the K-Inject function: t = K-Inject(s, E); In the embedding layer, the vocabulary provided by Google-BERT is used to convert each token in the entity relationship graph into an embedding vector of dimension H through a trainable lookup table, and then [CLS] is used as the classification token, and [MASK] is used to mask the token. Then, the word vectors are sorted by embedding in the soft position, adding the lost structural information of the input sentence; finally, segment embedding is used to identify different sentences containing multiple sentences; In the noise filtering layer, a visibility matrix is ​​used to limit the visible area of ​​each vector: Among them, represents a visualization matrix, represents a word and the word are in the same branch; There are multiple stacks of Mask-self-attention in the mask learning layer, and the Mask-self-attention is: Among them, is the information to be queried, is the vector to be queried, is the value queried, is the i hidden state of the i-th masked self-attention block, is the attention score, M is the visible matrix calculated by the visual layer, is the scaling factor, , and are trainable model parameters, is the hidden state of the (i + 1)-th masked self-attention block.

6. A method for generating geography exam questions based on knowledge guidance according to claim 1, characterized in that, The decoder uses the GPT-2 language model to decode and generate geography test questions; the decoder is composed of a plurality of stacked decoding modules, each of which includes a position encoding layer, a multi-head attention mechanism layer, and a batch normalization layer; After obtaining the output vector h of the K-BERT text encoder and the output vector En of the graph encoder, h and En are concatenated to obtain the vector z=[h:En], and then z=[h:En] is input into the multi-head attention mechanism layer as the output of the encoding layer; By sampling the conditional probability p ( T | S ), high-quality geography exam questions are generated: Among them, T represents the target sequence of the output , S represents the sequence that has been generated , N represents the total length of the target output sequence, represents the conditional probability distribution for predicting the next word based on the sequence that has been generated.

7. A method for generating geography exam questions based on knowledge guidance according to claim 6, characterized in that, It also includes generating interference items by combining the negative sampling mechanism CTRL. The specific steps are: CTRL learns by inputting a text sequence with conditions c as follows : Decompose using the chain rule and inject it into the loss function for training. Use the answer as a condition to adjust the generated interference factors D : The obtained n sequences are represented by d dimensional vectors, i.e., the matrix is obtained, and it is injected into the multi-head attention mechanism layer and the batch normalization layer, and finally the best three distractors are selected by scoring: Among them, is the generated interference item score, is the normalization layer, is the vocabulary weight matrix.

8. A device for generating geography exam questions based on knowledge guidance, applied to the method for generating geography exam questions based on knowledge guidance according to any one of claims 1 to 7, characterized in that, It includes data construction module, geographic knowledge graph construction module, question generation module and interference item generation module; The data construction module is used to obtain unstructured geographic knowledge text corpus to construct a geographic text corpus; The geographic knowledge graph construction module is used to construct a structured geographic knowledge graph; wherein, it includes: The sentence recognition module is used to identify corresponding factual sentences from the geographic text corpus according to the syntactic template; Event extraction module, used to extract geographical events from logical sentences; Knowledge generalization module, used to generalize geographic events; The question generation module is used to generate geography test questions according to the sequence model guided by graph knowledge; The interference term generation module is used to generate interference terms in combination with the negative sampling mechanism CTRL.

Citation Information

Patent Citations

  • Chinese choice question interference item generation method based on free text

    CN112686025A

  • System and method for hybrid question answering over knowledge graph

    US20220292262A1