Answer type generation method and system, electronic equipment and storage medium
By constructing a knowledge graph and dependent syntax technical analysis, a question-and-answer question type that meets the requirements of educational evaluation is generated, which solves the problems of inefficiency and neglect of correlation of traditional question-type generation methods, and realizes the refined management of the test papers.
Patent Information
- Application Number
- CN202510358214.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The traditional method of question type generation is inefficient, and the correlation between question type and teaching content and student characteristics is ignored, making it difficult for the generated test papers to meet the refined requirements of educational assessment in terms of question type distribution, difficulty gradient, etc.
By extracting entity information, keyword information, text summary information and theme structure in the training document, a structured knowledge graph is constructed, and inquiry statements are generated for related questions, using dependent syntax technology to analyze the dependence relationship of word, calculate semantic similarity, determine the difficulty coefficient of the question, and generate the question type based on the difficulty coefficient and the preset test paper difficulty ratio.
Optimize the test paper structure to ensure that the test paper meets the requirements of educational assessment in terms of overall difficulty and question distribution, and improves the training effect.
Smart Images

Figure CN120336460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of question type generation, and in particular, to a method, system, electronic device, and storage medium for generating short-answer question types. Background Art
[0002] In the current training field, as a core component of training design and evaluation, question types play a crucial role in improving training effects. They are not only a key tool for measuring the knowledge mastery and skill application abilities of trainees, but also an important reference for guiding teaching content and optimizing teaching strategies. From basic multiple-choice questions and fill-in-the-blank questions, which focus on knowledge memorization and understanding, to short-answer questions and essay questions that focus on in-depth thinking, comprehensive analysis, and innovative expression, and then to highly practice-oriented question types such as case analysis, each question type carries a specific evaluation purpose and difficulty level.
[0003] However, traditional question type generation methods are often limited by fixed rules or teachers' empirical judgments. They manually screen and arrange question types. On the one hand, the efficiency is low; on the other hand, the relevance between question types, teaching content, and trainee characteristics is ignored, resulting in the generated test papers being difficult to meet the refined requirements of educational evaluation in terms of question type distribution, difficulty gradient, etc. Summary of the Invention
[0004] The main purpose of this application is to overcome the disadvantages and deficiencies of the prior art, and provide a method, system, electronic device, and storage medium for generating short-answer question types, which can flexibly customize appropriate question types according to the difficulty coefficient of the questions and the preset difficulty ratio of the test paper question types, ensuring that the test paper meets the requirements of educational evaluation in terms of overall difficulty, question type distribution, etc.
[0005] To achieve the above purpose, this application adopts the following technical solutions:
[0006] In the first aspect, this application provides a method for generating short-answer question types, including the following steps:
[0007] Extract the first entity information, first keyword information, text summary information, and information on the theme structure from the training document;
[0008] Construct a structured knowledge graph from the training document according to the first entity information, first keyword information, text summary information, and information on the theme structure;
[0009] Generate inquiry statements for related questions based on the knowledge graph;
[0010] Extract the second keyword information and second entity information from the inquiry statements of the related questions;
[0011] Retrieve the text related to the inquiry statement in the training document according to the second keyword information and the second entity information;
[0012] Use dependency syntax technology to analyze the dependency relationships between words in the text related to the inquiry statement to obtain candidate answers for the inquiry statement;
[0013] Calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement;
[0014] Determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement;
[0015] Generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the test paper question types, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
[0016] As a preferred technical solution, the extraction of the first entity information in the training document includes:
[0017] Segment the text in the training document into word groups, and perform part-of-speech tagging on the word groups to obtain word group texts with part-of-speech tags;
[0018] Extract the features of the word group texts with part-of-speech tags and initialize the features as feature vectors;
[0019] Use a preset feature function and weights to calculate the conditional probabilities of all tag sequences under the word group texts with part-of-speech tags according to the feature vectors;
[0020] Decode the conditional probabilities, select the tag sequence with the highest conditional probability, and use the highest tag sequence as the first entity information in the training text.
[0021] As a preferred technical solution, the extraction of the first keyword information in the training document includes:
[0022] Calculate the frequency of occurrence of the word group in the training document;
[0023] Calculate the inverse document frequency of the word group in the training document; the inverse document frequency is a statistic for evaluating the importance of the word group in the training document;
[0024] Calculate the weight value of the word group according to the frequency of occurrence of the word group in the training document and the inverse document frequency of the word group in the training document;
[0025] Sort each word group in descending order according to the weight value, and select the top N word groups as the first keyword information in the training document.
[0026] As a preferred technical solution, generating an inquiry statement for a related question based on the knowledge graph includes:
[0027] Filling the first entity information, relationship or attribute value of the knowledge graph into the placeholder of a preset question template to generate a preliminary question inquiry statement;
[0028] Encoding the preliminary question statement to obtain a deep semantic representation;
[0029] Decoding the deep semantic representation to generate an optimized question inquiry statement.
[0030] As a preferred technical solution, using the dependency syntax technology to analyze the dependency relationship between words in the text to obtain candidate answers for the inquiry statement includes:
[0031] Parsing the sentence structure in the text related to the inquiry statement and constructing a dependency syntax tree; wherein each node in the dependency syntax tree represents a phrase, and the edge between nodes represents the dependency relationship between words;
[0032] Traversing the dependency syntax tree to find the dependency relationship chain related to the second keyword information in the inquiry statement;
[0033] Extracting key information fragments from the dependency relationship chain according to the dependency relationship and semantic logic;
[0034] Combining the key information fragments into a complete sentence or phrase to form a candidate answer.
[0035] As a preferred technical solution, obtaining the difficulty coefficient of each question according to the inquiry statement type and the answer of the inquiry statement includes:
[0036] Analyzing the semantic content of the inquiry statement to determine the inquiry statement type;
[0037] Determining a first difficulty value according to the inquiry statement type;
[0038] Obtaining the information length, information type and logical relationship characteristics between information of the answer corresponding to the inquiry statement;
[0039] Determining a second difficulty value according to the information length, information type and logical relationship characteristics between information of the answer corresponding to the inquiry statement;
[0040] Performing weighting according to the first difficulty value, the second difficulty value, the preset difficulty weight of the inquiry statement and the preset difficulty weight of the answer to obtain the difficulty coefficient.
[0041] As a preferred technical solution, generating the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the test paper question types includes:
[0042] According to the preset difficulty coefficient range value, allocate the difficulty coefficient of the question to the corresponding question types to obtain the preliminary question types of the test paper;
[0043] Calculate the difficulty ratio in the preliminary question types of the test paper. If the difficulty ratio of the preliminary question types of the test paper meets the preset difficulty ratio of the test paper question types, then use the preliminary question types of the test paper as the final question types;
[0044] If the difficulty ratio of the preliminary question types of the test paper does not meet the preset difficulty ratio of the test paper question types, then adjust the preliminary question types of the test paper until the preset difficulty ratio of the test paper question types is met.
[0045] In a second aspect, the present application provides an open-ended question type generation system, which is applied to the described open-ended question type generation method, and includes a first extraction module, a knowledge graph construction module, an inquiry statement generation module, a second extraction module, a retrieval module, an analysis module, an answer acquisition module, a difficulty coefficient determination module, and a question type generation module;
[0046] The first extraction module is used to extract the first entity information, the first keyword information, the text summary information, and the information of the theme structure in the training document;
[0047] The knowledge graph construction module is used to construct the training document into a structured knowledge graph according to the first entity information, the first keyword information, the text summary information, and the information of the theme structure;
[0048] The inquiry statement generation module is used to generate inquiry statements for related questions based on the knowledge graph;
[0049] The second extraction module is used to extract the second keyword information and the second entity information in the inquiry statements of the related questions;
[0050] The retrieval module is used to retrieve the text related to the inquiry statement in the training document according to the second keyword information and the second entity information;
[0051] The analysis module is used to analyze the dependency relationship between words in the text related to the inquiry statement by using the dependency parsing technology to obtain candidate answers for the inquiry statement;
[0052] The answer acquisition module is used to calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement;
[0053] The difficulty coefficient determination module is used to determine the difficulty coefficient of each question according to the question type of the inquiry statement and the answer of the inquiry statement;
[0054] The question type generation module is used to generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the question types in the test paper, where the difficulty ratio refers to the expected proportion of the question types with different difficulty coefficients in the whole test paper.
[0055] In a third aspect, the present application provides an electronic device, which includes:
[0056] At least one processor; and a memory communicatively connected to the at least one processor;
[0057] Wherein, the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for generating question types of short-answer questions.
[0058] In a fourth aspect, the present application provides a computer-readable storage medium storing a program, which when executed by a processor, implements the method for generating question types of short-answer questions.
[0059] In summary, compared with the prior art, the effective effects brought by the technical solution provided by the present application at least include:
[0060] The present application proposes a method for generating question types of short-answer questions, which extracts the first entity information, the first keyword information, the text summary information and the information of the theme structure from the training document; constructs the training document into a structured knowledge graph according to the first entity information, the first keyword information, the text summary information and the information of the theme structure; generates inquiry statements of relevant questions based on the knowledge graph; extracts the second keyword information and the second entity information in the inquiry statements of relevant questions; retrieves the text related to the inquiry statement in the training document according to the second keyword information and the second entity information; analyzes the dependency relationship between words in the text related to the inquiry statement by using the dependency syntax technology to obtain candidate answers of the inquiry statement; calculates the semantic similarity between the candidate answers and the inquiry statement, and obtains the answer of the inquiry statement based on the similarity; determines the difficulty coefficient of each question according to the question type of the inquiry statement and the answer of the inquiry statement; generates the question type of each question based on the difficulty coefficient of the test question and the preset difficulty ratio of the question types in the test paper. It can be seen that the present application flexibly customizes appropriate question types according to the difficulty coefficient of the test questions and the preset difficulty ratio of the question types in the test paper, thereby optimizing the test paper structure and ensuring that the test paper meets the requirements of educational evaluation in terms of overall difficulty, question type distribution, etc. Description of the Drawings
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0062] Figure 1 Flow chart of a method for generating question-and-answer question types provided by an embodiment of the present application;
[0063] Figure 2 Flow chart of the steps for extracting first entity information provided by an embodiment of the present application;
[0064] Figure 3 Flow chart of the steps for extracting first keyword information provided by an embodiment of the present application;
[0065] Figure 4 Flow chart of the steps for generating inquiry statements for related questions provided by an embodiment of the present application;
[0066] Figure 5 Flow chart of the steps for obtaining candidate answers provided by an embodiment of the present application;
[0067] Figure 6 Flow chart of the steps for determining the difficulty coefficient of each question provided by an embodiment of the present application;
[0068] Figure 7 Flow chart of the steps for generating question types provided by an embodiment of the present application;
[0069] Figure 8 Block diagram of a question-and-answer question type generation system provided by an embodiment of the present application;
[0070] Figure 9 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0071] In order to enable those skilled in the art of this technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by the present application.
[0072] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment each time, nor are they independent or alternative embodiments mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0073] As a core component of training design and evaluation, the question type of examination questions plays a crucial role in improving the training effect. It is not only a key tool for measuring the knowledge mastery and skill application ability of trainees, but also an important reference for guiding teaching content and optimizing teaching strategies; each question type carries specific evaluation purposes and difficulty levels. However, traditional question type generation methods are often limited by fixed rules or the empirical judgment of teachers. They select and arrange question types manually. On the one hand, the efficiency is low; on the other hand, the relevance between question types, teaching content, and trainee characteristics is ignored, resulting in the generated test papers being difficult to meet the refined requirements of educational evaluation in terms of question type distribution, difficulty gradient, etc.
[0074] To solve the above problems, in the embodiments of this specification, first, key knowledge points (first entity information, first keyword information, text summary information, and information on the topic structure) in the training document are extracted, and the key knowledge points are constructed into a structured knowledge graph, which represents the connections between different knowledge points; and according to different knowledge points in the knowledge graph, inquiry statements related to relevant questions are generated; secondly, according to the second keyword information of the inquiry statements, texts related to the inquiry statements are retrieved in the training document, and the dependency relationships between words in the texts related to the inquiry statements are analyzed based on dependency syntax technology to obtain candidate answers, and then the final answers are obtained according to the semantic similarity between the candidate answers and the inquiry statements; according to the type of the inquiry statements and the answers to the inquiry statements, the difficulty coefficient of each question is determined; finally, based on the difficulty coefficient of the questions and the preset difficulty ratio of the test paper question types, the question type of each question is generated. Flexible customization of appropriate question types according to the difficulty coefficient of the questions and the preset difficulty ratio of the test paper question types can optimize the test paper structure, ensure that the test paper meets the requirements of educational evaluation in terms of overall difficulty, question type distribution, etc., and improve the training effect.
[0075] The following describes in detail the technical solutions provided by the embodiments in this application with reference to the accompanying drawings.
[0076] Please refer to Figure 1 , in an embodiment of this application, a method for generating question-and-answer question types is provided, including the following steps:
[0077] S1. Extract the first entity information, first keyword information, text summary information, and information on the topic structure in the training document.
[0078] Further, first obtain training documents, including multimedia training document materials such as text, pictures, audio and video; and clean, format-unify and preliminarily structure the training documents, aiming to extract useful information from the multimedia training document materials and convert it into text data that is easier to process and analyze.
[0079] Among them, the entity information in this embodiment refers to an object or concept with a clear meaning and independence in the training document; these entity information can be specific nouns, such as names of people, places, organizations, product names, etc., or abstract concepts, such as technical terms, theoretical models, event names, etc.
[0080] Please refer to Figure 2 , to extract the first entity information in the training document, the steps include:
[0081] S11.1. Segment the text in the training document into phrases, and perform part-of-speech tagging on the phrases to obtain phrase text with part-of-speech tags;
[0082] Further, before segmenting the text into phrases, it also includes preprocessing the text in the training document, including removing punctuation marks and stop words (such as common but meaningless words like "de", "le", etc.), in order to reduce noise in the subsequent processing and reduce the scale of the text data.
[0083] Segmenting the text in the training document into phrases aims to cut the continuous text stream into independent meaningful words or phrase units. Secondly, after word segmentation, use a part-of-speech tagging tool or model to perform part-of-speech tagging on each phrase, that is, assign corresponding part-of-speech tags (such as nouns, verbs, adjectives, etc.) according to its grammatical function in the sentence.
[0084] S11.2. Extract the features of the phrase text with part-of-speech tags and initialize the features as a feature vector;
[0085] Further, extracting the features of the phrase text with part-of-speech tags includes extracting the text content of the current phrase and the N phrases before and after the current phrase as lexical features; extracting the part-of-speech tags of the current phrase and the N phrases before and after the current phrase; the character composition of the current phrase, extracting character-level features such as prefixes, suffixes, and whether it is all uppercase, as well as context information features such as the position of the current phrase in the sentence and the relationship with the context before and after. Integrate the extracted features and convert them into a numerical vector to obtain a multi-dimensional feature vector.
[0086] S11.3. Use a preset feature function and weights to calculate the conditional probabilities of all possible tag sequences under the phrase text with part-of-speech tags according to the feature vector;
[0087] Further, input the feature vector into a pre-trained conditional random field model, and use the feature functions and weights preset in the conditional random field model to calculate the conditional probabilities of all possible tag sequences for the phrase text with part-of-speech tags. The calculation formula is as follows:
[0088]
[0089] where X is the phrase text with part-of-speech tags, y is the tag sequence, n is the sequence length, f k is the feature function, λ k is the weight of the feature function, and Z(x) is the normalization factor.
[0090] S11.4. Decode the conditional probabilities, select the tag sequence with the highest conditional probability, and use the highest tag sequence as the first entity information in the training text.
[0091] Further, a dynamic programming algorithm (such as the Viterbi algorithm) can be used to decode the conditional probabilities output by the conditional random field model to find the tag sequence with the highest conditional probability, that is, the most likely named entity recognition result. Specifically, the Viterbi algorithm is used to decode the conditional probabilities output by the conditional random field model to find the tag sequence with the highest conditional probability.
[0092] Please refer to Figure 3 to extract the first keyword information in the training document. The steps include:
[0093] S12.1. Calculate the frequency of the phrase in the training document;
[0094] S12.2. Calculate the inverse document frequency of the phrase in the training document; the inverse document frequency is a statistic used to evaluate the importance of the phrase in the training document.
[0095] Further, if a phrase appears frequently in many documents but rarely in other documents, its inverse document frequency will be low, indicating that this phrase is less important for distinguishing documents; conversely, if a phrase appears only in a few documents, its inverse document frequency will be high, indicating that this phrase is more important for the theme or content of the document.
[0096] S12.3. Calculate the weight value of the phrase according to the frequency of the phrase in the training document and the inverse document frequency of the phrase in the training document;
[0097] Further, multiply the frequency of the phrase in the training document by the inverse document frequency of the phrase in the training document to obtain the weight value of the phrase.
[0098] S12.4. Sort each phrase in descending order according to the weight value, and select the top N phrases as the first keyword information in the training document.
[0099] In this embodiment, the steps for extracting the text summary of the training document include:
[0100] (1) Cut the text in the training document into a list of individual sentences according to sentence delimiters (such as full stops, question marks, exclamation marks, etc.);
[0101] (2) Obtain the position information of the sentence in the document, and assign a first weight to the sentence according to its position information;
[0102] (3) Assign a second weight to the sentence according to the length of the sentence;
[0103] (4) Count the high frequency of the phrases in the sentence in the training document, and assign a third weight to the sentence according to its frequency;
[0104] (5) Calculate the similarity between the sentence and the document title, and assign a fourth weight to the sentence according to the similarity;
[0105] (6) Perform a weighted sum of the above first weight, second weight, third weight, and fourth weight to calculate the comprehensive importance score of each sentence; and sort the sentences according to the importance score, select the top m sentences with the highest scores as the text summary, and rearrange the selected sentences in the original order to form the final text summary information.
[0106] In this embodiment, the LDA (Latent Dirichlet Allocation) algorithm can be used to identify the topic structure of the document; by statistically analyzing the phrase distribution in the document, the LDA algorithm can infer the potential relationships between the document and the topics, and between the topics and the phrases, so as to identify the structural information of the topics.
[0107] S2. Construct a structured knowledge graph for the training document according to the first entity information, the first keyword information, the text summary information, and the information of the topic structure;
[0108] Further, after identifying the first entity information from the text, it also includes analyzing the relationships (such as parent-child relationship, subordination relationship, similarity relationship, etc.) between the first entity information according to the training document.
[0109] Furthermore, the knowledge graph represents the connections between different knowledge points and is a knowledge base that uses a graph structure to represent the relationships between entities, storing and representing information in a structured form. Specifically, the first entity information is used as the nodes of the knowledge graph, while the first keyword information can be used as the attributes or labels of these nodes; the relationships between the first entity information are used as the edges of the knowledge graph; using the text summary information and the theme structure information, a higher-level hierarchical structure or classification system is established in the knowledge graph. For example, multiple related first entity information and their relationships are organized into different subject areas or chapters.
[0110] In this embodiment, based on the first entity information, the first keyword information, the text summary information, and the theme structure information, the comprehensiveness and accuracy of the key knowledge points in the training document are ensured; the knowledge graph constructed based on these key knowledge points can more accurately reflect the core and key points of the training or teaching content.
[0111] S3. Based on the knowledge graph, generate inquiry statements for related questions.
[0112] Please refer to Figure 4 , the steps of generating inquiry statements for related questions based on the knowledge graph include:
[0113] S31. Fill the first entity information, relationships, or attribute values of the knowledge graph into the placeholders of the preset question template to generate preliminary question inquiry statements;
[0114] Furthermore, the first entity information, relationships, or attribute values in the knowledge graph are used as the cornerstones for generating questions; the first entity information, relationships, or attribute values are matched with predefined question templates, and appropriate templates are selected to initially generate a series of structured inquiry statements. For example, "What is the {relationship} of {Entity A}?" or "How is the {relationship} reflected between {Entity A} and {Entity B}?"
[0115] S32. Encode the preliminary question statements to obtain deep semantic representations;
[0116] However, questions generated only relying on templates often lack depth and diversity; to improve the quality of questions, semantic encoding of the initially generated question statements is required; the encoder part of the BART model in natural language processing is used to capture the deep semantic information in the statements.
[0117] S33. Decode the deep semantic representations to generate optimized question inquiry statements.
[0118] After obtaining the deep semantic representation of the question, the decoder of the BART model converts the encoded semantic information back into the form of natural language, aiming to improve the fluency and diversity of the query statement. That is, the decoder will adjust the structure, word order or vocabulary selection of the query statement to ensure that the question not only retains the core of the original graph information but is also easy for humans to understand and answer. For example, the decoder will optimize a query statement like "Who invented the light bulb?" to "Who is the inventor of the light bulb?". Although both convey the same information, the latter is more natural and fluent in expression.
[0119] S4. Extract the second keyword information and the second entity information in the query statement of the relevant question.
[0120] Furthermore, extract the second keyword information and the second entity information in the query statement of the relevant question that can accurately reflect the theme or query intention of the query statement.
[0121] S5. Retrieve the text related to the query statement in the training document according to the second keyword information and the second entity information.
[0122] Furthermore, use the extracted second keyword information and second entity information to match in the training document, and find all text paragraphs containing these second keyword information and second entity information. For the matched text paragraphs, filter out the text paragraphs with high relevance to the query statement; the relevance can be measured by various factors, such as the matching degree of keywords, the length of the text paragraph, the position (such as whether it appears at the beginning or end of the document), and semantic similarity, etc.
[0123] S6. Analyze the dependency relationship between words in the text related to the query statement by using dependency syntax technology to obtain candidate answers to the query statement.
[0124] Dependency syntax analysis technology reveals the grammatical structure of a sentence by analyzing the dependency relationship between words in the sentence.
[0125] Please refer to Figure 5 , use dependency syntax technology to analyze the dependency relationship between words in the text related to the query statement to obtain candidate answers to the query statement, and its steps include:
[0126] S61. Analyze the sentence structure in the text related to the query statement and construct a dependency syntax tree; where each node in the dependency syntax tree represents a phrase, and the edges between nodes represent the dependency relationship between words.
[0127] Among them, the dependency relationship usually includes the relationship type (such as subject-predicate relationship, verb-object relationship, etc.) and the directionality (that is, which word is the dependent word and which word is the governing word).
[0128] S62. Traverse the dependency syntax tree to find the dependency relationship chain related to the second keyword information in the inquiry statement;
[0129] Locate the node corresponding to the keyword in the dependency syntax tree. Subsequently, starting from the corresponding node, explore upward or downward along the dependency relationship chain to find words and relationships that may contain answer information or answer clues.
[0130] S63. Extract key information fragments from the dependency relationship chain according to the dependency relationship and semantic logic;
[0131] S64. Combine the key information fragments into a complete sentence or phrase to form a candidate answer.
[0132] Furthermore, it also includes correcting the candidate answer in terms of grammar and semantics to ensure that the statement is smooth and logical.
[0133] S7. Calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement.
[0134] Furthermore, for the inquiry statement and the candidate answer, extract the feature vectors of the inquiry statement and the candidate answer respectively; the feature vector is intended to reflect the semantic content of the text. Adopt the cosine similarity measurement method to calculate the similarity between the feature vector of the inquiry statement and the feature vector of each candidate answer; sort the candidate answers according to the similarity score, where the higher the similarity score of the candidate answer, the closer the semantics to the inquiry statement. According to the sorting result, select the candidate answer with the highest similarity score as the final answer to the inquiry statement.
[0135] S8. Determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement.
[0136] Specifically, please refer to Figure 6 , to determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement, the steps include:
[0137] S81. Analyze the semantic content of the inquiry statement to determine the inquiry statement type;
[0138] Furthermore, the inquiry statement types include comparison type (such as asking about the similarities and differences between two or more objects, concepts or situations), explanation type (such as asking to explain the meaning, reason or mechanism of a certain concept, phenomenon or process), fact type (such as asking about a specific fact, data or information), etc.
[0139] S82. Determine the first difficulty value according to the inquiry statement type;
[0140] Specifically, the difficulty value of the inquiry statement type will be preset. For example, if the inquiry statement type is a comparison type, the first difficulty value is 0.3, the interpretation type is 0.5, and the fact type is 0.2.
[0141] S83. Obtain the information length, information type, and logical relationship characteristics between the information of the answer corresponding to the inquiry statement;
[0142] Furthermore, the information length refers to the length of the answer or the amount of information it contains, which can be measured by indicators such as the number of characters, the number of words, or the information density. Information types, such as factual data, definition explanations, comparative analysis, etc. Understanding the logical relationship between the information is crucial for understanding the structure and logic of the answer, such as causal relationships, parallel relationships, progressive relationships, etc.
[0143] S84. Determine the second difficulty value according to the information length, information type, and logical relationship characteristics between the information of the answer corresponding to the inquiry statement.
[0144] S85. Perform weighted summation according to the first difficulty value, the second difficulty value, the preset difficulty weight of the inquiry statement, and the preset difficulty weight of the answer to obtain the difficulty coefficient.
[0145] Specifically, the difficulty coefficient = (the first difficulty value × the preset difficulty weight of the inquiry statement) + (the second difficulty value × the preset difficulty weight of the answer).
[0146] S9. Generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the question types in the test paper, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
[0147] Please refer to Figure 7 , generating the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the question types in the test paper, the steps include:
[0148] S91. According to the preset difficulty coefficient range value, allocate the difficulty coefficient of the question to the corresponding question types to obtain the preliminary question types of the test paper;
[0149] S92. Calculate the difficulty ratio in the preliminary question types of the test paper. If the difficulty ratio of the preliminary question types of the test paper meets the preset difficulty ratio of the question types in the test paper, then use the preliminary question types of the test paper as the final question types;
[0150] S93. If the difficulty ratio of the preliminary question types of the test paper does not meet the preset difficulty ratio of the question types in the test paper, then adjust the preliminary question types of the test paper until the preset difficulty ratio of the question types in the test paper is met.
[0151] Further, the question types include multiple-choice questions, matching questions, sorting questions, fill-in-the-blank questions, short-answer questions, essay questions, etc.
[0152] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be carried out in other sequences or simultaneously.
[0153] Based on the same idea as a method for generating essay question types in the above embodiments, this application also provides a system for generating essay question types, and this system can be used to execute the above method for generating essay question types. For the sake of convenience of description, in the structural schematic diagram of an embodiment of the system for generating essay question types, only the parts related to the embodiments of this application are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the system, and it may include more or fewer components than those illustrated, or combine some components, or arrange different components.
[0154] Please refer to Figure 8 In another embodiment of this application, a system for generating essay question types is provided. This system includes a first extraction module 101, a knowledge graph construction module 102, a query statement generation module 103, a second extraction module 104, a retrieval module 105, an analysis module 106, an answer acquisition module 107, a difficulty coefficient determination module 108, and a question type generation module 109;
[0155] The first extraction module 101 is used to extract the first entity information, the first keyword information, the text summary information, and the information of the topic structure in the training document;
[0156] The knowledge graph construction module 102 is used to construct the training document into a structured knowledge graph according to the first entity information, the first keyword information, the text summary information, and the information of the topic structure;
[0157] The query statement generation module 103 is used to generate query statements for related questions based on the knowledge graph;
[0158] The second extraction module 104 is used to extract the second keyword information and the second entity information in the query statements of the related questions;
[0159] The retrieval module 105 is used to retrieve the text related to the query statement in the training document according to the second keyword information and the second entity information;
[0160] The analysis module 106 is used to analyze the dependency relationship between words in the text related to the query statement by using the dependency parsing technique to obtain candidate answers to the query statement;
[0161] The answer acquisition module 107 is configured to calculate the semantic similarity between the candidate answer and the inquiry statement, and obtain the answer to the inquiry statement based on the similarity.
[0162] The difficulty coefficient determination module 108 is configured to determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement.
[0163] The question type generation module 109 is configured to generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the test paper question types, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
[0164] As a preferred technical solution, it further includes a first entity extraction module, and the first entity extraction module is specifically configured to:
[0165] Segment the text in the training document into word groups, perform part-of-speech tagging on the word groups, and obtain the word group text with part-of-speech tags.
[0166] Extract the features of the word group text with part-of-speech tags, and initialize the features as feature vectors.
[0167] Use a preset feature function and weights to calculate the conditional probability of all tag sequences under the word group text with part-of-speech tags according to the feature vectors.
[0168] As a preferred technical solution, it further includes a first keyword information extraction module, and the first keyword information extraction module is specifically configured to:
[0169] Calculate the frequency of occurrence of the word group in the training document.
[0170] Calculate the inverse document frequency of the word group in the training document; the inverse document frequency is a statistic for evaluating the importance of the word group in the training document.
[0171] Calculate the weight value of the word group according to the frequency of occurrence of the word group in the training document and the inverse document frequency of the word group in the training document.
[0172] Sort each word group in descending order according to the weight value, and select the top N word groups as the first keyword information in the training document.
[0173] As a preferred technical solution, the inquiry statement generation module 103 is specifically configured to:
[0174] Fill the first entity information, relationship or attribute value of the knowledge graph into the placeholder of the preset question template to generate a preliminary question inquiry statement.
[0175] Encode the preliminary question statement to obtain a deep semantic representation;
[0176] Decode the deep semantic representation to generate an optimized question query statement.
[0177] As a preferred technical solution, the analysis module 106 is specifically configured to:
[0178] Parse the sentence structure in the text related to the query statement and construct a dependency syntax tree; wherein, each node in the dependency syntax tree represents a phrase, and the edges between the nodes represent the dependency relationships between words;
[0179] Traverse the dependency syntax tree to find the dependency relationship chain related to the second keyword information in the query statement;
[0180] Extract key information fragments from the dependency relationship chain according to the dependency relationship and semantic logic;
[0181] Combine the key information fragments into a complete sentence or phrase to form a candidate answer.
[0182] As a preferred technical solution, the difficulty coefficient determination module 108 is specifically configured to:
[0183] Analyze the semantic content of the query statement to determine the type of the query statement;
[0184] Determine a first difficulty value according to the type of the query statement;
[0185] Obtain the information length, information type, and logical relationship characteristics between information of the answer corresponding to the query statement;
[0186] Determine a second difficulty value according to the information length, information type, and logical relationship characteristics between information of the answer corresponding to the query statement;
[0187] Perform weighting according to the first difficulty value, the second difficulty value, the preset difficulty weight of the query statement, and the preset difficulty weight of the answer to obtain a difficulty coefficient.
[0188] As a preferred technical solution, the question type generation module 109 is specifically configured to allocate the difficulty coefficient of the question to the corresponding question types according to the preset difficulty coefficient range value to obtain a preliminary question type of the test paper;
[0189] Calculate the difficulty ratio in the preliminary question type of the test paper. If the difficulty ratio of the preliminary question type of the test paper meets the preset test paper question type difficulty ratio, then use the preliminary question type of the test paper as the final question type;
[0190] If the difficulty ratio of the preliminary question types of the test paper does not meet the preset difficulty ratio of the test paper question types, adjust the preliminary question types of the test paper until the preset difficulty ratio of the test paper question types is met.
[0191] It should be noted that a question type generation system of the present application corresponds one-to-one with a question type generation method of the present application. The technical features and their beneficial effects described in the embodiments of the above-mentioned question type generation method are all applicable to the embodiments of a question type generation system. For specific content, reference can be made to the description in the method embodiments of the present application, which will not be repeated here. This is hereby declared.
[0192] In addition, in the implementation manner of a question type generation system in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the question type generation system is divided into different program modules to complete all or part of the functions described above.
[0193] Please refer to Figure 9 , in another embodiment, there is provided an electronic device for a question type generation method, including a processor, a memory, and a bus, and may further include a computer program stored in the memory and executable on the processor.
[0194] Exemplarily, in this embodiment, the computer program may be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present application. The one or more module elements may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device.
[0195] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor and a memory.
[0196] In some embodiments, the processor may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or it may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs). It may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, connecting various components of the entire device through various interfaces and lines; by running or executing programs or modules stored in the memory, and calling data stored in the memory, it performs various functions of the electronic device and processes data.
[0197] The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0198] Among them, in some embodiments, the memory may be an internal storage unit of the electronic device, such as a mobile hard disk of the electronic device. In some other embodiments, the memory may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory may also include both an internal storage unit and an external storage device of the electronic device. The memory can be used not only to store application software installed on the electronic device and various types of data, such as the code of a question-and-answer type generation program, etc., but also to temporarily store data that has been output or will be output.
[0199] Figure 9 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 9 The shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0200] A question-and-answer type generation program stored in the memory of the electronic device is a combination of multiple instructions. When running in the processor, it can implement:
[0201] Extract the first entity information, the first keyword information, the text summary information, and the information of the topic structure from the training document;
[0202] Construct a structured knowledge graph from the training document according to the first entity information, the first keyword information, the text summary information, and the information of the topic structure;
[0203] Generate inquiry statements for related questions based on the knowledge graph;
[0204] Extract the second keyword information and the second entity information from the inquiry statements of the related questions;
[0205] Retrieve the text related to the inquiry statement in the training document according to the second keyword information and the second entity information;
[0206] Use dependency syntax technology to analyze the dependency relationship between words in the text related to the inquiry statement to obtain candidate answers to the inquiry statement;
[0207] Calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement;
[0208] Determine the difficulty coefficient of each question according to the type of the inquiry statement and the answer to the inquiry statement;
[0209] Generate the question type for each question based on the difficulty coefficient of the question and the preset difficulty ratio of the question types in the test paper, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
[0210] Correspondingly, the present application further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute a method for generating question types as described in any one of the above embodiments.
[0211] The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, and so on.
[0212] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0213] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0214] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present application shall be equivalent replacement methods and are all included in the protection scope of the present application.
Claims
1. A method for generating question-and-answer type questions, characterized in that, Including the following steps: Extract the first entity information, first keyword information, text summary information, and subject structure information from the training document; Construct the training document into a structured knowledge graph according to the first entity information, first keyword information, text summary information, and subject structure information; Generate inquiry statements for relevant questions based on the knowledge graph; Extract the second keyword information and second entity information from the inquiry statements of the relevant questions; Retrieve the text related to the inquiry statement in the training document according to the second keyword information and second entity information; Use dependency syntax technology to analyze the dependency relationship between words in the text related to the inquiry statement to obtain candidate answers to the inquiry statement; Calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement; Determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement; Generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the test paper question types, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
2. The method for generating a question-and-answer type according to claim 1, wherein The extraction of the first entity information from the training document includes: Segment the text in the training document into phrases, and perform part-of-speech tagging on the phrases to obtain phrase texts with part-of-speech tags; Extract the features of the phrase texts with part-of-speech tags and initialize the features as feature vectors; Use a preset feature function and weights to calculate the conditional probability of all tag sequences under the phrase texts with part-of-speech tags according to the feature vectors; Decode the conditional probability, select the tag sequence with the highest conditional probability, and use the highest tag sequence as the first entity information in the training text.
3. The method for generating a question-and-answer type according to claim 2, wherein, The extraction of the first keyword information from the training document includes: Calculate the frequency of the phrase appearing in the training document; Calculate the inverse document frequency of the phrase in the training document; the inverse document frequency is a statistic for evaluating the importance of the phrase in the training document; Calculate the weight value of the phrase according to the frequency of the phrase appearing in the training document and the inverse document frequency of the phrase in the training document; Sort each phrase in descending order according to the weight value, and select the top N phrases as the first keyword information in the training document.
4. The method for generating a question-and-answer type according to claim 1, wherein The generation of inquiry statements for relevant questions based on the knowledge graph includes: Fill the first entity information, relationship, or attribute value of the knowledge graph into the placeholder of the preset question template to generate a preliminary question inquiry statement; Encode the preliminary question statement to obtain a deep semantic representation; Decode the deep semantic representation to generate an optimized question inquiry statement.
5. The method for generating a question-and-answer type according to claim 1, characterized in that, The use of dependency syntax technology to analyze the dependency relationship between words in the text to obtain candidate answers to the inquiry statement includes: Parse the sentence structure in the text related to the query statement and construct a dependency syntax tree; wherein, each node in the dependency syntax tree represents a phrase, and the edges between the nodes represent the dependency relationships between words; Traverse the dependency syntax tree to find the dependency relationship chain related to the second keyword information in the query statement; Extract key information fragments from the dependency relationship chain according to the dependency relationship and semantic logic; Combine the key information fragments into a complete sentence or phrase to form a candidate answer.
6. The method for generating a question-and-answer type according to claim 1, wherein Determine the difficulty coefficient of each question according to the query statement type and the answer to the query statement, including: Analyze the semantic content of the query statement to determine the query statement type; Determine the first difficulty value according to the query statement type; Obtain the information length, information type, and logical relationship characteristics between information of the answer corresponding to the query statement; Determine the second difficulty value according to the information length, information type, and logical relationship characteristics between information of the answer corresponding to the query statement; Perform weighting according to the first difficulty value, the second difficulty value, the preset difficulty weight of the query statement, and the preset answer difficulty weight to obtain the difficulty coefficient.
7. The method for generating a question-and-answer type according to claim 1, wherein Generate the question type of each question based on the difficulty coefficient of the question and the preset test paper question type difficulty ratio, including: Allocate the difficulty coefficient of the question to the corresponding question type according to the preset difficulty coefficient range value to obtain the preliminary question type of the test paper; Calculate the difficulty ratio in the preliminary question type of the test paper. If the difficulty ratio of the preliminary question type of the test paper meets the preset test paper question type difficulty ratio, then use the preliminary question type of the test paper as the final question type; If the difficulty ratio of the preliminary question type of the test paper does not meet the preset test paper question type difficulty ratio, then adjust the preliminary question type of the test paper until it meets the preset test paper question type difficulty ratio.
8. A question-and-answer type generation system, characterized in that, Applied to a method for generating a question-and-answer question type described in any one of claims 1-7, including a first extraction module, a knowledge graph construction module, a query statement generation module, a second extraction module, a retrieval module, an analysis module, an answer acquisition module, a difficulty coefficient determination module, and a question type generation module; The first extraction module is used to extract the first entity information, the first keyword information, the text summary information, and the information of the theme structure in the training document; The knowledge graph construction module is used to construct the training document into a structured knowledge graph according to the first entity information, the first keyword information, the text summary information, and the information of the theme structure; The query statement generation module is used to generate query statements for related questions based on the knowledge graph; The second extraction module is used to extract the second keyword information and the second entity information in the query statements of the related questions; The retrieval module is used to retrieve the text related to the query statement in the training document according to the second keyword information and the second entity information; The analysis module is used to analyze the dependency relationship between words in the text related to the query statement by using the dependency syntax technology to obtain candidate answers to the query statement; The answer acquisition module is used to calculate the semantic similarity between the candidate answer and the inquiry statement, and based on the similarity, obtain the answer to the inquiry statement; The difficulty coefficient determination module is used to determine the difficulty coefficient of each question according to the inquiry statement type and the answer to the inquiry statement; The question type generation module is used to generate the question type of each question based on the difficulty coefficient of the question and the preset difficulty ratio of the test paper question types, where the difficulty ratio refers to the expected proportion of question types with different difficulty coefficients in the entire test paper.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; Wherein, the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute a method for generating short answer question types as described in any one of claims 1-7.
10. A computer-readable storage medium stores a program, characterized in that, When the program is executed by the processor, it implements a method for generating short answer question types as described in any one of claims 1-7.