A method, device, equipment and computer-readable storage medium for generating a topic
By combining named entity recognition and natural language processing technology to generate the question stems and options, the problem of incorrect question generation in the existing technology is solved, and high-quality and diverse question generation is achieved.
Patent Information
- Application Number
- CN202510283454.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The correctness of the generated questions in the prior art is difficult to guarantee, and there are problems such as incorrect answers, indifferent answers, and inappropriate interference options.
By obtaining text materials containing knowledge points and pre-organized entities and entity labels, a named entity recognition model and natural language processing model are used for entity recognition, semantic analysis and part-of-speech analysis to generate the question stem and options for the question.
It improves the accuracy and quality of question generation, reduces artificial dependence, and can generate questions covering all aspects of life, avoiding repetition of questions on the market.
Smart Images

Figure CN119783635B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent question generation, and in particular to a question generation method, device, equipment and computer-readable storage medium. Background Art
[0002] There are many types of questions in the quiz game, and the number is large. Therefore, the questions generated by intelligent question generation must cover all aspects of life, and the generated questions must be of high quality, so as to meet the needs of users in the competition. At present, deep neural networks are generally used to batch produce questions. If a neural network is used directly, since the neural network model is a black box, it can only generate questions by learning certain rules through a large amount of data, but it is unknown how to achieve this. Therefore, the output of the model is uncontrolled, and there will be problems such as incorrect answers, non-unique answers, and inappropriate interference options.
[0003] Therefore, how to ensure the correctness of the final question generation is a technical problem that needs to be solved urgently. Summary of the invention
[0004] In view of this, an object of the present invention is to provide a question generating method, device, equipment and computer-readable storage medium, which solves the problem of incorrectly generated questions in the prior art.
[0005] In order to solve the above technical problems, the present invention provides a method for generating a topic, comprising:
[0006] Obtain text materials containing knowledge points and pre-organize the vocabulary formed by entities in various fields and corresponding entity labels;
[0007] Performing entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label;
[0008] Using a natural language processing model, and performing semantic analysis and part-of-speech analysis based on the recognition results in the text material, to generate a question stem;
[0009] Obtain interference items of the question stem from the vocabulary according to the target entity tag, and use the interference items and the target entity as options for the question.
[0010] Optionally, a named entity recognition model is used in combination with the vocabulary to perform entity recognition on the text material to obtain a recognition result, including:
[0011] Using the Jieba entity recognition model and combining the vocabulary to perform entity recognition on the text material, obtaining a first recognition result;
[0012] Performing entity recognition on the text material using the LCA entity recognition model in combination with the vocabulary to obtain a second recognition result;
[0013] An overlapping portion of the first recognition result and the second recognition result is used as the recognition result.
[0014] Optionally, after performing entity recognition on the text material by using the named entity recognition model in combination with the vocabulary to obtain a recognition result, the method further includes:
[0015] Performing entity label correction on the target entity label through the trained FastText model to obtain a corrected entity label; the corrected entity label and the target entity together constitute a corrected recognition result;
[0016] Accordingly, a natural language processing model is used to perform semantic analysis and part-of-speech analysis based on the recognition results in the text material to generate the stem of the question, including:
[0017] A natural language processing model is used to perform semantic analysis and part-of-speech analysis on the corrected recognition result to generate the stem of the question.
[0018] Optionally, a natural language processing model is used to perform semantic analysis and part-of-speech analysis based on the recognition result in the text material to generate the stem of the question, including:
[0019] The BERT model is used to understand the context of the recognition result. If the target entity is associated with other entities, and the word order of the question stem is judged to be reasonable according to the logical order of the parts of speech in the text material, the target entity in the text material is hollowed out to obtain a fill-in-the-blank question stem;
[0020] The Transform model is used to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem of the question.
[0021] Optionally, using a Transform model to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem of the question includes:
[0022] Using the Transform model to perform natural language processing on the fill-in-the-blank question stem to obtain an optimized question stem;
[0023] Analyze the optimized question stem using the LSTM-based question stem evaluation model to obtain question stem parameter values;
[0024] If the question stem parameter value is greater than a preset question stem threshold, the optimized question stem is used as the question stem of the question;
[0025] Otherwise, the fill-in-the-blank question stem will be used as the question stem of the question.
[0026] Optionally, after obtaining the distractor items of the question stem from the vocabulary according to the target entity tag, the method further includes:
[0027] The GAN model is used to diversify the interference items of the question stem to obtain multiple options;
[0028] Accordingly, the distractor and the target entity are used as options for the question;
[0029] The multiple options and the target entity are used as options for the question.
[0030] Optionally, using the multiple options and the target entity as options for the question includes:
[0031] Analyze the multiple options using the LSTM-based interference item evaluation model, sort the options according to the parameter values of the options, and obtain a preset number of target options according to the sorting results;
[0032] The target option and the target entity are used as options of the topic.
[0033] The present invention also provides a method for generating a topic, comprising:
[0034] The text material and vocabulary acquisition module is used to acquire text materials containing knowledge points and pre-organize the vocabulary formed by entities in various fields and corresponding entity tags;
[0035] An entity recognition module, used to perform entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label;
[0036] A question stem generation module, used to generate the question stem by using a natural language processing model and performing semantic analysis and part-of-speech analysis according to the recognition results in the text material;
[0037] An option generation module is used to obtain interference items of the question stem from the vocabulary according to the target entity label, and use the interference items and the target entity as options of the question.
[0038] The present invention also provides a topic generating device, comprising:
[0039] Memory for storing computer programs;
[0040] A processor is used to implement the steps of the above-mentioned question generation method when executing the computer program.
[0041] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the steps of the above-mentioned question generating method are implemented.
[0042] It can be seen that the present invention obtains text materials containing knowledge points and pre-sorts entities in various fields and corresponding entity labels to form a vocabulary; uses a named entity recognition model and combines the vocabulary to perform entity recognition on the text material to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label; uses a natural language processing model and performs semantic analysis and part-of-speech analysis based on the recognition results in the text material to generate the stem of the question; obtains interference items of the stem from the vocabulary according to the target entity label, and uses the interference items and the target entity as options for the question. The present invention obtains recognition results by adopting named entity recognition technology and a custom vocabulary, and then performs semantic and part-of-speech analysis through semantic understanding technology to generate the stem of the question; obtains question options based on the vocabulary and recognition results, thereby generating a large number of questions. Since the present invention decomposes the generation of questions into multiple steps, each step can achieve the purpose of question accuracy by training and optimizing different models, thereby ensuring the accuracy of question generation. In addition, the overall architecture is based on natural language processing technology and other neural network algorithms, which reduces the degree of dependence on manual labor in the process of question production. Only by collecting appropriate knowledge materials, a large number of questions can be produced through the above method.
[0043] In addition, the present invention also provides a question generating device, equipment and computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0045] Figure 1 A flowchart of a method for generating a topic provided by an embodiment of the present invention;
[0046] Figure 2 An example flow chart of a method for determining a recognition result provided by an embodiment of the present invention;
[0047] Figure 3 A flowchart of a method for generating a topic provided by an embodiment of the present invention;
[0048] Figure 4 A flowchart of a method for optimizing a topic provided by an embodiment of the present invention;
[0049] Figure 5 A flowchart of another method for generating a topic provided by an embodiment of the present invention;
[0050] Figure 6 A schematic diagram of the structure of a question generating device provided by an embodiment of the present invention;
[0051] Figure 7 A schematic diagram of the structure of a question generating device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] First, some terms used in this application are explained:
[0054] Jieba entity recognition model: Jieba is a Python library for Chinese word segmentation, and also supports Chinese named entity recognition;
[0055] LAC entity recognition model: It is a joint lexical analysis model, which aims to complete Chinese word segmentation, part-of-speech tagging, and proper name recognition tasks in an integrated manner;
[0056] FastText model: is a fast text classification model;
[0057] BERT model: is a typical pre-trained language model using bidirectional encoding;
[0058] Transform model: A deep learning model that is mainly used to process sequence data, especially in the field of natural language processing;
[0059] LSTM: Long Short-Term Memory, a time-recurrent neural network;
[0060] GAN model: Generative adversarial network is a deep learning model.
[0061] The present invention mainly solves the requirements of question categories and quantity in the question answering game by intelligently generating multiple-choice questions. The intelligent question generation requires that the generated questions can cover all aspects of life, and at the same time requires that the generated questions have high quality and large quantity, so as to meet the questions required by the user competition in the game. At present, the solution of directly using neural network cannot guarantee the quality of the questions, including the problems of incoherent and unreasonable expression of the question stem and inappropriate matching of options. Since the neural network model is a black box, it can only learn certain rules to generate questions through a large amount of data, but how to achieve it specifically is unknown, so the output result of the model is uncontrolled; in addition, the accuracy of the questions is difficult to be guaranteed. The logical completeness of a suitable question is very strong. The generation of multiple-choice questions by neural network technology alone will have problems such as incorrect answers, non-unique answers, and inappropriate interference options.
[0062] In the actual process of generating questions, the present invention mainly uses technologies such as named entity recognition, semantic analysis, and text generation. Please refer to Figure 1 , Figure 1 A flowchart of a method for generating a topic provided by an embodiment of the present invention. The method may include:
[0063] S101: Acquire text materials containing knowledge points and a vocabulary formed by pre-organizing entities in various fields and corresponding entity tags.
[0064] Specifically, step S101 is a data preparation step. Before the process of generating questions is started, entities and corresponding labels in various fields need to be collected and sorted to form a vocabulary. The finer the field division, the higher the quality of question generation. At the same time, short text materials containing knowledge points need to be collected as input for the entire question generation process.
[0065] S102: Perform entity recognition on the text material using a named entity recognition model in combination with a vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label.
[0066] This embodiment does not limit the named entity recognition model, nor does it limit the number of named entity models. Step S102 is the entity recognition part. When the text material is input into the named entity recognition model, this embodiment also loads a pre-organized vocabulary to constrain the named entity model, so as to obtain accurate recognition results.
[0067] Furthermore, in order to improve the accuracy of entity result recognition, the above-mentioned named entity recognition model is used in combination with the vocabulary to perform entity recognition on the text material to obtain the recognition result, which can specifically include the following steps:
[0068] Step 21: Use the Jieba entity recognition model and the vocabulary to perform entity recognition on the text material to obtain a first recognition result;
[0069] Step 22: Use the LCA entity recognition model and the vocabulary to perform entity recognition on the text material to obtain a second recognition result;
[0070] Step 23: taking the overlapping part of the first recognition result and the second recognition result as the recognition result.
[0071] This embodiment uses two entity recognition tools, specifically the Jieba entity recognition model and the LAC entity recognition model, and combines them with a pre-organized vocabulary for double recognition. Finally, the intersection of the two entity recognition models is selected as the final recognition result, reducing the probability of recognition errors by a single entity recognition model.
[0072] Further, the recognition result is corrected to improve the accuracy of the recognition result. After the named entity recognition model is used in combination with the vocabulary to perform entity recognition on the text material and obtain the recognition result, the following steps may be specifically included:
[0073] The target entity label is corrected by the trained FastText model to obtain a corrected entity label; the corrected entity label and the target entity together constitute the corrected recognition result;
[0074] Correspondingly, step S103 specifically utilizes a natural language processing model to perform semantic analysis and part-of-speech analysis on the corrected recognition result to generate the stem of the question.
[0075] This embodiment takes into account that different entities belong to different entity labels in different semantic environments, that is, different entity categories. For example, "Li Bai" belongs to the category of "poet" in historical literature materials, and "Li Bai" belongs to the category of "game character" in game materials. Relying solely on entity recognition may result in errors in the recognition of the categories to which some entities belong. Therefore, this embodiment trains an entity category corrector (predictor) through the FastText model. The input of the FastText model is the encoding vector composed of the material text and the entity. The output of the FastText model is the probability of belonging to each category. The category with the highest probability is selected as the final entity category, that is, the entity label. Through this step, the target entity and target entity label corresponding to each knowledge material will be obtained. The detailed process can be referred to Figure 2 , Figure 2 A flowchart of a method for determining a recognition result provided by an embodiment of the present invention.
[0076] S103: Using a natural language processing model, and performing semantic analysis and part-of-speech analysis based on the recognition results in the text material, a question stem is generated.
[0077] This embodiment does not limit the natural language processing model. Simple semantic analysis and part-of-speech analysis are performed based on the results of entity recognition in the text material to determine whether the recognized entity can be used as an answer to generate a question stem.
[0078] Furthermore, in order to improve the accuracy of question stem generation, the above-mentioned natural language processing model is used to perform semantic analysis and part-of-speech analysis based on the recognition results in the text material to generate the question stem, which may specifically include the following steps:
[0079] Step 31: Use the BERT model to understand the context of the recognition results. If the target entity is related to other entities, and the word order of the question stem is reasonable based on the logical order of the parts of speech in the text material, then the target entity in the text material is hollowed out to obtain a fill-in-the-blank question stem.
[0080] The semantic analysis in this embodiment mainly uses the BERT model to understand the context and determine whether the currently identified target entity has a certain relationship with other entities. If so, the word order of the generated question stem will be determined based on the logical order of the parts of speech in the text material. If all the above conditions are met, the target entity in the text material will be replaced with a horizontal line to generate a fill-in-the-blank question stem. Figure 3 The left half shows Figure 3 A flowchart of a method for generating a topic provided in an embodiment of the present invention.
[0081] Step 32: Use the Transform model to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem.
[0082] This embodiment also uses the Transform model to optimize the fill-in-the-blank question stem in natural language, and generates a question sentence in accordance with natural language rules through the trained Transformer model, such as: input: "The author of "Quiet Night Thoughts" is ___.", output: "Which poet is the author of "Quiet Night Thoughts"? ".
[0083] Furthermore, in order to ensure the accuracy of question stem generation, the Transform model is used to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem, which may specifically include:
[0084] Step 321: Use the Transform model to perform natural language processing on the fill-in-the-blank question stem to obtain an optimized question stem;
[0085] Step 322: Analyze the optimized question stem using the LSTM-based question stem evaluation model to obtain question stem parameter values;
[0086] Step 323: If the stem parameter value is greater than the preset stem threshold, the optimized stem is used as the stem of the question;
[0087] Step 324: Otherwise, the fill-in-the-blank question stem is used as the question stem.
[0088] This embodiment further tests the optimization results of the Transform model. Since the results generated by the Transform model are not completely accurate, an evaluation model is added after the results generated by the Transform model to score the optimization results. The model is trained using LSTM, which can generate a value between [0,1] by inputting the generated question stem and answer. The closer the value is to 1, the better the generated effect is. Therefore, a threshold is set to decide whether to keep the optimized question stem. The optimization of the question stem can refer to the left half of the figure. Figure 4 A flowchart of a method for optimizing a topic provided in an embodiment of the present invention.
[0089] S104: Obtain distractors of the question stem from the vocabulary according to the target entity tag, and use the distractors and the target entity as options for the question.
[0090] This embodiment does not limit the specific method of obtaining the stem interference items. For example, a clustering algorithm can be used; or other algorithms for calculating similarity can be selected. The method of generating options is as follows: Figure 3 The right half shows, Figure 3 A flowchart of a method for generating a topic provided in an embodiment of the present invention.
[0091] Furthermore, in order to generate distractors in a diverse and reasonable manner, after obtaining distractors of the question stem from the vocabulary according to the target entity label, the following steps may be further included:
[0092] Use the GAN model to diversify the interference items in the question stem and obtain multiple options;
[0093] Accordingly, distractors and target entities are used as options in the questions;
[0094] Include multiple choices and target entities as options for the question.
[0095] The optimization of options in this embodiment must ensure the diversity and rationality of the options. First, consider the diversity of options. Based on the option data in the existing question bank, a generative adversarial network is trained to generate options, and the first three generated options are automatically added to the initial option list. That is, the GAN model is used to diversify the interference items in the question stem to obtain multiple options. For details, please refer to Figure 4 The right half of Figure 4 A flowchart of a method for optimizing a topic provided in an embodiment of the present invention.
[0096] Furthermore, in order to ensure the accuracy and reliability of the generation of distractors, the above-mentioned method of using multiple options and target entities as options for the question may specifically include the following steps:
[0097] Step 41: Analyze the multiple options using the LSTM-based interference item evaluation model, sort the options according to the parameter values of the options, and obtain a preset number of target options according to the sorting results;
[0098] Step 42: Use the target options and target entities as options for the question.
[0099] This example uses LSTM to train a distractor evaluation model based on the question stem and distractor option data in the existing question bank, scores and sorts the initial multiple options, and finally selects them in order according to the required number of options. The purpose of this step is to select the distractor that best matches the current question stem from the initial option set through the evaluation model. For the optimization of options, please refer to Figure 4 The right half of Figure 4 A flowchart of a method for optimizing a topic provided in an embodiment of the present invention.
[0100] Furthermore, for the final generated multiple-choice questions, the specified large language model will be called to perform the final correctness verification. By comparing the answer results of the large language model with the correct answer to the current question, it is determined whether further manual verification is needed.
[0101] The method for generating a question provided by the embodiment of the present invention is applied, by obtaining text materials containing knowledge points and pre-organizing entities in various fields and corresponding entity labels to form a vocabulary; using a named entity recognition model and combining the vocabulary to perform entity recognition on the text materials to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label; using a natural language processing model, and performing semantic analysis and part-of-speech analysis according to the recognition result in the text material, to generate the stem of the question; obtaining interference items of the stem from the vocabulary according to the target entity label, and using the interference items and the target entity as options for the question. The present invention decomposes the generation of questions into multiple processes, each of which can achieve the purpose of question accuracy by training and optimizing different models, so that the accuracy of the generated questions is higher. In addition, the overall architecture is based on natural language processing technology and other neural network algorithms, such as sequence network models, generative adversarial networks, and diffusion generation models, which reduces the degree of dependence on manual labor in the process of producing questions. Only suitable knowledge materials need to be collected to produce a large number of questions through the above method; in addition, the questions after the stem and options are optimized ensure the diversity of the questions, and also avoid the duplication with a large number of questions on the market; finally, the large language model is called for machine proofreading, which saves manpower and improves efficiency.
[0102] In order to make the present invention easier to understand, please refer to Figure 5 , Figure 5 Another example flow chart of a method for generating a topic provided by an embodiment of the present invention may specifically include:
[0103] (1) Data preparation: natural language text materials containing specific knowledge points and classification thesaurus covering various fields.
[0104] (2) Entity recognition: Identify specific entities contained in the material; identify specific categories of entities based on the text material description.
[0105] (3) Generate knockout multiple-choice questions: Question stem: Use semantic analysis and part-of-speech analysis to determine the entities that can be used to ask questions and replace them with underscores; Options: Based on the current entity, use a clustering algorithm to obtain other entities of the same category as the initial interference items.
[0106] (4) Question stem optimization: Based on the trained natural language model, the generated hollowed-out question stem is converted into a question-like question. The generated question is then judged by a scoring model (i.e., evaluation model) in terms of grammar and fluency to decide whether to retain the optimized result.
[0107] (5) Question option optimization: Generate a batch of distractors outside the vocabulary based on generative adversarial networks, diffusion models, etc., and finally select the distractors that best match the current question based on the trained option scoring model.
[0108] (6) Question generation is completed: The generated questions are verified for accuracy through a large language model.
[0109] The following is an introduction to a question generating device provided in an embodiment of the present invention. The question generating device described below and the question generating method described above can be referenced to each other.
[0110] Please refer to Figure 6 , Figure 6 A schematic diagram of a structure of a question generating device provided by an embodiment of the present invention may include:
[0111] The text material and vocabulary acquisition module 100 is used to acquire text materials containing knowledge points and a vocabulary formed by pre-organizing entities in various fields and corresponding entity tags;
[0112] An entity recognition module 200 is used to perform entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label;
[0113] A question stem generation module 300 is used to generate a question stem by using a natural language processing model and performing semantic analysis and part-of-speech analysis according to the recognition result in the text material;
[0114] The option generation module 400 is used to obtain interference items of the question stem from the vocabulary according to the target entity tag, and use the interference items and the target entity as options of the question.
[0115] Based on the above embodiment, the entity identification module 200 may include:
[0116] A first recognition unit, configured to perform entity recognition on the text material by using the Jieba entity recognition model in combination with the vocabulary to obtain a first recognition result;
[0117] A second recognition unit is used to perform entity recognition on the text material by using the LCA entity recognition model in combination with the vocabulary to obtain a second recognition result;
[0118] The recognition result determining unit is used to use the overlapping part of the first recognition result and the second recognition result as the recognition result.
[0119] Based on the above embodiment, the question generating device may further include:
[0120] An entity label correction module is used to perform entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result, and then perform entity label correction on the target entity label using a trained FastText model to obtain a corrected entity label; the corrected entity label and the target entity together constitute a corrected recognition result;
[0121] Accordingly, the question stem generation module 300 may include:
[0122] The question stem generation unit is used to use a natural language processing model to perform semantic analysis and part-of-speech analysis on the corrected recognition result to generate the question stem of the question.
[0123] Based on any of the above embodiments, the question stem generation module 300 may include:
[0124] A fill-in-the-blank question stem generating unit is used to use a BERT model to understand the context of the recognition result. If the target entity is associated with other entities, and the word order of the question stem is judged to be reasonable according to the logical order of the parts of speech in the text material, the target entity in the text material is hollowed out to obtain a fill-in-the-blank question stem;
[0125] The natural language processing unit is used to perform natural language processing on the fill-in-the-blank question stem using a Transform model to obtain the question stem of the question.
[0126] Based on the above embodiment, the natural language processing unit may include:
[0127] A stem optimization subunit, used for performing natural language processing on the fill-in-the-blank question stem using the Transform model to obtain an optimized question stem;
[0128] A question stem evaluation subunit, used to analyze the optimized question stem using an LSTM-based question stem evaluation model to obtain a question stem parameter value;
[0129] A first result subunit is used to use the optimized stem as the stem of the question if the stem parameter value is greater than a preset stem threshold;
[0130] The second result sub-unit is used to otherwise use the fill-in-the-blank question stem as the question stem of the question.
[0131] Based on the above embodiment, the question generating device may further include:
[0132] A diversification processing module, configured to obtain interference items of the question stem from the vocabulary according to the target entity label, and then perform diversification processing on the interference items of the question stem using a GAN model to obtain multiple options;
[0133] Accordingly, the option generation module 400 may include:
[0134] An option generating unit is used to use the multiple options and the target entity as options for the question.
[0135] Based on the above embodiment, the option generating unit may include:
[0136] An option evaluation subunit is used to analyze the multiple options using an LSTM-based interference item evaluation model, sort the options according to the parameter values of the options, and obtain a preset number of target options according to the sorting results;
[0137] The option determination subunit is used to use the target option and the target entity as options for the topic.
[0138] It should be noted that the order of the modules and units in the above-mentioned question generating device can be changed without affecting the logic.
[0139] The question generation device provided by the embodiment of the present invention is applied, through the text material and vocabulary acquisition module 100, used to obtain text materials containing knowledge points and a vocabulary formed by pre-organizing entities in various fields and corresponding entity tags; the entity recognition module 200 is used to use the named entity recognition model and the vocabulary to perform entity recognition on the text material to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity tag; the question stem generation module 300 is used to use a natural language processing model and perform semantic analysis and part-of-speech analysis based on the recognition result in the text material to generate the question stem; the option generation module 400 is used to obtain interference items of the question stem from the vocabulary according to the target entity tag, and use the interference items and the target entity as options of the question. Since the present device decomposes the question generation into multiple steps, each step can achieve the purpose of question accuracy by training and optimizing different models, thereby ensuring the accuracy of question generation. In addition, the overall architecture is based on natural language processing technology and other neural network algorithms, such as sequence network models, generative adversarial networks, and diffusion generation models, which reduce the degree of dependence on manual labor in the question generation process. Only by collecting appropriate knowledge materials can a large number of questions be generated through the above-mentioned device; in addition, the questions are optimized through the stems and options to ensure the diversity of the questions and avoid duplication with a large number of questions on the market; finally, by calling a large language model for machine proofreading, manpower is saved and efficiency is improved.
[0140] The following is an introduction to a topic generating device provided by an embodiment of the present invention. The topic generating device described below and the topic generating method described above can be referred to each other.
[0141] Please refer to Figure 7 , Figure 7 A schematic diagram of a structure of a question generating device provided by an embodiment of the present invention may include:
[0142] A memory 10, used for storing computer programs;
[0143] The processor 20 is used to execute a computer program to implement the above-mentioned question generating method.
[0144] The memory 10 , the processor 20 , and the communication interface 31 all communicate with each other via the communication bus 32 .
[0145] In the embodiment of the present invention, the memory 10 is used to store one or more programs, and the program may include program code, and the program code includes computer operation instructions. In the embodiment of the present invention, the memory 10 may store programs for implementing the following functions:
[0146] The word library is formed by obtaining text materials containing knowledge points and pre-organizing entities in various fields and corresponding entity tags;
[0147] Use the named entity recognition model and the vocabulary to perform entity recognition on the text material to obtain the recognition result; the recognition result includes the target entity and the corresponding target entity label;
[0148] Using the natural language processing model, semantic analysis and part-of-speech analysis are performed based on the recognition results in the text material to generate the question stem;
[0149] According to the target entity label, the distractors of the question stem are obtained from the vocabulary, and the distractors and the target entity are used as options for the question.
[0150] In a possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function, etc.; the data storage area may store data created during use.
[0151] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include an NVRAM. The memory stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0152] The processor 20 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic device, a microprocessor or any conventional processor, etc. The processor 20 may call a program stored in the memory 10 .
[0153] The communication interface 31 may be an interface of a communication module, and is used to connect to other devices or systems.
[0154] Of course, it should be noted that Figure 7 The structure shown does not constitute a limitation on the topic generating device in the embodiment of the present invention. In actual applications, the topic generating device may include Figure 7 More or fewer components than shown, or combinations of certain components.
[0155] The computer-readable storage medium provided in an embodiment of the present invention is introduced below. The computer-readable storage medium described below and the question generating method described above can be referenced to each other.
[0156] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned question generating method are implemented.
[0157] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0158] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0160] Finally, it should be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0161] The above is a detailed introduction to a method, device, equipment and computer-readable storage medium for generating questions provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for generating a topic, characterized in that: include: Obtain text materials containing knowledge points and pre-organize the vocabulary formed by entities in various fields and corresponding entity labels; Performing entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label; Using a natural language processing model, and performing semantic analysis and part-of-speech analysis based on the recognition results in the text material, to generate a question stem; Obtain interference items of the question stem from the vocabulary according to the target entity tag, and use the interference items and the target entity as options for the question; The named entity recognition model is used in combination with the vocabulary to perform entity recognition on the text material to obtain recognition results, including: Using the Jieba entity recognition model and combining the vocabulary to perform entity recognition on the text material, obtaining a first recognition result; Performing entity recognition on the text material using the LCA entity recognition model in combination with the vocabulary to obtain a second recognition result; taking the overlapping part of the first recognition result and the second recognition result as the recognition result; After performing entity recognition on the text material using the named entity recognition model in combination with the thesaurus to obtain a recognition result, the method further includes: Performing entity label correction on the target entity label through the trained FastText model to obtain a corrected entity label; the corrected entity label and the target entity together constitute a corrected recognition result; Accordingly, a natural language processing model is used to perform semantic analysis and part-of-speech analysis based on the recognition results in the text material to generate the stem of the question, including: A natural language processing model is used to perform semantic analysis and part-of-speech analysis on the corrected recognition result to generate the stem of the question.
2. The topic generation method according to claim 1, characterized in that: Using a natural language processing model, and performing semantic analysis and part-of-speech analysis based on the recognition results in the text material, the stem of the question is generated, including: The BERT model is used to understand the context of the recognition result. If the target entity is associated with other entities, and the word order of the question stem is judged to be reasonable according to the logical order of the parts of speech in the text material, the target entity in the text material is hollowed out to obtain a fill-in-the-blank question stem; The Transformer model is used to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem of the question.
3. The topic generation method according to claim 2, characterized in that: The Transformer model is used to perform natural language processing on the fill-in-the-blank question stem to obtain the question stem of the question, including: Using the Transformer model to perform natural language processing on the fill-in-the-blank question stem to obtain an optimized question stem; Analyze the optimized question stem using the LSTM-based question stem evaluation model to obtain question stem parameter values; If the question stem parameter value is greater than a preset question stem threshold, the optimized question stem is used as the question stem of the question; Otherwise, the fill-in-the-blank question stem will be used as the question stem of the question.
4. The method for generating a topic according to claim 1, characterized in that: After obtaining the distractor items of the question stem from the vocabulary according to the target entity tag, the method further includes: The GAN model is used to diversify the interference items of the question stem to obtain multiple options; Accordingly, the distractor and the target entity are used as options for the question; The multiple options and the target entity are used as options for the question.
5. The topic generation method according to claim 4, characterized in that: Using the multiple options and the target entity as options for the topic includes: Analyze the multiple options using the LSTM-based interference item evaluation model, sort the options according to the parameter values of the options, and obtain a preset number of target options according to the sorting results; The target option and the target entity are used as options of the topic.
6. A question generating device, characterized in that: include: The text material and vocabulary acquisition module is used to acquire text materials containing knowledge points and pre-organize the vocabulary formed by entities in various fields and corresponding entity tags; An entity recognition module, used to perform entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result; the recognition result includes a target entity and a corresponding target entity label; A question stem generation module, used to generate the question stem by using a natural language processing model and performing semantic analysis and part-of-speech analysis according to the recognition results in the text material; An option generation module, used to obtain interference items of the question stem from the vocabulary according to the target entity tag, and use the interference items and the target entity as options for the question; The entity recognition module comprises: A first recognition unit, configured to perform entity recognition on the text material by using the Jieba entity recognition model in combination with the vocabulary to obtain a first recognition result; A second recognition unit is used to perform entity recognition on the text material by using the LCA entity recognition model in combination with the vocabulary to obtain a second recognition result; a recognition result determination unit, configured to use an overlapping portion of the first recognition result and the second recognition result as the recognition result; Also includes: An entity label correction module is used to perform entity recognition on the text material using a named entity recognition model in combination with the vocabulary to obtain a recognition result, and then perform entity label correction on the target entity label using a trained FastText model to obtain a corrected entity label; the corrected entity label and the target entity together constitute a corrected recognition result; Correspondingly, the question stem generation module includes: The question stem generation unit is used to use a natural language processing model to perform semantic analysis and part-of-speech analysis on the corrected recognition result to generate the question stem of the question.
7. A question generating device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the question generating method as claimed in any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the steps of the question generating method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Choice question generation model training method, choice question generation method, equipment and medium
CN112560443A
Chinese choice question interference item generation method based on free text
CN112686025A