Topic generation method, medium, computer equipment and program product

By generating questions using multimodal language models and text-based language models, this technology solves the problem of limited question materials in existing technologies, achieving high-quality and diverse question generation, and is suitable for education, online examinations, and scientific research.

CN120996049APending Publication Date: 2025-11-21ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511214407.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, question generation mainly relies on text modality, resulting in poor material richness and limited information, which cannot meet the complex and diverse question-generating needs. In particular, it is difficult to generate high-quality and accurately adapted questions in the fields of education, online examinations and scientific research.

Method used

Non-textual information is extracted using a multimodal language model, and the relationships between knowledge points are determined by combining this with a text-based language model to generate questions.

Benefits of technology

It has enriched the sources of materials and generation ideas for questions, improved the diversity and quality of questions, met the personalized and complex question generation needs, reduced labor costs, and improved generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996049A_ABST
    Figure CN120996049A_ABST
Patent Text Reader

Abstract

The invention discloses a question generation method, a medium, computer equipment and a program product. The method comprises the following steps: acquiring input data; the input data comprises non-text information; performing knowledge point extraction on the input data through a multi-modal language model to obtain text description information of a plurality of knowledge points in the input data; inputting the text description information of the plurality of knowledge points into a text-based language model, so that the text-based language model determines an association relationship among the plurality of knowledge points based on the text description information of the plurality of knowledge points; and inputting the plurality of knowledge points and the association relationship among the plurality of knowledge points into a question generation model, so that the question generation model generates a question based on the plurality of knowledge points and the association relationship among the plurality of knowledge points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a question generation method, medium, computer equipment, and program product. Background Technology

[0002] In today's era, there is an urgent need for a large number of high-quality questions that are accurately tailored to various needs across multiple fields, including education, online testing, and scientific research. Traditional methods primarily rely on text-based data to generate questions, resulting in limited material richness and a small amount of information that can be extracted from single-modality data. This leads to a limited range of question types and perspectives, failing to meet the complex and diverse needs of real-world question creation. Summary of the Invention

[0003] Firstly, embodiments of this specification provide a question generation method, the method comprising:

[0004] Obtain input data; the input data includes non-text information;

[0005] The input data is processed by a multimodal language model to extract knowledge points, thereby obtaining textual descriptions of multiple knowledge points in the input data.

[0006] The textual description information of the multiple knowledge points is input into a text-based language model so that the text-based language model can determine the relationship between the multiple knowledge points based on the textual description information of the multiple knowledge points.

[0007] The multiple knowledge points and their relationships are input into the question generation model so that the question generation model generates questions based on the multiple knowledge points and their relationships.

[0008] Secondly, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any embodiment of this specification.

[0009] Thirdly, embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the methods described in any embodiment of this specification.

[0010] Fourthly, embodiments of this specification provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any embodiment of this specification.

[0011] In the embodiments of this specification, a multimodal language model is used to extract knowledge points from input data including non-textual information, obtaining textual descriptions of multiple knowledge points in the input data. This aligns the non-textual information to a textual modality, allowing non-textual information that is not easily utilized by traditional text processing methods to be integrated into the question generation process. Based on this, a text-based language model analyzes these textually described knowledge points to determine the relationships between them, thereby enriching the knowledge base for question generation. These relationships between knowledge points can provide questions with more complex logical structures and diverse examination perspectives. After inputting the knowledge points and their relationships into the question generation model, the model can generate questions based on this information. This method utilizes non-textual information to generate questions, broadening the sources of question materials and generation approaches, and effectively improving the diversity of generated questions.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this specification and, together with the specification, serve to explain the technical solutions described herein.

[0014] Figure 1 This is a flowchart of the title generation method in an embodiment of this specification.

[0015] Figure 2 This is a schematic diagram of the knowledge graph in an embodiment of this specification.

[0016] Figure 3 This is a general flowchart of an embodiment of this specification.

[0017] Figure 4 This is a block diagram of a title generation apparatus according to an embodiment of this specification.

[0018] Figure 5 This is a schematic diagram of a computer device according to an embodiment of this specification. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0020] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Additionally, the term “at least one” herein means any combination of at least two of any one or more of a plurality.

[0021] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0022] To enable those skilled in the art to better understand the technical solutions in the embodiments of this specification, and to make the above-mentioned objectives, features and advantages of the embodiments of this specification more apparent and understandable, the technical solutions in the embodiments of this specification will be further described in detail below with reference to the accompanying drawings.

[0023] In today's era, education, online testing, and scientific research all urgently require a large number of high-quality and precisely tailored questions. These needs are primarily concentrated in the following areas: In education, with the advancement of smart education, students' demand for personalized learning resources is increasing, and challenging questions with broad knowledge coverage can help students better tap into their potential; in online testing, challenging questions are an important tool for assessing candidates' abilities; in scientific research, challenging questions, as "challenging tasks," can help models learn complex reasoning logic and improve reasoning depth, especially in fields such as natural language processing, mathematical reasoning, and multimodal AI, where high-quality challenging data is crucial for improving model performance. Traditional question generation methods mainly rely on manual methods, which are not only time-consuming but also difficult to meet the growing demand for personalization. With the development of artificial intelligence technology, using AI to assist in automatic question generation has become an important way to improve efficiency and quality. However, in the traditional model, questions are mainly generated based on text-based data, resulting in poor richness of generated materials and limited information that can be extracted from single-modality data. This leads to a relatively limited range of question types and perspectives, failing to meet the complex and diverse needs of actual question creation. For example, some geometry problems require images of geometric figures in their stems. However, related technologies generate problems based solely on text-based information, making it difficult to clearly describe the specific graphic features and positional changes of a geometric figure after transformations such as rotation, translation, and scaling. Consequently, it is difficult to design problems from multiple perspectives, including different visual angles, graphic constituent elements, and interactions between figures.

[0024] Based on this, embodiments of this specification provide a question generation method that aligns non-text information to a text modality, allowing non-text information that is not easily utilized by traditional text processing methods to be integrated into the question generation process. The implementation details of these embodiments are described below with reference to the accompanying drawings.

[0025] like Figure 1 As shown, the question generation method in the embodiments of this specification includes:

[0026] Step S12: Obtain input data; the input data includes non-text information;

[0027] Step S14: Extract knowledge points from the input data using a multimodal language model to obtain textual descriptions of multiple knowledge points in the input data;

[0028] Step S16: Input the text description information of multiple knowledge points into the text-based language model so that the text-based language model can determine the relationship between multiple knowledge points based on the text description information of multiple knowledge points;

[0029] Step S18: Input multiple knowledge points and the relationships between them into the question generation model so that the question generation model can generate questions based on the multiple knowledge points and the relationships between them.

[0030] In step S12, input data can be acquired. This input data can be seed questions, textbooks, subject knowledge system outlines, or other types of data. Seed questions can be pre-generated questions through manual processing or language models, serving as a benchmark and reference for question generation. The input data can include multiple knowledge points. A knowledge point refers to a basic unit in a subject that has a certain logical relationship and can systematically express one or more concepts, principles, methods, etc. It is the smallest unit of a knowledge system. For example, in mathematics, knowledge points include the definition and properties of functions (such as the graphical characteristics and changing patterns of linear and quadratic functions), the basic properties of geometric figures (such as the sum of the interior angles of a triangle, the Pythagorean theorem, etc.), and the rules of algebraic operations (such as factorization, methods for solving equations), etc. In physics, knowledge points include Newton's laws of motion in mechanics (the relationship between changes in the state of motion of an object and the forces acting on it), Ohm's law in electricity (the relationship between current, voltage, and resistance), and the laws of reflection and refraction of light in optics, etc.

[0031] Input data includes non-text information, such as images, videos, and / or audio. Additionally, input data may include text information. Non-text information provides non-textual material for the question generation process, resulting in generated questions containing non-textual content. For example, audio files can be used as material for listening comprehension questions, generating questions that include audio content; images can be used as material for geometry questions, generating questions that include images. Text information provides textual material for the question generation process, resulting in generated questions containing textual content. For example, an article can be used as material for reading comprehension questions, generating questions that include textual content. When input data includes both text and non-text information, the generated questions can include both textual and non-textual content simultaneously. For example, a geometry question might include an image describing a geometric figure and text describing the problem.

[0032] In step S14, a multimodal language model can be used to extract knowledge points from the input data, obtaining textual descriptions of multiple knowledge points within the input data. In this way, non-textual information in the input data can be aligned to the text modality for subsequent processing. For example, assuming the input data includes an image containing an equilateral triangle, the following textual description can be output: "The sum of the three interior angles of the triangle is 180 degrees."

[0033] The type of multimodal language model can be determined based on the modality to which the information in the input data belongs. For example, if the input data includes information from the image modality, the multimodal language model can be a Vision-Language Model (VLM); as another example, when the input data contains information from the audio modality, such as speech or music, the multimodal language model can be an Audio-Language Model (ALM).

[0034] In some embodiments, input data and first prompt information can be input into a multimodal language model, so that the multimodal language model outputs a knowledge point extraction strategy based on the first prompt information to extract knowledge points from the input data. The first prompt information includes guidance information to guide the multimodal language model in outputting the knowledge point extraction strategy. Then, a second prompt information is generated based on the knowledge point extraction strategy, and this second prompt information is input into the multimodal language model, so that the multimodal language model uses the knowledge point extraction strategy included in the second prompt information to extract knowledge points from the input data.

[0035] This embodiment employs a two-stage approach for knowledge point extraction. The first stage generates a knowledge point extraction strategy, clarifying the key directions, essential content, and corresponding rules for extraction. This ensures that the second stage of the knowledge point extraction process can accurately pinpoint the core knowledge points. For example, the knowledge point extraction strategy may include, but is not limited to, at least the following: determining the subject to which the knowledge points in the input data belong; analyzing the specific content of the known conditions in the input data (such as graphic features, annotation information, implicit logic, etc.) and their impact on the solution to the questions; proposing possible question directions or question types (such as calculation problems, reasoning problems, concept comprehension problems, etc.); and explaining the relationship between the proposed questions and the known conditions in the input data. This knowledge point extraction strategy provides a clear framework to guide the knowledge point extraction process.

[0036] In some embodiments, after extracting knowledge points from the input data, these knowledge points can be expanded to increase their richness, thereby enhancing the diversity and richness of subsequently generated questions. For example, a knowledge base including knowledge points can be established, and the knowledge points extracted from the input data can be expanded based on this knowledge base. The knowledge base can be updated regularly to ensure it keeps pace with industry standards or actual needs. To facilitate knowledge point expansion, knowledge points and their relationships can also be represented using a knowledge graph, and the knowledge points extracted from the input data can be expanded based on this knowledge graph. A knowledge graph includes nodes and edges, where nodes represent knowledge points, and edges between nodes represent the relationships between the knowledge points represented by the corresponding nodes. A knowledge graph provides a clear knowledge point structure, facilitating the generation of questions that meet specific requirements, while also ensuring that the generated questions conform to subject-specific logic, reducing the error rate. Furthermore, the knowledge graph can be continuously expanded to accommodate more subjects and knowledge points.

[0037] Figure 2 An exemplary knowledge graph is shown, whose nodes include: Newton's Second Law, uniformly accelerated linear motion, the relationship between mass and inertia, resultant force calculation, velocity-time relationship, displacement-time relationship, and velocity-displacement relationship. The relationships between nodes are shown in the text information on the edges connecting the nodes. For example, the relationship between the nodes Newton's Second Law and uniformly accelerated linear motion is: Newton's Second Law is the dynamic basis of uniformly accelerated linear motion, and uniformly accelerated linear motion is the application scenario of Newton's Second Law. Based on the knowledge graph, knowledge points related to the knowledge points in the input data can be determined. The relationship between two knowledge points can be defined as a connection between the corresponding nodes of these two knowledge points in the knowledge graph, but is not limited to this. For example, assuming the knowledge point extracted from the input data is Newton's Second Law, then related knowledge points may include the relationship between mass and inertia, resultant force calculation, and / or uniformly accelerated linear motion. Furthermore, when expanding knowledge points, the target discipline to which the knowledge points extracted from the input data belong can be obtained, and the knowledge points extracted from the input data can be expanded based on the knowledge points in the knowledge graph belonging to the target discipline.

[0038] Furthermore, the knowledge graph may include various knowledge points of different difficulty levels, such as elementary school level arithmetic, middle school level Newton's second law, and university level calculus. The probability of a node being selected can be set based on the difficulty of the knowledge point it corresponds to, thereby reducing the likelihood of selecting knowledge points that are too difficult or too easy, and thus allowing for more accurate control over the difficulty of the generated questions.

[0039] In step S16, the relationships between multiple knowledge points in the input data can be extracted using a text-based language model. These relationships may specifically include at least one of the following:

[0040] Cause and effect relationship, such as extracting the cause and effect relationship from the knowledge points "heat is generated when current flows through a conductor" and "the thermal effect of current can be used to heat heating equipment", that is, because heat is generated when current flows through a conductor (cause), the thermal effect of current can be applied to heating equipment, such as electric water heaters, electric heaters, etc. (result).

[0041] Logical relationships, such as parallel relationships ("Photosynthesis requires conditions such as carbon dioxide, water and light" and "Photosynthesis can produce glucose and oxygen" both describe the process of photosynthesis), and progressive relationships ("Elements are pure substances composed of the same element" and "Elements can undergo a variety of chemical reactions" provide a progressively deeper explanation of elements).

[0042] Complementary relationships, such as "the Earth's internal structure includes the crust, mantle, and core" and "the crust is a thin shell on the Earth's surface, mainly composed of various rocks; the mantle is located between the crust and the core, mainly composed of silicate rocks; the core is mainly composed of iron and nickel," respectively describe the Earth's internal structure from different perspectives, forming a complete knowledge system;

[0043] Related relationships, such as "the sum of the interior angles of a triangle is 180 degrees" and "the sum of the interior angles of a quadrilateral is 360 degrees", are both related to the sum of the interior angles of geometric figures;

[0044] Inclusion relationships, such as "the animal kingdom includes vertebrates and invertebrates" and "vertebrates include fish, amphibians, reptiles, birds, and mammals," demonstrate the inclusion relationship in classification.

[0045] Semantic analysis can be used to extract the relationships between knowledge points, or causal reasoning can be used to perform causal reasoning on the input knowledge points to determine the relationships between them.

[0046] In step S18, multiple knowledge points and their relationships can be input into the question generation model so that the model can generate questions based on these knowledge points and their relationships. Taking the above relationships as causal relationships as an example, in physics, the knowledge points "a conductor generates heat when current flows through it" (cause) and "the thermal effect of current can be used to heat devices" (effect) and their causal relationship can be input into the model to generate the question: "Why do heating devices such as electric water heaters generate heat when they work? Please explain from the perspective of the phenomenon when current flows through a conductor and its applications." In biology, "photosynthesis requires carbon dioxide and water" (cause) and "photosynthesis can produce glucose and oxygen" (effect) and their causal relationship can be input into the model to generate the question: "Briefly describe the important significance of photosynthesis for the maintenance of green plants and even the entire ecosystem, and explain what its raw materials and products are respectively."

[0047] In some embodiments, the associated knowledge points of each knowledge point among multiple knowledge points and the relationship between each knowledge point and its associated knowledge points can be determined. A third prompt information is generated based on the multiple knowledge points, the relationship between the multiple knowledge points, the associated knowledge points of each knowledge point among multiple knowledge points, and the relationship between each knowledge point and its associated knowledge points. The third prompt information is then input into the question generation model so that the question generation model generates questions based on the third prompt information.

[0048] In this embodiment, questions are not only generated based on the knowledge points extracted from the input data and their relationships, but the knowledge points extracted from the input data are further expanded to obtain related knowledge points, thereby making the knowledge points available for question selection richer and improving the richness and diversity of the synthesized questions.

[0049] A knowledge graph can be pre-built, and based on the knowledge graph, the associated knowledge points of each knowledge point and the relationships between each knowledge point and its associated knowledge points can be determined. Details regarding the knowledge graph are provided in the aforementioned embodiments and will not be repeated here.

[0050] In some embodiments, the generated questions may include text-based content and non-text-based content, wherein the non-text-based content can be generated based on non-textual information included in the input data. Optionally, the non-textual information in the input data can be directly used as the non-textual content of the questions. For example, if the non-textual information in the input data is an image, the question generation model can generate the text-based content of the questions and use the image in the input data as the non-textual content of the questions.

[0051] In some embodiments, the generated question includes a stem, a solution, and an answer. A multi-turn dialogue can be used to generate the information included in the question. For example, in the first round of dialogue, a question design strategy can be input into the question generation model, causing the model to generate the stem based on the strategy. After generating the stem, a next round of dialogue can proceed, generating a hint based on the stem, which then enables the model to generate the solution and answer. Alternatively, in the first round, the question design strategy can be input into the model, causing it to generate the stem; in the second round, a hint can be generated, causing the model to generate the solution; and in the third round, another hint can be generated, causing the model to generate the answer.

[0052] In some embodiments, after generating the questions, semantic consistency verification can be performed on the input data and the questions. Semantic consistency verification refers to checking whether there is a semantic contradiction or conflict between the input data and the generated questions. If there is no semantic contradiction or conflict, the verification passes; otherwise, the verification fails. In some embodiments, semantic consistency verification of the input data and questions can be performed based on at least one of the following:

[0053] The consistency between the elements in the input data and the elements in the question is verified. Elements in the input data may include, but are not limited to, at least one of the following: text, graphics in images, lines, sound clips in audio, etc. Verifying the consistency between the elements in the input data and the elements in the question can include element type consistency verification (checking whether the question contains elements not included in the input data) and element quantity consistency verification (checking whether the number of elements in the question matches the number of elements of the corresponding type in the input data). When performing element type consistency verification, if the question contains elements not included in the input data (e.g., the input data includes images of triangles, while the question includes images of quadrilaterals), then the types of elements in the input data and the question are inconsistent. If all elements in the question are included in the input data, then the types of elements in the input data and the question are consistent. When performing element quantity consistency verification, if the number of elements in the question is the same as the number of elements of the corresponding type in the input data, then the number of elements in the input data and the question are consistent; otherwise, the number of elements in the input data and the question are inconsistent.

[0054] The consistency between the relationships between elements in the input data and the corresponding relationships between elements in the problem. These relationships include, but are not limited to, at least one of the following: geometric relationships, numerical relationships, logical relationships, semantic relationships, etc. If the relationship between any two elements in the problem is the same as the relationship between corresponding elements in the input data, then the relationships between the elements in the input data are consistent with the relationships between corresponding elements in the problem; otherwise, the relationships between the elements in the input data are inconsistent with the relationships between corresponding elements in the problem. For example, assuming both the input data and the problem include two straight lines (denoted as line A and line B respectively), if line A and line B intersect in the input data, but line A and line B are parallel in the problem, then the relationships between the elements in the input data are inconsistent with the relationships between corresponding elements in the problem.

[0055] The consistency between the annotation information in the input data and the annotation information in the question. Annotation information may include, but is not limited to, length information, angle information, and / or name. If the annotation information for a certain element in the question is the same as the annotation information for the corresponding element in the input data, then the annotation information in the input data is consistent with the annotation information in the question; otherwise, the annotation information in the input data is inconsistent with the annotation information in the question.

[0056] Logical consistency between input data and the problem. If the reasoning process and implicit conditions in the problem can be derived from the input data, then the input data and the problem are logically consistent; otherwise, the input data and the problem are logically inconsistent.

[0057] In addition to the dimensions listed above, other dimensions can also be used to verify the consistency between the input data and the generated questions. When multiple dimensions are used to verify the consistency between the input data and the questions, if all dimensions pass the verification (i.e., the input data and questions are consistent under that dimension), the consistency verification is considered successful. If any dimension fails the consistency verification (i.e., the input data and questions are inconsistent under that dimension), the consistency verification is considered unsuccessful.

[0058] If the validation fails, you can return to step S12 to obtain new input data and regenerate the questions based on the new input data. Alternatively, you can return to step S14 to re-extract the knowledge points. Or, you can return to step S16 to re-extract the relationships between the knowledge points.

[0059] In some embodiments, the question stem information and input data can be input into a multimodal language model, so that the multimodal language model generates answer information corresponding to the question stem information based on the input data, and performs a consistency check between the answer information generated by the multimodal language model and the answer information included in the question. If the check fails, the process can return to steps S12, S14, or S16 in the aforementioned embodiments. In this embodiment, the multimodal language model is used to answer the question stem information in the generated question. If the answer information obtained is inconsistent with the answer information in the question, it indicates that the question may have certain quality problems (such as unreasonable question stem information, inconsistent text and image content, illogical reasoning, etc.). Therefore, the question can be regenerated. Through the above method, the quality of the generated questions can be improved.

[0060] The overall flow of the embodiments of this specification will be described below with reference to the accompanying drawings. The embodiments of this specification significantly improve the quality, diversity, and applicability of generated questions through multimodal collaboration and intelligent optimization mechanisms, solving problems such as poor semantic consistency and insufficient flexibility in existing technologies. This solution not only meets the market demands of online education, intelligent assessment, and other fields, but also reduces labor costs and improves production efficiency through automated generation, and has the potential to be extended to multiple scenarios such as intelligent question answering and vocational training. Figure 3 As shown in the embodiments of this specification, the solution includes four core modules: knowledge extraction, question synthesis, verification and evaluation, and strategy adjustment. The specific steps are as follows:

[0061] (1) Knowledge Extraction Module

[0062] Step 1.1 Material Extraction: Obtain image or text input data, use a Visual Language Model (VLM) to parse and extract the internal semantics, and obtain multimodal knowledge points and tags.

[0063] Step 1.2 Analysis Method Definition: Using domain-specific semantic analysis techniques, the input data is broken down and categorized into knowledge points to ensure the accuracy of the extracted content and the rationality of the contextual understanding.

[0064] Step 1.3 Convergence / In-depth Analysis: Conduct in-depth analysis of the extracted knowledge points, including confirmation of semantic relevance, analysis of knowledge breadth, and evaluation of their suitability for integration with the question generation module.

[0065] This module ensures that the knowledge base content of the generated questions is accurate and semantically consistent, providing support for the subsequent question synthesis module.

[0066] (2) Question Synthesis Module

[0067] Step 2.1 Knowledge point input: Based on the knowledge points output by the knowledge extraction module, a structured corpus is generated by combining the language model (LLM).

[0068] Step 2.2 Determine the question design approach: Based on the application scenario (such as education, self-test questionnaires, intelligent dialogue datasets, etc.), design question types (such as short answer questions, multiple choice questions, fill-in-the-blank questions, etc.) and set the question depth and difficulty levels.

[0069] Step 2.3 Question Information Generation: Use a Language Model (LLM) to generate questions and automatically complete the question stem and options to form a semantic correspondence with the knowledge points.

[0070] Step 2.4 Solution Generation: Based on the generated question stem, automatically generate question analysis information and answer information.

[0071] This module combines knowledge reasoning and logical processing to ensure the accuracy of answer information and the completeness of explanation.

[0072] (3) Verification and Evaluation Module

[0073] Step 3.1 Semantic Consistency Verification: Verify the semantic consistency between the generated questions and the knowledge points in the input data, including the relevance of the question content and the scope of knowledge coverage.

[0074] Step 3.2 Answer Reasonableness Verification: This step checks the accuracy and logical feasibility of the generated answer information and question explanation information. If the answer information is incorrect or the question explanation information is logically unclear, the process proceeds to the strategy adjustment module.

[0075] Step 3.3 Question Difficulty Adjustment: Using semantic analysis technology, the question difficulty is adjusted to meet specific application requirements. The mechanism for adjusting question difficulty can be determined through the quality assessment process.

[0076] The verification and evaluation module ensures that the quality of the generated questions meets the design expectations and provides feedback for further optimization.

[0077] In some embodiments, after the question generation model generates questions, a domain-rule-based post-processing method can be applied to validate the generated question content. For example, the generated questions may be required to meet certain knowledge point coverage or knowledge depth standards. If these standards are not met, the questions are regenerated. Alternatively, validation rules can be defined using a knowledge base. The generated questions can be manually validated to ensure they match the defined rules, and then annotated based on the manual validation results. The annotation results can be fed back to the question generation model to enhance its performance in subsequent question generation processes.

[0078] (4) Strategy Adjustment Module

[0079] When the verification module finds defects or room for optimization, it enters the strategy adjustment stage to determine the difficulty adjustment strategy for the question generation model.

[0080] Specifically, the expected difficulty of the questions can be preset. If the validation module passes (i.e., the semantic consistency of the input data with the questions and the reasonableness of the answers both pass), a difficulty adjustment strategy is determined based on the difference between the difficulty of the currently generated questions and the expected difficulty. Specifically, if the validation module passes and the difficulty of the currently generated questions is lower than the expected difficulty, the difficulty adjustment strategy is to increase the difficulty of the next generated questions relative to the currently generated questions by a first adjustment margin. If the validation module passes and the difficulty of the currently generated questions is not lower than the expected difficulty, the difficulty adjustment strategy is to increase the difficulty of the next generated questions relative to the currently generated questions by a second adjustment margin. The first adjustment margin is greater than the second adjustment margin.

[0081] In this embodiment, if the verification module passes and the difficulty of the currently generated question is lower than the expected difficulty, the difficulty of the next generated question is increased by a larger adjustment compared to the currently generated question. This allows for more efficient generation of questions with a difficulty level no lower than the expected difficulty, thus improving question generation efficiency. Conversely, if the verification module passes and the difficulty of the currently generated question is no lower than the expected difficulty, the difficulty of the next generated question is increased by a smaller adjustment compared to the currently generated question. This allows for a slight exploration of the upper limit of the difficulty level that the question generation model can generate, thereby generating questions that are as difficult as possible.

[0082] In some embodiments, if the verification module fails (i.e., at least one of the semantic consistency verification of the input data and the question, and the reasonableness verification of the answer fails), the expected difficulty can be adjusted. Specifically, if the verification module fails and the preset difficulty adjustment trigger condition is met, the adjustment strategy for the expected difficulty information includes reducing the expected difficulty. If the verification module fails and the preset difficulty adjustment trigger condition is not met, the adjustment strategy for the expected difficulty information includes maintaining the expected difficulty. Generally, the higher the expected difficulty, the greater the difficulty for the question generation model to generate high-quality questions. If the expected difficulty is set too high, the question generation model may still be unable to generate questions that meet the expected difficulty after multiple iterations of adjustment, thus causing the question generation process to enter an infinite loop. This embodiment can adjust the expected difficulty information in a timely manner when the verification module fails, thereby avoiding the question generation process from entering an infinite loop. Furthermore, by setting the difficulty adjustment trigger condition, the question generation model has a certain number of attempts to generate questions that meet the preset quality constraints before the preset difficulty adjustment trigger condition is met. If a question that meets the preset quality constraints is still not generated even when the preset difficulty adjustment trigger conditions are met, it means that the expected difficulty is too high. In this case, the expected difficulty can be reduced to prevent the waste of system resources by continuously trying at the same expected difficulty.

[0083] In some embodiments, the generated questions can be scored. If the verification module fails, points can be deducted from the generated questions; if the verification module passes, points can be added to the generated questions. The difficulty adjustment trigger condition can be that the score of the generated questions is less than or equal to a preset score threshold (e.g., 0).

[0084] Specifically, during the (i+1)th generation of a question, the question generation score after the i-th generation can be obtained. If the question generated in the (i+1)-th generation fails the verification module's check, points can be deducted from the question generation score after the i-th generation, resulting in the final question generation score for the (i+1)-th generation. If the question generated in the (i+1)-th generation passes the verification module's check, points can be added to the question generation score after the i-th generation, resulting in the final question generation score for the (i+1)-th generation. If the final question generation score after the (i+1)-th generation is less than or equal to 0, then the preset difficulty adjustment trigger condition is met.

[0085] In some embodiments, the deduction step size for generating questions can be determined based on the relationship between the difficulty of the generated questions and the expected difficulty. Specifically, if the question generated in the (i+1)th iteration fails the verification module's check and its difficulty is lower than the expected difficulty, a deduction step size can be applied to the question generation score after the i-th iteration. If the question generated in the (i+1)th iteration fails the verification module's check and its difficulty is not lower than the expected difficulty, a deduction step size can be applied to the question generation score after the i-th iteration. The first deduction step size is larger than the second deduction step size. This guides the question generation model to generate questions of at least the expected difficulty while ensuring question quality, preventing the final generated questions from being too easy in order to guarantee quality.

[0086] After determining the difficulty adjustment strategy, the strategy can be fed back to the question generation model so that the question generation model can adjust the difficulty relationship between the next generated question and the currently generated question based on the difficulty adjustment strategy, and regenerate the question based on the difficulty relationship.

[0087] The embodiments described in this specification have the following advantages:

[0088] (1) Combining the semantic understanding capabilities of multimodal language models and text-based language models, it dynamically parses multimodal input data and can generate new and complex questions without the need for fixed templates.

[0089] (2) By combining reasoning mechanisms, the modeling of the relationship between knowledge points (such as causal relationship) can be realized, which can generate more complex logical question types (such as "causal chain reasoning question").

[0090] (3) Improve the consistency between the questions and the knowledge points in the input data through semantic consistency verification, and evaluate the quality of the generated questions in multiple dimensions, including knowledge point coverage, answer accuracy, and logical verification, to achieve adaptive closed loop, automatically adjust the generation strategy, and reduce reliance on manual intervention.

[0091] (4) It supports in-depth analysis and dynamic expansion of knowledge points (convergence and knowledge point supplementation), breaks through the limitation of traditional templates that can only cover fixed knowledge points, and supports a variety of question types, from short answer questions, multiple choice questions, fill-in-the-blank questions to cross-modal listening questions and video questions.

[0092] (5) Adopt modular design to provide clear interpretability (such as knowledge point accuracy analysis and verification module details).

[0093] (6) Provides a dynamic strategy adjustment mechanism, which replaces the complex reward function design with a scoring / simplification strategy to ensure that the generation logic is transparent and reliable.

[0094] See Figure 4 This specification also provides a question generation device, the device comprising:

[0095] Data acquisition module 102 is used to acquire input data; the input data includes non-text information;

[0096] The knowledge point extraction module 104 is used to extract knowledge points from the input data through a multimodal language model to obtain text description information of multiple knowledge points in the input data.

[0097] The relationship extraction module 106 is used to input the text description information of the multiple knowledge points into a text-based language model, so that the text-based language model can determine the relationship between the multiple knowledge points based on the text description information of the multiple knowledge points.

[0098] The question generation module 108 is used to input the plurality of knowledge points and the relationships between the plurality of knowledge points into the question generation model, so that the question generation model generates questions based on the plurality of knowledge points and the relationships between the plurality of knowledge points.

[0099] This specification also provides a computer device, which includes at least a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the foregoing embodiments.

[0100] Figure 5This diagram illustrates a more specific hardware structure of a computer device provided in an embodiment of this specification. The device may include: a processor 202, a memory 204, an input / output interface 206, a communication interface 208, and a bus 210. The processor 202, memory 204, input / output interface 206, and communication interface 208 are interconnected internally via the bus 210.

[0101] The processor 202 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification. The processor 202 may also include a graphics card, such as an Nvidia Titan X graphics card or a 1080Ti graphics card.

[0102] The memory 204 can be implemented in the form of read-only memory (ROM), random access memory (RAM), static storage device, dynamic storage device, etc. The memory 204 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 204 and is called and executed by the processor 202.

[0103] Input / output interface 206 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0104] The communication interface 208 is used to connect the communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0105] Bus 210 includes a pathway for transmitting information between various components of the device, such as processor 202, memory 204, input / output interface 206, and communication interface 208.

[0106] It should be noted that although the above-described device only shows the processor 202, memory 204, input / output interface 206, communication interface 208, and bus 210, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0107] This specification provides a computer program product, including a computer program that, when executed by a processor, implements the methods described in any embodiment of this specification.

[0108] This specification also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods described in any of the foregoing embodiments.

[0109] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by computer devices. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0110] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. When implementing the embodiments of this specification, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0111] The above description is merely a specific implementation of the embodiments of this specification. It should be noted that, for those skilled in the art, many improvements and modifications can be made without departing from the principles of the embodiments of this specification, and these improvements and modifications should also be considered within the protection scope of the embodiments of this specification.

Claims

1. A method for generating questions, the method comprising: Get the input data; The input data includes non-text information; The input data is processed by a multimodal language model to extract knowledge points, thereby obtaining textual descriptions of multiple knowledge points in the input data. The textual description information of the multiple knowledge points is input into a text-based language model so that the text-based language model can determine the relationship between the multiple knowledge points based on the textual description information of the multiple knowledge points. The multiple knowledge points and their relationships are input into the question generation model so that the question generation model generates questions based on the multiple knowledge points and their relationships.

2. The method according to claim 1, wherein extracting knowledge points from the input data using a multimodal language model comprises: The input data and the first prompt information are input into the multimodal language model so that the multimodal language model can output the knowledge point extraction strategy adopted by the input data based on the first prompt information. The first prompt message includes guidance information used to guide the multimodal language model to output the knowledge point extraction strategy; A second prompt message is generated based on the knowledge point extraction strategy, and the second prompt message is input into the multimodal language model so that the multimodal language model uses the knowledge point extraction strategy included in the second prompt message to extract knowledge points from the input data.

3. The method according to claim 1, wherein inputting the plurality of knowledge points and the relationships between the plurality of knowledge points into the question generation model, so that the question generation model generates questions based on the plurality of knowledge points and the relationships between the plurality of knowledge points, comprises: Determine the associated knowledge points for each of the multiple knowledge points and the relationship between each knowledge point and its associated knowledge points; A third prompt message is generated based on the multiple knowledge points, the relationships between the multiple knowledge points, the related knowledge points of each of the multiple knowledge points, and the relationships between each knowledge point and its related knowledge points. The third prompt information is input into the question generation model so that the question generation model generates questions based on the third prompt information.

4. The method according to claim 3, further comprising: Obtain a pre-established knowledge graph; the knowledge graph includes multiple nodes, each node corresponds to a knowledge point, and the edge between any two nodes represents the association relationship between the knowledge points corresponding to the two nodes. Based on the knowledge graph, determine the associated knowledge points of each of the multiple knowledge points and the relationship between each knowledge point and its associated knowledge points.

5. The method according to claim 1, further comprising: Perform semantic consistency verification on the input data and the question; If the verification fails, return to the step of obtaining input data, or return to the step of extracting knowledge points from the input data using a multimodal language model, or return to the step of inputting the text description information of the multiple knowledge points into a text-based language model.

6. The method according to claim 5, wherein the semantic consistency check of the input data and the question includes: The semantic consistency of the input data and the question is verified based on at least one of the following: The consistency between the elements included in the input data and the elements included in the question; The relationship between the elements in the input data is consistent with the relationship between the corresponding elements in the question; The consistency between the annotation information in the input data and the annotation information in the question; The input data and the question have logical consistency.

7. The method according to claim 1, wherein the question includes question stem information and answer information; the method further includes: The question stem information and the input data are input into the multimodal language model, so that the multimodal language model generates the answer information corresponding to the question stem based on the input data; The consistency between the answer information generated by the multimodal language model and the answer information included in the question is verified. If the verification fails, return to the step of obtaining input data, or return to the step of extracting knowledge points from the input data using a multimodal language model, or return to the step of inputting the text description information of the multiple knowledge points into a text-based language model.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 7.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1 to 7.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-round iteration mathematical question generation method, system and device and storage medium

    CN122115169A