Multi-option question generation method and device, electronic equipment and storage medium

By receiving and processing multimodal contexts, combined with the most relevant examples in preset training data, using the multimodal large language model to generate multi-option problems and interference options, the problems of low quality of multi-option problems generation, poor domain correlation, and insufficient confusion of interference terms in the prior art are solved, and more efficient problem generation is achieved.

CN119938860AActive Publication Date: 2025-05-06SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510093799.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In the prior art, the multi-option problem generation quality is low, the domain correlation is poor, and the interference items are not confusing, which cannot meet the needs for problem diversity and scale.

Method used

By receiving the multimodal context input from the user, feature extraction is obtained to obtain the multimodal context embedding representation, retrieve the most relevant examples in the preset training dataset, and input them into the pre-trained multimodal large language model, outputting multi-option problems and interference options.

Benefits of technology

It improves the subject specificity and contextual relevance between generated questions, while also improving the confusion of interference items, and enhancing the challenge and value of the questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938860A_ABST
    Figure CN119938860A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-option question generation method and device, electronic equipment and a storage medium, which are used for solving the technical problems of low multi-option question generation quality, poor field correlation and insufficient confusion of interference terms in the prior art. The method comprises the following steps: receiving a multi-modal context input by a user; performing feature extraction on the multi-modal context to obtain a multi-modal context embedding representation; retrieving a most relevant example of the multi-modal context from a preset training data set; and inputting the context embedded representation and the most relevant example into a pre-trained multi-modal large language model, and outputting a multi-option problem and an interference option.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of question generation, and in particular to a method, device, electronic device and storage medium for generating multiple-choice questions. Background Art

[0002] Traditional multiple-choice question generation mostly relies on manual design. Designers in related fields need to spend a lot of time and energy to write questions based on specific fields and subject content. However, according to relevant research, even experienced experts can only design three to four high-quality multiple-choice questions per day. This inefficient and manual-dependent method cannot meet the needs of question diversity and scale.

[0003] In recent years, with the rapid development of deep learning and large language models, automatic question generation technology based on natural language processing has gradually emerged. Some studies (such as generation methods based on pre-trained language models) have made some progress in generating text-based multiple-choice questions, but these methods mainly rely on pure text input and ignore the multimodal information (such as pictures, charts, videos, etc.) that is widely present in the content. In fact, content often exists in a multimodal form, such as ecosystem diagrams in natural science courses and historical event maps in social sciences. This multimodal information is of great significance to the learning and understanding of sketching, but existing methods fail to make full use of this information.

[0004] In addition, when generating multiple-choice questions, the quality of distractor options directly affects the effectiveness of the questions. Low-quality distractor options are often too simple or obviously wrong, and cannot effectively assess students' cognitive level or stimulate their in-depth thinking. This problem has not been fully addressed in existing research. Although some methods can generate distractor options related to the question topic, these options lack sufficient confusion, thus reducing the challenge and value of the question. Summary of the invention

[0005] The present invention provides a multiple-choice question generation method, device, electronic device and storage medium, which are used to solve the technical problems in the prior art of low multiple-choice question generation quality, poor field relevance and insufficient confusion of interference items.

[0006] The present invention provides a method for generating a multiple-choice question, comprising:

[0007] receiving a multimodal context of user input;

[0008] Performing feature extraction on the multimodal context to obtain a multimodal context embedding representation;

[0009] Retrieving the most relevant examples of the multimodal context from a preset training dataset;

[0010] The context embedding representation and the most relevant examples are input into a pre-trained multimodal large language model, and a multi-option question and distractor options are output.

[0011] Optionally, the step of extracting features from the multimodal context to obtain a multimodal context embedding representation includes:

[0012] extracting text information and image information from the multimodal context;

[0013] Performing word segmentation processing on the text information to obtain a plurality of word segments;

[0014] Embedding the word segmentation to extract semantic features;

[0015] Converting the semantic features into high-dimensional semantic vectors to obtain text embedding representation of the text information;

[0016] Resizing the image information to obtain an image of a target size;

[0017] Performing pixel normalization on the target size image to obtain a normalized image;

[0018] extracting visual features from the normalized image by a visual transformer;

[0019] Encoding the visual features into a high-dimensional visual vector to obtain an image embedding representation of the image information;

[0020] The text embedding representation and the image embedding representation are fused to generate a multimodal context embedding representation.

[0021] Optionally, the step of retrieving the most relevant example of the multimodal context from a preset training data set comprises:

[0022] Extracting a plurality of samples from the preset training data set, each of the samples comprising a context, an answer and a question;

[0023] Using a feature encoder to sequentially encode the context, the answer, and the question to obtain a context vector, an answer vector, and a question vector;

[0024] Calculating feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample;

[0025] The feature similarity is used to filter the most relevant examples of the multimodal context.

[0026] Optionally, the pre-trained multimodal large language model includes a question generator, an intermediate reasoning generator, and a distractor generator; the step of inputting the context embedding representation and the most relevant example into the pre-trained multimodal large language model and outputting a multi-option question and distractor options includes:

[0027] Obtaining a target answer corresponding to the context embedding representation;

[0028] Constructing a first input using the context embedding representation, the target answer, and the most relevant example;

[0029] Inputting the first input into the question generator to generate a multiple-choice question;

[0030] Performing intermediate reasoning on the most relevant example to obtain a first reasoning result;

[0031] Adding the first reasoning result to the most relevant example in the first input, and adding the multiple-choice question to the first input to obtain a second input;

[0032] Inputting the second input into the intermediate reasoning generator to generate intermediate reasoning data;

[0033] Using a preset interference item demonstration to replace the most relevant example in the second input, and adding the intermediate reasoning data into the second input to obtain a third input;

[0034] The third input is input into a distractor generator to generate distractor options for the multiple-choice question.

[0035] The present invention also provides a multiple-choice question generating device, comprising:

[0036] A multimodal context receiving module, used to receive a multimodal context input by a user;

[0037] A feature extraction module, used to extract features from the multimodal context to obtain a multimodal context embedding representation;

[0038] A most relevant example retrieval module, used to retrieve the most relevant example of the multimodal context from a preset training data set;

[0039] The module for outputting multiple-choice questions and distracting options is used to input the context embedding representation and the most relevant examples into a pre-trained multimodal large language model, and output multiple-choice questions and distracting options.

[0040] Optionally, the feature extraction module includes:

[0041] A text information and image information extraction submodule, used to extract text information and image information from the multimodal context;

[0042] A word segmentation processing submodule is used to perform word segmentation processing on the text information to obtain a plurality of word segments;

[0043] A semantic feature extraction submodule, used to generate embedding for the word segmentation and extract semantic features;

[0044] A text embedding representation conversion submodule, used to convert the semantic features into high-dimensional semantic vectors to obtain a text embedding representation of the text information;

[0045] A size adjustment submodule, used to adjust the size of the image information to obtain an image of a target size;

[0046] A pixel normalization submodule, used for performing pixel normalization on the target size image to obtain a normalized image;

[0047] A visual feature extraction submodule, used for extracting visual features from the normalized image through a visual transformer;

[0048] An image embedding representation conversion submodule, used for encoding the visual features into high-dimensional visual vectors to obtain an image embedding representation of the image information;

[0049] The fusion submodule is used to fuse the text embedding representation and the image embedding representation to generate a multimodal context embedding representation.

[0050] Optionally, the most relevant example retrieval module includes:

[0051] A sample extraction submodule, used to extract a number of samples from the preset training data set, each of which includes context, answers and questions;

[0052] A vector generation submodule, used to use a feature encoder to sequentially encode the context, the answer, and the question to obtain a context vector, an answer vector, and a question vector;

[0053] A feature similarity calculation submodule, used to calculate the feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample;

[0054] The most relevant example screening submodule is used to screen the most relevant examples of the multimodal context using the feature similarity.

[0055] Optionally, the pre-trained multimodal large language model includes a question generator, an intermediate reasoning generator and a distractor generator; the multi-option question and distractor output module includes:

[0056] A target answer acquisition submodule, used to acquire a target answer corresponding to the context embedding representation;

[0057] A first input construction submodule, configured to construct a first input using the context embedding representation, the target answer, and the most relevant example;

[0058] A multiple-choice question generation submodule, used for inputting the first input into the question generator to generate a multiple-choice question;

[0059] A first reasoning result generating submodule, configured to perform intermediate reasoning on the most relevant example to obtain a first reasoning result;

[0060] A second input generating submodule, configured to add the first reasoning result to the most relevant example in the first input, and add the multiple-choice question to the first input, to obtain a second input;

[0061] an intermediate reasoning data generating submodule, configured to input the second input into the intermediate reasoning generator to generate intermediate reasoning data;

[0062] A third input generating submodule, configured to replace the most relevant example in the second input with a preset interference item demonstration, and add the intermediate reasoning data into the second input to obtain a third input;

[0063] The distractor option generation submodule is used to input the third input into the distractor item generator to generate distractor options for the multiple-choice question.

[0064] The present invention also provides an electronic device, the device comprising a processor and a memory:

[0065] The memory is used to store program code and transmit the program code to the processor;

[0066] The processor is used to execute any one of the above methods for generating multiple-choice questions according to the instructions in the program code.

[0067] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the multiple-choice question generating method as described in any one of the above items.

[0068] It can be seen from the above technical solutions that the present invention has the following advantages: The present invention provides a method for generating multiple-choice questions, and specifically discloses: receiving a multimodal context input by a user; extracting features from the multimodal context to obtain a multimodal context embedding representation; retrieving the most relevant example of the multimodal context from a preset training data set; inputting the context embedding representation and the most relevant example into a pre-trained multimodal large language model, and outputting a multiple-choice question and interference options. The present invention retrieves the most relevant example from a training data set formed by historical data, combines it with the multimodal context input by the user as input data, and analyzes it through a multimodal large language model to improve the subject specificity and context relevance between generated questions, while improving the confusingness of interference items. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0070] Figure 1 A flowchart of a method for generating multiple-choice questions provided by an embodiment of the present invention;

[0071] Figure 2 A flowchart of a method for generating multiple-choice questions provided by another embodiment of the present invention;

[0072] Figure 3 A schematic diagram of a method for generating multiple-option questions based on multi-modal chained demonstrative reasoning provided by an embodiment of the present invention;

[0073] Figure 4 A structural block diagram of a multiple-choice question generating device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0074] The embodiments of the present invention provide a multiple-choice question generation method, device, electronic device and storage medium, which are used to solve the technical problems in the prior art of low quality of multiple-choice question generation, poor field relevance, and insufficient confusion of interference items.

[0075] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0076] See also Figure 1 , Figure 1 A flowchart of the steps of a method for generating a multiple-choice question provided by an embodiment of the present invention.

[0077] The present invention provides a method for generating multiple-choice questions, which may specifically include the following steps:

[0078] Step 101, receiving a multimodal context input by a user;

[0079] Step 102, extracting features from the multimodal context to obtain a multimodal context embedding representation;

[0080] In the embodiment of the present invention, the multimodality may be any combination of text, chart, image, and video.

[0081] When a multimodal context of user input is received, features can be extracted to generate a multimodal context embedding representation.

[0082] Step 103, retrieving the most relevant examples of the multimodal context from a preset training data set;

[0083] In an embodiment of the present invention, historical data may be collected to generate a training data set, and the most relevant examples of the multimodal context may be retrieved therefrom.

[0084] Step 104, input the context embedding representation and the most relevant examples into a pre-trained multimodal large language model, and output a multi-option question and distractor options.

[0085] After obtaining the contextual embedding representation of the multimodal context and the most relevant examples, they can be input into the pre-trained multimodal large language model to generate multiple-choice questions and distractor options based on the reasoning strategy of the thought chain.

[0086] The present invention retrieves the most relevant examples from a training data set formed by historical data, combines them with the multimodal context input by the user as input data, and analyzes them through a multimodal large language model to improve the subject specificity and contextual relevance between generated questions, while also reducing the confusing nature of interference items.

[0087] See also Figure 2 , Figure 2 A flowchart of a method for generating multiple-choice questions provided by another embodiment of the present invention. Specifically, the following steps may be included:

[0088] Step 201, receiving a multimodal context input by a user;

[0089] Step 202, extracting text information and image information from the multimodal context;

[0090] Step 203, performing word segmentation processing on the text information to obtain a number of word segments;

[0091] Step 204, embedding the word and extracting semantic features;

[0092] Step 205, converting the semantic features into high-dimensional semantic vectors to obtain text embedding representation of the text information;

[0093] Word segmentation: Word segmentation is the process of dividing a continuous text into independent vocabulary units (tokens). These units can be words, phrases, or character-level elements.

[0094] Embedding generation: It maps discrete vocabulary units into a continuous vector space. Each word is represented as a high-dimensional vector that captures the semantic information and contextual relationship of the word.

[0095] In the embodiment of the present invention, the input text information is first segmented to form a number of segmented words, and then the segmented words are embedded to extract the semantic features in the text, such as keywords, logical relationships between concepts, and the overall meaning of the context. These semantic features are then further converted into high-dimensional semantic vectors, which are represented as text embedding representations of the text information.

[0096] Step 206, resizing the image information to obtain an image of a target size;

[0097] Step 207, normalizing the target size image to obtain a normalized image;

[0098] Step 208, extracting visual features from the normalized image through a visual transformer;

[0099] Step 209, encoding the visual features into a high-dimensional visual vector to obtain an image embedding representation of the image information;

[0100] In an embodiment of the present invention, the input image can be preprocessed, including resizing and pixel normalization, to obtain a normalized image; then the visual features of the image are extracted through a deep learning model such as a vision transformer (ViT). These visual features are then encoded into high-dimensional visual vectors, which can be used to represent image embedding representations of image information.

[0101] Step 210, fusing the text embedding representation and the image embedding representation to generate a multimodal context embedding representation;

[0102] In the implementation of the present invention, the text embedding representation and the image embedding representation can be fused through the multimodal attention mechanism to establish an association between text and image, capture the semantic connection between the two, and thus form a more comprehensive multimodal semantic representation as a multimodal contextual embedding representation. This fused representation contains the key information of the text and image, and can support more accurate question generation and reasoning processes in subsequent steps.

[0103] Step 211, retrieving the most relevant examples of the multimodal context from a preset training data set;

[0104] In an embodiment of the present invention, historical data may be collected to generate a training data set, and the most relevant examples of the multimodal context may be retrieved therefrom.

[0105] In one example, step 211 may include the following sub-steps:

[0106] S11, extracting a number of samples from a preset training data set, each sample including context, answer and question;

[0107] S12, using a feature encoder to sequentially encode the context, the answer, and the question to obtain a context vector, an answer vector, and a question vector;

[0108] S13, calculating the feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample;

[0109] S14, feature similarity is used to filter the most relevant examples of multimodal context.

[0110] In the specific implementation, a feature encoder is used to encode the attribute information of each sample, such as context T, answer A, and question Q, into a vector (including context vector, answer vector, and question vector), and all vectors are placed in a latent sample space with rich semantic features. The encoding method is shown in the following formula:

[0111]

[0112] in represents the feature encoder, Represents a feature vector. If two vectors are close in position in the latent space, it means that they are more likely to contain similar information in similar fields. Then calculate the cosine similarity of each attribute vector between the current test instance (the multimodal context mentioned above) and other samples in the dataset to determine the feature similarity between them. The calculation formula is as follows.

[0113]

[0114] in Represents the index number of the most similar sample retrieved in the sample pool. Subsequently, the current test instance and the most relevant retrieved example are concatenated into a formatted prompt as the input of the three generators. In the subsequent generation process, the retrieved examples will provide supplementary knowledge that may not exist in the context of the test instance, and flexibly control the output to make its style similar to the example. This method is very effective for question generation tasks with limited context. At the same time, the similar domain knowledge retrieved by the contextualized example retrieval module can also make the generated content closer to the current topic.

[0115] Step 212, input the context embedding representation and the most relevant examples into a pre-trained multimodal large language model, and output a multi-option question and distractor options.

[0116] After obtaining the contextual embedding representation of the multimodal context and the most relevant examples, they can be input into the pre-trained multimodal large language model to generate multiple-choice questions and distractor options based on the reasoning strategy of the thought chain.

[0117] In a specific implementation, the pre-trained multimodal large language model may include a question generator, an intermediate reasoning generator, and a distractor generator. The training of these three generators can be performed through multi-task training, specifically: all examples in the three tasks of question generation, intermediate reasoning generation, and distractor generation are combined and shuffled to assemble formatted data, and the basic question and intermediate reasoning are used as inputs for distractor generation to minimize the negative log-likelihood loss of the average word segmentation in the three generation tasks. The sum is used as the training target:

[0118]

[0119] in, is the maximum length of the output sequence, and Represent the first By training a multimodal large language model on these tasks simultaneously, the goal is to prevent intermediate errors during thought chaining training that could interfere with reasoning and to make the model more robust to the wording choices of the prompts.

[0120] In one example, step 212 may include the following sub-steps:

[0121] S21, obtain the target answer corresponding to the context embedding representation;

[0122] S22, constructs the first input using the contextual embedding representation, the target answer, and the most relevant example;

[0123] S23, inputting the first input into a question generator to generate a multiple-choice question;

[0124] S24, performing intermediate reasoning on the most relevant example to obtain a first reasoning result;

[0125] S25, adding the first reasoning result to the most relevant example in the first input, and adding the multiple-choice question to the first input, to obtain a second input;

[0126] S26, inputting the second input into the intermediate reasoning generator to generate intermediate reasoning data;

[0127] S27, replacing the most relevant example in the second input with a preset interference item demonstration, and adding the intermediate reasoning data to the second input to obtain a third input;

[0128] S28, inputting the third input into the distractor generator to generate distractor options for the multiple-choice question.

[0129] In the specific implementation, Figure 3 As shown, the context embedding representation, the most relevant example and the thought chain reasoning can be combined to generate multiple-choice questions and interference options. Specifically, it includes three stages: generating questions, generating intermediate reasoning and generating interference options. These three stages share the same model architecture and only differ in the input and output formats. It should be noted that the model architecture can select any large language model architecture in the field, and the embodiment of the present invention does not make specific restrictions on this.

[0130] Problem generation phase:

[0131] Get the contextual embedding to represent the target answer, and then provide the most relevant retrieved example to the question generator , target answer and contextual embedding representation (including text embedding representation and image embedding representation ), thereby outputting a multiple-choice question Q.

[0132]

[0133] in, For question generator.

[0134] Intermediate reasoning generation stage:

[0135] In the intermediate reasoning generation stage, the generated multiple-choice questions is added to the first input and in the most relevant examples Add the corresponding intermediate inference as , to construct further input for the second stage, i.e., the second input . Then, the updated second input is input into the intermediate inference generator to generate the intermediate inference data R.

[0136]

[0137] in, is the intermediate inference generator.

[0138] Interference option generation phase:

[0139] Demonstrate using appropriate distractors Replace the most relevant example and generate the intermediate inference data With the second input before Connect them, that is . Then, the modified third input is input into the distractor generator by the following formula:

[0140]

[0141] in Indicates distractor options for multiple-choice questions, is the distractor generator.

[0142] The present invention retrieves the most relevant examples from a training data set formed by historical data, combines them with the multimodal context input by the user as input data, and analyzes them through a multimodal large language model to improve the subject specificity and contextual relevance between generated questions, while also reducing the confusing nature of interference items.

[0143] See also Figure 4 , Figure 4 A structural block diagram of a multiple-choice question generating device provided in an embodiment of the present invention.

[0144] An embodiment of the present invention provides a multiple-choice question generating device, comprising:

[0145] The multimodal context receiving module 401 is used to receive the multimodal context input by the user;

[0146] A feature extraction module 402 is used to extract features from the multimodal context to obtain a multimodal context embedding representation;

[0147] The most relevant example retrieval module 403 is used to retrieve the most relevant examples of the multimodal context from the preset training data set;

[0148] The multiple-choice question and distraction option output module 404 is used to input the context embedding representation and the most relevant examples into the pre-trained multimodal large language model, and output the multiple-choice question and distraction options.

[0149] In this embodiment of the present invention, the feature extraction module 402 includes:

[0150] A text information and image information extraction submodule, used to extract text information and image information from a multimodal context;

[0151] The word segmentation processing submodule is used to perform word segmentation processing on text information to obtain a number of word segments;

[0152] The semantic feature extraction submodule is used to generate word embeddings and extract semantic features;

[0153] The text embedding representation conversion submodule is used to convert semantic features into high-dimensional semantic vectors to obtain text embedding representation of text information;

[0154] A size adjustment submodule is used to adjust the size of the image information to obtain an image of a target size;

[0155] A pixel normalization submodule is used to perform pixel normalization on the target size image to obtain a normalized image;

[0156] A visual feature extraction submodule, for extracting visual features from the normalized image through a visual transformer;

[0157] The image embedding representation conversion submodule is used to encode visual features into high-dimensional visual vectors to obtain image embedding representation of image information;

[0158] The fusion submodule is used to fuse text embedding representation and image embedding representation to generate multimodal contextual embedding representation.

[0159] In the embodiment of the present invention, the most relevant example retrieval module 403 includes:

[0160] The sample extraction submodule is used to extract several samples from the preset training data set, each sample includes context, answer and question;

[0161] A vector generation submodule is used to sequentially encode the context, the answer, and the question using a feature encoder to obtain a context vector, an answer vector, and a question vector;

[0162] The feature similarity calculation submodule is used to calculate the feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample;

[0163] The most relevant example filtering submodule is used to filter the most relevant examples of multimodal contexts using feature similarity.

[0164] In the embodiment of the present invention, the pre-trained multimodal large language model includes a question generator, an intermediate reasoning generator and a distractor generator; the multi-option question and distractor output module 404 includes:

[0165] The target answer acquisition submodule is used to obtain the target answer corresponding to the context embedding representation;

[0166] A first input construction submodule, for constructing a first input using a context embedding representation, a target answer, and a most relevant example;

[0167] A multiple-choice question generation submodule, used for inputting the first input into a question generator to generate a multiple-choice question;

[0168] A first reasoning result generating submodule, used for performing intermediate reasoning on the most relevant example to obtain a first reasoning result;

[0169] A second input generating submodule, configured to add the first reasoning result to the most relevant example in the first input, and add the multiple-choice question to the first input, to obtain a second input;

[0170] An intermediate reasoning data generating submodule, used for inputting the second input into the intermediate reasoning generator to generate intermediate reasoning data;

[0171] A third input generation submodule, used to replace the most relevant example in the second input with a preset interference item demonstration, and add the intermediate reasoning data into the second input to obtain a third input;

[0172] The distractor option generation submodule is used to input the third input into the distractor generator to generate distractor options for the multiple-choice question.

[0173] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:

[0174] The memory is used to store the program code and transmit the program code to the processor;

[0175] The processor is used to execute the multiple-choice question generating method of the embodiment of the present invention according to the instructions in the program code.

[0176] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the multiple-choice question generating method of the embodiment of the present invention.

[0177] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0178] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0179] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0180] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0181] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0183] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0184] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0185] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating multiple-choice questions, characterized in that: include: receiving a multimodal context of user input; Performing feature extraction on the multimodal context to obtain a multimodal context embedding representation; Retrieving the most relevant examples of the multimodal context from a preset training dataset; The context embedding representation and the most relevant examples are input into a pre-trained multimodal large language model, and a multi-option question and distractor options are output.

2. The method according to claim 1, characterized in that: The step of extracting features from the multimodal context to obtain a multimodal context embedding representation comprises: extracting text information and image information from the multimodal context; Performing word segmentation processing on the text information to obtain a plurality of word segments; Embedding the word segmentation to extract semantic features; Converting the semantic features into high-dimensional semantic vectors to obtain text embedding representation of the text information; Resizing the image information to obtain an image of a target size; Performing pixel normalization on the target size image to obtain a normalized image; extracting visual features from the normalized image by a visual transformer; Encoding the visual features into a high-dimensional visual vector to obtain an image embedding representation of the image information; The text embedding representation and the image embedding representation are fused to generate a multimodal context embedding representation.

3. The method according to claim 1, characterized in that The step of retrieving the most relevant example of the multimodal context from a preset training data set comprises: Extracting a plurality of samples from the preset training data set, each of the samples comprising a context, an answer and a question; Using a feature encoder to sequentially encode the context, the answer, and the question to obtain a context vector, an answer vector, and a question vector; Calculating feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample; The feature similarity is used to filter the most relevant examples of the multimodal context.

4. The method according to claim 1, characterized in that: The pre-trained multimodal large language model includes a question generator, an intermediate reasoning generator, and a distractor generator; the step of inputting the context embedding representation and the most relevant example into the pre-trained multimodal large language model and outputting a multi-option question and distractor options includes: Obtaining a target answer corresponding to the context embedding representation; Constructing a first input using the context embedding representation, the target answer, and the most relevant example; Inputting the first input into the question generator to generate a multiple-choice question; Performing intermediate reasoning on the most relevant example to obtain a first reasoning result; Adding the first reasoning result to the most relevant example in the first input, and adding the multiple-choice question to the first input to obtain a second input; Inputting the second input into the intermediate reasoning generator to generate intermediate reasoning data; Using a preset interference item demonstration to replace the most relevant example in the second input, and adding the intermediate reasoning data into the second input to obtain a third input; The third input is input into a distractor generator to generate distractor options for the multiple-choice question.

5. A multiple-choice question generating device, characterized in that: include: A multimodal context receiving module, used to receive a multimodal context input by a user; A feature extraction module, used to extract features from the multimodal context to obtain a multimodal context embedding representation; A most relevant example retrieval module, used to retrieve the most relevant example of the multimodal context from a preset training data set; The module for outputting multiple-choice questions and distracting options is used to input the context embedding representation and the most relevant examples into a pre-trained multimodal large language model, and output multiple-choice questions and distracting options.

6. The device according to claim 5, characterized in that The feature extraction module comprises: A text information and image information extraction submodule, used to extract text information and image information from the multimodal context; A word segmentation processing submodule is used to perform word segmentation processing on the text information to obtain a plurality of word segments; A semantic feature extraction submodule, used to generate embedding for the word segmentation and extract semantic features; A text embedding representation conversion submodule, used to convert the semantic features into high-dimensional semantic vectors to obtain a text embedding representation of the text information; A size adjustment submodule, used to adjust the size of the image information to obtain an image of a target size; A pixel normalization submodule, used for performing pixel normalization on the target size image to obtain a normalized image; A visual feature extraction submodule, used for extracting visual features from the normalized image through a visual transformer; An image embedding representation conversion submodule, used for encoding the visual features into high-dimensional visual vectors to obtain an image embedding representation of the image information; The fusion submodule is used to fuse the text embedding representation and the image embedding representation to generate a multimodal context embedding representation.

7. The device according to claim 5, characterized in that The most relevant example retrieval module comprises: A sample extraction submodule, used to extract a number of samples from the preset training data set, each of which includes context, answers and questions; A vector generation submodule, used to use a feature encoder to sequentially encode the context, the answer, and the question to obtain a context vector, an answer vector, and a question vector; A feature similarity calculation submodule, used to calculate the feature similarity between the multimodal context and the context vector, answer vector, and question vector of each sample; The most relevant example screening submodule is used to screen the most relevant examples of the multimodal context using the feature similarity.

8. The device according to claim 5, characterized in that The pre-trained multimodal large language model includes a question generator, an intermediate reasoning generator and a distractor generator; the multi-option question and distractor output module includes: A target answer acquisition submodule, used to acquire a target answer corresponding to the context embedding representation; A first input construction submodule, configured to construct a first input using the context embedding representation, the target answer, and the most relevant example; A multiple-choice question generation submodule, used for inputting the first input into the question generator to generate a multiple-choice question; A first reasoning result generating submodule, configured to perform intermediate reasoning on the most relevant example to obtain a first reasoning result; A second input generating submodule, configured to add the first reasoning result to the most relevant example in the first input, and add the multiple-choice question to the first input, to obtain a second input; an intermediate reasoning data generating submodule, configured to input the second input into the intermediate reasoning generator to generate intermediate reasoning data; A third input generating submodule, configured to replace the most relevant example in the second input with a preset interference item demonstration, and add the intermediate reasoning data into the second input to obtain a third input; The distractor option generation submodule is used to input the third input into the distractor item generator to generate distractor options for the multiple-choice question.

9. An electronic device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the multiple-choice question generating method according to any one of claims 1 to 4 according to the instructions in the program code.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to execute the multiple-choice question generating method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-option question and answer method and device, computer readable storage medium and terminal

    CN118051588A

  • File explorer system usable in an emulated integrated development environment (IDE)

    US20160239509A1

  • Question answering system-based generation of distractors using machine learning

    US20160292593A1