AI-based large model evaluation question bank automatic generation method and system
Through the automatic generation method of AI-based large-scale model evaluation question bank, we dynamically acquire knowledge and generate questions from different dimensions, solving the problems of low efficiency and low quality of manual question bank generation, and achieving efficient and diversified question bank generation.
Patent Information
- Application Number
- CN202510637331.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In the prior art, manual generation of evaluation question banks is low efficiency and low quality of questions, making it difficult to cover diverse test scenarios.
The automatic generation method of AI-based large-scale model evaluation question bank is adopted to dynamically acquire knowledge, establish a knowledge base, generate questions from different dimensions, and enhance the processing, difficulty calibration and hierarchical calibration of the questions to establish a high-quality and diverse question bank.
It has achieved efficient generation of high-quality and diverse evaluation question banks, improved generation efficiency and question quality, and is especially suitable for artificial intelligence evaluation scenarios.
Smart Images

Figure CN120179810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly to an automatic generation method and system for an evaluation question bank of a large model based on AI. Background Art
[0002] Machine learning, especially deep learning, is changing numerous industries with its powerful predictive ability and wide application scenarios. With the rapid development of artificial intelligence technology, large-scale machine learning models, such as GPT, BERT, etc., are more and more widely used in various fields. Considering that machine learning models highly depend on data and computing resources, and there is often a problem of insufficient model interpretability, it is necessary to evaluate the models. Model evaluation is an important link in machine learning, which can evaluate the performance of a trained model and make a more optimized selection for the final model to be deployed.
[0003] To evaluate the performance of these models, a large number of evaluation questions usually need to be constructed. However, the traditional method of manually constructing a question bank is inefficient and difficult to cover diverse test scenarios.
[0004] Therefore, there is an urgent need for an automated method to generate a high-quality and diverse evaluation question bank. Summary of the Invention
[0005] The present invention solves the problems existing in the prior art, and provides an automatic generation method and system for an evaluation question bank of a large model based on AI, which solves the problems of low efficiency and low question quality in manually generating an evaluation question bank in the prior art.
[0006] The technical solution adopted by the present invention is an automatic generation method for an evaluation question bank of a large model based on AI, and the method includes the following steps: S1 Dynamically obtain knowledge and establish a knowledge base; S2 Based on the knowledge base, generate questions from different dimensions and perform enhancement processing on the questions; S3 After completing the verification, perform difficulty calibration and grading calibration on the questions and establish a question bank; S4 If the conditions for updating the knowledge base or the question bank are met, repeat S1 or S2.
[0007] Preferably, in S1, a feature vector of input multi-source data is extracted by a self-supervised pre-training language model, concepts and relationships therein are obtained based on the feature vector, a concept relationship network is constructed, and a knowledge graph is established as the knowledge base; in the knowledge graph, an importance score is matched for any concept.
[0008] Preferably, the concepts and relationships in the knowledge graph match time stamps.
[0009] Preferably, S2 includes the following steps: S2.1 Screen the question dimensions and generate an input sequence based on the question dimensions; S2.2 Fix the position features of the input sequence with positional encoding, then input it into the hybrid attention module, dynamically calculate the importance weights of different positions in the input sequence, and generate new questions; S2.3 Perform enhancement processing on the questions.
[0010] Preferably, in S2.2, the hybrid attention module includes a parallel multi-head attention unit, a self-attention unit, and a cross-attention unit. An input layer is provided in front of the multi-head attention unit, the self-attention unit, and the cross-attention unit, and an importance review layer and an output layer are provided behind the multi-head attention unit, the self-attention unit, and the cross-attention unit.
[0011] Preferably, in S2.3, the enhancement processing includes adding interference items, inserting multi-modal elements, and generating adversarial variants.
[0012] Preferably, S3 includes the following steps: S3.1 Preprocess the questions generated in S2; S3.2 Use the pre-tuned T5 model to evaluate the performance of the preprocessed questions and determine whether they meet the requirements of the questions; S3.3 Calibrate the difficulty of the questions and perform hierarchical calibration, and then establish a question bank.
[0013] Preferably, in S3.3, the difficulty calibration is associated with the computational complexity, the capacity of the hybrid attention module, and the difficulty score of the data features. The difficulty score of the data features is associated with the sequence length and feature dimension of the questions.
[0014] Preferably, establish an adversarial generation model and train it. Combine the simulated questions generated by the generator in the trained adversarial generation model with the real questions for difficulty calibration review.
[0015] An AI-based large model evaluation question bank automatic generation system, the system includes: An input unit for inputting or updating knowledge to the system; A question bank generation unit, using the above-mentioned AI-based large model evaluation question bank automatic generation method, for dynamically obtaining knowledge, establishing a knowledge base, generating questions, and performing difficulty calibration and hierarchical calibration to generate a question bank; An interaction unit, configured with an interaction port for outputting different questions in the question bank.
[0016] The present invention relates to a method and system for automatically generating an evaluation question bank for large models based on AI, which dynamically obtains knowledge and establishes a knowledge base, generates questions from different dimensions and enhances the questions. After verification, the difficulty of the questions is calibrated and graded, and a question bank is established, and the knowledge base or the question bank is updated by condition triggering; the system inputs or updates knowledge through an input unit, and the question bank generation unit uses the method to dynamically obtain knowledge, establish a knowledge base, generate questions and perform difficulty calibration and grading calibration to generate a question bank, and an interaction unit with a configured interaction port is used to output different questions in the question bank.
[0017] The beneficial effects of the present invention are as follows: through dynamic knowledge extraction, multi-dimensional question generation, intelligent difficulty calibration and adversarial verification, an evaluation question bank for models is automatically generated, with high generation efficiency, high question quality and strong diversity, and is particularly suitable for artificial intelligence evaluation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of the method of the present invention; Figure 2 is a schematic block diagram of the system structure of the present invention; Figure 3 is a schematic diagram of the hybrid attention module in the present invention; Figure 4 is a schematic flowchart of establishing a question bank once in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] The present invention relates to a method for automatically generating an evaluation question bank for large models based on AI, and the method includes the following steps: S1 Dynamically obtain knowledge and establish a knowledge base; S2 Based on the knowledge base, generate questions from different dimensions and enhance the questions; S3 After verification, calibrate the difficulty of the questions and perform grading calibration to establish a question bank; S4 If the condition for updating the knowledge base or the question bank is met, repeat S1 or S2.
[0021] The following is a specific implementation description in combination with the steps.
[0022] S1 Dynamically obtain knowledge and establish a knowledge base; In S1, a self-supervised pre-trained language model is used to extract the feature vectors of the input multi-source data, concepts and relationships therein are obtained based on the feature vectors, a concept relationship network is constructed, and a knowledge graph is established as a knowledge base; in the knowledge graph, an importance score is matched for any concept.
[0023] In the implementation process of the present invention, the multi-source data comes from, including but not limited to, textbooks, papers, technical documents, forum discussions, etc., which are regarded as knowledge metadata. The self-supervised pre-trained language model used is RoBERTa-large, which is used for extracting concepts and relationships therebetween; when constructing the concept relationship network, ConceptNet is considered and its technical feature of being able to connect various concepts to each other and assign weights to these relationships to represent semantic knowledge is fully utilized.
[0024] In the present invention, in order to better implement the importance score, the concept of time stamp is introduced, that is, the concepts and relationships in the knowledge graph match time stamps, which are used to evaluate the importance of concepts based on a time decay factor. The time decay factor γ is defined as γ = △T -a , where △T is the duration of knowledge appearance and a is an adjustable parameter.
[0025] S2 Based on the knowledge base, questions are generated from different dimensions and the questions are enhanced. Specifically, S2 covers the generation and enhancement of basic questions, including the following steps: S2.1 Screen the question dimensions and generate input sequences based on the question dimensions. Generally, existing question banks can be applied as the dimensions of the seed question bank. The dimensions of the questions include content security, model security, data security, model robustness, application and infrastructure security, etc. A rejected question bank dimension is also set to generate the seed question bank together.
[0026] Here, "content security" is used as the dimension of the input sequence. The input sequence (such as "content security") is encoded as a vector (such as word embedding) to form a matrix X, whose dimension is sequence length × model dimension d model .
[0027] Since the attention mechanism itself does not contain sequence order information, position encoding is required to inject position features into the input. Therefore, S2.2 Fix the positional features of the input sequence with Positional Encoding, generate a positional encoding matrix P using a sine function or a learnable vector, add it to the input to get X = X + P, and then input it into the hybrid attention module. Through the core component of the Attention Mechanism in Transformer, dynamically calculate the importance weights of different positions in the input sequence, focus on the key information, and generate a new question accordingly. In S2.2, the hybrid attention module includes a parallel multi-head attention unit, a self-attention unit, and a cross-attention unit. An input layer is provided before the multi-head attention unit, the self-attention unit, and the cross-attention unit, and an importance review layer and an output layer are provided after them.
[0028] S2.2.1 Generate three key vectors of X through linear transformation: Query (Q): The currently focused element; Key (K): The element to be compared; Value (V): The element that actually provides information; Calculation formula: Q = XW Q , K = XW K , V = XW V , where W Q , W K , W V are learnable weight matrices; S2.2.2 Calculate the attention weights: Similarity calculation, measure the similarity between the query Q and the key K through the dot product, Attention Scores = QK T , T is the transpose, and the larger the value of Attention Scores, the stronger the correlation; Scaling, divide by , to get Scaled Scores = , to prevent the dot product from being too large and causing unstable gradients, where d k is the dimension of the key vector; Softmax normalization, convert the scores into a probability distribution with a weight sum of 1, to get ; S2.2.3 Weight and sum the value vector V with the attention weights to get the final attention output, Output = Attention Weights ⋅ V; S2.2.4 Subsequently, the obtained attention weights and corresponding output values are input into the input layer of the hybrid attention module, and are processed by the parallel multi-head attention unit, self-attention unit, and cross-attention unit respectively; S2.2.4.1 The multi-head attention is used to capture diverse dependencies in different subspaces, enabling the hybrid attention module to "analyze the same passage from multiple perspectives", while paying attention to information in different dimensions and then splicing the conclusions of each group to comprehensively obtain the final result; First, split Q, K, V into h "heads" head1,..., head h , which are used to represent each independent subspace, and calculate the attention in parallel; Calculate the attention independently, with each head paying attention to information in different aspects (such as keywords, semantics, etc.), and finally splice the results and fuse them through a linear layer, satisfying MultiHead(Q, K, V)=Concat(head1,..., head h )W o where head i =Attention (QW i Q , KW i K , VW i V ), and W o is the weight matrix; S2.2.4.2 The self-attention is used to process the same piece of information internally. In self-attention, Q, K, V all come from the same input sequence (such as inside the encoder), which is used to capture the internal relationships of the sequence. During the process, analyze the relationship between each word in the sentence and other words to capture the "relationships"; S2.2.4.3 The cross-attention is adopted. In cross-attention, Q comes from the decoder, and K, V come from the encoder output. For example, in machine translation, the decoder pays attention to the encoder information, enabling two pieces of information to obtain key information across sequences, which is the same as associating different modalities or languages; S2.2.5 For the content output by the parallel multi-head attention unit, self-attention unit, and cross-attention unit, the importance weights of different positions in the input sequence are dynamically calculated by the importance review layer, and the content with the highest importance weight is selected, and the calculation result (generating the question) is output by the output layer to achieve the focus on key information; the dynamic calculation here is associated with the importance score matching for any concept mentioned above; of course, manual fine-tuning can also be performed.
[0029] Through the hybrid attention mechanism, the Transformer can efficiently model complex dependencies, dynamically adjust the focus according to the input, and make full use of dynamic weights; it can achieve long-range dependencies, directly model the relationships between elements at any distance, and solve the gradient vanishing problem of RNNs; it does not require sequential processing and can implement parallel computing to improve training efficiency.
[0030] S2.3 Enhance the questions.
[0031] In S2.3, the enhancement process includes adding interference items, inserting multimodal elements, and generating adversarial variants.
[0032] In the present invention, the enhancement process includes, but is not limited to, adding interference items for multi-turn Q&A, inserting multimodal elements including but not limited to charts and code snippets, and generating adversarial variants. In particular, when the number of heads in multi-head attention approaches infinity, what phenomenon will occur to the model.
[0033] After completing the verification in S3, calibrate the difficulty and grade of the questions, and establish a question bank; S3 includes the following steps: S3.1 Preprocess the questions generated in S2, including: Clean the text to remove noise such as HTML tags, special characters, and extra spaces; Unify the text format, such as converting all text to lowercase; Use a tokenizer to tokenize the text, such as using the T5Tokenizer pre-trained by the T5 model.
[0034] S3.2 Use the pre-tuned T5 model to evaluate the performance of the preprocessed questions to determine whether they meet the requirements of the questions; The T5 (Text-to-Text Transfer Transformer) model here is used to detect question ambiguity and verify fact accuracy based on a knowledge graph; after loading, organize the prepared test dataset into a format suitable for model input, generally an input data-label pair, where the label is 0 or 1, corresponding to non-question or question respectively, and fine-tune it using the Trainer class in the transformers library; use the test dataset to evaluate the fine-tuned model, view metrics such as the accuracy, recall, and F1 value of the T5 model, and finally use the fine-tuned model to predict whether a new question is a question (the output label is 1).
[0035] Eliminate the questions that do not meet the requirements this time.
[0036] After calibrating the difficulty and grading of the questions, establish a question bank.
[0037] In S3.3, the difficulty calibration is associated with the computational complexity C, the capacity M of the hybrid attention module, and the difficulty score D of the data features, and the difficulty score of the data features is associated with the sequence length and feature dimension of the question.
[0038] Regarding the difficulty calibration of the computational complexity, when processing the question text sequence, the time complexity of dot product attention is O(n 2 d), where n is the sequence length of the question text and d is the feature dimension. For some more complex questions, an additive attention mechanism needs to be adopted. Due to the addition of extra linear transformation and activation function calculations, the time complexity will increase accordingly. Let the extra complexity of the linear transformation and activation function calculations in the additive attention be O(n 2 d 2 ). Then the total time complexity C of the additive attention is approximately O(n 2 d + n 2 d 2 ). The higher the complexity, the greater the difficulty of the question in computational processing. Furthermore, the model capacity reflects the ability of the model to learn complex patterns, which is also applicable to the model for processing question difficulty calibration. In the hybrid attention module, its capacity is evaluated by calculating the number of its parameters. When the question involves more complex semantic and logical relationships, a larger-capacity model is required to process it, that is, more parameters mean a larger capacity M of the model, and it may also correspond to a higher question difficulty. Even further, it is necessary to consider the characteristics of the question text data to evaluate the difficulty. When the sequence length of the question text is longer, it means that more information segments need to be processed. And when the feature dimension is higher, it indicates that the semantic, syntactic, etc. features contained in the question are more complex. Based on this, a difficulty index that comprehensively considers the sequence length and feature dimension is defined, such as the difficulty score D = sequence length × feature dimension. In summary, the computational complexity C, the capacity M of the hybrid attention module, and the difficulty score D of the data features are weighted and fused. Through a large amount of experimental data and machine learning methods, the weights w1, w2, w3 of each dimension are determined, so as to obtain the final question difficulty calibration value S, S = w1C + w2M + w3D.
[0039] In actual applications, the present invention also establishes and trains an adversarial generation model, and uses the simulated questions generated by the generator in the trained adversarial generation model together with the real questions for difficulty calibration review, that is, to complete adversarial verification. The adversarial generation model includes a generator and a discriminator. In the scenario of question difficulty calibration, the task of the generator is to generate simulated question feature data that is as close as possible to the distribution of real question data. For example, it generates question text vectors with different sequence lengths and feature dimension combinations, or simulates data with different computational complexities of attention mechanisms and model capacities. The discriminator is responsible for determining whether the input data is real question features or simulated data generated by the generator. And for question difficulty calibration, the discriminator distinguishes between real questions and those generated by the generator based on the difficulty calibration results of computational complexity, model capacity, and data features; it checks the stability and accuracy of the calibration results.
[0040] Here, the adversarial generation model is something easily understood by those skilled in the art, and those skilled in the art can set it according to their needs.
[0041] During the implementation process, observe the fluctuation of the difficulty calibration value before and after adding simulated data. If the fluctuation is small, it indicates that adversarial verification helps to improve the stability and generalization ability of the difficulty calibration algorithm.
[0042] S4 If the conditions for updating the knowledge base or the question bank are met, then repeat S1 or S2.
[0043] In the present invention, by real-time monitoring data sources such as academic papers and technical blogs, the knowledge base is automatically updated, and an incremental learning algorithm is used to keep the knowledge fresh.
[0044] When there is no need for update, the question bank is continuously output.
[0045] The present invention also relates to an AI-based large model evaluation question bank automatic generation system, and the system includes: An input unit for inputting or updating knowledge to the system; A question bank generation unit, using the above-mentioned AI-based large model evaluation question bank automatic generation method, for dynamically obtaining knowledge, establishing a knowledge base, generating questions, and generating a question bank after difficulty calibration and hierarchical calibration; An interaction unit configured with an interaction port for realizing the output of different questions in the question bank.
[0046] In the present invention, the input unit refers to a module that automatically inputs and updates the knowledge base by real-time monitoring data sources such as academic papers and technical blogs, which can be achieved by data crawling or actively input by technicians through the interaction terminal; the interaction unit is generally at the application layer and can provide functions including but not limited to display and download.
[0047] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0048] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0049] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0050] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0051] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0052] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A method for automatically generating a large model evaluation question bank based on AI, characterized by: The method comprises the following steps: S1 dynamically acquires knowledge and builds a knowledge base; S2 generates questions from different dimensions based on the knowledge base and enhances the questions; After S3 completes the verification, the difficulty and grading of the questions are calibrated and a question bank is established; S4 If the conditions for updating the knowledge base or question bank are met, repeat S1 or S2.
2. According to claim 1, a method for automatically generating a large model evaluation question bank based on AI is characterized by: In S1, a self-supervised pre-trained language model is used to extract feature vectors of input multi-source data, and concepts and relationships therein are obtained based on the feature vectors, a concept relationship network is constructed, and a knowledge graph is established as a knowledge base; in the knowledge graph, an importance score is given to any concept match.
3. The method for automatically generating a large model evaluation question bank based on AI according to claim 2 is characterized in that: The concepts and relationships in the knowledge graph match timestamps.
4. The method for automatically generating a large model evaluation question bank based on AI according to claim 1, characterized in that: S2 includes the following steps: S2.1 Screen the topic dimensions and generate input sequences based on the topic dimensions; S2.2 uses position encoding to fix the position features of the input sequence, and then inputs it into the hybrid attention module to dynamically calculate the importance weights of different positions in the input sequence and generate new questions; S2.3 Enhance the questions.
5. The method for automatically generating a large model evaluation question bank based on AI according to claim 4 is characterized in that: In S2.2, the hybrid attention module includes parallel multi-head attention units, self-attention units and cross-attention units, and an input layer is provided in front of the multi-head attention units, self-attention units and cross-attention units, and an importance review layer and an output layer are provided after the multi-head attention units, self-attention units and cross-attention units.
6. The method for automatically generating a large model evaluation question bank based on AI according to claim 4, characterized in that: In S2.3, the enhancement processing includes adding interference terms, inserting multimodal elements, and generating adversarial variants.
7. The method for automatically generating a large model evaluation question bank based on AI according to claim 1, characterized in that: S3 includes the following steps: S3.1 Preprocess the questions generated in S2; S3.2 Use the pre-adjusted T5 model to evaluate the performance of the pre-processed questions to determine whether they meet the requirements of the questions; S3.3 Establish a question bank after calibrating the difficulty and grading of the questions.
8. The method for automatically generating a large model evaluation question bank based on AI according to claim 7, characterized in that: In S3.3, the difficulty calibration is associated with the computational complexity, the capacity of the hybrid attention module, and the difficulty score of the data feature, and the difficulty score of the data feature is associated with the sequence length and feature dimension of the question.
9. The method for automatically generating a large model evaluation question bank based on AI according to claim 1, characterized in that: An adversarial generative model is established and trained, and the simulated questions generated by the generator in the trained adversarial generative model are compared with the real questions for difficulty calibration and review.
10. An AI-based large-model evaluation question bank automatic generation system, characterized by: The system comprises: An input unit, used to input or update knowledge into the system; A question bank generation unit, which uses the AI-based large model evaluation question bank automatic generation method according to any one of claims 1 to 9 to dynamically acquire knowledge, establish a knowledge base, generate questions, perform difficulty calibration and grade calibration, and then generate a question bank; An interactive unit is configured with an interactive port and is used to realize the output of different questions in the question bank.
Citation Information
Patent Citations
Method for automatically generating test questions based on case text, storage medium and electronic equipment
CN114969337A
Geographic examination question generation method and device based on knowledge guidance
CN115455167A
Construction method based on novel research and development institution scientific and technological innovation service knowledge graph system
CN116992042A
Problem generation method and system based on large language model, electronic equipment and medium
CN118468034A
Knowledge data joint-driven deep learning method
CN118861763A
Cited By
High-precision AI surface test question generation method based on self-distillation and industry knowledge base
CN120632089A
High-precision ai interview question generation method based on self-distillation and industry knowledge base
CN120632089B