Large model automatic proposition method, system and equipment for self-learning examination and medium

Through the large-scale automatic question setting method, the problems of teachers' workload and stable test quality in the traditional self-study test question setting method are solved, and an efficient and stable question setting process is achieved, which is suitable for educational scenarios in different subjects.

CN120045638AActive Publication Date: 2025-05-27QICHEN GUANGZHOU ELECTRONICS TECH CO LTD

Patent Information

Application Number
CN202510473875.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-27
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the traditional self-study examination question setting methods, teachers have a large workload and stable test quality.

Method used

The big model automatic question setting method is adopted, and the textbook documents-knowledge points-subject questions and knowledge points-cognitive levels-discipline questions are established by scanning and pre-processing the relevant documents of the self-study exam, and the training big model is used to generate the destined test questions, answers and question setting basis.

Benefits of technology

It improves the efficiency of self-study exam questions, reduces the workload of teachers, ensures the stability and fairness of the quality of the test questions, and can be applied to educational scenarios in different subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045638A_ABST
    Figure CN120045638A_ABST
Patent Text Reader

Abstract

The invention discloses a large-model automatic proposition method, system, device and medium for a self-learning examination, and is suitable for full-automatic proposition construction of the self-learning examination, and the proposition method comprises the following steps: scanning different self-learning examination subject test papers, self-learning examination exercises and self-learning examination teaching materials to obtain electronic document images; carrying out OCR (Optical Character Recognition) and archiving Establishing a mapping relationship between the textbook content and the subject test questions to obtain a textbook-test question document data set; training the large model by adopting a textbook-test question document data set; deploying a large model to realize availability of an offline scene; and inputting the textbook content to be subjected to question setting and the knowledge points into the trained large model, and outputting life test questions, answers and proposition bases related to the textbook content. According to the method, the data set of the textbook mapped to the test questions is constructed, and the large model is input for training, so that the life system efficiency of the self-learning test questions is effectively improved, and accurate proposition of different subjects in the self-learning test is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and educational technology, and particularly relates to a large model automatic proposition method, an automatic proposition system, a computer device, and a storage medium for self-study examinations. Background Art

[0002] In the past, the proposition process of self-study examinations mainly relied on the experience and intuition of educational experts, as well as manually retrieving materials to compile questions. Although this method can ensure the professionalism and accuracy of the questions, with the development of the times and the expansion of the education system, its limitations have become increasingly apparent. The continuous increase in the number of self-study examination subjects has made teachers face increasingly heavy tasks. They not only need to master a wide range of knowledge fields but also design examination questions that meet the requirements of the teaching syllabus and can effectively evaluate students' learning outcomes.

[0003] In traditional methods, teachers must invest a large amount of time and energy in repetitive tasks, such as searching for reference materials, determining appropriate difficulty levels, and ensuring that the questions cover all necessary knowledge points. This high-intensity workload not only brings a large workload to teachers but may also affect the quality and fairness of proposition due to human factors.

[0004] In recent years, the progress of artificial intelligence technology, especially large language models and other machine learning algorithms, has provided new possibilities for solving these problems. These large models, trained with a large amount of data, can understand complex language structures and demonstrate strong adaptability and innovation capabilities in different industries, and can replace humans to perform mechanical and tedious work to a certain extent. Therefore, it is of great significance to apply large model-related technologies to self-study examinations. Summary of the Invention

[0005] The main purpose of the present invention is to solve the problems of heavy workload of teachers and stable quality of test questions in the traditional self-study examination proposition method. To improve the proposition efficiency of self-study examinations, a large model automatic proposition method, an automatic proposition system, a computer device, and a storage medium for self-study examinations are provided, which are suitable for the proposition of test questions in different disciplines.

[0006] To achieve the above first purpose, the present invention discloses a large model automatic proposition method for self-study examinations, including the following steps:

[0007] S1. Scan test papers, self-study examination exercises, and self-study examination textbooks of different self-study examination disciplines, and preprocess to obtain clear electronic document images;

[0008] S2. Recognize the electronic document images through optical character recognition and archive the text content;

[0009] S3. Establish the mapping relationships between teaching material documents - knowledge points - subject test questions and knowledge points - cognitive levels - subject test questions, and obtain the teaching material document - knowledge point - subject test question dataset and the knowledge point - cognitive level - subject test question dataset;

[0010] S4. Use the teaching material document - knowledge point - subject test question dataset and the knowledge point - cognitive level - subject test question dataset to train the large model;

[0011] S5. Input the teaching material content and knowledge points to be tested into the trained large model, and output the test questions, answers, and proposition bases related to the teaching material content.

[0012] Furthermore, the large model includes a text encoding layer and a model backbone connected in sequence; among them, in the text encoding layer, considering the factor that the self-study teaching material content has a long length and the original model structure is difficult to adapt to long documents, rotary position encoding is introduced to improve the reasoning length and ability of the large model; the model backbone is to fully extract the global semantic features of the teaching material text through multiple network modules, effectively understand the teaching material content, and improve the proposition quality;

[0013] The text encoding layer uses byte pair encoding as the tokenization method to encode the input text, encodes different teaching materials and test questions with different lengths into initial text vectors of the same dimension, and denotes the text encoding layer as G( ) and is defined as follows:

[0014]

[0015] Among them, represents the initial input teaching material document content with a length of , represents the initial text vector after passing through the text encoding layer, contains word embedding vectors with a dimension of ;

[0016] On the basis of the initial text vector, rotary position encoding is introduced and defined as follows:

[0017]

[0018] Among them, represents the frequency parameter for rotary position encoding, represents the index number, = 1, 2…, d - 1, represents the floor function. Therefore, the rotary position encoding and the final text vector are defined as follows:

[0019]

[0020]

[0021] Among them, represents the position index of each element in the input text, and respectively represent the word embedding vectors at positions and in and respectively represent the cosine function and the sine function;

[0022] The model backbone includes 32 network modules connected in sequence. Each network module includes a first root mean square layer, an attention layer, a second root mean square layer, and a multi-layer perceptron. The first root mean square layer and the second root mean square layer are used to scale the input vector. The scaling operation formula is defined as follows:

[0023]

[0024] Among them, is the i-th element of the text vector the -th is the learnable parameter corresponding to the -th element, represents the number of elements of the input vector;

[0025] The attention layer adopts a grouped query attention mechanism. Each group of queries shares the same key and value, which is defined as follows:

[0026]

[0027] Among them, is the query matrix. The query matrix is divided into groups, represents the index of the group, g = 1, 2…, G, and respectively represent the key matrix and the value matrix of the -th group. Multiple query matrices within the same group correspond to the same key matrix and value matrix. represents the normalized exponential function, represents the dimension of the key, which is used to scale the dot product to stabilize the gradient;

[0028] The multi-layer perceptron consists of 3 linear layers and an activation function. The activation function is used to capture the non-linear relationships in the input data. The multi-layer perceptron weights the information at different positions in the output sequence of the attention layer to form a more informative vector representation, which is defined as follows:

[0029]

[0030] Among them, represents the vector output by the second root mean square layer, , and represent the weight matrices of the first, second, and third linear layers respectively, , and represent the bias vectors of the first, second, and third linear layers respectively, represents the output of the first linear layer, represents the output of the second linear layer, represents the SwiGLU activation function, which takes the output of the first linear layer as the input, represents the element-wise multiplication between matrices.

[0031] Furthermore, the text encoding layer receives the initial input textbook document content . The model backbone is connected to the text vector output by the text encoding layer. In the first layer network module of the model backbone, the output of the first root mean square layer is connected to the input of the attention layer. The input passes through the query matrix and key matrix in the attention layer, dot product operation, softmax function, and then performs matrix multiplication with the output of the input passing through the value matrix. The output of the attention layer is connected to the second root mean square layer, and the output of the second root mean square layer is connected to the input of the multi-layer perceptron. After passing through the first linear layer and activation function in the multi-layer perceptron, the output connected to the third linear layer after dot product operation with the output of the input passing through the second linear layer is obtained as the final output of the first layer network module; the input of the second layer network module is connected to the output of the first layer network module, and so on. After passing through 32 layer network modules, the final output text is obtained.

[0032] Furthermore, for the traditional method of manually establishing mappings, which is not only time-consuming and laborious but also difficult to establish cross-chapter associations of knowledge points among multiple textbook documents, we use the similarity information between documents and knowledge points, as well as the cognitive level information of subject questions in the examination syllabus to establish mapping relationships, which greatly speeds up the processing time of single textbook sheets and improves the mapping accuracy;

[0033] The process of establishing the mapping relationships among textbook content, knowledge points, and subject questions, as well as knowledge points, cognitive levels, and subject questions in step S3 is as follows:

[0034] S301. Extract the corresponding knowledge points of the textbook and the cognitive levels of the knowledge points from the examination outlines of different subjects in self-study examinations, and establish the mapping between the textbook and the knowledge points and the mapping between the knowledge points and the cognitive levels;

[0035] S302. Establish the mapping relationships of textbook document - knowledge point - subject question and knowledge point - cognitive level - subject question through keyword matching operations and text similarity matching operations:

[0036] Keyword matching operation: Divide the textbook into several documents according to sections, and preprocess the documents and real exam questions, including removing stop words, punctuation marks, special characters, and tokenization; Extract keywords from the preprocessed documents and real exam questions to obtain document keyword vectors and real exam question keyword vectors respectively, and calculate the keyword matching degree through similarity calculation;

[0037] Text similarity matching operation: Encode the documents and real exam questions respectively, and calculate the cosine similarity between the textbook fragment vectors and the real exam question vectors; Perform weighted summation of the keyword matching degree and the cosine similarity, select the weights of the keyword matching degree and the cosine similarity, calculate the comprehensive score of each real exam question with all documents, and select the document with the highest comprehensive score as the mapping result;

[0038] After obtaining the mapping relationships of textbook document - knowledge point - subject question and knowledge point - cognitive level - subject question, according to the question types of self-study examination real exam questions, the question types include single-choice questions, multiple-choice questions, fill-in-the-blank questions, short-answer questions, noun explanation questions, and case questions; Construct the above mapping relationships into the following four types of data formats to obtain the final textbook document - knowledge point - subject question dataset and knowledge point - cognitive level - subject question dataset:

[0039] (1) Input: Please formulate questions based on the following textbook content {textbook document}, Output: {subject question};

[0040] (2) Input: Please formulate questions based on the following textbook content and knowledge points {textbook document}{knowledge point}, Output: {subject question};

[0041] (3) Input: Please formulate questions based on the following knowledge points and cognitive levels {knowledge point}{cognitive level}, Output: {subject question};

[0042] (4) Input: Please judge the corresponding cognitive level according to the following subject question {subject question}, Output: {cognitive level}.

[0043] Further, the specific process of training the large model with the textbook - question document dataset in step S4 is as follows:

[0044] S401. Initialize the automatic proposition large model, set the number of training rounds, truncate the input teaching material document according to the specified maximum length, select the AdamW optimizer, set the batch size, set the gradient accumulation steps, and obtain the initial training settings;

[0045] S402. Take the teaching material - test question document dataset in batches, divide it into several batches, randomly select a batch of teaching material - test question document label pairs from the divided dataset, and input them into the large model for forward propagation to obtain the predicted text;

[0046] S403. Use the large model loss function to calculate the difference between the target text and the predicted text, and update the large model parameters using the error backpropagation algorithm;

[0047] S404. Save checkpoints every specified number of training steps, and evaluate the performance of the current training times. If the performance does not improve and reaches the early stopping threshold, stop training in advance;

[0048] S405. Continuously update the learning rate during the training process according to the training situation, and obtain the new learning rate according to the learning rate decay strategy.

[0049] Furthermore, for the proposition mode in self - taught examinations, use the constructed datasets in four formats to train the large model, effectively improving the proposition quality and efficiency of the large model in self - taught examinations. The large model loss function is defined as:

[0050]

[0051]

[0052]

[0053]

[0054]

[0055] where represents the large model loss function, represents the large model parameters, 、 、 and correspond to the above four types of datasets respectively, 、 、 and represent the number of samples used for training in each dataset, and in each dataset, the 、 , and training samples are input into the textbook text fragments, , , and for each dataset, the , , and disciplinary test questions corresponding to the training samples, and represent the , knowledge points corresponding to the training samples, and the and cognitive levels corresponding to the training samples, is a conditional probability distribution, given the condition , the probability of outputting is obtained.

[0056] The second object of the present invention is to provide a large model automatic proposition method and system for self-study examinations, which is used to execute the above-mentioned large model automatic proposition method for self-study examinations. The large model automatic proposition system includes:

[0057] A scanning and preprocessing module, which scans test papers, exercises, and textbooks of different self-study examination disciplines, and preprocesses them to obtain clear electronic document images;

[0058] An optical character recognition module, which recognizes the electronic document images through optical character recognition and archives the text content;

[0059] A mapping relationship and dataset construction module, which establishes the mapping relationships of textbook document - knowledge point - disciplinary test question and knowledge point - cognitive level - disciplinary test question, and obtains the textbook document - knowledge point - disciplinary test question dataset and the knowledge point - cognitive level - disciplinary test question dataset;

[0060] A large model training module, which trains the large model by using the textbook document - knowledge point - disciplinary test question dataset and the knowledge point - cognitive level - disciplinary test question dataset;

[0061] A self-study examination question proposition module, which inputs the textbook content and knowledge points to be set questions into the trained large model, and outputs the set questions, answers, and proposition bases related to the textbook content.

[0062] The third object of the present invention is to provide a computer device, including a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, the above-mentioned large model automatic proposition method for self-study examinations is realized.

[0063] The fourth object of the present invention is to provide a storage medium storing a program which, when executed by a processor, implements the above-mentioned large model automatic proposition method for self-study examinations.

[0064] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0065] 1. The present invention discloses a large model automatic proposition method and system for self-study examinations, and proposes a large model for fully automatic proposition in self-study examinations. This model is trained using the constructed textbook-question document data set, enabling the large model to master the proposition mode and characteristics of self-study examinations, ensuring the stability of the proposition quality of the model. At the same time, a paging method is adopted for model deployment, effectively accelerating the inference speed of the model and enhancing the adaptability of the model in practical applications.

[0066] 2. Compared with the existing proposition methods, in the face of the continuous increase in the number of self-study examination subjects, the present invention can continuously expand the capabilities of the model by continuously enriching the textbook-question document data set, ensuring the robustness of the model. The technical solution of the present invention has strong subject generality and can be applied to the educational scenarios of different subjects. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 is a flowchart of a large model automatic proposition method and system for self-study examinations disclosed in Embodiment 1 of the present invention;

[0069] Figure 2 is a model structure diagram of a large model automatic proposition method for self-study examinations disclosed by the present invention;

[0070] Figure 3 is a network structure diagram of the large model for self-study examination proposition disclosed by the present invention;

[0071] Figure 4 is a schematic diagram of the self-study examination proposition result disclosed by the present invention;

[0072] Figure 5 is a structural block diagram of a large model automatic proposition system for self-study examinations disclosed in Embodiment 3 of the present invention;

[0073] Figure 6It is the structural block diagram of the computer device in Embodiment 4 of the present invention. Detailed implementation manners

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0075] Embodiment 1

[0076] Figure 1 It is the flowchart of a large model automatic proposition method for self-study examinations disclosed by the present invention. As Figure 1 shown, a large model automatic proposition method for self-study examinations disclosed in this embodiment includes the following steps: scanning test papers, exercises, and textbooks of different self-study examination subjects, and preprocessing to obtain clear electronic document images; performing optical character recognition on the electronic document images and archiving the text content; establishing a mapping relationship between the textbook content and subject test questions to obtain a textbook document-question document dataset; training an automatic proposition large model using the textbook document-question document dataset; deploying the large model to enable it to be used in an offline scenario; inputting the textbook content and knowledge points to be set questions into the automatic proposition large model trained with the document dataset, and outputting the prepared test questions, answers, and proposition basis related to the textbook content, specifically as follows:

[0077] T1. Scan test papers, exercises, and textbooks of different self-study examination subjects, and preprocess to obtain clear electronic document images;

[0078] T2. Perform optical character recognition on the electronic document images and archive the text content;

[0079] T3. Establish mapping relationships of textbook document-knowledge point-subject test questions and knowledge point-cognitive level-subject test questions to obtain a textbook document-knowledge point-subject test question dataset and a knowledge point-cognitive level-subject test question dataset;

[0080] In this embodiment, the mapping relationship established in T3 adopts the method of keyword matching and text similarity weighting. The specific approach is to divide the textbook into several documents according to sections, and preprocess the documents and real exam questions, including removing stop words, punctuation marks, special characters, and word segmentation; extract keywords from the preprocessed documents and real exam questions to obtain document keyword vectors and real exam question keyword vectors respectively, and use methods such as TF-IDF calculation to obtain keyword matching degrees; encode the documents and real exam questions respectively, and calculate the cosine similarity between the textbook fragment vectors and real exam question vectors; perform weighted summation on the keyword matching degree and cosine similarity, select the weights of the keyword matching degree and cosine similarity, calculate the comprehensive score of each real exam question with all documents, and select the document with the highest comprehensive score as the mapping result. Taking the "Hotel Strategic Management Tutorial" in the self-study subject as an example, the textbook is "Hotel Strategic Management Tutorial", which is divided into 51 documents according to sections. The knowledge point is "Requirements for formulating hotel strategic goals", and the real exam question is a multiple-choice question: "The requirements for formulating hotel strategic goals are: A. Specificity B. Measurability C. Achievability D. Irrelevance". Through TF-IDF calculation of the textbook document and the real exam question, the weight of the keyword matching degree is 0.4, and the weight of the cosine similarity is 0.6. In the third chapter of the textbook document, "Hotel Strategic Mission and Goals", the second section, "Hotel Strategic Goals", has the highest score, with a keyword matching degree of 0.95 and a cosine similarity of 0.86. The weighted summation gives a final score of 0.896. Compared with the traditional manual mapping method, the similarity establishment method of the present invention has an 8% increase in mapping accuracy, a 98% reduction in single-chapter processing time, and a 65% increase in cross-chapter association.

[0081] T4. Use the textbook document - knowledge point - subject exam question dataset and the knowledge point - cognitive level - subject exam question dataset to train the large model;

[0082] In this embodiment, the specific training process in T4 is as follows:

[0083] T401. Initialize the automatic proposition large model, set the number of training rounds to 3, truncate the input textbook document according to the maximum length of 10,000 characters, select the AdamW optimizer, set the batch size to 8, and set the gradient accumulation step to 8 to obtain the initial training settings;

[0084] T402. Take batches as units, divide the textbook - exam question document dataset into several batches, randomly select a batch of textbook - exam question document label pairs from the divided dataset, and input them into the large model for forward propagation to obtain the predicted text;

[0085] T403. Use the large model loss function to calculate the difference between the target text and the predicted text, and use the error backpropagation algorithm to update the large model parameters;

[0086] T404. Save the checkpoint every 2000 training steps and evaluate the performance of the current training times. If the performance does not improve and reaches the early stopping threshold, stop the training in advance;

[0087] T405. During the training process, continuously update the learning rate according to the training situation, and obtain a new learning rate based on the learning rate decay strategy.

[0088] T5. Input the teaching material content and knowledge points to be set questions into the trained large model, and output the prepared test questions, answers and proposition bases related to the teaching material content.

[0089] Table 1. Comparison table of the accuracy of manual mapping, single chapter processing time and cross-chapter association of the present invention with traditional manual mapping under 5 self-study examination subjects

[0090]

[0091] In summary, as shown in Table 1, an automatic proposition method of a large model for self-study examination disclosed in this embodiment is verified and compared with the traditional manual proposition method under 5 self-study examination subjects. The artificial mapping accuracy of the mapping model of the present invention is 90%, and the accuracy of traditional manual mapping is 82%, which is 8 percentage points higher than that of traditional manual mapping; the single chapter processing time of the mapping model of the present invention is 8.2 seconds, and the single chapter processing time of traditional manual mapping is 7 minutes. The single chapter processing speed of the mapping model of the present invention has been improved by an order of magnitude. In addition, while the processing speed is improved, the cross-chapter association of the examination propositions obtained by the mapping model of the present invention is 70, while the cross-chapter association of traditional manual mapping is only 24. It can be seen from Table 1 that the present invention effectively improves the efficiency of preparing self-study examination questions and realizes accurate proposition for different subjects in self-study examinations.

[0092] Embodiment 2

[0093] Based on the automatic proposition method of a large model for self-study examination disclosed in Embodiment 1, this embodiment continues to refer to steps T1 to T5 of the automatic proposition method of the large model for self-study examination disclosed in Embodiment 1, as Figure 1 shown, the method includes the following steps:

[0094] T1. Scan the test papers, self-study examination exercises and self-study examination textbooks of different self-study examination subjects, and preprocess to obtain clear electronic document images;

[0095] T2. Recognize the electronic document images through optical character recognition and file the text content;

[0096] T3. Establish the mapping relationship between teaching material documents - knowledge points - subject test questions and knowledge points - cognitive levels - subject test questions to obtain the teaching material document - knowledge points - subject test question dataset and knowledge points - cognitive level - subject test question dataset;

[0097] T4. Use the teaching material document - knowledge points - subject test question dataset and knowledge points - cognitive level - subject test question dataset to train the large model;

[0098] T5. Input the teaching material content and knowledge points to be tested into the trained large model, and output the prepared test questions, answers, and proposition basis related to the teaching material content.

[0099] Among them, for the proposition in step T5, taking the self - taught examination subject "Educational Communication" as an example, the input teaching material is Chapter 6 "Educational Communication Environment" of "Educational Communication", Section 2 "Functions of Educational Communication Environment", and the knowledge point to be propositioned is "Basic Functions of Educational Communication Environment". Inputting the teaching material document and knowledge points into the model, the proposition result of the single - choice question is: "Which of the following belongs to the content of the basic functions of the educational communication environment? () A. Interaction function B. Environmental function C. Incentive function D. Regulation function ". Among them, the corresponding true question for this teaching material fragment is a multiple - choice question, examining the five basic functions of the educational communication environment: expansion function, incentive function, edification function, intelligence - enhancing function, and enhancement function. The knowledge points examined by the large - model proposition are close to the true - question content, and at the same time, the question types are more diverse, and the proposition quality is relatively high.

[0100] Among them, the experimental results of step T5 compared with the existing untrained model are shown in Table 2. Experiments are carried out on 5 self - taught examination subjects (Educational Communication, Hotel Strategic Mission and Goals, Fundamentals of Management, Human Resource Development and Management, and Introduction to Tour Guide). Let each model proposition according to a teaching material and knowledge points to form a complete test paper, and finally obtain the final knowledge - point coverage rate and difficulty coefficient results through blind evaluation by education experts.

[0101] The large model for self - taught examination proposition is superior to the untrained model in these two evaluation indicators and is similar to the true - question proposition quality.

[0102] Table 2. Comparison table of knowledge - point coverage rate and difficulty coefficient between the present invention and other methods under 5 self - taught examination subjects

[0103]

[0104] In summary, as described in Table 2, a large model automatic proposition method for self-study examinations disclosed in this embodiment is compared with a large model that has not been trained with a teaching material document - knowledge point - subject question dataset and a knowledge point - cognitive level - subject question dataset under 5 self-study examination subjects, and two core proposition indicators, namely knowledge point coverage rate and difficulty coefficient, are analyzed. The knowledge point coverage rate of the model of the present invention reaches 95.8%, which is 23.5 - 43.5 percentage points higher than that of the untrained model (such as Llama2-7B-Chat 52.3% and Baichuan-7B 75.6%), approaching the 100% coverage rate of the real exam papers; in terms of controlling the difficulty coefficient, the difficulty of the questions output by this model is stable at 0.61, which is highly consistent with 0.6 of the real exam papers, while the untrained model shows significant fluctuations (such as ChatGLM-6B 0.44 and Qwen2-7B 0.64). Experiments show that the model jointly trained with the three-way data of teaching materials - knowledge points - questions not only retains the relevance of knowledge points but also significantly improves the matching accuracy between the difficulty of the questions and the syllabus. Its proposition quality is improved by 2 - 3 orders of magnitude compared with general large models, providing technical support for building a standardized and normalized self-study examination proposition system.

[0105] Embodiment 3

[0106] As Figure 5 shown, this embodiment provides a large model automatic proposition system for self-study examinations. The large model automatic proposition system includes: a scanning preprocessing module 501, an optical character recognition module 502, a mapping relationship and dataset construction module 503, a large model training module 504, and a self-study examination question proposition module 505. The specific functions of each module are as follows:

[0107] The scanning preprocessing module 501 scans test papers, exercises, and teaching materials for different self-study examination subjects and preprocesses them to obtain clear electronic document images;

[0108] The optical character recognition module 502 archives the text content by optically recognizing the electronic document images;

[0109] The mapping relationship and dataset construction module 503 establishes the mapping relationships of teaching material document - knowledge point - subject questions and knowledge point - cognitive level - subject questions to obtain a teaching material document - knowledge point - subject question dataset and a knowledge point - cognitive level - subject question dataset;

[0110] The large model training module 504 trains the large model using the teaching material document - knowledge point - subject question dataset and the knowledge point - cognitive level - subject question dataset;

[0111] The self-study examination question creation module 505 inputs the teaching material content and knowledge points to be questioned into a trained large model, and outputs the created questions, answers, and question-setting basis related to the teaching material content.

[0112] Example 4

[0113] This embodiment provides a computer device, which can be a computer, such as Figure 6 shown, which is connected by a system bus 601 to a processor 602, a memory, an input device 603, a display 604, and a network interface 605. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 606 and an internal memory 607. The non-volatile storage medium 606 stores an operating system, computer programs, and a database. The internal memory 607 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. When the processor 602 executes the computer programs stored in the memory, it implements the method for automatically creating questions using a large model for self-study examinations proposed in the above-mentioned Embodiment 1. The method for automatically creating questions using a large model for self-study examinations includes the following steps:

[0114] T1. Scan test papers, exercises, and textbooks for different self-study examination subjects, and preprocess them to obtain clear electronic document images;

[0115] T2. Recognize the text content of the electronic document images through optical character recognition and file them;

[0116] T3. Establish the mapping relationships between teaching material documents - knowledge points - subject questions and knowledge points - cognitive levels - subject questions, and obtain the teaching material document - knowledge points - subject question dataset and the knowledge points - cognitive level - subject question dataset;

[0117] T4. Use the teaching material document - knowledge points - subject question dataset and the knowledge points - cognitive level - subject question dataset to train the large model;

[0118] T5. Input the teaching material content and knowledge points to be questioned into the trained large model, and output the created questions, answers, and question-setting basis related to the teaching material content.

[0119] Example 5

[0120] This embodiment provides a storage medium, which is a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it implements the method for automatically creating questions using a large model for self-study examinations in the above-mentioned Embodiment 1. The method for automatically creating questions using a large model for self-study examinations includes the following steps:

[0121] T1. Scan test papers, exercises, and textbooks of different self-study examination subjects, and preprocess them to obtain clear electronic document images;

[0122] T2. Recognize the text content of the electronic document images through optical character recognition and file it;

[0123] T3. Establish the mapping relationships between textbook documents - knowledge points - subject questions and knowledge points - cognitive levels - subject questions to obtain the textbook document - knowledge points - subject question dataset and the knowledge points - cognitive level - subject question dataset;

[0124] T4. Use the textbook document - knowledge points - subject question dataset and the knowledge points - cognitive level - subject question dataset to train the large model;

[0125] T5. Input the textbook content and knowledge points to be tested into the trained large model, and output the prepared questions, answers, and proposition bases related to the textbook content.

[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0127] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0128] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A large-model automatic question setting method for self-study examinations, characterized in that: The proposition method comprises the following steps: S1. Scan the test papers, exercises and teaching materials of different self-study examination subjects, and pre-process them to obtain clear electronic document images; S2, electronically image the document and archive the text content through optical character recognition; S3, establish the mapping relationship between textbook document-knowledge point-subject test question and knowledge point-cognitive level-subject test question, and obtain the textbook document-knowledge point-subject test question data set and the knowledge point-cognitive level-subject test question data set; S4, use the textbook document-knowledge point-subject test data set and the knowledge point-cognitive level-subject test data set to train the large model; S5. Input the textbook content and knowledge points to be set into the trained large model, and output the test questions, answers and basis related to the textbook content.

2. The large-scale automatic question setting method for self-study examination according to claim 1 is characterized in that: The large model includes a text encoding layer and a model backbone connected in sequence, wherein: The text encoding layer uses byte pair encoding as a word segmentation method to encode the input text, and encodes different textbooks and test questions of different lengths into initial text vectors of the same dimension. The text encoding layer is denoted as G( ), defined as follows: in, Represents the initial input textbook document content, length is , represents the initial text vector after the text encoding layer, Included The dimensions are The word embedding vector of The rotation position encoding is introduced based on the initial text vector and is defined as follows: in, represents the frequency parameter used for rotational position encoding, Indicates the index number, =1,2…,d-1,, represents the floor function, so the rotation position encoding And the final text vector The definition is as follows: in, Represents the position index of each element in the input text. and Respectively The middle position is and The word embedding vector of and denote the cosine function and the sine function respectively; The model backbone includes 32 layers of network modules connected in sequence, wherein each layer of the network module includes a first root mean square layer, an attention layer, a second root mean square layer and a multilayer perceptron, and the first root mean square layer and the second root mean square layer are used to scale the input vector, and the scaling operation formula is defined as follows: in, is a text vector No. elements, It is with The learnable parameters corresponding to the elements are Indicates the number of elements in the input vector; The attention layer adopts a group query attention mechanism, where each group of queries shares the same key and value, defined as follows: in, is the query matrix, which is divided into Groups, represents the index of the group, =1,2…,G, and Respectively represent The key matrix and value matrix of each group, multiple query matrices in the same group correspond to the same key matrix and value matrix. represents the normalized exponential function, Represents the dimension of the key, used to scale the dot product to stabilize the gradient; The multilayer perceptron consists of three linear layers and an activation function. The activation function is used to capture the nonlinear relationship in the input data. The multilayer perceptron performs weighted processing on the information at different positions in the output sequence of the attention layer to form a vector representation with richer information, which is defined as follows: in, represents the vector of the output of the second RMS layer, , and denote the weight matrices of the first, second and third linear layers respectively, , and denote the bias vectors of the first, second and third linear layers respectively, represents the output of the first linear layer, represents the output of the second linear layer, represents the SwiGLU activation function, which takes the output of the first linear layer as input, Represents element-wise multiplication between matrices.

3. The large-scale automatic question setting method for self-study examination according to claim 2 is characterized in that: The working process of the large model is as follows: The text encoding layer accepts the initial input textbook document content The model backbone is connected to the text vector output by the text encoding layer. In the first network module of the model backbone, the output of the first root mean square layer is connected to the input of the attention layer. The input passes through the query matrix and key matrix, dot multiplication, and normalized exponential function in the attention layer. The output of the attention layer is connected to the second root mean square layer. The output of the second root mean square layer is connected to the input of the multilayer perceptron. After passing through the first linear layer and the activation function in the multilayer perceptron, the output of the dot multiplication operation with the input through the second linear layer is connected to the third linear layer to obtain the final output of the first network module. The input of the second network module is connected to the output of the first network module, and so on. After passing through 32 layers of network modules, the final output text is obtained.

4. The large-scale automatic question setting method for self-study examination according to claim 1 is characterized in that: The process of establishing the mapping relationship between the teaching material content, knowledge points and subject questions, and between knowledge points, cognitive levels and subject questions in step S3 is as follows: S301, extracting knowledge points and cognitive levels corresponding to the textbooks from the examination syllabus of different subjects of the self-study examination, and establishing a mapping between the textbooks and the knowledge points and between the knowledge points and the cognitive levels; S302, establishing the mapping relationships of textbook document-knowledge point-subject test question and knowledge point-cognitive level-subject test question through keyword matching operation and text similarity matching operation: Keyword matching operation: Divide the textbook into several documents according to the sections, pre-process the documents and the real questions, including removing stop words, punctuation marks, special characters and word segmentation; extract keywords from the pre-processed documents and real questions, obtain document keyword vectors and real question keyword vectors respectively, and calculate the keyword matching degree through similarity calculation; Text similarity matching operation: Encode the document and the real question separately, calculate the cosine similarity between the textbook segment vector and the real question vector; perform weighted summation of the keyword matching and cosine similarity, select the keyword matching weight and cosine similarity weight, calculate the comprehensive score of each real question with all documents, and select the document with the highest comprehensive score as the mapping result; S303, after obtaining the mapping relationship between textbook document-knowledge point-subject test question and knowledge point-cognitive level-subject test question, according to the question types of the self-study exam real questions, the question types include single-choice questions, multiple-choice questions, fill-in-the-blank questions, short-answer questions, definition questions, and case questions; the above mapping relationship is constructed into the following four types of data formats to obtain the final textbook document-knowledge point-subject test question data set and knowledge point-cognitive level-subject test question data set: (1) Input: Please formulate test questions based on the following textbook content {textbook document}, Output: {subject test questions}; (2) Input: Please formulate test questions based on the following textbook content and knowledge points {textbook document} {knowledge points}, Output: {subject test questions}; (3) Input: Please formulate test questions based on the following knowledge points and cognitive levels {knowledge points} {cognitive levels}, Output: {subject test questions}; (4) Input: Please judge the corresponding cognitive level based on the following subject test questions {subject test questions}, output: {cognitive level}.

5. The large-scale automatic question setting method for self-study examination according to claim 1 is characterized in that: The specific process of using the textbook-test document dataset to train the large model in step S4 is as follows: S401, initialize the automatic question setting model, set the training rounds, input the textbook document and truncate it according to the specified maximum length, select the AdamW optimizer, set the batch size, set the gradient accumulation step, and obtain the initial training settings; S402, dividing the textbook-exam question document data set into several batches in batch units, randomly extracting a batch of textbook-exam question document label pairs from the divided data set, and inputting them into the large model for forward propagation to obtain predicted text; S403, using the large model loss function to calculate the difference between the target text and the predicted text, and using the error back propagation algorithm to update the large model parameters; S404, saving checkpoints every specified number of training steps, and evaluating the performance of the current number of training steps. If the performance does not improve and reaches the early stopping threshold, stop the training in advance; S405: During the training process, the learning rate is continuously updated according to the training situation, and a new learning rate is obtained according to the learning rate decay strategy.

6. The large-scale automatic question setting method for self-study examination according to claim 5 is characterized in that: The large model loss function Defined as: in, represents the large model loss function, represents the large model parameters, , , and Corresponding to the above four types of data sets, , , and Indicates the number of samples used for training in each data set, and In each data set , , and training sample input textbook text fragments, , , and In each data set , , and The subject test questions corresponding to the training samples are and Indicates , The knowledge points corresponding to the training samples are and No. and The cognitive level corresponding to the training samples is is a conditional probability distribution, given the condition In the case of probability.

7. A large-model automatic question setting system for self-study examinations, used to execute the large-model automatic question setting method for self-study examinations as described in any one of claims 1 to 6, characterized in that: The large model automatic proposition system comprises: Scanning preprocessing module, scanning different self-study examination papers, self-study examination exercises and self-study examination teaching materials, and pre-processing to obtain clear electronic document images; Optical character recognition module, which recognizes electronic document images and archives text content through optical character recognition; The mapping relationship and data set construction module establishes the mapping relationship between textbook document-knowledge point-subject test question and knowledge point-cognitive level-subject test question, and obtains the textbook document-knowledge point-subject test question data set and the knowledge point-cognitive level-subject test question data set; The large model training module uses the textbook document-knowledge point-subject test question data set and the knowledge point-cognitive level-subject test question data set to train the large model; The self-study exam question setting module inputs the textbook content and knowledge points to be set into the trained large model, and outputs the set questions, answers and question setting basis related to the textbook content.

8. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, it implements the large-model automatic question setting method for self-study examinations as described in any one of claims 1 to 6.

9. A storage medium storing a program, characterized in that: When the program is executed by the processor, the large-model automatic question setting method for self-study examinations as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intelligent workshop production control method and equipment based on scheduling knowledge self-learning updating

    CN114020861A

  • Exercise generation method and device based on language model and medium

    CN116561260A

  • Automatic test paper generation method and system, electronic equipment and medium

    CN117195867A

  • Model training method and apparatus, knowledge classification method and apparatus, and device and medium

    WO2023108991A1

Cited By

  • Examination proposition person behavior model construction method and system based on deep learning

    CN120822160A