Multi-model collaborative test item output methods, devices, electronic equipment, and media

By using a multi-model collaborative test question generation method, the quality and efficiency issues of generating test questions and test papers using a single deep neural network model are solved, achieving high-quality, low-resource-consumption test question generation and display.

CN121457589BActive Publication Date: 2026-04-03BEIJING MENGJIANXING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the generation of test papers using a single deep neural network model suffers from poor adaptability to different question types, low quality of different question types, incomplete coverage of knowledge points, and untimely updates. This results in low quality test papers, long generation and display times, and the failure to consider device information leads to low display quality and wasted video memory resources.

Method used

By employing a multi-model collaborative approach, through multi-source data fusion, knowledge point association analysis, and multi-hop graph reasoning query, initial question sets for different question types are generated. Multi-dimensional quality assessment and dynamic adjustment are then performed to adapt the display of question papers and dynamically adjust in response to user operations.

Benefits of technology

It improved the quality and efficiency of test paper generation, shortened the generation and display time, reduced the waste of video memory resources, and improved the display quality of test questions and the breadth of knowledge points covered.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121457589B_ABST
    Figure CN121457589B_ABST
Patent Text Reader

Abstract

This disclosure presents embodiments of a test question output method, apparatus, electronic device, and medium based on multi-model collaboration. One specific implementation of the method includes: performing multi-source data fusion on a multi-source test question association dataset and then conducting knowledge point association analysis to obtain a dynamic test question knowledge point graph; performing multi-hop graph reasoning query on the dynamic test question knowledge point graph to obtain a set of test question-related knowledge points; generating an initial test question set by model collaboration using test question generation request information and the set of test question-related knowledge points; evaluating the quality of the initial test question set to obtain a set of test question quality evaluation values; dynamically adjusting the initial test question set to obtain and store an adjusted test question set; and extending and adapting the adjusted test question set for test paper rendering and display to obtain and print the test paper. This implementation can improve the quality of test questions and test papers, increase the efficiency of test question generation, shorten the test question and test paper generation and display time, improve the quality of test question display, and reduce the waste of display resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to a method, apparatus, electronic device, and medium for outputting test questions based on multi-model collaboration. Background Technology

[0002] With the rapid development of deep neural network models, model-based automatic test question generation methods are widely used in educational fields such as assessment, talent selection, and knowledge evaluation. This reduces the workload of test creators and improves test creation efficiency. The typical approach to test question generation is as follows: a test question generation request and individual test question data are sent to a single deep neural network model to obtain an initial test question set. Then, the initial test question set is processed to create a test paper, which is then displayed.

[0003] However, in practice, it has been found that when generating test questions using the above method, the following technical problems often occur: Since test questions need to contain different types of entities, using only a single deep neural network model for test question generation results in poor adaptability to different question types, low quality of test questions of different types, and incomplete coverage and untimely updates of knowledge points in test question material data from a single source. This leads to low quality of test questions, long generation and display times, and failure to consider the device information of the test question output device, resulting in low display quality and wasted video memory resources.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose a method, apparatus, electronic device, and medium for outputting test questions based on multi-model collaboration to solve one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a test question output method based on multi-model collaboration, including: in response to receiving device information groups of various test question output devices, performing multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data, wherein the aforementioned test question output devices include: a display and a printer; performing knowledge point association analysis on the aforementioned fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps; in response to detecting test question generation request information, performing multi-hop graph reasoning query on the aforementioned dynamic test question knowledge point graph to obtain a test question association knowledge point set; and using a test question generation model set, The above-mentioned question generation request information and the above-mentioned question-related knowledge point set are used to generate an initial question set for different question types through model collaboration. A multi-dimensional question quality assessment is performed on the initial question set to obtain a question quality assessment value set. Based on the question quality assessment value set, the initial question set is dynamically adjusted to obtain an adjusted question set, and the adjusted question set is stored in the question database. Based on the above-mentioned device information group, the adjusted question set is displayed for test paper adaptation to obtain the test paper. In response to the detection of target user operation behavior information, the test paper and the adjusted question set are dynamically displayed and adjusted.

[0008] Secondly, some embodiments of this disclosure provide a test question output device based on multi-model collaboration, including: a multi-source data fusion unit configured to, in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data, wherein the aforementioned test question output devices include: a display and a printer; a knowledge point association analysis unit configured to perform knowledge point association analysis on the aforementioned fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps; a multi-hop graph reasoning query unit configured to, in response to detecting test question generation request information, perform multi-hop graph reasoning query on the aforementioned dynamic test question knowledge point graph to obtain a test question association knowledge point set; and a generation unit configured to generate test questions... The system comprises: a model set generation unit, configured to collaboratively generate initial question sets for different question types based on the aforementioned question generation request information and the aforementioned question-related knowledge point set; a multi-dimensional question quality assessment unit, configured to perform multi-dimensional question quality assessment on the aforementioned initial question sets to obtain a question quality assessment value set; a dynamic adjustment unit, configured to dynamically adjust the aforementioned initial question sets based on the aforementioned question quality assessment value set to obtain an adjusted question set, and to store the adjusted question set in the question database; and a test paper adaptation display unit, configured to perform test paper adaptation display on the aforementioned adjusted question set based on the aforementioned device information set to obtain a test paper, and to dynamically adjust the display of the aforementioned test paper and the aforementioned adjusted question set in response to the detection of target user operation behavior information.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0011] The above embodiments of this disclosure have the following beneficial effects: The test question output method based on multi-model collaboration in some embodiments of this disclosure can improve the quality of test questions and test papers, improve the efficiency of test question generation, shorten the test question and test paper generation time and display time, improve the test question display quality, and reduce the waste of display resources. Specifically, the reasons for the low quality of related test questions and test papers, the long time for generating and displaying test questions and test papers, the low display quality, and the waste of video memory resources are as follows: Since test questions and test papers need to contain different types of entities, using only a single deep neural network model for test question generation results in poor adaptability to question types, low quality of test questions of different types, and incomplete coverage and untimely updates of knowledge points in test question material data from a single source. This leads to low quality of test questions and test papers, long time for generating and displaying test questions and test papers, and the failure to consider the device information of the test question output device results in low display quality and waste of video memory resources when the test questions and test papers are displayed. Based on this, some embodiments of the test question output method based on multi-model collaboration disclosed herein can first, in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data. The aforementioned test question output devices include: a display and a printer. This can improve the knowledge point coverage and timeliness of the fused multi-source test question association dataset. Secondly, perform knowledge point association analysis on the aforementioned fused multi-source test question association dataset to obtain a dynamic test question knowledge point graph with timestamps. Here, the dynamic test question knowledge point graph with timestamps introduces time tags, which can improve the historical tracing and real-time updating of knowledge points, and improve the quality and knowledge point coverage of the dynamic test question knowledge graph. Thirdly, in response to detecting test question generation request information, perform multi-hop graph reasoning query on the aforementioned dynamic test question knowledge point graph to obtain a set of test question association knowledge points. This can improve the scope of knowledge points covered by the queried entity association knowledge point set and the accuracy of the queried knowledge point associations, so as to subsequently improve the coverage of more knowledge points in the initial test question set. Subsequently, using a question generation model set, the aforementioned question generation request information and the aforementioned set of related knowledge points are collaboratively generated to produce initial question sets for different question types. Here, multiple question generation models work collaboratively in parallel, each responsible for generating different types of question sections. This improves the efficiency and shortens the time required for initial question set generation. Furthermore, the set of related knowledge points can provide knowledge point guidance to the question generation model set, effectively reducing the model illusion problem and further improving the accuracy and quality of the generated initial question set. Next, a multi-dimensional question quality assessment is performed on the initial question set, resulting in a set of question quality assessment values. Conducting the initial question quality assessment from different aspects improves the accuracy of the assessment values, facilitating subsequent optimization and adjustment of the initial questions.Then, based on the aforementioned test question quality assessment value set, the initial test question set is dynamically adjusted to obtain an adjusted test question set, which is then stored in the test question database. This further improves the quality of the test questions and reduces the waste of database storage resources. Finally, based on the aforementioned device information group, the adjusted test question set is adapted for test paper display to obtain the test paper. In response to the detected user's operational behavior information, the test paper and the adjusted test question set are dynamically adjusted for display. This increases the diversity of different types of test questions included in the test paper, improves the quality of the test paper and the breadth of knowledge points covered, shortens display time, better adapts to the display requirements of the test question output device, improves the display quality of the test paper, and reduces the waste of video memory resources. Therefore, this test question output method based on multi-model collaboration can improve the quality of test questions and test papers, increase the efficiency of test question generation, shorten the test question and test paper generation time and display time, improve the test question display quality, and reduce the waste of display resources by adopting the collaborative work of multi-source test question association datasets and test question generation models. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the test item output method based on multi-model collaboration according to the present disclosure;

[0014] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the test item output device based on multi-model collaboration according to this disclosure;

[0015] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] Figure 1 A flowchart 100 is shown, illustrating some embodiments of a multi-model collaborative test item output method according to this disclosure. This multi-model collaborative test item output method includes the following steps:

[0023] Step 101: In response to receiving the device information group of each test question output device, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data.

[0024] In some embodiments, the execution entity (e.g., an electronic device) of the above-described multi-model collaborative test question output method can, in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data. The various test question output devices include: a display and a printer. The test question output devices can be devices used for visually displaying the generated test questions. The displays can include devices of different sizes. Each test question output device corresponds to one or more users, allowing the test questions to consider not only the personalization of the corresponding user but also the display requirements of the test question output devices during the generation process, thereby improving the quality of subsequent test question display and reducing the waste of video memory resources. The multi-source test question association data in the above-described multi-source test question association dataset can be data collected from different data sources that is related to the test questions. The above-described multi-source test question association dataset can include, but is not limited to, at least one of the following: past exam question data, media news datasets about the test questions, and domain-specific knowledge bases related to the test questions. The aforementioned fused multi-source test item association data can be the test item association data after removing duplicate and conflicting data from the original multi-source test item association dataset. In practice, the executing entity can first perform data preprocessing on the aforementioned multi-source test item association dataset to obtain a preprocessed multi-source test item association dataset. Secondly, using a fuzzy logic data fusion algorithm, the preprocessed multi-source test item association dataset is fused to obtain the fused multi-source test item association data.

[0025] Step 102: Perform knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point map with timestamps.

[0026] In some embodiments, the aforementioned executing entity may perform knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps. This dynamic test question knowledge point graph may be a directed acyclic graph that represents the relationships between various test question knowledge points in the form of a graph (e.g., question type labels, difficulty level labels) and is updated periodically.

[0027] In some optional implementations of certain embodiments, the above-mentioned knowledge point association analysis of the fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps may include the following steps:

[0028] The first step is to perform semantic block segmentation on the above-mentioned fused multi-source test item association data to obtain a test item association semantic block set. The test item association semantic blocks in this set can be data blocks that divide the fused multi-source test item association data into sentences and paragraphs.

[0029] The second step involves generating knowledge extraction prompts and data source relationship extraction prompts based on the aforementioned set of semantic blocks associated with the test questions. The knowledge extraction prompts can be text used to guide the large language model generated from the graph for knowledge extraction. Similarly, the data source relationship extraction prompts can be text used to guide the large language model generated from the graph for knowledge point relationship extraction. In practice, the executing entity can input the aforementioned set of semantic blocks associated with the test questions into the prompt template to obtain the knowledge extraction prompts and data source relationship extraction prompts.

[0030] The third step involves inputting the aforementioned knowledge extraction prompts, data source relationship extraction prompts, and question-related semantic block sets into a large-scale language model for graph generation, resulting in initial knowledge point graph sets for different data sources. The initial knowledge point graph in this set can be a graph containing the entity sets and relationships between entity sets within the question-related semantic blocks, and includes knowledge point update iteration times. The large-scale language model for graph generation can perform named entity recognition and association extraction on the input question-related semantic block sets based on the input knowledge extraction prompts and data source relationship extraction prompts, outputting the initial knowledge point graph set.

[0031] The fourth step involves performing knowledge point fusion processing on the initial knowledge point graph set to obtain an initial knowledge point fusion graph. This initial knowledge point fusion graph can be obtained by fusing different entities or relationships that represent the same meaning. The knowledge point fusion processing can also involve linking entities and relationships representing the same knowledge point but with different expressions to the same entity or relationship.

[0032] The fifth step involves identifying entity-relationship conflicts in the initial knowledge point fusion graph to obtain a set of knowledge point conflict triples. These conflict-resolving triples can be identified triples where entities or attributes are in conflict and represented in the form of <entity, relation / attribute, entity>.

[0033] The sixth step is to perform entropy conflict disambiguation on the above set of conflicting knowledge points to obtain a disambiguated set of knowledge point triplets. This entropy conflict disambiguation can be performed using the entropy filtering algorithm in the TruthfulRAG (Truthful Retrieval-AugmentedGeneration) model.

[0034] Step 7: Input the disambiguated knowledge point triplet set into the initial knowledge point fusion graph to obtain the knowledge point fusion graph, which serves as a dynamic test question knowledge point graph with timestamps. Then, perform reinforcement learning dynamic optimization on the dynamic test question knowledge point graph. This reinforcement learning dynamic optimization can be an environment-based update optimization of the dynamic test question knowledge point graph using reinforcement learning algorithms.

[0035] Step 103: In response to the detection of the question generation request information, perform multi-hop graph reasoning query on the dynamic question knowledge point graph to obtain the question-related knowledge point set.

[0036] In some embodiments, the aforementioned executing entity may, in response to detecting a question generation request, perform a multi-hop graph reasoning query on the aforementioned dynamic question knowledge point graph to obtain a set of question-related knowledge points. The question generation request information may be user-inputted request information for generating questions. This request information may include information such as exam type, question format, question type distribution, score and difficulty, and exam syllabus. The question-related knowledge points in the set of question-related knowledge points may be the basic units constituting the exam syllabus in the question generation request information. For example, the question-related knowledge points may include, but are not limited to, at least one of the following: common sense judgment, verbal comprehension and expression, quantitative reasoning, and judgment theory.

[0037] In some optional implementations of certain embodiments, the above-mentioned multi-hop graph reasoning query of the dynamic test question knowledge point graph to obtain the test question-related knowledge point set may include the following steps:

[0038] The first step involves performing a thought chain decomposition on the set of test point selection information corresponding to the aforementioned test question generation request information, resulting in a sequence of test point selection information. This sequence includes test point selection information with logical dependencies. Specifically, the test point selection information in this set can be questions related to knowledge points within the test point domain of the test question generation request information. For example, the test point selection information could be "Please provide knowledge points related to common sense judgment." This thought chain decomposition can be performed using the Auto-CoT (Automatic Chain of Thought Prompting) method.

[0039] The second step involves selecting the information sequence for the test points mentioned above and performing the following generation steps:

[0040] Sub-step 1: Based on the previous initial query knowledge point set and the dynamic test question knowledge point graph, generate the current query knowledge point set for the test question selection information. The current query knowledge points in the current query knowledge point set can be the knowledge points corresponding to the test question selection information in the dynamic test question knowledge point graph. The previous initial query knowledge points in the previous initial query knowledge point set can be the current query knowledge point set of the previous test question selection information. If the test question selection information is the selection information at the initial position in the test question selection information sequence, then the previous initial query knowledge point set is the entity set included in the test question selection information. As an example, the execution entity can first determine the question type information of the test question selection information. Then, in response to determining that the test question selection information is the selection information at the initial position, perform entity recognition on the test question selection information to obtain the test point selection entity set, which serves as the previous initial query knowledge point set. Convert the previous initial query knowledge point set and question type information into a graph database query statement. The graph database query statement mentioned above can be a Cypher query statement. Finally, the dynamic test question knowledge point graph is queried using the graph database query statement to obtain the current query knowledge point set. In response to the determination that the test question point selection information is not the selection information located in the initial position, the previous initial query knowledge point set and question type information are retrieved.

[0041] Sub-step 2 involves performing contextual dynamic encoding on the test point selection information to obtain a query context embedding vector set. The query context embedding vectors in this set represent the relationship between each test point selection word segment and other test point selection word segments in the test point selection information. This dynamic context encoding can be performed using the BERT (Bidirectional Encoder Representation from Transformers) model.

[0042] Sub-step 3 involves determining the inference step association confidence set of the query context embedding vector set. The inference step association confidence set represents the non-linear relationship between the test point selection word segmentation and the entities included in the dynamic test point knowledge graph. In practice, the execution entity can first input the query context embedding vector set into a linear transformation layer to obtain a dynamic test point query vector set. Next, it determines the dot product similarity between each dynamic test point query vector set and each query context embedding vector in the query context embedding vector set, obtaining a dot product similarity set. Then, it inputs the dot product similarity set into a Softmax function for normalization, obtaining a selection word segmentation confidence set. Finally, it multiplies and superimposes each selection word segmentation confidence set with its corresponding query context embedding vector to obtain the target word segmentation dynamic vector set. The target word segmentation dynamic vector in this set represents the importance of each word in the test point selection information to the current determination step. Finally, the target word segmentation dynamic vector set is input into a multilayer perceptron including a Sigmoid activation function to obtain the inference step association confidence set.

[0043] Sub-step 4 involves generating a knowledge point entity transition matrix based on the association confidence set of the reasoning steps and the current query knowledge point set. This knowledge point entity transition matrix represents the probability weights of the connections between entities corresponding to the current query knowledge points included in the current query knowledge point set. For example, the executing entity can determine whether there is an association relationship between the current query knowledge point set and the dynamic test question knowledge point graph. If no association relationship exists, the corresponding element in the knowledge point entity transition matrix is ​​0; if an association relationship exists, the corresponding element in the knowledge point entity transition matrix is ​​the association confidence score of the reasoning steps.

[0044] Sub-step 5 involves generating a set of multi-hop inference paths for entities based on the knowledge point entity transition matrix and the previous initial query knowledge point set. These multi-hop inference paths characterize the propagation of the inference step association confidence within the dynamic question knowledge point graph. As an example, the executing entity can first encode the previous initial query knowledge point set to obtain a query feature vector set. Then, after determining the product of the query feature vector set and the knowledge point entity transition matrix, normalization is performed to obtain an entity update state vector set. Subsequently, the top preset number of knowledge points with the highest values ​​are selected from the entity update state vector set to obtain the target knowledge point set. This preset number can be pre-set and determined according to specific circumstances; it is not limited here. Then, a beam search algorithm is used to backtrack the knowledge point entity transition matrix to obtain multi-hop inference paths for the target knowledge point set.

[0045] Sub-step 6, in response to determining that the test point selection information is located at the termination position, generates a set of test-related knowledge points based on the entity multi-hop reasoning path set. The hierarchical prompt information mentioned above can include prompt words from the test-related multi-hop reasoning path set. In practice, the executing entity can first convert the entity multi-hop reasoning path set into natural language to obtain a path reasoning text information set. Then, the path reasoning text information set is filled into the prompt word template to obtain hierarchical prompt information. Finally, the hierarchical prompt information is input into the large language model to obtain the test-related knowledge point set.

[0046] Optionally, the above method may further include the following steps:

[0047] In response to the determination that the test point selection information is not the selection information located at the termination position, the test knowledge point set corresponding to the entity multi-hop reasoning path set is determined as the previous initial query knowledge point set, so that the above generation steps are executed again. Thus, the multi-hop graph reasoning query implementation starts with the decomposition of the thought chain, combined with dynamic encoding and the knowledge point entity transfer matrix, which can solve the problems of fragmented demand logic, semantic and graph entity disconnection and irrelevant information interference in the query, and can effectively avoid the problems of entity query deviation, redundancy or omission. The entity multi-hop reasoning path set breaks through the limitation of traditional single-hop query that can only mine surface knowledge points, and realizes the efficient mining of deep-related knowledge points across levels and branches in the dynamic test knowledge point graph, which can improve the comprehensiveness and relevance of knowledge point coverage.

[0048] Step 104: Through the question generation model set, the question generation request information and the set of knowledge points associated with the question are generated collaboratively by the model to generate initial question sets for different question types.

[0049] In some embodiments, the aforementioned executing entity can use a question generation model set to collaboratively generate the question generation request information and the question-related knowledge point set, generating initial question sets for different question types. The question generation models in the aforementioned question generation model set can be deep neural network models that generate different types and numbers of questions based on the input question generation request information and question-related knowledge point set. The aforementioned question generation model set can include, but is not limited to, at least one of the following: a question stem generation model, a distractor option generation model, an answer parsing generation model, and a question difficulty assessment model. The aforementioned question stem generation model can be a model whose input is the question-related knowledge point set and the question generation request information, and whose output is the question stem information. For example, the aforementioned question stem generation model can be the GPT-4o (Generative Pre-trained Transformer 4 omni) model. The aforementioned question answer generation model can be a large language model whose input is the question-related knowledge point set and the question stem information, and whose output is the question answer information. For example, the aforementioned question answer generation model can be the Gemini 2.0 model. The aforementioned distractor generation model can be a large language model that takes a set of knowledge points related to the question and the question stem as input, and outputs a set of correct options and multiple distractor options. This distractor generation model could be the Doubao large language model. The aforementioned answer parsing generation model can be a large language model that takes a question stem, option information, and answer information as input, and outputs answer parsing information about the answer generation process (e.g., the reasons why correct and distractor options are right or wrong). For example, this answer parsing generation model could be the Gemini 2.0 model. The aforementioned question difficulty assessment model can be a large language model that takes a question stem, answer information, a set of question options, and answer parsing information as input, and outputs a predicted question difficulty value. For example, this question difficulty assessment model could be the LLaMa 3 (Large Language Model Meta AI) model. The aforementioned different question types can include subjective questions and objective questions. The aforementioned objective questions can include, but are not limited to, at least one of the following: multiple choice questions, fill-in-the-blank questions, true / false questions, and analogy reasoning questions. The subjective questions mentioned above may include, but are not limited to, at least one of the following: case analysis questions, essay questions, material analysis questions, and writing questions.

[0050] In some optional implementations of certain embodiments, the above-mentioned method of generating initial question sets for different question types by using a question generation model set to collaboratively generate the question generation request information and the question-related knowledge point set may include the following steps:

[0051] The first step is to perform structured parsing on the above-mentioned question generation request information to obtain structured question requirement information. This structured question requirement information can be obtained by converting the text-based question generation request information into JSON format.

[0052] The second step involves performing semantic enhancement detection on the aforementioned set of knowledge points related to the test questions, based on the structured information of the test question requirements, to obtain a set of knowledge point text information. This set of knowledge point text information can be multi-source test question association data that is fused with the structured information of the test question requirements and is part of the test question association knowledge points. In practice, the executing entity can utilize GraphRAG (Graph Retrieval-Augmented Generation) to perform semantic enhancement detection on the aforementioned set of knowledge points related to the test questions, based on the structured information of the test question requirements, to obtain the set of knowledge point text information.

[0053] The third step involves identifying entities to be filled in for questions that are determined to be fill-in-the-blank questions. This process involves recognizing entities to be filled in from the aforementioned knowledge point text information set. These entities can be key knowledge points or concepts corresponding to the knowledge points being tested within the text information. For example, if the text information states "The largest palace in the world is the Forbidden City," then the entity to be filled in could be "Forbidden City."

[0054] The fourth step involves invoking the question stem generation model to generate different types of question stem information sets based on the aforementioned entity set to be filled, the aforementioned structured information of the test question requirements, and the aforementioned set of knowledge point text information. The question stem information in these different types of question stem information sets can be information from the question stems of objective question types such as fill-in-the-blank, multiple-choice, and true / false questions. The question stem generation model can be a large language model that constructs a prompt word project from the input entity set to be filled, the aforementioned structured information of the test question requirements, and the aforementioned set of knowledge point text information, allowing the question stem generation model to generate question stem information sets. As an example, the executing entity can first fill the aforementioned entity set to be filled, the aforementioned structured information of the test question requirements, and the aforementioned set of knowledge point text information into the prompt word project template to generate a question stem generation prompt word set. Then, the question stem generation prompt word set is input into the question stem generation model to obtain different types of question stem information sets.

[0055] The fifth step involves invoking the answer generation model to generate different types of question answer information sets based on the aforementioned question stem information set and knowledge point text information set. The question answer information in these sets can be the information corresponding to the question stem information. The answer generation model can be a large language model that constructs prompt word engineering from the input question stem information set and knowledge point text information set to generate the question answer information corresponding to the question stem information set, as well as the solution ideas and logical information related to the question stem. As an example, the executing entity can first fill the question stem information set and knowledge point text information set into the prompt word engineering template to generate a question answer prompt word set and an answer analysis prompt word set. Then, the question answer prompt word set and the answer analysis prompt word set are input into the answer generation model to obtain different types of question answer information sets.

[0056] Step 6: In response to the question being determined to be a multiple-choice question, the distractor option generation model is invoked. This model performs multi-strategy generation on the question answer information set corresponding to the multiple-choice question, resulting in a distractor option text information set. The distractor option text information in this set can be information from the multiple-choice questions other than the correct options, used to confuse and mislead test-takers. The distractor option generation model can be a large language model that applies different strategies to the input question answer information set to generate different distractor option text information. These different strategies can be: a concept / knowledge point confusion strategy, i.e., extracting relevant but not corresponding to the initial question, correct knowledge points from the knowledge point text information set as distractors; an attribute manipulation strategy, i.e., changing a correct attribute in the question answer information to an incorrect attribute as an option; or an answer logic inversion strategy, i.e., reversing the logical relationship of the question answer information as an option.

[0057] Step 7: Combine the above-mentioned question stem information set, question answer information set, and distractor option text information set to obtain initial question sets for different question types. This combination process can involve combining the question stem information set, question answer information set, and corresponding distractor option text information set into complete multiple-choice questions, or combining the question stem information set and corresponding question answer information set into complete fill-in-the-blank or true / false questions.

[0058] In addressing the aforementioned technical problems in the application scenario of large-scale standardized examinations (e.g., the National College Entrance Examination), the following technical issues often arise: The generation of distractor options in multiple-choice questions during large-scale standardized examinations can assess the quality of the questions. However, existing technologies that automatically generate multiple-choice questions using large language models suffer from disconnects between distractor options and the question stem information and correct answer, resulting in either excessively low or high deceptiveness, homogenization, and illusion problems. This leads to low-quality generated options, further reducing the overall quality of the multiple-choice questions. Repeated question generation is necessary, extending the question generation and display time, resulting in lower display quality and wasting device memory resources. Based on the characteristics of this application scenario—multi-strategy generation of distractor options, de-homogenization, diversity, and a knowledge graph linking knowledge points—we have decided to adopt the following solution:

[0059] In some optional implementations of certain embodiments, the above-mentioned method of generating initial question sets for different question types by using a question generation model set to collaboratively generate the question generation request information and the question-related knowledge point set may include the following steps:

[0060] The first step involves using a question generation model set to collaboratively generate question sets for fill-in-the-blank, multiple-choice, and true / false questions based on the aforementioned question generation request information and the associated knowledge point set. This collaborative generation can be achieved by calling the API (Application Programming Interface) of the question generation model set in a parallel or serial manner, allowing the question generation models to mutually call and complement each other for collaborative generation.

[0061] The second step is to perform the following sorting steps for each multiple-choice answer in the multiple-choice answer information set included in the above multiple-choice question set:

[0062] Sub-step 1: Based on the multiple-choice question answer information, perform a proximity search on the dynamic test question knowledge point map to obtain a group of candidate answer information. The candidate answer information in this group can be information from knowledge points located in the dynamic test question knowledge point map that have one hop and two adjacency relationships with the multiple-choice question answer information. This group of candidate answer information can also contain answer information that has a storage inclusion relationship, a proximity relationship, or a functional similarity relationship with the multiple-choice question answer information. For example, if the multiple-choice question answer information is "thylakoid membrane of chloroplasts," the candidate answer information group can include: "chloroplasts" and "grana" with a membership relationship, "chloroplast stroma" with an adjacency relationship, and "mitochondrial inner membrane" with a functional similarity relationship.

[0063] Sub-step 2 involves determining the aforementioned candidate answer information group and the aforementioned multiple-choice answer information, along with the corresponding multiple-choice question's option semantic relevance group, option filter value group, stem semantic relevance group, and conditional semantic relevance group. Specifically, the option semantic relevance group characterizes the semantic matching degree between the candidate answer information and the multiple-choice answer information. The option filter value group uses 0 or 1 to represent whether the candidate answer information and the multiple-choice answer satisfy the conditions of being the same type of option (e.g., knowledge point corresponding to vocabulary, option categories being both numerical or verbal), having an inclusion relationship, or being synonyms of the multiple-choice answer; 1 indicates satisfaction, and 0 indicates dissatisfaction. The stem semantic relevance group characterizes the semantic matching degree between the candidate answer information and the multiple-choice question. The conditional semantic relevance group characterizes the semantic relevance between the candidate answer information and the multiple-choice answer information under given multiple-choice question conditions. In practice, the aforementioned execution entity can perform the following determination steps for each candidate answer information: First, using the WordNet semantic database and the GloVe (Global Vectors for Word Representation) algorithm, determine the candidate answer information, multiple-choice answer information, the similarity of the first option, the similarity of the second option, the similarity of the first stem, and the similarity of the second stem for each multiple-choice question. Second, perform a weighted sum of the first and second option similarities to obtain the option semantic relevance, and perform a weighted sum of the first and second stem similarities to obtain the stem semantic relevance. Next, represent the candidate answer information, multiple-choice answer information, and multiple-choice questions using vectors to obtain candidate option vectors, answer vectors, and question vectors. Subsequently, determine the answer vector difference between the question vector and the answer vector, and the candidate vector difference between the question vector and the candidate option vector. Then, determine the cosine similarity between the answer vector difference and the candidate vector difference to obtain conditional semantic relevance. Finally, use a named entity recognition model to determine the option filtering value.

[0064] Sub-step 3 involves performing initial screening of the above-mentioned candidate answer information group based on the semantic relevance groups of the options, the semantic relevance groups of the question stem, the filter value groups of the options, and the semantic relevance groups of the conditions, to obtain the screened candidate answer information group of the question.

[0065] As an example, the aforementioned implementing entity can, in the first step, perform the following filtering steps for each candidate answer information: First, using the entropy weight method, determine the first, second, and third weight values ​​for the corresponding option semantic relevance, question stem semantic relevance, and conditional semantic relevance. Second, determine the sum of the products of the first weight value and the option semantic relevance set, the second weight value and the question stem semantic relevance set, and the difference between the third weight value and the target conditional semantic relevance set, as the first option decoy value. Here, the target conditional semantic relevance can be the difference between 1 and the conditional semantic relevance. Then, determine the product of the first option decoy value and the corresponding option filtering value, as the second option decoy value. In the second step, filter out the top 3 candidate answer information with the largest corresponding second option decoy values ​​from the candidate answer information set, as the filtered candidate answer information group.

[0066] Sub-step 4 involves using a distractor generation strategy rule engine to execute multiple strategies in parallel on the multiple-choice answer information, resulting in a strategy answer information set. The strategy answer information in this set can be option information that is misleading and distracting compared to the multiple-choice answers. The distractor generation strategy rule engine can be a rule engine composed of distractor option generation strategy information. These distractor option generation strategies can include: conceptual knowledge point confusion strategies, attribute manipulation strategies, answer logic inversion strategies, and common error simulation strategies—strategies based on errors commonly made by test takers statistically analyzed from past exam data.

[0067] Sub-step 5 involves performing diversity ranking processing on the above-mentioned strategy answer information group and the above-mentioned filtered candidate answer information group to obtain the interference option information group. The interference option information in the interference option information group can be the three candidate answer information that are most misleading and distracting compared to the multiple-choice answer information. In practice, the implementing entity can first perform deduplication processing on the above-mentioned strategy answer information group and the above-mentioned filtered candidate answer information group to obtain a deduplicated candidate answer information set. Then, using a large language model, the deduplicated candidate answer information set is subjected to option quality evaluation processing to obtain an option evaluation value set. The option quality evaluation processing can include evaluations based on dimensions such as option rationality, misleadingness, ambiguity, and distinguishability from the multiple-choice answer information. Finally, the top three deduplicated candidate answer information with the highest corresponding option evaluation values ​​are selected from the deduplicated candidate answer information set as the interference option information group.

[0068] The third step involves combining the aforementioned set of fill-in-the-blank questions, multiple-choice questions, true / false questions, and the obtained set of distractor options to obtain an initial question set. Then, based on the aforementioned device information set, the initial question set is displayed in an adaptive manner.

[0069] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem that "the quality of the generated multiple-choice options is low, which leads to a decrease in the quality of the multiple-choice questions, requiring repeated generation of multiple-choice questions to extend the question generation and display time, resulting in low display quality and wasting device memory resources." The factors that cause the low quality of the generated multiple-choice options, leading to a decrease in the quality of the multiple-choice questions, requiring repeated generation of multiple-choice questions to extend the question generation and display time, and wasting device memory resources are often as follows: In large-scale standardized examinations, the generation of distractor options for multiple-choice questions can assess the quality of the multiple-choice questions. However, in existing technologies, the automatic generation of multiple-choice questions through large language models suffers from a disconnect between the distractor options and the question stem information and the correct answer, resulting in either too low or too high deceptiveness, homogenization, and illusion problems. This leads to low quality of the generated multiple-choice options, which in turn leads to a decrease in the quality of the multiple-choice questions, requiring repeated generation of multiple-choice questions to extend the question generation and display time, and wasting device memory resources. If the above factors are addressed, the quality of generated multiple-choice options and the quality of the multiple-choice questions can be improved, the number of multiple-choice question generation times can be reduced, the question generation and display time can be shortened, the display quality of the questions can be improved, and the waste of video memory resources can be reduced. To achieve this effect, this disclosure firstly, through a question generation model set, collaboratively generates fill-in-the-blank question sets, multiple-choice question sets, and true / false question sets, which can improve the efficiency and quality of generating different types of questions. Secondly, for each multiple-choice answer information, the first step is to perform a proximity search on the dynamic question knowledge point graph, which can efficiently search for knowledge points that are structurally directly related to the multiple-choice answer information (containing, adjacent, functionally similar), which can improve the quality of the candidate answer information group and avoid blindly searching in the entire knowledge space. The second step is to determine the option semantic relevance group, option filter value group, question stem semantic relevance group, and condition semantic relevance group, which considers the relevance between the candidate answer information and the multiple-choice answer information and the multiple-choice question stem information, which can improve the deceptiveness of the candidate answers, and the filtering of the candidate answer information can reduce the computational load of subsequent option generation. The third step utilizes a distractor generation strategy rule engine to generate strategy answer information groups. This engine directly simulates high-frequency error points, generating highly deceptive distractors that are difficult to detect through association alone, thus enriching the diversity and relevance of multiple-choice options. The fourth step involves diversity ranking. Through diversity ranking and global evaluation using a large language model, the best options are selected from the deduplicated candidate answer set, ensuring the diversity and discriminative power of the options. This avoids homogenization of options, improves the quality of the options, and ultimately enhances the overall quality of the multiple-choice questions.Finally, by adapting the initial question set to the device information group, the number of generation times can be reduced, the efficiency of question generation can be improved, and the time for question generation and display can be shortened. By considering the device information of the question output device for display, the quality and adaptability of question display can be improved, and the waste of video memory resources can be reduced.

[0070] In some optional implementations of certain embodiments, the above-mentioned method of generating initial question sets for different question types by using a question generation model set to collaboratively generate the question generation request information and the question-related knowledge point set may include the following steps:

[0071] The first step involves generating a set of prompts for multiple types of subjective questions based on the aforementioned question generation request information and the set of related knowledge points. These prompts can guide the question generation model set to generate different types of subjective questions (e.g., essay questions, material analysis questions, short answer questions, compound questions, etc.) while maintaining the diversity of subjective questions—that is, generating prompts from different perspectives. In practice, the executing entity can input the question generation request information and the set of related knowledge points into a subjective question generation prompt template to obtain the set of prompts for multiple types of subjective questions. This template can be composed of the user-input question generation request information, the set of related knowledge points, and system prompts from the question generation model. These system prompts may include, but are not limited to, at least one of the following: multiple argumentation angles for subjective question generation, argumentation rigor, subjective question output format, and model role information.

[0072] The second step involves performing reinforcement serialization processing on the aforementioned set of multiple types of prompts for subjective questions to obtain a target set of multiple types of prompts. The target multiple types of prompts in this target set can be prompt word information obtained by optimizing the environment of the set of multiple types of prompts for subjective questions using a reinforcement learning algorithm.

[0073] The third step is to determine the call path set of the question generation model set based on the aforementioned multi-type prompt information set. The call paths in this set represent the order in which the question generation model set is invoked to generate subjective questions. As an example, the execution entity can use Docker orchestration technology to determine the call path set of the question generation model set based on the aforementioned multi-type prompt information set.

[0074] The fourth step involves calling the question generation model set corresponding to the aforementioned call path set to generate a set of subjective question stem text information, a set of subjective question answer text information, and a set of answer score text information. The answer score text information in the answer score text information set can be information derived from parsing the subjective question answer text information, i.e., the knowledge points included in the subjective question answer and the score for each knowledge point.

[0075] The fifth step involves performing a consistency check on the aforementioned subjective question stem text information set and the aforementioned subjective question answer text information set to obtain a question stem answer detection result set. The question stem answer detection results in this set can represent whether the subjective question answer text information truly answers the question posed by the subjective question stem text information and whether the logic is self-consistent. This consistency check can be performed using the self-consistency (SC) method of a large language model.

[0076] Step 6: Based on the above set of question stem and answer detection results, the above set of subjective question stem text information and the above set of subjective question answer text information are adjusted to remove homogenization, resulting in adjusted subjective question stem text information and adjusted subjective question answer text information. In practice, the above-mentioned implementing entity can first use a diversity reinforcement perception learning model to filter out the subjective question stem text information and subjective question answer text information sets corresponding to consistent detection results from the above set of question stem and answer detection results, and then perform homogenization adjustment to obtain adjusted subjective question stem text information sets and adjusted subjective question answer text information sets. The above-mentioned diversity reinforcement perception learning model can be a reinforcement learning model that uses the accuracy and degree of difference of the subjective question stem text information set and the subjective question answer text information set as the reward function.

[0077] Step 7: Determine the adjusted subjective question stem text information set, the adjusted subjective question answer text information set, and the answer scoring text information set as the initial question sets for different question types.

[0078] Step 105: Perform a multi-dimensional quality assessment on the initial question set to obtain a numerical set of question quality assessments.

[0079] In some embodiments, the executing entity may perform a multi-dimensional test quality assessment on the initial test question set to obtain a test quality assessment numerical set. The test quality assessment values ​​in this numerical set can characterize the difficulty level of the initial test questions.

[0080] In some optional implementations of certain embodiments, the above-mentioned multi-dimensional test item quality assessment of the initial test item set to obtain a test item quality assessment numerical set may include the following steps:

[0081] The first step involves using a question difficulty prediction model to assess the difficulty of the initial question set, resulting in a set of question difficulty assessment values. These values ​​characterize the difficulty level of the initial questions. The question difficulty prediction model can be a Transformer model that assesses the difficulty level of the input initial question set and outputs question difficulty assessment values.

[0082] The second step involves using a question similarity prediction model to evaluate the knowledge point similarity of the initial question set, resulting in a set of question similarity evaluation values. These value sets characterize the degree of similarity between the knowledge points tested by the initial questions. The question similarity prediction model can be a deep neural network model that calculates similarity between the input initial question set and outputs a set of question similarity evaluation values. Alternatively, the question similarity prediction model can be a BERT model.

[0083] The third step involves extracting multimodal features from the initial test item set to obtain a test item feature vector set. The test item feature vectors in this set represent the semantic, syntactic, and syntactic information of the initial test items using vectors.

[0084] The fourth step involves evaluating the content quality of the initial question set based on the aforementioned feature vector set, resulting in a numerical set of question content quality assessments. These numerical values ​​characterize the basic quality of the initial questions (e.g., fluency, coherence, formatting), factual accuracy, completeness, and model illusion prevention. In practice, the executing entity can first perform a semantic quality assessment on the initial question set using regular expressions, obtaining a numerical set of semantic assessments. Then, it can use a large language model to perform a factual quality assessment on the initial question set, obtaining a numerical set of factual assessments. Finally, the average of the semantic assessment value and the corresponding factual assessment value for each question is determined, resulting in the mean assessment value, which serves as the question content quality assessment value.

[0085] The fifth step involves performing a knowledge point matching assessment on the initial question set based on the aforementioned question feature vector set, resulting in a knowledge point matching assessment numerical set. The knowledge point matching assessment values ​​in this set can characterize the cognitive level of the initial questions (e.g., Bloom's goal matching degree, depth of thinking) and the degree of knowledge point relevance. This knowledge point matching assessment can be conducted using IRT (Item Response Theory) theory.

[0086] Step 6: Based on the aforementioned feature vector set of test items, perform bias detection processing on the initial test item set to obtain a set of test item bias evaluation values. The test item bias evaluation values ​​in this set can characterize the bias, content sensitivity, and ethical risk of the initial test items. This bias detection processing can be performed using an LLM-as-a-Judge model.

[0087] Step 7: Perform multi-feature fusion on the above-mentioned numerical sets of test question difficulty assessment, test question similarity assessment, test question content quality assessment, knowledge point matching assessment, and test question bias assessment to obtain the test question quality assessment numerical set. In practice, the executing entity can perform a weighted summation of the above-mentioned numerical sets of test question difficulty assessment, test question similarity assessment, test question content quality assessment, knowledge point matching assessment, and test question bias assessment to obtain the test question quality assessment numerical set.

[0088] In addressing the technical problems mentioned above, the application scenario of constructing an educational institution's exam question database for question quality assessment often presents the following challenges: The database requires a large number of diverse and high-quality questions. However, relying solely on model evaluation for initial question quality assessment suffers from low accuracy due to issues like illusions, factual errors, and logical jumps. This results in low-quality questions and necessitates multiple question displays, wasting memory resources and further compromising memory quality. Based on the characteristics of this application scenario—question type diversity, fairness in question quality assessment, knowledge point adaptability, and human-computer collaborative assessment—we have decided to adopt the following solution:

[0089] Optionally, the above-mentioned multi-dimensional assessment of the quality of the initial test item set to obtain a numerical set of test item quality assessments may include the following steps:

[0090] The first step involves extracting features from the initial test item set, resulting in a semantic feature vector set, a grammatical feature vector set, and a syntactic feature vector set. The semantic feature vectors in the semantic feature vector set represent the semantic information of the vocabulary included in the initial test items. The grammatical feature vectors in the grammatical feature vector set represent the grammatical structure of the initial test items (e.g., part-of-speech distribution, noun and verb ratios). The syntactic feature vectors in the syntactic feature vector set represent the syntactic structure information and dependency relationships between words in the initial test items. This feature extraction can be achieved by using a word embedding model to extract the semantic feature vector set, which is then input into a syntactic parser to obtain the grammatical and syntactic feature vector sets.

[0091] The second step involves inputting the aforementioned semantic feature vector set, syntactic feature vector set, and syntactic feature vector set into a multi-dimensional quality prediction model to obtain a numerical set of predicted test item quality. This multi-dimensional quality prediction model includes: a bidirectional long short-term memory neural network, a fact consistency verification network, a test item logic verification network, a test item content density prediction network, and a two-layer attention mechanism network. The predicted test item quality values ​​in this numerical set characterize the quality level of the initial test items predicted by the model. This multi-dimensional quality prediction model can be a deep neural network model that concatenates the input semantic, syntactic, and syntactic feature vector sets to obtain a concatenated feature vector set, then performs multi-dimensional evaluation and prediction to output the numerical set of predicted test item quality. The bidirectional long short-term memory neural network can be a model used to extract sequence features and contextual information from the input concatenated feature vector set. The fact consistency verification network can be a deep neural network model that matches the input concatenated feature vector set with the knowledge points included in the dynamic test item knowledge point graph to identify and output information about potential factual errors. For example, the fact consistency verification network could be the Factcheck-GPT model. The aforementioned question logic validation network can be a model that checks for logical jumps or conflicts in questions, semantic similarity between questions, and difficulty levels based on the dependency relations included in the input syntactic feature vector set. For example, the aforementioned question logic validation network can be the ERNIE 3.0 model. The aforementioned question content density prediction network can be a model that predicts the information density of the input linguistic concatenation feature vector set to effectively reduce the problem of describing simple content with complex language, and outputs a question information density score. The aforementioned question content density prediction network can be a model composed of a BERT model and a multilayer perceptron connected in series. The aforementioned two-layer attention mechanism network can be an attention mechanism layer with a first layer for analyzing the causal logic keywords of the initial question and a second layer for evaluating the effectiveness of the options in the initial question. The aforementioned multi-dimensional quality prediction model can be a model composed of a bidirectional long short-term memory neural network, a parallel fact consistency validation network, a question logic validation network, and a question content density prediction network, as well as a two-layer attention mechanism network connected in series.

[0092] The third step involves generating a multi-objective human-machine allocation threshold function. This function includes: minimizing the expert cost function and maximizing the evaluation fairness function. The multi-objective human-machine allocation threshold function can be a function used to allocate the evaluation workload between the multi-dimensional quality prediction model evaluation and the expert manual evaluation. The minimizing expert cost function can be a function that minimizes the expert evaluation workload while maintaining the accuracy of the initial test item quality evaluation. The maximizing evaluation fairness function can be a function that uses the Gini coefficient to measure the fairness between the predicted test item quality data set and the benchmark scores given by experts.

[0093] The fourth step involves performing Pareto multi-attribute decision-making on the aforementioned multi-objective human-machine allocation threshold function to obtain the first objective credibility threshold, the second objective credibility threshold, the first prediction credibility threshold, and the second prediction credibility threshold. The first objective credibility threshold characterizes the maximum value required to measure the credibility of the predicted test quality values ​​(i.e., the difference between the predicted test quality values ​​and the standard scores from expert criteria), thus determining whether expert intervention is needed. The second objective credibility threshold characterizes the minimum value required to measure the credibility of the predicted test quality values, thus determining whether expert intervention is needed. The first prediction credibility threshold characterizes the maximum value required to determine whether to accept the predicted test quality values ​​as the true test quality assessment. The second prediction credibility threshold is the minimum value required to determine whether to accept the predicted test quality values ​​as the true test quality assessment. In practice, the implementing entity can first use a third-generation non-dominated sorting genetic algorithm to solve the Pareto optimal threshold problem for the aforementioned multi-objective human-machine allocation threshold function, obtaining a Pareto optimal solution set. Then, TOPSIS (Technique for Order Preference by Similarity to Ideal Solution) is used to sort the Pareto optimal solution set to obtain the first target confidence threshold, the second target confidence threshold, the first prediction confidence threshold, and the second prediction confidence threshold.

[0094] The fifth step involves generating a human-machine assessment granularity decision allocation model based on the aforementioned first target credibility threshold, second target credibility threshold, first predicted credibility threshold, second predicted credibility threshold, and the predicted test item quality value set. This human-machine assessment granularity decision allocation model can be a sequential three-branch decision model with a two-level granularity structure, allocating the human-machine assessment workload among the input first target credibility threshold, second target credibility threshold, first predicted credibility threshold, second predicted credibility threshold, and predicted test item quality value set. As an example, the executing entity can first determine the difference function between the predicted test item quality value set and the preset expert scoring standard as the first-level assessment function value, and determine the predicted test item quality value set as the second-level assessment function value. The preset expert scoring standard can be assessment values ​​set by experts for evaluating different types of initial test items. Then, the initial test item set, the first-level assessment function, the first target credibility threshold, and the second target credibility threshold are used as the first-level human-machine allocation model. Subsequently, the initial question set within the reliable region output by the first-level human-machine allocation model, the second-level evaluation function value, the first prediction confidence threshold, and the second prediction confidence threshold are used as the second-level human-machine allocation model. Finally, the first-level human-machine allocation model and the first-level human-machine allocation model are determined as the human-machine evaluation granularity decision allocation model.

[0095] Step 6: Based on the aforementioned human-machine evaluation granularity decision allocation model, the above-mentioned test item quality prediction numerical set is adjusted collaboratively by humans and machines to obtain the test item quality evaluation numerical set. The test item quality evaluation numerical values ​​in this set can represent the initial test item quality evaluation values ​​obtained from the multi-dimensional quality prediction model and expert collaborative evaluation.

[0096] As an example, the aforementioned implementing entity can first input the predicted test item quality value set into the first-level evaluation function included in the aforementioned human-machine evaluation granularity decision allocation model to obtain the first evaluation function value set. Secondly, using the first-level human-machine allocation model, the initial test item set is divided: when the first evaluation function value set contains values ​​less than the aforementioned second target reliability threshold, the corresponding initial test item set is assigned to a high-reliability region, and acceptance is performed, along with outputting the first evaluation function value set as the target evaluation function value set; when the first evaluation function value set contains values ​​greater than the first target reliability threshold, the corresponding initial test item set is assigned to a low-reliability region, and expert re-evaluation is performed, along with outputting the first test item quality value set after expert re-evaluation; when the first evaluation function value set contains values ​​greater than or equal to the second target reliability threshold and less than or equal to the first target reliability threshold, the corresponding initial test item set is assigned to a medium-reliability region and input into the second-level human-machine allocation model. Subsequently, using the second-level human-machine allocation model, the initial question set included in the medium reliability region is divided into target initial question sets. Specifically, if any predicted question quality value set exceeds the first prediction confidence threshold, the corresponding target initial question set is assigned to the high-score region, and expert re-evaluation is performed, resulting in the output of a second set of re-evaluated question quality values. If any predicted question quality value set falls below the second prediction confidence threshold, the corresponding target initial question set is assigned to the low-score region, and expert re-evaluation is performed, resulting in the output of a third set of re-evaluated question quality values. If any predicted question quality value set is greater than or equal to the second prediction confidence threshold but less than or equal to the first prediction confidence threshold, the corresponding target initial question set is assigned to the medium-score region, and acceptance is performed, resulting in the output of the target question quality predicted value set. Finally, the target evaluation function value set, the first question quality value set, the second question quality value set, the third question quality value set, and the target question quality predicted value set are determined as the question quality evaluation value set.

[0097] Step 7: Based on the aforementioned set of test question quality assessment values, dynamically adjust the initial test question set to obtain an adjusted test question set. Then, based on the aforementioned device information group and the set of test question quality assessment values, perform adaptive display on the initial test question set. In practice, the execution entity may first select at least one initial test question from the initial test question set whose corresponding test question quality assessment value is less than a preset test question assessment threshold. Then, repeatedly regenerate the at least one initial test question until the test question quality assessment value corresponding to the generated initial test question is greater than or equal to the preset test question assessment threshold. Finally, using video memory access technology, based on the device information group, perform adaptive display on the initial test question set that is greater than or equal to the preset test question assessment threshold.

[0098] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "low accuracy in test question quality assessment, low test question quality, and the need for multiple test question displays resulting in wasted video memory resources and low video memory quality." The factors leading to low accuracy in test question quality assessment, low test question quality, and the need for multiple test question displays resulting in wasted video memory resources and low video memory quality are often as follows: Educational institutions' exam question databases require a large number of diverse and high-quality test questions. Using only model evaluation to assess the quality of initial test questions results in low accuracy in test question quality assessment due to problems such as illusions, factual errors, and logical jumps in the model. This necessitates multiple test question displays, leading to wasted video memory resources and low video memory quality. Solving these factors can improve the accuracy of test question quality assessment, improve test question quality, reduce wasted video memory resources, and improve video memory quality. To achieve this effect, this disclosure firstly extracts semantic feature vector sets to represent the deep meaning of the vocabulary included in the test questions, grammatical feature vector sets to reflect the structural information of part-of-speech distribution, and syntactic feature vector sets to represent the dependency relationships of test question sentences. This comprehensively covers the feature information of the initial test questions, providing foundational data for subsequent model-based quality assessment. Secondly, automatic multi-dimensional quality assessment through a multi-dimensional quality prediction model can improve the accuracy and comprehensiveness of the model assessment. Next, a multi-objective human-machine allocation threshold function is generated and subjected to Pareto multi-attribute decision-making. The multi-objective human-machine allocation threshold function, employing a multi-objective optimization form, can balance cost minimization and quality assessment accuracy. Pareto multi-attribute decision-making provides the accuracy of the solution, allowing for a reasonable allocation of the human-machine assessment workload in subsequent determinations. Subsequently, a human-machine collaborative adjustment is achieved using a granular decision-making allocation model for human-machine evaluation. This model employs a two-layer granularity division for human-machine task allocation, clearly defining allocation boundaries to avoid ambiguity and improve allocation efficiency. Human-machine collaborative adjustment effectively avoids the subjectivity of single-person evaluation and the illusion problems of single-model evaluation. Combining these two evaluation methods while minimizing manual costs improves the accuracy and reliability of test item quality assessment. Finally, the initial test item set is dynamically adjusted based on the test item quality assessment numerical set, resulting in an adjusted test item set which is then stored. Furthermore, the initial test item set is adaptively displayed based on device information groups and the test item quality assessment numerical set. This not only improves the quality of the initial test items but also considers the display requirements of the test item output devices, improving the quality of test item display and reducing the waste of video memory resources.

[0099] Step 106: Based on the test quality assessment numerical set, dynamically adjust the initial test set to obtain the adjusted test set, and store the adjusted test set in the test database.

[0100] In some embodiments, the executing entity can dynamically adjust the initial question set based on the question quality assessment value set to obtain an adjusted question set, and store the adjusted question set in a question database. The question database can be a database for storing generated question sets that are greater than or equal to a preset question assessment threshold. The preset question assessment threshold can be a pre-set critical value for evaluating the quality of the initial questions. The adjusted question set can include a collection of questions that meet the condition of being greater than or equal to the preset question assessment threshold before and after dynamic adjustment. The dynamic adjustment can be an adjustment made to different types of question stem information, question answer information, question option sets, and answer explanation information. For example, for multiple-choice questions, the dynamic adjustment can be an adjustment to add more confusing options to the distractors in the question option set, an adjustment to increase the range of knowledge points covered by the question option set, or an adjustment to optimize the accuracy of the question stem description. As an example, the executing entity can first select at least one initial question from the initial question set that is less than or equal to the preset question assessment threshold to obtain a target initial question set. Secondly, the initial test question set is adjusted to obtain the adjusted test question set, and the adjusted test question set is stored in the test question database.

[0101] Step 107: Based on the device information group, perform test paper adaptation display on the adjusted test question set to obtain the test questions and test papers. In response to the detected operation behavior information of the target user, dynamically adjust the display of the test questions and test papers and the adjusted test question set.

[0102] In some embodiments, the execution entity can perform test paper adaptation display on the adjusted question set according to the device information group to obtain the test paper, and dynamically adjust the display of the test paper and the adjusted question set in response to the detected operation behavior information of the target user. The test paper can be a test paper composed of initial questions of different types. The target user can be a user corresponding to the test paper output device who manually checks the displayed test paper. The operation behavior information can be operation information of the target user modifying the test paper's questions or layout. For example, if the test paper output device is a high-precision tablet or terminal display screen, the test paper adaptation display can use hardware acceleration to render the formulas and icons in the test paper, and utilize a layout engine to ensure clear and high-quality display of each parameter in the formulas. When the test paper output device is a mobile phone or a narrow-screen terminal display device, the test paper adaptation display can simplify the formula structure in the test paper through distributed loading, increase line spacing to avoid font overlap, or use a single-column layout. In practice, the aforementioned executing entity can first use a genetic algorithm to assemble the adjusted question set into test papers. Then, using double buffering technology, the test papers are displayed adaptively based on the device information set. Finally, in response to the detection of the target user's operation behavior information, the test papers are dynamically modified and displayed based on the user's operation behavior information, and the adjusted question set is regenerated into test papers using a question generation model set before being displayed.

[0103] The aforementioned test paper assembly request information may be a request to assemble the adjusted test question set into a single test paper. This request information may include: JSON-structured information such as the various question types included in the test paper, the number of questions for each question type, the distribution information of the various question types, the difficulty level of each question, the required examination time for each question, the total number of questions, and the distribution of question scores.

[0104] In addressing the aforementioned technical problems in the application scenario of large-scale standardized examinations (e.g., the National College Entrance Examination), the following technical issues often arise: In large-scale standardized examinations, test paper generation needs to meet multiple constraints and adapt to the majority of examinees. However, existing particle swarm optimization (PSO) algorithms suffer from uneven initial particle distribution and susceptibility to local optima. Furthermore, they fail to recognize the knowledge connections between questions, resulting in low-quality test papers. This necessitates repeated test paper generation, extending the time required for test paper assembly and rendering, reducing display quality, and wasting video memory resources. Based on the characteristics of this application scenario—high-quality test papers, diverse question types, high level of examinee understanding, broad knowledge coverage, low question repetition, and simultaneous satisfaction of multiple constraints—we have decided to adopt the following solution:

[0105] In some optional implementations of certain embodiments, the process of adapting the adjusted test question set to a test paper based on the device information group to obtain the test paper, and dynamically adjusting the display of the test paper and the adjusted test question set in response to the detection of the target user's operation behavior information, may include the following steps:

[0106] The first step involves generating an initial test paper assembly particle swarm and a test paper difficulty and question type distribution function based on the aforementioned test question generation request information. The initial test paper assembly particle swarm can be defined as follows: the initial test paper assembly particles are individuals that use real-number encoding of the selected initial test questions' IDs and question types in the test question database as encoding genes, and the number of test questions included in the aforementioned test question generation request information as the encoding length. Each initial test paper assembly particle can include an initialization velocity and an initialization position. Each initial test paper assembly particle can represent a feasible test paper assembly scheme based on the test paper difficulty and question type distribution function. The initialization velocity can refer to the direction of movement of the initial test paper assembly particle from its current position to the next position, i.e., the search direction of the initial test paper assembly particle in the solution space. The initial position can represent the numerical values ​​corresponding to each parameter in the test paper difficulty and question type distribution function, i.e., a feasible test paper assembly information. The test paper difficulty and question type distribution function can represent the variance of the difficulty level of the selected initial test questions and the preset difficulty level, the proportion of knowledge points, and the question type distribution. As an example, the aforementioned execution entity can utilize the Circle chaotic mapping algorithm to generate an initialization particle swarm and a question type distribution function for the test paper based on the aforementioned test question generation request information.

[0107] The second step involves performing the following volume assembly steps based on the initialized volume assembly particle swarm:

[0108] Sub-step 1 involves performing crossover and mutation processing on the initial test paper particle swarm to obtain the mutated test paper particle swarm. In practice, the aforementioned execution entity can first use a tournament selection algorithm to select the target test paper particle swarm. Secondly, two random numbers are randomly determined, where the value range is [0, number of questions]. Next, a gene crossover operation is performed on the target test paper particle swarm to obtain a crossover test paper particle swarm. The crossover particle swarm includes individuals whose gene encoding is inherited from the mother particle in the target test paper particle swarm, specifically from the gene encoding at the two random numbers, and from the father particle, specifically from the gene encoding at the other two random numbers. Then, a question duplication detection and replacement process is performed on the crossover particle swarm to obtain a replaced test paper particle swarm. This question duplication detection and replacement process can involve first checking if there are duplicate questions in the three most recently generated test papers; if a duplication is detected, a question with the same knowledge point and score as the duplicate question is randomly selected from the question database. Then, the Gaussian mutation algorithm is used to mutate the particle swarm of the replacement test papers, resulting in the mutated particle swarm. This mutation process can involve replacing the initial test questions with questions of the same type, corresponding knowledge points that are one or two hops away from the replaced knowledge points, and keeping the question scores and difficulty levels unchanged.

[0109] Sub-step 2: Based on the aforementioned question difficulty distribution function, determine the individual target position set and the population target position of the mutated test paper particle swarm. The individual target positions in the individual target position set can be the positions corresponding to the highest fitness obtained by inputting the position information of the mutated test paper particle from the initial loop iteration to the current iteration number into the question difficulty distribution function. The population target position can be the position corresponding to the highest fitness value in the mutated test paper particle swarm in each loop iteration. As an example, the execution entity can first determine the individual target position set obtained by each mutated test paper particle from the start of the loop iteration to the current loop iteration number. Secondly, for each individual target position set in the above individual target position set, determine the maximum value of the question difficulty distribution function corresponding to the individual target position set, as the initial target position. Then, input the mutated test paper particle swarm into the question difficulty distribution function to obtain the current particle fitness value set. Finally, determine the individual target position corresponding to the largest current particle fitness value in the above current particle fitness value set, as the cluster target position.

[0110] Sub-step 3: Based on the individual target position set and the population target position, perform quantum velocity position update processing on the mutated group particle swarm to obtain the updated group particle swarm.

[0111] As an example, the aforementioned execution entity can utilize a particle swarm optimization algorithm based on Gaussian quantum improvement to perform quantum velocity position update processing on the mutated grouped particle swarm according to the individual target position set and the population target position, thereby obtaining the updated grouped particle swarm.

[0112] Sub-step 4: Based on the aforementioned question type difficulty distribution function, the updated question set particle swarm is filtered to obtain a filtered question set particle swarm. The filtered question set particle swarm can be a group of particles from the updated question set particle swarm and the initial question set particle swarm with the highest corresponding question type difficulty distribution function values, representing a pre-defined population size. This pre-defined population size can be the number of updated question set particles included in the updated question set particle swarm.

[0113] Sub-step 5: In response to the determination that the number of times the test paper assembly step has been executed is greater than or equal to a preset execution threshold, the test paper is generated and rendered based on the filtered test paper assembly particle group. The preset execution threshold can be a pre-defined maximum value for executing the test paper assembly step. In practice, the test paper corresponding to the individual particle with the highest distribution function value of the corresponding test paper difficulty question type is selected from the filtered test paper assembly particle group and determined as the test paper.

[0114] The third step involves determining that the number of executions is less than a preset execution threshold, identifying the filtered group of volume-building particles as the initial volume-building particle group, and determining the sum of the number of executions and a preset value as the number of executions, so as to execute the volume-building steps again. The preset value can be a pre-defined value, typically 1.

[0115] The fourth step involves adapting the test questions and papers to the aforementioned device information group, and dynamically adjusting the display of the test questions and papers and the adjusted test question set in response to the detected user's operational behavior. The specific implementation of this step can be found in step 107.

[0116] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "low-quality generated test papers, requiring repeated generation, extending the time for test paper assembly and rendering, reducing display quality, and wasting video memory resources." The factors leading to low-quality generated test papers, requiring repeated generation, extending the time for test paper assembly and rendering, and wasting video memory resources are often as follows: In large-scale standardized examinations, test paper generation needs to meet multiple constraints and adapt to most candidates as much as possible. However, existing particle swarm optimization algorithms for test paper assembly suffer from uneven initial particle distribution and are prone to getting trapped in local optima. Furthermore, they cannot perceive the knowledge connections between test questions, resulting in low-quality generated test papers, requiring repeated generation, extending the time for test paper assembly and rendering, and wasting video memory resources. Solving these factors can improve the quality of generated test papers, reduce the number of test paper generation attempts, shorten the time for test paper assembly and rendering, improve display quality, and reduce the waste of video memory resources. To achieve this effect, this disclosure first utilizes the Circle chaotic mapping algorithm to generate an initial test paper particle swarm that improves the uniformity of particle swarm distribution, effectively preventing the test paper from getting stuck in localized regions at the outset. The test paper difficulty and question type distribution function considers factors such as difficulty, knowledge point coverage, and question type distribution, enabling a comprehensive and accurate assessment of the test paper's quality, thus facilitating subsequent improvements. Second, the initial test paper particle swarm undergoes crossover and mutation processing, fully leveraging the correlations between knowledge points corresponding to the questions. This makes mutation a fine-tuning or proximity exploration rather than a destructive random jump, improving the optimization efficiency of the test paper. Furthermore, the repeated replacements in the crossover and mutation process ensure that duplicates are repaired without disrupting the overall knowledge point, question type, and total score distribution of the test paper. Finally, a particle swarm algorithm based on Gaussian quantum optimization is used for quantum velocity position updates, which better adapts to the search in discrete space. After the update, individuals in the test paper particle swarm can appear at any position in the search space with a certain probability, exhibiting a strong ability to escape local optima and achieving rapid convergence, thus shortening the test paper generation time. Finally, based on the filtered test paper particle group, test questions and test papers are generated, and the display is adapted according to the device information group. In response to the detection of the target user's operation behavior information, the test questions and test papers and the adjusted test question set are dynamically adjusted for display. This can improve the display quality of test questions and test papers, reduce problems such as incomplete display and low display resolution, reduce the waste of video memory resources, and shorten the display time.

[0117] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a test item output device based on multi-model collaboration. These device embodiments are similar to... Figure 1Corresponding to the method embodiments shown, this multi-model collaborative test output device can be specifically applied to various electronic devices.

[0118] like Figure 2 As shown, a test question output device 200 based on multi-model collaboration includes: a multi-source data fusion unit 201, a knowledge point association analysis unit 202, a multi-hop graph reasoning query unit 203, a generation unit 204, a multi-dimensional test question quality assessment unit 205, a dynamic adjustment unit 206, and a test paper adaptation display unit 207. The multi-source data fusion unit 201 is configured to: in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data. The various test question output devices include: a display and a printer. The knowledge point association analysis unit 202 is configured to: perform knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps. The multi-hop graph reasoning query unit 203 is configured to: in response to detecting test question generation request information, perform multi-hop graph reasoning query on the dynamic test question knowledge point graph to obtain a set of test question association knowledge points. The generation unit 204 is configured to: generate initial question sets for different question types by collaboratively generating the question generation request information and the set of related knowledge points using a question generation model set. The multi-dimensional question quality assessment unit 205 is configured to: perform multi-dimensional question quality assessment on the initial question set to obtain a question quality assessment value set. The dynamic adjustment unit 206 is configured to: dynamically adjust the initial question set based on the question quality assessment value set to obtain an adjusted question set, and store the adjusted question set in the question database. The test paper adaptation display unit 207 is configured to: perform test paper adaptation display on the adjusted question set based on the device information set to obtain a test paper, and dynamically adjust the display of the test paper and the adjusted question set in response to detected target user operation behavior information.

[0119] It is understandable that the units and references recorded in the multi-model collaborative test output device 200 are related to... Figure 1 The steps described in the method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the multi-model collaborative test output device 200 and the units contained therein, and will not be repeated here.

[0120] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0121] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0122] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0123] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0124] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0125] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0126] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data, wherein the aforementioned test question output devices include: a display, a printer; perform knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point graph with timestamps; in response to detecting test question generation request information, perform multi-hop graph reasoning query on the aforementioned dynamic test question knowledge point graph to obtain a set of test question association knowledge points; and through the test... The question generation model set is used to collaboratively generate initial question sets for different question types based on the aforementioned question generation request information and the aforementioned set of knowledge points associated with the questions. Multi-dimensional question quality assessment is then performed on the initial question sets to obtain a set of question quality assessment values. Based on the aforementioned question quality assessment values, the initial question sets are dynamically adjusted to obtain an adjusted question set, which is then stored in the question database. Based on the aforementioned device information group, the adjusted question set is displayed for test paper adaptation to obtain the test paper. Finally, in response to the detection of target user operation behavior information, the test paper and the adjusted question set are dynamically displayed and adjusted.

[0127] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0129] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including: a multi-source data fusion unit, a knowledge point association analysis unit, a multi-hop graph reasoning query unit, a generation unit, a multi-dimensional test question quality assessment unit, a dynamic adjustment unit, and a test paper adaptation display unit. The names of these units do not necessarily limit the unit itself; for example, the multi-source data fusion unit may also be described as "a unit that, in response to receiving device information groups from various test question output devices, performs multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data."

[0130] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0131] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A test item output method based on multi-model collaboration, comprising: In response to receiving the device information group of each test question output device, the acquired multi-source test question association dataset is subjected to multi-source data fusion processing to obtain fused multi-source test question association data, wherein each test question output device includes: a display and a printer; The fused multi-source test question association data is subjected to knowledge point association analysis to obtain a dynamic test question knowledge point map with timestamps; In response to the detection of a question generation request, a multi-hop graph reasoning query is performed on the dynamic question knowledge point graph to obtain a set of question-related knowledge points. This includes: performing a thought chain decomposition process on the question test point selection information set corresponding to the question generation request to obtain a sequence of question test point selection information, wherein the sequence of question test point selection information includes information with logical dependencies on the question test point selection information; for the sequence of question test point selection information, the following generation steps are performed: based on the previously initial query knowledge point set and the dynamic question knowledge point graph, a sequence of questions is generated for the question test point selection information set. The system selects the current set of knowledge points for the selected information; performs contextual dynamic encoding on the selected information to obtain a set of query context embedding vectors; determines the set of confidence scores associated with the reasoning steps of the set of query context embedding vectors; generates a knowledge point entity transition matrix based on the set of confidence scores associated with the reasoning steps and the current set of knowledge points; generates a set of entity multi-hop reasoning paths based on the knowledge point entity transition matrix and the previous initial set of query knowledge points; and generates a set of knowledge points associated with the selected information based on the set of entity multi-hop reasoning paths, in response to determining that the selected information is selected at the termination position. By using a question generation model set, the question generation request information and the set of knowledge points associated with the question are collaboratively generated to generate initial question sets for different question types. A multi-dimensional test quality assessment is performed on the initial test item set to obtain a test quality assessment numerical set. Based on the set of test item quality assessment values, the initial test item set is dynamically adjusted to obtain an adjusted test item set, and the adjusted test item set is stored in the test item database. Based on the device information group, the adjusted test question set is displayed to adapt to the test paper, and in response to the detection of the target user's operation behavior information, the test question set and the adjusted test question set are dynamically displayed and adjusted.

2. The method according to claim 1, wherein, The test question generation system model set includes: a question stem generation model, an answer generation model, and a distractor option generation model; and The step involves using a question generation model set to collaboratively generate initial question sets for different question types by combining the question generation request information and the set of knowledge points associated with the questions. The test question generation request information is processed by structured parsing to obtain structured test question requirement information; Based on the structured information of the test question requirements, semantic enhancement detection is performed on the set of knowledge points associated with the test questions to obtain a set of knowledge point text information. In response to a question that is determined to be a fill-in-the-blank question, the knowledge point text information set is subjected to entity recognition to obtain a set of entities to be filled. The question stem generation model is invoked to generate different types of question stem information sets based on the entity set to be filled, the structured information of the test question requirements, and the knowledge point text information set. The answer generation model is invoked to generate different types of question answer information sets based on the question stem information set and the knowledge point text information set; In response to a question that is determined to be a multiple-choice question, the distractor option generation model is invoked to generate a multi-strategy set of question answer information for the multiple-choice question, thereby obtaining a set of distractor option text information. The question stem information set, the question answer information set, and the distractor option text information set are combined and processed to obtain an initial question set for different question types.

3. The method according to claim 1, wherein, The step involves using a question generation model set to collaboratively generate initial question sets for different question types by combining the question generation request information and the set of knowledge points associated with the questions. Based on the question generation request information and the set of knowledge points associated with the question, generate a set of prompt information for multiple types of subjective questions; The set of multiple types of prompts for the subjective questions is subjected to enhanced serialization processing to obtain the target set of multiple types of prompts. Based on the target multi-type prompt information set, determine the call path set of the test question generation model set; The question generation model set corresponding to the call path set is invoked to generate a set of subjective question stem text information, a set of subjective question answer text information, and a set of answer score text information. A consistency check is performed on the set of subjective question stem text information and the set of subjective question answer text information to obtain a set of question stem and answer detection results. Based on the question stem and answer detection result set, the subjective question stem text information set and the subjective question answer text information set are adjusted to remove homogenization, resulting in the adjusted subjective question stem text information set and the adjusted subjective question answer text information set; The adjusted subjective question stem text information set, the adjusted subjective question answer text information set, and the answer score text information set are determined as the initial question sets for different question types.

4. The method according to claim 1, wherein, The process of performing a multi-dimensional quality assessment on the initial question set to obtain a numerical set of question quality assessments includes: Using a test question difficulty prediction model, the initial test question set is evaluated to obtain a set of test question difficulty evaluation values; Using a question similarity prediction model, the initial question set is evaluated for knowledge point similarity to obtain a question similarity evaluation numerical set. Multimodal feature extraction is performed on the initial question set to obtain a question feature vector set; Based on the set of feature vectors of the test questions, the initial set of test questions is evaluated for content quality to obtain a set of test question content quality evaluation values; Based on the set of feature vectors of the test questions, the initial set of test questions is evaluated for knowledge point matching to obtain a set of knowledge point matching evaluation values. Based on the test item feature vector set, the initial test item set is subjected to bias detection processing to obtain a test item bias evaluation numerical set; The test question quality assessment set is obtained by multi-feature fusion of the test question difficulty assessment set, the test question similarity assessment set, the test question content quality assessment set, the knowledge point matching assessment set, and the test question bias assessment set.

5. The method according to claim 1, wherein, The method further includes: In response to the determination that the test point selection information is not the selection information located at the termination position, the test knowledge point set corresponding to the entity multi-hop reasoning path set is determined as the previous initial query knowledge point set, so as to execute the generation step again.

6. The method according to claim 1, wherein, The step of performing knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point map with timestamps includes: The fused multi-source test item association data is segmented into semantic blocks to obtain a set of test item association semantic blocks; Based on the set of semantic blocks associated with the test questions, generate knowledge extraction prompts and data source relationship extraction prompts; The knowledge extraction prompts, the data source relationship extraction prompts, and the test question association semantic block set are input into the graph generation large language model to obtain the initial knowledge point graph set for different data sources; The initial knowledge point graph set is subjected to knowledge point fusion processing to obtain an initial knowledge point fusion graph; Entity relationship conflict identification is performed on the initial knowledge point fusion graph to obtain a set of knowledge point conflict triples; Entropy conflict disambiguation is performed on the set of knowledge point conflict triples to obtain the disambiguated set of knowledge point triples. The disambiguated knowledge point triplet set is input into the initial knowledge point fusion graph to obtain the knowledge point fusion graph, which serves as a dynamic test question knowledge point graph with timestamps. The dynamic test question knowledge point graph is then dynamically optimized through reinforcement learning.

7. A test item output device based on multi-model collaboration, comprising: The multi-source data fusion unit is configured to, in response to receiving device information groups from various test question output devices, perform multi-source data fusion processing on the acquired multi-source test question association dataset to obtain fused multi-source test question association data, wherein each test question output device includes: a display and a printer; The knowledge point association analysis unit is configured to perform knowledge point association analysis on the fused multi-source test question association data to obtain a dynamic test question knowledge point map with timestamps. The multi-hop graph reasoning query unit is configured to, in response to detecting a question generation request, perform a multi-hop graph reasoning query on the dynamic question knowledge point graph to obtain a set of question-related knowledge points. This includes: performing a thought chain decomposition process on the question test point selection information set corresponding to the question generation request to obtain a sequence of question test point selection information, wherein the sequence of question test point selection information includes information with logical dependencies on the question test point selection information; for the sequence of question test point selection information, the following generation steps are performed: based on the previously initial query knowledge point set and the dynamic question knowledge point graph, Generate a current set of knowledge points for the selected test points; perform context-based dynamic encoding on the selected test points to obtain a set of query context embedding vectors; determine the set of inference steps associated with the set of query context embedding vectors; generate a knowledge point entity transition matrix based on the set of inference steps associated with the set of current query knowledge points; generate a set of entity multi-hop inference paths based on the knowledge point entity transition matrix and the previous initial set of query knowledge points; in response to determining that the selected test points are selected at the termination position, generate a set of test point associated knowledge points based on the set of entity multi-hop inference paths. The generation unit is configured to use a test question generation model set to collaboratively generate the test question generation request information and the test question-related knowledge point set to generate an initial test question set for different test question types. The multi-dimensional test item quality assessment unit is configured to perform a multi-dimensional test item quality assessment on the initial test item set to obtain a test item quality assessment numerical set. The dynamic adjustment unit is configured to dynamically adjust the initial test question set according to the test question quality assessment value set to obtain an adjusted test question set, and to store the adjusted test question set in the test question database. The test paper adaptation display unit is configured to perform test paper adaptation display on the adjusted test paper set according to the device information group to obtain test paper, and to dynamically adjust the display of the test paper and the adjusted test paper set in response to the detection of the target user's operation behavior information.

8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Large model knowledge retrieval method based on knowledge graph enhancement

    CN121256062A

  • Middle and primary school test question intelligent generation method based on education big model

    CN121257488A