Open type question automatic scoring method and system based on multi-modal large model and target detection

By combining a multimodal large language model and target detection technology, key text information is extracted from test paper images and modal fusion is performed, which solves the accuracy problem of existing automatic scoring technology when processing complex content and achieves more efficient and stable scoring results.

CN120635927AActive Publication Date: 2025-09-12BEIJING NORMAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510731207.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing automatic scoring technology is prone to problems such as character misrecognition and format loss when processing complex content such as handwritten text, formulas, and charts, resulting in reduced scoring accuracy.

Method used

A method based on a multimodal large language model and object detection is adopted. The DETR object detector is used to extract key text information from the test paper images, generate image token sequences of variable length, and perform modal fusion through a cross-modal attention mechanism. Finally, a prompt model optimized by reinforcement learning is used to generate prompt words for scoring.

Benefits of technology

It significantly improves the accuracy and stability of scoring, reduces information loss, enhances the fusion effect of multimodal information, and improves the adaptability and scoring efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635927A_ABST
    Figure CN120635927A_ABST
Patent Text Reader

Abstract

The invention discloses an open question automatic scoring method and system based on a multi-modal large model and target detection, and the method comprises the steps: obtaining a test paper picture containing an answer to an open question, and carrying out the end-to-end scoring through a multi-modal large language model; processing the test paper picture by using a DETR-based target detector, extracting key text information and generating a corresponding image token sequence; inputting the image token sequence and the key text information into a multi-modal large language model, and performing modal fusion through a cross-modal attention mechanism to obtain a fused representation; and generating cue words by using a prompt model of reinforcement learning optimization, combining the cue words with the fused representation, and inputting the combined cue words and the fused representation into a scoring model to obtain a scoring result. According to the invention, the accuracy and efficiency of open type question automatic scoring can be obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and educational scoring, and in particular relates to a method and system for automatic scoring of open-ended questions based on a multimodal large model and target detection. Background Art

[0002] With the rapid development of educational informatization, the demand for automatic scoring technology in scenarios such as large-scale examinations and online education is growing. Traditional automatic scoring methods mainly rely on optical character recognition (OCR) technology to convert handwritten or printed text into a processable text format, and then use natural language processing (NLP) technology to score. However, existing technologies have the following shortcomings:

[0003] OCR introduces noise: When processing complex content such as handwritten text, formulas, and charts, OCR technology is prone to problems such as character misrecognition, typographical errors, and format loss, resulting in reduced scoring accuracy.

[0004] Insufficient multimodal information fusion: Existing large multimodal models typically use image encoders based on Visual Transformer (ViT). Fixed-length image tokens make it difficult to fully express rich text images containing complex information such as handwritten text, formulas, and tables, resulting in information loss.

[0005] Limitations of prompt optimization and model optimization: Traditional prompt optimization methods mainly rely on continuous prompt learning, which is prone to over-reliance on model parameters. This leads to poor generalization ability of prompt words on different tasks or model versions, making it difficult to adapt to diverse scoring tasks. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method and system for automatically scoring open-ended questions based on a multimodal large model and target detection. The method includes:

[0007] Obtain test paper images containing answers to open-ended questions and use a multimodal large language model for end-to-end scoring;

[0008] The test paper image is processed using a DETR-based object detector to extract key text information and generate a corresponding image token sequence;

[0009] Inputting the image token sequence and key text information into a multimodal large language model, performing modal fusion through a cross-modal attention mechanism, and obtaining a fused representation;

[0010] A prompt word is generated using a prompt model optimized by reinforcement learning, and the prompt word is combined with the fused representation and input into a scoring model to obtain a scoring result.

[0011] Preferably, the process of obtaining the test paper image containing the answers to the open-ended questions includes:

[0012] Collect students' real handwritten answer pictures and manually annotate them;

[0013] The LaTeX typesetting tool is used to generate synthetic data in rich text format, the existing text scoring data is converted into handwriting style images, and different handwriting styles are simulated through handwriting fonts.

[0014] Preferably, the multimodal large language model is trained using a step-by-step transfer learning strategy, and the training process includes:

[0015] A DETR-based object detector is used as the image encoder and continuously pre-trained on text recognition and TextVQA tasks;

[0016] Use existing text scoring data to fine-tune the backbone of the multimodal large language model;

[0017] The entire model is fine-tuned end-to-end using the constructed multimodal ratings dataset.

[0018] Preferably, the process of processing the test paper image using a DETR-based object detector includes:

[0019] Use DETR to detect key areas in the test image and generate anchor boxes containing text information and corresponding locations;

[0020] Filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor boxes to form image tokens;

[0021] MLP is used to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

[0022] Preferably, the process of filtering out low-confidence anchor boxes includes:

[0023] Low-confidence target areas are filtered out by a preset confidence threshold, and only anchor boxes with confidence higher than the confidence threshold are retained. The prior information of the scoring task is integrated into the DETR query.

[0024] Preferably, the process of performing modal fusion through a cross-modal attention mechanism to obtain a fused representation includes:

[0025] Treat the contextual information as a prompt word and use the BERT model to encode the prompt word into a DETR query;

[0026] The visual-text joint representation technology is introduced. Based on the idea of ​​contrastive learning, the model is trained using image-text matching tasks and text recognition tasks. The text and the corresponding image areas are aligned in the latent semantic space to obtain a fused representation.

[0027] Preferably, the process of generating prompt words using the prompt model optimized by reinforcement learning includes:

[0028] Design two generative models, a prompt optimization model and a scoring model, that do not share parameters, and train them separately; the prompt optimization model is used to generate discrete prompt words, and the scoring model is used to score input data based on the prompt words; the prompt optimization model and the scoring model are connected directly or after being transformed through several layers of neural networks;

[0029] When training the prompt optimization model, the parameters of the scoring model are frozen, the prompt words generated by the prompt optimization model are concatenated with the input and then input into the scoring model for evaluation, and the error between the model output and the label is used to generate reward information which is then passed back to the prompt optimization model;

[0030] When training the scoring model, the parameters of the prompt optimization model are frozen, the prompt words generated by the prompt optimization model are concatenated with the input and input into the scoring model for evaluation, and the error between the model output and the label is back-propagated into the scoring model to optimize the parameters of the scoring model.

[0031] Preferably, the prompt optimization model is trained using a Proximal Policy Optimization algorithm.

[0032] The present invention also provides an automatic scoring system for open-ended questions based on a multimodal large model and target detection, comprising:

[0033] The image acquisition module is used to obtain test paper images containing answers to open-ended questions and perform end-to-end scoring using a multimodal large language model;

[0034] An object detection module is used to process the test paper image using a DETR-based object detector, extract key text information and generate a corresponding image token sequence;

[0035] A modality fusion module is used to input the image token sequence and key text information into a multimodal large language model, perform modality fusion through a cross-modal attention mechanism, and obtain a fused representation;

[0036] The prompt and scoring module is used to generate prompt words using the prompt model optimized by reinforcement learning, and combine the prompt words with the fused representation, input them into the scoring model, and obtain the scoring results.

[0037] Preferably, the target detection module includes:

[0038] Anchor frame generation unit, used to use DETR to detect key areas in the test volume image and generate anchor frames containing text information and corresponding positions;

[0039] The anchor box screening unit is used to filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor box to form an image token; and use MLP to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

[0040] Compared with the prior art, the present invention has the following advantages and technical effects:

[0041] Reduce information loss: This invention avoids the problems of character misrecognition and format loss caused by OCR conversion by directly processing the original test paper image, and significantly improves the accuracy and stability of scoring.

[0042] Enhanced information expression capability: This paper adopts a DETR-based object detector, which can detect key areas in rich text and generate image tokens of variable length, better express the fine-grained content in the image, and enhance the fusion effect of multimodal information.

[0043] Improved generalization ability: The prompt words generated by the reinforcement learning-based prompt model joint optimization method of the present invention have better generalization ability, can maintain stable scoring ability in different models and tasks, and improve the adaptability of the system.

[0044] Improved scoring efficiency: The end-to-end multimodal scoring method of the present invention reduces preprocessing steps and combines efficient parameter fine-tuning technology to reduce computing costs and enhance the system's deployment potential and practical application value.

[0045] The present invention can significantly improve the accuracy and efficiency of automatic scoring of open-ended questions and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0047] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0049] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0050] Example 1

[0051] This embodiment intends to construct an automatic scoring method based on a multimodal large language model, combining technologies such as target detection, prompt engineering, and reinforcement learning to improve scoring results. In scenarios such as large-scale examinations and online education, manual scoring is time-consuming and easily affected by subjective factors. The automatic scoring system can improve scoring efficiency and fairness, enable students to obtain feedback more quickly, and optimize the learning experience. To overcome the problem of information loss in OCR conversion, this embodiment adopts a multimodal large model so that it can score directly based on the test paper image, reducing information loss. At the same time, in view of the limitation that the existing multimodal model is difficult to capture complex text image information, this embodiment introduces a DETR-based target detector to extract key information from rich text test papers and enhance modal fusion. In addition, in order to improve the generalization ability of prompt optimization, this embodiment combines reinforcement learning to dynamically adjust the prompt words, making the scoring model more stable and more adaptable. This embodiment not only promotes the application of large language models in educational assessment, but also provides a methodological reference for other text processing tasks.

[0052] like Figure 1 As shown, this embodiment provides an automatic scoring method for open-ended questions based on a multimodal large model and target detection, including:

[0053] Obtain test paper images containing answers to open-ended questions and use a multimodal large language model for end-to-end scoring;

[0054] Use the DETR-based object detector to process the test paper image, extract key text information and generate the corresponding image token sequence;

[0055] The image token sequence and key text information are input into a multimodal large language model, and modal fusion is performed through a cross-modal attention mechanism to obtain a fused representation;

[0056] The prompt model optimized by reinforcement learning is used to generate prompt words, which are then combined with the fused representation and input into the scoring model to obtain the scoring results.

[0057] Furthermore, the process of obtaining the test paper image containing the answers to the open-ended questions includes:

[0058] Collect students' real handwritten answer pictures and manually annotate them;

[0059] The LaTeX typesetting tool is used to generate synthetic data in rich text format, the existing text scoring data is converted into handwriting style images, and different handwriting styles are simulated through handwriting fonts.

[0060] Furthermore, the multimodal large language model is trained using a step-by-step transfer learning strategy. The training process includes:

[0061] A DETR-based object detector is used as the image encoder and continuously pre-trained on text recognition and TextVQA tasks;

[0062] Use existing text scoring data to fine-tune the backbone of the multimodal large language model;

[0063] The entire model is fine-tuned end-to-end using the constructed multimodal ratings dataset.

[0064] Furthermore, traditional automatic scoring systems are mainly based on natural language processing (NLP) technology, which requires optical character recognition (OCR) tools to convert handwritten answers into text information. However, existing OCR tools may not be able to process information such as formulas and charts, and may produce recognition errors, resulting in information loss, which in turn affects the accuracy of scoring. This embodiment designs an end-to-end short-answer question automatic scoring method based on a multimodal large language model. The multimodal large language model can recognize the original test paper image and perform end-to-end scoring directly based on the original test paper image without relying on OCR conversion, thereby reducing information loss and improving scoring results.

[0065] Specifically, existing automatic scoring methods are mainly based on text processing and rely on OCR technology to convert handwritten answers into readable text. However, the OCR process easily introduces noise such as character misrecognition and formatting errors, which leads to the accumulation of scoring errors. To solve this problem, this embodiment uses a multimodal large language model to directly input the original handwritten answer image for end-to-end scoring. In order to improve the adaptability of the model, this embodiment introduces transfer learning to migrate the existing text-based scoring method to multimodal tasks, so that the model can not only understand text information, but also effectively parse the complex structure in rich text images.

[0066] More specifically, to train a multimodal model that can effectively perform automatic scoring tasks, it is first necessary to construct a high-quality multimodal scoring dataset. This embodiment collects and constructs training data from multiple dimensions, including real handwritten answer data and synthetic data.

[0067] During the collection of real handwritten data, this example collected students' actual handwritten answers and annotated them using a manual scoring method. Furthermore, to enhance data diversity and interpretability, some handwritten answers were converted to text using OCR and supplemented with manual scoring comments. This allows the model to simultaneously learn the correspondence between images, text, and scoring criteria during training.

[0068] At the same time, in order to expand the data scale and improve the generalization ability of the model, this embodiment uses typesetting tools such as LaTeX to generate synthetic data in rich text format, converts existing text scoring data into handwriting-style images, and simulates different handwriting styles through handwritten fonts, thereby improving the model's adaptability to multiple handwriting styles.

[0069] More specifically, during the model training phase, this embodiment adopts a gradual transfer learning strategy to fine-tune different components in stages to fully utilize existing text scoring knowledge while enhancing the model's ability to understand rich text images.

[0070] First, this example selects DETR as the image encoder and performs continuous pre-training on text recognition and TextVQA tasks to enable it to perceive structured information such as text, formulas, and tables in rich text images. Next, this example uses existing text scoring data to fine-tune the backbone of the MLLM, enabling it to learn the basic scoring patterns. Finally, this example uses the constructed multimodal scoring dataset to perform end-to-end fine-tuning on the entire model, adapting it to the handwritten answer scoring task, thereby improving the accuracy and stability of scoring.

[0071] Furthermore, the process of processing the test paper image using the DETR-based object detector includes:

[0072] Use DETR to detect key areas in the test image and generate anchor boxes containing text information and corresponding locations;

[0073] Filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor boxes to form image tokens;

[0074] MLP is used to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

[0075] Furthermore, existing large multimodal models usually use an image encoder based on Visual Transformer (ViT) to convert low-resolution images into fixed-length image tokens. However, this method often fails to capture all key information when processing rich text images (such as test papers or documents containing formulas, tables, and handwritten text), resulting in information loss. In order to more comprehensively express the fine-grained content in the image, this embodiment intends to use a target detector based on DETR (DEtectionTRansformer) as an image encoder to generate an image token sequence of variable length. This can more flexibly extract important areas in the image, improve information expression capabilities, and enhance the fusion effect with the text modality.

[0076] Specifically, in multimodal scoring tasks, relying solely on ViT-based Image Encoders to extract fixed-length image tokens often fails to fully express the complex information of text-rich images. Therefore, this embodiment uses a DETR-based object detector to enhance image encoding capabilities, enabling it to detect key areas in handwritten answers and efficiently integrate this information into the large language model to improve scoring quality.

[0077] More specifically, this embodiment proposes a DETR-based image tokenizer. DETR is used to detect the handwritten answer image and generate anchor boxes containing text information and corresponding positions. After filtering out low-confidence anchor boxes, a CNN+Pooling algorithm is used to encode the text information in the anchor boxes into image tokens. An MLP algorithm is then used to encode the two-dimensional positions into a token sequence. The tokens are then arranged in sequence to form a token stream.

[0078] More specifically, key anchor box screening: Although DETR can detect a large number of regions, not all regions are valuable for the scoring task. Therefore, this embodiment further designs a series of screening mechanisms to ensure that the model focuses on information relevant to scoring. This embodiment filters out low-confidence target areas through confidence thresholds, retaining only reliable detection results. In addition, this embodiment incorporates prior information such as the title of the scoring task into the DETR query, allowing the model to pay more attention to task-related text information.

[0079] More specifically, in the process of fusing text and image information, this embodiment adopts a cross-modal attention mechanism to enable deep interaction between text input and image token, ensuring that key information in the image can be effectively utilized by the large language model. Specifically, this embodiment regards contextual background information as prompt words, and uses the BERT model to encode the prompt words into DETR Query, so that the DETR model can understand the contextual information. At the same time, this embodiment introduces visual-text joint representation technology, based on the idea of ​​contrastive learning, and uses image-text matching tasks and text recognition tasks to train the model, so that the text and the corresponding image area are more closely aligned in the latent semantic space. Specifically, this embodiment uses typesetting tools such as LaTeX to create a text-typesetting image pair dataset. The image-text matching task requires the model to find the text corresponding to the typesetting; the text recognition task requires the model to use the typesetting image to generate the original text.

[0080] Furthermore, the process of filtering out low-confidence anchor boxes includes:

[0081] Low-confidence target areas are filtered out through a preset confidence threshold, and only anchor boxes with confidence higher than the confidence threshold are retained. The prior information of the scoring task is integrated into the DETR query.

[0082] Furthermore, the process of performing modal fusion through the cross-modal attention mechanism to obtain the fused representation includes:

[0083] Treat the contextual information as a prompt word and use the BERT model to encode the prompt word into a DETR query;

[0084] The visual-text joint representation technology is introduced. Based on the idea of ​​contrastive learning, the model is trained using image-text matching tasks and text recognition tasks. The text and the corresponding image areas are aligned in the latent semantic space to obtain a fused representation.

[0085] Furthermore, the process of generating prompt words using the prompt model optimized by reinforcement learning includes:

[0086] Two generative models, a prompt optimization model and a scoring model, that do not share parameters are designed and trained separately. The prompt optimization model is used to generate discrete prompt words, and the scoring model is used to score input data based on the prompt words. The prompt optimization model and the scoring model are connected directly or after being transformed through several layers of neural networks.

[0087] When training the prompt optimization model, freeze the parameters of the scoring model, concatenate the prompt words generated by the prompt optimization model with the input, and input them into the scoring model for evaluation. The error between the model output and the label generates reward information and sends it back to the prompt optimization model.

[0088] When training the scoring model, the parameters of the prompt optimization model are frozen, the prompt words generated by the prompt optimization model are concatenated with the input and input into the scoring model for evaluation, and the error between the model output and the label is back-propagated to the scoring model to optimize the parameters of the scoring model.

[0089] Furthermore, the optimization model is prompted to be trained using the Proximal Policy Optimization algorithm.

[0090] Furthermore, existing prompt optimization methods mainly rely on prompt learning, that is, adapting to specific tasks by optimizing continuous prompt vectors. However, the generalization ability of prompt learning is poor, making it difficult to apply to different types of problems. At the same time, optimizing the prompts and the model itself is relatively difficult. In order to improve the generalization ability of prompts and optimize model performance, this embodiment adopts a prompt model joint optimization method based on reinforcement learning, so that prompts can be more effectively adapted to different tasks. At the same time, combined with model fine-tuning, the overall effect of the scoring task is improved. Through reinforcement learning, the prompt words are dynamically adjusted so that they can maintain stability in different scenarios, and are coordinated with the model parameters to enhance the accuracy and adaptability of the automatic scoring system.

[0091] More specifically, the prompt model joint optimization method uses two generative models (the structures can also be different) that do not share parameters, namely the prompt optimization model and the scoring model. The prompt optimization model is used to generate (discrete) prompt words, while the scoring model is responsible for scoring the input data (based on the prompt words). The two models can be directly connected, that is, the output of the prompt optimization model is used as the input of the scoring model, or the output of the prompt optimization model can be transformed through several layers of neural networks and used as the input of the scoring model. During reasoning, the prompt optimization model generates prompt words based on the information of the question and the reference answer. The prompt words are concatenated with the student's answer and used as the input of the scoring model. The first token output by the scoring model is linearly transformed to obtain the score of the student's answer. During the training process, the two models are trained separately. When training the prompt optimization model, the parameters of the scoring model are frozen. The prompt words generated by the prompt optimization model are concatenated with the input and fed into the scoring model for evaluation. The error between the model output and the label generates reward information and is passed back to the prompt optimization model. When training the scoring model, the parameters of the prompt optimization model are also frozen. Similarly, the prompt words generated by the prompt optimization model are concatenated with the input and fed into the scoring model for evaluation. However, the error between the model output and the label is backpropagated to the scoring model to optimize the scoring model parameters. To save computing resources, both models can freeze some parameters or use LoRA to reduce the number of trainable parameters.

[0092] More specifically, regarding the training of the prompt optimization model: A variety of reinforcement learning algorithms can be used for the training of the prompt optimization model. This embodiment intends to select the Proximal Policy Optimization (PPO) algorithm to ensure the stability of the training. During the training process, the scoring model scores the student's answer based on the prompt words generated by the prompt optimization model. The mean squared error between the output score and the label is used as the reward, that is:

[0093]

[0094] in is the predicted value generated by the scoring model, y t are the labels of the training data.

[0095] In order to provide reward information to the prompt optimization model, this embodiment adopts a generalized advantage estimation approach, using:

[0096]

[0097] Computational advantages, including is the cumulative reward through the entire answer generation process, V(*) is the value function,

[0098] Fitted by a multi-layer neural network. According to the advantage function A, the optimized objective function can be obtained:

[0099]

[0100] in

[0101] More specifically, regarding the training of the scoring model: Unlike continuous prompts that rely heavily on specific parameters, the prompt words generated by the above method remain effective across different model parameters and even architectures. Therefore, after freezing the prompt optimization model, this embodiment can simply fine-tune the scoring model using backpropagation to enable it to learn more effectively for new tasks. Through repeated iterations, the model gradually adapts and optimizes its parameters, ultimately achieving a complete fit for the target task.

[0102] This embodiment proposes an end-to-end short-answer question scoring method based on a large language model, which reduces reliance on preprocessing steps such as OCR and improves scoring stability.

[0103] This embodiment is based on modality fusion of object detectors, using object detection instead of traditional image tokenizers to achieve variable-length image token representation, which can better grasp the details of the image.

[0104] This embodiment combines the prompt model joint optimization method with reinforcement learning, adopts a dual-model architecture (prompt generation model + scoring model), uses PPO to optimize prompts, generates prompt words suitable for different models, and optimizes the scoring model in combination with prompt words to improve the scoring effect.

[0105] This example constructs an automatic scoring method based on a large language model and conducts an in-depth analysis of its performance in scoring tasks. This not only expands the application scope of large language models in education, but also helps understand their capabilities and limitations in complex language tasks. Furthermore, this example explores the impact of different model architectures, fine-tuning strategies, and prompt engineering on scoring performance, providing a theoretical basis for the subsequent application of large language models in automatic scoring and related tasks.

[0106] In scenarios such as large-scale examinations, online education, and intelligent assessments, the scoring of open-ended questions has always been an important factor restricting teaching quality and efficiency. Manual scoring is not only time-consuming and labor-intensive, but may also be affected by subjective factors of the scorer, resulting in inconsistent scoring standards and affecting fairness. The application of an automatic scoring system can greatly improve the efficiency and stability of scoring, enabling students to obtain feedback more quickly, thereby optimizing the learning experience. The scoring method based on the large language model in this embodiment has stronger generalization and scalability than traditional methods. While reducing the cost of manual annotation, it can adapt to the scoring needs of different subjects and different question types, thereby providing a more intelligent solution for educational assessment.

[0107] Furthermore, the techniques involved in this research, such as prompt engineering and efficient parameter fine-tuning, not only help optimize the performance of the automatic scoring system but also provide a reference for the application of large language models in other vertical fields, such as medical Q&A, legal text analysis, and intelligent customer service. The results of this embodiment will not only promote the development of educational technology but also provide valuable methodological references for intelligent text processing tasks in other industries.

[0108] Example 2

[0109] Based on the same inventive concept, this embodiment also provides an automatic scoring system for open-ended questions based on a multimodal large model and target detection, including:

[0110] The image acquisition module is used to obtain test paper images containing answers to open-ended questions and perform end-to-end scoring using a multimodal large language model;

[0111] The object detection module is used to process the test paper images using a DETR-based object detector, extract key text information and generate corresponding image token sequences;

[0112] The modality fusion module is used to input the image token sequence and key text information into the multimodal large language model, perform modality fusion through the cross-modal attention mechanism, and obtain the fused representation;

[0113] The prompt and scoring module is used to generate prompt words using the prompt model optimized by reinforcement learning, and combine the prompt words with the fused representation, input them into the scoring model, and obtain the scoring results.

[0114] Furthermore, the target detection module includes:

[0115] Anchor frame generation unit, used to use DETR to detect key areas in the test volume image and generate anchor frames containing text information and corresponding positions;

[0116] The anchor box screening unit is used to filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor box to form an image token; and use MLP to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

[0117] The automatic scoring system for open-ended questions based on a multimodal large model and target detection provided in this embodiment has all the advantages of the automatic scoring method for open-ended questions provided in the first embodiment.

[0118] Example 3

[0119] This embodiment further discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first embodiment.

[0120] Example 4

[0121] This embodiment further discloses a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first embodiment are implemented.

[0122] Example 5

[0123] This embodiment further discloses a computer program product, including a computer program, which implements the steps of the method described in the first embodiment when executed by a processor.

[0124] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for automatic scoring of open-ended questions based on a multimodal large model and target detection, characterized in that: include: Obtain test paper images containing answers to open-ended questions and use a multimodal large language model for end-to-end scoring; The test paper image is processed using a DETR-based object detector to extract key text information and generate a corresponding image token sequence; Inputting the image token sequence and key text information into a multimodal large language model, performing modal fusion through a cross-modal attention mechanism, and obtaining a fused representation; A prompt word is generated using a prompt model optimized by reinforcement learning, and the prompt word is combined with the fused representation and input into a scoring model to obtain a scoring result.

2. The method according to claim 1, characterized in that The process of obtaining an image of the test paper containing the answers to the open-ended questions includes: Collect students' real handwritten answer pictures and manually annotate them; The LaTeX typesetting tool is used to generate synthetic data in rich text format, the existing text scoring data is converted into handwriting style images, and different handwriting styles are simulated through handwriting fonts.

3. The method according to claim 1, characterized in that The multimodal large language model is trained using a step-by-step transfer learning strategy. The training process includes: A DETR-based object detector is used as the image encoder and continuously pre-trained on text recognition and TextVQA tasks; Use existing text scoring data to fine-tune the backbone of the multimodal large language model; The entire model is fine-tuned end-to-end using the constructed multimodal ratings dataset.

4. The method according to claim 1, wherein The process of processing the test paper image using the DETR-based object detector includes: Use DETR to detect key areas in the test image and generate anchor boxes containing text information and corresponding locations; Filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor boxes to form image tokens; MLP is used to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

5. The method according to claim 4, characterized in that The process of filtering out low-confidence anchor boxes includes: Low-confidence target areas are filtered out by a preset confidence threshold, and only anchor boxes with confidence higher than the confidence threshold are retained. The prior information of the scoring task is integrated into the DETR query.

6. The method according to claim 1, wherein The process of performing modal fusion through the cross-modal attention mechanism to obtain the fused representation includes: Treat the contextual information as a prompt word and use the BERT model to encode the prompt word into a DETR query; The visual-text joint representation technology is introduced. Based on the idea of ​​contrastive learning, the model is trained using image-text matching tasks and text recognition tasks. The text and the corresponding image areas are aligned in the latent semantic space to obtain a fused representation.

7. The method according to claim 1, characterized in that The process of generating prompt words using the prompt model optimized by reinforcement learning includes: Design two generative models, a prompt optimization model and a scoring model, that do not share parameters, and train them separately; the prompt optimization model is used to generate discrete prompt words, and the scoring model is used to score input data based on the prompt words; the prompt optimization model and the scoring model are connected directly or after being transformed through several layers of neural networks; When training the prompt optimization model, the parameters of the scoring model are frozen, the prompt words generated by the prompt optimization model are concatenated with the input and then input into the scoring model for evaluation, and the error between the model output and the label is used to generate reward information which is then passed back to the prompt optimization model; When training the scoring model, the parameters of the prompt optimization model are frozen, the prompt words generated by the prompt optimization model are concatenated with the input and input into the scoring model for evaluation, and the error between the model output and the label is back-propagated into the scoring model to optimize the parameters of the scoring model.

8. The method according to claim 5, characterized in that The prompt optimization model is trained using the Proximal Policy Optimization algorithm.

9. An automatic scoring system for open-ended questions based on a multimodal large model and target detection, characterized by: include: The image acquisition module is used to obtain test paper images containing answers to open-ended questions and perform end-to-end scoring using a multimodal large language model; An object detection module is used to process the test paper image using a DETR-based object detector, extract key text information and generate a corresponding image token sequence; A modality fusion module is used to input the image token sequence and key text information into a multimodal large language model, perform modality fusion through a cross-modal attention mechanism, and obtain a fused representation; The prompt and scoring module is used to generate prompt words using the prompt model optimized by reinforcement learning, and combine the prompt words with the fused representation, input them into the scoring model, and obtain the scoring results.

10. The system according to claim 9, characterized in that The target detection module includes: Anchor frame generation unit, used to use DETR to detect key areas in the test volume image and generate anchor frames containing text information and corresponding positions; The anchor box screening unit is used to filter out low-confidence anchor boxes and use CNN+Pooling to encode the text information of the anchor box to form an image token; and use MLP to encode the two-dimensional position into a token sequence, which is arranged in sequence to form a token stream.

Citation Information

Patent Citations

  • Visual positioning method and device based on hierarchical cross-modal context attention mechanism

    CN116152810A

  • Visual positioning method based on multi-modal feature alignment

    CN117934803A

  • Artificial intelligence psychological assessment method and system based on multi-modal input

    CN119108112A

  • Defect detection method and system for unmanned aerial vehicle inspection equipment in open domain

    CN119942378A

  • Systems and methods for question-answering using a multi-modal end to end learning system

    US20240095455A1