Test paper structured storage method and device based on multi-modal large model
Through the structured storage method of test papers based on multimodal large models, the existing technology is solved by the problem of changing diversity and high maintenance costs of test papers, and the efficient structured extraction and complete storage of test paper information is achieved.
Patent Information
- Application Number
- CN202510193933.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing rules-based structured test papers are difficult to cope with the diversity of test papers in the field of education, with high maintenance costs and difficult to deal with complex typesetting.
The structured storage method of the test paper based on the multimodal large model is adopted, and the figures and tables in the test paper are stored and retrieved by databases, and the generalization ability of the structured extraction of the test paper content is improved by combining the multimodal large model.
It significantly improves the generalization ability of structured extraction of test papers, reduces maintenance costs, realizes the integrity and traceability of test paper information, and can efficiently handle complex typesetting.
Smart Images

Figure CN120216746A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent education technology and relates to a method and device for structured storage of test papers based on a multimodal large model. Background Art
[0002] Currently, most methods for structuring test paper content adopt rule-based methods. Rule-based methods have many limitations in test paper structuring:
[0003] 1. Poor rule adaptability, making it difficult to cope with the diverse changes in test papers in the education field:
[0004] Modern educational examinations are constantly innovating, and new question types emerge in an endless stream. For example, in some emerging disciplines or interdisciplinary examinations, there will be question types that integrate multiple knowledge fields and have unique answering requirements. The structures of such question types are very different from traditional question types, and the rules set based on past experience are difficult to effectively analyze them. Secondly, there are also differences in the question-setting styles of different regions and educational institutions. In order to highlight their own characteristics, some regions or schools will adopt unique question expressions, including the use of specific terms, dialect words, or cases with local cultural characteristics. This makes it difficult for the structured processing system based on general rules to understand the semantic connotations, and thus unable to correctly divide the question structure and extract key information. Moreover, the layout of test papers is becoming increasingly diverse. For example, some art design test papers may have a large number of mixed layouts of irregular pictures and text, or even text wrapping around pictures, or in math test papers, there are complex formula derivations that span multiple pages and are not associated with question numbers and question stems in a conventional format. The established rules cannot accurately identify the question paragraphs and logical relationships.
[0005] 2. High rule maintenance costs, requiring continuous manual update and adjustment of rules, lacking flexibility and intelligence:
[0006] The rule-based test paper structuring method is trapped in the dilemma of rule maintenance due to various dynamic changes in the education field. The update of curriculum standards, the innovation of teaching methods, the progress of educational technology, and the transformation of educational evaluation concepts all prompt the continuous innovation of test paper question types, structures, and content. Maintenance personnel need to deeply explore the essence and characteristics of various changes, manually reconstruct and debug rules to adapt to new test paper forms. This process is extremely complex and time-consuming, resulting in high rule maintenance costs and the need for continuous investment of a large amount of manual effort for update and adjustment.
[0007] In summary, the rule-based test paper structuring method has obvious drawbacks. These limitations seriously restrict the effectiveness and accuracy of the rule-based method in test paper structuring. Future research needs to focus on improving the generalization ability of the model and exploring how to combine chart detection with in-depth understanding and content matching to improve the practical application value of chart detection technology.
[0008] Therefore, how to provide a method and device for structured storage of test papers that can effectively improve the generalization ability of structured extraction of test paper content is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0009] In view of this, the present invention proposes a method and device for structured storage of test papers based on a multimodal large model. By using the method of database storage and retrieval extraction for the graphs and tables in the test papers, while realizing the integrity of the stored content of the test papers, it effectively improves the generalization ability of structured extraction of test paper content in combination with the multimodal large model.
[0010] In order to achieve the above object, the present invention adopts the following technical solutions:
[0011] The present invention discloses a method for structured storage of test papers based on a multimodal large model, including the following steps:
[0012] Multimodal large model training step:
[0013] Obtain a test paper training data set, where the test paper training data set includes a test paper document and corresponding text label data in the test paper document, and perform primary training based on the test paper training data set to obtain a pre-trained multimodal large model;
[0014] Test paper structured storage step:
[0015] Obtain a test paper document to be stored, extract n chart areas in the test paper document to be stored and crop them to generate n chart pictures, and replace and fill n identifiers at the corresponding cropping positions to generate a new test paper document, where n≥1; the identifiers are unique, and store the cropped chart pictures in the database;
[0016] Input the new test paper document into the pre-trained multimodal large model for text content extraction;
[0017] Bind the identifiers in the text content correspondingly and embed retrieval information to obtain a structured text of the test paper and store it in the database. The retrieval information is used to index the storage address of the chart pictures in the database; the structured text of the test paper and the corresponding chart pictures are used as the final storage results.
[0018] Preferably, the text label data includes a tree structure identifier, and the tree structure identifier is used to distinguish the question layout structure in the test paper document.
[0019] Preferably, the multimodal large model training step further includes a continued pre-training step:
[0020] Execute the above-mentioned test paper structured storage step on multiple test paper documents to be stored in the target field, and obtain the test paper structured text in the target field;
[0021] Perform post-processing and revision on the test paper structured text in the target field according to preset conditions to obtain the revised test paper recognition result in the target field;
[0022] Use the multiple test paper documents to be stored in the target field and the corresponding revised test paper recognition results in the target field as new test paper training data, merge them into the test paper training data set, and execute the above-mentioned multi-modal large model training step again.
[0023] Preferably, the post-processing and revision step includes: the step of revising the text content in the test paper structured text in the target field, and / or, the step of revising the output format of the text content in the test paper structured text in the target field.
[0024] Preferably, the test paper structured storage step further includes:
[0025] Decompose the test paper to be stored into multiple test paper pictures according to its arrangement order;
[0026] Perform size preprocessing operations on the test paper pictures, and extract n chart areas in the preprocessed test paper pictures.
[0027] Preferably, the step of generating the chart picture in the test paper structured storage step includes:
[0028] Locate the figure or table in the test paper document to be stored, and obtain the regional position information of the figure or table in the test paper document;
[0029] Extract the pixel data of the figure or table from the test paper document to be stored according to the regional position information, and delete the pixel data in the corresponding regional position;
[0030] Fill the identifier in the regional position.
[0031] Preferably, the identifier is a text identifier, and the text identifier contains a chart serial number, and the chart serial number is generated according to the typesetting order of the chart area in the test paper document to be stored.
[0032] Preferably, the output format of the text content is a tree structure, and the text content includes: multi-level question information; embedding the level identifier of each level in the question information; the level is determined according to the question typesetting order in the test paper document.
[0033] Preferably, when receiving an instruction from a user to retrieve text content containing an identifier, the retrieved information is used to index and obtain corresponding chart pictures in a database, and the obtained chart pictures are used to be restored to the position in the text content where the identifier is located.
[0034] The present invention also provides a test paper structured storage device according to the above-mentioned test paper structured storage method based on a multimodal large model, including:
[0035] A data preparation module, configured to obtain a test paper document to be stored, extract n chart areas from the test paper document to be stored and crop them to generate n chart pictures, and replace and fill n identifiers at the corresponding cropping positions to generate a new test paper document, where n≥1; the identifiers are unique, and the cropped chart pictures are stored in a database;
[0036] A multimodal large model module, configured to perform text content extraction on the new test paper document;
[0037] A structured storage module, configured to bind and embed retrieval information corresponding to the identifiers in the text content, obtain a structured text of the test paper and store it in the database, where the retrieval information is used to index the storage address of the chart pictures in the database; the structured text of the test paper and the corresponding chart pictures are used as the final storage result.
[0038] It can be seen from the above technical solutions that, compared with the prior art, the beneficial effects of the present invention include:
[0039] 1. Improve generalization ability: By introducing a multimodal large model, it can effectively cope with the diversity changes of test paper content and form. Compared with traditional rule-based methods, the present invention can better adapt to new question types, interdisciplinary questions, and the question-setting styles of different regions and institutions, thus significantly improving the generalization ability of test paper structured extraction.
[0040] 2. Reduce maintenance costs: Traditional rule-based methods require frequent manual updates and adjustments of rules, while the present invention optimizes the model performance through model training and data accumulation, reduces the dependence on manual intervention, and significantly reduces the maintenance costs.
[0041] 3. Achieve storage integrity: For the chart content in the test paper, the method of database storage and retrieval extraction is adopted to ensure the integrity and traceability of the test paper information. The charts and text content are bound by identifiers, realizing structured storage while retaining the relevance of the original information.
[0042] 4. Enhance flexibility and intelligence: The present invention supports dynamic updates and iterations. By post-processing and revising the test papers in the target field and feeding the revision results back to model training, the accuracy and adaptability of the model are further improved. This closed-loop optimization mechanism enables the system to continuously improve and meet the diverse needs of the education field.
[0043] 5. Efficiently handle complex layout: For test papers containing complex formulas, irregular picture layouts, or text wrapping around pictures, the present invention can accurately locate the chart area and extract relevant information, solving the limitations of traditional methods in dealing with complex layouts.
[0044] In summary, the present invention not only improves the efficiency and accuracy of test paper structured processing but also provides a new solution for the development of intelligent education technology, having important practical application value and promotion significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings;
[0046] Figure 1 It is a schematic flow chart of a test paper structured storage method based on a multi-modal large model provided by an embodiment of the present invention;
[0047] Figure 2 It is a schematic diagram of the low-rank matrix fine-tuning principle of a multi-modal large model provided by an embodiment of the present invention;
[0048] Figure 3 It is a schematic diagram of a test paper picture provided by an embodiment of the present invention;
[0049] Figure 4 It is a schematic diagram of the process of generating text content according to a test paper picture provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0051] Embodiment 1:
[0052] As Figure 1As shown in the figure, the first aspect of the embodiment of the present invention provides a method for structured storage of test papers based on a multimodal large model, as Figure 1 shown, including the following steps:
[0053] Multimodal large model training step:
[0054] Obtain a test paper training data set, which includes test paper documents and corresponding text label data in the test paper documents. Based on the test paper training data set, conduct primary training to obtain a pre-trained multimodal large model;
[0055] Test paper structured storage step:
[0056] Obtain the test paper document to be stored, extract n chart areas in the test paper document to be stored and crop them to generate n chart pictures, and replace and fill n identifiers at the corresponding cropping positions to generate a new test paper document, where n≥1; the identifiers are unique, and store the cropped chart pictures in the database;
[0057] Input the new test paper document into the pre-trained multimodal large model for text content extraction;
[0058] Bind and embed the retrieval information corresponding to the identifiers in the text content to obtain the structured text of the test paper and store it in the database. The retrieval information is used to index the storage address of the chart pictures in the database; the structured text of the test paper and the corresponding chart pictures are used as the final storage results.
[0059] In one embodiment, obtain the existing question data from the ES database or the SQL database as the test paper training data set. These data are all derived from various carefully compiled test papers and have undergone meticulous annotation processing. The so-called annotation is to perform text recognition on the test paper pictures converted from PDF format test papers by means of advanced optical character recognition technology. Subsequently, according to specific rules, accurately extract the tree structure contained in the test paper data and set it as the label data corresponding to the test paper pictures.
[0060] In this embodiment, the text label data includes a tree structure identifier, which is used to distinguish the question layout structure in the test paper document.
[0061] In one embodiment, the architecture of the multimodal large model includes a visual encoder, a projector, and an image decoder. The training steps of the multimodal large model include:
[0062] In the initial training stage of the model, the image decoder and the projector parts in the multimodal large model will be frozen, so that their parameters do not participate in the training process in this stage, but only focus on conducting targeted training on the backbone large language model.
[0063] To further effectively reduce the stringent requirements for hardware facilities, this embodiment adopts the Lora (Low-Rank Adaptation) method to train the model. The core principle of the Lora method lies in introducing the low-rank matrix factorization technology. By adding a small number of trainable low-rank matrices on the basis of the original model architecture, the model is finely tuned and optimized in a lightweight manner. As Figure 2 shown, the formula for its main idea is:
[0064] h = W0x + γΔWx = W0x + γBAx;
[0065] Fine-tuning is carried out by adding the product of two low-rank matrices on the basis of the pre-trained model weights. Specifically, the weight W of the model is divided into two parts, where the weight matrix W is decomposed into the product of matrix B and matrix A, so as to achieve the purpose of reducing the rank of the matrix, thereby saving training resources.
[0066] The Lora method cleverly utilizes the special properties of low-rank matrices, enabling the model to adapt to and learn new data distributions with less computational resources and time costs during the training process. Thus, the smooth progress of model training can still be ensured under limited hardware conditions, and the performance and generalization ability of the model can be improved to a certain extent, laying a solid foundation for the effective deployment of subsequent multi-modal large models in various practical application scenarios.
[0067] In one embodiment, the multi-modal large model training steps further include a continued pre-training step. Continued pre-training is a further in-depth training process carried out on the basis of the model that has completed pre-training. The specific steps include:
[0068] Execute the test paper structured storage step on multiple test paper documents to be stored in the target domain to obtain the structured text of the test papers in the target domain;
[0069] Post-process and revise the structured text of the test papers in the target domain according to preset conditions to obtain the revised test paper recognition results in the target domain;
[0070] Use multiple test paper documents to be stored in the target domain and the corresponding revised test paper recognition results in the target domain as new test paper training data, merge them into the test paper training data set, and execute the multi-modal large model training steps again.
[0071] When this embodiment is specifically executed, during the initial pre-training period, the model will fully explore large-scale image and text data and study the general characteristics of language from them, such as core elements like accurate word vector representations and complex grammar structure knowledge. The continued pre-training step of this embodiment then implements additional targeted training measures for the pre-trained model according to the unique requirements of specific tasks or the distinct characteristics of new data.
[0072] For example, a model may have been pre-trained in a general corpus environment. When it needs to be applied to a specific field of medical literature, it is further pre-trained using rich medical literature data. In this way, the model can better fit the semantic subtleties, unique professional vocabulary, and possible special image requirements involved in medical literature, and thus can accomplish complex tasks such as medical text classification and key information extraction with higher accuracy. The continued pre-training step is undoubtedly an extremely efficient strategy that can fully utilize the results of existing models and cleverly adapt to new scenarios by fine-tuning model parameters. It can significantly reduce the huge computing resources and massive data investment required to train the model from scratch, and effectively improve the actual performance of the model in new tasks.
[0073] This step uses the model obtained in the initial stage to gradually accumulate a considerable amount of new test paper training data during its continuous operation. Subsequently, these accumulated new test paper training data are converted into a data format that meets the requirements of continued pre-training. When the accumulated data volume reaches a certain scale, a comprehensive training is carried out using the new test paper training data and the preliminary pre-trained test paper training data set. This training method can not only significantly enhance the model's adaptability and processing level when dealing with new test paper types, so that it can still analyze and answer with ease when facing unprecedented question structure and knowledge test points; at the same time, it can also steadily maintain the model's excellent processing ability when dealing with those test paper types that have been encountered, ensuring that it remains efficient and accurate when facing familiar question types and knowledge points. Through this cyclical process, the model can continue to adapt to the increasingly diverse and complex situation of test paper types caused by the dynamic changes in education policies, and always maintain excellent performance and adaptability in the intelligent application scenarios in the field of education, providing solid and reliable technical support for the quality improvement and intelligent transformation of education and teaching.
[0074] In the initial model training phase, the first batch of reliable training data was obtained through data preprocessing and annotation, laying the foundation for the subsequent model training phase. In this process, the vectorized storage of graphs and tables in the test paper provides tools for subsequent online processing of the test paper to ensure the integrity of the test paper. During the initial model operation, a large amount of test paper data was obtained, and the model was used to continue pre-training, so that the model can not only ensure the expansion capability when facing new data, but also stably handle the test paper structure that has already appeared. These two steps together ensure the efficiency and stability of the model when processing diverse data and adapting to new information.
[0075] In this embodiment, the Qwen-VL model is used as a multimodal large model. The steps of processing the test paper image of this model are as follows:
[0076] Image preprocessing: Qwen-VL receives multiple test paper images and scales them to a resolution of 448x448 to ensure their clarity and size are suitable for subsequent analysis.
[0077] Extract image features: In this step, the features of the image are extracted through a vision encoder, specifically the VisionTransformer. During the feature extraction process, ViT first divides the input image into multiple patches of a fixed size, just like disassembling the image into individual "words". For example, an image of 224x224 pixels will be divided into 14x14 image patches of 16x16 pixels. These patches are then mapped into a high-dimensional space to form a serialized image representation. Each image patch is also assigned a position encoding during the mapping process to preserve its spatial position information in the original image, ensuring that the model can understand the relative position relationship between the image patches. Next, this sequence of encoded image patches is fed into the core part of the Transformer model. Through the multi-head self-attention mechanism, the model can capture the complex dependencies between the image patches and achieve efficient integration of the global information of the image. At the same time, the feed-forward network performs further feature transformation and extraction on each image patch, and the addition of layer normalization and residual connections optimizes the training process of the model, improving the stability and convergence speed of training. After sufficient training, ViT can accurately extract the rich features of the image.
[0078] Fusion of visual and text information: The core module in this process is the vision-language adapter, that is, the projector. The simplest linear layer method can be used to align the multi-modal information and send it to the downstream large language model for decision-making.
[0079] Understanding of the large language model: After receiving the upstream vision-text information, these inputs are serialized and input into a network composed of multiple Transformer layers together with the position encoding. At each layer, the model analyzes each element in the sequence through the multi-head self-attention mechanism and uses the feed-forward network for non-linear transformation. Layer normalization and residual connections ensure the stability of the information flow and the training efficiency of the network. After multiple layers of processing, the model finally generates an output sequence that combines visual and language information. According to the specific task, the final text response is generated through a decoding strategy. In the process of processing the test paper, it can be decomposed into two processes. First, through the text extraction ability of the model itself, the text in the test paper is extracted completely and accurately. Then, according to the format of the training data, the content of the test paper is organized into a tree structure for convenient subsequent storage process. As Figures 3-4 shown, it shows the effect of the multi-modal large model extracting the text content of the sixth grade (volume 1) mathematics mid-term test paper.
[0080] It should be noted that the multimodal large model is not limited to a single model architecture, and replaceable algorithms include InternVL2, GLM4V, etc.
[0081] In one embodiment, the post-processing revision step includes: a step of revising the text content in the structured text of the test paper in the target field, and / or a step of revising the output format of the text content in the structured text of the test paper in the target field.
[0082] In this embodiment, in order to improve the output quality of the model, the post-processing revision can be performed in the form of prompt words, and the prompt words can be decomposed into two processes, including: content extraction and output in the format of the training data.
[0083] In one embodiment, the test paper structured storage step further includes:
[0084] In this process, first, the PDF test papers to be stored are disassembled into multiple test paper pictures in an orderly manner;
[0085] Perform size preprocessing operations on the test paper pictures, extract n chart areas from the preprocessed test paper pictures, and input them into the multimodal large model together. Use the multimodal large model obtained in the training stage to extract the content of the test paper in PDF format passed in online.
[0086] In one embodiment, the database stores the final storage results in partitions in units of test paper documents.
[0087] In one embodiment, the tables and picture elements in the test paper pictures require a special processing process. The steps of generating chart pictures in the test paper structured storage step include:
[0088] Locate the figures or tables in the test paper document to be stored, and obtain the regional position information of the figures or tables in the test paper document;
[0089] According to the regional position information, extract the pixel data of the figures or tables from the test paper document to be stored, and delete the pixel data in the corresponding regional position;
[0090] Fill the regional position with identifiers.
[0091] When this embodiment is specifically executed, relying on the processing results of the previous multi-modal large model and the database constructed and stored for each set of test paper figures and tables during the data preparation process in the model preparation stage, the complete structured storage of the test paper is achieved. First, use a chart detection model that has been trained and has reliable performance to accurately locate the figures and tables within the test paper image. Then, according to the position information determined by the positioning, carefully extract the pixel data of the figures and tables from the image, and at the same time fill appropriate retrieval information at the original positions of the figures and tables to build a strong basis for finding the corresponding figures and tables in the database later. The data of these figures and tables will be stored in an orderly manner in the database independently established for each set of test papers. By implementing the above series of processes, the inherent defect that the current multi-modal large model cannot accurately locate figures and tables in the test paper scenario can be effectively compensated, and finally the goal of comprehensively and completely extracting test paper information can be achieved, providing a solid data foundation and strong support for subsequent related applications and analyses.
[0092] The above process first extracts the text information in the test paper through a multi-modal large model and performs structured output on the test paper. When storing the structured content, to ensure the integrity of the test paper content, the method of establishing a chart database is used to embed the picture and table data into the corresponding positions of the structured test paper.
[0093] In one embodiment, the identifier is a text identifier, and the text identifier contains a chart serial number, and the chart serial number is generated according to the layout order of the chart area in the test paper document to be stored.
[0094] When this embodiment is specifically executed, pixel indexes will be filled in the corresponding positions of the figures and tables. For example, on the first figure of a certain set of test papers, it is marked with black text " Figure 1 ", which indicates that it is the first figure in this set of test papers. And the figure with the index "1" in the database is exactly the original test paper image extracted during the data preparation stage. When performing structured storage operations, the original image data is embedded into its original position by virtue of the index, so as to ensure the integrity and coherence of the test paper content, and enable the structured storage of the test paper to be accurately and effectively realized.
[0095] In one embodiment, the output format of the text content is a tree structure, and the text content includes: multi-level question information; embedding the level identifier of its corresponding level for each level of question information; the level is determined according to the question layout order in the test paper document.
[0096] In one embodiment, the retrieval information is used to index and obtain the corresponding chart picture in the database when receiving an instruction from the user to retrieve the text content containing the identifier, and the retrieved chart picture is used to be restored to the position in the text content where the identifier is located.
[0097] Embodiment Two:
[0098] In the second aspect of the embodiments of the present invention, a test paper structured storage device is further provided, which applies the test paper structured storage method based on a multi-modal large model provided in the first aspect of the embodiments. The device includes:
[0099] A data preparation module, configured to obtain a test paper document to be stored, extract n chart areas from the test paper document to be stored and crop them to generate n chart pictures, and replace and fill n identifiers at the corresponding cropping positions to generate a new test paper document, where n≥1; the identifiers are unique, and the cropped chart pictures are stored in a database;
[0100] A multi-modal large model module, configured to extract text content from the new test paper document;
[0101] A structured storage module, configured to bind and embed retrieval information corresponding to the identifiers in the text content, obtain a structured test paper text and store it in the database, where the retrieval information is used to index the storage address of the chart pictures in the database; the structured test paper text and the corresponding chart pictures are used as the final storage results.
[0102] The above has introduced in detail the test paper structured storage method and device based on a multi-modal large model provided by the present invention. In this embodiment, specific examples are used to illustrate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0103] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined in this embodiment can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this embodiment, but will conform to the widest scope consistent with the principles and novel features disclosed in this embodiment.
Claims
1. A method for storing test papers in a structured manner based on a multimodal large model, characterized in that: The steps include: Multimodal large model training steps: Acquire a test paper training data set, wherein the test paper training data set includes a test paper document and corresponding text label data in the test paper document, and perform primary training based on the test paper training data set to obtain a pre-trained multimodal large model; Steps for structured storage of test papers: Obtain the test paper document to be stored, extract n chart areas in the test paper document to be stored and cut them to generate n chart images, and replace and fill n identifiers corresponding to the cut positions to generate a new test paper document, where n≥1; the identifier is unique, and the cut chart images are stored in the database; Inputting the new test paper document into the pre-trained multimodal large model to extract text content; The identifiers in the text content are bound to and embedded in the search information to obtain the test paper structured text and store it in the database. The search information is used to index the storage address of the chart image in the database; the test paper structured text and the corresponding chart image are used as the final storage results.
2. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The text tag data includes a tree structure identifier, and the tree structure identifier is used to distinguish the question layout structure in the test paper document.
3. The test paper structured storage method based on a multimodal large model according to claim 1 is characterized in that: The multimodal large model training step also includes continuing the pre-training step: Execute the test paper structured storage step on a plurality of test paper documents to be stored in the target field to obtain the test paper structured text in the target field; Post-process and revise the structured text of the test paper in the target field according to the preset conditions to obtain the revised recognition result of the test paper in the target field; The multiple test paper documents to be stored in the target field and the corresponding revised target field test paper recognition results are used as new test paper training data, merged into the test paper training data set, and the multimodal large model training step is performed again.
4. The method for storing test papers in a structured manner based on a multimodal large model according to claim 3 is characterized in that: The post-processing revision step includes: a step of revising the text content in the structured text of the test paper in the target field, and / or a step of revising the output format of the text content in the structured text of the test paper in the target field.
5. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The test paper structured storage step also includes: Decomposing the test paper to be stored into a plurality of test paper images according to their arrangement order; A size preprocessing operation is performed on the test paper image, and n chart areas in the preprocessed test paper image are extracted.
6. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The step of generating the chart image in the step of storing the test paper in a structured manner comprises: Locate the graph or table in the test paper document to be stored, and obtain the regional position information of the graph or table in the test paper document; According to the regional position information, pixel data of the graph or table are extracted from the test paper document to be stored, and the pixel data in the corresponding regional position is deleted; The identifier is populated in the region position.
7. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The identifier is a text identifier, and the text identifier includes a chart serial number, and the chart serial number is generated according to the layout order of the chart area in the test paper document to be stored.
8. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The output format of the text content is a tree structure, and the text content includes: multi-level test question information; for each level of the test question information, the level identifier of the level to which it belongs is embedded; and the level is determined according to the layout order of the test questions in the test paper document.
9. The method for storing test papers in a structured manner based on a multimodal large model according to claim 1, characterized in that: The retrieval information is used to index the corresponding chart image in the database when receiving a user's instruction to retrieve the text content containing the identifier, and the indexed chart image is used to restore to the position in the text content where the identifier is located.
10. A test paper structured storage device according to any one of claims 1 to 9, characterized in that: include: The data preparation module is used to obtain the test paper document to be stored, extract n chart areas in the test paper document to be stored and cut them to generate n chart images, and replace and fill n identifiers corresponding to the cut positions to generate a new test paper document, where n≥1; the identifier is unique, and the cut chart images are stored in the database; A multimodal large model module, used for extracting text content from the new test paper document; A structured storage module is used to bind and embed the identifiers in the text content into retrieval information, obtain the test paper structured text and store it in the database, the retrieval information is used to index the storage address of the chart image in the database; the test paper structured text and the corresponding chart image are used as the final storage result.
Citation Information
Cited By
Test paper splitting method, computer program product, equipment and storage medium
CN121074926A