Test paper structure element type identification method based on contextual information fusion
By using a method based on context information fusion in the identification of the structural element of the test paper, and using a large language model to process the context information in the test paper document, the problem of low accuracy in the identification of structural element of the test paper in the prior art is solved, and more efficient and accurate identification of structural element of the test paper is achieved.
Patent Information
- Application Number
- CN202510304145.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems of low conversion efficiency and low accuracy when identifying structural elements of the test paper, especially when facing test papers with irregular formats or non-standard typesettings, the recognition accuracy rate will be greatly reduced.
The test paper structure element type recognition method based on context information fusion is adopted. By pre-processing the input test paper document, the context information of the target line text is extracted, and then spliced with the current line text and input it into the fine-tuned large language model to output the category to which the target line text belongs.
It improves the accuracy of identification of structural elements of the test paper, can better understand the overall structural logic of the test paper, adapt to a variety of test question formats and text styles, and maintains high robustness.
Smart Images

Figure CN120236299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text recognition, and more specifically, to a method for identifying the types of test paper structure elements based on context information fusion. Background Art
[0002] With the continuous progress of informatization and intelligent technologies, traditional education evaluation methods are gradually shifting from paper-based test papers to electronic and intelligent directions. In this process, how to efficiently convert a large number of existing paper or scanned test papers into electronic versions with clear structure markings has become an important technical problem.
[0003] Traditional test paper digitization mainly relies on manual input, that is, manually segmenting, marking, and classifying the test question text. This is not only time-consuming and laborious, but also prone to human omissions and inconsistencies. For example, for a test paper containing dozens of multiple-choice questions, fill-in-the-blank questions, or short-answer questions, it is necessary to identify the stem, options, reference answers, and even hint explanations or analysis content of each question. If the entire process relies on manual operation, the efficiency is often extremely low when dealing with a large number of test papers.
[0004] Some early automation technologies attempted to identify the test paper structure through rule matching. These rules are based on the specific formats and layout features of elements such as stems, options, answers, and others in the test paper. However, when faced with test papers with non-standard formats or layouts, the recognition accuracy of this method will drop significantly. The rule matching technology lacks sufficient flexibility and adaptability to cope with the diversity and complexity of the test paper structure.
[0005] With the development of machine learning technology, some studies have begun to attempt to use simple machine learning algorithms to identify the test paper structure. These algorithms extract the features of the test paper text and then use a classifier to classify each line of text. Although this method has improved the automation level to a certain extent, due to the limitations of the algorithms, they still have problems with insufficient accuracy when dealing with complex test paper structures.
[0006] In recent years, natural language processing technology has developed rapidly. In particular, the emergence of large-scale pre-trained language models (such as BERT, GPT, GLM series models, etc.) has brought new opportunities for tasks such as text classification, named entity recognition, and text structured analysis. By fine-tuning the pre-trained language model for specific domains and tasks, the model can be made to have the ability to quickly and accurately identify the semantic structure of text.
[0007] However, there are still the following difficulties in education evaluation and test paper processing:
[0008] 1. Diversified test paper formats: Different question-setting teachers or institutions often adopt unique test paper formats, including differences in line spacing, delimiters, and option marking formats;
[0009] 2. Text feature ambiguity: In the test paper content, there may be a high similarity in text features between the question stem, options, and answers. Especially when there is a lack of context, it is very difficult to determine its category based on the current line of text alone;
[0010] 3. Context association requirement: There is often a context association between the question stem, its options, and answers. It is difficult to accurately judge the element category solely relying on the text features of the current line. The model needs to consider the text information of the previous and next lines.
[0011] Due to the above difficulties, when converting a paper or scanned test paper into an electronic version, there are defects such as low conversion efficiency and low accuracy in identifying the structural elements of the test paper.
[0012] Therefore, how to improve the accuracy of identifying the structural elements of the test paper is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0013] In view of the above problems, the present invention provides a method for identifying the type of test paper structural elements based on context information fusion to at least solve some of the technical problems mentioned in the above background technology.
[0014] To achieve the above object, the present invention adopts the following technical solutions:
[0015] A method for identifying the type of test paper structural elements based on context information fusion, including the following steps:
[0016] Preprocess the input test paper document to obtain multiple lines of text in the test paper document;
[0017] For the target line text in the multiple lines of text, extract the first N lines of text and the last N lines of text of the target line text as the context information of the target line text;
[0018] Concatenate the first N lines of text, the target line text, and the last N lines of text, and input the concatenated result into a fine-tuned large language model to output the category to which the target line text belongs.
[0019] Further, the preprocessing specifically includes:
[0020] Extract the text line sequence from the test paper document;
[0021] Delete the spaces and line breaks in each line of text;
[0022] Uniformly standardize the full-width punctuation and half-width punctuation in each line of text;
[0023] Store the processed multiple lines of text in a list in order.
[0024] Further, extracting the text line sequence from the test paper document specifically includes: recognizing the test paper document through OCR to obtain the text line sequence.
[0025] Further, the preprocessing further includes: automatically identifying the question number lines in the test paper document using regular expressions.
[0026] Further, if the upper context information or lower context information of the target line text is less than N lines, padding processing is performed.
[0027] Further, the large language model uses glm4-9b-chat.
[0028] Further, the fine-tuning steps of the large language model include:
[0029] Determine the training dataset, where the training samples in the training dataset are multiple lines of text information with real labels; the real label is the category of the line text information;
[0030] For each line of text information, construct the input representation as: "[First N lines of text][Special separator][Current line of text][Special separator][Last N lines of text]";
[0031] Load the weights of the large language model and set the hyperparameters of the large language model;
[0032] Batch the training samples and input them into the large language model, and obtain the prediction distribution through forward propagation;
[0033] Based on the prediction distribution, combined with the real label, obtain the cross-entropy loss;
[0034] According to the cross-entropy loss, use the optimizer to update the hyperparameters until convergence to obtain the fine-tuned large language model.
[0035] Further, the [Special separator] is [SEP].
[0036] Further, for each input current text line, retain its corresponding real label as the supervision signal.
[0037] Further, the real label is expressed as: (<Question stem>, <Options>, <Answer>, <Others>).
[0038] From the above technical solutions, it can be seen that compared with the prior art, the present invention discloses a method for identifying the types of test paper structure elements based on context information fusion, which has the following beneficial effects:
[0039] 1. The present invention combines the text of the previous and next lines of the context window with the text of the current line and inputs them into a pre-trained large language model to form a context-enhanced line-level classification method. This context-aware strategy breaks the limitation of traditional single-line independent classification, enabling the model to better understand the overall structural logic of the test paper and improving the accuracy of identifying the structural elements of the test paper.
[0040] 2. The present invention fine-tunes a large-scale pre-trained model such as glm4-9b-chat, enabling the model to automatically adapt to various test question formats and text styles in contexts-rich scenarios.
[0041] 3. In response to the noise and uncertainties in the test paper scenario, the present invention proposes a strategy for dynamically adjusting the size of the context window, enabling the model to maintain high robustness when facing test questions of different complexities.
[0042] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on the provided accompanying drawings without creative efforts.
[0044] Figure 1 It is a schematic flow chart of a method for identifying the structural elements of a test paper based on context information fusion provided by an embodiment of the present invention.
[0045] Figure 2 It is a schematic flow chart of fine-tuning a large language model provided by an embodiment of the present invention. Detailed Embodiments
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0047] An embodiment of the present invention discloses a method for identifying the type of structural elements of a test paper based on context information fusion. Refer to Figure 1 As shown, the method includes the following steps
[0048] S1. Preprocess the input test paper document to obtain multiple lines of text in the test paper document;
[0049] S2. For the target line text in the multi-line text, extract the first N lines of text and the last N lines of text of the target line text as the context information of the target line text;
[0050] S3. Concatenate the first N lines of text, the target line text, and the last N lines of text, and input the concatenation result into the fine-tuned large language model to output the category to which the target line text belongs.
[0051] Next, each of the above steps will be described in detail.
[0052] In the above step S1, preprocess the input test paper document to obtain the multi-line text in the test paper document; the preprocessing includes:
[0053] (1) Extract the text line sequence from the test paper document; for example, the text line sequence can be obtained by OCR recognition of PDF or scanned images; this OCR (Optical Character Recognition) is a technology that converts the text in the image into machine-readable text;
[0054] (2) Delete the spaces and line breaks in each line of text;
[0055] (3) Uniformly standardize the full-width punctuation and half-width punctuation in each line of text;
[0056] (4) Store the processed multi-line text in a list in order to prepare for the subsequent operation of inputting into the model.
[0057] The above preprocessing may also include:
[0058] Use regular expressions to automatically identify the question number lines in the test paper document; for example, the lines starting with "1.", "1)", or "(1)" are potential markers for the question stem lines, or markers such as Arabic numeral numbers, Chinese numeral numbers, or Roman numeral numbers. Although this step does not make a final prediction classification, it can provide auxiliary clues for subsequent analysis.
[0059] In the above step S2, for the target line text in the multi-line text, extract the first N lines of text and the last N lines of text of the target line text as the context information of the target line text; where N is a hyperparameter that can be adjusted according to the actual situation; when the above context information or the following context information of the target line text is less than N lines, padding is performed.
[0060] In the above step S3, concatenate the first N lines of text, the target line text, and the last N lines of text, and input the concatenation result into the fine-tuned large language model to output the category to which the target line text belongs;
[0061] In the embodiments of the present invention, taking the concatenation result of the target line text and its context information as the model input can enable the model to refer to the context relationship between the question stem, options, and answers during the semantic understanding process of the current line, thereby improving the classification accuracy.
[0062] In the embodiments of the present invention, the large language model uses glm4-9b-chat. This model is pre-trained on a large amount of text data and has strong natural language understanding and context reasoning capabilities. In the fine-tuning stage, instead of changing the basic structure of the model, it is adapted to this task through training with a small amount of task-related data. During the training process, to ensure the diversity of the dataset, it is necessary to collect test papers from different academic segments and different disciplines, annotate them, and fine-tune the model based on the annotated dataset so that it can predict the element categories for the test paper line text.
[0063] Specifically, as shown in Figure 2 the fine-tuning steps of the above large language model include:
[0064] (1) Determine the training dataset. The training samples in the training dataset are multiple lines of text information with real labels; the real label is the category of the line text information; among them, the real label is represented as (<question stem>, <options>, <answer>, <others>);
[0065] (2) For each line of text information, construct the input representation as: "[the first N lines of text][special separator][the current line of text][special separator][the last N lines of text]"; where, [special separator] is [SEP], which helps the model distinguish different parts; for each input current text line, its corresponding real label (<question stem>, <options>, <answer>, <others>) is retained as the supervision signal.
[0066] (3) Load the weights of the large language model and set the hyperparameters of the large language model; among them, the hyperparameters include learning rate, batch size, number of training epochs, truncation length (total context length), the value of N (number of context lines), etc.;
[0067] (4) Batch the training samples and input them into the large language model, and obtain the prediction distribution through forward propagation;
[0068] (5) Based on the prediction distribution, combined with the real label, obtain the cross-entropy loss;
[0069] (6) According to the cross-entropy loss, update the hyperparameters through the optimizer until convergence, and obtain the fine-tuned large language model.
[0070] After training is completed, an independent validation set or test set is evaluated. By comparing the predicted results of each line of text with the manually annotated results, metrics such as accuracy, precision, recall, and F1 value are calculated to test the classification ability of the model.
[0071] In practical applications, when it is necessary to process an unannotated new test paper text, the specific method is as follows:
[0072] (1) Process line by line: Starting from the first line of the test paper, take the current line, and extract the first N lines (fill with blanks if less than N lines) and the last N lines (fill with blanks if less than N lines) of text, and splice them together with the current line to form an input sequence;
[0073] (2) Model prediction: Feed the input sequence into the fine-tuned large language model, and the model outputs the category to which the current line of text belongs;
[0074] (3) Result recording: After storing the prediction result of the current line, continue to process the next line. Traverse all lines of text in the entire test paper in turn, and finally obtain the structured classification result of the entire test paper;
[0075] (4) Error correction: Simple rule checking and correction can be performed on the basis of the automatic classification result. For example, if several option lines follow the question stem line and the answer line follows immediately, the consistency and logic of the model prediction result can be verified through post-processing rules.
[0076] Next, a method for identifying the type of test paper structure elements based on context information fusion provided by the present invention is verified:
[0077] During the verification process, 1000 test papers containing various types of questions were collected from multiple sources, 800 of which were annotated as the training and validation data sets, and 200 were used as the test set. In the experiment:
[0078] 1. Comparison with the baseline method: Compared with the baseline model that does not consider context, the method of the present invention has a significant improvement in both accuracy and F1 value. The baseline model may only rely on the text features of the current line and ignore the context clues; while this method can more accurately reflect the structural relationship between the question stem, options, and answers by fusing context.
[0079] 2. Influence of the context window size: Different N values (such as N = 2, N = 5, N = 8) were tried in the experiment, and it was found that a moderate N value (such as N = 5) achieved a better balance between performance and computational efficiency. When N is too small, the model has insufficient context; when N is too large, although there is more information, it may cause input redundancy and increase the computational overhead.
[0080] 3. Generalizability test: When processing test papers of different grades and subjects, this method can still maintain high robustness and accuracy.
[0081] In summary, the present invention provides an efficient and accurate method for identifying the types of test paper structure elements based on context information fusion to achieve rapid electronic processing of the test paper structure. This method uses natural language processing (NLP) technology and large language models, especially the pre-trained model glm4-9b-chat, to identify and classify the questions, options, answers, and other elements in the test paper.
[0082] In another embodiment, a vision large model can be used. The test paper document picture is input into the vision large model, and the vision large model is allowed to identify the content of the test paper and process the redundant information. At the same time, through the way of prompt input, the element category of each line can be obtained.
[0083] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0084] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying test paper structure element types based on context information fusion, characterized in that: The steps include: Preprocessing the input test paper document to obtain multiple lines of text in the test paper document; For a target line of text in the multiple lines of text, extract the first N lines of text and the last N lines of text of the target line of text as context information of the target line of text; The first N lines of text, the target line of text and the last N lines of text are concatenated, and the concatenation result is input into the fine-tuned large language model, and the category to which the target line of text belongs is output.
2. According to the method for identifying the type of test paper structure elements based on context information fusion according to claim 1, it is characterized in that: The pre-processing specifically includes: Extracting a text line sequence from the test paper document; Remove spaces and line breaks from each line of text; Standardize the full-width and half-width punctuation marks in each line of text; Store the processed multiple lines of text in order as a list.
3. According to claim 2, a method for identifying test paper structure element types based on context information fusion is characterized in that: The extracting of the text line sequence from the test paper document specifically includes: recognizing the test paper document by OCR to obtain the text line sequence.
4. According to the method for identifying the type of test paper structure elements based on context information fusion according to claim 2, it is characterized in that: The preprocessing also includes: using regular expressions to automatically identify question number rows in the test paper document.
5. According to the method for identifying the type of test paper structure elements based on context information fusion according to claim 1, it is characterized in that: If the previous information or the following information of the target line of text is less than N lines, a blank filling process is performed.
6. The method for identifying the type of test paper structural elements based on context information fusion according to claim 1, characterized in that: The large language model adopts glm4-9b-chat.
7. The method for identifying the type of test paper structural elements based on context information fusion according to claim 1, characterized in that: The fine-tuning steps of the large language model include: Determine a training data set, wherein the training samples in the training data set are multiple lines of text information with real labels; the real labels are categories described in the line text information; For each line of text information, construct the input representation as: "[first N lines of text][special separator][current line of text][special separator][next N lines of text]"; Loading the weight of the large language model and setting hyperparameters of the large language model; Inputting the training samples into the large language model in batches, and obtaining the predicted distribution through forward propagation; Based on the predicted distribution and the true label, a cross entropy loss is obtained; According to the cross entropy loss, the hyperparameters are updated by an optimizer until convergence, thereby obtaining a fine-tuned large language model.
8. A method for identifying test paper structure element types based on context information fusion according to claim 7, characterized in that: The [special separator] is [SEP].
9. The method for identifying the type of test paper structural elements based on context information fusion according to claim 7, characterized in that: For each input current text line, its corresponding true label is retained as a supervisory signal.
10. The method for identifying the type of test paper structural elements based on context information fusion according to claim 7, characterized in that: The true label is expressed as: (<question>, <options>, <answer>, <others>).