Topic detection method, apparatus and device, and readable storage medium
By employing an adaptive correction method on the image of the answer content, the problem of question detection under different teacher annotation methods is solved, achieving high-precision question alignment and grading, and improving the robustness of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing question detection technologies cannot handle the different annotation methods teachers use for complex questions when creating test question courseware. This makes it difficult for question detection models to accurately identify and match question structures, affecting the accuracy of automatic question grading and error analysis.
By performing question detection on the answer content image from the dimensions of individual sub-questions and the entire question of a compound question, the question box containing both the question stem and the answer content is accurately identified. After mapping the standard answer image, adaptive correction is performed to ensure that the mapping area is consistent with the annotation granularity.
It achieves high-precision, adaptive recognition of user answers in the answer content image, improves the accuracy of question alignment and the robustness of the grading system, and ensures the accuracy and reliability of grading.
Smart Images

Figure CN121861683A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational technology, and in particular to a method, apparatus, device, and readable storage medium for question detection. Background Technology
[0002] In actual teaching applications, some test questions in the courseware are created by teachers themselves. For example, in the Bixin homework correction scenario, when teachers create test questions, they often break down complex questions in different ways according to their personal habits: some teachers treat the entire complex question as a whole and only mark it with one question box; while other teachers break it down into multiple sub-questions and mark each sub-question with a separate question box.
[0003] However, existing question detection and alignment technologies all assume that the question box annotations are uniform, that is, each question has a clear and consistent boundary definition. This cannot cope with the scenario where the same question has different annotation methods. Such inconsistency in annotation granularity makes it difficult for traditional question detection models to accurately identify and match question structures, thereby affecting the accuracy of automatic question grading and error analysis. Summary of the Invention
[0004] In view of this, in order to solve the above-mentioned technical problems, this application provides a question detection method, apparatus, device and readable storage medium.
[0005] Specifically, this application is implemented through the following technical solution: According to a first aspect of the embodiments of this application, a question detection method is provided, the method comprising: Detect composite questions in the answer content image to obtain sub-question boxes corresponding to each question in the composite question and the composite question title box corresponding to the composite question; the sub-question box corresponding to any question question includes at least the question question and the answer content for the question question; Map each marked question box in the standard answer image of the compound question to the answer content image to obtain the mapping area in the answer content image corresponding to each marked question box; For each mapped region on the answer content image, find the question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. Based on the question box that matches the mapping area, the mapping area is adjusted so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein, the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
[0006] Optionally, the mapping area can be adjusted based on the question box that matches it, including: If there is only one question box that matches the mapped area, then that question box will be used as the target question box. If there are multiple question boxes that match the mapped area, the question box with the smallest area among the matched question boxes will be taken as the target question box. The boundaries of the mapping region are adjusted based on the boundaries of the target title box.
[0007] Optionally, adjusting the boundary of the mapped region includes: In the horizontal direction of the answer content image, if the left boundary of the target question box is located to the left of the left boundary of the mapping area, then the left boundary of the mapping area is adjusted to the left boundary of the target question box; and if the right boundary of the target question box is located to the right of the right boundary of the mapping area, then the right boundary of the mapping area is adjusted to the right boundary of the target question box. In the vertical direction of the answer content image, if the upper boundary of the target question frame is located above the upper boundary of the mapping area, then the upper boundary of the mapping area is adjusted to be the upper boundary of the target question frame; and if the lower boundary of the target question frame is located below the lower boundary of the mapping area, then the lower boundary of the mapping area is adjusted to be the lower boundary of the target question frame.
[0008] Optionally, the mapping area can be adjusted based on the question box that matches it, including: Based on the question box that matches the mapped area, the mapped area is adjusted to obtain the first adjusted area corresponding to the mapped area; If the question in the first adjustment area is not a question-solving type, for each handwritten text box in the answer content image, it is detected whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box. If so, then based on the boundary of the handwritten text box, the boundary of the first adjustment area within the range of the handwritten text box is adjusted to obtain the adjusted mapping area.
[0009] Optionally, the boundaries of the first adjustment area within the handwritten text box are adjusted, including: If the upper boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the upper boundary of the handwritten text box. If the lower boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the lower boundary of the handwritten text box.
[0010] Optionally, the method further includes: If the question in the first adjustment area is a question of the answer type, detect whether there is a printed text box located below the first adjustment area in the vertical direction of the answer content image from each printed text box included in the answer content image. If it exists, then based on the printed text boxes located below the first adjustment area, the lower boundary of the first adjustment area in the vertical direction is adjusted to obtain the adjusted mapping area.
[0011] Optionally, adjusting the lower boundary of the first adjustment area in the vertical direction includes: If there is a text box for printed text located below the first adjustment area, then the lower boundary of the first adjustment area in the vertical direction is adjusted to the upper boundary of the text box for printed text. If multiple printed text boxes are located below the first adjustment area, the lower boundary is adjusted to the upper boundary of each of the multiple printed text boxes that is closest to the first adjustment area.
[0012] Optionally, after adjusting the lower boundary of the first adjustment region in the vertical direction to obtain the second adjustment region, the method further includes: For each handwritten line text box on the image of the answer content, it is detected whether the lower boundary of the second adjustment area in the vertical direction is within the range of the handwritten line text box; If so, the lower boundary of the mapped area is further adjusted to the lower boundary of the handwritten text box to obtain the adjusted mapped area.
[0013] According to a second aspect of the embodiments of this application, a question detection device is provided, the device comprising: The question frame detection module is configured to detect composite questions in the answer content image to obtain sub-question frames corresponding to each question in the composite question and the composite question frame corresponding to the composite question; the sub-question frame corresponding to any question includes at least the question and the answer content for the question; The question box mapping module is configured to map each marked question box in the standard answer image of the composite question to the answer content image, so as to obtain the mapping area in the answer content image corresponding to each marked question box; The question box matching module is configured to, for each mapped region on the answer content image, search for a question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. The adjustment module is configured to adjust the mapping area based on the question box that matches the mapping area, so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein, the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
[0014] Optionally, when the adjustment module is configured to adjust the mapped region based on the question box that matches the mapped region, it includes: The target question frame determination module is configured such that if there is only one question frame that matches the mapped area, then that question frame is taken as the target question frame; if there are multiple question frames that match the mapped area, then the question frame with the smallest area among the matched question frames is taken as the target question frame. The boundary adjustment module is configured to adjust the boundary of the mapped region based on the boundary of the target title box.
[0015] Optionally, the boundary adjustment module is configured as follows: In the horizontal direction of the answer content image, if the left boundary of the target question box is located to the left of the left boundary of the mapping area, then the left boundary of the mapping area is adjusted to the left boundary of the target question box; and if the right boundary of the target question box is located to the right of the right boundary of the mapping area, then the right boundary of the mapping area is adjusted to the right boundary of the target question box. In the vertical direction of the answer content image, if the upper boundary of the target question frame is located above the upper boundary of the mapping area, then the upper boundary of the mapping area is adjusted to be the upper boundary of the target question frame; and if the lower boundary of the target question frame is located below the lower boundary of the mapping area, then the lower boundary of the mapping area is adjusted to be the lower boundary of the target question frame.
[0016] Optionally, when the adjustment module is configured to adjust the mapped region based on the question box that matches the mapped region, it includes: The first adjustment module is configured to adjust the mapping region based on the question box that matches the mapping region, so as to obtain the first adjustment region corresponding to the mapping region; The second adjustment module is configured to, when the question in the first adjustment area is not a question-solving type, detect whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box for each handwritten text box; if so, adjust the boundary of the first adjustment area within the range of the handwritten text box according to the boundary of the handwritten text box to obtain the adjusted mapping area.
[0017] Optionally, when the second adjustment module is configured to adjust the boundary of the handwritten text box within the first adjustment area, it includes: If the upper boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the upper boundary of the handwritten text box. If the lower boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the lower boundary of the handwritten text box.
[0018] Optionally, when the adjustment module is configured to adjust the mapped region based on the question box that matches the mapped region, it also includes: The third adjustment module is configured to, when the question in the first adjustment area is an answer question, detect whether there is a printed text box located below the first adjustment area in the vertical direction of the answer content image from each printed text box included in the answer content image; if so, adjust the lower boundary of the first adjustment area in the vertical direction based on each printed text box located below the first adjustment area to obtain the adjusted mapping area.
[0019] Optionally, when the third adjustment module is configured to adjust the lower boundary of the first adjustment region in the vertical direction, it includes: If there is a text box for printed text located below the first adjustment area, then the lower boundary of the first adjustment area in the vertical direction is adjusted to the upper boundary of the text box for printed text. If multiple printed text boxes are located below the first adjustment area, the lower boundary is adjusted to the upper boundary of each of the multiple printed text boxes that is closest to the first adjustment area.
[0020] Optionally, after adjusting the lower boundary of the first adjustment area in the vertical direction to obtain the second adjustment area, the device further includes: The fourth adjustment module is configured to detect whether the lower boundary of the second adjustment area in the vertical direction is within the range of the handwritten text box for each handwritten text box on the answer content image; if so, the lower boundary of the mapping area is further adjusted to the lower boundary of the handwritten text box to obtain the adjusted mapping area.
[0021] According to a third aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a memory and a processor; the memory being used to store a computer program; the processor being used to execute the above-described question detection method by invoking the computer program.
[0022] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, wherein the program, when executed by a processor, implements the above-described question detection method.
[0023] The technical solutions provided in this application embodiment may include the following beneficial effects: In the technical solution provided in this application, by performing question detection on the answer content image from the dimensions of individual sub-questions and the entire question of a compound question, the question boxes containing the question stem and the answer content for that question stem are accurately identified. After mapping each marked question box pre-annotated by the teacher on the standard answer image to the answer content image, the mapped area on the answer content image is further adaptively corrected using the question boxes detected on the answer content image. Thus, while maintaining the same annotation granularity as the marked question boxes on the standard answer image, the mapped area on the answer content image contains both the question stem and the user's answer content for that question stem. This ensures high-precision and adaptive recognition of the user's answer content in the answer content image while maintaining consistent annotation granularity, thereby improving the question alignment accuracy and the robustness of the grading system.
[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Furthermore, no embodiment in this application needs to achieve all the effects described above. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0026] Figure 1A This is a schematic diagram of a question detection method illustrated in an exemplary embodiment of this application; Figure 1B This is a schematic diagram illustrating an image of a response content in an exemplary embodiment of this application; Figure 1C This is a schematic diagram of a question frame after question frame detection and annotation on a composite question on an image of the answer content, as shown in an exemplary embodiment of this application; Figure 1D This is a schematic diagram of a labeled question box on a standard answer image of a compound question, as illustrated in an exemplary embodiment of this application; Figure 1E This is a schematic diagram illustrating an exemplary embodiment of the present application, showing how each marked question box on a standard answer image is mapped onto the answer content image to obtain a mapped area; Figure 1FThis is a schematic diagram illustrating a comparison of a sub-question frame obtained by question detection and a composite question frame on the same answer content image, as well as the mapped area obtained by mapping, according to an exemplary embodiment of this application. Figure 1G This is a schematic diagram of an adjusted mapping region B2' on a response content image, as shown in an exemplary embodiment of this application. Figure 2A This is a schematic diagram illustrating an exemplary embodiment of the present application of a process for adjusting a mapping area containing questions of the non-answerable question type; Figure 2B This is a comparative schematic diagram showing, in an exemplary embodiment of this application, a question box, a mapping area, a first adjustment area, and a handwritten text box detected on an image of a response content for a compound question; Figure 3A This is a schematic diagram illustrating an exemplary embodiment of the present application of a process for adjusting a mapping area containing questions of the problem-solving type; Figure 3B This is a flowchart illustrating another step in adjusting the mapping area of a compound question containing answer-type questions, as shown in an exemplary embodiment of this application. Figure 4 This is an overall flowchart illustrating an exemplary embodiment of the question detection method of this application; Figure 5 This is a schematic diagram of the structure of a question detection device according to an exemplary embodiment of this application; Figure 6 This is a hardware schematic diagram of an electronic device illustrated in an exemplary embodiment of this application. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another.
[0028] In actual teaching applications, some test questions in the courseware are created by teachers themselves. For example, in the Bixin homework correction scenario, when teachers create test questions, they often break down complex questions containing multiple sub-questions in different ways according to their personal habits: some teachers treat the entire complex question as a whole and only mark it with one question box; while other teachers break it down into multiple sub-questions and mark each sub-question with a separate question box.
[0029] However, current mainstream question detection models are typically trained based on uniform annotation standards (e.g., forcing each independent question to be treated as a single question unit), resulting in a fixed way of dividing the output question boxes. When this model is applied to the answer content image, the detected question box structure often fails to maintain consistency with the annotation granularity on the standard answer image. This mismatch in question box granularity directly leads to the failure of subsequent question alignment and answer comparison: for example, the model detects three sub-question boxes in the answer content image, but there is only one overall question box in the standard answer image, making it impossible to establish a one-to-one correspondence. This results in missed judgments, misjudgments, or confusion in the scoring logic, affecting the accuracy and reliability of automatic question grading.
[0030] Related technologies attempt to align questions by mapping the question boxes marked by teachers on the answer images to the answer content images. However, since the answer content images themselves do not contain the user's answers, the question boxes marked by teachers usually only cover the question stem area. For question types such as calculation problems and problem-solving questions that require the answer area to be included, the user's answers often extend below the question stem or even span across lines. The directly mapped question boxes cannot completely cover the user's answer area, thus affecting the accuracy of subsequent automatic grading and error analysis.
[0031] In view of this, this application provides a question detection method. This method no longer relies on a single answer image question box mapping strategy, but combines the question structure of the answer content image and the standard answer image. First, it performs question detection on the answer content image from the dimensions of individual sub-questions of compound questions and the entire question dimension of compound questions to accurately identify question boxes containing the question stem and the answer content for that question stem. After mapping each question box pre-annotated by the teacher on the standard answer image to the answer content image, the mapped area on the answer content image has the same annotation granularity as the annotation question boxes on the standard answer image. On this basis, the detected question boxes on the answer content image are further used to adaptively correct the mapped area on the answer content image, so that the mapped area on the answer content image contains the question stem and the user's answer content for that question stem. This achieves that the mapped area fits the real answer context. While ensuring consistent annotation granularity, it achieves high-precision and adaptive recognition of user's answer content in the answer content image, improving the question alignment accuracy and the robustness of the grading system.
[0032] See Figure 1AThe flowchart illustrating an exemplary question detection method shows that the question detection method provided in this application may include at least the following steps: S101, detect the composite question in the answer content image to obtain the sub-question box corresponding to each question in the composite question and the composite question box corresponding to the composite question; the sub-question box corresponding to any question question includes at least the question question and the answer content for the question question; The answer content image is an image uploaded by the user after completing their answers on an answer sheet / answer panel or other answering tool. It includes the complete question content and handwritten answer information. For example, see Figure 1B The illustration shows a schematic diagram of an answer content image, which represents an answer page of a math exam paper, with multiple consecutively arranged questions and their corresponding handwritten answer areas. The handwritten answer areas contain the answer content entered by the user after completing the answer.
[0033] A question box represents a structured bounding box that encloses the entire text content of a complete question (including the question number, question stem, and user answer area). Its location and extent can be obtained through a question detection model or a rule-based layout analysis method. In this embodiment, the question box can be automatically identified through question-level semantic segmentation, clustering and merging based on question number keywords, or by using a pre-trained deep learning detection network.
[0034] A compound question is a question composed of multiple related sub-questions. These sub-questions revolve around a common theme or scenario, are logically connected, and share some information from the question stem. Each sub-question in a compound question constitutes a single question. For example, see... Figure 1B The compound question Q shown is: "Xiaoming's family has a rectangular vegetable garden, which is 18 meters long and 12 meters wide."
[0035] (1) If a fence is to be built around this vegetable garden, how long will the fence need to be at least? (2) If 3.5 kg of vegetables can be harvested per square meter, how many kilograms of vegetables can be harvested from this plot of land in total? The composite question consists of two sub-questions, question (1) and question (2). When performing question frame detection on the composite question, the sub-question frame is obtained for each question included in the composite question, that is, the sub-question frame is determined for each question (1) and question (2). The scope of any sub-question frame includes a question and the user's answer to the question. In addition, the composite question frame is obtained for the entire composite question, that is, the composite question is marked as a question frame as a whole. The scope of the composite question frame includes the question (1) and question (2) of the composite question and the user's answer to each question.
[0036] For example, see Figure 1C This example illustrates a schematic diagram of a question frame after detecting and annotating the question frame on a composite question in an image of the answer content. Figure 1B Taking the answer content image shown as an example, the answer content image includes the above-mentioned composite question Q. The question frame of the composite question includes two types: one is the composite question frame that labels the entire composite as a single question, as shown in the figure, the whole question frame A0, which covers the question stem, question (1) and (2) and their answer content; the other is the sub-question frame that is based on each sub-question (i.e., question) included in the composite question, as shown in the figure, the sub-question frame A2 corresponding to question (1), which contains question (1) and the answer content, and the sub-question frame A3 corresponding to question (2), which contains question (2) and the answer content. Since the questions in the composite question share the same question stem, in order to ensure that the question stem content is completely preserved and accurately associated in structural analysis and subsequent mapping, a separate question frame can be labeled for the shared question stem, as shown in the figure, question frame A1.
[0037] In this embodiment, the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question can be implemented using a pre-trained question detection model. By inputting the answer content image into the question detection model, the position information of the sub-question boxes and the composite question box output by the question detection model is obtained. For example, the YOLO model can be used as the question detection model, and other models can also be used in practical applications.
[0038] S102, map each marked question box in the standard answer image of the compound question to the answer content image to obtain the mapping area in the answer content image corresponding to each marked question box; The standard answer image is a reference image that includes the correct answer and the question structure, featuring clear text layout, complete solution steps, and well-defined question boundaries. Each marked question box on the standard answer image is a rectangle pre-marked by the user (such as a teacher) to indicate the question and / or corresponding answer area, and its position and extent reflect the logical organization of the questions in the standard answer.
[0039] For the same complex question, the annotation box may employ various annotation strategies depending on the teacher's annotation habits, layout style, or scoring requirements. For example, see... Figure 1DThe example shows a schematic diagram of the labeled question boxes on the standard answer image of a compound question. In the left figure, the entire compound question is labeled as a whole, and only one labeled question box B0 covering all sub-questions and answers is output. In the right figure, each question included in the compound question is labeled in detail, and multiple sub-question boxes are output respectively: B1 corresponds to the common question stem, B2 corresponds to sub-question (1), and B3 corresponds to sub-question (2).
[0040] When mapping each labeled question box in the standard answer image of the composite question to the answer content image, spatial alignment across images can be achieved through methods such as image registration, optical flow estimation, or alignment networks built based on deep learning. Based on the geometric transformation relationships such as translation, scaling, and rotation between the answer content image and the standard answer image, combined with keypoint matching or semantic consistency constraints, each labeled question box in the standard answer image is projected onto its corresponding position in the answer content image, thereby obtaining the mapped region in the answer content image corresponding to each labeled question box. In addition, the answer content image may have writing offsets, paper deformation, etc., and non-rigid deformation correction methods can be combined during the mapping process to improve mapping accuracy and robustness.
[0041] For example, see Figure 1E This example illustrates a schematic diagram of mapping each labeled question box on a standard answer image to a mapped area on the answer content image. The labeled question boxes B1, B2, and B3 of the compound questions in the standard answer image are mapped to corresponding areas on the answer content image, forming mapped areas B1', B2', and B3'. Each mapped area maintains the same spatial position as the original labeled box, and its boundaries are determined by the labeling information of the standard answer image and the mapping operation between the two images, ensuring a structural correspondence between the mapped area and the original labeled question box. In the example given, when the labeled question boxes on the standard answer image are mapped to the answer content image, factors such as inconsistent image resolution and scaling deviations may cause the mapped area on the answer content image to shift from the expected position of the original labeled question box, failing to completely cover the question. This embodiment effectively solves this problem through subsequent operations.
[0042] This mapping operation transfers the structured annotations in the standard answer image to the answer content image, ensuring that the mapped area on the answer content image is consistent with the question boxes pre-annotated by the teacher in the standard answer image in terms of spatial layout and logical granularity. This effectively avoids misjudgment or omissions caused by mismatched annotation granularity.
[0043] S103, for each mapped region on the answer content image, find the question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. The question box that matches the mapping region represents the question box that is closest to the mapping region in position and size among the sub-question boxes and the question box of the composite question detected on the answer content image. Alternatively, it can be understood as the question box and the mapping region having the greatest overlap or the closest alignment in space.
[0044] In this embodiment, the degree of matching between two regions on the same image can be measured by calculating the IoU (Intersection over Union). When searching for question boxes that match the mapped region, the matching question boxes can be determined based on the ratio of the intersection area to the union area between the sub-question boxes, composite question boxes, and the mapped region. Specifically, for each mapped region on the answer content image, the intersection area between each sub-question box and composite question box included in the answer content image and the mapped region can be determined first. Further, question boxes whose IoU with respect to the mapped region is greater than a set threshold are considered as question boxes that match the mapped region. Here, the intersection area represents the rectangular or irregular area formed by the overlapping part of two regions on the image plane, and the area of the intersection area can be used to measure the actual degree of overlap between the two regions.
[0045] For example, see Figure 1F The example illustration shows a comparison of sub-question frames obtained through question detection, composite question frames, and mapped regions on the same answer content image. For each mapped region (e.g., B1', B2', B3'), a question frame matching the mapped region is searched from the question frames A0, A1, A2, A3 obtained from question detection of the composite question. Taking mapped region B2' as an example, the IoU value between B2' and each question frame (A0, A1, A2, A3) is calculated, and question frames with IoU values greater than a preset threshold are selected as the question frames matching B2'. For example, if the IoU of the intersection of the detected composite question frame A0, sub-question frame A2, and mapped region B2' with B2' is 100% and 95% respectively, both greater than the set threshold of 75%, then question frames A2 and A0 are determined to be the question frames matching mapped region B2'.
[0046] In the process of finding matching question boxes based on the above-mentioned mapping region, some question boxes detected and identified on the answer content image may not be covered or successfully mapped by the standard answer image, for example, due to image deformation causing mapping failure. Therefore, this embodiment also provides a method for missing question boxes. After completing the operation of finding question boxes that match the mapping region, the following operation can be performed on each question box (i.e., sub-question boxes and composite question boxes) corresponding to the compound question on the answer content image, thereby improving the recall rate and robustness of question region recognition. This operation includes at least the following steps: Determine the IoU of the intersection of the question box with each mapped region relative to the specified region. The specified region is the smaller area among the question box and the mapped regions in the current intersection calculation. If the IoU of the intersection of the question box with all mapped regions relative to the specified region is less than a set threshold, it means that the question box is not effectively covered by any mapped region, and the question box can be used as a mapped region on the answer content image.
[0047] S104, Based on the question box that matches the mapping area, adjust the mapping area so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein, the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
[0048] The question boxes detected in the answer content image include the actual writing layout and the user's answer content. However, the mapped area of the marked question boxes on the standard answer image onto the answer content image may be offset or incomplete due to image distortion, registration errors, or differences in teacher annotation granularity. Therefore, this embodiment uses the question boxes detected in the answer content image as preliminary experience. For each mapped area on the answer content image, the question box that matches the mapped area is used as the boundary optimization basis for the mapped area. This ensures that the optimized mapped area can effectively cover the question and the answer content for the question, avoiding the truncation of key question information (such as the question and handwritten answer) due to the mapped area being too small or offset, thereby improving the accuracy of question grading.
[0049] Based on the fact that the question boxes corresponding to compound questions include the compound question box at the whole question level (e.g., A0) and the sub-question boxes at the individual sub-question level (e.g., A1, A2, A3), the compound questions on the standard answer image are therefore broken down into multiple sub-questions and their question boxes are labeled separately (e.g., ...). Figure 1DAs shown in the right-hand annotation diagram (containing question boxes B1, B2, B3, etc.), theoretically, the mapped area (e.g., B2') of any question box (e.g., B2) on the answer content image can match the entire-question-level composite question box (e.g., A0) detected on the answer content image, and also match the sub-question box (e.g., A2) detected on the answer content image that points to the same sub-question as the mapped area. In this case, directly adjusting the mapped area using the composite question box would result in an excessively large adjusted area, incorporating content from other sub-questions and affecting the scoring granularity. Therefore, adjustments to the mapped area should prioritize sub-question boxes with consistent semantic granularity.
[0050] Based on this, when adjusting the mapping region according to the question boxes that match the mapping region, if there is only one question box matching the mapping region, then that question box is taken as the target question box; if there are multiple question boxes matching the mapping region, then the question box with the smallest area among the matching question boxes is taken as the target question box, so as to accurately reflect the question and its answer range corresponding to the current mapping region. Furthermore, the boundary of the mapping region is adjusted according to the boundary of the target question box.
[0051] by Figure 1D In the example shown, the composite question on the standard answer image is treated as a whole question and labeled with a question box (as shown in the left image, which includes the labeled question box B0). Assuming that the labeled question box is mapped to the corresponding mapping area B0' in the answer content image, the composite question box A0 detected on the user's answer image matches the mapping area B0'. At this time, the composite question box A0 is taken as the target question box, and the mapping area B0' on the answer content image is adjusted according to the area boundary of the composite question box A0 on the answer content image.
[0052] by Figure 1D The standard answer image shown illustrates how a complex question is broken down into multiple sub-questions, each with its own labeled question box (e.g., the right-hand image includes labeled question boxes B1, B2, and B3). Taking question box B2 as an example... Figure 1F As shown, it is mapped onto the answer content image to obtain the mapping region B2'. The composite question frame A0 and the sub-question frame A2 detected on the user's answer image are matched with the mapping region B2' on the answer content image. The smaller sub-question frame A2 is taken as the target question frame. The mapping region B2' on the answer content image is adjusted using the area boundary of the sub-question frame A2 on the answer content image.
[0053] In this implementation, adjusting the boundary of the mapping area based on the boundary of the target question frame essentially expands or corrects the boundary of the mapping area to effectively include all content covered by the target question frame. This avoids the mapping area being affected by factors such as registration errors, paper wrinkles, or writing offsets, which could prevent it from effectively covering the question and its answer. Therefore, when adjusting the boundary of the mapping area, an outer bounding strategy is used, taking the outermost boundary between the target question frame and the mapping area in each direction.
[0054] Therefore, when adjusting the boundary of the mapping region based on the boundary of the target title box, it can be achieved in the following way: In the horizontal direction of the answer content image, if the left boundary of the target question box is located to the left of the left boundary of the mapping area, then the left boundary of the mapping area is adjusted to the left boundary of the target question box; and if the right boundary of the target question box is located to the right of the right boundary of the mapping area, then the right boundary of the mapping area is adjusted to the right boundary of the target question box. In the vertical direction of the answer content image, if the upper boundary of the target question frame is located above the upper boundary of the mapping area, then the upper boundary of the mapping area is adjusted to be the upper boundary of the target question frame; and if the lower boundary of the target question frame is located below the lower boundary of the mapping area, then the lower boundary of the mapping area is adjusted to be the lower boundary of the target question frame.
[0055] For example, with Figure 1F Taking the sub-question box A2 on the answer content image as the target question box, and adjusting the mapped region B2' on the answer content image as an example, assuming that the left boundary of the sub-question box A2 in the horizontal direction is Bound1_x1 and the right boundary is Bound1_x2, and the upper boundary in the vertical direction is Bound1_y1 and the lower boundary is Bound1_y2, correspondingly, assuming that the left boundary of the mapped region B2' in the horizontal direction is Bound2_x1 and the right boundary is Bound2_x2, and the upper boundary in the vertical direction is Bound2_y1 and the lower boundary is Bound2_y2, then: If the left boundary Bound1_x1 of sub-topic box A2 is located to the left of the left boundary Bound2_x1 of the mapping region B2', then the left boundary of the mapping region B2' will be adjusted to Bound1_x1; otherwise, the left boundary Bound2_x1 of the mapping region B2' will remain unchanged. If the right boundary Bound1_x2 of sub-topic box A2 is located to the right of the left boundary Bound2_x2 of the mapping region B2', then the right boundary of the mapping region B2' will be adjusted to Bound1_x2; otherwise, the left boundary Bound2_x2 of the mapping region B2' will remain unchanged. If the upper boundary Bound1_y1 of sub-topic box A2 is located above the upper boundary Bound2_y1 of the mapping region B2', then the upper boundary of the mapping region B2' will be adjusted to Bound1_y1; otherwise, the upper boundary Bound2_y1 of the mapping region B2' will remain unchanged. If the lower boundary Bound1_y2 of sub-topic box A2 is located below the lower boundary Bound2_y2 of the mapped region B2', then the lower boundary of the mapped region B2' will be adjusted to Bound1_y2; otherwise, the lower boundary Bound2_y2 of the mapped region B2' will remain unchanged.
[0056] Using sub-question box A2 on the answer content image as the target question box, adjust the mapped area B2' on the answer content image. See [link / reference]. Figure 1G An exemplary diagram of the adjusted mapping region B2' on the answer content image is shown. Compared to the case where the mapping region B2' before adjustment only included the question (1) and did not completely cover the question (1), the adjusted mapping region B2'' not only completely covers the question (1) but also includes the user's answer to the question (1). Similarly, by using the sub-question box A1 on the answer content image as the target question box, the mapping region A2' on the answer content image is adjusted to obtain the corresponding mapping region B1'', and by using the sub-question box A3 on the answer content image as the target question box, the mapping region A3' on the answer content image is adjusted to obtain the corresponding mapping region B3''.
[0057] Through the above mapping area adjustment process, the adjusted mapping area retains the logical structure of the question box annotation in the standard answer image maintained by the mapping operation, and also integrates the spatial information of the question-answer content layout in the answer content image based on the boundary adjustment of the target question box, so as to provide a reliable, complete and semantically consistent question area for automatic grading.
[0058] In this embodiment, question box detection is performed simultaneously on the answer content image at both the individual sub-question level and the overall composite question level. After mapping each annotated question box in the standard answer image to the answer content image, the mapped region is adaptively corrected using the question boxes detected on the answer content image. By combining the question structure information of the answer content image and the standard answer image, the annotation granularity of the mapped region on the answer content image is kept consistent with that of the annotated question boxes on the standard answer image. Furthermore, the mapped region on the answer content image includes the question stem and the user's answer to the question stem, making the mapped region fit the actual answer context. This achieves high-precision and adaptive recognition of the user's answer in the answer content image, effectively addressing the situation where different teachers use different annotation granularities for the same composite question, and improving the accuracy of question alignment and automatic grading.
[0059] In some embodiments, users may modify or supplement their answers, causing the original answer to be crossed out and a new answer to be handwritten above, below, or in other adjacent areas. This results in the user's answer exceeding the original pre-defined answer area of the question. When obtaining the sub-question boxes and the composite question box on the answer image through the aforementioned steps, the detection is often based only on the question stem and the initial answer layout, failing to cover the handwritten content added due to the modification. As a result, even after adjusting the mapping area based on the question box that matches the mapping area, the adjusted mapping area still fails to completely cover all of the user's valid answer content.
[0060] To address this type of situation, this embodiment proposes that, during the process of adjusting the mapping region based on the question box that matches the mapping region, after adjusting the mapping region based on the question box that matches the mapping region to obtain the first adjustment region corresponding to the mapping region, the boundary information of the handwritten text box included on the answer content image is further introduced to further adjust the first adjustment region, so that all the answers actually written by the user can be included in the final output adjusted mapping region.
[0061] This embodiment addresses the aforementioned overflow issues in both open-ended and non-open-ended compound questions. Since the answer area for non-open-ended compound questions (such as fill-in-the-blank and multiple-choice questions) is typically concentrated in a fixed location with limited space, user modifications can easily cause new answers to extend beyond the original question frame. Therefore, for non-open-ended compound questions, it is necessary to detect adjacent handwritten text boxes in the vertical direction and expand the mapping area boundary accordingly. In contrast, open-ended compound questions generally allow for a larger answer space, but users may still continue writing derivations or supplementary explanations outside the answer area of the current sub-question and around the stem of the next sub-question. Therefore, for open-ended compound questions, adjustments need to be made by combining printed text and the user's handwritten text.
[0062] Therefore, see Figure 2A The exemplary diagram illustrates a process for adjusting a mapping area containing non-solution questions. This adjustment process can be implemented through the following steps: S201, if the question in the first adjustment area is not a question type, for each handwritten text box in the answer content image, detect whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box. Printed text refers to standard font text displayed by printing equipment (such as printers and copiers) or electronic screens, such as Song, Hei, and Times New Roman, characterized by clarity, neatness, no ligatures, and uniform character shapes. In contrast, handwritten text, or text written by hand, typically has irregular shapes, diverse styles, and is difficult to recognize. In scenarios such as learning machines, homework correction, and test paper analysis, the questions, stems, and options are mostly printed, while the user's answers are mostly handwritten.
[0063] For the answer content image, a deep learning-based text detection model can be used to detect handwritten line text, identifying the handwritten text region at each line level in the answer content image and representing it in a structured form as line text boxes, thus obtaining each handwritten line text box included in the answer content image. Each handwritten line text box corresponds to a complete line of handwritten text content, and the positional information of the handwritten line text box reflects the spatial coordinate range of that line of handwritten text in the answer content image.
[0064] For any handwritten text box, when detecting whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box, it can be determined in the following way: For any of the upper and lower boundaries of the first adjustment area in the vertical direction of the answer content image, if the boundary of the first adjustment area is located within the interval defined by the upper and lower boundaries of the handwritten text box, that is, the boundary of the first adjustment area is located below the upper boundary of the handwritten text box and above the lower boundary of the handwritten text box, then the boundary of the first adjustment area is determined to be within the range of the handwritten text box.
[0065] See Figure 2B This example illustrates a comparison diagram of a question box, a mapping area, a first adjustment area, and a handwritten text box detected on an image of a response content for a compound question. The response content image includes handwritten text boxes S1, S2, S3, S4, and S5. A question box marked on the standard answer image is mapped to a corresponding mapping area D2' on the response content image. This mapping area is adjusted according to the boundary of the sub-question box C2 detected on the response content image to obtain the first adjustment area D2''. Taking the first adjustment area D2'' as an example, the upper and lower boundaries of the first adjustment area D2'' are detected respectively to determine whether it is located below the upper boundary of any one of the handwritten text boxes S1, S2, S3, S4, and S5, and is located below the upper boundary of that text box. As shown in the figure, assuming that the lower boundary of the first adjustment area D2'' is below the upper boundary of the handwritten text box S3 and above the lower boundary of S3, then it is determined that the lower boundary of the first adjustment area D2'' in the vertical direction of the answer content image is within the range of the handwritten text box S3.
[0066] S202, if so, then based on the boundary of the handwritten text box, adjust the boundary of the first adjustment area within the range of the handwritten text box to obtain the adjusted mapping area.
[0067] This step adjusts the boundary of the first adjustment area within the handwritten text box. The purpose is to ensure that the first adjustment area includes any user responses it does not cover. Therefore, when adjusting the boundary of the first adjustment area within the handwritten text box, the vertical boundary of the first adjustment area needs to be expanded outwards, that is: If the upper boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the upper boundary of the handwritten text box; if the lower boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the lower boundary of the handwritten text box.
[0068] See Figure 2BAs shown, when it is determined that the lower boundary of the first adjustment area D2'' in the vertical direction of the answer content image is within the range of the handwritten text box S3, the lower boundary of the first adjustment area D2'' needs to be adjusted to the upper boundary of the handwritten text box S3. This ensures that the adjusted first adjustment area D2''' includes the answer content within the handwritten text box S3, thus solving the problem of answer overflow and the inability of the mapped area to cover the answer in non-answer question types.
[0069] This embodiment iterates through each handwritten text box in the answer content image. If it is detected that neither of the two vertical boundaries of the first adjustment region in the answer content image is within the range of the handwritten text box, the process continues to traverse the next handwritten text box in the answer content image. After determining that any boundary (such as the upper boundary) in the first adjustment region is located in a handwritten text box, the subsequent detection process can only detect whether the other boundary (such as the lower boundary) is within the subsequently traversed handwritten text box, without needing to detect both boundaries of the first adjustment region again.
[0070] In the above embodiments, by introducing boundary detection and dynamic expansion of handwritten text boxes during the adjustment of the mapping area for non-answer questions, the actual user's answer content is fully captured in cases of unconventional answering behaviors such as correction, supplementation, or offset writing. This ensures that the final output adjusted mapping area effectively covers the spatial distribution of the actual answer, avoiding grading errors caused by truncation or omission of the answer content, thereby improving the accuracy of question grading and user experience in complex handwriting scenarios.
[0071] In some embodiments, for compound questions of the problem-solving type, see [link to relevant documentation]. Figure 3A The exemplary diagram illustrates a process for adjusting the mapping area of questions containing problem-solving questions. This adjustment process can be achieved through the following steps: S301, if the question in the first adjustment area is a question of the answer type, detect from each printed text box in the answer content image whether there is a printed text box located below the first adjustment area in the vertical direction of the answer content image. The printed text boxes in the answer content image can be obtained by performing printed text detection on the answer content using a deep learning-based text detection model. The text detection model identifies the printed text region at the line level in the answer content image and represents it in a structured form as line text boxes, resulting in each printed text box in the answer content image. Each printed text box corresponds to a complete line of printed text content. The positional information of the printed text box reflects the spatial coordinate range of that line of printed text in the answer content image, for example, represented as a rectangular bounding box, including the coordinates of the top-left and bottom-right corners (x1, y1, x2, y2) or the center point, width, height (cx, cy, w, h), and other geometric parameters.
[0072] When detecting the printed text box located vertically below the first adjustment area in the image of the answer content, it can be determined as follows: For each printed text box, it is checked whether the printed text box simultaneously meets the following two conditions: In the vertical direction of the answer content image, the upper boundary of the printed text box is below the lower boundary of the first adjustment area; in the horizontal direction of the answer content image, the horizontal interval defined by the left and right boundaries of the printed text box intersects with the horizontal interval defined by the left and right boundaries of the first adjustment area. If both conditions are met, the printed text box is designated as the printed text box located below the first adjustment area. By limiting the overlap or partial alignment of the printed text box and the first adjustment area in the horizontal direction, the printed text box located below the first adjustment area is designed to be the next question immediately following the current question, such as the stem or options of the next sub-question, thereby avoiding excessive expansion during area adjustment and the introduction of irrelevant content.
[0073] For example, suppose an image of the answer content includes printed text boxes P1, P2, P3, P4, and P5. A question box on the standard answer image is mapped to a corresponding mapped region D3' on the answer content image. This mapped region is adjusted according to the boundary of the detected sub-question box C3 on the answer content image to obtain a first adjusted region D3''. Taking this first adjusted region D3'' as an example, we detect whether there is a printed text box located below the first adjusted region D3'' among the printed text boxes P1, P2, P3, P4, and P5. Assuming that the upper boundary of the printed text box P5 is below the lower boundary of the first adjusted region D3'', and the horizontal interval defined by the left and right boundaries of the printed text box P5 overlaps with the horizontal interval defined by the left and right boundaries of the first adjusted region D3'', then the printed text box P5 is considered to be a printed text box located below the first adjusted region D3''.
[0074] S302, if present, then based on the printed text boxes located below the first adjustment area, adjust the lower boundary of the first adjustment area in the vertical direction to obtain the adjusted mapping area.
[0075] There may be one or more text boxes located below the first adjustment area. Therefore, when adjusting the lower boundary of the first adjustment area in the vertical direction, the adjustment is handled according to the number of text boxes located below the first adjustment area: if there is one text box located below the first adjustment area, the lower boundary of the first adjustment area in the vertical direction is adjusted to the upper boundary of the text box; if there are multiple text boxes located below the first adjustment area, the lower boundary is adjusted to the upper boundary of the multiple text boxes that is closest to the first adjustment area.
[0076] Taking the example of using the printed text box P5 as the printed text box located below the first adjustment area D3'', when adjusting the lower boundary of the first adjustment area D3'', the lower boundary of the first adjustment area D3'' is adjusted to the upper boundary of the printed text box P5.
[0077] When there are no printed text boxes located below the first adjustment area, step S303 can be executed to keep the boundary of the first adjustment area unchanged, or the following steps S304-S305 can be executed on the first adjustment area to adjust the handwritten text boxes, and the processed first adjustment area can be used as the final output adjusted mapping area.
[0078] In this embodiment of the disclosure, by introducing the detection and lower boundary adjustment method of the printed text box located below the first adjustment area during the adjustment process of the question mapping area for the question type, it is possible to accurately cover the user's answer content even when the user goes beyond the answer area and answers near the question stem of the next question.
[0079] Based on the printed text boxes located below the first adjustment area, the lower boundary of the first adjustment area in the vertical direction is adjusted to obtain the second adjustment area. Further steps may include boundary detection and fusion of the handwritten text boxes to address situations where the user continues to add or modify their answer below the original answer area, causing the handwritten content to exceed the range of the second adjustment area. Figure 3B The exemplary flowchart illustrating another step in adjusting the mapping area for compound questions containing problem-solving questions may further include the following steps: S304, for each handwritten line text box on the answer content image, detect whether the lower boundary of the second adjustment area in the vertical direction is within the range of the handwritten line text box; Since the answer area for compound questions (including open-ended questions) is located below the question itself, users typically write their solutions below the question and rarely above it. Therefore, in open-ended question scenarios, only the lower boundary of the second adjustment area needs to be detected and adjusted.
[0080] For any handwritten text box, when detecting whether the lower boundary of the second adjustment area is within the range of the handwritten text box, it is determined whether the lower boundary of the second adjustment area satisfies the following condition: it is located below the upper boundary of the handwritten text box and above the lower boundary of the handwritten text box. If so, it indicates that the lower boundary of the handwritten text box and the lower boundary of the second adjustment area overlap, and it is determined that the lower boundary of the second adjustment area is within the range of the handwritten text box. The detection method of the handwritten text box is the same as in the previous embodiment, and will not be repeated in this embodiment.
[0081] S305, if so, then the lower boundary of the mapped area is further adjusted to the lower boundary of the handwritten text box to obtain the adjusted mapped area.
[0082] This step adjusts the boundary of the second adjustment area within the range of the handwritten text box. The purpose is to include the user's answer content that the second adjustment area does not cover. Therefore, when adjusting the lower boundary of the second adjustment area within the range of the handwritten text box, the lower boundary of the first adjustment area in the vertical direction needs to be expanded outward. Thus, the lower boundary of the mapping area is further adjusted to the lower boundary of the handwritten text box, resulting in the adjusted mapping area.
[0083] To enable those skilled in the art to better understand the question detection method provided in this embodiment, please refer to... Figure 4 The illustrated flowchart represents an overall process flow of a question detection method, which may include at least the following process modules: The question detection module 410 is used to preprocess the answer content image 400, such as scaling it to a fixed size and standardizing the image data, and input the preprocessed answer content image into the question detection model, outputting the question box for each detected question. The question box for any composite question includes: a composite question box representing the entire composite question, and sub-question boxes representing the individual sub-question boxes included in the composite question.
[0084] The question alignment module 420 takes into account the answer content image 400, the standard answer image 401, and each labeled question box 402 in the standard answer image 401, and performs preprocessing, such as scaling the image data to a fixed size and standardizing the image data. The preprocessed data is then input into the question alignment model for forward inference, outputting a mapping region where each labeled question box in the answer content image 400 corresponds one-to-one with the corresponding area in the standard answer image 401. This embodiment can use the RoMa dense matching model as the question alignment model; however, other models can also be used in practical applications.
[0085] Question frame fusion module 430: Used to adjust the boundaries of each mapped region output by the question alignment module 420 based on the question frames output by the question detection module 410; (1) Traverse the mapping area on the answer content image output by the question alignment module. For the current mapping area, traverse the question boxes on the answer content image output by the question detection module, calculate the intersection area of the current mapping area and each question box, and divide the intersection area by the area of the current mapping area to obtain the IOU of the intersection of the current mapping area and each question box relative to the current mapping area. (2) Count the number of mapping regions currently being traversed whose Interchange of Interest (IOU) with all question frames is greater than the threshold: If there are no values greater than the threshold, continue traversing the next mapped region; If the number of boxes with an IOU greater than the threshold is 1, then retrieve the box whose IOU is greater than the threshold corresponding to the currently traversed mapping region, let's say it's box B. Assuming the currently traversed mapping region is box A, then adjust the four boundaries of box A. If we define the horizontal boundary values to increase sequentially from left to right, and the vertical boundary values to decrease sequentially from top to bottom, then the adjustment method is as follows: Left boundary: The smaller left boundary value between boxA and boxB is taken as the left boundary of boxA.
[0086] Upper boundary: The smaller of the upper boundary values of boxA and boxB is taken as the upper boundary of boxA.
[0087] Right boundary: The larger of the right boundary values of boxA and boxB is taken as the right boundary of boxA.
[0088] Lower boundary: The lower boundary of boxA is the larger of the lower boundary values of boxA and boxB.
[0089] If the number of boxes with an IOU greater than the threshold is 2, it indicates that the currently traversed mapping region corresponds to a question box for a sub-question within a composite question. The teacher breaks down a composite question into multiple sub-questions on the standard answer image and marks the question boxes for each sub-question. At this point, two question boxes with an IOU greater than the threshold are selected. One of these question boxes corresponds to the entire composite question, and the other corresponds to a sub-question box for one of the sub-questions included in the composite question. The areas of these two question boxes are calculated, and the question box with the smaller area is selected as the question box that matches the currently traversed mapping region. This is then adjusted according to the method described above. (3) Continue until all mapped regions have been traversed, and perform corresponding operations on each mapped region according to the process described in (1)-(2) above. In this embodiment, the threshold in the above steps can be set to 0.6, and it can also be adjusted according to the application scenario in actual application.
[0090] Question box resizing module 440: Perform the following operations on each mapped region output by the question box merging module: (1) For non-solution questions To address the possibility that students might cross out their answers in the current answer area (e.g., a blank in a fill-in-the-blank question) and then write their answers above or below the crossed-out content, the question frame might truncate the user's answer when the answer area is on the first or last line of the question. Therefore, the question frame needs to be adjusted based on the handwritten content. Iterate through each mapped region output by the question box fusion module. For the currently traversed mapped region, iterate through all handwritten text boxes on the answer content image, and determine whether the upper or lower boundary of the currently traversed mapped region overlaps with the handwritten text box. If the vertical boundary value is defined to increase sequentially from top to bottom, the determination method is as follows: Determine whether the upper boundary value of the currently traversed mapping region is greater than the upper boundary of the handwritten text box and less than the lower boundary of the handwritten text box; if so, determine that the upper boundary of the currently traversed mapping region is on top of the handwritten text box, and adjust the upper boundary of the currently traversed mapping region to be the upper boundary of the handwritten text box.
[0091] Determine whether the lower boundary value of the currently traversed mapping region is greater than the upper boundary of the handwritten text box and less than the lower boundary of the handwritten text box; if so, determine that the lower boundary of the currently traversed mapping region is on top of the handwritten text box, and adjust the lower boundary of the currently traversed mapping region to be the lower boundary of the handwritten text box.
[0092] (2) For problem-solving questions In practical applications, when answering problem-solving questions, students often extend their handwritten content along the blank lines to the top boundary of the next question or even into the range of the next question. Therefore, it is necessary to further expand the question frame for problem-solving questions. ① Extend the lower boundary of the mapping area, which includes the answer questions, to the upper boundary of the first line of printed text below it: Iterate through all printed text boxes on the image containing the answer key, and determine if the printed text box is directly below the mapped area. Specifically, the upper boundary of the printed text box must be greater than the lower boundary of the mapped area including the answer key, and the left-to-right boundary of the printed text box must intersect with the left-to-right boundary of the mapped area including the answer key. If the printed text box meets these conditions, it is considered a candidate text box.
[0093] After traversing all printed text boxes, if the number of candidate text boxes is 0, the lower boundary of the mapping area including the answer question remains unchanged; if the number of candidate text boxes is 1, the lower boundary of the mapping area is adjusted to the upper boundary of the candidate text box; if the number of candidate text boxes is greater than 1, the topmost candidate text box is selected, and the lower boundary of the mapping area is adjusted to the upper boundary of the topmost candidate text box.
[0094] ② Further adjustments were made to the handwritten text box: For each handwritten text box included in the image of the answer content, determine whether the lower boundary of the mapped area including the answer question overlaps with that handwritten text box; if the vertical boundary value is defined to increase sequentially from top to bottom, the determination method is as follows: Determine whether the lower boundary value of the currently traversed mapping region is greater than the upper boundary of the handwritten text box and less than the lower boundary of the handwritten text box; if so, determine that the lower boundary of the currently traversed mapping region is on top of the handwritten text box, and adjust the lower boundary of the currently traversed mapping region to be the lower boundary of the handwritten text box.
[0095] In this embodiment, traditional question detection models rely on fixed detection standards. However, different teachers handle complex questions differently, treating complex questions as a whole or subdividing them into multiple independent questions for separate annotation. This embodiment uses the above mapping operation to make the mapping area on the answer content image consistent with the annotation granularity on the standard answer image. The mapping area is adjusted based on the question box detected for the complex question on the answer content image, so that the mapping area can effectively cover the question and the user's answer to the question. This solves the grading problem caused by inconsistent annotation granularity and improves the accuracy of question grading.
[0096] For an example corresponding to the aforementioned question detection method, see [link to example]. Figure 5 As shown, this application also provides an embodiment of a question detection device, the device comprising: The question frame detection module 501 is configured to detect composite questions in the answer content image to obtain sub-question frames corresponding to each question in the composite question and the composite question frame corresponding to the composite question; the sub-question frame corresponding to any question includes at least the question and the answer content for the question. The question box mapping module 502 is configured to map each marked question box in the standard answer image of the composite question to the answer content image, so as to obtain the mapping area in the answer content image corresponding to each marked question box; The question box matching module 503 is configured to, for each mapped region on the answer content image, search for a question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. The adjustment module 504 is configured to adjust the mapping area based on the question box that matches the mapping area, so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
[0097] In some embodiments, when the adjustment module is configured to adjust the mapped region based on a question box that matches the mapped region, it includes: The target question frame determination module is configured such that if there is only one question frame that matches the mapped area, then that question frame is taken as the target question frame; if there are multiple question frames that match the mapped area, then the question frame with the smallest area among the matched question frames is taken as the target question frame. The boundary adjustment module is configured to adjust the boundary of the mapped region based on the boundary of the target title box.
[0098] In some embodiments, the boundary adjustment module is configured to: In the horizontal direction of the answer content image, if the left boundary of the target question box is located to the left of the left boundary of the mapping area, then the left boundary of the mapping area is adjusted to the left boundary of the target question box; and if the right boundary of the target question box is located to the right of the right boundary of the mapping area, then the right boundary of the mapping area is adjusted to the right boundary of the target question box. In the vertical direction of the answer content image, if the upper boundary of the target question frame is located above the upper boundary of the mapping area, then the upper boundary of the mapping area is adjusted to be the upper boundary of the target question frame; and if the lower boundary of the target question frame is located below the lower boundary of the mapping area, then the lower boundary of the mapping area is adjusted to be the lower boundary of the target question frame.
[0099] In some embodiments, when the adjustment module is configured to adjust the mapped region based on a question box that matches the mapped region, it includes: The first adjustment module is configured to adjust the mapping region based on the question box that matches the mapping region, so as to obtain the first adjustment region corresponding to the mapping region; The second adjustment module is configured to, when the question in the first adjustment area is not a question-solving type, detect whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box for each handwritten text box; if so, adjust the boundary of the first adjustment area within the range of the handwritten text box according to the boundary of the handwritten text box to obtain the adjusted mapping area.
[0100] In some embodiments, when the second adjustment module is configured to adjust the boundary of the first adjustment region within the range of the handwritten text box, it includes: If the upper boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the upper boundary of the handwritten text box. If the lower boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the lower boundary of the handwritten text box.
[0101] In some embodiments, when the adjustment module is configured to adjust the mapped region based on a question box that matches the mapped region, it further includes: The third adjustment module is configured to, when the question in the first adjustment area is an answer question, detect whether there is a printed text box located below the first adjustment area in the vertical direction of the answer content image from each printed text box included in the answer content image; if so, adjust the lower boundary of the first adjustment area in the vertical direction based on each printed text box located below the first adjustment area to obtain the adjusted mapping area.
[0102] In some embodiments, when the third adjustment module is configured to adjust the lower boundary of the first adjustment region in the vertical direction, it includes: If there is a text box for printed text located below the first adjustment area, then the lower boundary of the first adjustment area in the vertical direction is adjusted to the upper boundary of the text box for printed text. If multiple printed text boxes are located below the first adjustment area, the lower boundary is adjusted to the upper boundary of each of the multiple printed text boxes that is closest to the first adjustment area.
[0103] In some embodiments, after adjusting the lower boundary of the first adjustment region in the vertical direction to obtain the second adjustment region, the device further includes: The fourth adjustment module is configured to detect whether the lower boundary of the second adjustment area in the vertical direction is within the range of the handwritten text box for each handwritten text box on the answer content image; if so, the lower boundary of the mapping area is further adjusted to the lower boundary of the handwritten text box to obtain the adjusted mapping area.
[0104] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0106] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 6 As shown, the electronic device 600 includes at least one processor 601, a memory 602, and a bus 603. At least one processor 601 is electrically connected to the memory 602. The memory 602 is configured to store at least one computer-executable instruction, and the processor 601 is configured to execute the at least one computer-executable instruction to perform the steps of any title detection method provided in any embodiment or optional implementation of this application.
[0107] Furthermore, the processor 601 can be an FPGA (Field-Programmable Gate Array) or other devices with logic processing capabilities, such as an MCU (Microcontroller Unit) or a CPU (Central Processing Unit).
[0108] This application also provides another readable storage medium storing a computer program that, when executed by a processor, implements the steps of any question detection method provided in any embodiment or optional implementation of this application.
[0109] The readable storage media provided in this application include, but are not limited to, any type of disk (including floppy disk, hard disk, optical disk, CD-ROM, and magneto-optical disk), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, readable storage media include any medium by which a device (e.g., a computer) stores or transmits information in a readable form.
[0110] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0111] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0112] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method of title detection, the method comprising: The method includes: Detect composite questions in the answer content image to obtain sub-question boxes corresponding to each question in the composite question and the composite question title box corresponding to the composite question; the sub-question box corresponding to any question question includes at least the question question and the answer content for the question question; Map each marked question box in the standard answer image of the compound question to the answer content image to obtain the mapping area in the answer content image corresponding to each marked question box; For each mapped region on the answer content image, find the question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. Based on the question box that matches the mapping area, the mapping area is adjusted so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein, the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
2. The method of claim 1, wherein, Based on the question box that matches the mapped area, adjust the mapped area, including: If there is only one question box that matches the mapped area, then that question box will be used as the target question box. If there are multiple question boxes that match the mapped area, the question box with the smallest area among the matched question boxes will be taken as the target question box. The boundaries of the mapping region are adjusted based on the boundaries of the target title box.
3. The method of claim 2, wherein, Adjusting the boundaries of the mapped region includes: In the horizontal direction of the answer content image, if the left boundary of the target question box is located to the left of the left boundary of the mapping area, then the left boundary of the mapping area is adjusted to the left boundary of the target question box; and if the right boundary of the target question box is located to the right of the right boundary of the mapping area, then the right boundary of the mapping area is adjusted to the right boundary of the target question box. In the vertical direction of the answer content image, if the upper boundary of the target question frame is located above the upper boundary of the mapping area, then the upper boundary of the mapping area is adjusted to be the upper boundary of the target question frame; and if the lower boundary of the target question frame is located below the lower boundary of the mapping area, then the lower boundary of the mapping area is adjusted to be the lower boundary of the target question frame.
4. The method of claim 1, wherein, Based on the question box that matches the mapped area, adjust the mapped area, including: Based on the question box that matches the mapped area, the mapped area is adjusted to obtain the first adjusted area corresponding to the mapped area; If the question in the first adjustment area is not a question-solving type, for each handwritten text box in the answer content image, it is detected whether at least one boundary of the first adjustment area in the vertical direction of the answer content image is within the range of the handwritten text box. If so, then based on the boundary of the handwritten text box, the boundary of the first adjustment area within the range of the handwritten text box is adjusted to obtain the adjusted mapping area.
5. The method of claim 4, wherein, Adjusting the boundaries of the handwritten text box within the first adjustment area includes: If the upper boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the upper boundary of the handwritten text box. If the lower boundary of the first adjustment area in the vertical direction is within the range of the handwritten text box, then the upper boundary is adjusted to be the lower boundary of the handwritten text box.
6. The method of claim 4, wherein, The method also includes: If the question in the first adjustment area is a question of the answer type, detect whether there is a printed text box located below the first adjustment area in the vertical direction of the answer content image from each printed text box included in the answer content image. If it exists, then based on the printed text boxes located below the first adjustment area, the lower boundary of the first adjustment area in the vertical direction is adjusted to obtain the adjusted mapping area.
7. The method of claim 6, wherein, Adjusting the lower boundary of the first adjustment area in the vertical direction includes: If there is a text box for printed text located below the first adjustment area, then the lower boundary of the first adjustment area in the vertical direction is adjusted to the upper boundary of the text box for printed text. If multiple printed text boxes are located below the first adjustment area, the lower boundary is adjusted to the upper boundary of each of the multiple printed text boxes that is closest to the first adjustment area.
8. The method according to claim 6 or 7, characterized in that, After adjusting the lower boundary of the first adjustment region in the vertical direction to obtain the second adjustment region, the method further includes: For each handwritten line text box on the image of the answer content, it is detected whether the lower boundary of the second adjustment area in the vertical direction is within the range of the handwritten line text box; If so, the lower boundary of the mapped area is further adjusted to the lower boundary of the handwritten text box to obtain the adjusted mapped area.
9. A document inspection apparatus characterized by comprising: The device includes: The question frame detection module is configured to detect composite questions in the answer content image to obtain sub-question frames corresponding to each question in the composite question and the composite question frame corresponding to the composite question; the sub-question frame corresponding to any question includes at least the question and the answer content for the question; The question box mapping module is configured to map each marked question box in the standard answer image of the composite question to the answer content image, so as to obtain the mapping area in the answer content image corresponding to each marked question box; The question box matching module is configured to, for each mapped region on the answer content image, search for a question box that matches the mapped region from the sub-question boxes corresponding to each question in the composite question and the composite question box corresponding to the composite question. The adjustment module is configured to adjust the mapping area based on the question box that matches the mapping area, so that the adjusted mapping area includes at least the question in the matching question box and the answer to the question; wherein, the adjusted mapping area and the corresponding marked question box are used to grade the answer in the adjusted mapping area.
10. An electronic device, comprising: include: Memory, processor; The memory is used to store computer programs; The processor is configured to invoke the computer program to implement the method as described in any one of claims 1-8.
11. A readable storage medium, having stored thereon a computer program, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.