Test paper image labeling method and device, storage medium and electronic equipment

CN114943406BActive Publication Date: 2026-08-28BEIJING ZHIYUAN HANGCHENG SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210346565.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2026-08-28
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

[0003]目前进行数据标注时,一般是通过线下沟通进行管理,包括试卷图像数据的导入和标注数据的导出都是通过移动存储设备(硬盘、U盘等)进行传输,当标注量和标注团队规模较大的时候就很难适用

Benefits of technology

[0042] In the above technical solution, a preset node is selected from the first sub-window and dragged to the second sub-window to create a corresponding type of work node. Multiple work nodes are then connected sequentially according to task type to generate a workflow for a real annotation task. This workflow facilitates the breakdown of complex annotation tasks into multiple simpler sub-tasks, with each annotation work node only needing to complete one simple sub-task. On one hand, the task executor for each sub-task only needs to focus on their own sub-task, making the annotation task more focused, more efficient, and of higher quality. On the other hand, after all work nodes are completed, the annotation data from each work node can be simply merged to obtain a complete set of annotation data for the complex annotation task. Since all annotation operations are completed online, data security is ensured, the inconvenience of importing and exporting data via mobile storage devices is reduced, and annotation efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943406B_ABST
    Figure CN114943406B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a test paper image labeling method and device, a storage medium and an electronic device, wherein the method comprises: determining a plurality of preset nodes and displaying the plurality of preset nodes in a first sub-window; in response to selecting a preset node from the first sub-window and dragging it to a second sub-window, a task configuration window configured based on the selected preset node is popped up; obtaining task information and task performers configured in the task configuration window, and creating a work node corresponding to the task type of the preset node; after creating a plurality of work nodes, a work flow is generated; the work nodes in the work flow are sequentially assigned to the corresponding task performers; when all the work nodes in the work flow are completed, the data of each labeling work node is merged to obtain the final labeling data of the test paper image. The present disclosure can ensure data security, reduce the inconvenience of importing and exporting data through a mobile storage device, and improve labeling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data annotation technology, specifically to a method, apparatus, storage medium, and electronic device for annotating test paper images. Background Technology

[0002] With the continuous development of technology, more and more artificial intelligence (AI) technologies are being applied in the education field. AI is being used in many tasks, such as text recognition, intelligent grading, automatic test paper generation, student learning analysis, and electronic test papers. Utilizing AI technology can greatly improve teachers' productivity and work efficiency. To enable computers to better possess these capabilities, a large amount of high-quality manually labeled data is needed as supervisory information to train models and improve the accuracy of machine learning. The efficiency and quality of manual labeling are crucial aspects of model training.

[0003] Currently, data annotation is generally managed through offline communication. This includes importing test paper image data and exporting annotation data, which are transferred via mobile storage devices (hard drives, USB flash drives, etc.). This approach is not suitable when the amount of annotation and the size of the annotation team are large. Summary of the Invention

[0004] The purpose of this disclosure is to provide a method, apparatus, storage medium, and electronic device for annotating test paper images to solve the aforementioned technical problems.

[0005] To achieve the above objectives, this disclosure provides a method for annotating test paper images, including:

[0006] Multiple preset nodes are identified and displayed in a first sub-window; each preset node corresponds to a task type.

[0007] In response to selecting a preset node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up to configure the preset node based on the selected node.

[0008] Retrieve the task information and task executor configured in the task configuration window, and create a work node corresponding to the task type of the preset node;

[0009] After creating multiple work nodes, the multiple work nodes are connected sequentially according to their corresponding task types, and a workflow is generated in response to the confirmation of creation. The workflow includes at least one annotation work node generated according to the preset node of the annotation type, and the task corresponding to each annotation work node is to annotate the test paper image accordingly.

[0010] The work nodes in the workflow are sequentially assigned to the corresponding task executors.

[0011] Once all work nodes in the workflow are completed, the data from each annotation work node are merged to obtain the final annotation data for the exam paper image.

[0012] Optionally, each task executor is visible to task executors who are themselves in the workflow, but not visible to task executors who are not themselves in their work nodes.

[0013] Optionally, the workflow also includes an audit work node generated based on a preset node of the audit type, a quality inspection work node generated based on a preset node of the quality inspection type, and an acceptance work node generated based on a preset node of the acceptance type.

[0014] The task corresponding to the review work node is to review the labeled data of the previous work node; the task corresponding to the quality inspection work node is to conduct random quality inspection on the labeled data that has been reviewed and approved; and the task corresponding to the acceptance work node is to accept the labeled data that has been reviewed and approved if the random quality inspection is passed.

[0015] Optionally, the method further includes:

[0016] After generating the workflow, save the workflow;

[0017] When a new workflow is created, in response to selecting a workflow from at least one saved historical workflow as the base workflow, the base workflow is displayed in a second sub-window;

[0018] In response to selecting any work node in the basic workflow, a task configuration window pops up to configure based on the selected work node;

[0019] Receive new task information and / or new task executors entered in the task configuration window, and update the work node;

[0020] In response to the confirmation of the creation operation, a new workflow is generated.

[0021] Optionally, merging the data from each annotation work node to obtain the final annotation data for the exam paper image includes:

[0022] By merging the data from each annotation work node, multiple annotation boxes and corresponding annotation labels are obtained for the test paper image.

[0023] Calculate the hierarchical relationship between the multiple annotation boxes;

[0024] Based on the multiple annotation boxes, the annotation labels corresponding to each annotation box, and the hierarchical relationship between the multiple annotation boxes, multi-level annotation data corresponding to the test paper image is generated.

[0025] Optionally, calculating the hierarchical relationship between the plurality of annotation boxes includes:

[0026] Based on the intersection relationship between the multiple annotation boxes and the area of ​​each annotation box, determine the parent-child relationship between any two annotation boxes;

[0027] The hierarchical relationship between the multiple annotation boxes is determined based on the parent-child relationship between any two annotation boxes.

[0028] Optionally, determining the parent-child relationship between any two annotation boxes based on the intersection relationship between the plurality of annotation boxes and the area of ​​each annotation box includes:

[0029] For any two of the plurality of annotation boxes, calculate the ratio of the intersection area between the two annotation boxes to the area of ​​the annotation box with the smaller area, and obtain the intersection-smaller ratio between the two annotation boxes;

[0030] When the intersection ratio is greater than a preset threshold, the label box with the smaller area of ​​the two label boxes is determined to be the child label box, and the label box with the larger area is determined to be the parent label box.

[0031] This disclosure also provides a test paper image annotation device, including:

[0032] The preset node determination module is used to determine a variety of preset nodes and display the variety of preset nodes in a first sub-window; wherein, each preset node corresponds to a task type;

[0033] The node drag-and-drop module is used to pop up a task configuration window based on the selected preset node in response to selecting a preset node from the first sub-window and dragging it to the second sub-window.

[0034] The work node creation module is used to obtain the task information and task executor configured in the task configuration window, and create a work node corresponding to the task type of the preset node.

[0035] The workflow generation module is used to connect multiple work nodes sequentially according to their corresponding task types after creating multiple work nodes, and generate a workflow in response to the confirmation of creation operation; wherein, the workflow includes at least a labeling work node generated according to a preset node based on the labeling type, and the task corresponding to the labeling work node is to label the test paper image;

[0036] The work node allocation module is used to sequentially allocate the work nodes in the workflow to the corresponding task executors.

[0037] The annotation data acquisition module is used to merge the data of each annotation work node after all work nodes in the workflow are completed, so as to obtain the final annotation data of the test paper image.

[0038] This disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.

[0039] This disclosure also provides an electronic device, including:

[0040] A memory on which computer programs are stored;

[0041] A processor for executing the computer program in the memory to implement the steps of the method described in the first aspect.

[0042] In the above technical solution, a preset node is selected from the first sub-window and dragged to the second sub-window to create a corresponding type of work node. Multiple work nodes are then connected sequentially according to task type to generate a workflow for a real annotation task. This workflow facilitates the breakdown of complex annotation tasks into multiple simpler sub-tasks, with each annotation work node only needing to complete one simple sub-task. On one hand, the task executor for each sub-task only needs to focus on their own sub-task, making the annotation task more focused, more efficient, and of higher quality. On the other hand, after all work nodes are completed, the annotation data from each work node can be simply merged to obtain a complete set of annotation data for the complex annotation task. Since all annotation operations are completed online, data security is ensured, the inconvenience of importing and exporting data via mobile storage devices is reduced, and annotation efficiency is improved.

[0043] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0044] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:

[0045] Figure 1 A flowchart of an exemplary embodiment of a test paper image annotation method is shown;

[0046] Figure 2 This diagram illustrates how selecting a preset node from the first sub-window and dragging it to the second sub-window generates a working node.

[0047] Figure 3This is another illustration showing how selecting a preset node from the first sub-window and dragging it to the second sub-window generates a working node;

[0048] Figure 4 A schematic diagram of the workflow corresponding to the generation and marking of application questions is shown;

[0049] Figure 5 A flowchart illustrating a specific implementation of generating a new workflow is shown in an exemplary embodiment.

[0050] Figure 6 A flowchart illustrating a specific implementation of step S160 provided in an exemplary embodiment is shown;

[0051] Figure 7 This diagram illustrates how to determine the parent-child relationship between any two annotation boxes in a set of multiple annotation boxes.

[0052] Figure 8 A block diagram of an examination paper image annotation apparatus provided in an exemplary embodiment is shown;

[0053] Figure 9 A block diagram of an electronic device provided in an exemplary embodiment is shown. Detailed Implementation

[0054] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.

[0055] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.

[0056] In the education field, intelligent grading and digitized exam papers rely heavily on artificial intelligence (AI) technology. The core and foundation of these technologies are document layout analysis and OCR (Optical Character Recognition) technology. Document layout analysis requires computers to logically and hierarchically interpret various information on the exam paper, such as student ID, class, name, subject information, question number, question area, answer area, question content, answer content, formulas, and charts. OCR technology needs to analyze the text line position information, language type, and content information in detail. Depending on actual business needs, the final result is output using a single-level or multi-level approach.

[0057] To enable computers to better perform the aforementioned capabilities, a large amount of high-quality manually labeled data is needed as supervised information to train models and improve the accuracy of machine learning. The efficiency and quality of manual labeling are crucial aspects of model training; therefore, an efficient and easy-to-use data labeling system is particularly important. Currently, there are no publicly available data labeling tools specifically designed for labeling tasks in the education field. Typically, public or commercial data labeling tools are selected based on the specific labeling task requirements (timeframe, data volume, labeling difficulty, etc.).

[0058] When choosing to use publicly available data annotation tools, for some simple tasks in the education field (such as text detection in a single language or font), you can directly select an annotation tool for a specific image task, such as existing annotation tools for object detection, image classification, and contour annotation. However, for complex annotation tasks in the education field (such as annotation of structured data in layout), this approach cannot be used.

[0059] Using publicly available image annotation tools has several drawbacks:

[0060] (1) Adaptability and flexibility are difficult to reconcile. Software that can adapt to multiple environments is often highly encapsulated and not open source, and does not support customized modifications.

[0061] (2) Lack of task management, process management and personnel management, and can only be managed through offline communication. For example, the import of test paper data and the export of annotation data are transmitted through mobile storage devices, which is difficult to apply when the amount of annotation and the size of the annotation team are large.

[0062] (3) It is difficult to guarantee the security of data when transmitting data offline.

[0063] Commercial data annotation tools are generally provided by crowdsourced annotation platforms. These platforms can basically annotate data such as images, videos, text, and audio, but each has its own focus. Some are good at image processing, while others are better at video annotation.

[0064] Using commercial data platforms for annotation still presents the following problems:

[0065] (1) The data annotation systems on the market are all general annotation systems that do not address the specific needs of educational scenarios, such as primary school math mental arithmetic recognition, test paper layout analysis, and English essay correction. From the perspective of functionality and efficiency, they cannot fully meet the needs of the education field.

[0066] (2) Platforms generally use a crowdsourcing model to allocate labeling tasks, resulting in inconsistent quality of labeled data, which affects the accuracy of algorithm models;

[0067] (3) Data labeling tasks based on the crowdsourcing model may result in a lack of security for user data and face the risk of privacy leakage.

[0068] In summary, existing methods for annotating test paper data in educational settings have certain problems in terms of data security, annotation quality, and annotation efficiency. Therefore, this disclosure provides a method for annotating test paper images. Figure 1 A flowchart of an exemplary embodiment of a test paper image annotation method is shown, such as... Figure 1 As shown, the method includes:

[0069] S110, determine multiple preset nodes and display these multiple preset nodes in the first sub-window; wherein, each preset node corresponds to a task type.

[0070] First, several preset nodes are identified, each corresponding to a specific task type. These preset nodes include, but are not limited to, annotation nodes, review nodes, quality inspection nodes, acceptance nodes, and pre-transcription nodes. Furthermore, custom preset nodes can be added for some special annotation tasks. These preset nodes are then displayed in the first sub-window of the page.

[0071] The page includes a first sub-window and a second sub-window. The first sub-window is used to display the various preset nodes, and the second sub-window is used to create work nodes based on the preset nodes in the first sub-window and display a workflow composed of multiple work nodes.

[0072] S120, in response to selecting a preset node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up to configure based on the selected preset node.

[0073] S130: Obtain the task information and task executors configured in the task configuration window, and create a worker node corresponding to the task type of the preset node.

[0074] After displaying various preset nodes in the first sub-window, the user can select one of these preset nodes and drag it to the second sub-window. In response to selecting a preset node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up, allowing configuration based on the selected preset node. After the task configuration window appears, the user can configure the corresponding task information and task executors, and confirm the configuration. Based on the task information and task executors configured by the user, a worker node corresponding to the task type of that preset node is then created.

[0075] As an example, the first sub-window displays various preset nodes, including annotation nodes, review nodes, quality inspection nodes, and acceptance nodes. Figures 2-3 This diagram illustrates how selecting a preset node from the first sub-window and dragging it to the second sub-window generates a working node.

[0076] Please refer to Figures 2-3 In response to selecting a labeling node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up, allowing users to configure the labeling task based on the labeling node. This window retrieves the task information and personnel configured by the user regarding the labeling task. Upon completion by the user, a labeling work node is created. For example, the task information for this labeling work node could be to label key objectives on the exam paper, label handwritten equations on the exam paper, or label handwritten Chinese characters on the exam paper, etc.

[0077] In response to selecting an audit node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up, allowing users to configure tasks based on the audit node. This window retrieves the task information and personnel configured by the user for the audit task. Upon completion by the user, an audit work node is created. For example, the task information for this audit work node could be auditing the labeled data of key objectives on the exam paper, auditing the labeled data of handwritten equations, or auditing the labeled data of handwritten Chinese characters, etc.

[0078] In response to selecting a quality inspection node from the first sub-window and dragging it to the second sub-window, a task configuration window based on the quality inspection node pops up. The task information and task executors configured by the user in the task configuration window are obtained, and a quality inspection work node is created.

[0079] In response to selecting an acceptance node from the first sub-window and dragging it to the second sub-window, a task configuration window based on the acceptance node pops up. The task information and task executors configured by the user in the task configuration window are obtained, and an acceptance node is created.

[0080] S140, after creating multiple work nodes, the multiple work nodes are connected sequentially according to their corresponding task types, and a workflow is generated in response to the confirmation of creation; wherein, the workflow includes at least one annotation work node generated according to a preset node based on the annotation type.

[0081] Each workflow includes at least one annotation node. Within each workflow, there can be one or more annotation nodes and one or more review nodes. It should be noted that the number of review nodes in each workflow is the same as the number of annotation nodes. The next node after each annotation node is a review node, whose task is to review the annotation data from the previous node. The next node after a review node can be either an annotation node or a quality inspection node. There can be one quality inspection node and one acceptance node. The next node after a quality inspection node is an acceptance node, whose task is to perform random quality checks on the approved annotation data. The task of the acceptance node is to accept the approved annotation data if the random quality checks pass.

[0082] When connecting multiple work nodes according to their corresponding task types, they should be connected in the following order: the next work node after the labeled work node is the audit work node, the next work node after the audit work node is the labeled work node or the quality inspection work node, and the next work node after the quality inspection work node is the acceptance work node.

[0083] S150 assigns the work nodes in the workflow to the corresponding task executors in sequence.

[0084] After the workflow is generated, the first work node in the workflow is assigned to the task executor corresponding to the first work node. After the first work node is completed, the next work node in the workflow is assigned to the task executor corresponding to the next work node. This process continues until the last work node in the workflow is completed, at which point step S160 is executed.

[0085] S160: After all work nodes in the workflow are completed, merge the data from each annotation work node to obtain the final annotation data for the test paper image.

[0086] It is worth noting that in practical application scenarios, for complex annotation tasks in educational settings, such as annotation of structured information on a page, this disclosure can break down complex annotation tasks into multiple simple tasks for annotation, with each annotation work node used to complete one of the decomposed simple tasks. For example, for the task of grading and annotating word problems on a math test paper, it can be broken down into multiple simple tasks such as annotating key targets on the test paper, annotating handwritten equations, and annotating handwritten Chinese characters.

[0087] Taking the task of grading and annotating word problems as an example, Figure 4 This diagram illustrates the workflow for generating and annotating the application problem grading task. For example... Figure 4As shown, this workflow includes the following nodes connected in sequence: test paper key target annotation node, test paper key target review node, handwritten equation annotation node, handwritten equation review node, handwritten Chinese annotation node, handwritten Chinese review node, word problem grading and quality inspection node, and word problem grading and acceptance node.

[0088] The "Key Objectives Annotation" node is used to annotate key objectives on the exam paper. Key objectives include, but are not limited to, question numbers, question stems, images, and answer areas for each application problem. After completing the Key Objectives Annotation node, the process moves to the Key Objectives Review node. This node reviews the annotated data for the key objectives. If the review is successful, the Key Objectives Annotation node is complete, and the process moves to the Handwritten Equation Annotation node. If the review fails, the process returns to the previous node, and the Key Objectives are annotated again.

[0089] The handwritten equation annotation node is used to annotate handwritten equations in the answer area of ​​the test paper. After the handwritten equation annotation node is completed, the process moves to the handwritten equation review node. The handwritten equation review node reviews the annotated data of the handwritten equations. If the review is passed, the handwritten equation review node is completed, and the process moves to the handwritten Chinese annotation node. If the review is failed, the process returns to the previous node and the handwritten equation annotation is redone.

[0090] The handwritten Chinese annotation work node is used to annotate the handwritten Chinese characters in the answer area of ​​the test paper. After the handwritten Chinese annotation work node is completed, the process moves to the handwritten Chinese review work node. The handwritten Chinese review work node is used to review the handwritten Chinese annotation data. If the review is passed, the handwritten Chinese review work node is completed, and the process moves to the application problem grading and quality inspection work node. If the review is failed, the process returns to the previous work node, and the handwritten Chinese annotation is redone.

[0091] The word problem grading quality inspection node is used to randomly inspect the labeled data of key objectives, handwritten equations, and handwritten Chinese characters in the approved test papers. If the quality inspection passes, the word problem grading quality inspection node is completed, and the process moves to the word problem grading acceptance node. If the quality inspection fails, the process returns to the previous node. The word problem grading acceptance node is used to accept the labeled data that has already been approved.

[0092] according to Figure 4The workflow shown assigns each work node to the corresponding task executor in sequence. When all work nodes in the workflow are completed, the data of each labeled work node in the workflow are merged to obtain the final labeled data of the test paper image. The final labeled data includes the labeled data of the key targets of the test paper, the labeled data of the handwritten formulas, and the labeled data of the handwritten Chinese characters.

[0093] Understandably, in the above technical solution, a preset node is selected from the first sub-window and dragged to the second sub-window to create a corresponding type of work node. Then, multiple created work nodes are connected sequentially according to task type to generate a workflow for a real annotation task. Establishing a workflow facilitates the decomposition of complex annotation tasks into multiple simple annotation sub-tasks, with each annotation work node only needing to complete one simple sub-task. On the one hand, the task executor for each sub-task only needs to focus on their own sub-task, making the annotation task more focused, more efficient, and of higher quality. On the other hand, after all work nodes are completed, simply merging the annotation data from each work node yields a complete set of annotation data for the complex annotation task.

[0094] Optionally, to further ensure data security, each task executor is visible to their own work nodes in the workflow, but not to work nodes where they are not their own work nodes.

[0095] Optionally, the workflow can be saved after each workflow is generated.

[0096] Figure 5 A flowchart illustrating a specific implementation of generating a new workflow according to an exemplary embodiment is shown. Figure 5 As shown, the method also includes the following steps:

[0097] S210, when creating a new workflow, in response to selecting a workflow from at least one saved historical workflow as the base workflow, and displaying the base workflow in a second sub-window.

[0098] S220, in response to selecting any work node in the basic workflow, pops up a task configuration window that configures based on the selected work node.

[0099] S230 receives new task information and / or new task executors entered in the task configuration window and updates the work node.

[0100] S240, in response to the confirmation of the creation operation, generates a new workflow.

[0101] Understandably, in the above technical solution, after each workflow is generated, the generated workflow is saved as a template. When a similar labeled task needs to create a workflow, a workflow can be selected from at least one saved historical workflow as the base workflow, and this base workflow is displayed in the second sub-window. Selecting any work node in the base workflow displayed in the second sub-window will automatically pop up a task configuration window based on the selected work node. The user can configure new task information and / or new task executors in this task configuration window and confirm the configuration. The task information and task executors for that work node will then be updated. The user confirms creation in the second sub-window, and a new workflow is automatically generated. Therefore, this disclosure can generate new workflows with one click from historical workflows, avoiding repetitive operations and improving workflow generation efficiency.

[0102] Figure 6 A flowchart illustrating step S160 in an exemplary embodiment, showing the merging of data from various annotation work nodes to obtain the final annotation data for the exam paper image, is shown. Figure 6 As shown, step S160 includes:

[0103] S310, merge the data of each annotation work node to obtain multiple annotation boxes for the test paper image and the annotation label corresponding to each annotation box.

[0104] By merging the annotation data from each annotation work node, complete annotation data for the exam paper image can be obtained. For example, this complete annotation data includes annotation data for key targets of the exam paper from the exam paper key target annotation work node, annotation data for handwritten equations from the handwritten equation annotation work node, and annotation data for handwritten Chinese characters from the handwritten Chinese character annotation work node. This complete annotation data includes multiple annotation boxes and corresponding annotation labels for each annotation box.

[0105] S320, calculate the hierarchical relationship between the multiple annotation boxes.

[0106] S330, Based on the multiple annotation boxes, the annotation labels corresponding to each annotation box, and the hierarchical relationship between the multiple annotation boxes, generate multi-level annotation data corresponding to the test paper image.

[0107] This disclosure supports multi-level annotation of exam paper images to obtain multi-level structured annotation data, which is beneficial for complex layout analysis tasks. Multi-level annotation is typically done manually by dragging and dropping annotation elements, which is inefficient and prone to errors. The above-mentioned technical solution automatically calculates the hierarchical relationship between multiple annotation boxes, automatically generating multi-level structured annotation data. This avoids manually dragging and dropping annotation elements to determine hierarchical classification, improving annotation efficiency and reducing the probability of errors.

[0108] Optionally, step S320 includes: determining the parent-child relationship between any two annotation boxes based on the intersection relationship between the plurality of annotation boxes and the area of ​​each annotation box; and determining the hierarchical relationship between the plurality of annotation boxes based on the parent-child relationship between any two annotation boxes.

[0109] In the specific implementation, for any two annotation boxes among the multiple annotation boxes, the ratio of the intersection area between the two annotation boxes to the area of ​​the annotation box with the smaller area among the two annotation boxes is calculated to obtain the intersection ratio between the two annotation boxes. It is then determined whether the intersection ratio is greater than a preset threshold. When the intersection ratio is greater than the preset threshold, the annotation box with the smaller area among the two annotation boxes is determined to be the child annotation box and the annotation box with the larger area is determined to be the parent annotation box.

[0110] Figure 7 This diagram illustrates how to determine the parent-child relationship between any two annotation boxes, such as... Figure 7 As shown, the test paper image is labeled with bounding boxes A and B. For bounding boxes A and B, determine the area of ​​the intersection A∩B, denoted as S(A∩B), and determine the area of ​​the bounding box with the smaller area between A and B, denoted as S(min[A,B]). Then calculate the intersection ratio P between bounding boxes A and B:

[0111] P = S(A∩B) / S(min[A,B]);

[0112] Assuming that the smaller of the two label boxes A and B is labeled box B, when the intersection ratio P is greater than a preset threshold (such as 0.9), the parent-child relationship between label boxes A and B is determined as follows: label box B is the child label box, and label box A is the parent label box.

[0113] In this disclosure, the parent-child relationship between any two annotation boxes is determined by automatically calculating the intersection ratio between each pair of annotation boxes, thereby obtaining the hierarchical relationship between the multiple annotation boxes, wherein the level of the parent annotation box is the level above the level of the child annotation box.

[0114] Figure 8A block diagram of an exemplary embodiment of a test paper image annotation device 400 is shown. Please refer to... Figure 8 The device 400 includes:

[0115] The preset node determination module 410 is used to determine multiple preset nodes and display the multiple preset nodes in a first sub-window; wherein, each preset node corresponds to a task type;

[0116] The node drag-and-drop module 420 is used to pop up a task configuration window based on the selected preset node in response to selecting a preset node from the first sub-window and dragging it to the second sub-window.

[0117] The work node creation module 430 is used to obtain the task information and task executor configured in the task configuration window, and create a work node corresponding to the task type of the preset node.

[0118] The workflow generation module 440 is used to connect the multiple work nodes sequentially according to their corresponding task types after creating multiple work nodes, and generate a workflow in response to the confirmation of creation operation; wherein, the workflow includes at least one annotation work node generated according to the preset node of the annotation type, and the task corresponding to each annotation work node is to annotate the test paper image accordingly.

[0119] The work node allocation module 450 is used to sequentially allocate the work nodes in the workflow to the corresponding task executors.

[0120] The annotation data acquisition module 460 is used to merge the data of each annotation work node after all work nodes in the workflow are completed, so as to obtain the final annotation data of the test paper image.

[0121] Optionally, each task executor is visible to task executors who are themselves in the workflow, but not visible to task executors who are not themselves in their work nodes.

[0122] Optionally, the workflow further includes an audit work node generated based on a pre-set node according to the audit type, a quality inspection work node generated based on a pre-set node according to the quality inspection type, and an acceptance work node generated based on a pre-set node according to the acceptance type; wherein, the task corresponding to the audit work node is to audit the labeled data of the previous work node, the task corresponding to the quality inspection work node is to perform random quality inspection on the approved labeled data, and the task corresponding to the acceptance work node is to accept the approved labeled data if the random quality inspection passes.

[0123] Optionally, the device 400 also includes:

[0124] A workflow saving module is used to save the workflow after it has been generated;

[0125] The workflow display module is used to select a workflow as the base workflow from at least one saved historical workflow when a new workflow is created, and display the base workflow in a second sub-window.

[0126] The node reconfiguration module is used to pop up a task configuration window based on the selected work node in response to the selection of any work node in the basic workflow.

[0127] The node update module is used to receive new task information and / or new task executors entered in the task configuration window and update the work node;

[0128] The workflow regeneration module is used to generate a new workflow in response to the confirmation of creation.

[0129] Optionally, the labeled data acquisition module 460 includes:

[0130] The node data merging module is used to merge the data of each annotation working node to obtain multiple annotation boxes annotated on the test paper image and the annotation label corresponding to each annotation box.

[0131] The hierarchy calculation module is used to calculate the hierarchy relationship between the multiple annotation boxes;

[0132] The annotation data generation module is used to generate multi-level annotation data corresponding to the test paper image based on the multiple annotation boxes, the annotation labels corresponding to each annotation box, and the hierarchical relationship between the multiple annotation boxes.

[0133] Optionally, the hierarchical relationship calculation module is used for:

[0134] Based on the intersection relationship between the multiple annotation boxes and the area of ​​each annotation box, determine the parent-child relationship between any two annotation boxes;

[0135] The hierarchical relationship between the multiple annotation boxes is determined based on the parent-child relationship between any two annotation boxes.

[0136] Optionally, the hierarchical relationship calculation module is used for:

[0137] For any two of the plurality of annotation boxes, calculate the ratio of the intersection area between the two annotation boxes to the area of ​​the annotation box with the smaller area, and obtain the intersection-smaller ratio between the two annotation boxes;

[0138] When the intersection ratio is greater than a preset threshold, the label box with the smaller area of ​​the two label boxes is determined to be the child label box, and the label box with the larger area is determined to be the parent label box.

[0139] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0140] Figure 9 This is a block diagram illustrating an electronic device 500 according to an exemplary embodiment. For example... Figure 9 As shown, the electronic device 500 may include a processor 501 and a memory 502. The electronic device 500 may also include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.

[0141] The processor 501 controls the overall operation of the electronic device 500 to complete all or part of the steps in the above-described test paper image annotation method. The memory 502 stores various types of data to support the operation of the electronic device 500. This data may include, for example, instructions for any application or method operating on the electronic device 500, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 503 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 502 or transmitted via communication component 505. The audio component also includes at least one speaker for outputting audio signals. I / O interface 504 provides an interface between processor 501 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 505 is used for wired or wireless communication between the electronic device 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 505 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0142] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described test paper image annotation method.

[0143] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the above-described test paper image annotation method. For example, the computer-readable storage medium may be the memory 502 including the program instructions, which may be executed by the processor 501 of the electronic device 500 to complete the above-described test paper image annotation method.

[0144] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a programmable device, the computer program having a code portion for performing the above-described test paper image annotation method when executed by the programmable device.

[0145] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0146] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0147] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A method for annotating test paper images, characterized in that, include: Multiple preset nodes are identified and displayed in a first sub-window; each preset node corresponds to a task type. In response to selecting a preset node from the first sub-window and dragging it to the second sub-window, a task configuration window pops up to configure the preset node based on the selected node. Retrieve the task information and task executor configured in the task configuration window, and create a work node corresponding to the task type of the preset node; After creating multiple work nodes, the multiple work nodes are connected sequentially according to their corresponding task types, and a workflow is generated in response to the confirmation of creation. The workflow includes at least one annotation work node generated according to the preset node of the annotation type, and the task corresponding to each annotation work node is to annotate the test paper image accordingly. The work nodes in the workflow are sequentially assigned to the corresponding task executors. Once all work nodes in the workflow are completed, the data from each annotation work node are merged to obtain the final annotation data for the exam paper image; The process of merging the data from each annotation work node to obtain the final annotation data for the exam paper image includes: By merging the data from each annotation work node, multiple annotation boxes and corresponding annotation labels are obtained for the test paper image. For any two of the plurality of annotation boxes, calculate the ratio of the intersection area between the two annotation boxes to the area of ​​the annotation box with the smaller area, and obtain the intersection-smaller ratio between the two annotation boxes; When the intersection ratio is greater than a preset threshold, the label box with the smaller area of ​​the two label boxes is determined to be the child label box, and the label box with the larger area is determined to be the parent label box. The hierarchical relationship between the multiple annotation boxes is determined based on the parent-child relationship between any two annotation boxes; Based on the multiple annotation boxes, the annotation labels corresponding to each annotation box, and the hierarchical relationship between the multiple annotation boxes, multi-level annotation data corresponding to the test paper image is generated.

2. The method according to claim 1, characterized in that, Each task executor is visible to their own work nodes in the workflow, but not to work nodes where they are not task executors.

3. The method according to claim 1, characterized in that, The workflow also includes audit work nodes generated from preset nodes based on audit type, quality inspection work nodes generated from preset nodes based on quality inspection type, and acceptance work nodes generated from preset nodes based on acceptance type. The task corresponding to the review work node is to review the labeled data of the previous work node; the task corresponding to the quality inspection work node is to conduct random quality inspection on the labeled data that has been reviewed and approved; and the task corresponding to the acceptance work node is to accept the labeled data that has been reviewed and approved if the random quality inspection is passed.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: After generating the workflow, save the workflow; When a new workflow is created, in response to selecting a workflow from at least one saved historical workflow as the base workflow, the base workflow is displayed in a second sub-window; In response to selecting any work node in the basic workflow, a task configuration window pops up to configure based on the selected work node; Receive new task information and / or new task executors entered in the task configuration window, and update the work node; In response to the confirmation of the creation operation, a new workflow is generated.

5. A test paper image annotation device, characterized in that, include: The preset node determination module is used to determine a variety of preset nodes and display the variety of preset nodes in a first sub-window; wherein, each preset node corresponds to a task type; The node drag-and-drop module is used to pop up a task configuration window based on the selected preset node in response to selecting a preset node from the first sub-window and dragging it to the second sub-window. The work node creation module is used to obtain the task information and task executor configured in the task configuration window, and create a work node corresponding to the task type of the preset node. The workflow generation module is used to connect multiple work nodes sequentially according to their corresponding task types after creating multiple work nodes, and generate a workflow in response to the confirmation of creation operation; wherein, the workflow includes at least one annotation work node generated according to the preset node of the annotation type, and the task corresponding to each annotation work node is to annotate the test paper image accordingly; The work node allocation module is used to sequentially allocate the work nodes in the workflow to the corresponding task executors. The annotation data acquisition module is used to merge the data of each annotation work node after all work nodes in the workflow are completed, so as to obtain the final annotation data of the test paper image. The labeled data acquisition module is specifically used for: By merging the data from each annotation work node, multiple annotation boxes and corresponding annotation labels are obtained for the test paper image. For any two of the plurality of annotation boxes, calculate the ratio of the intersection area between the two annotation boxes to the area of ​​the annotation box with the smaller area, and obtain the intersection-smaller ratio between the two annotation boxes; When the intersection ratio is greater than a preset threshold, the label box with the smaller area of ​​the two label boxes is determined to be the child label box, and the label box with the larger area is determined to be the parent label box. The hierarchical relationship between the multiple annotation boxes is determined based on the parent-child relationship between any two annotation boxes; Based on the multiple annotation boxes, the annotation labels corresponding to each annotation box, and the hierarchical relationship between the multiple annotation boxes, multi-level annotation data corresponding to the test paper image is generated.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-4.

7. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Data marking method and device, electronic device and storage medium

    CN108984490A

  • Method, system and device for managing image frame and medium

    CN111814885A

  • Data labeling method and device

    CN114090534A