Test question image processing method, device and equipment and computer readable storage medium

CN120032382APending Publication Date: 2025-05-23TENCENT DIGITAL TIANJIN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311557570.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is less efficient when correcting homework or test papers, mainly due to the need for manual operations.

Method used

Provide a test question image processing method, displays the test question image by displaying data marking page, and displays the marking box surrounding the target area in the image, including the answer area, the question stem area, the question area and the test question area. This method allows the user to mark these areas, thereby generating labeled data for the test questions.

Benefits of technology

The efficiency of marking and correcting test questions or homework has been improved, and the automatic correction of test questions or homework has been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032382A_ABST
    Figure CN120032382A_ABST
Patent Text Reader

Abstract

The invention discloses a test question image processing method, device and equipment and a computer readable storage medium. The embodiment of the invention can be applied to scenes such as cloud technology, smart traffic, aided driving and education. A data annotation page can be displayed, and the uploaded test question image is displayed on the data annotation page; in response to a marking operation for a target area in the test question image, displaying a marking box surrounding the target area in the test question image, the target area including an answering area, a question stem area, a question area and a test question area, the types of the mark boxes comprise an answer box surrounding the answer area, a question stem box surrounding the question stem area, a question box surrounding the question area and a test question box surrounding the test question area; and displaying target contents associated with the mark box in a mark analysis area in the data mark page, wherein the target contents are used for generating mark data of the test questions. Therefore, the annotation data generated by the target content associated with the tag box can be used for automatically correcting the test questions or the homework, and the correction efficiency of the homework or the test questions is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a test question image processing method, apparatus, device and computer-readable storage medium. Background Art

[0002] Image and text recognition technology has been widely used in various fields. For example, it is used in the fields of education for homework and test grading. Specifically, images of homework and test papers are obtained by taking photos or scanning, and the image and text information of the homework and test papers is extracted through image and text recognition technology, and then the homework and test papers are graded based on the extracted image and text information.

[0003] During the research and practice of related technologies, the inventors of the present application found that when grading homework or test papers in related technologies, the homework or test papers are generally graded manually, which is inefficient. Summary of the Invention

[0004] The embodiments of the present application provide a test question image processing method, apparatus, device, and computer-readable storage medium, which can be applied to the marking of test questions or assignments of any question type and structure, thereby improving the marking efficiency of test questions or assignments, and using the marked data for automatic grading of test questions or assignments, thereby improving the grading efficiency of test questions or assignments.

[0005] The present application provides a method for processing test image data, including:

[0006] Displaying a data annotation page and displaying the uploaded test question image on the data annotation page;

[0007] In response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes an answer area, a question stem area, a question area, and a test question area consisting of at least the answer area and the question stem area, and the types of the marking frame include an answer frame surrounding the answer area, a question stem frame surrounding the question stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area;

[0008] Displaying target content associated with the mark box in the annotation analysis area of ​​the data annotation page, wherein the target content is used to generate annotation data of the test question;

[0009] Among them, when the marking box is the answer box, the associated target content is the answer information; when the marking box is the question stem box, the associated target content is the question stem information; when the marking box is the question frame, the associated target content is the question content.

[0010] Accordingly, an embodiment of the present application further provides a test question image processing device, comprising:

[0011] A display unit, configured to display a data annotation page and present the uploaded test question image on the data annotation page;

[0012] a marking unit, configured to, in response to a marking operation on a target area in the test question image, display a marking frame surrounding the target area in the test question image, wherein the target area includes an answer area, a question stem area, a question area, and a test question area consisting of at least the answer area and the question stem area, and the types of the marking frame include an answer frame surrounding the answer area, a question stem frame surrounding the question stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area;

[0013] A display unit, configured to display target content associated with the mark box in the annotation analysis area of ​​the data annotation page, wherein the target content is used to generate annotation data for the test question;

[0014] Among them, when the marking box is the answer box, the associated target content is the answer information; when the marking box is the question stem box, the associated target content is the question stem information; when the marking box is the question frame, the associated target content is the question content.

[0015] In some embodiments, the data annotation page further displays an uploaded answer image, and the marking unit is further configured to:

[0016] In response to a marking operation on an answer area where each piece of answer information is located in the answer image, displaying an answer frame surrounding the answer area in the answer image;

[0017] The display unit is further configured to display the answer information associated with the answer box in the annotation analysis area of ​​the data annotation page.

[0018] In some embodiments, the test question image device further includes:

[0019] a determining unit, configured to determine, based on position information of each mark frame in the question image, a subordinate relationship between the answer frame and the sub-question frame, a subordinate relationship between the question stem frame and the sub-question frame, a subordinate relationship between the question frame and the question frame, and a subordinate relationship between the sub-question frame and the question frame;

[0020] a sorting unit, configured to sort the multiple marked boxes belonging to the same level based on the subordinate relationship and the position information of each marked box to obtain a marked box sorting relationship;

[0021] A generating unit is used to combine the answer information associated with each answer area in the test question image, the question stem content of each question stem area, and the question content of each question area according to the subordinate relationship and the mark box sorting relationship to generate annotation data for the test question.

[0022] In some embodiments, the determining unit is further configured to:

[0023] For the answer box and the question box to be detected, calculating the overlapping area of ​​the answer box and the question box according to the position information of the answer box and the question box in the question image;

[0024] Calculating the overlap ratio between the answer frame and the question frame according to the overlap area;

[0025] The subordinate relationship between the answer box and the question box is determined according to the overlapping ratio.

[0026] In some embodiments, the sorting unit is further configured to:

[0027] Based on the subordinate relationship, a target markup frame set having the same parent object at each level is classified from the answer frame, the question frame, and the test frame, where the parent object is any of the test frame;

[0028] For each target mark frame set, determining at least one target mark frame subset based on position information of each mark frame in the target mark frame set, each target mark frame subset including at least one target mark frame in the same row in the test question image;

[0029] For each target mark frame set, multiple target mark frame subsets are sorted, and the target mark frames within each target mark frame subset are sorted to obtain a mark frame sorting relationship of the test question image.

[0030] In some embodiments, the sorting unit is further configured to:

[0031] For each target mark frame set, determining the upper boundary position and the lower boundary position of each target mark frame according to the position information of each mark frame;

[0032] According to the first condition, at least one target mark frame in the same row in the test question image is respectively classified from the target mark frame set, and the set of at least one target mark frame in the same row is used as a target mark frame subset;

[0033] The first condition is for the first target mark box and the second target mark box belonging to the same row position, wherein the upper boundary position of the first target mark box is higher than the lower boundary position of the second target mark box, and the lower boundary position of the first target mark box is lower than the upper boundary position of the second target mark box.

[0034] In some embodiments, the sorting unit is further configured to:

[0035] For each target mark frame set, determining the overlapping length of the longitudinal boundaries between any two target mark frames in the target mark frame set according to the position information of each mark frame;

[0036] determining a longitudinal overlap coefficient according to the overlap length and a reference longitudinal side length; the reference longitudinal side length refers to the longitudinal side length of the target mark frame with the smaller longitudinal boundary length between the two target mark frames;

[0037] If the longitudinal overlap coefficient is greater than a preset threshold, it is determined that the two target mark frames are in the same row in the test question image; and a set of at least one target mark frame in the same row in the test question image is taken as a target mark frame subset.

[0038] In some embodiments, the sorting unit is further configured to:

[0039] For each target mark frame set, sorting the target mark frame subsets in order from upper row to lower row according to the position of the row where the target mark frame in the target mark frame subset is located, to obtain a row sorting relationship;

[0040] According to the position information of each target mark frame in each target mark frame subset, the multiple target mark frames are sorted in order from left to right to obtain a horizontal sorting relationship between the multiple target mark frames in the same row position;

[0041] According to the row sorting relationship and the horizontal sorting relationship, the sorting relationship between the target mark frames in the corresponding target mark frame set is determined to obtain the mark frame sorting relationship of the test question image.

[0042] In some embodiments, the sorting unit is further configured to:

[0043] Adding a start mark to the target mark frame on the leftmost side of the test question image in the target mark frame subset, and adding an end mark to the target mark frame on the rightmost side of the test question image;

[0044] Starting from the target mark frame containing the start mark, the target mark frames in the target mark frame subset are numbered incrementally from left to right until the target mark frame containing the end mark is numbered, thereby obtaining a horizontal sorting relationship between multiple target mark frames belonging to the same row position.

[0045] In some embodiments, the test question image processing apparatus further includes an adding unit configured to:

[0046] Identifying answer boxes in the text annotation data;

[0047] Performing semantic analysis on the answer information in each answer box to determine key information in the answer information and the answer category corresponding to each key information;

[0048] A key answer mark box is marked at the key information in each answer box, and the corresponding answer category is marked for each key answer mark box.

[0049] In some embodiments, the test question image processing apparatus further includes an adjustment unit configured to:

[0050] Determining the target content whose subordinate relationship is to be adjusted and the test question area whose content is to be expanded in the text annotation data;

[0051] The position of the mark frame corresponding to the target content is adjusted to the question frame corresponding to the question area of ​​the content to be expanded in the text annotation data, and updated target text annotation data is generated.

[0052] In some embodiments, the generating unit is further configured to:

[0053] Based on the subordinate relationship and the order relationship of the mark boxes, constructing initial hierarchical structure information including the answer box, the question box and the test box;

[0054] If it is detected that the initial hierarchical structure information includes a target branch structure with an insufficient number of levels, the target branch structure in the initial hierarchical structure information is padded to obtain padded target hierarchical structure information, where the target branch structure is a branch structure with a number of levels less than a preset level threshold;

[0055] If it is detected that the initial hierarchical structure information does not include the target branch structure, determining the initial hierarchical structure information as the target hierarchical structure information;

[0056] According to the target hierarchical structure information, the answer information associated with each answer area in the test question image, the question stem content of each question stem area, and the question content of each question area are combined to generate annotation data for the test question.

[0057] In addition, an embodiment of the present application also provides a computer device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in any one of the test question image processing methods provided in the embodiment of the present application.

[0058] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any test question image processing method provided in the embodiment of the present application.

[0059] In addition, an embodiment of the present application also provides a computer program product, including computer instructions, which, when executed, implement the steps of any test question image processing method provided in the embodiment of the present application.

[0060] An embodiment of the present application provides a data annotation page of a data annotation platform, which can display a test image of any test question on the data annotation page, and divide the test question image into target areas. The target areas include a minimum-level answer area, a question stem area associated with the answer area, and a minimum-level test question area consisting of a combination of the answer area and the question stem area. In addition, each minimum-level test question area is associated with a question area, and one or more minimum-level test question areas and commonly associated question areas are combined to obtain an increasing-level test question area, and so on. Then, for the marking operation of each target area, a marking box surrounding the target area is displayed in the test question image. Finally, the target content associated with each marking box is displayed in the annotation analysis area in the data annotation page. Specifically, when the marking box is an answer box surrounding the answer area, the target content is answer information; when the marking box is a question stem box surrounding the question stem area, the target content is question stem information; when the marking box is a question box, the target content is question content. The above target content is used to generate annotation data for the test question. In this way, it can be applied to the marking process of test questions or homeworks of any question type structure, so as to automatically generate marking data for test questions or homeworks according to the target content associated with the marking box, thereby improving the marking efficiency of test questions or homeworks, so as to realize automatic grading of test questions or homeworks based on the marked data, thereby improving the grading efficiency of homeworks or test questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0062] Figure 1 This is a schematic diagram of a scenario of a test question image processing system provided by an embodiment of the present application;

[0063] Figure 2 This is a schematic diagram of the steps of the test question image processing method provided in an embodiment of the present application;

[0064] Figure 3 This is an example diagram of a marking scenario provided by an embodiment of the present application;

[0065] Figure 4 This is another example diagram of the annotation scene provided in the embodiment of the present application;

[0066] Figure 5This is a schematic diagram of the structure of the annotation platform provided in the embodiment of the present application;

[0067] Figure 6 This is a schematic diagram of the structure of text annotation data provided in an embodiment of the present application;

[0068] Figure 7 This is a schematic diagram of an information supplement scenario provided by an embodiment of the present application;

[0069] Figure 8 This is another step flow chart of the test question image processing method provided in an embodiment of the present application;

[0070] Figure 9 This is a flow chart of the marking phase provided in an embodiment of the present application;

[0071] Figure 10 This is a schematic diagram of the process flow of the reorganization stage provided in the embodiment of the present application;

[0072] Figure 11 Schematic diagram of the structure of the test question image processing device provided in an embodiment of the present application;

[0073] Figure 12 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.

[0075] Some processes described in the specification, claims, and accompanying figures include multiple steps that appear in a specific order. However, it should be understood that these steps may be performed in a different order or in parallel. Step numbers are used solely to distinguish between different steps and do not inherently indicate an order of execution. Furthermore, terms such as "first" and "second" are used to distinguish similar items and do not necessarily describe a specific order or precedence.

[0076] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0077] The embodiments of the present application provide a test question image processing method, apparatus, device and computer-readable storage medium. Specifically, the embodiments of the present application will be described from the dimension of a test question image processing apparatus, which can be specifically integrated into a computer device, which can be a server or a user terminal or other device. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Among them, the user terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart home appliance, car terminal, intelligent voice interaction device, aircraft, etc., but is not limited to this.

[0078] It is understandable that in the specific implementation of this application, data related to user information, user usage records, user status, etc. (for example, information related to the "test questions" below, such as "test question files", "test question answers", etc.) are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0079] It should be noted that the test image processing method provided in the embodiments of the present application is applicable to scenarios where advertisements are inserted into live streaming data. These scenarios are not limited to being implemented through cloud services, cloud technologies, big data, or a combination thereof. The following embodiments will be used to illustrate this:

[0080] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0081] Cloud technology is a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identifier, requiring transmission to backend systems for logical processing. Different levels of data will be processed separately, and data from various industries will require a strong system backend, which can only be achieved through cloud computing.

[0082] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a per-use basis.

[0083] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0084] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging tools. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0085] The test image processing service involved in the embodiments of this application can be implemented through cloud computing. This is specifically described through the following embodiments:

[0086] For example, see Figure 1 , is a scenario diagram of the test question image processing system provided in an embodiment of the present application. The devices in this scenario system may include a server and / or a terminal; when the devices in the system only include a server or a terminal, the server or the terminal may directly execute the test question image processing method of the embodiment of the present application; when the system is a combination of a terminal and a server, the terminal and the server may cooperate with each other to execute the test question image processing method of the embodiment of the present application.

[0087] Specifically, when the devices in the system only include servers or terminals, the server or terminal can display a data annotation page and display the uploaded test question image on the data annotation page; in response to the marking operation on the target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes the answer area, the question area, the stem area and the test question area consisting of at least the answer area and the stem area, and the types of marking frames include the answer box surrounding the answer area, the stem box surrounding the stem area, the question box surrounding the question area and the test question box surrounding the test question area; the target content associated with the marking frame is displayed in the annotation analysis area in the data annotation page, and the target content is used to generate annotation data for the test question; wherein, when the marking frame is an answer frame, the corresponding target content is answer information, when the marking frame is a stem frame, the corresponding target content is stem information, and when the marking frame is a question frame, the corresponding target content is question content.

[0088] For another example, taking a system composed of a terminal and a server as an example, a communication connection is established between the terminal and the server. A target client is installed on the terminal, and the terminal can send a test question image to the server through the target client. The server can display a data annotation page and display the uploaded test question image on the data annotation page; in response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes an answer area, a stem area, a question area, and a test question area consisting of at least an answer area and a stem area, and the types of marking frames include an answer frame surrounding the answer area, a stem frame surrounding the stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area; the target content associated with the marking frame is displayed in the annotation analysis area of ​​the data annotation page, and the target content is used to generate annotation data for the test question; wherein, when the marking frame is an answer frame, the corresponding target content is answer information, when the marking frame is a stem frame, the corresponding target content is stem information, and when the marking frame is a question frame, the corresponding target content is question content.

[0089] For example, the test image processing method of the embodiment of the present application can be applied to test image processing scenarios in the fields of education, government agencies, etc. The test questions can be standard teaching materials for teaching, examinations, exercises, etc. For example, taking the test image processing scenario in teaching materials as an example, combined with Figure 1To describe the scenario, the terminal or server can mark the target areas (such as the answer area, stem area, question area, test question area, etc.) contained in the test question image of the teaching material in order from small to large, such as the answer area, stem area and test question area, on the data annotation page, to obtain the answer box corresponding to the answer area, the stem box corresponding to the stem area, and the test question box corresponding to the test question area, etc., wherein the test question area includes the question area where the corresponding question is located and the answer area and stem area under the question. In addition, there can also be higher-level test question areas. It is assumed that the answer area, the stem area of ​​the stem to which the answer area belongs, and the question area of ​​the question to which the stem belongs are regarded as a target combination. When the question to which the stem belongs also has a larger question of a higher level, the target combination and the larger question can also be understood as a test question area; for higher-level test question areas, this can be deduced by analogy and will not be elaborated here. Furthermore, based on the position information of each marking box in the test question image, the subordinate relationship between the answer box and the test question box, the subordinate relationship between the question stem box and the test question box, and the subordinate relationship between test question frames at different levels are determined, so that multiple marking boxes belonging to the same level can be sorted according to the subordinate relationship and the position information of the marking box to obtain the marking box sorting relationship; finally, according to the subordinate relationship and the marking box sorting relationship, the answer information associated with each answer area in the test question image, the question stem content corresponding to each question stem, and the question content corresponding to each question are combined to generate annotation data for the test question.

[0090] Among them, the labeled data can be used in subsequent artificial intelligence algorithms. Specifically, it can be used as model training data, such as as sample labels output by the model, and the test image is input into the preset model so that the preset model outputs a predicted label for the test image, determines the difference between the sample label and the predicted label, and constructs the prediction loss. Then, the parameters of the preset model are adjusted according to the prediction loss, and iterative training is performed until the preset conditions are met. The preset conditions can be that the number of training times reaches a certain number, the prediction loss is minimized, or the predicted label output by the model is consistent with the sample label. In this way, the trained target model is obtained, and the target model can be used in intelligent automatic correction scenarios. It should be noted that the classification layer of the model can be set according to the application scenario. For example, when the scenario requires the model to output a score, a softmax or sigmoid activation function can be set. A softmax or tanh activation function can also be selected for the correctness or error judgment scenario of each answer information. The specific conditions can be determined according to the actual situation and are not limited here. For example, after obtaining the trained target model, the image of the test question to be corrected corresponding to the corresponding teaching material or test paper can be obtained by photographing or scanning, and the image of the test question to be corrected can be input into the target model to output the score, accuracy, correct and incorrect options, correction comments, etc. for the image of the test question to be corrected. In this way, intelligent automatic correction of answer documents of the test question type can be achieved, improving efficiency and reliability. The above is only an example and does not constitute a specific limitation for the implementation of this application.

[0091] It should be noted that the above are only examples and can also be applied to other test image processing scenarios, which will not be described in detail here.

[0092] For ease of understanding, each step of the test question image processing method will be described in detail below. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0093] In the embodiment of the present application, the description will be made from the perspective of a test question image processing device, which can be integrated into a computer device, such as a terminal or a server. Figure 2 , Figure 2 This is a schematic flow chart of the steps of the test image processing method provided in an embodiment of the present application. In this embodiment of the present application, the test image processing device is specifically integrated into a server as an example. When the processor on the server executes the program instructions corresponding to the test image processing method, the specific process is as follows:

[0094] 101. Display the data annotation page, and display the uploaded test question image on the data annotation page.

[0095] In an embodiment of the present application, in order to realize intelligent automatic grading and review of examination question files such as examinations, tests, exercises, etc. in any field (such as education, training, and government agencies), it is necessary to first obtain text annotation data of the standard answers to the examination questions for training of the neural network model, so that the trained neural network model can be used to subsequently grade and review the user's (such as students, examination participants, etc.) version of the examination question answer information to realize automated grading of examination question files.

[0096] Therefore, it is necessary to obtain the test question image and annotate the relevant content of the test question on the data annotation page, so that the text annotation data of the standard answer to the test question can be generated based on the annotated information.

[0097] Among them, the data annotation page can be the interface of the data annotation service provided by the target platform. Specifically, the operating logic of the data annotation service can be located on the server, and the terminal can be connected to the server through an interface to call the data annotation service, thereby displaying the data annotation page on the terminal interface. In addition, an application (client) can be installed on the terminal. The application can be run in a stand-alone or networked form, and the user can open the application to display the data annotation page on the terminal interface. Furthermore, relevant personnel can perform marking operations on the test image uploaded and displayed on the data annotation page.

[0098] Among them, the data annotation page is not limited to including the question column, answer column and annotation analysis area. The question column can be understood as the title and stem column of the test question. The answer column is for displaying the answer information of the user's answer or the standard answer information of the textbook. The annotation analysis area displays the relevant information corresponding to each mark box.

[0099] In some embodiments, after the data annotation page is displayed, relevant personnel can also select, open, and upload a test question image. The test question image may include prompt information (stem content and question content) for prompting answers, as well as answer information, which may be the standard answer information of the test question. It should be noted that when uploading the test question image, the prompt information image and the answer information image can be uploaded separately and independently, which is not limited here. It should be noted that the prompt information (stem content and question content) is displayed in the corresponding area under the test question column in the data annotation page, and the answer information is displayed in the corresponding area under the answer column. For examples of displaying test question images, please refer to Figure 5 , Figure 5 This is only an example and is not intended to limit the specific implementation.

[0100] 102. In response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image.

[0101] In an embodiment of the present application, the test question image is displayed in the data annotation page, and a marking operation can be performed on the test question image. Specifically, the marking operation can be a marking operation for any target area in the test question image, so that the computer responds to the instruction of the marking operation and displays a marking box for the target area in the test question image to achieve marking.

[0102] The target area includes the answer area, the question area, the question area, and a test question area consisting of at least the answer area and the question area. The types of marking frames include an answer frame surrounding the answer area, a question frame surrounding the question area, a question frame surrounding the question area, and a test question frame surrounding the test question area. It should be noted that the test question area can be divided according to the level. For example, the first-level test question area is composed of the corresponding answer area and the question area, the second-level test question area is composed of the corresponding one or more first-level test question areas and the question area, the third-level test question area is composed of the corresponding one or more second-level test question areas and the question area, and so on.

[0103] The marking operation can be a frame selection operation on the target area, which mainly selects the content of the target area by marking the frame. Figure 5 As shown, the user can perform a frame selection operation on the target area where the question "1. Word accumulation" is located, so as to frame and mark the question "1. Word accumulation" through the marking box (which can be understood as a level 3 question box); for example, the user can perform a frame selection operation on the target area where the question "1. Read pinyin and write words" is located, so as to frame and mark the question "1. Read pinyin and write words" through the marking box (which can be understood as a level 2 question box) to obtain the question box for "1. Read pinyin and write words"; for example, the marking operation is performed on the target area where the question stem "mei miao" is located, so as to frame and mark "mei miao" through the marking box (which can be understood as a level 1 question stem box) to obtain the question stem box for "mei miao"; in addition, it also includes a marking operation on the answer area, such as performing a marking operation on the answer area below "mei miao", so as to frame and mark the answer blank below "mei miao" through the marking box (which can be understood as the answer box) to obtain the answer box for "mei miao".

[0104] Regarding the process of marking the target area in the test image, the specific implementation process is as follows:

[0105] In an embodiment of the present application, in order to realize the intelligent automatic correction and review of test files such as examinations, tests, exercises, etc. in any field (such as education, training, and government agencies), it is necessary to first obtain the text annotation data of the standard answers to the test questions for use in the training of the neural network model, so that the trained neural network model can be used to correct and review the user's (such as students, test takers, etc.) version of the test answer information to realize the automatic correction of the test question files. In this regard, in order to obtain the text annotation data of the standard answers to the test questions, after obtaining the test question image corresponding to the test question file, it is necessary to frame and mark the relevant information in the test question image to obtain the marking box of each type of information. Among them, the types of the marking box include the answer box corresponding to the answer area, the question stem box corresponding to the question stem area, and the question box corresponding to the test question area.

[0106] In some embodiments, the answer area in the test image, the stem area of ​​each answer area, the question area of ​​the question to which the stem belongs, and the test question area corresponding to the question to which the stem belongs are marked to obtain multiple marking boxes.

[0107] Among them, the test question image can be a test question image of a page corresponding to a file such as a textbook, test paper, or exercise book, which contains information such as the question, question stem, and standard answer. In the test question image, multiple levels of test question areas can be included. The lowest level test question area can be composed of a question stem area and an answer area, and the upper level test question area of ​​the lowest level test question area can include the question area where the corresponding question is located and the lower level test question area (answer area and question stem area) under the question. In addition, there can also be higher level test question areas. Assuming that the answer area, the question stem area of ​​the question stem to which the answer area belongs, and the question area of ​​the question stem to which the question stem belongs are taken as a target combination, when the target combination also has a higher level large question, then the target combination and the large question can also be understood as a test question area; for higher level test question areas, this can be deduced by analogy, and will not be elaborated here.

[0108] For example, for ease of understanding, the description is made in the order of the levels from the question to the stem and the answer area. Assuming that the test image contains multiple questions, the test range area covered by each question can be understood as the largest test area or the highest level test area; for each highest level test area, it can contain multiple mid-level questions, and the test range area covered by each mid-level question can be understood as a medium-sized test area or a mid-level test area. In addition, each highest level test area can also directly contain a low-level test area; for each mid-level test area, it can contain multiple low-level questions, and the test range area covered by each low-level question can be understood as a small test area or a low-level test area. Each low-level test area contains a stem and an answer area (i.e., an answer filling area). It should be noted that for each mid-level test area, it can also contain multiple test areas with decreasing levels, until it is located in the upper level test area of ​​the low-level test area. For example, assuming that the low-level test question area is the level 1 test question area and the high-level test question area is the level 5 test question area, it can be understood as including 3 middle-level test question areas, such as the level 4 test question area, the level 3 test question area and the level 2 test question area. The above is only an example and is not a limiting method for implementing this application.

[0109] The stem area corresponds to the stem content, which describes the question content and is the information basis for answering in the answer area. Each stem area has an independent answer area, which can be represented by a "line blank" (such as "_______"), brackets, rectangular boxes or "blank areas". For example, the test image contains multiple large questions, each large question has a corresponding question, and the question contains one or more question areas of the next level of test questions, see Figure 3 As shown, the dotted box represents the question area, and the "1. Listen and choose what you hear" in the box can be understood as the title of the first question. The question range area covered by the question of this question is the outermost bold solid wireframe in the figure. In this scenario, the question of the big question can be understood as the question of the level 2 question; the question area of ​​the four lowest-level questions directly included in the big question can be understood as the question area of ​​the level 1 question. Each question area contains the question stem content in the question stem box and the answer information in the answer box. For example, the first question "1. Can you____a bike? A. ride B. bird" is the question stem content of the level 1 question, which serves as the information basis for the user when answering, and "(A)" is the answer information of the level 1 question. It can be understood that the answer information is attached to the question stem content. The above is only an example and is not a limiting method for implementing this application.

[0110] The types of the mark boxes may include answer boxes corresponding to the answer area, question boxes corresponding to the question area, and question boxes corresponding to the question area. Figure 3As shown, the answer area is "(A)", "(B)", "(B)", "(B)", "(A)", and the answer boxes corresponding to the answer area are boxes for selecting "(A)", "(B)", "(B)", "(B)", "(A)" respectively. The stem area is the area containing the stem content of "1. Can you____a bike? A. ride B. bird", and the stem box selects the stem content of "1. Can you____a bike? A. ride B. bird". The question area of ​​the lowest level entity is composed of the answer area and the stem area. The question box of the lowest level question area selects "(A)" and "1. Can you____a bike? A. ride B. bird". B.bird"; further, the lowest-level question frame also includes the question frame of the upper level to which it belongs, that is, the question frame of the Level 2 question. The question area of ​​the Level 2 question includes the question area and 5 Level 1 question areas. The question area includes the question content of "I. Listen and choose what you hear". Each Level 1 question area includes an answer area and a question stem area. The question frame corresponding to the question area of ​​the Level 2 question selects the question area and the 5 Level 1 question areas. It should be noted that for the "low-level question area", "middle-level question area" and "high-level question area" introduced above, if the question image contains these "question areas", they need to be marked one by one.

[0111] In some embodiments, when marking each area in the test question image, the target area contained in the test question image is mainly marked in order from low to high hierarchical order of the answer area, the stem area, and the test question area to obtain multiple marking boxes, wherein the type of the target area can be one of the answer area, the stem area, the question area, and the test question area, and the test question area is at least a first-level (low-level) test question area. Specifically, the marking process of the target area contained in the test question image is as follows: detect the answer area in the test question image, and mark each answer area in the test question image to obtain the answer frame corresponding to the answer area; detect the stem area of ​​the question stem to which each answer area in the test question image belongs, and mark each stem area in the test question image to obtain the stem frame corresponding to the stem area; detect each first-level test question area composed of the answer area and the stem area to which it belongs, and mark each first-level test question area in the test question image to obtain the first-level test question frame corresponding to the first-level test question area; detect the question area associated with the first-level test question area in the test question image, mark the question area, and obtain the question frame corresponding to each question area; detect the second-level test question area composed of at least one first-level test question area and the question area to which it belongs, and mark each second-level test question area in the test question image respectively to obtain the second-level test question frame corresponding to each second-level test question area; for test question areas of higher levels, mark the higher-level test frames. It should be noted that each marking box can carry a label or number. Specifically, when marking various areas in the test image, the marking boxes belonging to different levels and types can be numbered separately. For example, if the answer areas all belong to the same type or the same level, all the answer areas in the test image can be numbered uniformly (without considering whether they belong to the same question). Similarly, all the question stem areas in the test image can be numbered uniformly, and all the test question areas belonging to the same level in the test image can be numbered uniformly. The labels of all the above marking boxes can be used to refer to and distinguish the explanatory information of each marking box in the annotation explanation.

[0112] For example, it is assumed that the question hierarchy in the question image includes first-level questions, second-level questions, and third-level questions. The question area of ​​the first-level question consists of the answer area and the question stem area, see Figure 4As shown, each first-level question frame corresponds to one of the first-level question areas. For example, "4. Can you sing? (Answer negatively) No, I / We can't." is marked by a first-level question frame, representing a first-level question area. The first-level question frame contains a first-level question stem frame and a first-level answer frame. The first-level question stem frame marks the question stem content of "4. Can you sing? (Answer negatively)", and the first-level answer frame marks "No, I / We can't."; each first-level question frame is numbered 12, 13, 14, and 15 respectively. It can be understood that before this, the first-level question frames under other questions may be numbered, such as numbers 1 to 11 to mark the first-level question frames under other questions; similarly, all answer boxes in the question image are numbered uniformly, including 4 answer boxes, numbered 26, 27, 28, and 29 respectively; in addition, all stem frames in the question image are numbered uniformly, including 4 stem frames, numbered 31, 32, 33, and 34 respectively. Among them, the question area of ​​the second-level question is composed of 4 first-level question areas and the question area of ​​the question to which it belongs. The question content of "2. Complete the following questions as required" contained in the question area can be understood as the question stem to which each first-level question area belongs, the question box corresponding to the second-level question area is marked, the question area of ​​the second-level question, and the 4 first-level question areas. Among them, the question area of ​​the second-level test questions is composed of one or more second-level test question areas and the question area of ​​the corresponding questions. Figure 4 The example of the marking frame and range of the third-level test question area is not shown in the figure. You can refer to the relationship between the marking frames of the first-level test question area and the second-level test question area to determine it. In addition, after marking each area in the test question image, each marked frame has a corresponding label. The label is used to indicate the explanation information of each marked frame in the annotation analysis area. It can be understood that for any answer frame, question frame, title frame, and test question frame, there is a corresponding box annotation explanation in the annotation analysis area. For example, Figure 4 Take the question box (marked box) of the level 2 question in the question image as an example. The question box is labeled "2" in all the second-level question boxes in the question image, and there is corresponding explanation information in the marked analysis area. The explanation information specifically includes the level (level) of the marked box, the box label, the question type in the box, the answer method (requirements) and other information, such as "level 2 question box, box number 2, the question type is a type with a standard answer, and the answer method is to manually fill in the answer on a specific line"; in addition, other question boxes, answer boxes, stem boxes, and question boxes all have corresponding explanation information in the marked analysis area, which will not be detailed here. It should be noted that after marking each area in the question image, the marked question image can be referred to Figure 5 The above is only an example and is not intended to be a specific limitation for implementing this application.

[0113] Through the above method, after obtaining the test question image corresponding to the test question file, it is necessary to select and mark the relevant information in the test question image to obtain marked boxes for various types of information, so as to subsequently determine the subordinate relationship and sorting relationship of each marked box, thereby generating the labeled data of the test question.

[0114] 103. The target content associated with the mark box is displayed in the mark analysis area of ​​the data mark page.

[0115] In an embodiment of the present application, the data annotation page also includes an annotation analysis area, which is used to display relevant information of each marking box, and the relevant information can specifically be the attribute information of the marking box and the associated target content. For example, the annotation analysis area contains multiple annotation explanation items, each annotation explanation item corresponds to a marking box, and the annotation explanation item contains explanation information of the corresponding marking box, such as the type of marking box, associated content information, information display type, etc. Exemplarily, taking the marking box type of answer box as an example, the annotation explanation item of the answer box will be displayed below the annotation analysis area, and the annotation explanation item can include the type of marking box (i.e., "answer box"), the hierarchical depth of the answer box (such as "level 1"), and the answer information associated with the answer box (such as "wonderful") displayed in the information filling area.

[0116] The annotation parsing area is used to explain each marked box. The annotation parsing area contains multiple annotation parsing items, each of which represents the explanation of a marked box. Specifically, each annotation parsing item includes multiple information items, such as hierarchical information, title format, text type, content filling, etc. For example, see Figure 5 , the rightmost side of the data annotation page is the annotation analysis area, in which four marked box annotation analysis items are currently displayed. Among them, the second one is the annotation analysis item of the question box, which includes the hierarchical information of "Level 2 question", the option button of information format (text or image), the text type (formula and symbol, special symbol, preview, etc.), the fill-in content item ("1. Read pinyin and write words"), etc.; for example, the third one is the annotation analysis item of the question stem box, which includes the hierarchical information of "Level 1 question stem", the option button of information format (text or image), the text type (formula and symbol, special symbol, preview, etc.), the fill-in content item ("mei miao"), etc.; for example, the third one is the annotation analysis item of the answer box, which includes the hierarchical information of "Level 1 answer", the option button of information format (text or image), the text type (formula and symbol, special symbol, preview, etc.), the fill-in content item ("wonderful"), etc.; the above are only some examples, and other annotation analysis items can also be included according to actual conditions, which are not listed here one by one.

[0117] Among them, when the marked box is an answer box, the associated target content is answer information; when the marked box is a question stem box, the associated target content is question stem information; when the marked box is a question box, the associated target content is question content.

[0118] In some embodiments, the data annotation page further displays the uploaded answer image. For the marking process of the answer image, before step 103, it may include: in response to a marking operation on the answer area where each answer information in the answer image is located, displaying an answer box surrounding the answer area in the answer image; then step 103 may include: displaying the answer information associated with the answer box in the annotation parsing area on the data annotation page.

[0119] Among them, the answer image may contain answer information for each answer box in the test question image. Exemplarily, the answer image may be an image of the standard answer information associated with the question stem in the test question image, see Figure 5 As shown, the answer image is in the middle column on the data annotation page. For example, for the first-level question stem "mei miao" in the test question image (the leftmost column on the data annotation page), the answer information associated with its corresponding answer box in the answer image is "美妙", and for the first-level question stem "yanzou", the answer information associated with its corresponding answer box in the answer image is "演奏", and so on. It should be noted that the area where each answer information is located in the answer image is the answer area.

[0120] It should be noted that each answer information will be filled into the information filling area of the annotation parsing item of the corresponding answer box in the annotation parsing area.

[0121] Exemplarily, the answer image may be an image of the standard answer information associated with the question stem in the test question image, see Figure 5 As shown, the answer image is in the middle column on the data annotation page. For example, for the first-level question stem "mei miao" in the test question image (the leftmost column on the data annotation page), the answer information associated with its corresponding answer box in the answer image is "美妙", and for the first-level question stem "yan zou", the answer information associated with its corresponding answer box in the answer image is "演奏", and so on. The area where each of these answer information is located can be understood as the answer area; for the marking operation triggered by any answer area where an answer information in the answer image is located, box selection marking is respectively performed through the marked box, so that multiple answer boxes appear in the answer image, and each answer box marks one answer information; further, the answer information is filled into the information filling area of the annotation parsing item of the corresponding answer box in the annotation parsing area, such as Figure 5 the annotation parsing area in. One of the items displayed in the annotation parsing area is "1st-level answer". For the answer box associated with the question stem "mei miao", the filled content is "美妙". The above is only an example.

[0122] In the embodiment of the present application, the target content is used to generate the annotation data of the test question. The generation process of the annotation data of the test question is described in the subsequent description.

[0123] In some embodiments, the process (A) of generating the annotated data for the test question includes the following:

[0124] (A.1) Based on the position information of each marked frame in the test question image, determine the subordinate relationship between the answer frame and the test question frame, the subordinate relationship between the question frame and the test question frame, the subordinate relationship between the question frame and the test question frame, and the subordinate relationship between different test question frames.

[0125] In the embodiment of the present application, the test image contains the layout relationship of various information (such as answer information, question stem content, and questions), and the position information of each marked box in the test image is determined by the position of the content information in the test image. In order to determine the subordinate relationship of the marked boxes, the information layout relationship of various information (such as answer information, question stem content, and questions) in the test image can be determined first. This information layout relationship reflects the distribution of various information in the test image, and the position information of the marked boxes can be determined based on this information layout relationship, so that the subordinate relationship of the marked boxes can be determined based on the position information of the marked boxes.

[0126] Among them, the subordinate relationship can refer to the upper and lower hierarchical relationship or the parent-child relationship between the mark boxes. It can be understood that if there is a parent-child relationship between two mark boxes, the level of the child mark box is lower than the level of the parent mark box, and the child mark box is subordinate to the parent mark box, that is, the child mark box is one of the branch mark boxes under the parent mark box.

[0127] Among them, each marking box has corresponding information selected in the test image, and the distribution position of each selected information in the test image is fixed. After the marking box is selected for each information box in the test image, each marking box has fixed position information. The position information of each marking box in the test image is specifically determined according to the content information it selects. Specifically, the position information of each marking box can be represented by coordinate information, with the upper left corner or lower left corner of the test image as the origin, and the position information of each marking box is represented in sequence by coordinates.

[0128] For example, see Figure 5As shown in the figure, in the leftmost page of the test question image, where "I. Word Accumulation" is the title of the first major question, the range of the test question area covered by this first major question is the range outlined by the outermost test question frame (the "third-level test question frame" in the figure), and "1. Read the pinyin and write the words" is a sub-question under this first major question. This sub-question has a corresponding question frame, and the range of the test question area covered by this sub-question is the range outlined by the "second-level test question frame" in the figure. Under this sub-question, there are multiple lowest-level test question frames (i.e., first-level test question frames). Each first-level test question frame contains a question stem frame and an answer area frame. The question stem frame contains pinyin. For example, the content of the question stem in the first lowest-level test question frame is the pinyin of "mei miao", and the blank area below this pinyin is the answer area, which is used to guide the user to answer the word corresponding to this pinyin. Therefore, for each of the above-mentioned marked frames (such as the question frame, test question frame, answer area frame, question stem frame, etc.), their position information is fixed and can be determined according to the layout position of the corresponding information in the test question image. It should be noted that since the test question image is a test question image of the textbook version, the middle column of this test question image is the answer column, and this answer column contains the standard answers for each answer area. During the marking process, for the answer area frame of each answer area, the standard answer information corresponding to the answer area in the answer image can be filled into a specific area in the marking and analysis item for this answer area frame in the marking and analysis area. For example, in the filling area of the marking and explanation item corresponding to this answer area frame in the marking and analysis area, specifically, it can be the information filling area corresponding to the marking and explanation item. For example, fill in "wonderful" in the answer column into the information filling area of the marking and explanation item corresponding to the corresponding answer area frame in the marking and analysis area. It should be noted that this test question image can generally be displayed on the display interface of any device during the marking stage for relevant personnel to understand or operate the marking process of the test question image. In addition to displaying the test question image, this display interface also displays a marking and analysis area in the right area of the test question image. This marking and analysis area contains marking and explanation items corresponding to each marked frame, and the marking and explanation items contain explanatory information for this marked frame.

[0129] In some embodiments, if there is an inclusion relationship between two marked frames, then one of the marked frames is within the selection range of the other marked frame. At this time, there will be an overlap in terms of area between these two marked frames, indicating that there is a subordinate relationship between this marked frame and the other marked frame. Therefore, the overlapping area ratio between the marked frames can be calculated to determine whether there is a subordinate relationship. For example, taking the answer area frame and the test question frame as two types of marked frames, when determining the subordinate relationship between the answer area frame and the test question frame, step (A.1) may include:

[0130] (A.1.1) For the answer area frame and the test question frame to be detected, according to the position information of the answer area frame and the test question frame in the test question image, calculate the overlapping area between the answer area frame and the test question frame;

[0131] (A.1.2) Calculate the overlap ratio between the answer box and the question box based on the overlapping area;

[0132] (A.1.3) Determine the subordinate relationship between the answer box and the question box based on the overlap ratio.

[0133] Specifically, when determining the subordinate relationship between marked frames, the subordinate relationship between marked frames of adjacent levels is generally determined. Taking the lowest level question frame as an example, the direct subordinate relationship between the answer frame or question stem frame and the question frame of the lowest level question frame type is determined, that is, to which question frame the answer frame or question stem frame directly belongs. In addition, the subordinate relationship between the answer frame or question stem frame and the question frame of a higher level can also be determined, which is not limited here. For example, taking the subordinate relationship determination between the answer frame and the lowest level (level 1) type question frame as an example, whether there is an area overlap between the answer frame and any lowest level question frame is detected. Specifically, based on the position information of the answer frame and the position information of the lowest level question frame, whether there is an area overlap between the two is determined. It can be understood that the area of ​​the answer frame is generally smaller than the area of ​​the question frame. If the answer frame is located within the coverage area of ​​the lowest level question frame, then there is an area overlap between the two answer frames and the question frame, and the subordinate answer frame is subordinate to the target question frame of the lowest level. It should be noted that due to the inconsistent coverage (i.e., area) of each answer area and the complex distribution of locations, as well as the influence of frame selection decision factors during the marking stage, the sizes and positions of the selected marked frames during the marking stage may be irregular. To accurately determine the frame inclusion (overlap) relationship between the marked frames, when two marked frames with overlapping areas (including complete and partial area overlap) are detected based on the position information of the marked frames, the area overlap ratio between the two marked frames can be calculated. For example, when an overlapping answer frame and a lowest-level question frame are detected based on the position information of the answer frame and the lowest-level question frame, the overlapping area between the overlapping answer frame and the lowest-level question frame is calculated. The area overlap ratio between the answer frame and the lowest-level question frame is then calculated based on the size of the overlapping area. Finally, the area overlap ratio is compared with a preset threshold. When the area overlap ratio is greater than the preset threshold, it is determined that there is a subordinate relationship between the two marked frames with the greater area overlap ratio (e.g., the answer frame and the current lowest-level question frame), specifically, that the answer frame is subordinate to the current lowest-level question frame. In this way, for two marked boxes with partial area overlap, it is still possible to determine whether there is a subordinate relationship between the two marked boxes, thereby improving the accuracy of determining the subordinate relationship of the marked boxes and effectively improving the robustness of determining the subordinate relationship of the boxes in the case of non-standard marked boxes.

[0134] For example, assuming that one marked frame (such as an answer frame) in a test image is frame A, and another marked frame (such as a test question frame) is frame B, to determine whether there is a subordinate relationship between frames A and B, it is possible to confirm whether there is a frame containment relationship between them. Specifically, the overlapping area of ​​frames A and B is calculated, and the overlapping area ratio is calculated, expressed as "overlapping area / total area of ​​frame A". A preset threshold for determining the existence of a subordinate relationship is set to 60%. If the overlapping area / area of ​​frame A is greater than 60%, frame A is considered to be contained in frame B and subordinate to frame B. Conversely, frame A is considered to be partially contained in frame B and not subordinate to frame B. It should be noted that to avoid determining that two marked frames are mutually contained due to irregularities in the marked frames, a preset threshold can be set according to actual conditions, such as increasing the preset threshold to 85%. In this way, the accuracy of the subordinate relationship determination between the two marked frames with partial area overlap can be improved, ensuring reliability.

[0135] It should be noted that, for the determination of the subordinate relationship of each marking frame in the entire test question image, the above implementation process of the determination of the subordinate relationship of "the answer frame and the lowest-level test question frame" can be referred to. In addition, the coverage area of ​​any highest-level test question frame in the test question image can be used as a judgment area. Under this judgment area, according to the area overlap detection method, it is detected whether there is an area overlap between any two marking frames in the multiple marking frames (such as the answer frame, the question frame, the test question frames of each level, the question frame, etc.) in the judgment area. For the two marking frames with overlapping areas, the size of the area overlap ratio is used to further determine whether there is a subordinate relationship between the two marking frames with overlapping areas. According to the above method, the judgment area under each highest-level test question frame is traversed to determine the subordinate relationship of the marking frames in each judgment area. At this point, the subordinate relationship of all marking frames in the test question image can be determined.

[0136] Through the above method, for the marked boxes selected in the test image, the subordinate relationship of the marked boxes can be accurately determined according to the area overlap between the boxes, so that the subsequent sorting relationship of the combined marked boxes can be used to accurately generate the annotation data of the test question.

[0137] (A.2) Based on the subordinate relationship and the position information of each marked box, multiple marked boxes belonging to the same level are sorted to obtain a marked box sorting relationship.

[0138] In the embodiments of the present application, the text annotation data must not only reflect the subordinate relationships between the question, question stem, and answer information in the test image, but also the ordering relationship between multiple questions at the same level within the same scope, the ordering relationship between multiple question stems within the same test scope, and the ordering relationship between multiple answer boxes within the same test scope. Therefore, after determining the subordinate relationships between the marked boxes in the test image, it is also necessary to determine the ordering relationship between the marked boxes.

[0139] The above-mentioned hierarchy refers to the structural hierarchy of the test content in the test image. It is understandable that the test content in the test image can include test content of multiple levels. For example, assume that the test image contains multiple large questions, one of which is titled "1. Word accumulation", which belongs to the highest level of questions; the large question contains two small questions, titled "1. Read pinyin and write words" and "2. Use words to make sentences", which belong to the middle level of questions; and the title "1. Read pinyin and write words" contains multiple small questions, such as "mei miao", "yan zou", "tan qin", etc. The stem corresponding to each of these small questions can be understood as the lowest level of objectives; the stem box of the small question can be understood as the first level of stem mark box, the title box of each small question belongs to the second level of mark box, and the title box of the large question belongs to the third level of mark box. In an embodiment of the present application, when sorting the mark boxes, the mark boxes belonging to the same level within each question range can be sorted.

[0140] Among them, the sorting relationship of the marked frames can be understood as the sorting information of the marked frames in the test question image, which includes the sorting relationship between the highest-level test question frames, for example, the sorting relationship between the test question frames of "I. Word Accumulation" and "II. Reading Comprehension"; it also includes the sorting relationship between the test question frames of each lower level in the coverage area of ​​the highest-level test frame, such as, "1. Read Pinyin and Write Words" and "2. Use Words to Make Sentences" belong to two small questions under the title of "I. Word Accumulation". The test frame belonging to "1. Read Pinyin and Write Words" and the test frame belonging to "2. Use Words to Make Sentences" are sorted to obtain the sorting relationship. In addition, the question frames of "1. Read Pinyin and Write Words" and "2. Use Words to Make Sentences" under the same parent object (the test frame of "I. Word Accumulation") can also be sorted; in addition, it also includes the sorting relationship between the lowest-level stem frames. Assume that "1. Read Pinyin and Write Words" contains multiple Pinyin stems, such as "mei miao", "yan zou", "tan qin", then these pinyin question frames are sorted with the same serial number. The above is only an example and is not intended to be a limiting method for implementing this application.

[0141] In some embodiments, for the multiple marked frames selected in the test question image, a target marked frame set within the test question area covered by each major question or sub-question may be first determined. For each target marked frame set, a target marked frame subset belonging to the same row position is searched to perform horizontal sorting on the target marked frame subsets. Furthermore, the target marked frame subsets are sorted vertically. Following this approach, the sorting of the marked frames at the next level of each major question is completed, thereby obtaining the marked frame sorting relationship of the test question image. For example, step (A.2) may include:

[0142] (A.2.1) Based on the subordinate relationship, classify the target markup frames at each level that have the same parent object from the answer frame, question frame, and test frame, where the parent object is any test frame;

[0143] (A.2.2) For each target marker frame set, determine at least one target marker frame subset based on the position information of each marker frame in the target marker frame set, where each target marker frame subset includes at least one target marker frame in the same row in the test question image;

[0144] (A.2.3) For each target marker box set, sort multiple target marker box subsets, and sort the target marker boxes within each target marker box subset to obtain a marker box sorting relationship for each marker box in the test image.

[0145] Among them, the parent object can be understood as the mark frame of the previous level, which is determined according to the actual situation. For example, suppose that the test image contains content information of a three-level structure in which one of the big questions contains a three-level structure, wherein the title in the third-level structure is "I. Word accumulation", and the title frame and the test frame to which it belongs are all third-level mark frames; within the coverage of the third-level test frame, it contains a small question with the title "1. Read pinyin and write words", and the title frame and the test frame to which it belongs are second-level mark frames, wherein the common parent object of the second-level test frame and the third-level question frame is the third-level test frame; furthermore, within the coverage of the second-level test frame, it contains multiple pinyin question stems, such as "mei miao", "yan zou", "tan qin", each pinyin question stem has a corresponding answer area, then the second-level question frame covers multiple question stem frames of pinyin question stems and the answer frame corresponding to each pinyin question stem. Each question stem frame and the corresponding answer frame are covered by a first-level question frame, and the question stem frame and the answer frame also belong to the first-level mark frame. The common parent object of multiple question stem frames is the second-level question frame, the common parent object of multiple answer frames also belongs to the second-level question frame, and the common parent object of the first-level question frame for selecting question stems and answers also belongs to the second-level question frame. It should be noted that when sorting the mark frames under the same parent object, different types of mark frames are sorted separately, such as sorting the answer frames under the same parent object, sorting the question stem frames under the same parent object, and sorting the question frames under the same parent object.

[0146] Specifically, first, according to the subordinate relationship of the marker boxes in the test image, the target marker box set with the same parent object in each hierarchical structure is classified. For example, assuming that the highest level contained in the test question image is level three, assuming that the title of the first question is "1. Accumulation of words and characters", the title contains two small questions, namely "1. Read pinyin and write words" and "2. Use words to make sentences", and the first small question "1. Read pinyin and write words" contains multiple small questions, and the title of each small question can be a pinyin stem, such as "mei miao", "yan zou", and "tan qin", and each pinyin stem has a corresponding answer box. In the above test question content, the answer box, the stem box of the pinyin stem, and each test box that selects the answer box and the stem box belong to the first-level marking box, the small question box that selects "1. Read pinyin and write words" and the test box that covers the current small question box and the first-level test box belong to the second-level marking box, and the large question box that selects the large question box and the test box that covers the large question box and the second-level test box belong to the third-level marking box. Based on this, according to the determined tag frame affiliation, assuming that the parent object is the third-level test frame, for all tag frames belonging to the second level, the second-level test frame with the same parent object includes the test frame with "1. Read pinyin and write words" and the test frame with "2. Use words to make sentences" selected. In addition, the second-level question frame with the same parent object includes the question frame with "1. Read pinyin and write words" and the question frame with "2. Use words to make sentences" selected; assuming that the parent object is the second-level test frame, the question frames with "mei miao", "yan zou", "tan zou" and "mei miao" selected respectively. The question frames with "qin" selected belong to the same first-level mark frames. The common parent object of these question frames is the second-level test frame. The answer frame corresponding to each question frame belongs to the first-level mark frame. The common parent object of these answer frames is the second-level test frame. In addition, each test frame that simultaneously selects a question frame and a corresponding answer frame also belongs to the first-level test frame. The parent object of these first-level test frames is the second-level test frame. Based on the above, when the parent object is the third-level test frame, the target mark frame set can be the set of second-level test frames and the set of second-level question frames.When the parent object is a second-level test frame, two scenarios can be included. In the first scenario, the target mark frame set can be a second-level question frame set, a first-level test frame set, a first-level question frame set, and a first-level answer frame set. In this case, the first-level question frame set can contain multiple question frames, and the first-level answer frame set can contain multiple answer frames. In the second scenario, the target mark frame set can be a first-level test frame set and a second-level question frame set. Furthermore, when the parent object is a first-level test frame, the target mark frame set can be a first-level question frame set and a first-level answer frame set in the first-level test frame. The first-level question frame set can contain only one question frame, and the first-level answer frame set can contain only one answer frame. This is not limited here. It should be noted that the above mark frame set is based on the coverage area of ​​a large question. The mark frames within the coverage area of ​​other large questions are determined separately in the above manner.

[0147] Furthermore, for each of the above target mark frame sets, the target mark frames in the target mark frame set can be classified according to the distribution relationship of the row position to determine the target mark frame subset at each row position. Specifically, for the position information of each target mark frame in each target mark frame set, the row position of each target mark frame in the test question image is determined to determine the target mark frame subset corresponding to each row position, wherein each target mark frame subset contains at least one target mark frame at the same row position. For example, see Figure 5 As shown, taking the pinyin question stem frame in the first question "1. Read pinyin and write words" under "I. Word accumulation" as an example, the pinyin question stems of "mei miao", "yan zou" and "tan qin" are in the same row, so one of the target mark frame subsets contains the question stem frames of "mei miao", "yan zou" and "tanqin". For another example, the pinyin question stems of "gan shou", "yue qi" and "shui di" are in the same row, so one of the target mark frame subsets contains the question stem frames of "gan shou", "yue qi" and "shui di". For the first-level answer frame and test question frame, please refer to the example of the question stem frame, which will not be listed here one by one.

[0148] Finally, for each target marker box set, the one or more target marker box subsets contained in it are sorted, and the target marker boxes within each target marker box subset are sorted, thereby obtaining the sorting relationship of the target marker boxes within each target marker box set; according to the above method, the marker box sorting relationship of all target marker boxes in the test image can be obtained.

[0149] In some embodiments, a subset of target marker frames in the same row within a target marker frame set may be determined based on the upper and lower boundaries of the marker frames. For example, step (A.2.2) may include: for each target marker frame set, determining the upper and lower boundary positions of each target marker frame based on the position information of each marker frame; and, based on a first condition, classifying at least one target marker frame in the same row in the test question image from the target marker frame set, and taking the set of at least one target marker frame in the same row as the target marker frame subset; wherein the first condition is that for a first target marker frame and a second target marker frame in the same row, the upper boundary position of the first target marker frame is higher than the lower boundary position of the second target marker frame, and the lower boundary position of the first target marker frame is lower than the upper boundary position of the second target marker frame.

[0150] Each marking frame can be a rectangle, polygon or other shapes. Taking a rectangle as an example, each marking frame includes an upper boundary, a lower boundary and two longitudinal boundaries connecting the upper boundary and the lower boundary respectively. The longitudinal boundaries can be understood as the side boundaries of the marking frame.

[0151] It should be noted that the first condition is used to determine whether any two target mark frames within the target mark frame set are located in the same row position in the test question image. When determining whether the target mark frames are located in the same row position based on the upper and lower boundaries of the mark frames, the first condition is used to compare and classify any two target mark frames. Specifically, the two target mark frames to be compared and classified are defined as the first target mark frame and the second target mark frame, respectively. When the upper boundary of the first target mark frame is higher than the lower boundary position of the second target mark frame, and the lower boundary position of the first target mark frame is lower than the upper boundary position of the second target mark frame, the first target mark frame and the second target mark frame are determined to be located in the same row position. According to the classification method of the first condition, a subset of target mark frames located in the same row position can be determined. It should be noted that the target mark frame subset may contain one target mark frame, indicating that the target mark frame is located in a single row position in the test question image. It should be noted that this application does not limit the vertical height of the row position. The vertical height of the row position can be determined based on the maximum vertical height of the target mark frames contained in the row position, or based on the vertical height occupied by all target mark frames in the same row position.

[0152] In some embodiments, because the size of the target marker box is determined by the size of the corresponding region space, and the marking decision for each region during the marking phase affects the position distribution of the target marker boxes, this may cause the target marker boxes that are actually located in the same row to have uneven sizes and distribution positions. To improve the accuracy and robustness of classifying the row positions of the target marker boxes, the row positions of the target marker boxes can also be classified according to the vertical height overlap ratio of the marker boxes. For example, step (A.2.2) may include: for each target marker box set, determining the overlap length of the longitudinal boundaries between any two target marker boxes in the target marker box set based on the position information of each marker box; determining a longitudinal overlap coefficient based on the overlap length and a reference longitudinal edge length; the reference longitudinal edge length refers to the longitudinal edge length of the target marker box with the smaller longitudinal edge length between the two target marker boxes; if the longitudinal overlap coefficient is greater than a preset threshold, determining that the two target marker boxes are in the same row in the test question image; and treating the set of at least one target marker box in the same row in the test question image as a target marker box subset.

[0153] Specifically, for all target marker frames in each target marker frame set, the longitudinal boundary length of each target marker frame is determined based on the position information of each target marker frame. For any two target marker frames to be compared in each target marker frame set, the two target marker frames are defined as a first target marker frame and a second target marker frame, respectively, to determine the longitudinal overlap length between the first target marker frame and the second target marker frame. Furthermore, the longitudinal boundary length of the first target marker frame is compared with the longitudinal boundary length of the second target marker frame to determine the target longitudinal boundary length with the smaller longitudinal boundary length between the first target marker frame and the second target marker frame as the reference longitudinal side length for comparison and classification. The ratio between the overlap length and the reference longitudinal side length is calculated to obtain a longitudinal overlap coefficient. Finally, the longitudinal overlap coefficient is compared with a preset threshold for row position classification of longitudinal boundary overlap. When the longitudinal overlap coefficient is greater than the preset threshold, it is determined that the first target marker frame and the second target marker frame are located in the same row position in the test question image. Conversely, when the longitudinal overlap coefficient is less than the preset threshold, it is determined that the first target marker frame and the second target marker frame are not located in the same row position in the test question image.

[0154] For example, assuming that a certain marked frame (such as an answer frame) in the test image is frame A, and another marked frame (such as a test question frame) is frame B, in order to determine whether frame A and frame B are in the same row, the vertical boundary overlap between frame A and frame B can be used to determine. Specifically, the overlapping length of frame A and frame B on the vertical boundary is calculated, and the minimum vertical boundary length of frame A and frame B is determined as the reference vertical edge length. The vertical overlap ratio is calculated and expressed as "overlap length / reference vertical edge length". The preset threshold for determining the vertical edge overlap ratio in the same row position is set to 0.5. If the overlap length / reference vertical edge length is greater than 0.5, frame A and frame B are considered to be in the same row position. Otherwise, frame A and frame B are considered not to be in the same row position. It should be noted that the longitudinal boundary overlap can be understood as a situation where there is a real overlap between the two target marking boxes on the longitudinal boundary, and can also be understood as a situation where there is a projection overlap between the two target marking boxes in the longitudinal projection direction of the longitudinal boundary, that is, the longitudinal boundaries of the two target marking boxes do not actually physically overlap; in addition, the preset threshold can be increased to values ​​such as 0.6, 0.65, 0.7, and 0.8. In this way, for two marking boxes with partial area overlap, the accuracy of the row position distribution judgment of the two marking boxes can be improved, which is reliable.

[0155] It should be noted that when classifying the target mark frames by row position, the row position of the target mark frames within the target mark frame set can be classified by combining the "first condition" and "vertical boundary overlap". For example, the row position classification "according to the first condition" in the previous implementation is first performed, and then the row position classification according to "vertical boundary overlap" is performed, thereby improving the accuracy of the row position classification of the target mark frames.

[0156] In some embodiments, after determining each target mark frame subset at the same row position, each target mark frame subset includes one or more target mark frames at the same row position. The target mark frames included in each target mark frame subset can be sorted, and the target mark frame subsets in different rows can be sorted to determine the sorting relationship of all target mark frames in each target mark frame subset, thereby obtaining the mark frame sorting relationship of all target mark frames in the test question image. For example, step (A.2.3) can include:

[0157] (A.2.3.1) For each target marker frame set, sort the target marker frame subsets in descending order based on the row positions of the target marker frames in the target marker frame subsets to obtain a row sorting relationship;

[0158] (A.2.3.2) Sort the target marker frames in a left-to-right order based on the position information of each target marker frame in each target marker frame subset to obtain a horizontal sorting relationship between the target marker frames in the same row.

[0159] (A.2.3.3) Determine the sorting relationship between target marker frames in the corresponding target marker frame set based on the row sorting relationship and the horizontal sorting relationship to obtain the marker frame sorting relationship of each marker frame in the test question image.

[0160] Specifically, within each target mark frame set, after row position classification processing, a target mark frame subset corresponding to each row position is obtained. At this point, multiple target mark frame subsets in different rows can be sorted to obtain a row ordering relationship for the target mark frames within each target mark frame set. Then, target mark frames in the same row position within each target mark frame subset are horizontally sorted. Specifically, for each target mark frame subset, based on the position information of each target mark frame in each target mark frame subset, multiple target mark frames in the current target mark frame subset are horizontally sorted in the same row position from left to right. This determines the horizontal ordering relationship between the multiple target mark frames in the same row position within each target mark frame subset. Finally, based on the row ordering relationship and horizontal ordering relationship of the mark frames within each target mark frame set, the ordering relationship between the target mark frames in the corresponding target mark frame set is determined. This, combined with the ordering relationship of the target mark frame sets contained in the test question image, determines the mark frame ordering relationship for the test question image.

[0161] In some embodiments, when sorting the target mark frames in the target mark frame subset that are in the same row position, each target mark frame may be numbered incrementally in the order of sorting, and the number is used to indicate the sorting position of the corresponding target mark frame in the row position. For example, step (A.2.3.2) may include: adding a start mark to the target mark frame on the leftmost side of the test question image in the target mark frame subset, and adding an end mark to the target mark frame on the rightmost side of the test question image; starting from the target mark frame containing the start mark, the target mark frames in the target mark frame subset are numbered incrementally from left to right until the target mark frame containing the end mark is numbered, thereby obtaining a horizontal sorting relationship between multiple target mark frames belonging to the same row position.

[0162] For example, taking the upper left corner of the test image as the coordinate origin, the coordinate information of each marking box can be determined according to its position in the test image. For each target marking box subset, a subscript "start" is set at the coordinate information of the first target marking box from left to right in the same row position, and a subscript "end" is set at the coordinate information of the rightmost (i.e., end) target marking box in the same row position. Then, the coordinate information is sorted in ascending order from left to right in the same row position. For each target marking box, a serial number can be added to the subscript of the marking box. For example, assuming that the same row position contains 6 answer boxes, the serial number can be expressed as "D1, D2, D3, D4, D5, D6", or when numbering, the serial number of the row indicated by the row sorting relationship of the target marking box subset is added, and the serial number is expressed as "D 11 、D 12 、D 13 、D 14 、D 15 、D 16 "; For example, taking the question stem frame as an example, assuming that there are 3 question stem frames in the same row, the serial numbers can be expressed as "T1, T2, T3", and the row number indicated by the row sorting relationship of the target mark frame subset is added, and the serial number is expressed as "T 11 、T 12 、T 13 ”. According to the above example, the numbering is sorted until the coordinate information contains the target mark frame position with the subscript “end”, so that the target mark frames in the target mark frame subset at the same row position are sorted, and the horizontal sorting relationship between multiple target mark frames at the same row position is completed.

[0163] Through the above method, the marked boxes selected in the test question image can be sorted within the divided area to determine the sorting relationship of the marked boxes in the test question image, so that the hierarchical structure and distribution of the marked boxes in the test question image can be accurately represented by combining the subordinate relationship and sorting relationship of the marked boxes in the test question image, which is reliable.

[0164] (A.3) According to the subordinate relationship and the mark box sorting relationship, the answer information associated with each answer area in the test image, the question stem content of each question stem area, and the question content of each question area are combined to generate annotation data for the test question.

[0165] In an embodiment of the present application, after determining the subordinate relationship and the sorting relationship of the mark boxes in the test question image, the information associated with all the mark boxes in the test question image can be combined according to the subordinate relationship and the sorting relationship of the mark boxes. It can be understood that the answer box is associated with the answer information corresponding to the answer area, the stem box is associated with the corresponding stem content, and the question box is associated with the corresponding question content. The above information is combined to generate text annotation information with a text hierarchical structure, so as to represent the question, stem and answer information in the test question image according to a specific structure, so as to accurately represent the text information in the test question image.

[0166] The answer information may be the standard answer information printed in the test image, see Figure 5 The middle column is the "Answer Area". The answer area contains the printed standard answer information. Each standard answer information corresponds to the corresponding answer box under the corresponding question in the left column of the figure.

[0167] The question content can be the question information associated with each answer box in the test image, which is the information basis for the user to answer, such as Figure 5 The pinyin questions include “mei miao”, “yan zou”, etc.

[0168] The title content can be the title content corresponding to each level, see Figure 5 As shown in the figure, "1. Read pinyin and write words" and "1. Accumulate words and phrases" are both part of the question content.

[0169] The text annotation data may be text information with a hierarchical structure, which is presented in a framework form, for example, represented in a tree form. Specifically, for the text annotation data, each marked box is used as a leaf, and the branches are used to represent the subordinate relationship and sorting relationship of the marked boxes. For example, see Figure 6As shown, assuming that the test image belongs to the page content of a certain page number in the textbook, the basic information of the page content is taken as the highest-level information, such as the textbook subject, subject, page number and third-level question (such as title and coverage) and other related information. The above information constitutes a third-level mark frame, which serves as the root of the structure tree; then, based on the third-level mark frame, according to the subordinate relationship and sorting relationship, the second-level mark frame is determined. The figure contains two second-level mark frames, each of which is a node, which contains the second-level question stem (or title), coordinate range (that is, the coverage of the current second-level question) and the page number information occupied by the second-level question. The page information can be a single page number (such as the first 32, 32 or 33, etc.), or page turning information (such as page 32, 32 and 33, etc.); on the basis of the secondary mark box, a primary mark box is determined, and each primary mark box contains the answer question, the question type of the question stem, the coordinate range (that is, the coverage of the current primary question), and the page number information occupied by the current primary question, etc.; on the basis of the primary mark box, multiple answer boxes belonging to each primary mark box are further determined, and the content expressed by each answer box includes the corresponding standard answer information, the correction range for the answer box, etc. In addition, the answer box may also include some additional information, which may be prompt information for specific answer information content.

[0170] In some embodiments, in order to ensure the depth of the text annotation data in the hierarchical structure, it is necessary to ensure that the text annotation data has sufficient hierarchical depth. For branch structures with insufficient levels, they can be hierarchically padded to generate text annotation data that meets the requirements. For example, step 104 may include: constructing initial hierarchical structure information including answer boxes, question boxes, and test box boxes based on the subordinate relationship and the mark box sorting relationship; if it is detected that the initial hierarchical structure information contains a target branch structure with insufficient levels, then the target branch structure in the initial hierarchical structure information is hierarchically padded to obtain the padded target hierarchical structure information, where the target branch structure is a branch structure with a level less than a preset level threshold; if it is detected that the initial hierarchical structure information does not contain the target branch structure, then the initial hierarchical structure information is determined as the target hierarchical structure information; according to the target hierarchical structure information, the answer information associated with each answer area in the test question image, the question content of each question area, and the question content of each question area are combined to generate annotation data for the test question.

[0171] Specifically, after obtaining the subordination and ordering relationships of the marked boxes in the test question image, initial hierarchical structure information containing all the marked boxes in the test question image can be constructed based on the subordination and ordering relationships. This initial hierarchical structure information can be in the form of a structure tree, which contains the hierarchical structure relationships between all answer boxes, question boxes, and test boxes, reflecting the connection relationship between the marked boxes. After obtaining the initial hierarchical structure information, each branch structure can be traversed in a root-to-leaf or leaf-to-root direction of the structure tree, and the number of levels contained in each branch structure can be determined. The number of levels of each branch structure is then compared with a preset hierarchical threshold to detect whether each branch structure meets the hierarchical depth requirement. If the initial hierarchical structure information contains a target branch structure with a number of levels less than the preset hierarchical threshold, the target branch structure is hierarchically padded according to the requirements of the preset hierarchical threshold to obtain the target hierarchical structure information. Conversely, if the initial hierarchical structure information does not contain a target branch structure with a number of levels less than the preset hierarchical threshold, the initial hierarchical structure information is directly determined as the target hierarchical structure information. Finally, according to the hierarchical structure relationship indicated by the target hierarchical structure information, the answer information associated with each answer area in the test image, the stem content of each stem area, and the question content of each question area are combined to generate annotation data for the test question.

[0172] It should be noted that for the target branch structure with a number of levels less than the preset level threshold, since each branch structure generally contains an answer box and a corresponding first-level question stem box, the phenomenon of insufficient number of levels is mainly due to the lack of a hierarchical structure for the second-level questions. For example, in the scenario of application questions in mathematics or composition questions in Chinese, there is usually only information about the stem content, and no higher-level questions. After the corresponding branch structure is generated, the stem content and answer information of this scenario lack the hierarchical structure of the second-level questions relative to other multi-level questions. Therefore, in order to ensure that the subsequently generated text annotation structure meets the hierarchical depth requirements, a hierarchical box for the second-level questions can be generated between the level of the first-level questions and the root (the level of the third-level questions) in the target branch structure to connect the level of the first-level questions and the root. It should be noted that the hierarchical box for the generated second-level questions does not contain the actual information in the test question image, and is only used to meet the hierarchical depth of the current branch structure, so that each branch of the subsequently generated text annotation data has a specific hierarchical depth, meets the data format requirements of the subsequent intelligent automatic correction, and is reliable.

[0173] In some embodiments, for the generated hierarchically structured text annotation data, the answer information within the answer box may be further annotated to indicate the category to which the answer information belongs, thereby improving the accuracy and efficiency of the subsequent automatic correction process. For example, after step (A.3), the following steps may be included: identifying the answer boxes in the text annotation data; performing semantic analysis on the answer information within each answer box to determine the key information in the answer information and the answer category corresponding to each key information; marking the key information in each answer box with a key answer mark box, and marking the corresponding answer category for each key answer mark box.

[0174] For example, take the word problems in mathematics test as an example, see Figure 7 As shown, assuming that the answer content of the answer box of the text annotation data is represented in the test image as follows Figure 7 As shown in the "answer box" area, the answer information in each answer box can be semantically recognized to determine the key information in the answer information. This key information can be the final structure of the question in the stem. For example, "2200 (kilometers)" in the figure is the key information in the answer information, and "Answer: It can travel about 2200 kilometers" is also the key information. These key information are classified to determine the answer category corresponding to each key information. Thus, the key information is marked with a marking box in the answer information, and the corresponding answer category is not marked on each key answer marking box. For example, the key information "2200 (kilometers)" belongs to the answer value category, and the key information "Answer: It can travel about 2200 kilometers" belongs to the answer text category. It can be understood that the key information is mainly determined by the target question in the question content.

[0175] In some embodiments, the hierarchical structure relationship of the text annotation data is mainly determined based on the subordinate relationship and sorting relationship of the mark box. The subordinate relationship and sorting relationship of any mark box can be adjusted according to needs to change the position of the mark box, thereby adjusting the hierarchical structure relationship corresponding to the text annotation data. In this way, the corresponding annotation data can be obtained by quickly adjusting the hierarchical structure relationship of the text annotation data without regenerating the text annotation data, which has high efficiency. For example, after step (A.3), it can also include: determining the target content of the text annotation data whose subordinate relationship is to be adjusted, and the test question area of ​​the content to be expanded; adjusting the position of the mark box corresponding to the target content to the test question box corresponding to the test question area of ​​the content to be expanded in the text annotation data, and generating updated target text annotation data.

[0176] Specifically, first determine the target content that needs to be adjusted and the test area that needs to be expanded in the text annotation data, move the mark box that selects the target content to the test area that selects the content that needs to be expanded, and adjust the position of the mark box of the target content to update the subordinate relationship between the mark box of the target content and the test box of the previous level at its current location. In addition, all mark boxes of the same type as the mark box of the target content in the current test box can be re-sorted, thereby generating updated target text annotation data. In this way, the subordinate relationship of the mark boxes is updated, thereby updating the hierarchical structure relationship, so as to achieve rapid adjustment of previously generated text annotation data without the need to regenerate additional annotation data, which is convenient.

[0177] Through the above method, the information associated with all the marked boxes in the test question image can be combined according to the subordinate relationship and the marked box sorting relationship to generate text annotation data to accurately represent the text information in the test question image; in addition, by adjusting the subordinate relationship, the text annotation data can be quickly adjusted, which is flexible and convenient.

[0178] As can be seen from the above, the embodiment of the present application can provide a data annotation page of the data annotation platform, which can display the test image of any test question on the data annotation page, and divide the test question image into target areas. The target areas include the minimum-level answer area, the stem area associated with the answer area, and the minimum-level test question area composed of the answer area and the stem area. In addition, each minimum-level test question area is associated with the question area, and one or more minimum-level test question areas and the commonly associated question areas are combined to obtain an increasing-level test question area, and so on; then, for the marking operation of each target area, a marking box surrounding the target area is displayed in the test question image; finally, the target content associated with each marking box is displayed in the annotation analysis area in the data annotation page. Specifically, when the marking box is an answer box surrounding the answer area, the target content is answer information; when the marking box is a stem box surrounding the stem area, the target content is stem information; when the marking box is a question box, the target content is question content. The above target content is used to generate annotation data for the test question. This can be applied to the marking process of test questions or homeworks of any question type and structure, so as to automatically generate marking data for test questions or homeworks based on the target content associated with the marking box, thereby improving the marking efficiency of test questions or homeworks, so as to realize automatic grading of test questions or homeworks based on the marked data, thereby improving the grading efficiency of homeworks or test questions.

[0179] This application does not specifically limit the method of automatic grading. For example, the test taker can upload an image of the completed test question. The annotation data can be used to locate any content in the image of the completed test question, and the test taker's answer can be extracted and then compared with the answer information in the annotation data to complete the automatic grading.

[0180] The method described in the above embodiment will be further described in detail below with examples.

[0181] The embodiment of the present application takes the test question image processing as an example to further describe the test question image processing method provided in the embodiment of the present application.

[0182] Figure 8 This is another step flow chart of the test question image processing method provided by the embodiment of the present application. Figure 8 Provide a description.

[0183] In the embodiments of this application, the present invention will be described from the perspective of a test image processing device, which can be integrated into a computer device such as a server. For example, when a processor on the computer device executes a program corresponding to the test image processing method, the specific process of the test image processing method is as follows:

[0184] 201. Obtain the test question image corresponding to the target test question file.

[0185] In the embodiment of the present application, the target test question file may be a textbook, test paper, exercise book or other file, and the test question image may be an image containing information of any content page in the file.

[0186] When obtaining the test image corresponding to the target test file, the acquisition method can be taking a photo, scanning, etc., so that text annotation data can be generated based on the obtained test image for use in training the neural network model. Then, the trained neural network model can be used to correct and review the test answer information of the user (such as students, test takers, etc.) to achieve automated correction of test files.

[0187] 202. Mark the answer area in the test image, the question stem area of ​​each answer area, and the test question area corresponding to the question to which the question stem belongs, to obtain multiple marking frames.

[0188] Among them, the test question image may contain test question areas of multiple levels, assuming that there are three levels of test question areas: high, medium, and low. Exemplarily, the description is made in the order of the levels from question to question stem and answer area from high to low. Assuming that the test question image contains multiple questions, the test question range area covered by each question can be understood as the largest test question area or the highest level test question area; for each highest level test question area, it can contain multiple middle-level questions, and the test question range area covered by each middle-level question can be understood as a medium-sized test question area or a middle-level test question area. In addition, each highest level test question area can also directly contain a low-level test question area; for each middle-level test question area, it can contain multiple low-level questions, and the test question range area covered by each low-level question can be understood as a small test question area or a low-level test question area. Each low-level test question area contains a question stem and an answer area (i.e., an answer filling area). It should be noted that for each mid-level question area, it can also include multiple question areas of decreasing levels, until it is located in the question area of ​​the next level above the low-level question area. For example, assuming that the low-level question area is the level 1 question area and the high-level question area is the level 5 question area, it can be understood as including three mid-level question areas, such as the level 4 question area, the level 3 question area, and the level 2 question area. The above is only an example and is not intended to be a limiting method for implementing this application.

[0189] The stem area corresponds to the stem content, which describes the question content and is the information basis for answering in the answer area. Each stem area has an independent answer area, which can be represented by a "line blank" (such as "_______"), brackets, rectangular boxes or "blank areas". For example, the test image contains multiple large questions, each large question has a corresponding question, and the question contains one or more question areas of the next level of test questions, see Figure 3 As shown, the dotted box represents the question area, and the "1. Listen and choose what you hear" in the box can be understood as the title of the first question. The question range area covered by the question of this question is the outermost bold solid wireframe in the figure. In this scenario, the question of the big question can be understood as the question of the level 2 question; the question area of ​​the four lowest-level questions directly included in the big question can be understood as the question area of ​​the level 1 question. Each question area contains the question stem content in the question stem box and the answer information in the answer box. For example, the first question "1. Can you____a bike? A. ride B. bird" is the question stem content of the level 1 question, which serves as the information basis for the user when answering, and "(A)" is the answer information of the level 1 question. It can be understood that the answer information is attached to the question stem content. The above is only an example and is not a limiting method for implementing this application.

[0190] In the embodiment of the present application, when marking each area in the test image, the target area contained in the test image is marked in order from low to high hierarchical order of the answer area, the stem area, and the test area, to obtain multiple marking frames, wherein the type of the target area can be one of the answer area, the stem area, and the test area, and the test area is at least a first-level (low-level) test area. The types of the marking frames may include an answer frame corresponding to the answer area, a stem frame corresponding to the stem area, and a test frame corresponding to the test area.

[0191] 203. Determine the subordinate relationship between the answer frame and the test frame, the subordinate relationship between the question frame and the test frame, and the subordinate relationship between different test frame according to the position information of each mark frame in the test question image.

[0192] It should be noted that each marking box has corresponding information selected in the test image, and the distribution position of each selected information in the test image is fixed. After the marking box is selected for each information box in the test image, each marking box has fixed position information. The position information of each marking box in the test image is determined according to the content information it selects. Specifically, the position information of each marking box can be represented by coordinate information, with the upper left corner or lower left corner of the test image as the origin, and the position information of each marking box is represented in sequence by coordinates.

[0193] Among them, the subordinate relationship can refer to the upper and lower hierarchical relationship or parent-child relationship between the mark boxes. It can be understood that if there is a parent-child relationship between two mark boxes, the level of the child mark box is lower than the level of the parent mark box, and the child mark box is subordinate to the parent mark box.

[0194] In an embodiment of the present application, if there is a containment relationship between two marking frames, then one of the marking frames is within the selection range of the other marking frame. At this time, the two marking frames will overlap in area, indicating that there is a subordinate relationship between the marking frame and the other marking frame. Therefore, the existence of a subordinate relationship can be determined by calculating the overlapping area ratio between the marking frames. For example, taking the two types of marking frames, the answer frame and the test question frame, as an example, in determining the subordinate relationship between the answer frame and the test question frame, the specific process is as follows: for the answer frame and the test question frame to be detected, according to the position information of the answer frame and the test question frame in the test question image, calculate the overlapping area of ​​the answer frame and the test question frame; calculate the overlapping ratio between the answer frame and the test question frame based on the overlapping area; and determine the subordinate relationship between the answer frame and the test question frame based on the overlapping ratio.

[0195] 204. Based on the subordinate relationship, a target mark frame set with the same parent object at each level is classified from the answer frame, question frame and test question frame.

[0196] In the embodiment of the present application, for the multiple marked frames selected from the test image, the target marked frame sets with the same parent object in each hierarchical structure can be classified according to the subordinate relationship of the marked frames in the test image. Specifically, based on the subordinate relationship, the target marked frame sets with the same parent object in each hierarchy are classified from the answer frame, question frame and test frame. The parent object is any test frame, which can be understood as the marked frame of the previous level. It is determined according to the actual situation.

[0197] For example, assuming that the highest level contained in the test image is level three, see Figure 5 As shown, assuming that the title of the first question is "1. Accumulation of words and characters", the title contains two sub-questions, namely "1. Read pinyin and write words" and "2. Use words to make sentences". The first sub-question "1. Read pinyin and write words" contains multiple sub-questions, and the title of each sub-question can be a pinyin stem, such as "mei miao", "yan zou", and "tan qin", and each pinyin stem has a corresponding answer box. In the above test content, the answer box, the stem box of the pinyin stem, and each test box that selects the answer box and the stem box belong to the first-level marking box, the sub-question title box that selects "1. Read pinyin and write words" and the test box that covers the current sub-question title box and the first-level test box belong to the second-level marking box, and the large question title box that selects "1. Accumulation of words and characters" and the test box that covers the large question title box and the second-level test box belong to the third-level marking box. Based on this, according to the determined tag frame affiliation, assuming that the parent object is the third-level test frame, for all tag frames belonging to the second level, the second-level test frame with the same parent object includes the test frame with "1. Read pinyin and write words" and the test frame with "2. Use words to make sentences" selected. In addition, the second-level question frame with the same parent object includes the question frame with "1. Read pinyin and write words" and the question frame with "2. Use words to make sentences" selected; assuming that the parent object is the second-level test frame, the question frames with "meimiao", "yan zou", "tan The question stem frames with "qin" selected belong to the first-level mark frames. The common parent object of these question stem frames is the second-level test question frame. The answer frames corresponding to each question stem frame belong to the first-level mark frames. The common parent object of these answer frames is the second-level test question frame. In addition, each test question frame that simultaneously selects a question stem frame and a corresponding answer frame also belongs to the first-level test question frame. The parent object of these first-level test question frames is the second-level test question frame.

[0198] Based on the above, when the parent object is a third-level question frame, the target mark frame set can be a set of second-level question frames and a set of second-level question frames. When the parent object is a second-level question frame, two scenarios can be included. The first scenario is that the target mark frame set can be a set of second-level question frames, a set of first-level question frames, a set of first-level question stem frames, and a set of first-level answer frames. In this case, the first-level question stem frame set can contain multiple question stem frames, and the first-level answer frame set can contain multiple answer frames. The second scenario is that the target mark frame set can be a set of first-level question frames and a set of second-level question frames. Furthermore, when the parent object is a first-level question frame, the target mark frame set can be a set of first-level question stem frames and a set of first-level answer frames in the first-level question frame. The first-level question stem frame set can contain only one question stem frame, and the first-level answer frame set can contain only one answer frame. This is not limited here. It should be noted that the above set of marking boxes is based on the coverage area of ​​one major question. The marking boxes within the coverage areas of other major questions are determined separately in the above manner.

[0199] 205. For each target marker frame set, determine at least one target marker frame subset according to position information of each marker frame in the target marker frame set, wherein each target marker frame subset includes at least one target marker frame in the same row in the test question image.

[0200] In an embodiment of the present application, the row position of each target mark frame in each target mark frame set is determined based on the position information of each target mark frame in the test question image, so that the target mark frames in the target mark frame set are classified according to the distribution relationship of the row positions to determine the target mark frame subset at each row position. Specifically, based on the position information of each target mark frame in each target mark frame set, the row position of each target mark frame in the test question image is determined to determine the target mark frame subset corresponding to each row position, wherein each target mark frame subset contains at least one target mark frame at the same row position. For example, see Figure 5As shown, taking the pinyin question stem frame in the first question "1. Read pinyin and write words" under "I. Word accumulation" as an example, the pinyin question stems of "mei miao", "yan zou" and "tan qin" are in the same row, so one of the target mark frame subsets contains the question stem frames of "mei miao", "yan zou" and "tanqin". For another example, the pinyin question stems of "gan shou", "yue qi" and "shui di" are in the same row, so one of the target mark frame subsets contains the question stem frames of "gan shou", "yue qi" and "shui di". For the first-level answer frame and test question frame, please refer to the example of the question stem frame, which will not be listed here one by one.

[0201] Among them, when classifying the target mark box subsets at the same row position, the target mark boxes in the target mark box set can be determined according to the upper and lower boundaries of the mark boxes for row position classification, and the target mark boxes can also be classified according to the vertical height overlap ratio of the mark boxes.

[0202] 206. For each target marked frame set, sort multiple target marked frame subsets, and sort the target marked frames within each target marked frame subset, to obtain a sorting relationship of the marked frames of the test question image.

[0203] In an embodiment of the present application, for each target marking frame set, after searching for a subset of target marking frames belonging to the same row position, the subset of target marking frames can be sorted horizontally, and the subset of target marking frames can be sorted vertically. According to the above method, the sorting of the target marking frames in each target marking frame set is completed, and the marking frame sorting relationship of the test question image is obtained.

[0204] After determining each target mark frame subset in the same row position, each target mark frame subset includes one or more target mark frames in the same row position. The target mark frames included in each target mark frame subset can be sorted, and the target mark frame subsets in different rows can be sorted to determine the sorting relationship of all target mark frames in each target mark frame subset, thereby obtaining the mark frame sorting relationship of all target mark frames in the test question image. It should be noted that when sorting target mark frames within a target mark frame set under the same parent object, different types of mark frames are sorted separately, such as sorting answer frames under the same parent object, sorting question stem frames under the same parent object, and sorting test question frames under the same parent object.

[0205] 207. According to the subordinate relationship and the mark box sorting relationship, the answer information associated with each answer area in the test image, the question stem content of each question stem area, and the question content of each question area are combined to generate annotation data for the test question.

[0206] In an embodiment of the present application, after determining the subordinate relationship and the sorting relationship of the mark boxes in the test question image, the information associated with all the mark boxes in the test question image can be combined according to the subordinate relationship and the sorting relationship of the mark boxes. It can be understood that the answer box is associated with the answer information corresponding to the answer area, the stem box is associated with the corresponding stem content, and the question box is associated with the corresponding question content. The above information is combined to generate text annotation information with a text hierarchical structure, so as to represent the question, stem and answer information in the test question image according to a specific structure, so as to accurately represent the text information in the test question image.

[0207] Furthermore, in order to ensure the depth of the text annotation data in the hierarchical structure, it is necessary to ensure that the text annotation data has sufficient hierarchical depth. For branch structures with insufficient number of levels, the levels can be padded to generate text annotation data that meets the requirements.

[0208] In an embodiment of the present application, for the generated text annotation data with a hierarchical structure, the answer information in the answer box can be further annotated to indicate the category to which the answer information belongs, so as to improve the accuracy and efficiency of the subsequent automatic correction process.

[0209] In an embodiment of the present application, the hierarchical structure relationship of the text annotation data is mainly determined based on the subordinate relationship and sorting relationship of the marking box. The subordinate relationship and sorting relationship of any marking box can be adjusted according to needs to change the position of the marking box, thereby adjusting the hierarchical structure relationship corresponding to the text annotation data. In this way, the corresponding annotation data can be obtained by quickly adjusting the hierarchical structure relationship of the text annotation data, without the need to regenerate the text annotation data, which is highly efficient.

[0210] To facilitate understanding of the embodiments of the present application, the embodiments of the present application will be described using a specific application scenario example. Specifically, the application scenario example is described by executing the above steps 201-207.

[0211] It should be noted that this test image processing method is mainly used in test image processing scenarios in the fields of education and government. Taking the test image processing scenario in the education field as an example, the specific example of this scenario is as follows:

[0212] 1. Overview of the image processing scenario example of this question:

[0213] This scenario example primarily applies to the data annotation process for the intelligent photo-based grading feature. The result of this process is information such as the question hierarchy, type, question stem, grading location, and answer key for a single-page textbook test. This can be expanded to other scenarios involving text hierarchical structure annotation.

[0214] Combine Figure 5 As shown, on the page of the intelligent photo-taking and teaching aid annotation platform, the left column is the test question area, the middle column is the answer area, and the right column is the annotation and analysis area.

[0215] 2. The implementation process of this scenario example mainly includes the stage of marking the mark box corresponding to the message in the test question image and the stage of reorganizing the hierarchical structure of the content information in the test question image, as follows:

[0216] 1. Combination Figure 9 As shown, the processing process of this marking phase is as follows:

[0217] (1) Mark the answer box in the answer area. Unlike the current technical solution, the question marking in this solution starts from the smallest level and proceeds from bottom to top. That is, the answer area that needs to be corrected (correction points) is marked directly first. Because the correction points are marked first, the level depth of the correction points can be considered consistent, such as level 0.

[0218] (2) Fill in the answer area information. This solution process also displays the parsing annotation area page on the right side of the data annotation page. You can directly select the corresponding answer area for each answer area in the test image in the answer column in the middle column of the data annotation page to perform OCR on the text (answer information) in the answer area, reducing the workload of manual input and reducing input errors.

[0219] (3) Mark the question stem box and fill in the question stem information. All places with text and images in the textbook should be marked as question stems. We also use OCR to reduce manual work and support image capture, which will not be explained in detail.

[0220] (4) Mark the outer question frame. Mark the overall scope of the question. For example, take Question (2) of Question 3 of Question 1 as an example. First, mark the scope of Question (2), frame its answer area and the stem of the question; then mark the scope of Question 3 outward, frame all the questions; then mark the scope of Question 1 outward, frame all the questions, and so on.

[0221] After the answer area, question area, and test area are marked, the textbook presents a nested structure of area boxes (such as Figure 5 Left column), but the area boxes themselves do not have a rigid set hierarchical relationship, and they are flat structures (such as Figure 5At this point, the marking phase is over, and all the information required for manual annotation is available. The subsequent reassembly phase will assemble the information generated in the marking phase according to the business logic.

[0222] 2. Combination Figure 10 As shown in the figure, the processing process of the reorganization phase is as follows:

[0223] (1) Annotation frame standardization. This processing step may include preprocessing of the annotation data frame. For example, if the grading process requires that the level depth of all test questions must be at least 2, then in this step, a level 2 question frame of exactly the same size can be automatically added to the questions with a historical annotation depth of 1 (for example, questions with only a level 1 question frame and no level 2 question frame) to align the levels.

[0224] (2) Confirmation of frame inclusion relationship. Specifically, the flat-structured annotation frame is constructed as a hierarchical tree of the test questions. The key is to confirm the inclusion relationship between different frames. Because manual annotation has inevitable errors, the confirmation of frame inclusion relationship needs to be robust enough to solve the problem of frame offset. For example, based on the overlapping area ratio, if the inclusion relationship between frame A and frame B is confirmed, the overlapping area of ​​frame A and frame B is calculated. If the overlapping area (IOU area) / frame A area>60%, frame A is considered to be included in frame B. If the IOU area / frame B area>60%, frame B is considered to be included in frame A. In order to avoid mutual inclusion between frames, when comparing frames of the same type, the inclusion threshold will be raised to 85% (because non-similar frames must have a subordinate relationship, such as a level 2 question frame can only be included by a level 2 test question frame, but cannot include a level 2 test question frame in turn, so this processing is not required). If mutual inclusion and inability to judge still occur, it will be automatically checked during annotation to remind the annotator to adjust the frame range.

[0225] Among them, the confirmation of the frame inclusion relationship needs to distinguish the test question frame, the question stem frame and the answer frame and execute them separately. Specifically, for the test question frame, it is necessary to judge the inclusion relationship between the test question frames in pairs, and the upper test question frame contains the lower test question frame; for the question frame, it is only necessary to judge the inclusion relationship between the question frame and the test question frame to which it belongs (usually the minimum depth is 2-level test question frame); for the question stem frame, it is only necessary to judge the inclusion relationship between the question stem frame and the test question frame; for the answer frame, it is only necessary to judge the inclusion relationship between the answer frame and the test question frame). In this way, a primary structure tree of the test questions on the current page is generated based on the inclusion relationship. For example, taking the test question depth of 2 levels as an example, the structure tree has several 2-level test question frames on the page, each of which has 2-level questions and several 1-level test question frames, and the 1-level test question frame has several 1-level stems and answer frames; it should be noted that the test question depth can also be 3, 45 levels, which can be deduced in the above way, and will not be repeated here.

[0226] (3) Marking boxes to classify the positions of the same box. Specifically, boxes with the same parent object at the same level need to be divided into pairs to determine the positions of the same box. For example, the order of the answer boxes under a level 1 test question box needs to be clearly distinguished, otherwise the answer may be confused when displayed, causing misunderstanding.

[0227] The method of dividing the same-layer frames uses the following logical judgment:

[0228] (3.1) If the lower bound of box a is smaller than the upper bound of box b or the upper bound of box a is larger than the lower bound of box b, then the two boxes are not in the same row (the page origin is in the upper left).

[0229] (3.2) If the length of the longitudinal overlap between frame a and frame b (i.e., the minimum of their upper bounds minus the maximum of their lower bounds) / the minimum of the heights of frame a and frame b is greater than 0.5, then frame a and frame b are considered to be on the same layer; otherwise, they are not on the same layer.

[0230] (4) Sorting the marked boxes at the same position. Specifically, after drawing the marked boxes at the same position, the method of sorting at the same level adopts the following logic:

[0231] (4.1) Sort all boxes to be sorted in ascending order by upper bound coordinates (the origin of the page is at the upper left).

[0232] (4.2) Set two subscripts, start and end, and initialize them to 0.

[0233] (4.3) Compare the end position element with the next element to see if they are in the same row. If so, move the end right and repeat step (4.3). If not, sort all elements from start to end in ascending order by left boundary coordinate.

[0234] (4.4) Move from start to end+1 and repeat step C until all boxes are sorted.

[0235] (5) Structure tree generation. Specifically, the final structure tree is formed from high to low levels. The specific process is as follows:

[0236] (5.1) Find all Level 2 question boxes;

[0237] (5.2) For each Level 2 question frame, traverse all Level 2 question frames. Since the Level 2 question frame is subordinate to the Level 2 question frame, the Level 2 question frame is the parent node of the Level 2 question frame. Find the Level 2 question frames subordinate to it in each Level 2 question frame, sort all Level 2 question frames, and form the complete content of the Level 2 question.

[0238] (5.3) For each Level 2 question frame, traverse all Level 1 question frames, find the Level 1 question frames that are subordinate to it, sort all Level 1 question frames, and form the complete content of all Level 1 question frames under the Level 2 question frame;

[0239] (5.4) For each Level 1 question frame, traverse all Level 1 question stem frames, find the Level 1 question stem frames that are subordinate to it, and sort all Level 1 question stem frames to form the complete content of the Level 1 question stem;

[0240] (5.5) For each level 1 question frame, traverse all answer frames, find the answer frames that are subordinate to it, sort all answer frames, and form the complete content of the answer.

[0241] In addition, it also includes the information supplement stage. Specifically, after the tree structure is formed, more information can be added to the existing tree structure according to additional processing logic. This method is suitable for complex structure annotation or scenarios where annotation rules are modified later in the annotation process. For example, see Figure 7 As shown, when marking answers to math problems, the algorithm requires that, starting in the next version, additional answer values ​​and related "Answer" text be added, while already marked historical data information cannot be modified. In this case, you can choose to select a small answer box within the answer box to record the answer value and "Answer" information. Then, after step E of step 5, add traversal logic for the answer boxes to find the information supplement boxes under the large answer box. This supplementary information can be used to supplement the question structure tree.

[0242] By executing the above scenario steps, the following effects can be achieved: the hierarchical information of the test questions is generated in the later stage according to the frame relationship, which means that the annotators do not need to annotate redundant hierarchical information during the annotation process, saving annotation time and improving the annotation speed; the hierarchical relationship generated in the later stage according to the frame relationship allows for rapid readjustment of the data according to business needs, even including comprehensive adjustments to large amounts of historically annotated data (only historical data needs to be retrieved and a new reorganization process executed), which is very critical for the early stages of the project where product functions are being polished and the required data structure has not yet been fully finalized; the bottom-up question frame annotation method combined with the top-down question structure tree generation method determines its unified underlying question level depth, which creates great convenience for grading-type products that require consistent and standardized display.

[0243] From the above, it can be seen that the embodiment of the present application can provide a data annotation page of a data annotation platform, which can display the test image of any test question on the data annotation page, and divide the test question image into target areas. The target areas include the minimum-level answer area, the stem area associated with the answer area, and the minimum-level test question area composed of the answer area and the stem area. In addition, each minimum-level test question area is associated with the question area, and one or more minimum-level test question areas and the commonly associated question areas are combined to obtain an increasing-level test question area, and so on; then, for the marking operation of each target area, a marking box surrounding the target area is displayed in the test question image; finally, the target content associated with each marking box is displayed in the annotation analysis area in the data annotation page. Specifically, when the marking box is an answer box surrounding the answer area, the target content is answer information; when the marking box is a stem box surrounding the stem area, the target content is stem information; when the marking box is a question box, the target content is question content. The above target content is used to generate annotation data for the test question. This can be applied to the marking process of test questions or homeworks of any question type and structure, so as to automatically generate marking data for test questions or homeworks based on the target content associated with the marking box, thereby improving the marking efficiency of test questions or homeworks, so as to realize automatic grading of test questions or homeworks based on the marked data, thereby improving the grading efficiency of homeworks or test questions.

[0244] In order to better implement the above method, the embodiment of the present application also provides a test image processing device. Figure 11 As shown, the test question image processing device may include a display unit 301 , a marking unit 302 and a presentation unit 303 .

[0245] The display unit 301 is used to display the data annotation page and display the uploaded test image on the data annotation page;

[0246] a marking unit 302 for displaying a marking frame surrounding the target area in the test question image in response to a marking operation on the target area in the test question image, wherein the target area includes an answer area, a question stem area, a question area, and a test question area consisting of at least the answer area and the question stem area; and types of marking frames include an answer frame surrounding the answer area, a question stem frame surrounding the question stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area;

[0247] A display unit 303 is used to display the target content associated with the mark box in the annotation analysis area of ​​the data annotation page, and the target content is used to generate the annotation data of the test question;

[0248] Among them, when the marked box is an answer box, the associated target content is the answer information, when the marked box is a question stem box, the associated target content is the question stem information, and when the marked box is a question box, the associated target content is the question content.

[0249] In some embodiments, the data annotation page further displays an uploaded answer image, and the marking unit is further configured to:

[0250] In response to a marking operation on an answer area where each piece of answer information is located in the answer image, displaying an answer frame surrounding the answer area in the answer image;

[0251] The display unit is further configured to display the answer information associated with the answer box in the annotation analysis area of ​​the data annotation page.

[0252] In some embodiments, the test question image device further includes:

[0253] a determination unit, configured to determine, based on position information of each mark frame in the test question image, the subordinate relationship between the answer frame and the sub-test question frame, the subordinate relationship between the question stem frame and the sub-test question frame, and the subordinate relationship between the sub-test question frame and the test question frame;

[0254] a sorting unit, configured to sort the multiple marked boxes belonging to the same level based on the subordinate relationship and the position information of each marked box to obtain a sorting relationship of the marked boxes;

[0255] The generation unit is used to combine the answer information associated with each answer area in the test image, the question stem content of each question stem area, and the question content of each question area according to the subordinate relationship and the mark box sorting relationship to generate annotation data for the test question.

[0256] In some embodiments, the determination unit is further used to: calculate the overlapping area of ​​the answer box and the question frame to be detected based on the position information of the answer box and the question frame in the question image; calculate the overlapping ratio between the answer box and the question frame based on the overlapping area; and determine the subordinate relationship between the answer box and the question frame based on the overlapping ratio.

[0257] In some embodiments, the sorting unit is further used to: based on the subordinate relationship, classify the target mark frame sets with the same parent object at each level from the answer frame, the question frame and the test question frame, where the parent object is any test question frame; for each target mark frame set, determine at least one target mark frame subset based on the position information of each mark frame in the target mark frame set, and each target mark frame subset contains at least one target mark frame in the same row in the test question image; for each target mark frame set, sort multiple target mark frame subsets, and sort the target mark frames in each target mark frame subset to obtain the mark frame sorting relationship of the test question image.

[0258] In some embodiments, the sorting unit is further configured to: determine, for each target marker frame set, the upper boundary position and the lower boundary position of each target marker frame based on the position information of each marker frame; and classify, according to a first condition, at least one target marker frame in the same row in the test question image from the target marker frame set, and take the set of at least one target marker frame in the same row as the target marker frame subset; wherein the first condition is for the first target marker frame and the second target marker frame belonging to the same row position, wherein the upper boundary position of the first target marker frame is higher than the lower boundary position of the second target marker frame, and the lower boundary position of the first target marker frame is lower than the upper boundary position of the second target marker frame.

[0259] In some embodiments, the sorting unit is further used to: determine, for each target marker frame set, the overlapping length of the longitudinal boundaries between any two target marker frames in the target marker frame set based on the position information of each marker frame; determine the longitudinal overlap coefficient based on the overlapping length and the reference longitudinal edge length; the reference longitudinal edge length refers to the longitudinal edge length of the target marker frame with the smaller longitudinal boundary length among the two target marker frames; if the longitudinal overlap coefficient is greater than a preset threshold, determine that the two target marker frames are in the same row in the test question image; and treat the set of at least one target marker frame in the same row in the test question image as a target marker frame subset.

[0260] In some embodiments, the sorting unit is further used to: for each target marking frame set, sort multiple target marking frame subsets in order from upper to lower rows according to the positions of the rows in which the target marking frames in the target marking frame subsets are located, to obtain a row sorting relationship; for the position information of each target marking frame in each target marking frame subset, sort multiple target marking frames in order from left to right, to obtain a horizontal sorting relationship between multiple target marking frames in the same row position; and determine the sorting relationship between each target marking frame in the corresponding target marking frame set according to the row sorting relationship and the horizontal sorting relationship, to obtain the marking frame sorting relationship of the test image.

[0261] In some embodiments, the sorting unit is further used to: add a start mark at the target mark frame at the leftmost target mark frame in the test question image in the target mark frame subset, and add an end mark at the rightmost target mark frame in the test question image; starting from the target mark frame containing the start mark, the target mark frames in the target mark frame subset are incrementally numbered from left to right until the target mark frame containing the end mark is numbered, thereby obtaining a horizontal sorting relationship between multiple target mark frames belonging to the same row position.

[0262] In some embodiments, the test question image processing device also includes an adding unit for: identifying answer boxes in text annotation data; performing semantic analysis on the answer information in each answer box to determine the key information in the answer information and the answer category corresponding to each key information; marking a key answer marking box at the key information in each answer box, and marking the corresponding answer category for each key answer marking box.

[0263] In some embodiments, the test question image processing device also includes an adjustment unit, which is used to: determine the target content whose subordinate relationship is to be adjusted in the text annotation data, and the test question area of ​​the content to be expanded; adjust the position of the marking box corresponding to the target content to the test question box corresponding to the test question area of ​​the content to be expanded in the text annotation data, and generate updated target text annotation data.

[0264] In some embodiments, the generation unit is also used to: construct initial hierarchical structure information including answer boxes, question boxes and test question boxes based on subordinate relationships and mark box sorting relationships; if it is detected that the initial hierarchical structure information contains a target branch structure with insufficient number of levels, the target branch structure in the initial hierarchical structure information is hierarchically padded to obtain the padded target hierarchical structure information, and the target branch structure is a branch structure with a number of levels less than a preset level threshold; if it is detected that the initial hierarchical structure information does not contain the target branch structure, the initial hierarchical structure information is determined as the target hierarchical structure information; according to the target hierarchical structure information, the answer information associated with each answer area in the test question image, the question content of each question box area, and the question content of each question area are combined to generate annotation data for the test question.

[0265] From the above, it can be seen that the embodiment of the present application provides a data annotation page of a data annotation platform, which can display the test image of any test question on the data annotation page, and divide the test question image into target areas. The target areas include the minimum-level answer area, the stem area associated with the answer area, and the minimum-level test question area composed of the answer area and the stem area. In addition, each minimum-level test question area is associated with the question area, and one or more minimum-level test question areas and the commonly associated question areas are combined to obtain an increasing-level test question area, and so on; then, for the marking operation of each target area, a marking box surrounding the target area is displayed in the test question image; finally, the target content associated with each marking box is displayed in the annotation analysis area in the data annotation page. Specifically, when the marking box is an answer box surrounding the answer area, the target content is answer information; when the marking box is a stem box surrounding the stem area, the target content is stem information; when the marking box is a question box, the target content is question content. The above target content is used to generate annotation data for the test question. This can be applied to the marking process of test questions or homeworks of any question type and structure, so as to automatically generate marking data for test questions or homeworks based on the target content associated with the marking box, thereby improving the marking efficiency of test questions or homeworks, so as to realize automatic grading of test questions or homeworks based on the marked data, thereby improving the grading efficiency of homeworks or test questions.

[0266] The present application also provides a computer device, such as Figure 12 , which shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:

[0267] The computer device may include one or more processing core processors 601, one or more computer readable storage media memories 602, a power supply 603, an input unit 604 and other components. Those skilled in the art will understand that Figure 12 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0268] Processor 601 is the control center of the computer device. It connects all components of the computer device using various interfaces and circuits. It executes software programs and / or modules stored in memory 602 and accesses data stored in memory 602 to perform various computer functions and process data. Optionally, processor 601 may include one or more processing cores. Preferably, processor 601 integrates an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 601.

[0269] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and test image processing processes by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0270] The computer device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 603 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0271] The computer device may further include an input unit 604, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0272] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in the embodiment of the present application, the processor 601 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602 to implement various functions as follows:

[0273] A data annotation page is displayed, and the uploaded test question image is displayed on the data annotation page; in response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes an answer area, a stem area, a question area, and a test question area consisting of at least an answer area and a stem area, and the types of marking frames include an answer box surrounding the answer area, a stem box surrounding the stem area, a question box surrounding the question area, and a question box surrounding the test question area; target content associated with the marking frame is displayed in the annotation analysis area in the data annotation page, and the target content is used to generate annotation data for the test question; wherein, when the marking frame is an answer frame, the associated target content is answer information, when the marking frame is a stem frame, the associated target content is stem information, and when the marking frame is a question frame, the associated target content is question content.

[0274] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.

[0275] It can be seen that this solution can provide a data annotation page for a data annotation platform, which can display the test image of any test question on the data annotation page, and divide the test question image into target areas. The target areas include the minimum-level answer area, the stem area associated with the answer area, and the minimum-level test question area composed of the answer area and the stem area. In addition, each minimum-level test question area is associated with the question area, and one or more minimum-level test question areas and the commonly associated question areas are combined to obtain an increasing-level test question area, and so on; then, for the marking operation of each target area, a marking box surrounding the target area is displayed in the test question image; finally, the target content associated with each marking box is displayed in the annotation analysis area in the data annotation page. Specifically, when the marking box is an answer box surrounding the answer area, the target content is answer information; when the marking box is a stem box surrounding the stem area, the target content is stem information; when the marking box is a question box, the target content is question content. The above target content is used to generate annotation data for the test question. This can be applied to the marking process of test questions or homeworks of any question type and structure, so as to automatically generate marking data for test questions or homeworks based on the target content associated with the marking box, thereby improving the marking efficiency of test questions or homeworks, so as to realize automatic grading of test questions or homeworks based on the marked data, thereby improving the grading efficiency of homeworks or test questions.

[0276] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the test image processing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0277] A data annotation page is displayed, and the uploaded test question image is displayed on the data annotation page; in response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes an answer area, a stem area, a question area, and a test question area consisting of at least an answer area and a stem area, and the types of marking frames include an answer box surrounding the answer area, a stem box surrounding the stem area, a question box surrounding the question area, and a question box surrounding the test question area; target content associated with the marking frame is displayed in the annotation analysis area in the data annotation page, and the target content is used to generate annotation data for the test question; wherein, when the marking frame is an answer frame, the associated target content is answer information, when the marking frame is a stem frame, the associated target content is stem information, and when the marking frame is a question frame, the associated target content is question content.

[0278] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0279] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0280] Since the instructions stored in the computer-readable storage medium can execute the steps in any test question image processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any test question image processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0281] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations provided in the above embodiments.

[0282] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0283] The above is a detailed introduction to a test image processing method, device, equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A test image processing method, It is characterized in that include: Displaying a data annotation page, and displaying the uploaded test question image on the data annotation page; In response to a marking operation on a target area in the test question image, a marking frame surrounding the target area is displayed in the test question image, wherein the target area includes an answer area, a question stem area, a question area, and a test question area at least composed of the answer area and the question stem area, and the types of the marking frame include an answer frame surrounding the answer area, a question stem frame surrounding the question stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area; Displaying target content associated with the mark box in the annotation analysis area of ​​the data annotation page, the target content is used to generate annotation data of the test question; Among them, when the marked box is the answer box, the associated target content is the answer information, when the marked box is the question stem box, the associated target content is the question stem information, and when the marked box is the question frame, the associated target content is the question content.

2. The method according to claim 1, It is characterized in that The data annotation page also displays the uploaded answer image, and the method further includes: In response to a marking operation on an answer area in the answer image where answer information is located, displaying an answer frame surrounding the answer area in the answer image; The target content associated with the mark box displayed in the mark analysis area of ​​the data mark page includes: The answer information associated with the answer box is displayed in the annotation analysis area in the data annotation page.

3. The method according to claim 1, It is characterized in that The method further comprises: Determine, according to the position information of each mark frame in the test question image, the subordinate relationship between the answer frame and the test question frame, the subordinate relationship between the question frame and the test question frame, the subordinate relationship between the question frame and the test question frame, and the subordinate relationship between different test question frames; Based on the subordinate relationship and the position information of each of the mark boxes, a plurality of mark boxes belonging to the same level are sorted to obtain a sorting relationship of the mark boxes; According to the subordinate relationship and the mark box sorting relationship, the answer information associated with each answer area in the test image, the question stem content of each question stem area, and the question content of each question area are combined to generate annotation data for the test question.

4. The method according to claim 3, It is characterized in that The determining the subordinate relationship between the answer box and the test question box according to the position information of each mark box in the test question image includes: For the answer box and the question box to be detected, according to the position information of the answer box and the question box in the question image, the overlapping area of ​​the answer box and the question box is calculated; Calculate the overlap ratio between the answer frame and the question frame according to the overlap area; According to the overlapping ratio, the subordinate relationship between the answer box and the question box is determined.

5. The method according to claim 3, It is characterized in that The step of sorting the multiple marked boxes belonging to the same level based on the subordinate relationship and the position information of each marked box to obtain the marked box sorting relationship includes: Based on the subordinate relationship, a target mark frame set having the same parent object at each level is classified from the answer frame, the question frame and the test frame, where the parent object is any of the test frame; For each target mark frame set, at least one target mark frame subset is determined according to the position information of each mark frame in the target mark frame set, each of the target mark frame subsets includes at least one target mark frame in the same row in the test question image; For each target mark frame set, a plurality of target mark frame subsets are sorted, and the target mark frames within each target mark frame subset are sorted to obtain a mark frame sorting relationship of the test question image.

6. The method according to claim 5, It is characterized in that The step of determining, for each target marking frame set, at least one target marking frame subset according to position information of each marking frame in the target marking frame set includes: For each target mark frame set, determine the upper boundary position and the lower boundary position of each target mark frame according to the position information of each mark frame; According to the first condition, at least one target mark frame in the same row in the test question image is respectively classified from the target mark frame set, and the set of at least one target mark frame in the same row is used as the target mark frame subset; Among them, the first condition is that for the first target mark box and the second target mark box belonging to the same row position, the upper boundary position of the first target mark box is higher than the lower boundary position of the second target mark box, and the lower boundary position of the first target mark box is lower than the upper boundary position of the second target mark box.

7. The method according to claim 5, It is characterized in that The step of determining, for each target marking frame set, at least one target marking frame subset according to position information of each marking frame in the target marking frame set includes: For each target mark frame set, determining the overlapping length of the longitudinal boundaries between any two target mark frames in the target mark frame set according to the position information of each mark frame; Determine a longitudinal overlap coefficient according to the overlap length and a reference longitudinal edge length; the reference longitudinal edge length refers to the longitudinal edge length of the target mark frame with the smaller longitudinal boundary length of the two target mark frames; If the longitudinal overlap coefficient is greater than a preset threshold, it is determined that the two target mark frames are in the same row in the test question image; and a set of at least one target mark frame in the same row in the test question image is taken as a target mark frame subset.

8. The method according to claim 5, It is characterized in that The step of sorting the plurality of target mark frame subsets for each target mark frame set, and sorting the target mark frames in each target mark frame subset to obtain a mark frame sorting relationship of the test question image includes: For each target mark frame set, according to the position of the row where the target mark frame in the target mark frame subset is located, the plurality of target mark frame subsets are sorted in order from the upper row to the lower row to obtain a row sorting relationship; According to the position information of each target mark frame in each target mark frame subset, the multiple target mark frames are sorted from left to right to obtain a horizontal sorting relationship between the multiple target mark frames in the same row position; According to the row sorting relationship and the horizontal sorting relationship, the sorting relationship between the target mark frames in the corresponding target mark frame set is determined to obtain the mark frame sorting relationship of the test question image.

9. The method according to claim 8, It is characterized in that The step of sorting the target mark boxes from left to right for the position information of each target mark box in each target mark box subset to obtain a horizontal sorting relationship between the target mark boxes in the same row includes: In the target mark frame subset, a start mark is added to the leftmost target mark frame in the test question image, and an end mark is added to the rightmost target mark frame in the test question image; Taking the target mark frame containing the start mark as the starting point, the target mark frames in the target mark frame subset are numbered incrementally from left to right until the target mark frame containing the end mark is numbered, thereby obtaining a horizontal sorting relationship between multiple target mark frames belonging to the same row position.

10. The method according to claim 3, It is characterized in that After combining the answer information associated with each answer area in the test image, the stem content of each stem area, and the question content of each question area to generate the annotation data for the test question, the method further includes: Identifying an answer box in the text annotation data; Performing semantic analysis on the answer information in each answer box to determine key information in the answer information and the answer category corresponding to each key information; A key answer mark box is marked at the key information in each answer box, and the corresponding answer category is marked for each key answer mark box.

11. The method according to claim 3, It is characterized in that After combining the answer information associated with each answer area in the test image, the stem content of each stem area, and the question content of each question area to generate the annotation data for the test question, the method further includes: Determining the target content whose subordinate relationship is to be adjusted in the text annotation data, and the test question area whose content is to be expanded; The position of the mark frame corresponding to the target content is adjusted to the question frame corresponding to the question area of ​​the content to be expanded in the text annotation data, and updated target text annotation data is generated.

12. The method according to claim 3, It is characterized in that The step of combining the answer information associated with each answer area in the test image, the stem content of each stem area, and the question content of each question area according to the subordinate relationship and the mark box sorting relationship to generate the annotation data for the test question includes: Based on the subordinate relationship and the ordering relationship of the mark boxes, constructing initial hierarchical structure information including the answer box, the question box and the test box; If it is detected that the initial hierarchical structure information includes a target branch structure with an insufficient number of levels, the target branch structure in the initial hierarchical structure information is level-filled to obtain the level-filled target hierarchical structure information, where the target branch structure is a branch structure with a level number less than a preset level threshold; If it is detected that the initial hierarchical structure information does not include the target branch structure, determining the initial hierarchical structure information as the target hierarchical structure information; According to the target hierarchical structure information, the answer information associated with each answer area in the test question image, the question stem content of each question stem area, and the question content of each question area are combined to generate annotation data for the test question.

13. A test image processing device, It is characterized in that include: A display unit, used to display a data annotation page and display the uploaded test question image on the data annotation page; a marking unit, for displaying a marking frame surrounding the target area in the test question image in response to a marking operation on the target area in the test question image, wherein the target area includes an answer area, a stem area, a question area, and a test question area at least composed of the answer area and the stem area, and the types of the marking frame include an answer frame surrounding the answer area, a stem frame surrounding the question stem area, a question frame surrounding the question area, and a test question frame surrounding the test question area; A display unit, used to display the target content associated with the mark box in the annotation analysis area of ​​the data annotation page, wherein the target content is used to generate the annotation data of the test question; Among them, when the marked box is the answer box, the associated target content is the answer information, when the marked box is the question stem box, the associated target content is the question stem information, and when the marked box is the question frame, the associated target content is the question content.

14. A computer device, It is characterized in that It includes a processor and a memory, the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the test question image processing method according to any one of claims 1 to 12.

15. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the test question image processing method according to any one of claims 1 to 12.

16. A computer program product, It is characterized in that The computer program product comprises computer instructions, which, when executed, implement the steps in the test question image processing method described in any one of claims 1 to 12.