A question content recognition and processing system and method for test paper entry
Through the question content recognition and processing system for test paper entry, the p-field identification in the html text format automatically detects and marks the question number, solving the problem of fast and accurate and high operational cost of the test paper entry system, and achieving efficient test paper entry and test bank management.
Patent Information
- Application Number
- CN202111074655.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-14
AI Technical Summary
The rapid accuracy and high operating cost of the test question entry system in the prior art lead to high work intensity of teachers and low efficiency in question bank management.
A question content recognition and processing system for exam paper entry is adopted, including text conversion, field identification extraction, judgment and division modules, automatically detect and mark the question number and question content, and process it through the p-field identification in the html text format, correct and manually check to ensure accuracy.
It realizes automated, fast and accurate test paper entry, reduces the intensity of teachers' work, improves the entry efficiency and accuracy, and ensures the seriousness, authority and fairness of the test paper question bank.
Smart Images

Figure CN113821598B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of test paper question bank management, and in particular, to a question content recognition and processing system and method for test paper input. Background Art
[0002] A question bank is "a collection of questions in a certain subject implemented in a computer system according to a certain educational measurement theory". Establishing a question bank system can standardize, systematize, and program the work of question selection and management. Utilizing the intelligent functions of the question bank, automatic scientific test paper compilation can be achieved, teaching resource sharing can be realized, the examination process can be standardized, and the work intensity of teachers can be reduced. Test paper compilation depends on question input, but many studies focus on test paper compilation technology and ignore the research on question input. As is well known, for a relatively complete question bank system, in order to ensure the scientificity and effectiveness of questions, thousands of questions are required, and the workload of writing and inputting these questions is extremely huge.
[0003] Therefore, in view of the above scenario, how to make the test paper question input system faster, more accurate, and lower in operation cost is an urgent problem to be solved at present. Summary of the Invention
[0004] In view of the technical problems in the prior art, the present invention provides a question content recognition and processing system and method for test paper input.
[0005] A question content recognition and processing system for test paper input, the system includes a text conversion module, a first field identifier extraction module, a first judgment module, a second field identifier extraction module, a second judgment module, and a division module, wherein:
[0006] The text conversion module is connected to the first field identifier extraction module; the text conversion module is used to convert the test paper text file into a set-format text with p field identifiers; the p field identifiers include Field start, with field end;
[0007] The first field identifier extraction module is connected to the text conversion module and the first judgment module; the first field identifier extraction module is used to extract field, and correspondingly form the first Label queue;
[0008] The first judgment module is connected to the first field identifier extraction module and the second field identifier extraction module; the first judgment module is used to judge the first for each in the label queue Whether the first text content after the field is a non - zero positive integer. If so, mark the non - zero positive integer as the corresponding Numeric label of the field label;
[0009] The second field identifier extraction module is connected to the first judgment module and the second judgment module; the second field identifier extraction module is used for the first Extract the fields with numeric labels from the label queue field, and correspondingly form the second Label queue;
[0010] The second judgment module is connected to the second field identifier extraction module and the division module; the second judgment module is used to judge the second Whether the numeric labels in the label queue are arranged in increasing order with a step size of 1. If so, use the fields corresponding to the increasing numeric labels to form the third Label queue;
[0011] Partition module, connected to the text conversion module, the first field identifier extraction module, and the second judgment module; the partition module is used to according to the third in the label queue ID of the field, find this field in the first position in the label queue, and correspond A non-zero positive integer after the field is marked as the question number, and the question content of each question is divided according to the position of the question number in the text in the set format.
[0012] Furthermore, the system further includes a storage module, and the storage module is connected to the division module; the storage module is used to store the question number and the question content after the question number.
[0013] Furthermore, the test paper text file is a word format text that only retains the question number and the question content.
[0014] Furthermore, the text in the set format is html text.
[0015] Furthermore, the system further includes a correction module, and the correction module is connected to the second judgment module; the correction module is used for the second When the numeric labels in the label queue are not arranged in increasing order with a step size of 1, check for numeric labels that do not have an increasing relationship with the previous and next numeric labels with a step size of 1, and list them as pending processing values. If the ratio of the number of pending processing values to the total number of numeric labels is less than the preset abnormal ratio value, then the pending processing values and the corresponding field in the second Deleted from the label queue.
[0016] Further, the system further includes a manual inspection module, which is connected to the second judgment module and the correction module; the manual inspection module is used for the second When the numeric labels in the label queue are not arranged in increasing order with a step size of 1 and there is no numeric label that does not have an increasing relationship with the previous and next numeric labels with a step size of 1, for the second numeric labels in the label queue and the corresponding Delete or add fields.
[0017] The present invention further includes a method for identifying and processing the content of questions for test paper entry, including:
[0018] Convert the test paper text file into a set-format text with p-field identifiers; the p-field identifiers include Field start, with Field completion;
[0019] Extract from the formatted text field, and correspondingly form the first Label queue;
[0020] Judge the first for each in the label queue Whether the first text content after the field is a non - zero positive integer. If so, mark the non - zero positive integer as the corresponding Numeric label of the field label;
[0021] In the first Extract the fields with numeric labels from the label queue field, and correspondingly form the second Label queue;
[0022] Judge the second Whether the numeric labels in the label queue are arranged in increasing order with a step size of 1. If so, use the fields corresponding to the increasing numeric labels to form the third Label queue;
[0023] According to the third in the label queue ID of the field, find this field in the first position in the label queue, and correspond A non-zero positive integer following a field is marked as the question number;
[0024] Based on the position of the question number in the text of the set format, the question content of each question is divided.
[0025] Furthermore, the method further includes:
[0026] Store the question number and the question content after the question number.
[0027] Furthermore, the method further includes:
[0028] In the second When the numerical tags in the tag queue are not arranged in increasing order with a step size of 1, check for numerical tags that do not have an increasing relationship with a step size of 1 with both the previous numerical tag and the next numerical tag, and list them as numerical values to be processed;
[0029] If the ratio of the number of numerical values to be processed to the total number of numerical tags is less than the preset abnormal ratio value, then the numerical values to be processed and their corresponding field in the second Deleted from the label queue.
[0030] The question content recognition and processing system and method for test paper entry of the present invention can automatically detect the test paper content and question numbers, and then accurately and quickly enter them into the question bank, reducing the workload of teachers, without a large amount of manual intervention, and can greatly improve work efficiency and accuracy, thus effectively ensuring the quality of test paper entry in question bank management, and also ensuring the smooth progress of test paper question bank management work, effectively guaranteeing the solemnity, authority and fairness of the test paper question bank. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is a structural composition diagram (one) of a question content recognition and processing system for test paper entry according to an embodiment of the present invention;
[0033] Figure 2 It is a structural composition diagram (two) of a question content recognition and processing system for test paper entry according to an embodiment of the present invention;
[0034] Figure 3 It is a structural composition diagram (three) of a question content recognition and processing system for test paper entry according to an embodiment of the present invention;
[0035] Figure 4 It is a structural composition diagram (four) of a question content recognition and processing system for test paper entry according to an embodiment of the present invention;
[0036] Figure 5 It is a step flow chart (one) of a question content recognition and processing method for test paper entry according to an embodiment of the present invention;
[0037] Figure 6 It is a step flow chart (two) of a question content recognition and processing method for test paper entry according to an embodiment of the present invention;
[0038] Figure 7 It is a step flow chart (three) of a question content recognition and processing method for test paper entry according to an embodiment of the present invention. Detailed Embodiments
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the protection scope of the present invention.
[0040] The present invention provides an embodiment of a question content recognition and processing system for test paper entry, as Figure 1 shown, the system includes a text conversion module 10, a first field identifier extraction module 20, a first judgment module 30, a second field identifier extraction module 40, a second judgment module 50, and a division module 60, where:
[0041] The text conversion module 10 is connected to the first field identifier extraction module 20; the text conversion module 10 is used to convert the test paper text file into a set-format text with p field identifiers.
[0042] The embodiments of the present invention do not limit the format of the test paper text file. However, in order to realize the recognition of the test paper question content by the present invention, the content of the test paper text file only includes the question number and the question content, without other content (such as the test paper title, test paper description, exam description, precautions, question types, etc., which are irrelevant to the questions). Preferably, the test paper text file in the embodiments of the present invention is a word format text that only retains the question number and the question content. And the set-format text with p field identifiers in this embodiment can be an html text with p field identifiers. The p field identifier in the html text is used for text segmentation. For example, if text needs to be segmented and start a new line, there is a p field identifier between the two paragraphs. The p field identifier in the embodiments of the present invention includes Field start, with Field completion. For example:
[0043] Question content 1
[0044] Question content 2
[0045] The first field identifier extraction module 20 is connected to the text conversion module 10 and the first judgment module 30; the first field identifier extraction module 20 is used to extract from the formatted text field, and correspondingly form the first Label queue.
[0046] Each text paragraph in the formatted text starts with field, and each field has a unique ID, and the question content after the field can be found according to the ID. For example, the beginning of the question text and option text in html has field, and each question (including the stem, options, questions) has at least one field. The first - field identification extraction module 20 combines these fields to form the first Label queue.
[0047] The first judgment module 30 is connected to the first field identifier extraction module 20 and the second field identifier extraction module 40; the first judgment module 30 is used to judge the first for each in the label queue Whether the first text content after the field is a non - zero positive integer. If so, mark the non - zero positive integer as the corresponding Numeric label for field marking.
[0048] The first determination module 30 determines the first for each in the label queue Whether the first text content after the field is a non - zero positive integer to filter out the fields belonging to the question number. If it is judged not to be a non - zero positive integer (such as Chinese characters, letters, Roman numerals, negative numbers, etc.), then for this field, if the judgment is not a non - zero positive integer (such as Chinese characters, letters, Roman numerals, negative numbers, etc.), then for this The field does not mark numerical labels.
[0049] The second field identification extraction module 40 is connected to the first judgment module 30 and the second judgment module 50; the second field identification extraction module 40 is used for the first Extract the fields with numeric labels from the label queue field, and correspondingly form the second Label queue.
[0050] The second judgment module 50 is connected to the second field identifier extraction module 40 and the division module 60; the second judgment module 50 is used to judge the second Whether the numeric labels in the label queue are arranged in increasing order with a step size of 1. If so, use the fields corresponding to the increasing numeric labels to form the third Label queue.
[0051] The second determination module 50 of this embodiment pairs the second Check the label queue, check whether the numeric labels in the second label queue are arranged in increasing order with a step size of 1, such as 1, 2, 3, ……, n. If the numeric labels in the second label queue meet the above requirements, then use the fields corresponding to the increasing numeric labels to form the third label queue, the third of the label queue The field is the starting field of each question. The partitioning module 60 is connected to the text conversion module 10, the first field identification extraction module 20, and the second judgment module 50; the partitioning module 60 is used to according to the third in the label queue ID of the field, find the position of the field in the first label queue, and correspond A non-zero positive integer following the field is marked as the question number, and the question content of each question is divided according to the position of the question number in the formatted text.
[0052] The division module 60 of this embodiment first depends on the thirdin the label queue ID of the field, find the position of the field in the first label queue, so as to determine which fields in the first label queue are the starting fields of the questions. Therefore, when the numerical labels in the corresponding A non-zero positive integer following a field is marked as the question number, and the text is divided in the html text to separate the content of each question.
[0053] Specifically, for example Figure 2 As shown, the system of the embodiment of the present invention further includes a storage module 70, and the storage module 70 is connected to the division module 60; the storage module 70 is used to store the question number and the question content after the question number. After the division module 60 divides the content of each question, the storage module 70 stores the question number and the corresponding question content.
[0054] The storage module 70 of this embodiment can be a server running software to implement the storage of relevant information.
[0055] Specifically, for example Figure 3 As shown, the system of this embodiment further includes a correction module 80, and the correction module 80 is connected to the second judgment module 50; the correction module 80 is used for the second label queue are not incremented by a step size of 1, check the numerical labels that are not incremented by a step size of 1 with both the previous numerical label and the next numerical label, and list them as the numerical labels to be processed. If the ratio of the number of numerical labels to be processed to the total number of numerical labels is less than the preset abnormal ratio value, then the numerical labels to be processed and the corresponding fields in the second Deleted from the label queue.
[0056] If the second label queue are not incremented by a step size of 1, then the second label queue needs to be corrected. Specifically, find a numerical label that satisfies that the numerical label is not incremented by a step size of 1 with both the previous numerical label and the next numerical label, and list these numerical labels as the numerical labels to be processed. Before this step, an abnormal ratio value is preset. The meaning of the abnormal ratio value is the ratio of the abnormal numerical labels to the total numerical labels, and it can be specifically set according to the number of total numerical labels. If the ratio of the number of numerical labels to be processed to the total number of numerical labels is less than the preset abnormal ratio value, it means that the number of numerical labels to be processed is small, then the numerical labels to be processed and the corresponding fields in the second Deleted from the label queue so that the remaining numerical labels are arranged in increasing order with a step size of 1.
[0057] By the correction module 80 for the second content of the label queue are corrected, and then the second judgment module 50 of the embodiment of the present invention corrects the second Determine whether the numerical tags in the tag queue are incremented by a step size of 1, and implement other corresponding functions.
[0058] Specifically, as Figure 4 shown, the system of this embodiment further includes a manual inspection module 90, and the manual inspection module 90 is connected to the second judgment module 50 and the correction module 80; the manual inspection module 90 is used for the second label queue, and when the numerical labels in the label queue are not incremented by a step size of 1 and there is no numerical label that is not incremented by a step size of 1 with both the previous numerical label and the next numerical label, for the numerical labels in the second label queue and the corresponding Delete or add fields.
[0059] In this embodiment, the manual inspection module 90 pairs the second When the numerical labels in the label queue are not incremented by a step size of 1 and there is no numerical label that is not incremented by a step size of 1 with both the previous numerical label and the next numerical label, for the numerical labels in the second label queue and the corresponding fields, delete or add them so that the modified second The numerical tags in the tag queue are arranged in ascending order with a step size of 1.
[0060] The present invention also provides an embodiment of a method for identifying and processing the content of questions for test paper entry, as Figure 5 shown. This embodiment includes the following steps:
[0061] Step S10: Convert the test paper text file into a set-format text with a p-field identifier, and the p-field identifier includes starting from the field, with Field completion.
[0062] Step S20: Extract from the formatted text field, and correspondingly form the first Label queue.
[0063] Step S30: Determine the first each Whether the first text content after the field is a non-zero positive integer. If so, execute step S40.
[0064] Step S40: Mark the non-zero positive integer as corresponding Numeric label of the field mark.
[0065] Step S50: At the first Extract the fields with numerical labels from the label queue, and correspondingly form the second fields, and correspondingly form the third Label queue.
[0066] Step S60: Determine the second Whether the numerical tags in the tag queue are arranged in increasing order with a step size of 1. If so, execute step S70.
[0067] Step S70: Corresponding to the increasing numerical tags fields corresponding to form the third Label queue.
[0068] Step S80: According to the third in the label queue ID of the field, find the position of the field in the first label queue, and correspond A non-zero positive integer after the field is marked as the question number.
[0069] Step S90: Divide the content of each question according to the position of the question number in the text in the set format.
[0070] As Figure 6 shown, the method for identifying and processing the question content for test paper entry according to the embodiment of the present invention further includes:
[0071] Step S100: Store the question number and the question content after the question number.
[0072] As Figure 7 shown, the method for identifying and processing the question content for test paper entry according to the embodiment of the present invention further includes:
[0073] When step S60 determines the second Whether the numerical tags in the tag queue are incrementally arranged with a step size of 1. If not, execute step S110.
[0074] Step S110: Check out the numerical tags that are not incrementally related with a step size of 1 to both the previous numerical tag and the next numerical tag, and list them as the numerical values to be processed.
[0075] Step S120: If the ratio of the number of numerical values to be processed to the total number of numerical tags is less than the preset abnormal ratio value, then the numerical values to be processed and the corresponding ones fields in the second Deleted from the label queue.
[0076] After that, step S60 is executed again for the second Judge the numerical tags in the tag queue to see if they are arranged in ascending order with a step size of 1.
[0077] For the method for identifying and processing the question content for test paper entry in the embodiments of the present invention, the execution of related steps can refer to the explanations of the related embodiments of the system for identifying and processing the question content for test paper entry in the present invention, which will not be elaborated here.
[0078] The system and method for identifying and processing the question content for test paper entry in the embodiments of the present invention can automatically detect the test paper content and question numbers, and then accurately and quickly enter them into the question bank, reducing the workload of teachers, without the need for a large amount of manual intervention, and can greatly improve work efficiency and accuracy, thus effectively ensuring the quality of test paper entry in question bank management, and also ensuring the smooth progress of the work of test paper question bank management, effectively guaranteeing the solemnity, authority and fairness of the test paper question bank.
[0079] The above further describes the present invention with the aid of specific embodiments. However, it should be understood that this specific description should not be construed as a limitation on the essence and scope of the present invention. Various modifications made by those of ordinary skill in the art to the above embodiments after reading this specification all fall within the scope protected by the present invention.
Claims
1. A question content recognition and processing system for test paper entry, characterized in that, The system includes a text conversion module, a first field identifier extraction module, a first judgment module, a second field identifier extraction module, a second judgment module, and a division module, where: The text conversion module is connected to the first field identifier extraction module; the text conversion module is used to convert the test paper text file into a set-format text with a p field identifier; the p field identifier includes The start of the field, with field end; The first field identifier extraction module is connected to the text conversion module and the first judgment module; the first field identifier extraction module is used to extract the the field, and correspondingly form the first tag queue; The first judgment module is connected to the first field identifier extraction module and the second field identifier extraction module; the first judgment module is used to judge the first For each in the tag queue whether the first text content after the field is a non-zero positive integer. If so, mark the non-zero positive integer as the corresponding numeric tag of the field; The second field identifier extraction module is connected to the first judgment module and the second judgment module; the second field identifier extraction module is used for the first Extract the fields with the said numeric tags from the tag queue, and correspondingly form the second tag queue; Whether the numeric tags in the tag queue are arranged in an increasing order with a step size of 1. If so, use the fields corresponding to the increasing numeric tags to form the third The second judgment module is connected to the second field identifier extraction module and the division module; the second judgment module is used to judge the second tag queue; tag queue; Find the ID of the field in the tag queue The division module is connected to the text conversion module, the first field identifier extraction module, and the second judgment module; the division module is used to according to the third tag queue, find the position of this field in the first tag queue, mark the non-zero positive integer after the corresponding field as the question number, and according to the position of the question number in the set format text, divide the question content of each question; It further includes a correction module, and the correction module is connected to the second judgment module; The correction module is used to check out the numeric tags that are not in an increasing relationship with a step size of 1 with both the previous numeric tag and the next numeric tag when the numeric tags in the second tag queue are not arranged in an increasing order with a step size of 1, list them as the numeric tags to be processed. If the ratio of the number of the numeric tags to be processed to the total number of numeric tags is less than the preset abnormal ratio value, then delete the numeric tags to be processed and the corresponding fields in the second tag queue; When the numeric tags in the tag queue are not arranged in an increasing order with a step size of 1, and there is no numeric tag that is not in an increasing relationship with a step size of 1 with both the previous numeric tag and the next numeric tag, delete or add the numeric tags in the second It further includes a manual inspection module, which is connected to the second judgment module and the correction module; the manual inspection module is used for the second tag queue and the corresponding fields; It further includes a storage module, and the storage module is connected to the division module; The storage module is used to store the question number and the question content after the question number.
2. The question content recognition and processing system for test paper entry according to claim 1, characterized in that The test paper text file is a word format text that only retains the question number and the question content.
3. The question content recognition and processing system for test paper entry according to claim 1, characterized in that The set format text is html text.
4. The question content recognition and processing system for test paper entry according to claim 1, characterized in that Including:
5. A recognition processing method for the question content recognition processing system for test paper entry described in claim 1, characterized in that, The start of the field, with Convert the test paper text file into a text in a set format with the p-field identifier; the p-field identifier includes the field, and correspondingly form the first Field end; Extract the tag queue; For each in the tag queue Determine the first whether the first text content after the field is a non-zero positive integer. If so, mark the non-zero positive integer as the corresponding numeric tag of the field; tag queue; In the first Extract the fields with the said numeric tags from the tag queue, and correspondingly form the second tag queue; Whether the numeric tags in the tag queue are arranged in an increasing order with a step size of 1. If so, use the fields corresponding to the increasing numeric tags to form the third Determine the second tag queue; tag queue; Find the ID of the field in the tag queue According to the third tag queue, find the position of this field in the first tag queue, mark the non-zero positive integer after the corresponding field as the question number; According to the position of the question number in the set format text, divide the question content of each question; It further includes: Storing the question number and the question content after the question number; It further includes: In the second When the numerical tags in the tag queue are not arranged in ascending order with a step size of 1, check for numerical tags that do not have an ascending relationship with a step size of 1 with both the previous numerical tag and the next numerical tag, and list them as numerical tags to be processed; If the ratio of the number of the values to be processed to the total number of the numerical tags is less than a preset abnormal ratio value, then the values to be processed and the corresponding field in the second Delete from the tag queue.
Citation Information
Patent Citations
Test paper splitting method and system
CN110674722A