Test paper answering text processing method and device, equipment and storage medium
By grouping and clustering exam answer texts, semantically similar subgroups are generated, solving the problem of low efficiency in traditional marking and achieving a highly efficient marking process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional manual marking methods are inefficient, requiring a large number of teachers to spend a lot of time to complete the marking work, which is difficult to meet the needs of high efficiency.
By grouping all the answer texts to be graded, and using a preset clustering algorithm to further group the answer texts for the same question, sub-answer text groups with similar text semantics are generated, and these groups are uploaded to the grading terminal for manual grading.
It enables batch application of scoring for answers from the same group, saving grading time and manpower and improving grading efficiency.
Smart Images

Figure CN121809412A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a test paper answer text processing method and device, equipment and a storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, more and more examinations have tried to use artificial intelligence to improve the efficiency and accuracy of examinations. For example, in the test paper marking, the test paper answer content written by the examinee is first scanned to form an answer image, and objective questions can be scored and judged. However, in the traditional artificial marking process, the marking teacher needs to judge each test paper; for high-stakes and large-scale examinations, each test paper needs to be reviewed by different marking teachers multiple times. This review method requires a large number of marking teachers and a long time to complete the test paper marking work, which is difficult to meet the demand for efficiency. SUMMARY
[0003] The main purpose of the present application is to provide a test paper answer text processing method, device, equipment and readable storage medium, which can at least solve the problem of low marking efficiency in the related art.
[0004] To achieve the above purpose, the first aspect of the present application provides a test paper answer text processing method, which comprises: grouping target answer texts of all to-be-marked answer papers to obtain a plurality of answer text groups; wherein the answer text group comprises the target answer texts of the same question in all the to-be-marked answer papers; grouping the target answer texts in the answer text group based on a preset clustering algorithm to obtain a plurality of sub-answer text groups; and uploading the plurality of sub-answer text groups to a marking terminal.
[0005] The second aspect of the present application provides a test paper answer text processing device, which comprises: a first grouping module configured to group answer texts of all to-be-marked answer papers to obtain a plurality of answer text groups; wherein the answer text group comprises the target answer texts of the same question in all the to-be-marked answer papers; a second grouping module configured to group the target answer texts in the answer text group based on a preset clustering algorithm to obtain a plurality of sub-answer text groups; and an uploading module configured to upload the plurality of sub-answer text groups to a marking terminal.
[0006] The third aspect of the present application provides an electronic device, which comprises a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory, and the processor implements each step of the test paper answer text processing method provided in the first aspect of the present application when executing the computer program.
[0007] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, each step of the test paper answer text processing method provided by the first aspect of the present application is implemented.
[0008] As can be seen from the above, according to the test paper answer text processing method, device, equipment and readable storage medium provided by the present application, the answer texts of all the to-be-evaluated answer sheets are grouped to obtain a plurality of answer text groups; wherein the answer text groups include the answer texts of the same question in all the to-be-evaluated answer sheets; the answer texts in the answer text groups are grouped based on a preset clustering algorithm to obtain a plurality of sub-answer text groups; and the plurality of sub-answer text groups are uploaded to an evaluation terminal. Through the implementation of the present application, the answer texts corresponding to each question are further clustered and grouped, and sub-answer text groups with similar text semantics can be obtained, so that the remaining answer texts in the same group can be applied in batches after a few samples in the sub-answer text groups are scored, thereby saving the time and manpower of reading the answer sheets and effectively improving the reading efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0010] Figure 1 The basic flowchart of the test paper answer text processing method provided by an embodiment of the present application is shown in the figure. Figure 2 The detailed flowchart of the test paper answer text processing method provided by an embodiment of the present application is shown in the figure. Figure 3 The module diagram of the test paper answer text processing device provided by an embodiment of the present application is shown in the figure. Figure 4 The structure diagram of the electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0011] In order to make the purposes, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0012] In addition, the terms "first", "second", etc. are used only for descriptive purposes and are not to be construed as indicating or implying relative importance or an ordered ranking of the indicated technical features. Thus, features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise explicitly and specifically limited.
[0013] To solve the problem of low efficiency of related art test paper review, an embodiment of the present application provides a test paper answer text processing method, which comprises the following steps: Figure 1 The basic flowchart of the test paper answer text processing method provided by the present embodiment comprises the following steps: Step 101, grouping the target answer texts of all the to-be-evaluated answer sheets to obtain a plurality of answer text groups.
[0014] Specifically, in the present embodiment, the to-be-evaluated answer sheet can be the examination answer sheet of all the examinees in a certain region in a single examination, which can be a test card or an original test paper, and the questions of all the to-be-evaluated answer sheets are the same. Each to-be-evaluated answer sheet can contain a plurality of answer texts, which can be objective question answers or subjective question answers. Since the objective questions (such as multiple-choice questions) can be scored by recognizing the answers and comparing them with the standard answers, only the answer texts of the subjective questions can be grouped for manual scoring. The target answer texts in the present embodiment can be fill-in-the-blank questions, short-answer questions, etc. Since each answer text corresponds to a question, the answer texts of a plurality of to-be-evaluated answer sheets can be grouped according to the differences in the test questions, i.e., the answer texts of the same question are grouped together, so that each answer text group contains the answer texts of the same question in all the to-be-evaluated answer sheets.
[0015] In some embodiments of the present embodiment, the test paper answer text processing method further comprises: obtaining an image of the answer text in the to-be-evaluated answer sheet; and converting the image into a digital text to obtain the answer text of the to-be-evaluated answer sheet.
[0016] Specifically, in the present embodiment, the answer region of the subjective question, such as the handwritten answer region, is scanned by a scanner to obtain an image of the answer text. Alternatively, the entire answer sheet can be scanned completely to retain the original layout information, so that the image completely presents all the elements on the paper, such as the question region, the answer region, the examinee information region, and the paper edge positioning mark. Then, the answer region image of each question is extracted from the full image by recognizing the layout rules or positioning marks. Then, the image can be converted into a digital text by a handwriting recognition technology such as OCR (Optical Character Recognition), so as to obtain the target answer text of the to-be-evaluated answer sheet. In addition, after the image and text conversion, the abnormal text such as the illegible handwriting can be marked, and the subsequent operation can be performed after manual review.
[0017] Step 102, grouping the target answer text in the answer text group based on a preset clustering algorithm to obtain a plurality of sub-answer text groups.
[0018] Specifically, in the present embodiment, for each answer text group, the clustering algorithm can be used to further group the answer texts in the group to obtain a plurality of sub-answer text groups. Each sub-answer text group can correspond to a sub-question in the question corresponding to the answer text group (for example, the answer text group corresponds to the first question, and the sub-answer text group can correspond to the first question of the first question), or a scoring point or a type of answer of the first question (or the first question). Thus, the answer texts in the answer text group can be further grouped into a plurality of sub-groups, facilitating subsequent manual scoring.
[0019] In some embodiments of the present embodiment, grouping the target answer text in the answer text group based on a preset clustering algorithm to obtain a plurality of sub-answer text groups comprises: performing feature extraction and vector conversion on the target answer text in the answer text group based on a preset feature extraction algorithm to obtain a plurality of feature vectors; grouping all feature vectors based on a preset clustering algorithm to obtain a plurality of sub-answer text groups.
[0020] Specifically, in the present embodiment, after converting the answer texts in each answer text group into vectors recognizable by the algorithm, the vectors can be input into a preset clustering algorithm (such as K-Means, DBSCAN), which will automatically group the answer texts according to their semantic similarity. For example, for all answer texts of a single sub-question (such as 1000 answers to question 2 (3) of Chinese reading), the features of each answer text (such as keywords, semantics, structure) can be extracted first, and the extracted features can be converted into vectors, and then the clustering algorithm can be used for grouping to generate sub-answer text groups (such as 10 sub-groups) for the sub-question. The scoring criteria can be disassembled into quantifiable dimensions (such as "clarity of opinion" and "sufficiency of argument") before clustering, so that the algorithm can fit the scoring logic during clustering and avoid large classification differences with similar semantics.
[0021] Specifically, when clustering and grouping, the content theme can be used for grouping, and the theme similarity of the text vector can be used for clustering, such as using LDA to extract keyword features of the answer text and quantifying them into vectors, or using BERT or other models to generate semantic features and quantifying them into vectors, and then using K-Means clustering algorithm to group according to vector distance. For example, for a question about analyzing the advantages of new energy vehicles, the answer texts can be divided into a first group (focusing on environmental protection), a second group (focusing on policy support), and a third group (mixed angle).
[0022] The clustering grouping manner can be aligned with the dimension of the scoring standard. For example, if the viewpoint relevance accounts for 60% of the weight in the scoring standard, the grouping should be preferentially divided according to whether the topic is on topic, and if the structural integrity is a core scoring point, the structural features should be extracted as the clustering basis. Therefore, the scoring points can be clustered as core features, such as decomposing the scoring standard of a small question into quantifiable scoring points, and defining the identification features (such as keywords) for each scoring point. When clustering and grouping, the features can be extracted and converted into feature vectors, and the clustering algorithm can be used to classify the similar answers into a class, so that the obtained groups can be the first group (containing three scoring points A, B, and C), the second group (containing two scoring points A and B), and the third group (containing one scoring point A). Thus, it can be ensured that after subsequent manual scoring, the score interval of the answers in the same group is reasonable and consistent with the scoring rules.
[0023] Further, in some embodiments of the present embodiment, based on a preset feature extraction algorithm, the target answer text in the answer text group is subjected to feature extraction and vector conversion, including: determining the text type of the target answer text according to the character length of the target answer text in the answer text group; if the text type is short text, a preset short text feature encoding model is used to extract features and convert vectors of the target answer text; if the text type is long text, a preset topic model is used to extract features and convert vectors of the target answer text.
[0024] Specifically, in the present embodiment, in order to make the extracted features more consistent with the expression characteristics of the text and improve the accuracy of subsequent clustering, the corresponding algorithm can be selected according to the length of the character number of the answer text of each small question. For example, an answer text less than 50 words can be determined as short text, such as fill-in-the-blank questions and short answer questions; an answer text greater than 200 words can be determined as long text, such as essays and discussion questions.
[0025] For short texts, the system inputs the answer text into a preset semantic encoding model (such as BERT, SimCSE, RoBERTa, etc.), to obtain a high-dimensional dense semantic vector of each answer text, generally 768 dimensions. The semantic vector is extracted through the Transformer structure in the deep neural network, and the core is to use the attention mechanism (Self-Attention) to establish the dependency between the context, so that the model can capture the semantic consistency and logical association of the text when encoding, and avoid keyword sparsity.
[0026] For long texts, a topic model such as LDA can be used to extract topic features to extract macro topic words and ignore redundant information. LDA calculates the probability of each word belonging to each topic by jointly modeling the document-topic distribution and the topic-word distribution, and the several words with the highest probability in the topic are the macro topic words. By extracting topic words, the system can compress long texts into low-dimensional topic distribution vectors to provide interpretable features (topic -> score point correspondence) after semantic aggregation for manual scoring. For example, the topic words of candidate 1 are technology, development, and artificial intelligence; the topic words of candidate 2 are technology, development, medical care, and health. Based on the topic words of different candidates, the corresponding topic distribution vectors are generated to provide high-quality vector input for subsequent clustering.
[0027] Further, in some embodiments of the present embodiment, all feature vectors are grouped based on a preset clustering algorithm, including: if the text type is short text, the feature vectors are grouped based on a center iterative clustering algorithm; and if the text type is long text, the feature vectors are grouped based on a graph cut clustering algorithm.
[0028] Specifically, in the present embodiment, the vector distribution features of answer texts of different length types are different, and the answer texts can be grouped by using answer texts adapted to their length types. For example, the core information of short text answer texts is concentrated, the vectors are usually spherical after being converted by a semantic model (such as BERT), and most dimensions of the vectors carry information. A center iterative clustering algorithm such as K-Means can be used to process these low-redundancy vectors. K-Means clustering algorithm is a center iterative clustering method based on distance minimization. By randomly selecting initial center points, the distance between each sample vector and each center point is calculated, and the sample is assigned to the nearest center cluster. Then the center point position of each cluster is recalculated until the center point no longer changes significantly or the iteration count reaches the upper limit. Fast grouping can be achieved with high processing efficiency. The topics of long text answer texts are diverse, and the topic distribution feature vectors of the documents are obtained after being converted by a topic model (such as LDA), which are usually non-spherical. A graph cut clustering algorithm such as spectral clustering (Spectral Clustering) algorithm can be used to calculate the similarity between all texts to obtain a similarity matrix , and these texts are regarded as nodes in the graph, and the similarity is the weight of the edge. By using graph theory, the graph is divided into several "weakly connected" subgraphs (i.e. topic clusters), so that the connection within the cluster is strong and the connection between the clusters is weak. Local dense relationships are captured by the similarity matrix, which is not affected by global sparsity, and flexible and accurate grouping can be achieved. By using different clustering methods, the clustering results can be more consistent with the human judgment of similarity to improve the accuracy of clustering.
[0029] In some embodiments of this example, the test paper answer text processing method further includes: calculating the intra-cluster average dispersion of the sub-answer text group; if there is a target sub-answer text group whose intra-cluster average dispersion exceeds a preset threshold, then splitting the target sub-answer text group to obtain multiple sub-answer text groups; or, if there is a target sub-answer text group whose intra-cluster average dispersion exceeds a preset threshold, then marking the target sub-answer text group for manual review.
[0030] Specifically, the intra-cluster average dispersion can be used to measure the "tightness" between samples within the same cluster, reflecting the semantic consistency of samples within the group. The smaller the dispersion, the more similar the samples within the same group (such as test takers' answers), and the more reliable the grouping logic; conversely, there may be classification confusion. The intra-cluster average dispersion can be calculated by averaging the distances (i.e., individual dispersion) of all samples within the group to the cluster center (which can be understood as the average feature of the samples in the group). The intra-cluster average dispersion can be calculated using the following formula:
[0031] in, For the first Semantic vectors of individual response texts Let be the center vector of this group. This is the sample size for this group.
[0032] Therefore, after obtaining multiple sub-response text groups, the dispersion of each sub-response text group can be calculated. If the dispersion is too large (exceeding a preset threshold), the sub-response text group can be split (e.g., splitting one sub-response text group with excessive dispersion into two), or the sub-response text group can be marked as "requiring manual review". This ensures that the response texts within the sub-response text group are sufficiently similar, avoiding subsequent manual scoring bias.
[0033] Step 103: Upload multiple sub-response text groups to the marking terminal.
[0034] Specifically, in this embodiment, after grouping all the answer sheets to be graded into groups, they are uploaded to the grading terminal, where manual scoring is performed on each sub-group of answer sheets. By summing the scores of each sub-group of answer sheets, the score for that sub-question can be obtained. By combining the scores of each sub-question of all subjective questions and the scores of all objective questions, the total score for each answer sheet to be graded can be obtained.
[0035] By clustering and grouping the to-be-evaluated answer sheets, artificial scoring can be applied in batches after scoring each group, for example, the scoring teacher only needs to check 1-2 representative answers in each type of grouping, gives the reference score of this type of answer (such as A type is 8-10 points, B type is 5-7 points) according to the scoring standard, and gives the score in batches to all answers in the same group without checking each answer. For example, 100,000 fill-in-the-blank questions need to be evaluated by traditional artificial evaluation, which needs to be evaluated by 278 hours if each answer needs about 10 seconds. After clustering and grouping evaluation, for medium difficulty questions, about 80% of the examinees' answers will be in the top 10 categories, and it takes about 10 minutes to evaluate each category. These answer texts can be evaluated in nearly 2 hours; greatly improving work efficiency, saving artificial evaluation time and quantity. At the same time, it also reviews the marked groups that need to be manually reviewed to ensure the accuracy of the scoring judgment.
[0036] In some embodiments of the present embodiment, uploading a plurality of sub-answer text groups to the evaluation terminal includes: obtaining the individual dispersion of each target answer text in the sub-answer text group; uploading the plurality of sub-answer text groups, the individual dispersion, and the intra-cluster average dispersion to the evaluation terminal.
[0037] Specifically, after completing the intra-question clustering (i.e., generating sub-answer text groups), the distance (such as Euclidean distance or cosine distance) of each sample (target answer text) in the sub-answer text group to the cluster center to which it belongs can be calculated, and the individual dispersion of each sample can be obtained. In the uploading stage, the grouping result, the individual dispersion of each sample, the intra-cluster average dispersion of each sub-group, and the synchronization to the evaluation terminal can be uploaded to provide the user with the maximum dispersion sample and the dispersion of each sub-group, so as to facilitate the review of the answer texts that may have grouping deviation, thereby improving the accuracy of grouping and scoring.
[0038] Among them, the several answer texts with the highest dispersion (i.e., the answers farthest from the cluster center) can be marked to facilitate the rapid positioning of possible classification errors by artificial. If the average dispersion of a sub-group is low (such as <0.2), the representative sample can be directly scored without checking more answers; if the average dispersion is high (such as >0.5), or there are multiple high-dispersion samples, these samples can be reviewed first to determine whether the cluster needs to be split or the grouping needs to be adjusted, and then the scoring is performed to avoid misjudgment.
[0039] Based on the technical solution of the above-described embodiments of this application, the answer texts of all answer sheets to be graded are grouped to obtain multiple answer text groups; wherein, the answer text group includes the answer texts of the same question in all answer sheets to be graded; based on a preset clustering algorithm, the answer texts in the answer text group are further grouped to obtain multiple sub-answer text groups; the multiple sub-answer text groups are uploaded to the grading terminal. Through the implementation of the solution of this application, the answer texts corresponding to each question are further clustered and grouped to obtain sub-answer text groups with similar text semantics. Therefore, by grading a few samples in the sub-answer text group, the results can be applied in batches to the remaining answer texts in the same group, saving grading time and manpower and effectively improving grading efficiency.
[0040] Figure 2 The method described in this application is a refined method for processing exam paper answer text, which includes: Step 201: Obtain an image of the target answer text from the answer sheet to be graded; Step 202: Convert the image into digital text to obtain the target answer text for the questionnaire to be graded; Step 203: Group all the target response texts of the questionnaires to be graded to obtain multiple response text groups; Step 204: Based on the preset feature extraction algorithm, perform feature extraction and vector transformation on the target answer text in the answer text group to obtain multiple feature vectors; Step 205: Based on the preset clustering algorithm, group all feature vectors to obtain multiple sub-response text groups; Step 206: Obtain the individual dispersion of each target response text in the sub-response text group and the intra-cluster average dispersion of the sub-response text group; Step 207: Upload multiple sub-response text groups, individual dispersion, and intra-cluster average dispersion to the marking terminal.
[0041] Specifically, in this embodiment, a high-speed scanner is used to image the candidate's answers. Then, handwriting recognition technology such as OCR is used to convert the candidate's handwritten answers into electronic text information. Text clustering technology is then used to cluster the scoring points of each question. Finally, the clustering results are submitted to the examiners for grouping and scoring of different categories.
[0042] Specifically, after candidates answer the questions, a high-speed scanner can be used to convert their handwritten answers into electronic images. Then, based on the questions and grading requirements, recognition areas are set, and OCR recognition is performed on the answers to generate electronic text information. Different clustering algorithms can be used to cluster the answers depending on the question type. For example, for short questions such as fill-in-the-blank questions, short text clustering methods can be used, such as a combination of BERT and K-Means; for longer questions such as essay questions, long text clustering methods can be used, such as a combination of LDA and spectral clustering. The answers from different types of candidates are clustered and sent as a group, and the graders score the answers for the entire group. Considering concerns about clustering accuracy, the system interface can simultaneously display the dispersion of each answer sample from the cluster center and the average dispersion of the entire subgroup, reminding graders to pay attention to samples with large discrepancies and to understand the overall dispersion of the group. After all the clustering scores for the questions are completed, the scores of all subgroups can be merged to obtain the associated candidate's exam score.
[0043] It should be understood that the sequence number of each step in this embodiment does not imply the order in which the steps are executed. The execution order of each step should be determined by its function and internal logic, and should not constitute a unique limitation on the implementation process of this application embodiment.
[0044] Based on the above technical solution of this application embodiment, the answer texts of all answer sheets to be graded are grouped to obtain multiple answer text groups; wherein, the answer text group includes the answer texts of the same question in all answer sheets to be graded; based on a preset clustering algorithm, the answer texts in the answer text group are further grouped to obtain multiple sub-answer text groups; the multiple sub-answer text groups are uploaded to the grading terminal. Through the implementation of this application solution, the answer texts corresponding to each question are further clustered and grouped to obtain sub-answer text groups with similar text semantics. Therefore, by grading a few samples in the sub-answer text group, the results can be applied in batches to the remaining answer texts in the same group, saving grading time and manpower, and effectively improving grading efficiency.
[0045] Figure 3 This application provides an embodiment of a test paper answer text processing device, which can be applied to the aforementioned test paper answer text processing method. For example... Figure 3 As shown, the test paper answer text processing device mainly includes: The first grouping module 301 is used to group the answer texts of all the answer sheets to be evaluated into multiple answer text groups; wherein, the answer text group includes the answer texts of the same question in all the answer sheets to be evaluated; The second grouping module 302 is used to group the answer texts in the answer text group based on a preset clustering algorithm to obtain multiple sub-answer text groups; Upload module 303 is used to upload multiple sub-answer text groups to the marking terminal.
[0046] In some embodiments of this example, the test paper answer text processing device further includes: an image-to-text conversion module, used to acquire an image of the target answer text in the answer sheet to be graded; and to convert the image into digital text to obtain the target answer text of the answer sheet to be graded.
[0047] In some embodiments of this example, the second grouping module is specifically used to: extract features and transform vectors from the target answer text in the answer text group based on a preset feature extraction algorithm to obtain multiple feature vectors; and group all feature vectors based on a preset clustering algorithm to obtain multiple sub-answer text groups.
[0048] In some embodiments of this example, the test paper answer text processing device further includes: a calculation module, used to calculate the intra-cluster average dispersion of the sub-answer text group; if there is a target sub-answer text group whose intra-cluster average dispersion exceeds a preset threshold, then the target sub-answer text group is split to obtain multiple sub-answer text groups; or, if there is a target sub-answer text group whose intra-cluster average dispersion exceeds a preset threshold, then the target sub-answer text group is marked for manual review.
[0049] In some implementations of this embodiment, the upload module is specifically used to: obtain the individual dispersion of each target answer text in the sub-answer text group; and upload multiple sub-answer text groups, individual dispersion, and intra-cluster average dispersion to the marking terminal.
[0050] In some embodiments of this example, the second grouping module is further configured to: determine the text type of the target response text based on the character length of the target response text in the response text group; if the text type is short text, then use a preset short text feature encoding model to extract features and perform vector conversion on the target response text; if the text type is long text, then use a preset topic model to extract features and perform vector conversion on the target response text.
[0051] In some embodiments of this example, the second grouping module is further used to: group the feature vectors based on a center-based iterative clustering algorithm if the text type is short text; and group the feature vectors based on a graph-based segmentation clustering algorithm if the text type is long text.
[0052] It should be noted that the test paper answer text processing methods in the foregoing embodiments can all be implemented based on the test paper answer text processing device provided in this embodiment. Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the test paper answer text processing device described in this embodiment can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0053] Based on the technical solution of the above embodiments of this application, the answer texts of all answer sheets to be graded are grouped to obtain multiple answer text groups; wherein, the answer text group includes the answer texts of the same question in all answer sheets to be graded; based on a preset clustering algorithm, the answer texts in the answer text group are further grouped to obtain multiple sub-answer text groups; the multiple sub-answer text groups are uploaded to the grading terminal. Through the implementation of the solution of this application, the answer texts corresponding to each question are further clustered and grouped to obtain sub-answer text groups with similar text semantics. Therefore, by grading a few samples in the sub-answer text group, the results can be applied in batches to the remaining answer texts in the same group, saving grading time and manpower and effectively improving grading efficiency.
[0054] Figure 4 An electronic device is provided as an embodiment of this application. This electronic device can be used to implement the test paper answer text processing method in the foregoing embodiments, mainly including: The system includes a memory 401, a processor 402, and a computer program 403 stored on the memory 401 and executable on the processor 402. The memory 401 and the processor 402 are communicatively connected. When the processor 402 executes the computer program 403, it implements the method described in the foregoing embodiments. The number of processors can be one or more.
[0055] The memory 401 can be a high-speed random access memory (RAM) or a non-volatile memory, such as a disk storage device. The memory 401 is used to store executable program code, and the processor 402 is coupled to the memory 401.
[0056] Furthermore, embodiments of this application also provide a computer-readable storage medium, which may be disposed in the aforementioned electronic device, and the computer-readable storage medium may be as described above. Figure 4 The memory in the illustrated embodiment.
[0057] The computer-readable storage medium stores a computer program that, when executed by a processor, implements the test paper answer text processing method described in the foregoing embodiments. Furthermore, the computer-readable storage medium can also be a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk, or any other medium capable of storing program code.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0059] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0060] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0061] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, external hard drives, ROM, RAM, magnetic disks, or optical disks.
[0062] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0063] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0064] The above is a description of the test paper answer text processing method, apparatus, device and readable storage medium provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for processing exam paper answer text, characterized in that, include: The target answer texts of all the answer sheets to be evaluated are grouped to obtain multiple answer text groups; wherein, the answer text group includes the target answer texts of the same question in all the answer sheets to be evaluated; Based on a preset clustering algorithm, the target response text in the response text group is grouped to obtain multiple sub-response text groups; Upload multiple sub-response text groups to the marking terminal.
2. The test paper answer text processing method according to claim 1, characterized in that, Also includes: Obtain an image of the target answer text from the answer sheet to be graded; The image is converted into digital text to obtain the target answer text of the questionnaire to be graded.
3. The test paper answer text processing method according to claim 1, characterized in that, The method involves grouping the target response text within the response text group using a pre-defined clustering algorithm to obtain multiple sub-response text groups, including: Based on a preset feature extraction algorithm, feature extraction and vector transformation are performed on the target response text in the response text group to obtain multiple feature vectors; Based on a preset clustering algorithm, all the feature vectors are grouped to obtain multiple sub-response text groups.
4. The test paper answer text processing method according to claim 1, characterized in that, Also includes: Calculate the intra-cluster average dispersion of the sub-response text group; If there is a target sub-response text group whose average dispersion within the cluster exceeds a preset threshold, then the target sub-response text group is split to obtain multiple sub-response text groups; Alternatively, if there is a target sub-response text group whose average dispersion within the cluster exceeds a preset threshold, then the target sub-response text group is marked for manual review.
5. The test paper answer text processing method according to claim 4, characterized in that, Uploading multiple sub-answer text groups to the marking terminal includes: Obtain the individual discreteness of each target response text in the sub-response text group; The multiple sub-response text groups, the individual dispersion, and the average dispersion within the cluster are uploaded to the marking terminal.
6. The test paper answer text processing method according to claim 3, characterized in that, The step of extracting features and transforming vectors from the target response text in the response text group based on a preset feature extraction algorithm includes: The text type of the target response text is determined based on the character length of the target response text in the response text group; If the text type is short text, then a preset short text feature encoding model is used to extract features and transform vectors in the target response text; If the text type is long text, then a preset topic model is used to extract features and transform vectors from the target response text.
7. The test paper answer text processing method according to claim 6, characterized in that, The step of grouping all the feature vectors based on a preset clustering algorithm includes: If the text type is short text, the feature vectors are grouped based on a center-based iterative clustering algorithm; If the text type is long text, the feature vectors are grouped based on a graph-cutting clustering algorithm.
8. A test paper answer text processing device, characterized in that, include: The first grouping module is used to group the answer texts of all the answer sheets to be evaluated into multiple answer text groups; wherein, the answer text group includes the answer texts of the same question in all the answer sheets to be evaluated; The second grouping module is used to group the answer texts in the answer text group based on a preset clustering algorithm to obtain multiple sub-answer text groups; The upload module is used to upload multiple sub-answer text groups to the marking terminal.
9. An electronic device, characterized in that, Includes memory and processor, of which: The processor is used to execute computer programs stored in the memory; When the processor executes the computer program, it implements the steps in the test paper answer text processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the test paper answer text processing method as described in any one of claims 1 to 7.