A subjective question intelligent marking method and system supporting multi-text mixing
By employing multi-granularity alignment and non-linear fusion strategies, combined with differentiated quality assessment for Chinese and English, this approach addresses the shortcomings of existing intelligent marking systems in handling mixed multi-text answers. It enables scoring of the coherence and structural integrity of Chinese essays and sentence-by-sentence detailed feedback on English essays, providing personalized comments and optimized model essays, thereby improving the accuracy and efficiency of scoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG SHIJIJINBANG SCI & EDUCATION & CULTURE
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-22
Smart Images

Figure CN121599806B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of online education technology, specifically to a method and system for intelligent grading of subjective questions that supports multi-text mixing. Background Technology
[0002] With the increasing demand for automated marking technology in the education sector, the development of intelligent subjective question marking systems has gradually become a research hotspot. Traditional subjective question marking methods rely on manual marking or semi-automated marking tools, which are inefficient and susceptible to human bias when marking large numbers of papers. To address this issue, intelligent marking methods based on natural language processing and machine learning have emerged in recent years. Intelligent systems can not only automatically assign scores but also provide detailed comments and suggestions to help students better understand and improve their answers. However, existing technologies still have many limitations, especially in handling complex text types and generating personalized comments and optimized model answers.
[0003] Currently, most intelligent scoring systems for subjective questions rely on keyword matching, semantic vector similarity, or large-scale model scoring methods. For example, scoring methods based on sentence or paragraph similarity can quickly identify key points and elements of the answer, but these methods often fail to effectively capture subtle differences in the text or understand the deeper meaning of the language. Furthermore, existing intelligent grading systems are mostly limited to simple score calculations and fail to truly generate personalized comments or model essays. For English essays, although grammar checking tools exist, few systems can combine syntactic analysis and vocabulary collocation for detailed, sentence-by-sentence feedback and generate optimized model essays for students to reference.
[0004] While existing intelligent marking systems perform well on some simple tasks, they still have significant shortcomings: First, they lack the ability to handle complex multi-text mixed answers, making it difficult to accurately assess the integration and logic of information across paragraphs; second, most existing systems cannot generate personalized comments or optimized model essays, remaining only at the "scoring" level; and finally, they lack in-depth analysis and optimization suggestions for sentence-by-sentence critique of English essays (especially syntactic errors and collocations). Therefore, how to provide accurate scoring, personalized comment generation, optimized model essays, and sentence-by-sentence critique and collocation optimization for subjective questions involving multi-text mixed answers has become an urgent technical problem to be solved. Summary of the Invention
[0005] In order to solve the above-mentioned technical problems, this application proposes the following technical solution:
[0006] In a first aspect, embodiments of this application provide a method for intelligent grading of subjective questions that supports multi-text mixing, including:
[0007] The system retrieves the question stem, reference answer, and scoring procedure, and structures the scoring procedure into scoring points and weights. It then segments the candidate's answer and generates text segments and sentence representations.
[0008] The text segment and the scoring point are determined and fused to obtain a soft alignment result. The contributions of multiple segments to the same point are non-linearly saturated and fused to obtain the point coverage and calculate the point score.
[0009] The relevance of the text segment to the prompt, inter-segment contradictions, language quality, and format compliance scores are calculated; for Chinese essays, coherence and structural integrity scores are further calculated; for English essays, grammatical errors are checked sentence by sentence and vocabulary collocation quality scores are calculated.
[0010] The key points, quality score, and each penalty item are weighted, merged, and calibrated to be mapped to the full score range to output the final score;
[0011] The system outputs the location of key paragraphs and sentences for each scoring point as interpretable evidence, generates comments based on the evidence, and generates optimized model essays for students to refer to; it provides detailed feedback on English essays sentence by sentence, outputting the original sentence, grammatical error annotations, and revised sentences, and updates calibration parameters based on manual results when a review is triggered.
[0012] In one possible implementation, the steps of obtaining the question stem, reference answer, and scoring procedure, structuring the scoring procedure into scoring points and weights, segmenting the candidate's answer, and generating text segments and sentence representations include:
[0013] Obtain the question stem information of the subjective questions to be graded Reference answer information and scoring procedure information;
[0014] The scoring procedure information is structured into a set of scoring points. and corresponding weights ;
[0015] The reference answer is broken down into a library of evidence fragments associated with each scoring point. The Contains multiple pieces of evidence ;
[0016] The test takers' answers were divided into text segments based on paragraphs, bullet points, and quotation blocks. :
[0017]
[0018] Each text segment is then segmented into sentences to obtain a set of sentences. ;
[0019] For the above , , and The encoding process generates the question stem information, evidence fragments, paragraph-level representations of each text segment, and sentence-level representations of each sentence.
[0020] , , ,
[0021]
[0022] The process involves identifying the language type of the test taker's answers to determine whether the essay is in Chinese or English.
[0023]
[0024] in: This represents the set of text segments after the answer has been segmented. For the number of segments, Representation segment The Middle One sentence. Key points The A fragment of evidence, The information vector in the question stem, For evidence fragment vectors, For segment vectors, For sentence vectors, For text encoding functions, For any text unit to be encoded, For L2 normalization operation, As the core module for context semantic encoding, The effective length of the text. For text The k-th basic unit in Indicates will Embedded, The positional encoding vector supplements the word order information for the text unit at the k-th position; This indicates the language type identification result. For the test takers' answers, This is the language recognition function. This represents Chinese characters. In English, Take the language type label corresponding to the highest probability. The posterior probability represents the confidence level that, given a candidate's answer A, the answer belongs to the language type, and its value ranges from 0 to 1.
[0025] In one possible implementation, determining the multi-granularity alignment scores of the text segment and scoring points and fusing them to obtain a soft alignment result, performing non-linear saturation fusion on the contributions of multiple segments to the same point to obtain point coverage and calculating the point score includes:
[0026] Based on the segment-level and sentence-level representations, multi-granularity alignment scores between each text segment and each scoring point are calculated and fused to obtain the soft alignment results of each text segment with respect to each scoring point, including:
[0027]
[0028]
[0029]
[0030] For the same scoring point, the soft alignment results of multiple text segments are non-linearly saturated and fused to obtain the coverage of the scoring point, and the score of the scoring point is calculated according to the weight of each scoring point:
[0031]
[0032]
[0033] in: For section Key points Segment-level matching score, This is the fusion coefficient between paragraph and sentence levels. For section Key points The multi-granularity fusion matching score, For section Key points The probability of soft alignment, For temperature parameters, Key points coverage, For the Sigmoid function, For the slope parameter, As the trigger threshold, This is the total score for each key dimension, summarized according to its weight. For the first The weight of each scoring point This represents the number of scoring criteria.
[0034] In one possible implementation, the calculation of the relevance of the text segment to the prompt, inter-segment contradictions, language quality, and format conformity scores; for Chinese essays, further calculation of coherence and structural integrity scores; and for English essays, sentence-by-sentence grammatical errors detection and vocabulary collocation quality scores, including:
[0035] Perform hybrid text risk control and quality constraint processing on the aforementioned text segment set;
[0036] When writing in Chinese, scores for coherence and structural integrity are calculated:
[0037]
[0038]
[0039] in: To score for fluency in Chinese, is the cosine similarity function, used to measure the semantic similarity between two vectors. The score is given for the structural integrity of the Chinese text. For a set of structural tags, Size of the structure tag set For structural tags, For paragraph type tags, For indicator functions, As an existence condition, there exists a certain paragraph index. This makes the structure tags of the paragraph... Equal to the current structure type ;
[0040] When writing English essays, the system checks each sentence within a paragraph for syntactic errors, generates grammatical penalty items, and calculates a quality score for vocabulary collocation.
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047] in: This is a syntax error. For each syntax error severity, This is the weighting coefficient for the grade. Indicates if Points are awarded for valid grammatical errors; otherwise, no points are awarded. For sentence The strength of syntax errors For the first Duan Di The set of grammatical errors in a sentence Punishment for English grammar For section The number of sentences, To score points for word collocation, To match the rationality scoring function, To output the corrected sentence. Indicates to Perform error correction output. This indicates the precise character-level location of each syntax error within the sentence. This means that all content preceding the incorrect word in the candidate's sentence should be preserved as is. This means replacing the incorrect word or phrase with its correct spelling. This indicates that all content following the incorrect word in the candidate's sentence will be retained as is. Indicates statement length;
[0048] When the question contains materials or a set of evidence, calculate the support degree between the text segment and the said materials or set of evidence, and generate a fact support penalty item:
[0049] ,
[0050] in: For the first The degree of factual support between each text segment and the set of materials for the question. Indicates the index of material fragments Perform the maximum value operation. For the first A fragment of material The vector representation of , To support the punishment with facts, This is a threshold for factual support, used to determine whether a certain passage lacks supporting evidence. This indicates that the support of a single segment is insufficient. Below the threshold A penalty will be applied if the penalty is applied; otherwise, the penalty is 0.
[0051] In one possible implementation, performing hybrid text risk control and quality constraint processing on the text segment set includes:
[0052] Calculate the relevance of the text segment to the topic of the question and generate off-topic penalty items:
[0053] ,
[0054] Calculate the probability of contradictions between text segments and generate a contradiction penalty term:
[0055]
[0056] Computational language quality score and formatting score:
[0057] ,
[0058] in: For section Relevance to the topic of the question stem The severity of the penalty for going off-topic is determined by the number of points; the higher the number, the more off-topic the The lower limit of relevance threshold, below which Included in the off-topic penalty. For contradictory penalty items, For section Section The probability of contradiction, For section Section The weight of contradictions, For language quality score, Language quality characteristics This is a weighted coefficient vector for language quality features. Score for format conformity. As a format specification feature, This is a weighted coefficient vector for the format specification features.
[0059] In one possible implementation, the step of weightedly fusing the key score, quality score, and each penalty item, and then calibrating and mapping them to the full score range to output the final score includes:
[0060] The original score is obtained by combining the key point score, language quality score, format conformity score, coherence score, structural integrity score, or lexical collocation quality score with each penalty item according to preset weights.
[0061]
[0062] The original score is constrained to a preset full score range by calibration mapping, and the final score is output.
[0063]
[0064] in: For the original score, This is the weighting coefficient for the language quality score. The format specification score weighting coefficient, This is the weighting coefficient for the Chinese coherence score. This represents the weighting coefficient for the Chinese structural integrity score. The weighting coefficient for vocabulary collocation scores. The weighting coefficient for the intensity of the off-topic penalty. The weighting coefficient for the contradiction penalty term. The weighting coefficient of the penalty item is used to support the facts. This is the weighting coefficient for English grammar penalties. To output the final score, This is the maximum possible score. For linear calibration proportional coefficient, For linear calibration bias, For calibration, This means taking the maximum value between the calibration score and 0. This means taking the minimum of the previous result and the full score, ensuring that the maximum score does not exceed the full score.
[0065] In one possible implementation, the location of key paragraphs and sentences for each scoring point is output as interpretable evidence. Comments are generated based on this evidence, and optimized model essays are produced for student reference. English essays are meticulously reviewed sentence by sentence, outputting the original sentence, grammatical error annotations, and corrected sentences. When a review is triggered, calibration parameters are updated based on manual results, including:
[0066] Based on the soft alignment results and coverage, interpretable evidence is generated, and the location information of the text segment with the largest contribution and the sentence with the largest contribution within the segment corresponding to each scoring point is output, including: outputting the evidence segment and sentence with the largest contribution for each point: , Output As a basis for awarding marks, among them: To the key points The segment index that contributed the most. This means selecting the paragraph with the highest score as the most relevant to the scoring criteria. This section contains an index of the sentences that best support the main points. This indicates the selected paragraph that contributed the most and is consistent with the reference evidence. The most relevant sentence, To indicate the first One scoring point, To indicate the paragraphs most relevant to the scoring criteria, This refers to the sentence in the paragraph that best supports the scoring criteria;
[0067] Subjective comments are generated based on the interpretable evidence and the results of risk control and quality constraint processing.
[0068] Based on the set of scoring points, evidence fragment library, and preset generation constraints, model essays with optimized quality that cover key points are generated for students' reference. When the essay is in English, a sentence-by-sentence detailed feedback result is further output. The sentence-by-sentence detailed feedback result includes at least the original sentence, syntactic error markings, and the corrected sentence.
[0069] When the review triggering conditions are met, a manual review is submitted, and the calibration mapping parameters and / or threshold parameters are updated based on the review results for self-calibration in subsequent marking tasks.
[0070] In one possible implementation, generating subjective comments based on the interpretable evidence and the results of risk control and quality constraint processing includes:
[0071] Construct comment slots for each dimension, and use features to trigger and cite evidence paragraphs or sentences. The comment slots include: key points, structure, coherence, digression, contradiction, language, norms, English grammar, and English collocations.
[0072] Select excerpts from the comments: ;
[0073] The format for providing detailed feedback on each sentence in English is to output a triple for each sentence: ;
[0074] in: For dimension The final output comment fragment, This is a template or excerpt for a candidate comment. For dimension The set of candidate comments Feature set used to generate comments Indicates candidate comments With features Match score, This represents the confidence level calculated based on feature F.
[0075] In one possible implementation, the step of generating high-quality model essays covering key points based on the set of scoring criteria, the evidence fragment library, and preset generation constraints is provided for student reference. When the essay is in English, a sentence-by-sentence detailed critique is further output. This detailed critique includes at least the original sentence, syntactic error annotations, and the corrected sentence, including:
[0076] Define the objective function for candidate model texts, satisfying the requirements of covering key points, high coherence, and low error, and controlled by length and rewriting extent:
[0077]
[0078] Setting constraints ensures that the sample essay must cover key points and that the extent of rewriting is controlled:
[0079] , ,
[0080] in: To generate an optimized sample document, As candidate sample texts, To select the optimal solution, select an operator from all candidate texts. In the middle, select the one that maximizes the index within the parentheses. To score points for covering the key points of the sample essay, To score points for the coherence of the sample essay, To score the completeness of the sample essay's structure, To score the quality of vocabulary collocation in the sample essay, This is a grammar penalty item for model essays. Penalty for essays going off-topic Weighting of the score for the coherence of the sample essay. Weighting of the score for the structural completeness of the sample essay. Weighting the quality score of vocabulary matching in sample essays. Weighting of the grammar penalty item for model essays. Weighting of the penalty item for essays going off-topic. Minimum length threshold, The maximum length threshold, For the length of the candidate sample essay, This indicates that constraints are applied to each of the key indexes individually. Indicates the key points of the candidate sample essay. coverage, Key points The coverage lower limit threshold, Indicates candidate sample Compared with the original answer Measurement of the degree of difference This represents the maximum permissible rewrite range.
[0081] Secondly, embodiments of this application provide an intelligent marking system for subjective questions that supports multi-text mixing, including:
[0082] The information processing module is used to obtain the question stem, reference answer and scoring procedure, and to structure the scoring procedure into scoring points and weights, and to segment the candidate's answer and generate text segments and sentence representations;
[0083] The text alignment and scoring calculation module is used to determine the multi-granularity alignment scores of the text segment and the scoring points and fuse them to obtain the soft alignment result. It performs non-linear saturation fusion on the contributions of multiple segments of the same point to obtain the point coverage and calculates the point score.
[0084] The quality constraint module is used to calculate the relevance of the text segment to the question stem, inter-segment contradictions, language quality, and format compliance scores; for Chinese essays, it further calculates the coherence and structural integrity scores; for English essays, it checks for grammatical errors sentence by sentence and calculates the vocabulary collocation quality score.
[0085] The scoring calibration module is used to weight and merge the key points, quality scores, and various penalty items, and then map them to the full score range to output the final score.
[0086] The comments and sample essay generation module is used to output the location of key paragraphs and sentences for each scoring point as interpretable evidence, generate comments based on the evidence, and generate optimized sample essays for students to refer to; the English essays are meticulously reviewed sentence by sentence, outputting the original sentences, grammatical error marks and revised sentences, and the calibration parameters are updated based on the manual results when a review is triggered.
[0087] In this embodiment, a multi-granularity alignment and non-linear fusion strategy effectively captures cross-paragraph information fusion and logic, significantly improving the accuracy of scoring complex answers. Combined with differentiated quality assessment in Chinese and English, it achieves accurate scoring of Chinese coherence and structural integrity, while providing sentence-by-sentence detailed feedback on English, accurately marking grammatical errors and optimizing vocabulary collocation. Simultaneously, it outputs interpretable scoring evidence, generating targeted personalized comments and optimized model essays to help students improve precisely. A calibration mechanism and dynamic parameter updates through manual review ensure scoring reliability, improving marking efficiency, reducing human bias, and promoting the upgrade of intelligent marking from simple scoring to a combination of accurate assessment and personalized guidance. This effectively solves the problems of insufficient multi-text mixed answer processing, lack of personalized comments and optimized model essays, and missing sentence-by-sentence detailed feedback in existing systems. Attached Figure Description
[0088] Figure 1 A flowchart illustrating an intelligent marking method for subjective questions that supports multi-text mixing, provided as an embodiment of this application;
[0089] Figure 2 A schematic diagram of the original Chinese essay written by a student as provided in this application embodiment;
[0090] Figure 3 A schematic diagram of the original English essay written by a student as provided in this application embodiment;
[0091] Figure 4 The embodiments provided in this application are for Figure 2 The final score given for the Chinese essay;
[0092] Figure 5 The embodiments provided in this application are for the purpose of Figure 3 The final score for the English essay in the text;
[0093] Figure 6 The embodiments provided in this application are for Figure 2 A diagram illustrating the analysis of flaws in Chinese compositions and suggestions for improvement;
[0094] Figure 7 The embodiments provided in this application are for Figure 3 The English essays provided include optimization suggestions and detailed sentence-by-sentence feedback diagrams;
[0095] Figure 8 This is a schematic diagram of an intelligent marking system for subjective questions that supports multi-text mixing, provided as an embodiment of this application. Detailed Implementation
[0096] The present solution will now be described in conjunction with the accompanying drawings and specific embodiments.
[0097] See Figure 1 The intelligent marking method for subjective questions that supports mixed text provided in this embodiment includes:
[0098] S101 retrieves the question stem, reference answer, and scoring procedure, and structures the scoring procedure into scoring points and weights. It then segments the candidate's answer and generates text segments and sentence representations.
[0099] To achieve accurate and automated grading of subjective questions, the first step is to collect and standardize core basic information. The first step involves comprehensively acquiring the three core pieces of information for the subjective questions to be graded: the question stem, the reference answer, and the grading procedure. The question stem clarifies the question requirements and the scope of the answer; the reference answer serves as the basis for grading; and the grading procedure includes key criteria such as grading points, the weight of each point, and deduction rules.
[0100] The scoring procedure information needs to be broken down into a clear set of scoring points and corresponding weight allocations through structured processing to ensure that the scoring criteria are quantifiable. At the same time, the reference answers should be further decomposed into a library of evidence fragments that are associated with each scoring point, with each evidence fragment serving as the core basis for judging whether the candidate's answer covers that point.
[0101] Specifically, the scoring procedure information is structured into a set of scoring points. and corresponding weights The reference answer is decomposed into a library of evidence fragments associated with each scoring point. The Contains multiple pieces of evidence .
[0102] Next, the submitted answers are analyzed using a multi-text hybrid approach: first, text segments are generated based on semantic logic, such as paragraph-level division. Then, each text segment is further broken down into sentence sets, such as simple sentences and complex sentences. Subsequently, paragraph-level semantic representations capturing the overall core meaning of the paragraph and sentence-level semantic representations capturing the detailed expressions of sentences are constructed.
[0103] Specifically, in this embodiment, the candidate's answer is divided into text segment sets according to paragraphs, bullet points, and quotation blocks. :
[0104]
[0105] Each text segment is then segmented into sentences to obtain a set of sentences. Then, respectively, the above , , and The encoding process generates the question stem information, evidence fragments, paragraph-level representations of each text segment, and sentence-level representations of each sentence.
[0106] , , ,
[0107]
[0108] in: This represents the set of text segments after the answer has been segmented. For the number of segments, Representation segment The Middle One sentence. Key points The A fragment of evidence, The information vector in the question stem, For evidence fragment vectors, For segment vectors, For sentence vectors, For text encoding functions, For any text unit to be encoded, For L2 normalization operation, As the core module for context semantic encoding, The effective length of the text. For text The k-th basic unit in Indicates will Embedded, The positional encoding vector supplements the word order information for the text unit at the k-th position.
[0109] To accommodate the differentiated scoring requirements of essays written in different languages, it is also necessary to identify the language type of the candidates' answers, clarifying whether they are Chinese or English essays, in order to determine the corresponding specific processing branch:
[0110]
[0111] in: This indicates the language type identification result. For the test takers' answers, This is the language recognition function. This represents Chinese characters. In English, Take the language type label corresponding to the highest probability. The posterior probability represents the confidence level that, given a candidate's answer A, the answer belongs to the language type, and its value ranges from 0 to 1.
[0112] S102, determine the multi-granularity alignment score between the text segment and the scoring point and fuse them to obtain the soft alignment result, perform nonlinear saturation fusion on the contributions of multiple segments of the same point to obtain the point coverage and calculate the point score.
[0113] Based on the segment-level and sentence-level representations generated by S101, multi-dimensional calculations are needed to achieve accurate matching between the examinee's answer and the scoring points, thereby deriving the core scoring points. First, considering the different granularities of semantic expression in the text, the alignment scores between each text segment and each scoring point need to be calculated separately and then fused.
[0114]
[0115] The paragraph alignment score focuses on whether the overall text paragraph conforms to the core meaning of the scoring criteria, and is achieved by calculating the maximum cosine similarity between the text paragraph vector and the corresponding evidence fragment vector.
[0116] Based on this, sentence-level alignment scores are incorporated, which are the maximum similarity between each sentence in the text segment and the evidence fragment. Through weighted fusion, multi-granularity fusion alignment scores are obtained, which not only ensures overall semantic matching but also does not ignore the detailed contributions of key sentences.
[0117]
[0118] To more accurately reflect the correlation between text segments and various scoring criteria, the multi-granularity fusion alignment scores need to be normalized to obtain a soft alignment probability. This probability can intuitively reflect which scoring criterion a particular text segment is more likely to match.
[0119]
[0120] For the same scoring criterion, the soft alignment results of all text segments need to be combined, and the criterion coverage needs to be calculated using a non-linear saturation fusion algorithm:
[0121]
[0122] This integrated approach avoids the extreme influence of a single text segment, and more objectively reflects the overall coverage of the key point in the candidate's answer. The higher the coverage, the more fully the candidate understands and expresses the key point.
[0123] Finally, based on the preset weights of each scoring point, the coverage of all points is weighted and summed to obtain the point score. This score directly reflects whether the core content of the candidate's answer meets the requirements. The relevant calculation formula is as follows:
[0124]
[0125] in: For section Key points Segment-level matching score, This is the fusion coefficient between paragraph and sentence levels. For section Key points The multi-granularity fusion matching score, For section Key points The probability of soft alignment, For temperature parameters, Key points coverage, For the Sigmoid function, For the slope parameter, As the trigger threshold, This is the total score for each key dimension, summarized according to its weight. For the first The weight of each scoring point This represents the number of scoring criteria.
[0126] S103, calculate the relevance of the text segment to the question stem, inter-segment contradictions, language quality, and format standardization scores; for Chinese essays, further calculate the coherence and structural integrity scores; for English essays, check for grammatical errors sentence by sentence and calculate the vocabulary collocation quality score.
[0127] The scoring of key points only reflects the matching degree of core content. The scoring of subjective questions also needs to take into account multiple dimensions such as logical rationality, language expression quality, and format standardization. Therefore, it is necessary to perform mixed text risk control and quality constraint processing on the collection of text segments, comprehensively investigate the answer problems and quantitatively evaluate the quality level.
[0128] The text segment set is subjected to hybrid text risk control and quality constraint processing. First, at the risk control level: off-topic risk detection is performed by calculating the cosine similarity between each text segment vector and the question stem vector to determine the relevance of the text segment to the question stem topic. If the relevance is lower than a preset threshold, an off-topic penalty is generated. The lower the relevance, the heavier the penalty, to prevent candidates' answers from deviating from the requirements of the question.
[0129] ,
[0130] in: For section Relevance to the topic of the question stem The severity of the penalty for going off-topic is determined by the number of points; the higher the number, the more off-topic the question. The lower limit of relevance threshold, below which This will be included in the off-topic penalty.
[0131] Logical contradiction risk detection analyzes semantic conflicts between all text segments, calculates the probability of contradiction, and generates a contradiction penalty term to ensure that the candidate's answer is logically consistent and free of inconsistencies.
[0132]
[0133] in: For contradictory penalties, For section Section The probability of contradiction, For section Section The contradictory weights, in this embodiment, and The values are all between 0 and 1.
[0134] for The test taker's answers are divided into paragraphs to obtain a set of text segments. Then, any two adjacent paragraphs are... and The construction section for the input can adopt the "premise-assumption" form, that is, based on... As a premise, with As an assumption, and simultaneously constructing reverse segment pairs ( As a premise, (As assumed). Semantic consistency is determined for each segment pair, outputting a probability distribution of "implied / neutral / contradictory". The probability corresponding to "contradictory" is taken as the contradiction probability of the segment pair. If bidirectional calculation is used, the positive and negative contradiction probabilities are fused. The fusion method can be either taking the maximum value or taking the average value, thus obtaining the final result. To reduce computational cost, we can first calculate the semantic relevance between the two segments. If the relevance is below a threshold, then we can directly set... .
[0135] This indicator reflects the degree to which different paragraphs influence the overall score, and its weighting can be determined according to the following rules: Weighting based on paragraph position: adjacent paragraphs have higher weights, and distant paragraphs have lower weights. Alternatively, only adjacent paragraph pairs within a preset range can be assigned non-zero weights. Weighting based on paragraph function: first, identify the paragraph type for each paragraph; assign higher weights when a paragraph pair contains at least an argument paragraph or a conclusion paragraph; assign lower weights when both paragraph pairs are background description paragraphs. Weighting based on relevance to key points: calculate the alignment strength between the paragraph and the scoring point; if both paragraphs are highly relevant to the same scoring point, increase the weighting. Assignment is used to emphasize the penalty for "self-contradictions under the same point".
[0136] Fact Support Risk Detection: This involves calculating the semantic match between the text segment and the question material. If the support is insufficient, it indicates that the answer lacks factual basis, and a fact support penalty item needs to be generated. Language quality score and format compliance score are also calculated.
[0137]
[0138] in: For language quality score, Language quality characteristics This is a weighted coefficient vector for language quality features.
[0139] In this embodiment, Characterizing the first The quality of language expression in a paragraph is extracted from the paragraph text by the language quality assessment module. Specifically, it may include: Fluency features: Calculating perplexity or average log-likelihood using a language model to reflect the naturalness of word choice and sentence structure. Readability / Sentence Complexity features: Statistically analyzing average sentence length, the proportion of subordinate clauses, and the proportion of conjunctions used. Lexical Richness features: Statistically analyzing the ratio of vocabulary type to total vocabulary, the proportion of low-frequency words, and the proportion of repeated words. Standardization features: Detecting the number or proportion of misspellings, grammatical errors, and punctuation misuse (Chinese), or the number of spelling errors (English). Arranging these indicators into a vector along a fixed dimension yields the... .
[0140] To The weighted average of the features can be obtained by: pre-setting the importance of each language quality indicator based on scoring procedures or teaching research experience, for example, giving more weight to fluency than to sentence complexity; collecting training samples with human scoring; calculating the correlation between each feature and the human "language expression score"; or fitting the weights through a regression learning ranking model to obtain the final score. After accumulating manually verified samples, they are periodically reviewed. Perform incremental updates to better align the automatic language quality score with the human language score. Ultimately, this will... Normalize or crop to a preset range to ensure stability.
[0141] Secondly, there is the quality constraint assessment: all essays are uniformly assessed for language quality scores, such as word accuracy and sentence fluency. Format compliance scores are also given, such as paragraph division, punctuation usage, and word count adherence.
[0142]
[0143] in: Score for format conformity. As a format specification feature, This is a weighted coefficient vector for the format specification features.
[0144] Characterizing the first The format and answer standardization of the paragraphs specifically include: Paragraph structure characteristics: whether there are paragraphs, the ratio of excessively long / short paragraphs, and whether there are blank lines, etc. Organizational characteristics: the presence of serial numbers / partitions, and the use of headings / subheadings. Punctuation and layout characteristics: continuous punctuation, the proportion of unusual symbols, the ratio of full-width and half-width characters mixed, and abnormal indentation / alignment, etc. Answer requirement compliance characteristics: whether the word count is exceeded / falls short, and whether irrelevant symbols or pasting traces are included, etc. Arranging the above indicators into a vector according to a fixed dimension yields the... .
[0145] Used for The weighted sum of the features in each dimension can be obtained as follows: Scoring procedure mapping: If the scoring procedure specifies the proportions of format, writing, and organization, then that proportion is allocated to... Formed on each sub-indicator Fitting through historical samples And based on manually reviewed samples Perform periodic updates to ensure the format scores align with human scoring. Normalize or crop to a preset range to avoid extreme weights that could cause scoring instability.
[0146] For Chinese essays, additional assessments are given for coherence and structural integrity. The coherence score indicates the naturalness of semantic connection between paragraphs, while the structural integrity score assesses whether the essay includes necessary structural elements such as an introduction, body, and conclusion, and whether it conforms to the conventional structural logic of Chinese essays.
[0147] The formulas for calculating the coherence score and the structural integrity score are as follows:
[0148]
[0149]
[0150] in: To score for fluency in Chinese, is the cosine similarity function, used to measure the semantic similarity between two vectors. Score for Chinese structural integrity. A set of structural tags, Size of the structure tag set For structural tags, For paragraph type tags, For indicator functions, As an existence condition, there exists a certain paragraph index. This makes the structure tags of the paragraph... Equal to the current structure type .
[0151] When writing English essays, the system checks each sentence in a paragraph for syntactic errors and generates grammatical penalty items. It also calculates vocabulary collocation quality scores, judges the rationality of vocabulary collocation between adjacent sentences, and outputs the sentence-by-sentence correction results, making it convenient for test takers to see specific errors and directions for correction.
[0152] The relevant calculation formulas are as follows:
[0153]
[0154] in: This is a syntax error. For each syntax error severity, For error type mapping, This represents the grade weighting coefficient. In this embodiment, Match the severity of each syntax error, including minor errors, basic errors, core errors, and critical errors. These are predefined fixed values, representing a mathematical quantification of the manual marking standards, based on... Different severity levels of the match have different values.
[0155] For example: minor errors that do not affect semantic understanding (such as spelling mistakes or capitalization errors), then The value is 0.5. If the error is fundamental and slightly affects local semantics (e.g., misuse of prepositions, or confusion between adjectives and adverbs), then... The value is 1. This is the core error, severely impacting the sentence's semantics (e.g., subject-verb disagreement, tense errors). The value is 2. If the error is severe and the sentence is completely incomprehensible (e.g., no predicate, broken sentence structure), then... The value is 3. Indicates if A valid syntax error is scored as 1; otherwise, it is scored as 0. The output values include 0, 0.5, 1, 2, 3, using... and To determine the severity of each grammatical error so that a single error does not distort the score, and to adapt to the graded deduction rules for subjective question scoring.
[0156]
[0157]
[0158] in: For sentence The strength of syntax errors For the first Duan Di The set of grammatical errors in a sentence Punishment for English grammar For section The number of sentences. The above determines... After the output result, The summation result It is the total score for grammatical errors in a single sentence, and the final score is based on different... Syntax penalty items aggregated into the whole text . The score design ensures There will be no situation where the penalty is too large and results in a negative total score.
[0159]
[0160] in: To score points for word collocation, To match the rationality scoring function, The value is determined by and The judgment is based on the coherence of the combinations. If and If two sentences share a common core collocation, then The value is 1; if and The two sentences do not share a common core collocation, but contain some logical connectors. The value is 0.5; if and If there are no logical connectors, then The value is 0. Once determined, the quality scores of each pair of adjacent sentences are added together, and then the average of the scores for each pair of adjacent sentences is normalized. .
[0161]
[0162]
[0163] in: To output the corrected sentence. Indicates to Perform error correction output. This indicates the precise character-level location of each syntax error within the sentence. This means that all content preceding the incorrect word in the candidate's sentence should be preserved as is. This means replacing the incorrect word or phrase with its correct spelling. This indicates that all content following the incorrect word in the candidate's sentence will be retained as is. Indicates the length of the statement.
[0164] When the question contains materials or a set of evidence, calculate the support degree between the text segment and the said materials or set of evidence, and generate a fact support penalty item:
[0165] ,
[0166] in: For the first The degree of factual support between each text segment and the set of materials for the question. Indicates the index of material fragments Perform the maximum value operation. For the first A fragment of material The vector representation of , To support the punishment with facts, This is a threshold for factual support, used to determine whether a certain passage lacks supporting evidence. This indicates that the support of a single segment is insufficient. Below the threshold A penalty will be applied if the penalty is applied; otherwise, the penalty is 0.
[0167] S104 weights and merges the key points score, quality score, and each penalty item, and then maps them to the full score range after calibration to output the final score.
[0168] Having obtained key point scores, language quality scores, format compliance scores, coherence scores (Chinese), structural integrity scores (Chinese), vocabulary collocation quality scores (English), and various penalty items (off-topic, contradiction, lack of factual support, grammatical errors) through S102 and S103, the next step is to obtain the raw scores through reasonable fusion rules and perform calibration processing to ensure the rationality and consistency of the scores.
[0169] The original score is obtained by combining the key point score, language quality score, format conformity score, coherence score, structural integrity score, or lexical collocation quality score with each penalty item according to preset weights.
[0170]
[0171] The weights of each scoring item and each penalty item are predetermined according to the scoring criteria for subjective questions. For example, the implicit weight of key point scoring is 1, reflecting the importance of core content; the weight of language quality scoring can be adjusted according to the essay type, and the weight of grammar penalty items in English essays can be appropriately increased. The raw score is the sum of all scoring items minus the sum of all penalty items, comprehensively reflecting the candidate's overall performance in core content, language expression, logical coherence, and other dimensions.
[0172] Since the original scores may exceed the preset full score range or the score distribution may be unreasonable, adjustments need to be made through calibration mapping:
[0173]
[0174] After adjusting the original fractions through a linear transformation, they are restricted to [0, ... Within the specified range, the final output is the calibrated score, ensuring that the score meets the actual requirements of the marking scenario.
[0175] See Figure 2 and Figure 3 Examples of Chinese and English essays written by students are provided below. Figure 4 In order to Figure 2 The final score given for the Chinese essay in the article. Figure 5 In order to Figure 3 The final score given for the English essay is shown in the image. It can be seen that the output not only provides a score but also analyzes and displays the essay's weaknesses; however, there are differences in the evaluation criteria for Chinese and English essays.
[0176] in: For the original score, This is the weighting coefficient for the language quality score. The format specification score weighting coefficient, This is the weighting coefficient for the Chinese coherence score. This represents the weighting coefficient for the Chinese structural integrity score. The weighting coefficient for vocabulary collocation scores. The weighting coefficient for the intensity of the off-topic penalty. The weighting coefficient for the contradiction penalty term. The weighting coefficient of the penalty item is used to support the facts. This is the weighting coefficient for English grammar penalties. To output the final score, This is the maximum possible score. For linear calibration proportional coefficient, For linear calibration bias, For calibration, This means taking the maximum value between the calibration score and 0. This means taking the minimum of the previous result and the full score, ensuring that the maximum score does not exceed the full score.
[0177] S105 outputs the location of key paragraphs and sentences for each scoring point as interpretable evidence, generates comments based on the evidence, and generates optimized model essays for students to refer to; it provides detailed sentence-by-sentence feedback on English essays, outputs the original sentence, grammatical error annotations, and revised sentences, and updates calibration parameters based on manual results when a review is triggered.
[0178] To improve the transparency and usability of automated marking, and to continuously optimize the performance of the marking model, it is necessary to complete interpretable output, sample paper generation, and self-calibration optimization.
[0179] First, interpretable evidence is generated: based on the soft alignment results and key point coverage in step S102, the text segment with the greatest contribution corresponding to each scoring point is located, that is, the text segment with the highest score aligned with that key point, and the sentence with the greatest contribution within that text segment, that is, the sentence with the highest matching degree with the evidence fragment of that key point. Specific location information is then output, allowing candidates and examiners to clearly understand the scoring basis for each key point and resolving the question of "how are the scores determined?"
[0180] Specifically, for each key point, output the paragraph and sentence that contributes the most to the evidence: , Output As a basis for scoring. Among them: To the key points The segment index that contributed the most. This means selecting the paragraph with the highest score as the most relevant to the scoring criteria. This section contains an index of the sentences that best support the main points. This indicates the selected paragraph that contributed the most and is consistent with the reference evidence. The most relevant sentence, To indicate the first One scoring point, To indicate the paragraphs most relevant to the scoring criteria, This indicates the sentence in the paragraph that best supports the scoring criteria.
[0181] Subjective question comments are generated based on the interpretable evidence and the results of risk control and quality constraint processing. The comments should clearly point out the candidate's strengths, such as comprehensive coverage of core points and fluent language, and their weaknesses, such as grammatical errors and incomplete structure, and provide targeted suggestions for improvement.
[0182] See Figure 6 In order to Figure 2 The Chinese essay in question is analyzed for its flaws and suggestions for improvement are given. Figure 7 In order to Figure 3 The text provides optimization suggestions and detailed sentence-by-sentence feedback for English essays.
[0183] First, comment slots are constructed for each dimension, triggering and referencing evidence paragraphs or sentences using features. These comment slots include: key points, structure, coherence, digression, contradiction, language, standardization, English grammar, and English collocation. Then, comment fragments are selected: For sentence-by-sentence detailed feedback of English text, the output format is a triplet for each sentence: .
[0184] in: For dimension The final output comment fragment, This is a template or excerpt for a candidate comment. For dimension The set of candidate comments Feature set used to generate comments Indicates candidate comments With features Match score, This represents the confidence level calculated based on feature F.
[0185] Based on the set of scoring points, evidence fragment library, and preset generation constraints, model essays covering key points and optimized in quality are generated for students' reference. When the essay is in English, a sentence-by-sentence detailed feedback result is further output. The sentence-by-sentence detailed feedback result includes at least the original sentence, syntactic error markings, and the corrected sentence, so that candidates can accurately correct their problems.
[0186] Define the objective function for candidate model texts, satisfying the requirements of covering key points, high coherence, and low error, and controlled by length and rewriting extent:
[0187]
[0188] Setting constraints ensures that the sample essay must cover key points and that the extent of rewriting is controlled:
[0189] , ,
[0190] Finally, there is self-calibration optimization: Review trigger conditions are set (such as scores falling within a critical range, high probability of inconsistencies, excessive penalties for off-topic grading, etc.), and grading results meeting these conditions are submitted for manual review. Based on the corrected scores from the manual review, the calibration mapping parameters and threshold parameters in each step are updated using the least squares method, making the grading results closer to the human grading standards. This enables self-calibration iteration for subsequent grading tasks, continuously improving the accuracy and reliability of automatic grading.
[0191] The self-calibration iteration formula is:
[0192]
[0193] in: To generate an optimized sample document, As candidate sample texts, To select the optimal solution, select an operator from all candidate texts. In the middle, select the one that maximizes the index within the parentheses. To score points for covering the key points of the sample essay, To score points for the coherence of the sample essay, To score for the completeness of the sample essay's structure, To score the quality of vocabulary collocation in the sample essay, This is a grammar penalty item for model essays. Penalty for essays going off-topic Weighting of the score for the coherence of the sample essay. The score is based on the structural completeness of the sample essay. Weighting the quality score of vocabulary matching in sample essays. Weighting of the grammar penalty item for model essays. Weighting of the penalty item for essays going off-topic. Minimum length threshold, The maximum length threshold, The length of the candidate sample essay, This indicates that constraints are applied to each of the key indexes individually. Indicates the key points of the candidate sample essay. coverage, Key points The lower limit threshold of coverage, Indicates candidate sample Compared with the original answer Measurement of the degree of difference To the maximum permissible rewrite range, Indicates to and Find the minimum. To verify the sample index, This is the total number of verification samples used for calibration. For the first The original score of each sample, Indicates the first The manual review score of each sample For the first The calibration prediction score for each sample.
[0194] When the review triggering conditions are met, a manual review is submitted, and the calibration mapping parameters and / or threshold parameters are updated based on the review results for self-calibration in subsequent marking tasks.
[0195] Corresponding to the above embodiment of the intelligent marking method for subjective questions that supports multi-text mixing, this application also provides an embodiment of an intelligent marking system for subjective questions that supports multi-text mixing.
[0196] See Figure 8 The intelligent marking system 20 for subjective questions that supports multi-text mixing in this embodiment includes:
[0197] The information processing module 201 is used to obtain the question stem, reference answer and scoring procedure, and to structure the scoring procedure into scoring points and weights, and to segment the candidate's answer and generate text segments and sentence representations.
[0198] The text alignment and scoring calculation module 202 is used to determine the multi-granularity alignment scores of the text segment and the scoring points and fuse them to obtain the soft alignment result. It performs nonlinear saturation fusion on the contributions of multiple segments of the same point to obtain the point coverage and calculates the point score.
[0199] The quality constraint module 203 is used to calculate the relevance of the text segment to the question stem, inter-segment contradictions, language quality, and format standardization scores; for Chinese essays, it further calculates the coherence and structural integrity scores; for English essays, it checks for grammatical errors sentence by sentence and calculates the vocabulary collocation quality score.
[0200] The scoring calibration module 204 is used to weight and merge the key points, quality scores and each penalty item, and then map them to the full score range to output the final score.
[0201] The 205 module for generating sample comments is used to output the location of key paragraphs and sentences for each scoring point as interpretable evidence, generate comments based on the evidence, and generate optimized sample comments for students to refer to; the module provides detailed sentence-by-sentence feedback on English essays, outputting the original sentence, grammatical error annotations, and revised sentences, and updates the calibration parameters based on the manual results when a review is triggered.
[0202] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0203] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for intelligent grading of subjective questions that supports mixed multi-text format, characterized in that, include: The system retrieves the question stem, reference answer, and scoring criteria, and structures the scoring criteria into scoring points and weights. It then segments the candidate's answer and generates text segments and sentence representations, including: Obtain the question stem information of the subjective questions to be graded Reference answer information and scoring procedures; The scoring procedure information is structured into a set of scoring points. and corresponding weights ; The reference answer is broken down into a library of evidence fragments associated with each scoring point. The Contains multiple pieces of evidence ; The test takers' answers were divided into text segments based on paragraphs, bullet points, and quotation blocks. : Each text segment is then segmented into sentences to obtain a set of sentences. ; For the above , , and The encoding process generates the question stem information, evidence fragments, paragraph-level representations of each text segment, and sentence-level representations of each sentence. , , , The process involves identifying the language type of the test taker's answers to determine whether the essay is in Chinese or English. in: This represents the set of text segments after the answer has been segmented. For the number of segments, Representation segment The Middle One sentence. Key points The A fragment of evidence, The information vector in the question stem, For evidence fragment vectors, For segment vectors, For sentence vectors, For text encoding functions, For any text unit to be encoded, For L2 normalization operation, As the core module for context semantic encoding, The effective length of the text. For text The k-th basic unit in Indicates will Embedded, The positional encoding vector supplements the word order information for the text unit at the k-th position; This indicates the language type identification result. For the test takers' answers, For text language recognition functions, This represents Chinese characters. In English, Take the language type label corresponding to the highest probability. The posterior probability represents the confidence level that, given a candidate's answer A, the answer belongs to the language type, and its value ranges from 0 to 1. The process involves determining and fusing multi-granularity alignment scores between the text segment and the scoring points to obtain a soft alignment result. Non-linear saturation fusion is then performed on the contributions of multiple segments to the same point to obtain the point coverage and calculate the point score, including: Based on the segment-level and sentence-level representations, multi-granularity alignment scores between each text segment and each scoring point are calculated and fused to obtain the soft alignment results of each text segment with respect to each scoring point, including: For the same scoring point, the soft alignment results of multiple text segments are non-linearly saturated and fused to obtain the coverage of the scoring point, and the score of the scoring point is calculated according to the weight of each scoring point: in: For section Key points Segment-level matching score, This is the fusion coefficient between paragraph and sentence levels. For section Key points The multi-granularity fusion matching score, For section Key points The soft alignment probability, For temperature parameters, Key points coverage, For the Sigmoid function, For the slope parameter, As the trigger threshold, This is the total score for each key dimension, summarized according to its weight. For the first The weight of each scoring point The number of scoring criteria; Calculate the relevance of the text segment to the prompt, inter-segment contradictions, language quality, and format compliance scores; for Chinese essays, further calculate coherence and structural integrity scores; for English essays, check for grammatical errors sentence by sentence and calculate vocabulary collocation quality scores, including: Perform hybrid text risk control and quality constraint processing on the aforementioned text segment set; When writing in Chinese, scores for coherence and structural integrity are calculated: in: To score for fluency in Chinese, is the cosine similarity function, used to measure the semantic similarity between two vectors. The score is given for the structural integrity of the Chinese text. For a set of structural tags, Size of the structure tag set For structural tags, For paragraph type tags, For indicator functions, As an existence condition, there exists a certain paragraph index. This makes the structure tags of the paragraph... Equal to the current structure type ; When writing English essays, the system checks each sentence within a paragraph for syntactic errors, generates grammatical penalty items, and calculates a quality score for vocabulary collocation. in: This is a syntax error. For each syntax error, the severity This is the grade weighting coefficient. Indicates if Points are awarded for valid grammatical errors; otherwise, no points are awarded. For sentence The strength of syntax errors For the first Duan Di The set of grammatical errors in a sentence Punishment for English grammar For section The number of sentences, To score points for word collocation, To match the rationality scoring function, To output the corrected sentence. Indicates to Perform error correction output. This indicates the precise character-level location of each syntax error within the sentence. This means that all content preceding the incorrect word in the candidate's sentence should be preserved as is. This means replacing the incorrect word or phrase with its correct spelling. This indicates that all content following the incorrect word in the candidate's sentence will be retained as is. Indicates statement length; When the question contains materials or a set of evidence, calculate the support degree between the text segment and the said materials or set of evidence, and generate a fact support penalty item: , in: For the first The degree of factual support between each text segment and the set of materials for the question. Indicates the index of material fragments Perform the maximum value operation. For the first A fragment of material The vector representation of , To support the punishment with facts, This is a threshold for factual support, used to determine whether a certain passage lacks supporting evidence. This indicates that the support of a single segment is insufficient. Below the threshold A penalty will be applied if the penalty is applied, otherwise the penalty will be zero. The key points, quality score, and each penalty item are weighted, merged, and calibrated to be mapped to the full score range to output the final score; The system outputs the location of key paragraphs and sentences for each scoring point as interpretable evidence, generates comments based on the evidence, and generates optimized model essays for students to refer to; it provides detailed feedback on English essays sentence by sentence, outputting the original sentences, grammatical error annotations, and revised sentences, and updates calibration parameters based on manual results when a review is triggered.
2. The intelligent marking method for subjective questions supporting multi-text mixing according to claim 1, characterized in that, The process of performing hybrid text risk control and quality constraint processing on the text segment set includes: Calculate the relevance of the text segment to the topic of the question and generate off-topic penalty items: , Calculate the probability of contradictions between text segments and generate a contradiction penalty term: Computational language quality score and formatting score: , in: For section Relevance to the topic of the question stem The severity of the penalty for going off-topic is determined by the number of points; the higher the number, the more off-topic the The lower limit of relevance threshold, below which Included in the off-topic penalty. For contradictory penalties, For section Section The probability of contradiction, For section Section The weight of contradictions, For language quality score, Language quality characteristics This is a weighted coefficient vector for language quality features. Score for format conformity. As a format specification feature, This is a weighted coefficient vector for the format specification features.
3. The intelligent marking method for subjective questions supporting multi-text mixing according to claim 2, characterized in that, The process of weightedly combining key score points, quality score, and each penalty item, and then calibrating and mapping them to the full score range to output the final score includes: The original score is obtained by combining the key point score, language quality score, format conformity score, coherence score, structural integrity score, or lexical collocation quality score with each penalty item according to preset weights. The original score is constrained to a preset full score range by calibration mapping, and the final score is output. in: For the original score, This is the weighting coefficient for the language quality score. The format specification score weighting coefficient, This is the weighting coefficient for the Chinese coherence score. This represents the weighting coefficient for the Chinese structural integrity score. The weighting coefficient for vocabulary collocation scores. The weighting coefficient for the intensity of the off-topic penalty. The weighting coefficient for the contradiction penalty term. The weighting coefficients of the penalty items are used to support the facts. This is the weighting coefficient for English grammar penalties. To output the final score, This is the maximum possible score. For linear calibration proportional coefficient, For linear calibration bias, For calibration, This means taking the maximum value between the calibration score and 0. This means taking the minimum of the previous result and the full score, ensuring that the maximum score does not exceed the full score.
4. The intelligent marking method for subjective questions supporting multi-text mixing according to claim 3, characterized in that, The key paragraphs and sentences for each scoring point are located as interpretable evidence. Comments are generated based on the evidence, and optimized sample essays are generated for students' reference. The English essay is meticulously critiqued sentence by sentence, outputting the original sentence, grammatical error annotations, and corrected sentences. When a review is triggered, the calibration parameters are updated based on the manual feedback, including: Based on the soft alignment results and coverage, interpretable evidence is generated, and the location information of the text segment with the largest contribution and the sentence with the largest contribution within the segment corresponding to each scoring point is output, including: outputting the evidence segment and sentence with the largest contribution for each point: , Output As a basis for awarding marks, among them: To the key points The segment index that contributed the most, This means selecting the paragraph with the highest score as the most relevant to the scoring criteria. This section contains an index of the sentences that best support the main points. This indicates the selected paragraph that contributed the most and is consistent with the reference evidence. The most relevant sentence, To indicate the first One scoring point, To indicate the paragraphs most relevant to the scoring criteria, This refers to the sentence in the paragraph that best supports the scoring criteria; Subjective comments are generated based on the interpretable evidence and the results of risk control and quality constraint processing. Based on the set of scoring points, evidence fragment library, and preset generation constraints, model essays with optimized quality that cover key points are generated for students' reference. When the essay is in English, a sentence-by-sentence detailed feedback result is further output. The sentence-by-sentence detailed feedback result includes at least the original sentence, syntactic error annotations, and the corrected sentence. When the review triggering conditions are met, a manual review is submitted, and the calibration mapping parameters and / or threshold parameters are updated based on the review results for self-calibration in subsequent marking tasks.
5. The intelligent marking method for subjective questions supporting multi-text mixing according to claim 4, characterized in that, The generation of subjective comments based on the interpretable evidence and the results of risk control and quality constraint processing includes: Construct comment slots for each dimension, and use features to trigger and cite evidence paragraphs or sentences. The comment slots include: key points, structure, coherence, digression, contradiction, language, norms, English grammar, and English collocations. Select excerpts from the comments: ; The format for providing detailed feedback on each sentence in English is to output a triple for each sentence: ; in: For dimension The final output comment fragment, This is a template or excerpt for a candidate comment. For dimension The set of candidate comments, Feature set used to generate comments Indicates candidate comments With features Match score, This represents the confidence level calculated based on feature F.
6. The intelligent marking method for subjective questions supporting multi-text mixing according to claim 5, characterized in that, The system generates high-quality model essays covering key points based on the set of scoring criteria, the evidence fragment library, and preset generation constraints for students' reference. When the essay is in English, it further outputs a sentence-by-sentence detailed critique, which includes at least the original sentence, syntactic error annotations, and the corrected sentence, including: Define the objective function for candidate model texts, satisfying the requirements of covering key points, high coherence, and low error, and controlled by length and rewriting extent: Setting constraints ensures that the sample essay must cover key points and that the extent of rewriting is controlled: , , in: For the generated optimized sample, As candidate sample texts, To select the optimal solution, select an operator from all candidate texts. In the middle, select the one that maximizes the index within the parentheses. To score points for covering the key points of the sample essay, To score points for the coherence of the sample essay, To score for the completeness of the sample essay's structure, To score the quality of vocabulary collocation in the sample essay, This is a grammar penalty item for model essays. Penalty for essays going off-topic Weighting of the score for the coherence of the sample essay. The score is based on the structural completeness of the sample essay. Weighting the quality score of vocabulary matching in sample essays. Weighting of the grammar penalty item for model essays. Weighting of the penalty item for essays going off-topic. Minimum length threshold, The maximum length threshold, For the length of the candidate sample essay, This indicates that constraints are applied to each of the key indexes individually. Indicates the key points of the candidate sample essay. coverage, Key points The coverage lower limit threshold, Indicates candidate sample Compared with the original answer Measurement of the degree of difference This represents the maximum permissible rewrite range.
7. A subjective question intelligent marking system supporting multi-text mixing, characterized in that, To implement the method of claim 1, the method comprises: The information processing module is used to obtain the question stem, reference answer and scoring procedure, and to structure the scoring procedure into scoring points and weights, and to segment the candidate's answer and generate text segments and sentence representations; The text alignment and scoring calculation module is used to determine the multi-granularity alignment scores of the text segment and the scoring points and fuse them to obtain the soft alignment result. It performs non-linear saturation fusion on the contributions of multiple segments of the same point to obtain the point coverage and calculates the point score. The quality constraint module is used to calculate the relevance of the text segment to the question stem, inter-segment contradictions, language quality, and format compliance scores; for Chinese essays, it further calculates the coherence and structural integrity scores; for English essays, it checks for grammatical errors sentence by sentence and calculates the vocabulary collocation quality score. The scoring calibration module is used to weight and merge the key points, quality scores, and various penalty items, and then map them to the full score range to output the final score. The comments and sample essay generation module is used to output the location of key paragraphs and sentences for each scoring point as interpretable evidence, generate comments based on the evidence, and generate optimized sample essays for students to refer to; the English essays are meticulously reviewed sentence by sentence, outputting the original sentences, grammatical error marks and revised sentences, and the calibration parameters are updated based on the manual results when a review is triggered.
Citation Information
Patent Citations
HanLP-based paperless examination subjective question automatic reviewing method
CN118378626A
Automatic paper marking and scoring method and system based on subjective questions
CN120047954A