Multi-dimensional interpretable subjective question scoring method based on large model
By adopting a multi-dimensional scoring system and dynamic weight adjustment mechanism of a large model in the scoring of subjective questions, the shortcomings of the existing scoring methods in the integration of explanatory, transparency and multi-dimensional factors are solved, and more accurate, fair and transparent scoring results are achieved, and dynamic optimization is supported, which improves the efficiency and practicality of the education system.
Patent Information
- Application Number
- CN202510007526.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The existing subjective question scoring method still has room for improvement in the explanatory, transparency and the integration of multi-dimensional factors of the scoring, which leads to the opaque scoring process, making it difficult to gain user trust, and the scoring dimensions are single, and the language quality, content relevance and expression fluency of the answers are not comprehensively evaluated.
A multi-dimensional explanation of subjective questions based on large models is adopted. By uniform data cleaning and structured processing of diverse data sources, a multi-dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation and time sensitivity evaluation is designed. A three-layer structure weighting mechanism is used to calculate the comprehensive score, generate a scoring basis report, and optimize the scoring rules through a dynamic feedback mechanism.
It significantly improves the accuracy and fairness of subjective questions, enhances the transparency and interpretability of the scoring results, supports dynamic optimization and continuous improvement, and improves the efficiency and practicality of the education system.
Smart Images

Figure CN120068840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and educational evaluation, and particularly to a multi-dimensional interpretable subjective question scoring method based on a large model. Background Art
[0002] With the development of artificial intelligence, automatic scoring technology in the field of education has become an important tool for improving teaching efficiency. Existing automatic scoring systems are mostly based on natural language processing technology, but mainly for objective questions, and are inadequate for subjective question types that require in-depth understanding and creative thinking, lacking effective scoring methods for subjective questions.
[0003] In the prior art, there are already various subjective question scoring methods. For example, the scoring method based on knowledge point annotation (Application No.: CN202310689136.7) accurately evaluates the knowledge point coverage of the answer content through the construction of a knowledge term table and weighted comparison, effectively improving the scoring accuracy and efficiency; the scoring system based on the ALBERT model and RPA technology (Application No.: CN202410025948.6) uses semantic similarity and keyword matching technology to extract text features for scoring, achieving more efficient feature extraction and accurate scoring; and the vertical domain subjective question scoring model selection method (Application No.: CN202411082745.7) screens the best model for subjective question scoring through multi-angle scoring templates and weighting strategies to ensure the rationality and applicability of the model in real applications.
[0004] Although these methods perform well in aspects such as knowledge point comparison and semantic similarity analysis, they mostly rely on similarity calculation or statistical feature analysis, lacking a detailed display of the scoring basis, resulting in poor transparency of the scoring process, difficulty in obtaining user trust, and when dealing with subjective questions, the scoring dimension is single, such as only focusing on semantic similarity or key point coverage, and failing to comprehensively evaluate the language quality, content relevance, and expression fluency of the answer, leaving room for improvement in the interpretability, transparency, and integration of multi-dimensional scoring elements of the scoring. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-dimensional interpretable subjective question scoring method based on a large model, which solves the problem that there is still room for improvement in the interpretability, transparency, and integration of multi-dimensional factors of existing subjective question scoring methods.
[0006] To achieve the above purpose, the present invention provides a multi-dimensional interpretable subjective question scoring method based on a large model, including the following steps:
[0007] Perform unified data cleaning and structured processing on diverse data sources, including data acquisition, data cleaning and verification, extraction of key data information, structured processing and storage of data;
[0008] Design a multi - dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time - sensitivity evaluation. The scoring results for each dimension are obtained based on specific algorithms and models;
[0009] Adopt a three - layer weighted mechanism to calculate the comprehensive score. Among them, the three - layer structure includes a task layer, a scenario layer, and a factor layer;
[0010] Generate a scoring basis report and build a scoring log, and continuously optimize the scoring rules through a dynamic feedback mechanism.
[0011] Among them, perform unified data cleaning and structuring processing on diverse data sources, including data acquisition, data cleaning and verification, extraction of key data information, data structuring processing and storage. The steps also include:
[0012] Directly read and parse the content of text files, use tools such as PyPDF2 or pdfplumber to parse PDF documents and extract text, crawl online question banks and answers through customized crawler scripts, and parse HTML structures to extract data text;
[0013] Perform denoising processing on the extracted data, delete useless characters, use NLP spelling - checking tools to correct errors, label fields, distinguish questions, standard answers, and user answers, and at the same time use a similarity algorithm to deduplicate the data;
[0014] Format the cleaned data into a unified JSON format, extract key information based on YAYI - UIE, use a pre - trained Embedding model to generate vector representations for retrieval, and use Milvus to build a vector database to achieve efficient similarity retrieval.
[0015] Among them, design a multi - dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time - sensitivity evaluation. The scoring results for each dimension are obtained based on specific algorithms and models. The steps also include:
[0016] Use a fine - tuned language model to extract knowledge points from the reference answers and user answers;
[0017] Map the extracted knowledge points to the vector space and calculate the similarity between knowledge points;
[0018] Assign weights according to the importance of knowledge points to generate scoring results.
[0019] Among them, design a multi - dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time - sensitivity evaluation. The scoring results for each dimension are obtained based on specific algorithms and models. The steps also include:
[0020] Vectorize the user answers and reference answers using pre-trained Embedding technology;
[0021] Calculate the similarity between text vectors through cosine similarity.
[0022] Among them, a multi-dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time sensitivity evaluation is designed. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include:
[0023] Use perplexity to evaluate the fitting degree of the language model of the sentence.
[0024] Among them, a multi-dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time sensitivity evaluation is designed. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include:
[0025] Calculate the answering duration;
[0026] And perform dynamic adjustment of time weight allocation according to the answering duration.
[0027] Among them, a three-layer structure weight mechanism is used to calculate the comprehensive score. Among them, the three-layer structure includes a task layer, a scenario layer, and an element layer. The steps also include:
[0028] In the scenario layer and the element layer, compare the scoring elements pairwise to construct an importance matrix;
[0029] Normalize the matrix columns to ensure that the sum of each column is 1;
[0030] Take the mean of each row to obtain the weight of each scoring element.
[0031] A multi-dimensional interpretable subjective question scoring method based on a large model according to the present invention comprehensively evaluates user answers from multiple perspectives by integrating multi-dimensional elements such as key content matching, text similarity analysis, and sentence fluency evaluation. Especially for subjective questions involving multiple knowledge points and complex expressions, it can dynamically adjust the weights according to the importance of the knowledge points to ensure that the scoring results not only reflect the overall quality of the answers but also highlight the examination of the core content, significantly improving the accuracy and fairness of scoring.
[0032] A traceability mechanism for the scoring process is designed, and the generated scoring report includes detailed score sources, scores of each dimension, and weight allocation ratios. Through these clear scoring bases, teachers and students can intuitively understand the logic behind the scoring.
[0033] Through the user feedback mechanism and the dynamic weight adjustment mechanism, it supports optimizing the scoring criteria according to actual needs. It can continuously improve in different educational scenarios to meet diverse needs.
[0034] Combined with retrieval enhancement and model fine-tuning technologies, it automatically completes the scoring process, reduces manual intervention, and saves educational resources. Through a multi-dimensional scoring system, a three-level weight mechanism, and a feedback optimization strategy, it comprehensively improves the scientificity, transparency, and flexibility of scoring, providing an efficient, transparent, and sustainable optimization solution for subjective question scoring in the education field, which helps to improve the practicality and reliability of the automatic scoring system. Brief Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art.
[0036] Figure 1 It is a flowchart of the present invention for unified data cleaning and structuring processing of diverse data sources, including data acquisition, data cleaning and verification, extraction of key data information, data structuring processing and storage.
[0037] Figure 2 It is an example diagram of the results of using the YAYI-UIE model to extract the above-structured data in the present invention.
[0038] Figure 3 It is an example diagram of the example data of the key content matching in the present invention.
[0039] Figure 4 It is an example diagram of the generated scoring basis report in the present invention.
[0040] Figure 5 It is an example diagram of the information generated during the scoring process in the present invention.
[0041] Figure 6 It is an example diagram of the working process of the overall system in the present invention.
[0042] Figure 7 It is a flowchart of the steps of the multi-dimensional interpretable subjective question scoring method based on a large model in the present invention. Detailed Embodiments
[0043] The following details the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0044] The first embodiment of the present application is:
[0045] Please refer to Figures 1 to 7 , wherein,Figure 1 It is a flow chart of the present invention for uniformly cleaning and structuring diverse data sources, including data acquisition, data cleaning and verification, extraction of key data information, data structuring processing and storage. Figure 2 It is an example diagram of the result of using the YAYI-UIE model to extract the above structured data in the present invention. Figure 3 It is an example diagram of the example data for key content matching in the present invention. Figure 4 It is an example diagram of the generated scoring basis report in the present invention. Figure 5 It is an example diagram of the generated information during the scoring process in the present invention. Figure 6 It is an example diagram of the working process of the overall system in the present invention. Figure 7 It is a step flow chart of the multi-dimensional interpretable subjective question scoring method based on a large model in the present invention.
[0046] The present invention provides a multi-dimensional interpretable subjective question scoring method based on a large model, including the following steps:
[0047] S101: Uniformly clean and structure diverse data sources, including data acquisition, data cleaning and verification, extraction of key data information, data structuring processing and storage;
[0048] Specifically, for diverse data sources, the present invention designs a unified data cleaning and structuring process to ensure data quality and format consistency:
[0049] 1. Data extraction:
[0050] Text data: Directly read and parse the file content.
[0051] PDF data: Use a library (such as PyPDF2 or pdfplumber) to parse the PDF document and extract the text content.
[0052] Online data: Through a customized crawler script, crawl the question bank and answers on the public open platform, and parse the HTML structure to extract the required data.
[0053] 2. Data cleaning and verification:
[0054] Denosing processing: Delete useless characters (such as line breaks, white space characters, special symbols).
[0055] Error correction: Use an NLP spelling check tool (such as SymSpell or Hunspell) to correct spelling mistakes and grammar errors.
[0056] Label fields: Add corresponding annotation information according to the content type (question, standard answer, user answer).
[0057] Data deduplication: Use similarity algorithms (such as cosine similarity, SimHash) to compare the currently read data with the data existing in the library, and set a threshold for duplicate detection.
[0058] 3. Data structuring, processing, and storage:
[0059] The data is formatted into a unified JSON format, with fields including question (topic), standard_answer (standard answer), related_knowpoints (related knowledge points), etc.
[0060] Extract key information based on YAYI-UIE.
[0061] Generate vector representations through a pre-trained Embedding model (such as xiaobu-embedding-v2 / gte-Qwen2-7B-instruct) for use by the retrieval module.
[0062] Use Milvus to build a vector database for efficient similarity retrieval.
[0063] The following is a detailed example:
[0064] Original data: The following information about "the concept of false sharing and measures to avoid this problem" was collected from a technical document (after processing some spaces and special characters).
[0065] "Oh, false sharing, you know, it's when you have several threads playing with the same cache line. Even if they're operating on different variables, due to the cache coherence protocol, the cache line will be invalidated frequently, and this phenomenon can really drag down your performance! Want to avoid it? Try adding some useless data between variables to fill the cache line, or use atomic operations to reduce contention. #Performance Optimization#Cache Line#False Sharing"
[0066] Key information extraction: Use the YAYI-UIE model to extract the above structured data, and the result is as Figure 2 .
[0067] Vector representation generation and storage:
[0068] Use a pre-trained Embedding model to generate vector representations for the above "question" field and store them in the Milvus vector database.
[0069] Vector representation generation:
[0070] Technology: Use a pre-trained Embedding model to generate vector representations.
[0071] Example:
[0072] Input text: "What is false sharing and how to avoid it?"
[0073] Vector representation: Assume that the vector representation after the model processing is [0.3, -0.4, 0.2,...] (the actual vector dimension is very high, and it is simplified here for representation).
[0074] Vector database construction:
[0075] Technology: Milvus is used to store and retrieve vector data.
[0076] Operation: Store the vectors generated above and their associated questions into Milvus.
[0077] Retrieval: When the user asks a new question, such as "How to solve the false sharing problem", retrieve the vector most similar to the new question through Milvus, and then find the most matching standard answer.
[0078] S102: Design a multi-dimensional scoring system including key content matching, text similarity analysis, sentence fluency evaluation, and time sensitivity evaluation. The scoring results of each dimension are obtained based on specific algorithms and models;
[0079] Specifically, the multi-dimensional scoring system is the core technology of the present invention, aiming to comprehensively evaluate the answers to subjective questions by dynamically integrating multiple evaluation dimensions. The scoring results of each dimension are obtained based on specific algorithms and models, and finally the comprehensive score is calculated through a weight mechanism:
[0080] Formula description:
[0081]
[0082] Where:
[0083] n represents the number of scoring dimensions (dynamically changing),
[0084] w i is the weight of the i-th scoring dimension,
[0085] Score i is the score of the i-th scoring dimension.
[0086] 1. Key content matching:
[0087] Purpose: Detect whether the user's answer covers the core knowledge points in the reference answer.
[0088] Implementation method:
[0089] Use the fine-tuned language model (such as YAYI-UIE or ChatGLM) to extract knowledge points from the reference answer and the user's answer. Here, the reference answer may be the answer extracted when collecting question-answer pairs, or it may be the answer generated by the large model and verified later when it did not exist during collection.
[0090] Map the extracted knowledge points to the vector space and calculate the similarity between the knowledge points.
[0091] Assign weights according to the importance of the knowledge points to generate the scoring result.
[0092] Formula description:
[0093]
[0094] Among them, w i is the knowledge point weight, and Sim(u i , s i ) is the similarity between the user's answer u i and the reference answer s i on the knowledge point i. Example data is as Figure 3 .
[0095] 2. Text similarity analysis:
[0096] Purpose: Analyze the overall text semantic similarity between the user's answer and the standard answer.
[0097] Implementation method:
[0098] Use the pre-trained Embedding technology to vectorize the user's answer and the reference answer.
[0099] Calculate the similarity between the text vectors through cosine similarity.
[0100] Formula description:
[0101]
[0102] Example:
[0103] Standard answer vector: [0.2, 0.5, 0.3]
[0104] User answer vector: [0.3, 0.4, 0.3]
[0105] Similarity score: CosSim = 0.98
[0106] 3. Sentence fluency evaluation:
[0107] Purpose: Measure the language expression quality of the user's answer, including sentence structure and logic.
[0108] Implementation method:
[0109] The language model fitting degree of a sentence is evaluated using Perplexity.
[0110] A low Perplexity indicates that the sentence is fluent and grammatically correct, resulting in a high score.
[0111] Formula description:
[0112]
[0113] where P(w i ) is the predicted probability of the i-th word, and N is the sentence length.
[0114] 4. Time sensitivity evaluation:
[0115] Purpose: For specific scenarios such as time-limited tests or real-time Q&A systems, it is necessary to quickly evaluate the completion status and content validity of the user's answer within the specified time, and evaluate the timeliness of the answer and the accuracy of time-related content.
[0116] Implementation method:
[0117] Response time calculation: ResponseTime = T submit -T start
[0118] Time weight assignment: It can be dynamically adjusted according to the actual situation. For example, the specified answering time is set, and the weight is 1.0 within the interval, and the weight decreases when exceeding the interval:
[0119]
[0120] where T limit is the specified time, and α is the time weight decay coefficient.
[0121] S103: A three-layer structure weight mechanism is adopted to calculate the comprehensive score. Among them, the three-layer structure includes a task layer, a scenario layer, and a factor layer;
[0122] Specifically, the weight mechanism solves the flexibility and scientificity problems in multi-task and multi-scenario scoring through hierarchical design and dynamic adjustment strategies, and provides a dynamic and accurate weight assignment method for the comprehensive score. The weight mechanism is based on a three-layer structure, and the assignment rules of the scoring weights are clarified layer by layer:
[0123] 1. Task layer:
[0124] The task layer determines the overall goal of scoring and its weight ratio. For example, "knowledge mastery assessment" (70%) and "innovation ability assessment" (30%).
[0125] The weights of this layer are configured by the user or automatically assigned by the system according to the task type to ensure that the global scoring objectives are clear and definite.
[0126] 2. Scenario layer:
[0127] The scenario layer adjusts the weights for specific application scenarios in the task. For example:
[0128] In the "time-limited test" scenario, time sensitivity (40%) and key content matching (30%) carry the main weights;
[0129] In the "open discussion" scenario, innovation (40%) and language fluency (30%) are more important.
[0130] 3. Factor layer:
[0131] The factor layer further allocates weights to specific scoring factors (such as key content matching, text similarity, time sensitivity, etc.) within the scenario.
[0132] The initial values of the weights for each factor can be determined by the Analytic Hierarchy Process (AHP), supporting dynamic adjustment and normalization.
[0133] Weighted Analytic Hierarchy Process (AHP)
[0134] In the scenario layer and the factor layer, the present invention uses the weighted Analytic Hierarchy Process (AHP) to calculate the initial weights. The specific steps are as follows:
[0135] Construct a judgment matrix: Experts or the system compare the scoring factors pairwise to construct an importance matrix.
[0136] Normalize the matrix: Normalize the columns of the matrix to ensure that the sum of each column is 1.
[0137] Calculate the weights: Take the mean of each row to obtain the weights of each scoring factor.
[0138] Formula description:
[0139]
[0140] where w i is the weight of the i-th factor, and a ij is the element in the judgment matrix.
[0141] S104: Generate a scoring basis report and construct a scoring log, and continuously optimize the scoring rules through a dynamic feedback mechanism.
[0142] Specifically, through the interpretability and scoring traceability mechanism, ensure that the scoring process is transparent and fair, and support dynamic adjustment and optimization based on user feedback. By refining the scoring steps, generating a scoring basis report, and constructing a scoring log, comprehensively improve the credibility and adaptability of the scoring results, and further improve the scoring rules through the dynamic feedback mechanism. It is mainly divided into the following parts:
[0143] 1. Interpretability of the scoring process:
[0144] By generating a detailed scoring basis report, provide a clear analysis of the scoring results. The report mainly includes the scores, weight distributions of each scoring dimension and their contributions to the final score, as well as the detailed calculation basis of the score. For example Figure 4 .
[0145] The contribution value of each scoring dimension is calculated by the formula Contribution i = w i ·Score i For example, the contribution value of key content matching is 0.4 × 0.85 = 0.34, accounting for 34% of the total score. In addition, the report also includes a transparent display of the weight distribution, enabling users to intuitively understand the impact of the scoring dimension on the total score, and explaining the calculation logic of each dimension by listing the refinement process of the scoring basis. For example, key content matching scores based on the knowledge point coverage rate and weight value, text similarity is calculated based on cosine similarity, and language fluency is evaluated through perplexity.
[0146] 2. Scoring traceability:
[0147] The scoring traceability mechanism supports the review and dynamic optimization of the scoring results by recording the key data of the entire scoring process.
[0148] The scoring log generates a detailed log during the scoring process, covering information such as user input, scoring steps, weight settings, model version numbers, and final scoring results. An example is Figure 5 .
[0149] Traceability query and review:
[0150] The log record provides a traceable basis for scoring. For example, students can query the scores and contribution values of each dimension of their own scores to understand the scientificity and fairness of the scoring; teachers or administrators can review the scoring steps based on the log data to verify whether the scoring rules and calculation processes meet the expectations.
[0151] 3. Dynamic feedback and rule optimization:
[0152] Combined with the user feedback mechanism, it can incorporate dynamic adjustment capabilities into the scoring rules and support the continuous optimization of the scoring mechanism. It is mainly divided into the following steps:
[0153] Collection and Analysis of User Feedback:
[0154] Allow users to provide feedback based on the scoring report, such as stating that "the time sensitivity score is too high" or "the language fluency score is unreasonable". Record this feedback and associate it with the scoring log for subsequent analysis.
[0155] Integration of Expert Opinions:
[0156] Experts review the scoring report and log and provide opinions on adjusting the rationality of the scoring rules. For example, in response to user doubts about a certain scoring dimension in the feedback, experts can recommend re-evaluating the weight of that dimension or optimizing the calculation formula.
[0157] Through an innovative multi-dimensional scoring system and a dynamic weight adjustment mechanism, combined with the powerful semantic analysis ability of the large language model (LLM), the present invention significantly improves the performance of the subjective question scoring system in practical applications, mainly reflected in the following aspects:
[0158] 1. Improve the accuracy and fairness of scoring
[0159] The present invention comprehensively evaluates the user's answer from multiple perspectives by integrating multi-dimensional elements such as key content matching, text similarity analysis, and sentence fluency evaluation. Especially for subjective questions involving multiple knowledge points and complex expressions, this system can dynamically adjust the weight according to the importance of the knowledge points to ensure that the scoring result not only reflects the overall quality of the answer but also highlights the examination of core content, significantly improving the accuracy and fairness of scoring.
[0160] In practical application scenarios, such as when an educational examination system needs to evaluate open-ended questions, this system can avoid biases caused by a single scoring dimension and fairly evaluate the quality of students' answers.
[0161] 2. Enhance the transparency and interpretability of scoring results
[0162] The present invention designs a traceability mechanism for the scoring process, and the generated scoring report includes detailed score sources, scores for each dimension, and weight allocation ratios. Through these clear scoring bases, teachers and students can intuitively understand the logic behind the scoring.
[0163] In practical applications, for example, when a student questions the score of a subjective question, the system can quickly provide the scoring basis, reduce disputes, and enhance the trust and fairness of the examination scoring.
[0164] 3. Dynamic optimization and continuous improvement capabilities
[0165] Through the user feedback mechanism and the dynamic weight adjustment mechanism, the present invention supports optimizing the scoring criteria according to actual needs. For example, users can put forward opinions on the scoring results, and the system incorporates reasonable feedback into model training to gradually optimize the scoring model. This dynamic adaptation ability enables the system to continuously improve in different educational scenarios and meet diverse needs.
[0166] In practical applications, such as customized tests in online learning platforms, the present invention can adjust the weights of scoring elements according to teachers' preferences to adapt to different disciplines and assessment objectives.
[0167] 4. Improve the efficiency and practicality of the education system
[0168] The present invention combines retrieval enhancement and model fine-tuning technologies to automatically complete the scoring process, reduce manual intervention, and save educational resources. In practical applications, such as online exam systems and intelligent education assessment platforms, the present invention can efficiently handle a large number of subjective question scoring tasks and meet the needs of large-scale exams and assessments. At the same time, the system supports scoring different question types (such as essay questions, case analysis questions), improving the practicality of the intelligent education system.
[0169] Through the multi-dimensional scoring system, the three-level weight mechanism, and the feedback optimization strategy, comprehensively improve the scientificity, transparency, and flexibility of scoring, providing an efficient, transparent, and sustainable optimization solution for subjective question scoring in the education field, which helps to improve the practicality and reliability of the automatic scoring system.
[0170] The above-disclosed are only one or more preferred embodiments of the present application, and the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A multi-dimensional interpretable subjective question scoring method based on a large model, characterized by: The following steps are involved: Perform unified data cleaning and structuring on various data sources, including data acquisition, data cleaning and verification, data key information extraction, data structuring processing and storage; Design a multi-dimensional scoring system including key content matching, text similarity analysis, sentence fluency assessment and time sensitivity evaluation. The scoring results of each dimension are based on specific algorithms and models. A three-layer weighting mechanism is used to calculate the comprehensive score, where the three-layer structure includes the task layer, the scenario layer, and the element layer; Generate scoring basis reports and build scoring logs, and continuously optimize scoring rules through a dynamic feedback mechanism.
2. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: Perform unified data cleaning and structuring on diverse data sources, including data acquisition, data cleaning and verification, data key information extraction, data structuring processing and storage, and the steps also include: Directly read and parse text file contents, use tools such as PyPDF2 or pdfplumber to parse PDF documents, extract text, crawl online question banks and answers through custom crawler scripts, and parse HTML structures to extract data text; De-noise the extracted data, delete useless characters, use NLP spell checking tools to correct errors, annotate fields, distinguish between questions, standard answers, and user answers, and use similarity algorithms to remove duplicate data; The cleaned data is formatted into a unified JSON format, key information is extracted based on YAYI-UIE, a pre-trained Embedding model is used to generate vector representation for retrieval, and a vector database is built using Milvus to achieve efficient similarity retrieval.
3. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: A multi-dimensional scoring system is designed, including key content matching, text similarity analysis, sentence fluency assessment, and time sensitivity evaluation. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include: Use the fine-tuned language model to extract knowledge points from reference answers and user answers; Map the extracted knowledge points to the vector space and calculate the similarity between the knowledge points; Assign weights according to the importance of knowledge points and generate scoring results.
4. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: A multi-dimensional scoring system is designed, including key content matching, text similarity analysis, sentence fluency assessment, and time sensitivity evaluation. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include: Use pre-trained Embedding technology to vectorize user answers and reference answers; The similarity between text vectors is calculated by cosine similarity.
5. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: A multi-dimensional scoring system is designed, including key content matching, text similarity analysis, sentence fluency assessment, and time sensitivity evaluation. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include: Use perplexity to evaluate the language model fit of a sentence.
6. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: A multi-dimensional scoring system is designed, including key content matching, text similarity analysis, sentence fluency assessment, and time sensitivity evaluation. The scoring results of each dimension are obtained based on specific algorithms and models. The steps also include: Calculate the time it takes to answer the questions; And dynamically adjust the time weight distribution according to the length of answering questions.
7. The multi-dimensional interpretable subjective question scoring method based on a large model as claimed in claim 1, characterized in that: A three-layer weight mechanism is used to calculate the comprehensive score, wherein the three-layer structure includes a task layer, a scenario layer, and an element layer. The steps further include: In the scenario layer and the factor layer, the scoring factors are compared pairwise to construct an importance matrix; Normalize the matrix columns to ensure that the sum of each column is 1; Take the mean of each row to get the weight of each scoring factor.
Citation Information
Patent Citations
A subjective question scoring method and device based on knowledge point annotation
CN116595129B
Subjective question scoring method and system based on ALBERT model and RPA technology
CN117540727B
Vertical domain subjective item scoring model selection method and vertical domain subjective item scoring method
CN118797031A
Cited By
Method for enhancing teaching evaluation credibility in large education model
CN120672216A
A method for enhancing teaching evaluation credibility in an educational large model
CN120672216B
Subjective question scoring method and device based on H5 terminal and related medium
CN121118909A
Data processing method and device
CN121706911A