Online paper marking method based on artificial intelligence
Through the vectorized auxiliary scoring model and vLLM inference framework combined with the finite state machine method, the problems of inefficient and unstable output of large language models in online evaluation are solved, and efficient and stable online marking scoring is achieved.
Patent Information
- Application Number
- CN202510501976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-11
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, when online evaluation is used to score using large language models, there are problems of inefficiency and unstable output.
The online marking method based on artificial intelligence is adopted, and the scoring is performed using vectorized auxiliary scoring model and large language model on the vLLM inference framework. Combined with the finite state machine to control the sampling and decoding of tokens, the scoring efficiency and stability are improved.
It improves the efficiency of online paper marking and the output stability of large language models, ensuring the accuracy and consistency of the scoring results.
Smart Images

Figure CN120387912A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of automatic marking, and specifically relates to an online marking method based on artificial intelligence. Background Art
[0002] In the process of online education, online assessment is an important part of online education. Currently, for online assessment, online test papers are used for testing, and users can answer questions and submit papers online, which is very convenient and efficient.
[0003] In the process of online assessment, in order to fully detect the user's mastery of educational content, educators often set various question types, such as fill-in-the-blank questions, noun explanation questions, essay questions, short answer questions, discrimination questions, discussion questions, translation questions, single-choice questions, multiple-choice questions, work appreciation questions, writing questions, and calculation questions, etc.
[0004] In the prior art, in order to efficiently and automatically score online test papers, large language models are used for intelligent scoring. However, when using large language models for scoring, there are often situations of low efficiency and unstable output structures of large language models. Summary of the Invention
[0005] In order to solve the deficiencies of the prior art, the purpose of this application is to provide an online marking method based on artificial intelligence, which can improve the efficiency and stability of using large language models for marking.
[0006] To achieve the above purpose, this application provides an online marking method based on artificial intelligence, including: To achieve the above purpose, the electronic device provided by this application includes: In response to the marking of the user's online test paper, determine the first type of subjective questions and the second type of subjective questions in the online test paper; wherein, the first type of subjective questions includes at least one of fill-in-the-blank questions and noun explanation questions, and the second type of subjective questions includes at least one of essay questions, short answer questions, discrimination questions, discussion questions, and translation questions; Establish a marking task corresponding to the online test paper into a task queue, and the marking task includes reading tasks corresponding to the first type of subjective questions and the second type of subjective questions; Poll the marking tasks in the marking task queue. For the reading tasks of the first type of subjective questions in the marking tasks, assign them to the vectorization-assisted scoring model for scoring; for the reading tasks of the second type of subjective questions in the marking tasks, use the large language model on the vLLM inference framework for scoring; wherein, in the process of using the large language model on the vLLM inference framework for scoring, use a preset finite state machine to control the sampling and decoding of tokens; In response to the completion of scoring all reading tasks in the marking task, the marking is completed.
[0007] Further, the specific steps of using a preset finite state machine to control the sampling and decoding of tokens include: Based on the logits tensor output by running the large language model on the vLLM inference framework, determine the model-predicted tokens; Based on the model-predicted tokens and the current state of the finite state machine, determine whether the model-predicted tokens meet the expectations of the current state and whether subsequent tokens can be determined based on the model-predicted tokens; If it is determined that the model-predicted tokens meet the expectations of the current state and subsequent tokens can be determined, output the model-predicted tokens and subsequent tokens, and update the state of the finite state machine; If it is determined that the model-predicted tokens do not meet the expectations of the current state, output the expected tokens or the model alternative tokens that meet the expectations, and update the state of the finite state machine; Decode the output tokens into text characters and add them to the output sequence.
[0008] Further, the model alternative tokens that meet the expectations are the tokens selected from the logits tensor that meet the expectations of the current state.
[0009] Further, for the question reading task of the second type of subjective questions in the marking task, the specific steps of using the large language model on the vLLM inference framework to score include: Use the prompt template corresponding to the question type to generate the prompt words for judging whether each scoring point in the reference answer is included in the user's question answer; The vLLM inference framework receives the prompt words and inference parameters and initializes the inference request; Based on the prompt words and inference parameters, the large language model deployed on the vLLM inference framework performs forward inference and outputs the logits tensor; Based on the logits tensor, use a preset finite state machine to control the sampling and decoding of tokens, and add the decoding result to the output sequence; Append the newly generated token IDs to the input token IDs sequence and perform the next round of forward inference until the termination condition is met.
[0010] Further, the prompt templates for different question types are configured with the response requirements and response examples corresponding to the question types.
[0011] Further, for the question reading task of the first type of subjective questions in the marking task, the specific steps of assigning the scoring to the vectorized auxiliary scoring model include: Use the vectorization-assisted scoring model to convert the user's question answers and the reference answers into high-dimensional vectors respectively; Calculate the cosine similarity between the high-dimensional vectors corresponding to the user's question answers and the reference answers to obtain a cosine similarity score; Generate a score based on the cosine similarity score and the total score of the question reading task, and update the question score in the user answer record table.
[0012] Further, the method further includes: In response to the timeout of the question reading task for the second type of subjective questions without processing, assign the question reading task to the vectorization-assisted scoring model for scoring.
[0013] Further, the method further includes: Determine the objective questions and the third type of subjective questions in the online test paper. The objective questions include single-choice questions and multiple-choice questions, and the third type of subjective questions include works appreciation questions, writing questions, and calculation questions; For each objective question, use the string comparison method to compare the user's question answer and the reference answer, generate a score and update the question score in the user answer record table; for each third type of subjective question, conduct manual marking, generate a score and update the question score in the user answer record table.
[0014] To achieve the above object, an electronic device of the present application includes: A processor; A memory, on which one or more computer instructions running on the processor are stored; When the processor runs the computer instructions, it executes the steps of the online marking method based on artificial intelligence as described above.
[0015] To achieve the above object, the present application provides a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are run by a processor, the steps of the online marking method based on artificial intelligence as described above are executed.
[0016] In the online marking method based on artificial intelligence of the present application, when using a large language model for scoring, a preset finite state machine is used to control the sampling and decoding process of tokens, which improves the marking efficiency while ensuring the output stability of the large language model.
[0017] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present application. Description of the Drawings
[0018] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the accompanying drawings: Figure 1 It is a schematic flowchart of the online marking method based on artificial intelligence in Embodiment 1 of the present application; Figure 2 It is a schematic flowchart of scoring using a large language model on the vLLM inference framework; Figure 3 It is a partial schematic diagram of the state transition table of the finite state machine in Embodiment 1. Detailed implementation manners
[0019] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the accompanying drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0020] It should be understood that the steps recited in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0021] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions of other terms will be given in the following description.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more". "Multiple" should be understood as two or more.
[0023] vLLM (Vectorized Large Language Model Serving System) is an efficient large language model inference framework developed by the team at the University of California, Berkeley. Its core innovation, the **PagedAttention technology**, achieves dynamic video memory allocation by managing key-value caches (KV Cache) in chunks, with a maximum video memory utilization rate of 99%, significantly improving throughput (e.g., LLaMA-7B can reach 1075 tokens / s on A100, 24 times higher than HuggingFace). This framework supports the deployment of mainstream large language models (such as LLaMA, Qwen, DeepSeek, Baichuan, etc.) and provides flexible deployment methods: starting an OpenAI-compatible API service through the command line or code, supporting multi-GPU distributed inference (such as deploying the QwQ-32B model with dual A6000 GPUs); it can also combine containerization technology (such as the Alibaba Cloud vLLM image) to quickly build an inference environment, simplify dependency configuration, and support out-of-the-box use in the cloud. In practical applications, vLLM optimizes the first token latency through dynamic batching and streaming output, is suitable for high-concurrency scenarios such as chatbots and text generation, and is also compatible with frameworks such as LangChain to achieve complex application integration. After developers install it via `pip install vllm`, they can call the `LLM` class to load the model and generate text, or connect to the service through the REST API, supporting advanced configurations such as adjusting the video memory block size and enabling dynamic batching, taking into account both performance and resource efficiency.
[0024] A finite state machine (FSM) is a formal computational model that describes the behavior of a system and consists of a finite number of states, transition rules between states, and events that trigger transitions. Its core idea is that the system is only in a specific state at any given moment and switches between states driven by external inputs or internal events. A typical FSM includes five elements: a set of states (such as "standby", "running", "fault"), an initial state, a set of input events, a transition function (defining how events trigger state changes), and a final state (optional). It can be divided into a deterministic finite state machine (DFA) (each state has only one transition path for the same input) and a non-deterministic (NFA) (allowing multi-path transitions) according to determinism. FSMs are widely used in the fields of computer science and engineering, such as compiler lexical analysis (identifying programming language symbols), automatic control systems (elevator state switching), communication protocol design (TCP connection state management), game AI (NPC behavior logic), and other scenarios.
[0025] Next, embodiments of the present application will be described in detail with reference to the accompanying drawings.
[0026] Embodiment 1 An embodiment of the present application provides an online marking method based on artificial intelligence, which will be described in detail below with reference to Figures 1-3 the online marking method based on artificial intelligence of the present application.
[0027] Figure 1 As shown in Figure 1 the flowchart of the online marking method based on artificial intelligence in Embodiment 1 of the present application, the online marking method includes: Step S101: In response to the marking of the user's online test paper, determine each subjective question in the online test paper that belongs to the first type of subjective questions and the second type of subjective questions.
[0028] In this embodiment, the user conducts an online exam and submits the answer sheet through a web interface provided by a Java web project.
[0029] In this embodiment, the question types of the online test paper include: The first type of subjective questions: fill-in-the-blank questions, noun explanation questions; The second type of subjective questions: essay questions, short answer questions, discrimination questions, discussion questions, translation questions; Objective questions: single-choice questions, multiple-choice questions; The third type of subjective questions: work appreciation questions, writing questions, calculation questions; In this embodiment, when the user submits the test paper, relevant information will be stored.
[0030] It can be understood that according to the actual question type settings on the test paper, the first type of subjective questions may also only include fill-in-the-blank questions or noun explanation questions, and the second type of subjective questions may also only include one or more of essay questions, short answer questions, discrimination questions, discussion questions, and translation questions.
[0031] Step S102: Establish a marking task corresponding to the online test paper into a task queue, and the marking task includes marking tasks corresponding to the first type of subjective questions and the second type of subjective questions; In this embodiment, when establishing a marking task corresponding to the online test paper, the marking task data of each subjective question will be assembled into marking task data.
[0032] Step S103: Poll the marking tasks in the marking task queue. For the marking tasks of the first type of subjective questions in the marking task, assign them to the vectorized auxiliary scoring model for scoring; for the marking tasks of the second type of subjective questions in the marking task, use the large language model on the vLLM inference framework for scoring.
[0033] In this embodiment, the specific steps of assigning the marking tasks of the first type of subjective questions in the marking task to the vectorized auxiliary scoring model for scoring include: According to the question-reading task data corresponding to the question-reading task, use the vectorization-assisted scoring model to convert the user's question answers and reference answers in the question-reading task data into high-dimensional vectors respectively; In this embodiment, the vectorization-assisted scoring model used is the BAAI / bge-large-zh-v1.5 model. This vectorization-assisted scoring model uses text embedding technology, also known as word embedding or sentence embedding, which is a method of converting text into a dense vector representation. The core idea of this technology is to map discrete text information into a continuous vector space, so that texts with similar semantics are close in this space. Specifically, text embedding technology encodes each word or the entire sentence in the text into a vector of a fixed dimension through a deep learning model. This vector not only encodes the literal information of the word, but also contains rich information such as its semantics, grammatical functions, and even context relationships. Using the BAAI / bge-large-zh-v1.5 model can analyze the answers to the first type of subjective questions such as fill-in-the-blank questions and noun explanation questions at the semantic level, thereby making the subsequent scoring more accurate.
[0034] In some other embodiments, other vectorization-assisted scoring models using text embedding technology can also be used, such as the moka-ai / m3e-base model, the BAAI / bge-large-zh model, and the openAi / text-embedding-ada-002 model, etc.
[0035] Calculate the cosine similarity between the high-dimensional vectors corresponding to the user's question answer and the reference answer to obtain a cosine similarity score; Cosine similarity is a commonly used similarity measurement method in vector space. It measures the angle between two vectors and can better reflect semantic similarity, that is, better judge the similarity between the user's question answer and the reference answer. The similarity score is a value between 0 and 1.
[0036] Based on the cosine similarity score and the total score of the question-reading task, generate a score and update the question score in the user answer record table.
[0037] Specifically, according to the calculated similarity score and the total score of the question, the similarity score is converted into a specific score through a preset mapping function. The mapping function is set artificially according to actual needs.
[0038] Figure 2 It is a schematic diagram of the process of scoring using a large language model on the vLLM inference framework. As Figure 2 shown, for the question-reading task of the second type of subjective questions in this marking task, the specific steps of scoring using a large language model on the vLLM inference framework include: S201: Generate a prompt word using the prompt template corresponding to the question type to determine whether each scoring point in the reference answer is included in the user's question answer; In this embodiment, for the question reading task of the second type of subjective questions, a prompt template corresponding to the question type is used to generate a prompt word for determining whether each scoring point in the reference answer is included in the user's question answer; In this embodiment, an agent is used to assemble and generate the prompt word. Different prompt templates and agents are configured for the second type of subjective questions of different question types. For questions of different question types, the agent corresponding to the question type disassembles the question information and the standard answer and assembles them into the prompt template of the prompt word corresponding to the question type to generate the prompt word.
[0039] In this embodiment, the prompt templates of different question types are configured with response requirements and response examples corresponding to the question types.
[0040] In this embodiment, for translation questions and short answer questions, since their subjectivity is relatively low, reference answers are set in the question bank; for essay questions, discussion questions, and discrimination questions, since they all tend to subjective analysis and have a high degree of subjectivity, setting reference answers is not very meaningful, and reference answers are usually not set in the question bank, only reasonable answers are required. For some special essay questions, discussion questions, and discrimination questions, reference answer scoring points are also set in the question bank.
[0041] In this embodiment, there is one prompt template for translation questions, one prompt template for short answer questions, and the same prompt template for essay questions, discussion questions, and discrimination questions.
[0042] Generally, the prompt template includes response requirements, response examples, and input. The response requirements are the requirements for the large language model when processing the prompt word, the response example is an example, and the input is the question information to be assembled.
[0043] It should be noted that usually, response examples may not be set in the prompt template corresponding to essay questions, discussion questions, and discrimination questions.
[0044] It can be understood that the response requirements and response examples of the prompt templates of different question types can also be adaptively modified according to the actual marking requirements.
[0045] S202: The vLLM inference framework receives the prompt word and inference parameters and initializes the inference request; The vLLM inference framework receives the inference request from the user, which usually includes the input text prompt (prompt) and some inference parameters (such as the maximum number of generated tokens, temperature, etc.).
[0046] Prompt preprocessing and Tokenization: The vLLM inference framework also preprocesses the received text prompts, such as cleaning, formatting, etc. Then, it uses the tokenizer corresponding to the pre-trained model to convert the text prompts into a sequence of token IDs, which is the numerical representation that the model can understand.
[0047] Then, a RequestContext is created. A context object is created for each inference request to store the state information of the request, including the sequence of token IDs, model parameters, and intermediate states during the inference process.
[0048] Allocate Page Table: The vLLM inference framework uses the PagedAttention technique to efficiently manage the attention key and value caches. During the request initialization phase, a Page Table is allocated for the request to store the attention cache subsequently.
[0049] S203: Based on the prompt and inference parameters, the large language model deployed on the vLLM inference framework performs forward inference and outputs a logits tensor. In this embodiment, the large language model used is the Qwen-14B large language model. Qwen-14B is one of many open-source pre-trained LLMs, and it has the best measured performance among large language models with the same number of parameters (14B, that is, about 14 billion parameters).
[0050] In some other embodiments, the Meta / Llama-3.1-8B-Instruct large language model, the baichuan-inc / Baichuan-13B-Chat large language model, and the ZhipuAI / glm-4-9b-chat large language model can also be used as the large language model.
[0051] In this embodiment, the process of the large language model performing forward inference and outputting a logits tensor includes: Embedding Lookup: The input sequence of token IDs is converted into a word vector representation through the Embedding layer.
[0052] Transformer Decoder Layers Inference: Input the word vectors into the Transformer decoder layers for forward computation. This part is the core of LLM inference and usually consists of multiple stacked decoder layers. Each decoder layer is mainly composed of the following components: PagedAttention mechanism: Calculate attention scores and aggregate context information. PagedAttention is the core optimization of vLLM. It stores the attention key and value caches separately in Pages and dynamically manages the Page Table, thereby reducing memory fragmentation and copying and improving efficiency; Feed-Forward Network (FFN): Perform further non-linear transformation and feature extraction on the output of the attention layer; Residual connection and Layer Normalization: Ensure the stability and convergence speed of model training.
[0053] Logits Computation: After passing through the last decoder layer, output the logits tensor. Logits is the unnormalized representation of the probability distribution of the next possible token by the large language model.
[0054] S204: Based on the logits tensor, use a preset finite state machine to control the sampling and decoding of tokens, and add the decoding result to the output sequence; It should be noted that in the prior art, when sampling and decoding tokens, a predetermined sampling strategy (such as greedy sampling, Top-k sampling, Temperature sampling, etc.) is often used to generate a token to be generated, and then decoded into text characters or words and added to the output sequence. This way of token sampling and decoding is likely to result in a high degree of freedom in the output of the large language model, with an uncertain structure and low stability.
[0055] In this embodiment, a preset finite state machine is used to control the sampling and decoding of tokens. The specific steps include: Based on the logits tensor output by running the large language model on the vLLM inference framework, determine the model-predicted token; It should be noted that the logits tensor is the confidence of the large language model in each token in the vocabulary. Then it will go through softmax to turn it into probabilities, and then the large language model will select a token according to the probabilities, which is the model-predicted token.
[0056] Based on the token predicted by the model and the current state of the finite state machine, determine whether the token predicted by the model meets the expectations of the current state and whether the subsequent token can be determined based on the token. In this embodiment, according to the structured content (such as JSON) defined by the output requirements of the prompt, define and design the state set of the finite state machine. Each state in the state set covers all the key components and hierarchical relationships of the structured content. According to the target structure of each state, formulate a state transition table. The state transition table describes which inputs (tokens / characters) are allowed to be received in each state, and which state should be transferred to after receiving the input. The rules need to ensure that the generated text conforms to the expected structure.
[0057] Exemplarily, in this embodiment, the finite state machine has the following states: START: Initial state, waiting for the start of a JSON object.
[0058] EXPECT_KEY_OPEN_QUOTE: { has been received, waiting for the starting double quote " of the key name. IN_KEY: Receiving the key name "question_analysis". An internal counter or pointer is needed to track which character of the key name is currently being matched.
[0059] EXPECT_KEY_CLOSE_QUOTE: The key name "question_analysis" has been fully received, waiting for the ending double quote " of the key name. EXPECT_COLON: The key name "question_analysis" and its ending quote have been received, waiting for the colon :. EXPECT_VALUE_OPEN_QUOTE: The colon : has been received, waiting for the starting double quote " of the value string. IN_VALUE: Receiving the content of the value string. Any character is allowed, and escape characters need to be processed.
[0060] IN_VALUE_ESCAPE: An escape character \ has been encountered in the value string, waiting for the next character (the escaped character).
[0061] EXPECT_VALUE_CLOSE_QUOTE: The value string has been received (internal state), waiting for the ending double quote " of the value string. (In practice, IN_VALUE can directly detect the ending quote and transfer) EXPECT_OBJ_CLOSE: The value string and its closing quote have been received, waiting for the closing curly brace} of the JSON object. END: Final acceptance state, the JSON structure is complete and correct.
[0062] ERROR: Error state, indicating that the generated sequence does not conform to the expected structure.
[0063] Exemplarily, refer to Figure 3 , a partial schematic diagram of the state transition table of the finite state machine in Embodiment 1.
[0064] In some other embodiments, the state set and state transition table of the finite state machine can also be modified according to the actual model output requirements.
[0065] If it is determined that the model-predicted token meets the expectation of the current state and the subsequent token can be determined, output the model-predicted token and the subsequent token, and update the state of the finite state machine; If it is determined that the model-predicted token does not meet the expectation of the current state, output the expected token or output the model alternative token that meets the expectation, and update the state of the finite state machine; Among them, the model alternative token that meets the expectation is the token that meets the expectation of the current state selected from the logits tensor.
[0066] The decoded output token is a text character and is added to the output sequence.
[0067] Exemplarily, if the model-predicted token is T_wrong, T_wrong is decoded as "wrong", and the corresponding characters 'w', 'r', 'o', 'n', 'g' are placed in char_buffer. The finite state machine first takes out 'w' from char_buffer. The expectation of the FSM in the IN_KEY (idx = 0) state is 'q'. 'w' does not match the set expectation 'q', so the expected token or the model alternative token that meets the expectation is output for error correction and then output.
[0068] It should be noted that if the correct output of a state is unique, it is considered that the state can determine the token. What is meant by being able to determine the subsequent token in this application is to judge whether the subsequent state of the current state can determine the token according to the above judgment criteria.
[0069] Exemplarily, refer to Figure 3, for states such as START state, EXCEPT_KEY_OPEN_QUOTE, IN_KEY state, EXCEPT_KEY_CLOSE_QUOTE state, EXCEPT_COLON state, and EXCEPT_VALUE_CLOSE_QUOTE state, the correct output is unique. These states can determine the token. When the model outputs, the tokens of these states can be output together with the tokens of the previous state, avoiding using the model for prediction, achieving accelerated inference, and significantly reducing the computational and video memory overhead.
[0070] S205: Append the newly generated token ID to the input token IDs sequence and perform the next round of forward inference until the termination condition is met.
[0071] It can be understood that the token ID and the token are set in a one-to-one correspondence as parameters for representing the token.
[0072] Append the newly generated token ID to the input token IDs sequence, then return to step S103 to perform the next round of forward inference of the model, and perform loop iteration until the termination condition is met.
[0073] In this embodiment, the termination conditions include: Reaching the maximum number of generated tokens.
[0074] The model generates the end-of-sequence token (EOS token).
[0075] The user actively stops the request.
[0076] In this embodiment, using a finite state machine ensures that the model output strictly conforms to the predefined structure. For scenarios that require a specific format output (such as JSON, XML, or even code structure), the finite state machine significantly improves the structural correctness of the output, avoids format errors that may occur when the large language model generates freely, and at the same time reduces the number of tokens that the model actually needs to infer, thereby accelerating the overall inference process. The acceleration effect is particularly obvious when generating highly structured and repetitive content.
[0077] In some other embodiments, post-processing will also be performed on the finally generated token IDs sequence, such as removing special tokens and further formatting the output, etc.
[0078] In this embodiment, since a vectorized auxiliary scoring model is used for scoring, the speed and efficiency of scoring far exceed those of using a large language model. Therefore, in order to improve the scoring efficiency, in this embodiment, when the question reading task of the second type of subjective questions times out and remains unprocessed, the corresponding question reading task will be assigned to the vectorized auxiliary scoring model for scoring.
[0079] It can be understood that there are two states for the question reading task, namely the unprocessed state and the processed state. When polling the question reading tasks in the marking task queue, it is the question reading tasks in the unprocessed state that are polled, and a question reading task will only be scored once.
[0080] In this embodiment, the online marking method further includes: Determine the objective questions and the third type of subjective questions in the online test paper. The objective questions include single-choice questions and multiple-choice questions, and the third type of subjective questions include works appreciation questions, writing questions, and calculation questions; For each objective question, use the string comparison method to compare the user's question answer with the reference answer, generate a score, and update the question score in the user answer record form; for each third type of subjective question, conduct manual marking, generate a score, and update the question score in the user answer record form.
[0081] Step S104: In response to all the question reading tasks in the marking task being scored, the marking is completed.
[0082] When the status of all question reading tasks is the processed state, calculate the question scores of all questions, obtain the test paper score, and update the test paper score in the user answer data.
[0083] Embodiment 2 In this embodiment, an electronic device is further provided, including a processor and a memory. The memory is used to store non-temporary computer-readable instructions. The processor is used to run the non-temporary computer-readable instructions, and when the non-temporary computer-readable instructions are run by the processor, one or more steps of the above-mentioned online marking method based on artificial intelligence can be executed. The memory and the processor can be interconnected through a bus system and / or other forms of connection mechanisms.
[0084] For example, the processor can be a central processing unit (CPU), a digital signal processor (DSP), or other forms of processing units with data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) can be of the X86 or ARM architecture, etc.
[0085] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage media, and the processor may run one or more computer program modules to implement various functions of the electronic device. Various application programs and various data, as well as various data used and / or generated by the application programs, etc., may also be stored in the computer-readable storage media.
[0086] It should be noted that in the embodiments of the present application, the specific functions and technical effects of the electronic device may refer to the description of the online marking method based on artificial intelligence in the foregoing text, and will not be elaborated herein.
[0087] Embodiment 3 In this embodiment, a computer-readable storage medium is further provided, and the storage medium is used to store non-temporary computer-readable instructions. For example, when the non-temporary computer-readable instructions are executed by a computer, one or more steps of the online marking method based on artificial intelligence described above may be executed.
[0088] For example, the storage medium may be applied to the above-mentioned electronic device. For example, the storage medium may be the memory in the electronic device of Embodiment 3. For example, the relevant description of the storage medium may refer to the corresponding description of the memory in the electronic device of Embodiment 3, and will not be elaborated herein.
[0089] It should be noted that the above-mentioned storage medium (computer-readable medium) of the present application may be a computer-readable signal medium or a non-temporary computer-readable storage medium or any combination of the two. The non-temporary computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the non-temporary computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0090] In this application, a non-transitory computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a non-transitory computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0091] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0092] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include but are not limited to object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in a block can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0095] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0096] The above description is only for some embodiments of this application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.
[0097] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0098] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. An online marking method based on artificial intelligence, characterized in that, Including: In response to grading a user's online test paper, determine the first type of subjective questions and the second type of subjective questions in the online test paper; wherein, the first type of subjective questions includes at least one of fill-in-the-blank questions and noun explanation questions, and the second type of subjective questions includes at least one of question-and-answer questions, short-answer questions, discrimination questions, essay questions, and translation questions; Establish a grading task corresponding to the online test paper into a task queue, and the grading task includes grading tasks corresponding to the first type of subjective questions and the second type of subjective questions; Poll the grading tasks in the grading task queue. For the grading tasks of the first type of subjective questions in the grading tasks, assign them to a vectorized auxiliary scoring model for scoring; for the grading tasks of the second type of subjective questions in the grading tasks, use a large language model on the vLLM inference framework for scoring; wherein, during the process of using the large language model on the vLLM inference framework for scoring, use a preset finite state machine to control the sampling and decoding of tokens; In response to the completion of scoring all grading tasks in the grading task, the grading is completed.
2. The online marking method based on artificial intelligence according to claim 1, wherein, The specific steps of using the preset finite state machine to control the sampling and decoding of tokens include: Based on the logits tensor output by running the large language model on the vLLM inference framework, determine the model-predicted token; Based on the model-predicted token and the current state of the finite state machine, judge whether the model-predicted token meets the expectations of the current state and judge whether the subsequent token can be determined based on the model-predicted token; If it is judged that the model-predicted token meets the expectations of the current state and the subsequent token can be determined, output the model-predicted token and the subsequent token, and update the state of the finite state machine; If it is judged that the model-predicted token does not meet the expectations of the current state, output the expected token or output the model alternative token that meets the expectations, and update the state of the finite state machine; Decode the output token into a text character and add it to the output sequence.
3. The online marking method based on artificial intelligence according to claim 2, characterized in that The model alternative token that meets the expectations is the token selected from the logits tensor that meets the expectations of the current state.
4. The online marking method based on artificial intelligence according to claim 1, wherein The specific steps of using the large language model on the vLLM inference framework for scoring the grading tasks of the second type of subjective questions in the grading tasks include: Use the prompt template corresponding to the question type to generate a prompt word for judging whether each scoring point in the reference answer is included in the user's question answer; The vLLM inference framework receives the prompt word and inference parameters, and initializes an inference request; Based on the prompt word and inference parameters, the large language model deployed on the vLLM inference framework performs forward inference and outputs a logits tensor; Based on the logits tensor, use a preset finite state machine to control the sampling and decoding of tokens, and add the decoding result to the output sequence; Append the newly generated token ID to the input token IDs sequence and perform the next round of forward inference until the termination condition is met.
5. The online marking method based on artificial intelligence according to claim 1, characterized in that, The prompt template configuration for different question types has corresponding response requirements and response examples for the question types.
6. The online marking method based on artificial intelligence according to claim 1, characterized in that For the question reading task of the first type of subjective questions in the marking task, the specific steps assigned to the vectorization-assisted scoring model for scoring include: Using the vectorization-assisted scoring model to convert the user's question answer and the reference answer into high-dimensional vectors respectively; Calculating the cosine similarity between the high-dimensional vectors corresponding to the user's question answer and the reference answer to obtain a cosine similarity score; Generating a score based on the cosine similarity score and the total score of the question reading task, and updating the question score in the user answer record table.
7. The online marking method based on artificial intelligence according to claim 1, characterized in that, The method further includes: In response to the timeout of the question reading task for the second type of subjective questions without processing, assigning the question reading task to the vectorization-assisted scoring model for scoring.
8. The online marking method based on artificial intelligence according to claim 1, characterized in that, The method further includes: Determining the objective questions and the third type of subjective questions in the online test paper, where the objective questions include single-choice questions and multiple-choice questions, and the third type of subjective questions include work appreciation questions, writing questions, and calculation questions; For each objective question, using the string comparison method to compare the user's question answer and the reference answer, generating a score and updating the question score in the user answer record table; for each third type of subjective question, conducting manual marking, generating a score and updating the question score in the user answer record table.
9. An electronic device, characterized in that, Including: A processor; A memory storing one or more computer instructions running on the processor; When the processor runs the computer instructions, it executes the steps of the artificial intelligence-based online marking method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, Storing computer instructions thereon, and when the computer instructions are run by the processor, it executes the steps of the artificial intelligence-based online marking method according to any one of claims 1-8.