Open homework correction method and system supporting autonomous uploading of teacher
By supporting teachers to independently upload open-ended assignment grading methods, intelligent grading of assignments in various formats and types has been achieved, solving the problem of narrow coverage in existing systems and improving the universality and adaptability of the grading system.
Patent Information
- Application Number
- CN202511331539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing homework grading systems can only grade specific types of homework, resulting in a narrow scope and an inability to handle homework of various formats and types.
It supports teachers to independently upload open-ended homework grading methods. By acquiring homework questions and student answers in various formats, it extracts key information to generate structured data, generates reference answers based on the structured data, and achieves unified grading through cross-modal comparison technology. It adopts differentiated processing strategies for different question types, including logical analysis and semantic similarity calculation.
It breaks through the limitations of assignment formats and can grade assignments in various formats such as documents, images, and videos, covering all types of assignments in various scenarios, including theoretical derivation, practical operation, and oral expression, significantly improving the universality of the grading system and its adaptability to teaching scenarios.
Smart Images

Figure CN120833242A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of educational informatization, and in particular to an open homework correction method and system supporting teacher self-uploading. BACKGROUND
[0002] Under the background of continuous development of educational informatization, homework correction systems are gradually widely used. Teachers can obtain part of the teaching materials through the network, and students can also submit homework online.
[0003] The homework correction system of the related art is composed of a homework management module, an answer preset module, a correction module and a result feedback module. The teacher selects homework from the resource library provided by the system in the homework management module and publishes it to the students, or uploads a small amount of homework in a specific format specified by the system. After the students complete the homework, they submit it online, and the system corrects it according to the preset answer through simple text comparison. If the answer is completely consistent with the preset, it is correct, and if it is not consistent, it is wrong. Finally, the correction result is fed back to the students.
[0004] For the related art in the above, the homework correction system can only correct specific types of homework, and the coverage is narrow. SUMMARY
[0005] In order to expand the coverage of intelligent homework correction, the present application provides an open homework correction method and system supporting teacher self-uploading.
[0006] In a first aspect, the present application provides an open homework correction method, which adopts the following technical solution: An open homework correction method supporting teacher self-uploading, comprising: In response to a question uploading operation of a teacher terminal, obtaining a homework question, the homework question adopting at least one of a document format, an image format and a video format; Extracting key information from the homework question to generate structured data; Generating a reference answer based on the structured data; In response to a homework uploading operation of a student terminal, obtaining a student answer; Comparing the reference answer and the student answer to generate an answer score of the student answer; Pushing the answer score to the teacher terminal and the student terminal.
[0007] By adopting the technical scheme, the subject uploading in multiple formats such as documents, images and videos is supported, and the limitation of the traditional correction system that only supports a single text format is solved. By converting the unstructured subject into structured data, the system can automatically generate a reference answer adapted to multimedia content. When the students submit equally diversified answers, the system realizes unified correction through cross-modal comparison technology and pushes the scores to the terminal of teachers and students in real time. The design breaks through the form limitation of homework, covers all-scene homework types such as theoretical derivation, practical operation and oral expression, and significantly improves the universality and teaching scene adaptability of the correction system.
[0008] Optionally, a subject type of the homework subject is acquired. In a case where the subject type is an objective question, a number of subjects in the student answer consistent with the reference answer is counted. The answer score is generated based on the number of subjects. In a case where the subject type is a subjective question, the student answer is logically analyzed to obtain a logical score. A semantic similarity between the reference answer and the student answer is calculated to obtain a semantic score. The logical score and the semantic score are weighted to obtain the answer score.
[0009] By adopting the technical scheme, for the differentiated correction needs of objective questions and subjective questions, the scheme realizes branch processing through subject type identification: the accurate matching strategy is adopted for objective questions to ensure efficiency, and the logical analysis and semantic similarity calculation are innovatively integrated for subjective questions. The logical analysis can capture the reasonableness of the answer, and the semantic similarity is compatible with the oral and diversified expression modes. The weighting mechanism of the two allows teachers to adjust the scoring emphasis according to the characteristics of the subject, so that the system can not only efficiently process standardized homework such as multiple-choice questions, but also accurately evaluate the answers to subjective questions with strong openness, realizing intelligent correction covering all types of questions.
[0010] Optionally, in a case where the student answer is in a video format, an examination target is determined according to the homework subject. A reference answer feature is obtained by performing a feature extraction operation on the reference answer. Multimedia data corresponding to the examination target is extracted from the student answer, and the multimedia data includes at least one of image data and audio data. Multimedia features are extracted from the multimedia data. The multimedia features are processed by dimensionality increasing to obtain high-dimensional multimedia features according to the examination target. The high-dimensional multimedia features are processed by dimensionality reduction to obtain multimedia representative features. A similarity between the reference answer feature and the multimedia representative feature is calculated. mapping processing is performed on the similarity to obtain the answer score.
[0011] By adopting the technical scheme, for the video homework correction problem, firstly, the semantic features of the reference answer are extracted based on the assessment target, and the multimedia data related to the target is positioned from the student video. The key information expression is strengthened by feature dimensioning, and the redundant data is removed by dimension reduction, and finally the representative features are compared with the reference answer. This technology breaks through the bottleneck of quantifying video answers, enabling the system to correct dynamic homework such as physics experiment demonstration and language oral test, significantly expanding the coverage dimension of intelligent correction.
[0012] Optionally, in the case of an assessment target related to audio, the audio data is subjected to semantic recognition processing to obtain audio semantics; The semantic recognition processing is performed on the homework question to obtain question semantics; The target audio semantics corresponding to the question semantics is searched in the audio semantics; The target audio data corresponding to the target audio semantics is determined; The audio features of the target audio data are extracted; The target audio semantics and the audio features are spliced to obtain the multimedia features.
[0013] By adopting the technical scheme, the core semantics of the question is identified first, and then the target semantic fragment associated with it is accurately positioned from the student audio, excluding irrelevant voice interference. By splicing the target semantic content and its acoustic features, a multimedia feature with content and expression is formed. This method ensures that the system focuses on the effective answering content in a complex audio environment, and is suitable for the correction scene of language oral homework, music performance and other audio answers.
[0014] Optionally, student information corresponding to the student terminal is extracted; The geographical information and the historical pronunciation information are extracted from the student information; In the first parameter mapping table, the first model parameter corresponding to the geographical information is searched; According to the historical pronunciation information, the second model parameter is obtained; According to the first model parameter and the second model parameter, the model parameter in the audio feature extraction model is adjusted; The audio feature extraction model is called to perform feature extraction operation on the target audio data to obtain the audio features.
[0015] By adopting the technical scheme, the pronunciation model is matched according to the student regional information, and the feature extraction parameter is dynamically optimized in combination with historical pronunciation data. The adaptive model reduces the misjudgment caused by individual pronunciation differences, so that the system can fairly evaluate oral work of students with different backgrounds, and improve the correction accuracy in dialect areas and language learning initial stage.
[0016] Optionally, in the case that the examination target is related to the video image, the video data is subjected to content extraction processing to obtain video content; The question is subjected to semantic recognition processing to obtain question semantics; The target video content corresponding to the question semantics is searched in the video content; The video frame corresponding to the target video content is determined; The key region image is extracted from the video frame; The image feature of the key region image is extracted; The target video content and the image feature are spliced to obtain the multimedia feature.
[0017] By adopting the technical scheme, the target content in the video is located based on the question semantics, the key video frame is further screened, and the related region image is extracted. The multimedia feature with strong correlation is constructed by fusing the target content description and the image feature. This design avoids invalid analysis of the whole video, so that the system can accurately correct the work that needs visual verification, and the video correction efficiency is improved.
[0018] Optionally, student information corresponding to the student terminal is extracted; Body information is extracted from the student information; According to the body information, the human body parts in the key region image are labeled to obtain a labeled image; According to the labeled image, a behavior mode is determined; According to the behavior mode, the image feature is generated.
[0019] By adopting the technical scheme, the key region is dynamically calibrated according to the student body information, and then the behavior mode is identified by labeling the image. The behavior mode is converted into quantifiable image features, so that the system can evaluate the motion standardization of dance, sports and other motion type work, break through the dependence of traditional correction on static content, and realize intelligent coverage of dynamic behavior work.
[0020] In a second aspect, the application provides an open homework correction system supporting teacher self-uploading, which adopts the following technical scheme: An open homework correction system supporting teacher self-uploading, comprising: An acquisition module for acquiring homework questions and student answers; a memory for storing a program of the open homework grading method supporting the teacher to upload autonomously; a processor, the program in the memory can be loaded and executed by the processor to implement the open homework grading method supporting the teacher to upload autonomously.
[0021] By adopting the technical solution, the question uploading in various formats of documents, images and videos is supported, and the limitation of the traditional grading system that only supports a single text format is solved. By converting the unstructured questions into structured data, the system can automatically generate the reference answers adapted to the multimedia content. When the students submit the same diversified answers, the system realizes unified grading through cross-modal comparison technology and pushes the scores to the teacher and student terminals in real time. The design breaks through the limitation of homework forms, covers all-scene homework types such as theoretical derivation, practical operation and oral expression, and significantly improves the universality and teaching scene adaptability of the grading system.
[0022] In a third aspect, the application provides an intelligent terminal, which adopts the technical solution as follows: An intelligent terminal includes a memory and a processor, and the memory stores a computer program capable of being loaded and executed by the processor to implement the method in any one of the above.
[0023] In a fourth aspect, the application provides a computer storage medium capable of storing a corresponding program, having the characteristics of facilitating the expansion of the coverage of homework intelligent grading, and adopting the technical solution as follows: A computer readable storage medium stores a computer program capable of being loaded and executed by the processor to implement any one of the open homework grading methods.
[0024] In summary, the application includes at least one of the following beneficial technical effects: The question uploading in various formats of documents, images and videos is supported, and the limitation of the traditional grading system that only supports a single text format is solved. By converting the unstructured questions into structured data, the system can automatically generate the reference answers adapted to the multimedia content. When the students submit the same diversified answers, the system realizes unified grading through cross-modal comparison technology and pushes the scores to the teacher and student terminals in real time. The design breaks through the limitation of homework forms, covers all-scene homework types such as theoretical derivation, practical operation and oral expression, and significantly improves the universality and teaching scene adaptability of the grading system. In response to the differentiated grading needs of objective and subjective questions, this solution implements split-path processing through question type identification: a precise matching strategy is used to ensure efficiency for objective questions, and a creative integration of logical analysis and semantic similarity calculation is used for subjective questions. Logical analysis can capture the rationality of the answer's reasoning, while semantic similarity is compatible with colloquial and diverse expressions. The weighted mechanism of the two allows teachers to adjust the emphasis of grading according to the characteristics of the subject, so that the system can not only efficiently handle standardized assignments such as multiple-choice questions, but also accurately evaluate the answers to open-ended subjective questions, achieving intelligent grading covering all question types; To address the challenge of grading video-based assignments, the system first extracts semantic features from reference answers based on assessment objectives, while simultaneously locating multimedia data relevant to these objectives from student videos. By increasing the dimensionality of these features, key information is enhanced, and then dimensionality reduction is used to eliminate redundant data. Ultimately, representative features are generated and compared with the reference answers. This technology overcomes the bottleneck of quantifying video answers, enabling the system to grade dynamic assignments such as physics experiment demonstrations and oral language exams, significantly expanding the scope of intelligent grading. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of an open homework grading method that supports teachers to upload homework independently, provided in an embodiment of the present application.
[0026] Figure 2 This is a flowchart of a method for generating an answer score provided in an embodiment of the present application.
[0027] Figure 3 This is a flowchart of a method for obtaining an answer score provided in an embodiment of the present application.
[0028] Figure 4 This is a flowchart of a multimedia feature extraction method 1 provided in an embodiment of the present application.
[0029] Figure 5 This is a flow chart of a second method for extracting multimedia features provided in an embodiment of the present application.
[0030] Figure 6 This is a flow chart of a third method for extracting multimedia features provided in an embodiment of the present application.
[0031] Figure 7 This is a flow chart of a fourth method for extracting multimedia features provided in an embodiment of the present application.
[0032] Figure 8 This is a schematic diagram of an open homework grading system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of this application more clear, the followingFigure 1 To the attached Figure 8 It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0034] The embodiment of the present application discloses an open homework correction method that supports teachers to upload homework independently. Figure 1 , the method comprising: Step S101: In response to the topic upload operation of the teacher terminal, the homework topic is obtained, and the homework topic is in at least one of a document format, an image format, and a video format.
[0035] A teacher terminal is a terminal held by a teaching staff member. The teacher terminal includes but is not limited to at least one of a mobile phone, a computer, and a smartwatch.
[0036] The topic upload operation is used to upload the homework topic to the teacher terminal. For example, the teacher inputs the homework topic in text form into the teacher terminal.
[0037] The homework title can also be composed of files in multiple different formats. For example, the homework title includes a video file and a document file, where the video file is a dance and the document file is the text "Please learn the dance in the video."
[0038] Optionally, a format parsing algorithm is used to determine the format of the homework title.
[0039] Step S102: Extract key information from the assignment title and generate structured data.
[0040] Structured data is computer-readable data. Key information includes the title and question type, which can be at least one of multiple-choice, short-answer, and essay questions.
[0041] When the assignment question is in document format, key information can be obtained through natural language processing technology (for example, semantic understanding, text summarization, etc.), and then the key information can be converted into computer-recognizable data to obtain structured data.
[0042] When the assignment questions are in image format, you can first use optical character recognition (OCR) technology to extract the text information in the image, and then use natural language processing technology to obtain key information.
[0043] When the homework question is in the form of a video, the video can be regarded as a video frame image and audio, and the processing of the video frame image is consistent with the processing of the homework question in the form of an image. Therefore, the processing of the video frame image will not be described here again. In addition, key information of the video can be extracted from the video frame image. For the processing of the audio, the audio-to-text technology can be used to convert the audio into text, and then the natural language processing can be used to extract the key information of the audio from the text. The video key information and the audio key information are converted into computer-recognizable data to obtain the structured data.
[0044] Step S103: generating a reference answer based on the structured data.
[0045] In some embodiments, the structured data is input into a generative pre-training model (for example, a pre-training language model such as GPT) to generate the reference answer. The generative pre-training model can generate the reference answer according to the homework question and the knowledge reserve. The knowledge reserve is related to the grade of the student.
[0046] In other embodiments, the reference answer can be manually uploaded by teaching staff.
[0047] Step S104: obtaining a student answer in response to a homework uploading operation of a student terminal.
[0048] The student terminal is a terminal held by a student. The student terminal includes at least one of a mobile phone, a computer, and a smart watch.
[0049] The student answer is in at least one of a document format, an image format, and a video format.
[0050] For example, a submit button, a homework question, and an answering area corresponding to the homework question are displayed on the student terminal, and the student can answer in the answering area or directly upload the student answer. After the student completes the answer, the student can trigger the submit button to complete the homework uploading operation corresponding to the student answer. At this time, the student terminal uploads the student answer to the server, and the server obtains the student answer.
[0051] Step S105: comparing the reference answer and the student answer to generate an answer score of the student answer.
[0052] The answer score is a quantitative index for measuring the matching degree between the student answer and the reference answer.
[0053] In some embodiments, the method for generating the answer score includes steps S1051 to S1056, and the specific contents are as follows: Step S1051: obtaining a question type of the homework question.
[0054] The question type includes objective questions and subjective questions. The objective questions refer to the selection questions, and the subjective questions include short answer questions and discussion questions. The question type can be uploaded at the same time as the homework questions.
[0055] Step S1052: In the case of objective questions, the number of questions consistent with the reference answer in the student answer is counted.
[0056] Exemplarily, the student answer is compared with the reference answer one by one, the same student answer and reference answer are taken out, the number of the taken-out student answers is counted, and the number of questions is obtained.
[0057] Step S1053: The answer score is generated based on the number of questions.
[0058] The number of questions corresponds to the number of correct answers in the student answer. Therefore, according to the number of questions and the question score, the answer score is obtained. For example, the homework questions are 50 selection questions, including 20 single-choice questions of 1 point and 30 multiple-choice questions of 3 points. The number of questions is counted as 40, of which the single-choice questions correspond to 15 and the multiple-choice questions correspond to 25, and the answer score is 15x1+25x3=90.
[0059] Step S1054: In the case of subjective questions, the student answer is logically analyzed to obtain a logical score.
[0060] The logical score is a quantitative score of the logical relationship of the student answer.
[0061] The evaluation dimensions of the logical score include self-consistency, factual truth, context consistency, and rationality. Exemplarily, the factual truth of the student answer is verified using a knowledge graph to obtain a logical score. Exemplarily, the context consistency of the student answer is verified using a large model (e.g., a GPT model) to obtain a logical score.
[0062] Step S1055: The semantic similarity between the reference answer and the student answer is calculated to obtain a semantic score.
[0063] Exemplarily, a semantic extraction model is called to perform a semantic extraction operation on the reference answer to obtain reference answer semantic features. The semantic extraction model is called to perform a semantic extraction operation on the student answer to obtain student answer semantic features. The similarity between the reference answer semantic features and the student answer semantic features is calculated. The similarity is normalized to obtain a semantic score.
[0064] Step S1056: The logical score and the semantic score are weighted to obtain an answer score.
[0065] The weight values used by the logical score and the semantic score can be preset, for example, the weight value of the logical score is 0.35, and the weight value of the semantic score is 0.65.
[0066] Step S106: Push the answer score to the teacher terminal and the student terminal.
[0067] In some other embodiments, a homework comment can also be generated according to the answer score, and the homework comment is pushed to the student terminal.
[0068] Optionally, the answer scores of the student terminals in the student terminal set are obtained to obtain a student score set, the student terminals in the student terminal set are terminals held by students in the same class. A statistical value of the student scores in the student score set is calculated, the statistical value includes at least one of an average value, a maximum value, a minimum value, and a median value.
[0069] Further, the answering situations of the student terminals in the student terminal set are obtained, the answering situations include correct questions and incorrect questions in the student answers. The incorrect questions of the student terminals in the student terminal set are taken out to obtain an incorrect question set, and the number of incorrect questions in the incorrect question set is calculated to obtain an incorrect question number. If the incorrect question number is greater than a preset number threshold, the corresponding target incorrect question is sent to the teacher terminal, so that the teaching staff can master the weak links of the students in the class in knowledge mastery. For example, if it is found that multiple students frequently make mistakes on the questions of a certain mathematical knowledge point, the teacher will be prompted that the knowledge point needs to be highlighted, and relevant teaching suggestions and supplementary learning materials are provided.
[0070] By adopting the above technical solution, the uploading of questions in multiple formats such as documents, images, and videos is supported, and the limitation of the traditional correction system that only supports a single text format is solved. By converting unstructured questions into structured data, the system can automatically generate reference answers that adapt to multimedia content. When students submit equally diversified answers, the system realizes unified correction through cross-modal comparison technology and pushes the scores to the teacher and student terminals in real time. This design breaks through the form limitation of homework, covers all-scene homework types such as theoretical derivation, practical operation, and oral expression, and significantly improves the universality and teaching scene adaptability of the correction system.
[0071] In the following embodiments, when the homework arranged by the teacher is a video homework, the student answers need to be answered in a video format, at this time, effective information needs to be extracted from the video to score the student answers. Therefore, the embodiment of the application discloses an answer score acquisition method. Referring to Figure 3 , the method comprises: Step S301: In the case that the student answer is in a video format, according to the homework question, the examination target is determined.
[0072] The assessment target is determined according to the homework question to determine the scoring dimension. For example, the keywords in the homework question are extracted to obtain the assessment target, for example, the assessment target is pronunciation standard degree, student dance action accuracy, experimental step integrity, etc.
[0073] Step S302: performing feature extraction on the reference answer to obtain reference answer features.
[0074] In combination with the assessment target, the feature extraction operation is performed on the reference answer to obtain the reference answer features.
[0075] For example, when the assessment target is pronunciation standard degree, if the reference answer is a standard audio of pronunciation, the audio features in the reference answer are extracted to obtain the reference answer features, including at least one of the time domain features, frequency domain features, and timbre features of the audio.
[0076] For example, when the assessment target is student dance action accuracy, if the reference answer is a reference video, the action change of the character in the reference answer is extracted. The action change is mapped to a vector to obtain the reference answer features. For example, the character raises the left leg and makes a step, which can be mapped to a vector [4, 7, 85, 56], where "4" represents the left leg, "7" represents the lifting action, "85" represents the action angle, and "56" represents the action amplitude.
[0077] Step S303: extracting multimedia data corresponding to the assessment target from the student answer, the multimedia data including at least one of image data and audio data.
[0078] In the case where the student answer is in video format, not all data can be used for scoring. For example, when the assessment target is pronunciation standard degree, the image data in the video has little effect on correcting the student answer, and can not be extracted.
[0079] Step S304: extracting multimedia features from the multimedia data.
[0080] The multimedia features include image features and audio features.
[0081] Step S305: performing dimensionality elevation processing on the multimedia features according to the assessment target to obtain high-dimensional multimedia features.
[0082] The dimensionality elevation processing is used to extract high-dimensional features in the multimedia features.
[0083] For example, according to the assessment target, a feature channel is set. The dimensionality elevation processing is performed on the multimedia features to output the features of the corresponding feature channel to obtain the high-dimensional multimedia features. For example, when the assessment target is pronunciation standard degree, the feature channel is a channel related to audio, frequency spectrum, etc.
[0084] Step S306: Dimensionality reduction processing is performed on the high-dimensional multimedia feature to obtain a multimedia representative feature.
[0085] The dimensionality reduction processing is used to reduce the correlation in the high-dimensional multimedia feature and prevent overfitting.
[0086] Step S307: Similarity between the reference answer feature and the multimedia representative feature is calculated.
[0087] For example, the Euclidean distance between the reference answer feature and the multimedia representative feature in the vector space is calculated to obtain the similarity.
[0088] Step S308: Mapping processing is performed on the similarity to obtain an answer score.
[0089] Optionally, the sigmoid function is used to perform mapping processing on the similarity to obtain the answer score.
[0090] By using the above technical solution, for the video homework grading problem, the semantic feature of the reference answer is first extracted based on the assessment target, and the multimedia data related to the target is located from the student video. The key information expression is strengthened through feature dimensionality reduction, and the representative feature is finally generated by comparing with the reference answer after removing redundant data through dimensionality reduction. This technology breaks through the bottleneck of the difficulty of quantifying video answers, so that the system can grade dynamic homework such as physics experiment demonstration and language oral test, and significantly expands the coverage dimension of intelligent grading.
[0091] In the following embodiments, when the assessment target is related to audio, corresponding processing needs to be performed on the audio data to obtain the multimedia feature. Therefore, the present application discloses a multimedia feature extraction method I. Referring to Figure 4 , the method comprises the following steps. Step S401: When the assessment target is related to audio, semantic recognition processing is performed on the audio data to obtain audio semantics.
[0092] For example, when the assessment target is pronunciation standard, student's textbook recitation, or instrument performance, it is considered that the assessment target is related to audio.
[0093] Optionally, the audio-to-text technology can be used to convert the audio into text, and then the converted text is subjected to semantic recognition processing to obtain the audio semantics.
[0094] Step S402: Semantic recognition processing is performed on the homework question to obtain question semantics.
[0095] Optionally, the natural language processing technology is used to perform semantic recognition processing on the homework question to obtain the question semantics.
[0096] Step S403: The target audio semantics corresponding to the question semantics is searched in the audio semantics.
[0097] Exemplarily, when the examination target is pronunciation standard, the question semantics is to judge whether the pronunciation of the words is accurate. Then, the words in the question semantics are taken as the retrieval keywords, the audio semantics containing the retrieval keywords or related to the retrieval keywords are retrieved, and the target audio semantics is obtained.
[0098] Exemplarily, when the examination target is the student's reading of the textbook, the question semantics is to judge whether the student reads the textbook completely. Then, the content of the textbook is taken as the retrieval keyword, the audio semantics containing the retrieval keyword is retrieved, and the target audio semantics is obtained.
[0099] Step S404: determining target audio data corresponding to the target audio semantics.
[0100] Optionally, the timestamp of the target audio semantics in the audio is located. According to the audio segment corresponding to the timestamp, the target audio data is obtained. In an actual scenario, there is usually some useless audio in the video uploaded by the student, so the audio data needs to be screened to extract the useful target audio data.
[0101] Step S405: extracting the audio features of the target audio data.
[0102] The audio features of the target audio data include, but are not limited to, at least one of the time domain features, frequency domain features, and timbre features of the audio.
[0103] Step S406: splicing the target audio semantics and the audio features to obtain multimedia features.
[0104] By adopting the above technical solution, the core semantics of the question is first recognized, and then the target semantic segment associated with it is accurately located from the student audio, and irrelevant voice interference is excluded. By splicing the target semantic content and its acoustic features, multimedia features with both content and expression are formed. This method ensures that the system focuses on the effective answering content in a complex audio environment, and is suitable for the correction scene of audio answers such as language oral test and music performance.
[0105] Embodiments of the present application disclose a second method for extracting multimedia features. Referring to Figure 5 The method comprises the following steps: Step S501: extracting student information corresponding to the student terminal.
[0106] The student information includes, but is not limited to, at least one of the name, age, grade, gender, geographical information, body information, and historical homework record of the student.
[0107] Step S502: extracting geographical information and historical pronunciation information from the student information.
[0108] The geographical information refers to the current residence, the native address or the long-term residence of the student. The long-term residence refers to the geographical position where the student has continuously lived for more than one year.
[0109] The historical pronunciation information is extracted from the pronunciation audio submitted by the student in the past. The historical pronunciation information is used to describe the pronunciation characteristics of the student.
[0110] In step S503, the first model parameter corresponding to the geographical information is searched in the first parameter mapping table.
[0111] People living in a certain area usually have specific rules in oral pronunciation, that is, common accent problems. The first parameter mapping table records the correspondence between the geographical information and the model parameters used by the audio feature extraction model. Different model parameters will be reflected in the audio features extracted by the audio feature extraction model, which will weaken the influence of the accent on the audio features. Therefore, the first parameter mapping table records the relationship between the geographical information and the pronunciation habits.
[0112] In step S504, the second model parameter is obtained according to the historical pronunciation information.
[0113] The second model parameter is the model parameter obtained by training the audio feature extraction model using the historical pronunciation information. The second model parameter records the pronunciation habits and pronunciation characteristics of each student.
[0114] In step S505, the model parameter in the audio feature extraction model is adjusted according to the first model parameter and the second model parameter.
[0115] Optionally, the audio feature extraction model adopts a CNN model. Further, the audio feature extraction model can also adopt an RNN model, a Transformer model or a hybrid model.
[0116] In step S506, the audio feature extraction model is called to perform a feature extraction operation on the target audio data to obtain the audio features.
[0117] By using the above technical solution, the pronunciation model is matched according to the geographical information of the student, and the feature extraction parameter is dynamically optimized in combination with the historical pronunciation data of the student. Through the adaptive model, the misjudgment caused by individual pronunciation differences is reduced, so that the system can fairly evaluate the oral work of students with different backgrounds, and improve the accuracy of correction in dialect areas and language learning initial stage.
[0118] In the following embodiments, when the examination target is related to a video image, the video data needs to be processed correspondingly to obtain multimedia features. Embodiments of the present application disclose a third method for extracting multimedia features. Referring to Figure 1 The method comprises the following steps. Step S601: In the case that the assessment target is related to the video image, content extraction processing is performed on the video data to obtain video content.
[0119] For example, when the assessment target is the accuracy of a student's dance action or a sports action assessment, it is considered that the assessment target is related to the video image.
[0120] The video content is used to describe the semantic information of different video frames in the video.
[0121] Step S602: Perform semantic recognition processing on the homework question to obtain question semantics.
[0122] Optionally, the natural language processing technology is used to perform semantic recognition processing on the homework question to obtain question semantics.
[0123] Step S603: Search for target video content corresponding to the question semantics in the video content.
[0124] For example, when the assessment target is the accuracy of a student's dance action, the question semantics is to determine whether the student's dance action is accurate. Then, the dance action is used as a search keyword to search for video content containing the search keyword in the video content to obtain the target video content.
[0125] For example, when the assessment target is a sports action assessment, the question semantics is to determine whether the student has completed the sports action. Then, the sports action is used as a search keyword to search for audio semantics containing the search keyword in the video content to obtain the target video content.
[0126] Step S604: Determine the video frame corresponding to the target video content.
[0127] Optionally, the timestamp of the target video content in the video is located. The video frame is obtained according to the timestamp.
[0128] Step S605: Extract a key region image from the video frame.
[0129] The key region image refers to the region image in the video frame related to the assessment target.
[0130] For example, the region where the human body exists in the video frame is framed to obtain the key region image.
[0131] Step S606: Extract image features of the key region image.
[0132] The image features include, but are not limited to, at least one of color features, texture features, and shape features.
[0133] Step S607: Concatenate the target video content and the image features to obtain multimedia features.
[0134] By adopting the technical scheme, the target content in the video is positioned based on the question semantics, and the key video frame is further screened and the relevant region image is extracted. By fusing the target content description and the image feature, the strong-association multimedia feature is constructed. The design avoids invalid analysis on the whole video, so that the system can accurately correct the homework requiring visual verification, and the video correction efficiency is improved.
[0135] The embodiment of the application discloses a multimedia feature extraction method four. Referring to Figure 1 The method comprises the following steps. Step S701: Extracting student information corresponding to a student terminal.
[0136] The student information includes at least one of the name, age, grade, gender, geographical information, body information, and historical homework record of the student.
[0137] Step S702: Extracting body information from the student information.
[0138] The body information is used to describe at least one of the height, girth information, width information, longitudinal ratio, and transverse ratio of the student. The girth information includes at least one of the bust, waist, and hip. The width information includes at least one of the shoulder width, pelvic width, and chest width.
[0139] Step S703: Labeling the human body parts in the key region image according to the body information to obtain a labeled image.
[0140] For example, the human body parts in the key region image are identified. The human body parts are labeled in combination with the body information to obtain a labeled image.
[0141] Step S704: Determining a behavior mode according to the labeled image.
[0142] The behavior mode is used to describe the action or behavior of the human body. For example, the behavior mode describes the behavior of the student lifting the left foot.
[0143] For example, the spatial change of the labeled points in the labeled image is determined. According to the human body parts corresponding to the labeled points and the spatial change, the behavior mode is obtained.
[0144] Step S705: Generating an image feature according to the behavior mode.
[0145] For example, the behavior mode is vectorized to obtain a behavior feature vector. The behavior feature vector is encapsulated into an image feature format to obtain the image feature.
[0146] By adopting the technical scheme, the key region is dynamically calibrated according to the student body information, and then the image is labeled to identify the behavior mode. The behavior mode is converted into quantifiable image features, so that the system can evaluate the motion standardization of dance, sports and other motion type homework, break through the dependence of traditional correction on static content, and realize intelligent coverage of dynamic behavior homework.
[0147] Based on the same inventive concept, the embodiment of the present application provides an open homework correction system supporting teacher autonomous uploading, comprising: The acquisition module 801 is configured to acquire homework questions and student answers.
[0148] The memory 802 is configured to store the program of the open homework correction method supporting teacher autonomous uploading.
[0149] The processor 803 is configured to load and execute the program in the memory, and implement the open homework correction method supporting teacher autonomous uploading.
[0150] By adopting the technical scheme, the uploading of questions in multiple formats such as documents, images and videos is supported, and the limitation of the traditional correction system that only supports single text format is solved. By converting unstructured questions into structured data, the system can automatically generate reference answers adapted to multimedia content. When students submit equally diversified answers, the system realizes unified correction through cross-modal comparison technology, and pushes the scores to the teacher and student terminals in real time. This design breaks through the form limitation of homework, covers all-scene homework types such as theoretical derivation, practical operation and oral expression, and significantly improves the universality and teaching scene adaptability of the correction system.
[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0152] The embodiment of the present application provides a computer readable storage medium, which stores a computer program capable of being loaded and executed by a processor to support a teacher autonomous uploading open homework correction method.
[0153] The computer storage medium includes, for example, a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0154] Based on the same inventive concept, the embodiment of the present application provides a kind of intelligent terminal, including memory and processor, the computer program of open homework correction method capable of being loaded and executed by processor to support teacher autonomous upload is stored on memory.
[0155] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0156] The above are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, any feature disclosed in the specification (including the abstract and drawings) can be replaced by other equivalent or similar purpose alternative features, unless specifically described. That is, each feature is only an example of a series of equivalent or similar features.
Claims
1. An open homework grading method supporting teacher self-uploading, characterized by, The method comprises: in response to a question uploading operation of a teacher terminal, obtaining a homework question, the homework question being in at least one of a document format, an image format, and a video format; extracting key information from the homework question to generate structured data; generating a reference answer based on the structured data; in response to a homework uploading operation of a student terminal, obtaining a student answer; comparing the reference answer and the student answer to generate an answer score of the student answer; pushing the answer score to the teacher terminal and the student terminal.
2. The open homework grading method of claim 1, wherein, The comparing the reference answer and the student answer to generate an answer score of the student answer comprises: obtaining a question type of the homework question; in a case where the question type is an objective question, counting a number of questions in the student answer that are consistent with the reference answer; generating the answer score based on the number of questions; in a case where the question type is a subjective question, performing logical analysis on the student answer to obtain a logical score; calculating semantic similarity between the reference answer and the student answer to obtain a semantic score; weighting the logical score and the semantic score to obtain the answer score.
3. The open homework grading method of claim 2, wherein, The method further comprises: in a case where the student answer is in a video format, determining an examination target according to the homework question; performing feature extraction on the reference answer to obtain reference answer features; extracting multimedia data corresponding to the examination target from the student answer, the multimedia data comprising at least one of image data and audio data; extracting multimedia features from the multimedia data; performing dimensionality increasing processing on the multimedia features according to the examination target to obtain high-dimensional multimedia features; performing dimensionality reducing processing on the high-dimensional multimedia features to obtain multimedia representative features; calculating similarity between the reference answer features and the multimedia representative features; performing mapping processing on the similarity to obtain the answer score.
4. The open homework grading method of claim 3, wherein the teacher uploads the answer key and the grading rule to the server. The extracting multimedia features from the multimedia data comprises: in a case where the examination target is related to audio, performing semantic recognition processing on the audio data to obtain audio semantics; performing semantic recognition processing on the homework question to obtain question semantics; searching for target audio semantics corresponding to the question semantics in the audio semantics; determining target audio data corresponding to the target audio semantics; extracting audio features of the target audio data; concatenating the target audio semantics and the audio features to obtain the multimedia features.
5. The open homework grading method of claim 4, wherein the teacher is allowed to upload the answer key and the grading rule in the form of a spreadsheet file. The extracting audio features of the target audio data comprises: extracting student information corresponding to the student terminal; extracting regional information and historical pronunciation information from the student information; searching for first model parameters corresponding to the regional information in a first parameter mapping table; obtaining second model parameters according to the historical pronunciation information; adjusting model parameters in an audio feature extraction model according to the first model parameters and the second model parameters; calling the audio feature extraction model to perform feature extraction on the target audio data to obtain the audio features.
6. The open homework grading method of claim 3, wherein the teacher uploads the answer key and the grading rule to the server. The extracting multimedia features from the multimedia data comprises: In the case that the examination target is related to a video image, content extraction processing is performed on the video data to obtain video content; Semantic recognition processing is performed on the homework question to obtain question semantics; Target video content corresponding to the question semantics is searched in the video content; A video frame corresponding to the target video content is determined; A key region image is extracted from the video frame; Image features of the key region image are extracted; The target video content and the image features are spliced to obtain the multimedia features.
7. The open homework grading method of claim 6, wherein the teacher is allowed to upload the answer key and the grading rule in the server. The extraction of the image features of the key region image includes: Student information corresponding to the student terminal is extracted; Body information is extracted from the student information; According to the body information, a human body part in the key region image is labeled to obtain a labeled image; A behavior mode is determined according to the labeled image; The image features are generated according to the behavior mode.
8. An open homework grading system supporting teacher self-uploading, characterized by, The system is used to perform the open homework grading method as claimed in any one of claims 1 to 7, comprising: An acquisition module is configured to acquire a homework question and a student answer; A storage is configured to store a program of the open homework grading method supporting the teacher to upload independently; A processor, the program in the storage can be loaded and executed by the processor, and the open homework grading method supporting the teacher to upload independently is implemented.
Citation Information
Patent Citations
Data alignment method and device
CN107766376A
Test question mark method and system, and computer readable storage medium
CN109064814A
Data correction method, device and system
CN111950240A
Homework correcting method and system based on image recognition and intelligent terminal
CN112115736A
Homework correction processing method and device, computer equipment and readable storage medium
CN115457585A