Job data acquisition method and device applied to education platform, and medium
By constructing a knowledge defect graph and employing a hierarchical data collection strategy, the problems of static and standardized knowledge graph collection in educational platforms were solved, enabling personalized homework data collection and analysis, and improving the reliability of teaching and the relevance of data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-10
AI Technical Summary
The existing knowledge graph construction of educational platforms lacks dynamic adaptability, resulting in delayed teaching interventions, and the lack of tiered data collection strategies fails to meet the needs of personalized teaching analysis.
By constructing a knowledge defect graph, hierarchical placement items and hierarchical correction items are generated. Knowledge defect graph strategy packages are collected and generated, question-level collection is performed, collection evidence data packages are generated, and chain association and supplementary collection verification are performed to generate a set of homework collection results.
It enables personalized collection and analysis of homework data, improves the reliability of teaching decisions and the relevance of data, and solves the problems of delayed intervention caused by the static nature of knowledge graphs and the lack of personalized data caused by standardized collection standards.
Smart Images

Figure CN121639422A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational big data collection technology, and in particular to a method, device and medium for collecting homework data for educational platforms. Background Technology
[0002] With the acceleration of digitalization in education, big data collection technology is increasingly being applied in educational platforms. The core objective is to acquire students' learning behavior data through intelligent means to optimize the allocation of teaching resources and improve educational quality. Currently, mainstream educational platforms generally employ technologies such as image recognition, natural language processing, and sensor fusion to collect and analyze homework data. For example, methods such as OCR-based question recognition, NLP-based text-based grading feedback, and recording students' problem-solving trajectories through behavioral sensors have gradually replaced traditional manual data entry methods, improving homework processing efficiency.
[0003] However, existing technologies still have two limitations: First, knowledge graph construction lacks dynamic adaptability. Current technologies mostly rely on pre-set knowledge nodes for static association, making it difficult to automatically update the graph structure based on real-time changes in students' frequently missed questions and their causes. This results in teaching intervention suggestions lagging behind students' actual needs. Second, a tiered data collection strategy is lacking. Existing platforms typically use a uniform collection standard, failing to design differentiated collection tasks based on the individual knowledge gaps of students. Consequently, the collected data cannot meet the needs of personalized teaching analysis in terms of granularity and depth. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a method for collecting homework data for educational platforms, which solves the problems of delayed intervention caused by the static nature of knowledge graphs and the lack of personalized data caused by standardized collection standards in the prior art.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a method for collecting homework data applied to an education platform, comprising,
[0008] Based on the frequently missed questions and reasons for errors in students' historical homework data, a knowledge deficiency map is constructed, and tiered assignment items and tiered correction items are written into it to generate a knowledge deficiency map strategy package.
[0009] Based on the hierarchical arrangement items in the knowledge deficiency graph strategy package, select the hierarchical question set for this assignment, break it down into question-level collection tasks, establish a question-level collection index, and generate a collection arrangement instruction package;
[0010] Based on the collection and arrangement instruction package, the system collects image fragments and evidence fragment records of the questions, and performs peer review of assignments in back-to-back and one-to-many modes. At the same time, it collects the marking trajectory records and generates a data package of collected evidence.
[0011] Chain-link evidence fragment records in the collected evidence data package to obtain the question-level collected evidence chain record, and solidify it into a question-level data asset record to generate an asset evidence chain dataset;
[0012] By combining the asset evidence chain dataset with the set of task-level collection, we can identify collection gaps, convert them into supplementary collection instructions, and simultaneously perform supplementary collection and verification to generate a set of job collection results.
[0013] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the following steps are taken: Based on the frequently missed questions and frequently missed reasons in students' historical homework data, a knowledge deficiency map is constructed, and hierarchical assignment items and hierarchical correction items are written into it to generate a knowledge deficiency map strategy package.
[0014] The system iterates through students' historical homework data question by question, filters frequently missed questions and frequently missed reasons, establishes correlations, and generates a set of records mapping the correlation between questions and reasons for mistakes.
[0015] Based on the mapping record set of questions and error causes, high-frequency error cause entries are used as defect nodes, and high-frequency error questions are used as performance nodes.
[0016] By dividing defect nodes into superior and subordinate categories according to the hierarchical relationship of knowledge points, and attaching the manifestation nodes to the corresponding defect nodes, a knowledge defect graph is generated.
[0017] The defect nodes are divided into different learning level labels, the corresponding layered layout items and layered correction items are obtained, and written into the knowledge defect graph, which is then encapsulated into a knowledge defect graph strategy package.
[0018] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the step of selecting the hierarchical question set for the current homework based on the hierarchical arrangement items in the knowledge deficiency graph strategy package is as follows:
[0019] Based on the hierarchical arrangement of items in the knowledge defect graph strategy package, the hierarchical question selection boundaries and question filtering range boundaries are defined, and priority selection constraints are set to generate a hierarchical question selection condition set.
[0020] Based on the tiered question selection criteria set, tiered filtering, priority selection, question deduplication, and truncation are performed to generate a tiered question set.
[0021] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the specific steps for generating the collection and arrangement instruction package are as follows:
[0022] Expand the hierarchical question set according to the learning level label, and assign question-level collection tasks and question-level collection task identifiers to generate a question-level collection task set;
[0023] Create a question-level collection index for the question-level collection task set, and attach index sequence number and verification flag to generate collection and arrangement instruction package.
[0024] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the following steps are taken: Based on the collection and arrangement instruction package, problem image fragments and evidence fragment records are collected, and homework peer review is performed in back-to-back and one-to-many modes. Simultaneously, the grading and annotation trajectory records are collected to generate a collection evidence data package.
[0025] The question-level acquisition tasks are triggered sequentially according to the execution order in the acquisition and arrangement instruction package. The question resources are located and the question display area is captured to generate question image fragment records.
[0026] Based on the recorded image fragments of the questions, collect operation trajectory information and establish a connection with the corresponding image fragments of the questions to generate evidence fragment records;
[0027] Based on evidence fragment records, assignments are mutually reviewed in a back-to-back and one-to-many manner according to the hierarchical review items, and the entire assignment review process is tracked and recorded, and compiled into a review annotation trajectory record.
[0028] The question image fragment records, evidence fragment records, and correction annotation trajectory records are grouped together to generate a data package of collected evidence.
[0029] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the steps of chaining evidence fragment records in the collected evidence data package to obtain question-level collected evidence chain records and solidifying them into question-level data asset records to generate an asset evidence chain dataset are as follows.
[0030] The evidence fragment records in the collected evidence data package are sorted, missing fields are filled in and duplicates are removed to generate a question-level ordered sequence of evidence fragments;
[0031] Add evidence chain nodes to the ordered evidence fragment sequence at the question level, set pointers to adjacent evidence chain nodes, mark key anchor points, and generate a question-level collected evidence chain record;
[0032] The evidence chain records collected at the question level are solidified into question-level data asset records, and then grouped and verified for consistency to generate an asset evidence chain dataset.
[0033] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the steps of combining the asset evidence chain dataset with the question-level collection task set to locate collection gaps and convert them into supplementary collection instructions are as follows.
[0034] Align the asset evidence chain dataset with the question-level collection task set one by one, mark the missing evidence types and pointer entries, and generate a question-level collection gap list;
[0035] Based on the list of gaps in the problem-level data collection, supplementary data collection instructions are generated for the gaps. Supplementary data collection instructions under the same learning level are grouped into a level-level supplementary data collection instruction segment to generate a supplementary data collection instruction package.
[0036] As a preferred embodiment of the homework data collection method applied to an education platform according to the present invention, the steps for performing supplementary data collection and verification to generate a homework collection result set are as follows:
[0037] Based on the supplementary collection instruction package, supplementary collection of evidence is carried out for the topic, new evidence fragment records are added to the topic-level collection evidence chain record, and the supplementary collection source mark is registered to generate a supplementary collection append record set;
[0038] The supplementary collection of records will be merged and updated with the question-level data asset records, and a consistency check will be re-executed. Once the check passes, the results will be uniformly packaged into a job collection result set.
[0039] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the homework data collection method applied to an education platform as described in the first aspect of the present invention.
[0040] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the homework data collection method applied to an education platform as described in the first aspect of the present invention.
[0041] The beneficial effects of this invention are as follows: by generating a chain of related evidence, the anonymity of back-to-back mutual review and one-to-many allocation constraints are guaranteed, the fragmentation problem of existing mutual review data is solved, and the reliability of homework data grouping and teaching decisions is improved; by selecting hierarchical questions and generating arrangement instructions, differentiated collection of homework tasks and data granularity control are realized under the big data collection framework, which enhances the pertinence and effectiveness of educational data analysis. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a flowchart of a method for collecting homework data for use on an educational platform.
[0044] Figure 2 A flowchart for generating a knowledge defect graph strategy package.
[0045] Figure 3 A flowchart for generating a data package for collecting evidence.
[0046] Figure 4 A flowchart for generating the job data collection result set. Detailed Implementation
[0047] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0048] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0049] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0050] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a method for collecting homework data applied to an education platform, comprising the following steps:
[0051] S1: Based on the high-frequency error questions and high-frequency error reasons in students' historical homework data, construct a knowledge deficiency map, and write tiered assignment items and tiered correction items to generate a knowledge deficiency map strategy package.
[0052] S1.1: Iterate through the students' historical homework data question by question, filter out frequently missed questions and frequently missed reasons, establish the relationship, and generate a set of question and reason association mapping records;
[0053] Specifically, during the traversal process, the question identifier and error cause entry identifier corresponding to each question are read synchronously, and the question identifier and error cause entry identifier are written into the traversal record content as association keys. At the same time, the number of error records with the same question identifier in the student's historical homework data is counted to obtain the number of question errors, and the number of annotations with the same error cause entry identifier in the student's historical homework data is counted to obtain the number of error causes. Based on the preset high-frequency error question judgment threshold and high-frequency error cause entry judgment threshold, the number of question error occurrences and error cause occurrences are judged. Question identifiers that meet the high-frequency error question judgment threshold are identified as high-frequency error questions, and error cause entry identifiers that meet the high-frequency error cause entry judgment threshold are identified as high-frequency error causes. The corresponding error cause entry identifier is located based on the question identifier and an association pair of question identifier and error cause entry identifier is formed. Only the association relationship that simultaneously meets the high-frequency error question judgment and high-frequency error cause entry judgment is retained. The error cause entry identifiers in the association relationship are sorted according to the number of error causes and truncated according to the number of retentions (for example, the number of retentions is 3), generating a question and error cause association mapping record set.
[0054] It should be noted that the threshold for judging high-frequency error questions is defined based on the proportion of the cumulative number of error records of the same question identifier across different batches of student assignments in the student's historical assignment data to the total number of times the question identifier was answered (example range is 30% to 60%); the threshold for judging high-frequency error cause items is defined based on the proportion of the cumulative number of times the same error cause item identifier was marked in the student's historical assignment data to the total number of times all error cause items were marked (example range is 20% to 50%).
[0055] S1.2: Based on the problem and error cause association mapping record set, high-frequency error cause entries are used as defect nodes, and high-frequency error problems are used as performance nodes;
[0056] Specifically, the process involves reading the question identifier and error cause entry identifier from the question and error cause association mapping record set one by one. The error cause entry identifiers are deduplicated, summarized, and registered as a high-frequency error cause entry identifier set. Simultaneously, question identifiers are deduplicated, summarized, and registered as a high-frequency error question identifier set. For each error cause entry identifier in the high-frequency error cause entry identifier set, a node type tag is concatenated with it, and a fixed-length encoding is performed to assign a defect node identifier and register it as a defect node. Similarly, for each question identifier in the high-frequency error question identifier set, a performance node identifier is assigned and registered as a performance node. Based on the association between the question identifier and error cause entry identifier in the question and error cause association mapping record set, a correspondence is established between defect nodes and performance nodes, forming a node structure where high-frequency error cause entries are used as defect nodes and high-frequency error questions are used as performance nodes.
[0057] S1.3: Based on the hierarchical relationship of knowledge points, the defect nodes are divided into superior and subordinate levels, and the manifestation nodes are attached to the corresponding defect nodes to generate a knowledge defect graph;
[0058] Specifically, the system reads the registered error cause entry identifiers and, in conjunction with the knowledge point identifiers corresponding to the question identifiers in the question and error cause association mapping record set, organizes the knowledge point hierarchy. For error cause entry identifiers at higher levels in the knowledge point hierarchy, the corresponding defect nodes are marked as superior defect nodes; for error cause entry identifiers at lower levels, the corresponding defect nodes are marked as subordinate defect nodes. A hierarchy association marker is written between the superior and subordinate defect nodes to form a defect node hierarchy structure. Based on the correspondence between question identifiers and error cause entry identifiers in the question and error cause association mapping record set, the system locates the matching defect node position in the defect node hierarchy structure and attaches the representation node to the corresponding defect node according to the association relationship, generating a knowledge defect graph.
[0059] It should be noted that the knowledge point identification information may include subject identification, grade identification, textbook chapter path, and knowledge point code.
[0060] The hierarchical relationship of knowledge points is a summary of the parent-child relationships of the same knowledge point identifier from the mapping records of questions and error causes (connected into a hierarchical chain according to the inclusion relationship of the parent to the child). Knowledge points that are closer to the subject outline or chapter directory at the parent level are considered to be at a higher level, while knowledge points that are closer to specific test points or subdivided skill points at the child level are considered to be at a lower level.
[0061] S1.4: Divide the defect nodes into different learning level labels, obtain the corresponding hierarchical arrangement items and hierarchical correction items, and write them into the knowledge defect graph, which is then encapsulated into a knowledge defect graph strategy package.
[0062] Specifically, based on the hierarchical depth information (hierarchical relationship of knowledge points and hierarchical structure of defect nodes) registered in the hierarchical relationship of knowledge points corresponding to defect nodes, learning level marking is performed on defect nodes. Defect nodes in different depth ranges are marked as different learning levels (for example, based on the hierarchical position of the knowledge point corresponding to the defect node in the hierarchical chain of knowledge points: those closer to the root node and at a shallower level are marked as basic level, those at an intermediate level are marked as intermediate level, and those closer to the leaf node and at a deeper level are marked as promotion level). The corresponding learning level mark is written into the defect node record. Based on the learning level mark of the defect node, the hierarchical arrangement item that matches the learning level mark is located from the existing hierarchical arrangement item set, and the hierarchical correction item that matches the learning level mark is located from the hierarchical correction item set. The hierarchical arrangement item and the hierarchical correction item are written below the corresponding defect node and bound to the defect node identifier. The knowledge defect graph containing the learning level mark, hierarchical arrangement item and hierarchical correction item is encapsulated as a whole to generate a knowledge defect graph strategy package.
[0063] S2: Based on the hierarchical arrangement items in the knowledge deficiency graph strategy package, select the hierarchical question set for this assignment, break it down into question-level collection tasks, establish a question-level collection index, and generate a collection arrangement instruction package;
[0064] S2.1: Based on the hierarchical arrangement of items in the knowledge defect graph strategy package, divide the hierarchical question selection boundaries and question filtering range boundaries, set priority selection constraints, and generate a hierarchical question selection condition set;
[0065] Specifically, the content of each layered arrangement item is expanded according to the learning level markers registered in the layered arrangement items. The learning level markers, defect nodes, and performance nodes corresponding to each layered arrangement item are written into the selection condition draft item. The layered question selection boundary is set according to the learning level markers in the selection condition draft item. The knowledge point hierarchy relationship items associated with the defect nodes corresponding to the learning level markers are read from the knowledge defect graph strategy package, and the lower and upper bounds of the hierarchy depth are extracted. The lower and upper bounds of the hierarchy depth are combined into the hierarchy coverage boundary. The learning level markers are mapped to the hierarchy coverage boundary of the layered question set and written into the layered question selection boundary. The question filtering range boundary is set according to the defect nodes and performance nodes in the selection condition draft item. The question filtering range boundary is limited to the coverage range of the performance nodes corresponding to the defect nodes and written into the question filtering range boundary. The layered question quantity threshold and defect node coverage rate threshold are judged. The boundary items that meet the threshold judgment are written into the priority selection constraint to generate the layered question selection condition set.
[0066] It should be noted that the threshold for the number of questions in a tiered arrangement is defined based on the question capacity and task size constraints corresponding to the same learning level tag in the tiered arrangement entries (example range: the number of questions corresponding to each learning level tag is limited to 10 to 30); the threshold for defect node coverage is defined based on the proportion of the number of defect nodes associated with the performance nodes within the question selection range to the total number of defect nodes required to be covered in the tiered arrangement entries (example range: 60% to 90%).
[0067] S2.2: Based on the hierarchical question selection condition set, perform hierarchical filtering, priority selection, question deduplication, and truncation to generate a hierarchical question set;
[0068] Specifically, tiered filtering is performed within the question selection range based on the learning level markers. Only question identifiers that meet the tiered question selection boundary requirements are retained. Based on the tiered question quantity threshold and defect node coverage threshold registered in the priority selection constraints, threshold judgment is performed on the retained question identifiers, and question identifiers that simultaneously meet the tiered question quantity threshold and defect node coverage threshold are prioritized for retention. At the same time, the question identifiers are sorted according to the priority rules implicit in the priority selection constraints (that is, sorting according to the defect node coverage rate and the number of errors in the question identifier, with question identifiers with higher defect node coverage and more errors appearing first and being retained first). Question deduplication is performed on the sorted question identifiers, and only one valid record is retained for duplicate question identifiers. The question identifiers are truncated and retained based on the tiered question quantity threshold (for example, the number of questions corresponding to each learning level marker is limited to 10 to 30), generating a tiered question set.
[0069] S2.3: Expand the hierarchical question set according to the learning level label, and assign question-level collection tasks and question-level collection task identifiers to generate a question-level collection task set;
[0070] Specifically, the question identifiers are expanded according to the learning level tags already registered in the hierarchical question set. Question identifiers under the same learning level tag are summarized, registered, and written into the corresponding learning level entry. Question-level collection tasks are assigned to each question identifier in each learning level entry. The question identifier and the learning level tag are written together into the content of the question-level collection task. A unique question-level collection task identifier is generated for each question-level collection task. The question-level collection task identifier is numbered and registered according to the combination rules of the learning level tag and the question identifier (that is, the learning level tag is used as a prefix field and is sequentially concatenated with the question identifier to form the question-level collection task identifier. At the same time, an incremental sequence number is added according to the order in which the question identifiers appear within the same learning level tag to ensure uniqueness). This generates a question-level collection task set.
[0071] S2.4: Create a question-level collection index for the question-level collection task set, and attach an index number and verification flag to generate a collection and arrangement instruction package.
[0072] Furthermore, the set of question-level collection tasks is arranged according to the registered sequence number in the question-level collection task identifier. A question-level collection index is established for each question-level collection task. The question-level collection index is associated with the question-level collection task identifier, question identifier, and learning level tag and written into the index record. Index numbers are added to the question-level collection index in the order of arrangement. The index numbers are generated and written into the index record according to a preset starting value (e.g., the index numbers increment from 1). At the same time, a verification flag is written to the index record. The uniqueness of the question-level collection task identifier, the validity of the question identifier, and the consistency of the learning level tag are verified according to the verification conditions corresponding to the verification flag (used to determine whether the question-level collection task identifier is unique, whether the question identifier exists in the hierarchical question set, and whether the learning level tag is consistent with the learning level registered in the question-level collection task; the verification is considered successful when all three conditions are met). The uniqueness of the question-level collection task identifier, the validity of the question identifier, and the consistency of the learning level tag are verified. The index records that pass the verification are retained as valid records. The question-level collection tasks corresponding to the valid records are uniformly organized and packaged into a collection orchestration instruction package according to the index sequence number.
[0073] S3: Based on the collection and arrangement instruction package, collect the question image fragments and evidence fragment records, and perform mutual grading of assignments in back-to-back and one-to-many modes. At the same time, collect the grading and annotation trajectory records and generate the collection evidence data package.
[0074] S3.1: Trigger the question-level acquisition tasks sequentially according to the execution order in the acquisition and arrangement instruction package, locate the question resources and capture the question display area, and generate question image fragment records;
[0075] Specifically, the task identifiers at the question level are read sequentially according to the index numbers registered in the task acquisition instruction package. The corresponding question identifier and learning level marker are located based on the task identifier. The corresponding question resource is retrieved from the job release resources according to the question identifier and the question resource presentation content is loaded. The question display area is determined based on the question area boundary information registered in the question resource. The question display area is then checked for area location and boundary integrity. When the question display area meets the boundary integrity requirements, a cropping operation is performed on the question display area. The cropping result, along with the task identifier at the question level, the question identifier, and the index number, is written into the question image fragment record entry to generate a question image fragment record.
[0076] It should be noted that homework publishing resources refer to the "collection of resources used for displaying and answering homework after it has been published" in the education platform. These resources are used to retrieve and load the content of the homework by the homework identifier. They typically include: homework publishing records (homework identifier, class and student range, and publishing time), homework resource index (correspondence between homework identifier, resource address, and storage key), homework content resources (text of the homework stem, sub-question structure of options, and rich media attachments such as formula images), homework presentation configuration (layout template, rendering parameters, and homework area boundary information), and homework-related metadata (knowledge point identifier, difficulty and learning level markers, etc.).
[0077] S3.2: Based on the recorded image fragments of the questions, collect operation trajectory information and establish a connection with the corresponding image fragments of the questions to generate evidence fragment records;
[0078] Specifically, operation trajectory information collection is initiated within the question display area corresponding to the question image fragment. Interaction events occurring within the question display area are recorded and their occurrence timestamps and location markers are registered. The operation trajectory information is then evaluated for validity. Interaction events whose location markers fall within the boundaries of the question display area and meet the trajectory step count and dwell time thresholds are retained as valid operation trajectory information. The valid operation trajectory information is then organized into trajectory fragments according to the time marker order. A one-to-one correspondence is established between the trajectory fragments and the corresponding question image fragments using the question-level collection task identifier and index number. An association key and verification mark are written into the correspondence to generate evidence fragment records.
[0079] It should be noted that the trajectory step threshold is defined based on the cumulative number of consecutive position changes within the same question display area to distinguish between valid operation trajectories and occasional touch point movements (example range: 3 to 10 consecutive position changes); the dwell time threshold is defined based on the duration for which the operation trajectory maintains a relatively stable position within the question display area to distinguish between valid operation behavior and momentary dwell time (example range: dwell time is 200 to 800 milliseconds).
[0080] S3.3: Based on evidence fragment records, conduct mutual review of assignments in a back-to-back and one-to-many mode according to the hierarchical review items, and track and record the entire assignment review process, which is then compiled into a review annotation trajectory record.
[0081] Specifically, based on the evidence fragment records and corresponding hierarchical grading items, assignment peer review is conducted in back-to-back and one-to-many modes. Within the question display area corresponding to the question image fragment, the occurrence time, grading location, and grading status changes of the assignment peer review actions agreed upon in the hierarchical grading items are recorded. Each assignment peer review action is time-aligned with the corresponding operation trajectory information, and the sequence, duration, and grading result mark of the assignment peer review actions are written into the record content. Assignment peer review actions that do not meet the requirements of the hierarchical grading items are marked with an anomaly. The recorded assignment peer review actions are organized according to the time mark order, and the question-level collection task identifier, question identifier, assignment peer review action sequence, grading location, grading status changes, and anomaly mark are uniformly registered to generate a grading annotation trajectory record.
[0082] It should be noted that in the back-to-back and one-to-many models, back-to-back means that the real identity information of the other party is not presented during the mutual review process. Instead, anonymous code names are used to replace identity information, and the anonymous code names are kept consistent in the mutual review allocation and submission stages. This allows mutual review students to receive the question content and submit the grading trajectory based solely on the anonymous code name. One-to-many means that the same question image fragment record is randomly assigned to multiple mutual review students at the same time, and corresponding evidence fragment records and grading annotation trajectory records are formed respectively. This allows the same assignment to obtain grading records from multiple mutual review students under the same hierarchical grading item constraints for subsequent grouping and solidification.
[0083] The requirements for tiered grading entries refer to the clearly defined order of peer review actions, the scope of the questions that must be covered, and the set of grading steps that must be completed under the corresponding learning level mark. Each requirement is judged based on the existence of verifiable action records. For example, under a certain learning level mark, the tiered grading entry requires first marking the area corresponding to the key steps of the question (determined by aligning the question stem number, the writing position of the solution process, or the position of the annotation anchor point corresponding to the key solution steps in the question display area) before giving the result judgment. The marking action must cover the specified position in the question display area and form a corresponding operation trajectory record. If a complete marking trajectory does not appear, it is judged that the requirements of the tiered grading entry are not met.
[0084] S3.4: Group the question image fragment records, evidence fragment records, and correction annotation trajectory records to generate a data package of collected evidence.
[0085] Furthermore, the question-level collection task identifier and question identifier are read one by one from the question image fragment record, evidence fragment record, and correction annotation trajectory record. The question-level collection task identifier is used as the grouping key to perform a grouping operation on the question image fragment record, evidence fragment record, and correction annotation trajectory record. Within the same grouping range, the consistency of the question identifier and the correspondence with the index number of the question image fragment record, evidence fragment record, and correction annotation trajectory record are verified. The records are then arranged in order of index number and organized into a collection evidence data package.
[0086] S4: Chain-associate the evidence fragment records in the collected evidence data package to obtain the question-level collected evidence chain record, and solidify it into a question-level data asset record to generate an asset evidence chain dataset;
[0087] S4.1: Sort the evidence fragment records in the collected evidence data package, fill in missing fields and remove duplicates to generate a question-level ordered evidence fragment sequence;
[0088] Specifically, evidence fragment records are read according to the topic-level collection task identifier, and sorted according to the registered index number and timestamp in the evidence fragment records. Evidence fragment records under the same topic-level collection task identifier are arranged in order. The completeness of the topic-level collection task identifier, topic identifier, index number and verification mark is checked for the evidence fragment records. Missing fields are filled in with the corresponding grouping information in the collected evidence data package. Evidence fragment records with more than the missing field judgment threshold are marked as abnormal. The evidence fragment records are deduplicated. Evidence fragment records with the same topic-level collection task identifier, index number and timestamp are judged as duplicates and only one valid record is retained, generating a topic-level ordered evidence fragment sequence.
[0089] It should be noted that the definition is based on whether the registered question-level collection task identifier, question identifier and index sequence number in the evidence fragment record are complete. Example range: the number of missing fields is allowed to be 1 to 3 and does not affect the determination of the question-level collection task identifier and index sequence number.
[0090] S4.2: Add evidence chain nodes to the ordered evidence fragment sequence at the question level, set pointers to adjacent evidence chain nodes, mark key anchor points, and generate a question-level collected evidence chain record;
[0091] Specifically, the process involves reading each evidence fragment record in the ordered sequence of question-level evidence fragments and assigning an evidence chain node identifier to each record. A correspondence is established between the evidence chain node identifier and the question-level collection task identifier, question identifier, and index number. Based on the order of the evidence fragment records in the ordered sequence, forward and backward pointers are written to adjacent evidence chain nodes. The forward pointer points to the evidence chain node identifier corresponding to the previous index number, and the backward pointer points to the evidence chain node identifier corresponding to the next index number. Only a backward pointer is written to the first evidence chain node, and only a forward pointer is written to the last evidence chain node. Based on the operation trajectory information registered in the evidence fragment records and the key action positions in the correction annotation trajectory records, key anchor point markers are written to the corresponding evidence chain nodes. These key anchor point markers indicate representative node positions in the evidence chain (e.g., the position of the first answer or the first correction for the corresponding question), thus generating a question-level collection evidence chain record.
[0092] It should be noted that the order of the evidence fragments is arranged according to the order of the evidence fragments in the ordered sequence of evidence fragments at the question level. This order is based on the specific order in which each evidence fragment is recorded during the answering or grading process, and is arranged according to the progress of the question or grading.
[0093] S4.3: Solidify the problem-level evidence chain records into problem-level data asset records, and perform grouping and consistency verification to generate an asset evidence chain dataset.
[0094] Furthermore, the task identifier, question identifier, evidence chain node identifier, adjacent evidence chain node pointers, and key anchor point markers in each question-level evidence chain record are read one by one and written into persistent records according to a fixed field structure to form question-level data asset records. Grouping operations are performed on the question-level data asset records based on the question-level task identifier, organizing the question-level data asset records corresponding to the same question-level task identifier into grouped results. Consistency checks are performed on the grouped results, covering the continuity of evidence chain node pointers, the consistency between the number of evidence chain nodes and the number of question-level ordered evidence fragment sequences, and the validity of key anchor point markers. During the consistency check process, the integrity score of the question-level evidence chain is calculated and written into the score field of the grouped results. When both consistency checks and score writing are completed, the corresponding grouped results are marked as passed (e.g., allowing the number of evidence chain nodes to be completely consistent with the number of ordered evidence fragments). The grouped results that pass the consistency check are then uniformly summarized to generate an asset evidence chain dataset.
[0095] It should be noted that the fixed field structure refers to the combination of fields that must be included in the question-level data asset record and stored in a predetermined order. This is used to uniformly save the question-level collection task identifier, question identifier, evidence chain node identifier, adjacent evidence chain node pointers, and key anchor point markers, thereby ensuring the consistency and verifiability of fields between different question-level data asset records.
[0096] The formula for calculating the completeness score of the evidence chain collected at the question level is:
[0097] ;
[0098] in, This indicates the score for the completeness of the evidence chain collected at the question level. This indicates the number of image fragment records that have been grouped into the topic-level data collection task. This indicates the expected number of question image fragment records for the question-level data collection task. Indicates the number of valid evidence fragments recorded. This indicates the expected number of evidence fragment records for a question-level data collection task. This indicates the number of valid correction and annotation trajectory records. This indicates the number of correction annotation trace records that must be generated for each hierarchical correction item. This indicates the number of pointer entries that satisfy the pointer continuity requirement. Indicates the number of missing pointer entries. Indicates the number of abnormal records. This represents the total number of records used to calculate the percentage of anomalies. This represents the weighting coefficient for key anchor points. This indicates the number of key anchor points that passed verification. This indicates the number of critical anchor points that are expected to exist.
[0099] The weighting coefficient for key anchor points (example range: 0.1 to 0.5) is defined based on the upper limit constraint of the proportion of the contribution of key anchor points to the total contribution of integrity score.
[0100] Within the grouping results corresponding to the task identifier at the question level, the number of grouped question image fragment records, the number of valid evidence fragment records, and the number of valid correction annotation trajectory records are statistically analyzed to obtain three sets of ratios, which are then averaged. Simultaneously, in the question-level collected evidence chain records, the number of consecutive pointer entries and the number of missing pointer entries are statistically analyzed to obtain the pointer continuity ratio. The number of abnormal records is also statistically analyzed to obtain the abnormality percentage. Furthermore, the number of key anchor points that have passed verification is statistically analyzed to obtain the expected number, and combined with the key anchor point weighting coefficient, the anchor point weighting item is obtained. The average ratio is multiplied by the pointer continuity ratio and the abnormality deduction item, and then the anchor point weighting item is added to calculate the question-level collected evidence chain completeness score and written into the score field.
[0101] S5: Combine the asset evidence chain dataset with the question-level collection task set to locate collection gaps and convert them into supplementary collection instructions. At the same time, perform supplementary collection and supplementary collection verification to generate the job collection result set.
[0102] S5.1: Align the asset evidence chain dataset with the question-level collection task set item by item, mark the missing evidence types and pointer entries, and generate a question-level collection gap list;
[0103] Specifically, the task-level collection tasks are read one by one according to the task-level collection task identifier in the task-level collection task set. The task-level data asset records corresponding to the same task-level collection task identifier are retrieved in the asset evidence chain dataset. When a task-level data asset record exists, the task-level collection task identifier, the task identifier, and the evidence chain node identifier are aligned and registered. When a task-level data asset record is missing, a missing marker is written at the alignment registration position, and the missing evidence type (referring to any one of the task-level collection task's missing task-level image fragment record, evidence fragment record, or correction annotation trajectory record) is marked as evidence type missing. For the task-level data asset records that have completed the alignment registration, the pointers of adjacent evidence chain nodes are read, and the forward and backward pointers of adjacent evidence chain node pointers are checked to see if they meet the continuity threshold judgment. At the position that does not meet the continuity threshold judgment, the corresponding pointer entry is written with a missing marker, and the missing pointer entry is marked as pointer entry missing, generating a task-level collection gap list.
[0104] It should be noted that the continuity threshold is defined based on the requirement that there must be a traceable pointer relationship between adjacent index numbers in the evidence chain collected at the question level. Example range: Evidence chain nodes with index numbers differing by 1 must simultaneously satisfy that both the forward pointer and the backward pointer are valid pointer entries, and the forward pointer points to the evidence chain node identifier of the adjacent next index number, and the backward pointer points to the evidence chain node identifier of the adjacent previous index number; the first node of the chain only requires that the forward pointer satisfies that the evidence chain node identifier of the adjacent next index number is satisfied, and the last node of the chain only requires that the backward pointer satisfies that the evidence chain node identifier of the adjacent previous index number is satisfied.
[0105] S5.2: Generate supplementary collection instructions for collection gaps based on the list of collection gaps at the question level, and group supplementary collection instructions under the same learning level mark into a level supplementary collection instruction segment to generate a supplementary collection instruction package;
[0106] Specifically, based on the missing evidence type, a supplementary collection instruction is generated for the corresponding topic-level collection task. The supplementary collection instruction clearly states that the target of supplementation is the missing evidence type or the missing pointer entry, and writes the topic-level collection task identifier, topic identifier, and missing marker type into the supplementary collection instruction record. After the supplementary collection instruction record is generated, the learning level markers registered for the corresponding topic-level collection task identifier in the topic-level collection task set are read and added to the supplementary collection instruction record. Simultaneously, the topic-level collection evidence chain completeness score in the topic-level data asset record corresponding to the topic-level collection task identifier is read from the asset evidence chain dataset, and the topic-level collection evidence chain completeness score is written into the supplementary collection instruction record as the basis for prioritizing supplementary collection (example: the lower the topic-level collection evidence chain completeness score, the higher the supplementary collection priority). Based on the learning level markers, the supplementary collection instruction records are grouped, and supplementary collection instruction records corresponding to the same learning level marker are organized into hierarchical supplementary collection instruction segments. The supplementary collection instruction records in the hierarchical supplementary collection instruction segments are arranged in the order of the topic-level collection task identifiers, and all hierarchical supplementary collection instruction segments are uniformly packaged to generate a supplementary collection instruction package.
[0107] S5.3: Based on the supplementary collection instruction package, perform supplementary collection of the topic, append the new evidence fragment record to the topic-level collection evidence chain record, register the supplementary collection source mark, and generate a supplementary collection append record set;
[0108] Specifically, the supplementary sampling instruction records are read one by one in the order of the hierarchical supplementary sampling instruction segments in the supplementary sampling instruction package. Based on the question-level sampling task identifier, the corresponding evidence chain node position is located in the question-level sampling evidence chain record. At the same time, for evidence type missing markers, question supplementary sampling is performed in the question display area corresponding to the question identifier, and the supplementary sampling time marker and supplementary sampling location marker are recorded to form a new evidence fragment record. For pointer entry missing markers, the pointers of adjacent evidence chain nodes are written in the question-level sampling evidence chain record and the verification marker is registered. The new evidence fragment record is appended to the corresponding position in the question-level sampling evidence chain record, and a supplementary sampling source marker is written to the new evidence fragment record to generate a supplementary sampling append record set.
[0109] S5.4: Merge and update the supplementary collection of records with the topic-level data asset records, and re-execute the consistency check. After the check passes, package them into a unified job collection result set.
[0110] Specifically, the supplementary record set and the topic-level data asset record are aligned according to the topic-level acquisition task identifier. The new evidence fragment records registered in the supplementary record set are written into the corresponding topic-level data asset records in the order of their index numbers, and the evidence chain node pointer relationship is updated synchronously. At the same time, the supplementary acquisition source mark is written into the corresponding evidence chain node record to complete the merge update. The consistency check is re-executed on the updated topic-level data asset records. The consistency check includes the consistency judgment of the number of evidence chain nodes and the number of evidence fragment records, the continuity judgment of adjacent evidence chain node pointers, and the validity judgment of key anchor point marks. When the consistency check meets the verification conditions, the corresponding topic-level data asset record is marked as verified (for example, the number of evidence chain nodes is consistent with the number of evidence fragment records after supplementary acquisition). The verified topic-level data asset records are uniformly packaged into the job acquisition result set.
[0111] This embodiment also provides a computer device applicable to the homework data collection method for an education platform, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the homework data collection method for an education platform as proposed in the above embodiment.
[0112] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0113] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the homework data acquisition method for an educational platform as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0114] In summary, this invention, through the generation of chain-linked evidence chains, ensures the anonymity of back-to-back mutual review and one-to-many allocation constraints, solves the problem of fragmented mutual review data in existing systems, and improves the reliability of homework data grouping and teaching decisions; through hierarchical question selection and generation of arrangement instructions, it realizes differentiated collection of homework tasks and data granularity control within a big data collection framework, enhancing the pertinence and effectiveness of educational data analysis.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A homework data collection method applied to an education platform, characterized in that: Comprising, According to the high-frequency error question and high-frequency error reason item in the student historical homework data, the knowledge defect graph is constructed, and the hierarchical arrangement item and the hierarchical correction item are written to generate the knowledge defect graph strategy package; According to the hierarchical arrangement item in the knowledge defect graph strategy package, the hierarchical question set of this homework is selected, and is disassembled into question level collection task, question level collection index is established, and collection arrangement instruction package is generated; Based on the collection arrangement instruction package, the question image segment and the evidence segment record are collected, and the homework mutual correction is carried out in the back-to-back and one-to-many mode, and the correction mark trajectory record is collected, and the collection evidence data package is generated; The evidence segment record in the collection evidence data package is chain associated, the question level collection evidence chain record is obtained, and is solidified into question level data asset record, and the asset evidence chain data set is generated; The asset evidence chain data set and the question level collection task set are combined, the collection gap is positioned, and is converted into supplement collection instruction, and the supplement collection and supplement verification are carried out, and the homework collection result set is generated.
2. The assignment data collection method for an educational platform of claim 1, wherein: The specific steps of the knowledge defect graph strategy package generated by the high-frequency error question and high-frequency error reason item in the student historical homework data are as follows, The student historical homework data is traversed question by question, the high-frequency error question and high-frequency error reason item are screened, and the association relationship is established, and the question and error reason association mapping record set is generated; According to the question and error reason association mapping record set, the high-frequency error reason item is taken as the defect node, and the high-frequency error question is taken as the performance node; The defect node is divided into upper and lower parts through the knowledge point level relationship, and the performance node is hung below the corresponding defect node, and the knowledge defect graph is generated; The defect node is divided into different learning level marks, the corresponding hierarchical arrangement item and hierarchical correction item are obtained, and are written in the knowledge defect graph, and are packaged as the knowledge defect graph strategy package.
3. The assignment data collection method for an educational platform of claim 2, wherein: The specific steps of selecting the hierarchical question set of this homework according to the hierarchical arrangement item in the knowledge defect graph strategy package are as follows, Based on the hierarchical arrangement item in the knowledge defect graph strategy package, the hierarchical question selection boundary and the question screening range boundary are divided, and the priority selection constraint is set, and the hierarchical question selection condition set is generated; Based on the hierarchical question selection condition set, hierarchical filtering, priority selection, question deduplication and truncation reservation are carried out, and the hierarchical question set is generated.
4. The assignment data collection method for an educational platform of claim 3, wherein: The specific steps of generating the collection arrangement instruction package are as follows, The hierarchical question set is unfolded according to the learning level mark, and the question level collection task and the question level collection task identifier are distributed, and the question level collection task set is generated; The question level collection index is established for the question level collection task set, and the index serial number and the verification mark are attached, and the collection arrangement instruction package is generated.
5. The assignment data collection method for an educational platform of claim 4, wherein: The specific steps of collecting the question image segment and the evidence segment record based on the collection arrangement instruction package, and carrying out the homework mutual correction in the back-to-back and one-to-many mode, and collecting the correction mark trajectory record, and generating the collection evidence data package are as follows, According to the execution order in the collection arrangement instruction package, the question level collection task is triggered in sequence, the question resource is positioned and the question display area is intercepted, and the question image segment record is generated; According to the subject image segment record, operation track information is collected and associated with the corresponding subject image segment to generate evidence segment records; Based on the evidence segment records, in the back-to-back and one-to-many modes, the homework mutual correction is performed according to the hierarchical correction items, and the whole homework correction process is tracked and recorded to generate correction annotation track records; The subject image segment records, evidence segment records and correction annotation track records are grouped to generate collection evidence data packages.
6. The assignment data collection method for an educational platform of claim 5, wherein: The evidence segment records in the collection evidence data packages are chain-associated to obtain subject-level collection evidence chain records, which are solidified as subject-level data asset records to generate asset evidence chain data sets, and the specific steps are as follows, The evidence segment records in the collection evidence data packages are sorted, missing fields are supplemented and de-duplicated to generate subject-level ordered evidence segment sequences; The subject-level ordered evidence segment sequences are attached with evidence chain nodes, adjacent evidence chain node pointers are set, and key anchor points are marked to generate subject-level collection evidence chain records; The subject-level collection evidence chain records are solidified as subject-level data asset records, and grouping and consistency checking are performed to generate asset evidence chain data sets.
7. The assignment data collection method for an educational platform of claim 6, wherein: The asset evidence chain data sets are combined with the subject-level collection task sets to locate collection gaps and convert them into supplement collection instructions, and the specific steps are as follows, The asset evidence chain data sets are aligned with the subject-level collection task sets one by one, and missing evidence types and pointer items are marked to generate a subject-level collection gap list; According to the subject-level collection gap list, supplement collection instructions are generated for the collection gaps, and the supplement collection instructions under the same learning level are grouped into level supplement collection instruction segments to generate a supplement collection instruction package.
8. The assignment data collection method for an educational platform of claim 7, wherein: The supplement collection is collected and verified to generate a homework collection result set, and the specific steps are as follows, According to the supplement collection instruction package, the subject is supplemented and collected, new evidence segment records are added to the subject-level collection evidence chain records, and supplement source markers are registered to generate a supplement addition record set; The supplement addition record set is merged and updated with the subject-level data asset records, and consistency checking is performed again. After passing the consistency checking, the homework collection result set is uniformly packaged. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that: The processor executes the computer program to realize the steps of the homework data collection method for the educational platform according to any one of claims 1-8.
10. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to realize the steps of the homework data collection method for the educational platform according to any one of claims 1-8.