Data annotation training system
Patent Information
- Application Number
- CN202510871212.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
Smart Images

Figure CN120807235A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data labeling, and particularly relates to a data standard training system. BACKGROUND
[0002] With the development of intelligent transportation systems, video and image data play a crucial role in traffic management and monitoring. However, video cloud data labeling can only form structured data to support AI middle platform algorithm iteration. Video labeling mainly labels, assigns or annotates video content to provide training data for computer vision and artificial intelligence models, and its core tasks include target detection, behavior recognition, event labeling, etc., and is widely used in automatic driving, security monitoring, medical imaging and other fields.
[0003] Although there are some ways to automatically identify and standardize, the overall recognition accuracy is low, and manual verification and error correction are still needed, so the training of video labeling talents is particularly important. The core content of video standard training is to teach students how to accurately, consistently and efficiently add structured labels or annotations to video data to meet the needs of specific artificial intelligence model training or data analysis. SUMMARY
[0004] The technical problem solved by the present application is to provide a data standard training system that can improve the training efficiency of data labeling.
[0005] The basic scheme provided by the present application is a data standard training system, which comprises a server, the server comprising a video library module, a task distribution module, an exercise scoring module and a learning report module; The video library module pre-stores labeled video data; The task distribution module is used to retrieve video data from the video library and send the retrieved unlabeled version of the video data to the students; The exercise scoring module is used to obtain the labeled video data of the students, compare the students' labeling with the labeling in the video library, the comparison content including boundary box specification identification, label classification identification, target tracking identification and action event identification, and generate a score according to the comparison results of each comparison; The learning report module is used to store the comparison results and scores of each student, and generate a learning report for each student according to the comparison results and scores of each student's practice within a period.
[0006] Further, the exercise scoring module comprises a boundary specification identification module, and the boundary box specification identification module comprises an intersection-over-union calculation module and a size ratio detection module. The intersection-over-union calculation module is configured to calculate the spatial overlap of the labeled bounding box of the trainee and the pre-stored standard bounding box.
[0007] When , it is determined that the position labeling is qualified, is a pre-set intersection-over-union threshold value; The size ratio detection module calculates the size deviation rate of the area of the labeled bounding box and the area of the standard bounding box .
[0008] The size deviation rate within the pre-set deviation rate threshold value is determined to be qualified in size; The width-to-height ratio deviation rate of the labeled bounding box and the standard bounding box is calculated:
[0009] wherein is the width-to-height ratio of the labeled bounding box, is the width-to-height ratio of the standard bounding box; The boundary specification identification module is configured to output a score according to the intersection-over-union and the width-to-height ratio.
[0010] Further, the exercise scoring module further comprises a label classification identification module, which comprises a basic label matching verification module, a missing and redundant verification module, and a hierarchical relationship verification module: The basic label matching verification module is configured to compare the matching consistency of the labeled label of the trainee and the pre-stored standard label to obtain a basic matching accuracy, the basic matching accuracy = correct label number / total label number; The missing and redundant verification module is configured to identify all standard labels contained in the video data, identify redundant labels and missing labels, and obtain an integrity proportion, the integrity proportion = 1 - missing or redundant label number / total label number; The hierarchical relationship verification module is configured to verify whether the sub-class label matches the parent class label according to a pre-set classification number; The label classification identification module is configured to generate a score according to the basic matching accuracy, the integrity proportion, and the sub-class label matching degree.
[0011] Further, the exercise scoring module further comprises a target action identification module; the target action identification module comprises The start and end frame identification module is configured to identify the deviation of the action start frame and the end frame labeled by the trainee from the standard; The target action identification module is configured to generate a score according to the deviation of the start frame, the end frame, and the standard, and the key frame labeling integrity.
[0012] Further, the exercise scoring module further comprises a motion event recognition module, the motion event recognition module is used for analyzing the cooperative working relationship and event trigger condition of multiple targets, verifying the space-time continuity of event logic chain, detecting the integrity of event elements, and generating a score.
[0013] Further, the exercise scoring module further comprises a comprehensive scoring module; The comprehensive scoring module is used for calculating a comprehensive score according to the scores generated by each module and the preset weights of each module.
[0014] The principles and advantages of the present application are that: A set of data labeling training closed-loop system based on multi-dimensional intelligent evaluation and dynamic optimization is constructed, a standard reference system is formed by pre-storing a labeled video library, a task distribution module pushes a video segment to be labeled to a student, and an exercise scoring module adopts a hierarchical technical means to deeply analyze and quantitatively evaluate the labeling result of the student: a bounding box specification recognition module verifies the target positioning accuracy by calculating parameters such as IoU and size deviation rate, a label classification recognition module ensures the accuracy and logical consistency of semantic labeling by relying on a classification tree structure and a conflict rule library, a target motion recognition module verifies the motion continuity by taking an atomic motion as a unit and combining a time tolerance mechanism, and a motion event recognition module verifies the cause-effect relationship of a complex scene by multi-target space-time modeling and atomic event chain reconstruction technology; a comprehensive scoring module innovatively introduces a dynamic weight distribution mechanism driven by a capability portrait, normalizes the evaluation results of the above four dimensions, adjusts the weight coefficients in real time according to historical defect data of the student, further generates a three-dimensional labeling capability matrix and maps it into a visualized heat map, and simultaneously automatically matches and optimizes path instructions based on a pre-constructed course knowledge graph, forming an intelligent training closed loop of "evaluation-positioning-improvement".
[0015] The patent has the following advantages: first, the traditional manual evaluation mode is completely innovated, the subjective and inefficient problems of label quality evaluation are solved by using a machine quantifiable evaluation system: the spatial overlap calculation of the boundary box module controls the target positioning accuracy deviation at the pixel level, the hierarchical tree verification of the label classification module avoids the model training failure caused by cross-layer labeling, and the spatio-temporal interaction modeling of the action event module can accurately capture the logical errors in the multi-target cooperative relationship. Second, the dynamic weight mechanism breaks through the limitations of static scoring, and the system automatically strengthens the training weight of weak links by continuously tracking the evolution trend of student defects, so that the training resources can accurately focus on the ability short board; third, the deep coupling of the ability matrix and the knowledge graph realizes the revolution of individualized teaching, the three-dimensional ability model decomposes the abstract labeling skills into measurable indicators such as accuracy, efficiency and logic, and the heat map intuitively presents the current ability boundary, such as automatically pushing the atomic event splicing training course when detecting event logic chain defects, and finally, the system forms strong industry adaptability through the cooperation of the pre-defined rule base and the algorithm model, the verification mechanism of the turning light frame and the wheel pressure line frame key frame set for the vehicle lane change action in the traffic scene, the continuity analysis module of the organ motion trajectory in the medical scene, and the abnormal behavior detection engine based on the mutual exclusion action library in the security field, which proves the scalability of the system in multiple fields, and provides core technical support for the standardization of artificial intelligence industry. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The drawings illustrate the embodiments of the present application. DETAILED DESCRIPTION
[0017] The embodiments will be further described in detail below: The embodiments are basically as shown in the accompanying drawings: Figure 1 A data standard training system includes a server, the server includes a video library module, a task distribution module, an exercise scoring module, and a learning report module; The video library module pre-stores labeled video data; The task distribution module is used to retrieve video data from the video library and send the un-labeled version of the retrieved video data to the student; The exercise scoring module is used to obtain the student's labeling of the video data, compare the student's labeling with the labeling in the video library, the comparison content includes boundary box specification identification, label classification identification, target tracking identification, and action event identification, and generate a score according to the comparison results of each comparison; The learning report module is used to store the comparison results and scores of each student, and generate a learning report for each student according to the comparison results and scores of the student's exercises within a period.
[0018] Further, the exercise scoring module comprises a bounding box specification identification module, the bounding box specification identification module comprising an intersection over union calculation module, a size proportion detection module; The intersection over union calculation module is configured to calculate the spatial overlap between the labeled bounding box of the trainee and the pre-stored standard bounding box, i.e., an intersection over union (IoU) ratio.
[0019] When the IoU ratio is greater than or equal to a pre-set IoU threshold, it is determined that the position labeling is qualified. The pre-set IoU threshold is a pre-set intersection over union threshold. The size proportion detection module is configured to calculate the size deviation rate of the area of the labeled bounding box and the area of the standard bounding box.
[0020] When the size deviation rate is within a pre-set deviation rate threshold, it is determined that the size is qualified. The aspect ratio deviation rate of the labeled bounding box and the standard bounding box is calculated.
[0021] wherein is the aspect ratio of the labeled bounding box, is the aspect ratio of the standard bounding box. The bounding box specification identification module is configured to output a score according to the intersection over union and the aspect ratio.
[0022] Further, the exercise scoring module further comprises a label classification identification module, the label classification identification module comprising a basic label matching verification module, a missing and redundant verification module, and a hierarchical relationship verification module. The basic label matching verification module is configured to compare the matching consistency of the labeled label of the trainee and the pre-stored standard label, to obtain a basic matching accuracy, the basic matching accuracy = correct label number / total label number. The missing and redundant verification module is configured to identify all standard labels contained in the video data, to identify redundant labels and missing labels, and to obtain an integrity proportion, the integrity proportion = 1-missing or redundant label number / total label number. The hierarchical relationship verification module is configured to verify whether the sub-class label matches the parent class label according to a pre-set classification number. The label classification identification module is configured to generate a score according to the basic matching accuracy, the integrity proportion, and the sub-class label matching degree.
[0023] Further, the exercise scoring module further comprises a target action identification module; the target action identification module comprising The start and end frame identification module is configured to identify the deviation of the action start frame and the action end frame labeled by the trainee from the standard. The target action recognition module is used to generate scores based on the deviation of the start frame and end frame from the standard and the completeness of the key frame annotation.
[0024] Furthermore, the practice scoring module also includes an action event recognition module, which is used to analyze the collaborative working relationship and event triggering conditions of multiple targets, verify the spatiotemporal continuity of the event logic chain, detect the integrity of event elements, and generate a score.
[0025] Furthermore, the practice scoring module also includes a comprehensive scoring module; The comprehensive scoring module is used to calculate the comprehensive score based on the scores generated by each module and the preset weights of each module.
[0026] The principles and advantages of the present invention are: A closed-loop data labeling training system based on multi-dimensional intelligent evaluation and dynamic optimization has been constructed. A standard reference system is formed by pre-stored annotated video libraries. The task distribution module pushes video clips to be annotated to trainees. The practice scoring module uses hierarchical technical means to deeply analyze and quantitatively evaluate the trainees' annotation results. The bounding box specification recognition module verifies target positioning accuracy by calculating parameters such as the intersection over union (IoU) and size deviation rate. The label classification recognition module relies on a classification tree structure and a conflict rule library to ensure the accuracy and logical consistency of semantic annotation. The target action recognition module uses atomic actions as units and combines a timing tolerance mechanism to verify action continuity. The action event recognition module verifies causal relationships in complex scenarios through multi-target spatiotemporal modeling and atomic event chain reconstruction technology. The comprehensive scoring module innovatively introduces a dynamic weight allocation mechanism driven by capability profiling. After normalizing the evaluation results of the above four dimensions, the weight coefficients are adjusted in real time based on the trainees' historical deficiency data. This generates a three-dimensional labeling capability matrix and maps it into a visual heat map. Simultaneously, based on the pre-built course knowledge graph, it automatically matches optimization path instructions, forming an intelligent training closed loop of "assessment-positioning-improvement."
[0027] The obvious advantages of the patent are firstly embodied in the complete innovation of the traditional manual evaluation mode, and the subjective and low efficiency of the labeling quality evaluation is solved by using the machine quantifiable evaluation system: the spatial overlap degree calculation of the boundary box module controls the target positioning accuracy deviation at the pixel level, the hierarchical tree verification of the label classification module avoids the model training failure caused by cross-layer labeling, and the spatio-temporal interaction modeling of the action event module can accurately capture the logical errors in the multi-target cooperative relationship. Secondly, the dynamic weight mechanism breaks through the limitations of static scoring, and the system automatically strengthens the weak link training weight by continuously tracking the evolution trend of the student's defects, so as to make the training resources accurately focus on the ability short board; thirdly, the deep coupling of the ability matrix and the knowledge graph realizes the revolution of individualized teaching, the three-dimensional ability model decomposes the abstract labeling skills into measurable indicators such as accuracy, efficiency and logic, and the heat map intuitively presents the current ability boundary, such as automatically pushing the atomic event splicing training course when detecting the event logic chain defect, and finally, the system forms strong industry adaptability through the synergistic effect of the pre-defined rule library and the algorithm model, the verification mechanism of the turning light opening frame and the wheel line pressing frame key frame is set for the vehicle lane changing action in the traffic scene, the continuity analysis module of the organ motion trajectory in the medical scene, and the abnormal behavior detection engine based on the mutual exclusion action library in the security field, which proves the scalability of the system in multiple fields, and provides core technical support for the artificial intelligence industry to deliver standardized labeling talents.
[0028] The above is only an embodiment of the present application, and the common knowledge of specific structures and characteristics in the scheme is not described in detail, and the ordinary skilled person in the art knows all the ordinary technical knowledge in the field of the application before the application date or the priority date, can know all the prior art in the field, and has the ability to apply conventional experimental means before that date, and the ordinary skilled person in the art can improve and implement the present scheme based on their own ability under the guidance of the present application, and some typical known structures or known methods should not be an obstacle for the ordinary skilled person in the art to implement the present application. It should be pointed out that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application, which will not affect the effect and practicality of the patent. The protection scope claimed in the present application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.
Claims
1. A data annotation training system, characterized by: The server comprises a video library module, a task distribution module, an exercise scoring module and a learning report module; Video library module, pre-stored with labeled video data; The task distribution module is used to retrieve video data from the video library and send the unlabeled version of the retrieved video data to the students; The exercise scoring module is used to obtain the video data annotated by the students and compare the student's annotations with the annotations in the video library. The comparison content includes bounding box specification recognition, label classification recognition, target tracking recognition, and action event recognition, and generates a score based on the comparison results of each comparison; The learning report module is used to store the comparison results and scores of each student, and generate a learning report for each student based on the comparison results and scores of each exercise of the student within the period.
2. A data annotation training system according to claim 1, characterized in that: The practice scoring module includes a boundary specification recognition module, and the bounding box recognition specification module includes an intersection-over-union calculation module and a size ratio detection module; The intersection-over-union (IoU) calculation module is used to calculate the spatial overlap between the student's labeled bounding box and the pre-stored standard bounding box: when When , it is judged that the position marking is qualified. is the preset intersection-over-union ratio threshold; Size ratio detection module, calculates the area of the annotation bounding box With the standard bounding box area Dimensional deviation rate: If the size deviation rate is within the preset deviation rate threshold, the size is judged to be qualified; Calculate the aspect ratio deviation between the annotation bounding box and the standard bounding box: in is the aspect ratio of the annotation bounding box, is the aspect ratio of the standard bounding box; Boundary norm identification module, which outputs scores based on the intersection-over-union ratio and aspect ratio.
3. A data annotation training system according to claim 2, characterized in that: The exercise scoring module also includes a label classification and identification module, which includes a basic label matching verification module, a missing redundancy verification module, and a hierarchical relationship verification module: The basic label matching verification module is used to compare the matching consistency between the student's annotated labels and the pre-stored standard labels to obtain the basic matching accuracy rate, which is the number of correct labels / total number of labels. A missing redundancy verification module is used to identify all standard tags contained in the video data, identify redundant tags and missing tags, and obtain the integrity ratio, where the integrity ratio = 1-the number of missing or redundant tags / the total number of tags; The hierarchical relationship verification module is used to verify whether the sub-category label matches the parent category label based on the preset number of categories; The label classification and recognition module is used to generate scores based on the basic matching accuracy, completeness ratio, and sub-category label matching degree.
4. A data annotation training system according to claim 3, characterized in that: The practice scoring module also includes a target action recognition module; the target action recognition module includes a start and end frame recognition module, a key frame recognition module The start and end frame recognition module is used to identify the deviation between the start and end frames of the action marked by the students and the standard; Keyframe recognition module, used to verify the continuity of action and the integrity of keyframe annotations; The target action recognition module is used to generate scores based on the deviation of the start frame and end frame from the standard and the completeness of the key frame annotation.
5. A data annotation training system according to claim 4, characterized in that: The practice scoring module also includes an action event recognition module, which is used to analyze the collaborative working relationship and event triggering conditions of multiple targets, verify the spatiotemporal continuity of the event logic chain, detect the integrity of event elements, and generate a score.
6. The data annotation training system according to claim 5, characterized in that: The practice scoring module also includes a comprehensive scoring module; The comprehensive scoring module is used to calculate the comprehensive score based on the scores generated by each module and the preset weights of each module.