A method and system for generating a labeled dataset of safety training multi-source data
By uniformly identifying, aligning, and preprocessing the multi-source heterogeneous data from the security training platform, a labeled dataset is generated, which solves the problem of the difficulty in directly processing multi-source heterogeneous data, improves data utilization efficiency and reliability, and supports subsequent analysis tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-21
AI Technical Summary
Existing security training platforms suffer from difficulties in directly and reliably associating and uniformly processing multi-source heterogeneous data, and lack supporting mechanisms for tag generation and feature construction, resulting in low data utilization efficiency.
By acquiring multi-source heterogeneous data related to construction worker safety training, personnel-level association alignment is performed based on unified identification information. After preprocessing and quality verification, static attribute labels and dynamic learning behavior labels are generated. Static attribute feature vectors and dynamic learning behavior feature vectors are constructed and fused feature vectors are generated to finally form a labeled dataset.
It achieves standardized processing and unified identification alignment of multi-source heterogeneous data, improves the computability and reproducibility of data, enhances data utilization efficiency, and supports subsequent personnel profiling analysis and training resource optimization.
Smart Images

Figure CN122087723B_ABST