A method and system for generating a labeled dataset of safety training multi-source data

By uniformly identifying, aligning, and preprocessing the multi-source heterogeneous data from the security training platform, a labeled dataset is generated, which solves the problem of the difficulty in directly processing multi-source heterogeneous data, improves data utilization efficiency and reliability, and supports subsequent analysis tasks.

CN122087723BActive Publication Date: 2026-07-21CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-04-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing security training platforms suffer from difficulties in directly and reliably associating and uniformly processing multi-source heterogeneous data, and lack supporting mechanisms for tag generation and feature construction, resulting in low data utilization efficiency.

Method used

By acquiring multi-source heterogeneous data related to construction worker safety training, personnel-level association alignment is performed based on unified identification information. After preprocessing and quality verification, static attribute labels and dynamic learning behavior labels are generated. Static attribute feature vectors and dynamic learning behavior feature vectors are constructed and fused feature vectors are generated to finally form a labeled dataset.

Benefits of technology

It achieves standardized processing and unified identification alignment of multi-source heterogeneous data, improves the computability and reproducibility of data, enhances data utilization efficiency, and supports subsequent personnel profiling analysis and training resource optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087723B_ABST
    Figure CN122087723B_ABST
Patent Text Reader

Abstract

The present application relates to the field of intelligent construction site data management and safety training information technology, and discloses a safety training multi-source data labeled data set generation method and system. The method comprises: obtaining static attribute data and dynamic learning behavior data related to construction worker safety training and performing correlation alignment; performing processing, screening and sample reservation on the aligned multi-source heterogeneous data; performing job label semantic reclassification and derived behavior label calculation on the reserved samples to obtain static attribute labels and dynamic learning behavior labels; constructing the labels to generate static attribute feature vectors, dynamic learning behavior feature vectors and fusion feature vectors, and generate encoding parameters; and collecting and integrating the labels to form a labeled data set by taking unified identification information as a personnel-level primary key. The present application can realize unified organization and reproducible output of multi-source heterogeneous data under the premise of protecting the privacy information of construction workers, and improve the data utilization efficiency of the safety training platform and the engineering deployment adaptation ability.
Need to check novelty before this filing date? Find Prior Art