Incremental learning-based post-identification method and system for unsafe behaviors of employees

Through the incremental learning method, the YOLOv8 algorithm and adaptive KL divergence and weight are used to improve the model timing increment, which solves the problem of insufficient detection capabilities of the existing model in variable scenarios, and reduces the deployment cost through the front-end coexistence architecture, achieving efficient identification of employee insecure behavior.

CN120220243APending Publication Date: 2025-06-27ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510376762.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing employee insecurity behavior model is immutable and the detection capabilities in variable scenarios are insufficient, and the technologies and architectures used by multiple identification systems deployed by enterprises are different, resulting in high replacement and deployment costs and difficult.

Method used

The post-recognition method of employee unsafe behavior based on incremental learning is adopted, and the basic model is constructed using the YOLOv8 algorithm, and the timing incremental improvement is carried out through manual interactive data and adaptive KL divergence and weights to improve the recognition effect of the model in variable scenarios. At the same time, a back-recognition architecture for coexistence between front and back-end is designed, using the existing system as the front-end, and the basic model and the manual interaction module as the back-end to realize secondary recognition and reduce resource consumption.

Benefits of technology

It improves the recognition effect of employee insecure behavior in changing scenarios, reduces deployment and replacement costs, and achieves low-cost and efficient deployment in existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220243A_ABST
    Figure CN120220243A_ABST
Patent Text Reader

Abstract

The invention discloses an incremental learning-based post-identification method and system for unsafe behaviors of employees, and belongs to the field of video analysis. In order to solve the problems that an existing detection model is weak in target detection capability and poor in recognition effect in a variable scene, historical unsafe behavior videos are extracted and labeled, a data set is constructed, a basic target detection model is constructed by using a YOLOv8 model, and the target detection efficiency is improved. And in combination with data for carrying out manual interaction on a secondary identification result of the basic target detection model, carrying out timed increment lifting on the basic target detection model by utilizing self-adaptive KL divergence and weight. Meanwhile, an employee unsafe behavior post-identification framework with coexisting front and rear ends is constructed by taking an existing unsafe behavior identification system of an enterprise as a front end and taking a basic target detection model, human interaction and a model timing increment module as a rear end, so that the deployment cost and the implementation difficulty of the method are reduced, and the implementation efficiency is improved. And the detection effect of the model on variable scenes is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for post-identifying employees' unsafe behaviors based on incremental learning. Background Art

[0002] With the complication of the production and working environment, various unsafe behaviors may occur during the operation of employees, thus greatly increasing the risk of accidents. Therefore, enterprises need to promptly identify and correct employees' unsafe behaviors. With the development of computer vision technology, the traditional method of manually identifying employees' unsafe behaviors has gradually been replaced by the automatic identification method based on computer vision due to its high cost, low efficiency, and poor stability. This method obtains videos through monitoring devices and automatically identifies and alarms employees' unsafe behaviors in the videos by using computer vision technology, thereby effectively ensuring the safety of employees' lives and enterprise property.

[0003] Currently, most enterprises have deployed employee unsafe behavior recognition systems based on computer vision technology, which can automatically identify unsafe behaviors through monitoring videos and provide timely warning and intervention means for enterprises. For example, the patent publication number is CN 118470787A, the publication date is August 9, 2024, and the patent name is: Method, device, and storage medium for employee behavior recognition based on video understanding network; this application collects video data, constructs an employee violation behavior dataset; trains a video understanding network based on the employee violation behavior dataset; builds the trained video understanding network in an edge device; detects whether there is a behavior to be recognized in the video data by setting a fuzzy matching threshold based on the fitting degree; sets at least one employee violation behavior to be monitored; establishes a recognition record database, obtains real-time monitoring videos, and recognizes and records employee violation behaviors.

[0004] Another example is the patent publication number CN110414320A, the publication date is November 5, 2019, and the patent name is: A method and system for safety production supervision; this application obtains historical videos of monitored objects and extracts multiple images from the obtained historical videos at preset time intervals; marks each of the extracted multiple images with corresponding labels, and performs image processing and feature extraction on each image with an existing label; constructs an image recognition model based on convolutional neural network regression according to each image with an existing label after image processing and feature extraction, and uses the error backpropagation algorithm to train the image recognition model until convergence to obtain a trained image recognition model; finally, obtains the current image to be inspected of the monitored object, and imports the current image to be inspected of the monitored object into the trained image recognition model for recognition to determine whether there are potential safety hazards for the monitored object.

[0005] Although the above method realizes the automatic recognition of employees' unsafe behaviors, the video understanding network and image recognition model trained by the dataset are fixed. Therefore, in the face of changing scenarios, the recognition effect of employees' unsafe behaviors is often poor due to insufficient detection ability. At present, most enterprises have deployed multiple unsafe behavior recognition systems, and the recognition technologies and deployment architectures used by these systems are different. If directly replaced with a new model, all systems need to be re-adapted, resulting in high implementation difficulty and cost. Summary of the Invention

[0006] 1. Technical problems to be solved by the invention

[0007] First, aiming at the problem that the existing employees' unsafe behavior model has insufficient detection ability in the face of changing scenarios due to immutability and poor recognition effect of employees' unsafe behaviors. The present invention proposes a post-recognition method for employees' unsafe behaviors based on incremental learning. The present invention uses the YOLOv8 algorithm to construct a basic model, combines the data of human-computer interaction, and uses the adaptive KL divergence and weights to incrementally improve the basic model regularly, thereby improving the recognition effect of unsafe behaviors of the model in changing scenarios.

[0008] Second, aiming at the fact that existing enterprises have deployed multiple unsafe behavior recognition systems, and the recognition technologies and deployment architectures used by these systems are different. In order to reduce the implementation cost and difficulty of deploying the recognition method proposed by the present invention to all recognition systems for secondary recognition, the present invention proposes a post-recognition system for employees' unsafe behaviors based on incremental learning. This system is a post-recognition architecture with both front-end and back-end coexisting. The existing multiple recognition systems are used as front-end modules, and the basic object detection model, human-computer interaction, and model regular incremental module are used as back-end modules. The front-end module is used to complete resource-consuming operations such as real-time monitoring of the employee behavior video stream and saving alarm information, and the back-end module performs secondary recognition on the alarm information of the front-end module, etc., thereby reducing resource consumption.

[0009] 2. Technical solutions

[0010] To achieve the above object, the technical solutions provided by the present invention are as follows:

[0011] A post-recognition method for employees' unsafe behaviors based on incremental learning of the present invention includes the following steps:

[0012] Step 1, construct a basic object detection model

[0013] Extract unsafe behavior video data from the monitoring video stream, annotate and generate a text file containing behavior categories and bounding boxes, construct a dataset using the images and the corresponding text files, and divide the training set, validation set, and test set. Use the YOLOv8 model to train to obtain a basic object detection model;

[0014] Step 2: Perform manual interaction data annotation

[0015] Receive manual review feedback. When the model recognition result is marked as inaccurate, intercept multiple consecutive frames of images associated with unsafe behaviors, pre-annotate the images through a semi-automated annotation tool, and use the basic object detection model to automatically detect the images. Then, manually correct the position and category of the detection box and store them in the incremental dataset.

[0016] Step 3: Timed incremental improvement of the model

[0017] Regularly count the data volume of the incremental dataset. When the data volume reaches the set threshold, perform incremental improvement on the basic object detection model based on the adaptive KL divergence and weights, update the model parameters, and replace the basic object detection model.

[0018] Furthermore, in the step of constructing the basic object detection model, the dataset is randomly divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The generated text file for annotation contains information on the types, positions, and sizes of unsafe behaviors.

[0019] Furthermore, in the manual interaction step, create a manual interaction mechanism. This manual interaction mechanism defines an object to receive the feedback parameters of the manual review in the visualization interface. The object attributes include a flag indicating whether the recognition result is accurate, the time when the unsafe behavior occurs, and the corresponding video stream of the unsafe behavior. The manual interaction mechanism uses the judgment of the flag to perform corresponding operations. For example, when the recognition result is marked as inaccurate, intercept a total of 4 images before and after the time when the unsafe behavior occurs, and perform automatic detection and manual fine-tuning of the image targets through the annotation tool.

[0020] Furthermore, the timed incremental improvement of the model specifically includes the following steps:

[0021] 3.1 Extract the fine-tuning images and corresponding text files from the data folder after manual interaction to construct a new training dataset;

[0022] 3.2 Use the saved basic object detection model to construct a teacher network, and use the YOLOv8 original model to construct a student network. Input the new training set into the teacher network and the student network, and respectively output the original prediction scores logit of the unsafe behavior categories;

[0023] 3.3 Divide the logit of the teacher network and the student network by the temperature parameter T, and extract the softened probability distribution soft label values of each unsafe behavior category after softmax operation;

[0024] 3.4 Based on the soft label values of the teacher network and the student network, calculate the entropy difference of each unsafe behavior category, and generate a weighting factor through maximum normalization processing;

[0025] 3.5. Construct an adaptive KL divergence based on the weighting factor and KL divergence, and calculate the Soft loss between the teacher network and the student network;

[0026] 3.6. Generate hard label values by performing softmax operation on the logits output by the student network, and calculate the Hard loss by combining the classification and regression losses;

[0027] 3.7. Generate an adaptive weight coefficient λ according to the number of images in the new training set and the batch size, and sum the Soft loss and Hard loss after weighting to form the total loss;

[0028] 3.8. Train the student network using the total loss, and replace the original basic model with the updated object detection model.

[0029] Furthermore, the calculation formula for entropy in step 3.4 is:

[0030] H P,i =-p i log2p i

[0031] H Q,i =-q i log2q i

[0032] Among them, H P,i represents the entropy of the teacher network for each category, H Q,i represents the entropy of the student network for each category, p i and q i respectively represent the soft label values of the teacher network and the student network for each category.

[0033] Furthermore, the calculation formula for the weighting factor in step 3.4 is:

[0034]

[0035] Among them, w i represents the weight factor assigned to each category, and α is a hyperparameter.

[0036] Furthermore, the calculation formula for the adaptive weight coefficient λ in step 3.7 is:

[0037]

[0038] Among them, i represents the index of the current training round, and N represents the number of training batches.

[0039] Furthermore, the triggering condition for the timed incremental improvement of the model is to count the amount of newly labeled data in the incremental dataset every hour. When the amount of data exceeds a preset threshold, incremental improvement is performed on the basic object detection model; otherwise, no processing is done.

[0040] An employee unsafe behavior post-identification system based on incremental learning according to the present invention includes:

[0041] Front-end module: Using the existing enterprise unsafe behavior identification system, it identifies employees' behaviors in real time and stores the identified information of employees' unsafe behaviors in the databases of their respective identification systems.

[0042] Data transmission module: Connects to the multi-source databases of the existing enterprise identification systems through the DBAPI low-code tool, regularly obtains alarm information, and uniformly stores it in the specified database table.

[0043] Back-end module: Based on the method described in any one of claims 1-8, uses the constructed basic object detection model to perform secondary identification of unsafe behaviors, performs manual interactive annotation and timed incremental training of the model, and updates the object detection model.

[0044] Furthermore, the system sets up a visualization interface, displays the information related to the unsafe behaviors secondarily identified by the back-end module in a list form, and uses manual review operations to return the review results of the secondary identification to the back-end module in the form of parameters.

[0045] 3. Beneficial effects

[0046] Adopting the technical solution provided by the present invention, compared with the existing well-known technologies, it has the following remarkable effects:

[0047] (1) For the method for post-identifying employees' unsafe behaviors based on incremental learning of the present invention, by extracting and annotating historical unsafe behavior videos, constructing a dataset, using the YOLOv8 model to construct a basic object detection model, and combining the data of manual interaction by enterprise personnel on the secondary identification results of the basic object detection model, timed incremental improvement of the basic object detection model is performed using the adaptive KL divergence and weights. It can continuously incrementally improve the model, better solve the problem of insufficient model detection ability in variable scenarios, and thus improve the identification effect of employees' unsafe behaviors in variable scenarios.

[0048] (2) A post-identification system for employees' unsafe behaviors based on incremental learning in the present invention uses the existing unsafe behavior identification system in the enterprise as the front end, and the basic object detection model, human-computer interaction, and model timing increment module as the back end to construct a post-identification architecture for employees' unsafe behaviors with coexistence of the front and back ends. It can be deployed in the existing enterprise system with low cost and low difficulty to identify the post-identification method for employees' unsafe behaviors, so as to conduct secondary identification of employees' unsafe behaviors and further improve the identification effect of employees' unsafe behaviors. Description of the Drawings

[0049] Figure 1 It is a flowchart of a post-identification method for employees' unsafe behaviors based on incremental learning proposed by the present invention;

[0050] Figure 2 It is a flowchart for improving model timing increment proposed by the present invention;

[0051] Figure 3 It is an architecture diagram for post-identification of employees' unsafe behavior videos with coexistence of the front and back ends constructed by the present invention. Detailed Embodiments

[0052] To further understand the content of the present invention, the present invention will be described in detail in combination with the drawings and embodiments.

[0053] Embodiment 1

[0054] Combined with Figure 1 , a post-identification method for employees' unsafe behaviors based on incremental learning in this embodiment is as follows:

[0055] Step 1: Construct a basic object detection model

[0056] Obtain relevant unsafe behavior videos from the historical monitoring video stream of the manufacturing enterprise, and intercept multiple images with a resolution of 640*640 from the obtained videos according to timestamps; for the intercepted images, deploy an image automation annotation tool X-AnyLabeling on the enterprise server; use it to annotate the features of the images, and generate a text file with the same name as the image after annotation. The text file mainly contains the types, locations, and sizes of unsafe behaviors. Construct a data set using the images and the corresponding text files, and randomly divide it into a training set, a validation set, and a test set according to a ratio of 7:2:1. Use the YOLOv8 model to train on the training set and validate on the validation set and the test set to construct a high-precision basic object detection model, and save it in the basic object detection model folder set on the enterprise server.

[0057] Step 2: Conduct human-computer interaction data annotation

[0058] Create an artificial interaction mechanism in the incremental improvement module. This artificial interaction mechanism defines an object to receive the feedback parameters of the enterprise reviewers on the secondary recognition results on the visualization interface. According to the feedback of the enterprise reviewers on the secondary recognition results, obtain a small number of images from the unsafe behavior videos.

[0059] The object attributes mainly include the flag indicating whether the recognition result is accurate (for example, 0 represents accurate, 1 represents inaccurate), the time when the unsafe behavior occurs, and the corresponding unsafe behavior video stream. The artificial interaction mechanism uses the judgment of the flag to perform corresponding operations; if the value of the flag is 0, no operation is performed; otherwise, according to the time when the unsafe behavior occurs in the object, intercept 4 images with a size of 640*640 at and around the corresponding time in the unsafe behavior video.

[0060] Subsequently, automatically open the X-AnyLabeling annotation tool interface. The enterprise personnel click the automatic detection button, and this annotation tool will use the saved basic object detection model to automatically detect the intercepted images and identify the suspected targets; the enterprise personnel fine-tune the detection results on the interface to obtain the accurate targets. After closing the interface, the fine-tuned images and the corresponding text files will be stored in the data folder set by the enterprise server.

[0061] Step 3: Model Timed Incremental Improvement

[0062] Create a timing mechanism for model incremental improvement in the incremental improvement module. Define a variable in this timing mechanism to receive the threshold passed by the enterprise personnel on the visualization interface, and create a timing task to regularly count the number of pictures saved in the data folder after artificial interaction at a set time interval (such as every hour); use the counted number of images and the variable receiving the threshold to make a size judgment. If the counted number of pictures is less than this variable, no processing is done; otherwise, perform incremental improvement on the basic object detection model.

[0063] The process of the incremental improvement is as Figure 2 shown, and specifically includes the following steps:

[0064] Step 3.1: Extract the fine-tuned images and the corresponding text files from the data folder in Step 2 to construct a new training dataset;

[0065] Step 3.2: Use the basic object detection model saved in the enterprise server to construct a teacher network, and use the YOLOv8 original model to construct a student network; input the new training dataset into the teacher network and the student network, and respectively output the original prediction scores for the unsafe behavior categories, that is, logit;

[0066] Step 3.3: Divide the logits output by the teacher network and the student network by the temperature parameter T, and perform a softmax operation to respectively extract the softened probability distributions of the teacher network and the student network for the unsafe behavior categories, that is, the soft label values;

[0067] For example, set the categories of unsafe behaviors as 0: not wearing a safety helmet, 1: not wearing work clothes, 2: sleeping on the job, 3: leaving the post; if the logit values predicted by the teacher network for each category are [-5, 2, 7, 9], then after passing through the softmax function with the temperature parameter T = 3, the soft label values are [0.0058, 0.0599, 0.3170, 0.6174]; if the logit values predicted by the student network for each category are [-10, 0, 3, 12], after passing through the softmax function with the temperature parameter T = 3, the soft label values are [0.0006, 0.0171, 0.0466, 0.9357];

[0068] Step 3.4: According to the soft label values calculated in Step 3.3, use entropy to help the student model capture the uncertainty of the soft label values output by the teacher model for each category, and avoid it relying too much on the categories with higher soft label values of the teacher model; and use the calculated entropy to perform a normalization process with the maximum value of the denominator to uniformly standardize the entropy differences of all categories, so as to calculate a weighting factor for the unsafe behavior categories, enabling the student model to adjust the learning focus for each category according to the weight size of the category.

[0069] The calculation formula for the entropy is:

[0070] H P,i =-p i log2p i

[0071] H Q,i =-q i log2q i

[0072] Among them, H P,i represents the entropy of the teacher network for each category, H Q,i represents the entropy of the student network for each category, p i and q i respectively represent the soft label values of the teacher network and the student network for each category.

[0073] The calculation formula for the weighting factor is:

[0074]

[0075] Among them, w i represents the weight factor assigned to each category, and α is a hyperparameter used to adjust the influence degree of the soft label on the weight.

[0076] For example, the soft label values calculated according to the teacher network: [0.0058, 0.0599, 0.3170, 0.6174], then the entropy of each unsafe behavior category is expressed as:

[0077] H P,0 = -0.0058 log2 0.0058 = 0.0431

[0078] H P,1 = -0.0599 log2 0.0599 = 0.2433

[0079] H P,2 = -0.3170 log2 0.3170 = 0.5254H P,3 = -0.6174 log2 0.6174 = 0.4295 After calculation, the entropy of each category of the teacher network is [0.0431, 0.2433, 0.5254, 0.4295]; the soft label values calculated according to the student network: [0.0006, 0.0171, 0.0466, 0.9357], then the entropy of each unsafe behavior category is expressed as:

[0080] H Q,0 = -0.0006 log2 0.0006 = 0.0064

[0081] H Q,1 = -0.0171 log2 0.0171 = 0.1004

[0082] H Q,2 = -0.0466 log2 0.0466 = 0.2061

[0083] H Q,3 = -0.9357 log2 0.9357 = 0.0897

[0084] After calculation, the entropy of each category of the student network is [0.0064, 0.1004, 0.2061, 0.0897]; according to the calculation formula of entropy difference, using the hyperparameter α = 1 and the entropies calculated by the teacher network and the student network to calculate the weights of each category, the weights of each unsafe behavior category are expressed as:

[0085]

[0086] After calculation, the weight values of each unsafe behavior category are [1.11, 1.42, 1.94, 2], indicating that the student model pays more attention to the unsafe behaviors of sleeping on duty and leaving the post.

[0087] Step 3.5: Combine the KL divergence with the weight factor of entropy difference to form an adaptive KL divergence. Monitor the difference between the student model and the teacher model (i.e., Soft loss) through the adaptive KL divergence to ensure that the student model retains the knowledge of the old tasks while learning new tasks and avoids forgetting. The calculation formula of the adaptive KL divergence is as follows:

[0088]

[0089] where w i represents the weight size assigned to each category, p i represents the soft label value of the teacher network, and q i represents the soft label value of the student network.

[0090] For example: According to the weight sizes of the above-mentioned unsafe behavior categories and the soft labels of the teacher network and the student network, the result of the distillation loss of the new training set in the student network is as follows:

[0091]

[0092] Step 3.6: After the logit value output by the student network is not divided by the temperature parameter and passed through the softmax function operation, extract the true label of each category, that is, the hard label value, and represent it in the form of a One-hot vector, such as [1, 0, 0, 0]. Calculate the loss according to the hard label value using the classification and regression loss in the student network, that is, Hard loss;

[0093] Step 3.7: Combine the Soft loss and the Hard loss together to form the total loss in the training process of the student network; Use the number of images in the new training set and the batch size of the training parameters to form an adaptive weight coefficient λ to flexibly balance the Soft loss and the Hard loss, so that the student model can gradually transition from relying on soft labels to relying on hard labels in the training strategy.

[0094] The calculation formula of the weight coefficient is as follows:

[0095]

[0096] where i represents the index of the current training round, its value ranges from 0 to N - 1, and N represents the number of training batches, and its value is

[0097] For example, if the number of images is set to 100 and the batch size is 32, then the initial weight coefficient λ in the first round of training is 0.7911;

[0098] Step 3.8: Save the object detection model constructed after training the new training set through the student network in the model folder set on the enterprise server, replacing the original basic object detection model.

[0099] This embodiment can continuously improve the model incrementally, better solve the problem of insufficient model detection ability in variable scenarios, and thus improve the recognition effect of employees' unsafe behaviors in variable scenarios.

[0100] Embodiment 2

[0101] As Figure 3 shown, in this embodiment, according to the object detection model improvement method based on incremental learning described in Embodiment 1, an unsafe behavior video post-recognition system architecture with coexistence of front-end and back-end is constructed, including:

[0102] Front-end module: Utilize multiple existing enterprise unsafe behavior video recognition systems to undertake tasks with high resource consumption and high technical complexity. Real-time capture the behaviors of on-site employees through various acquisition devices, use the models in the existing enterprise recognition systems to identify employees' behaviors, and store the identified employees' unsafe behavior information in the databases of their respective recognition systems.

[0103] For example, the enterprise monitors the main code stream of employees' behaviors through cameras, uses the employees' unsafe behavior models in the existing recognition systems to identify the obtained main code stream, and saves information such as the types of identified employees' unsafe behaviors (such as not wearing a safety helmet, not wearing work clothes, etc.), the occurrence location, the occurrence time, the predicted picture address of the Http interface, and the employees' unsafe behavior video stream within 5 seconds before and after the occurrence time of the Http interface.

[0104] Data transmission module: Deploy an open-source DBAPI low-code tool, use it to connect to the multi-source databases of the existing enterprise recognition systems, and obtain the Http interfaces containing alarm information in the multi-source databases. If there are already Http interfaces containing alarm information in the existing enterprise recognition systems, there is no need to create them in the DBAPI low-code tool. Then, define a scheduled task to run the obtained Http interfaces and store the alarm information uniformly in a specified table in MySQL.

[0105] Specifically, first, in the DBAPI low-code tool, select the database type of the enterprise's existing recognition system, such as MySQL, Oracle, etc.; then, write SQL statements to obtain recognition information from the table, such as the types of employees' unsafe behaviors (action_code), occurrence locations (action_site), occurrence times (action_time), the predicted picture address of the Http interface (action_image_path), the video stream of employees' unsafe behaviors within 5 seconds before and after the occurrence time of the Http interface (action_video_path). After saving, generate individual Http interfaces; finally, create fields for the corresponding information in the MySQL table. At the same time, in the scheduled task, form an array with the generated Http interfaces and the existing Http interfaces in the enterprise's existing recognition system, and define an object containing recognition information. Loop through the array, run the Http interfaces according to the Http protocol, save the recognition information in the object, and write SQL statements to add the object to the specified table in MySQL.

[0106] Backend module: It is composed of the content included in the incremental improvement method of the target detection model and is used to undertake tasks with low resource consumption and low technical complexity. Use the constructed basic target detection model to perform secondary recognition on the unsafe behaviors in the table after data transmission, and run the human-computer interaction mechanism and the model timing increment mechanism improvement operations in the post-recognition method of employees' unsafe behaviors based on incremental learning.

[0107] Specifically, add two fields to the specified MySQL table. One is the result of secondary recognition, is_match, such as 0: it is an employee's unsafe behavior; 1: it is not an employee's unsafe behavior. The other is the predicted picture address after secondary recognition, save_image_path, that is, save the predicted picture in the enterprise's existing Minio object storage server and generate the specified Http interface. Obtain the data of three fields, the primary key id, occurrence time, and unsafe behavior video stream, from the MySQL table. Intercept key frame images from the video stream according to the occurrence time, use the basic target detection model to identify the images, save the recognition results to the field is_match corresponding to the primary key id, and upload the predicted images to the Minio server. The generated Http interface is saved to the field save_image_path corresponding to the id.

[0108] Visualization interface: Display the information related to the unsafe behaviors secondary-recognized by the backend module in an easy-to-understand list form, that is, the types of employees' unsafe behaviors, occurrence locations, occurrence times, results of secondary recognition, and predicted pictures after secondary recognition in the MySQL database table. Use the manual review operation to return the review results of the secondary recognition to the backend module in the form of parameters.

[0109] Specifically, the backend uses the SpringBoot and MyBatis-Plus frameworks. Among them, SpringBoot provides data interfaces for the visualization interface, that is, it provides data by constructing RESTful APIs; MyBatis-Plus simplifies the interaction with the database and provides functions such as CRUD operations, paging queries, and condition constructors, making the development process more efficient and convenient; the visualization interface is developed using the Vue2 framework and the Element-UI component library, and displays the information related to the unsafe behaviors identified secondly provided by the backend, that is, the types of employees' unsafe behaviors, the occurrence locations, the occurrence times, the results of the secondary identification, and the predicted pictures after the secondary identification in the MySQL database table; at the same time, two buttons are developed for manual review operations, and the review results of the enterprise personnel on the secondary identification are returned to the backend in the form of parameters; the button names are correct and wrong.

[0110] The above schematically describes the present invention and its implementation manners. This description is not restrictive, and only one of the implementation manners of the present invention is shown in the drawings. The actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention, they should all fall within the protection scope of the present invention.

Claims

1. A method for post-identification of employee unsafe behaviors based on incremental learning, characterized in that: The following steps are involved: Step 1: Build a basic target detection model Extract unsafe behavior video data from surveillance video streams, annotate and generate text files containing behavior categories and bounding boxes, build a dataset using images and corresponding text files, and divide them into training sets, validation sets, and test sets. Use the YOLOv8 model to train and obtain a basic target detection model. Step 2: Labeling of manual interaction data Receive manual review feedback. When the model recognition result is marked as inaccurate, capture multiple consecutive frames of images associated with the unsafe behavior, pre-label the images using a semi-automatic labeling tool, and automatically detect the images using the basic object detection model. Then, manually correct the position and category of the detection box and store them in the incremental data set. Step 3: Regular incremental improvement of the model The data volume of the incremental data set is counted regularly. When the data volume reaches the set threshold, the basic target detection model is incrementally improved based on the adaptive KL divergence and weight, the model parameters are updated, and the basic target detection model is replaced.

2. According to claim 1, a method for post-identification of unsafe behaviors of employees based on incremental learning is characterized by: In the step of building the basic target detection model, the data set is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1, and the annotated text file contains the type, location, and size information of the unsafe behavior.

3. According to claim 1, a method for post-identification of unsafe behaviors of employees based on incremental learning is characterized by: In the manual interaction step, a manual interaction mechanism is created, which defines an object to receive feedback parameters for manual review in a visual interface, wherein the object attributes include a flag indicating whether the recognition result is accurate, the time when the unsafe behavior occurs, and a corresponding unsafe behavior video stream; The manual interaction mechanism uses the judgment of the signs to perform corresponding operations. For example, when the recognition result is marked as inaccurate, a total of 4 images before and after the unsafe behavior occurs are captured, and the image targets are automatically detected and manually fine-tuned through the annotation tool.

4. A method for post-identification of unsafe behaviors of employees based on incremental learning according to any one of claims 1 to 3, characterized in that: The model is incrementally improved periodically, and specifically comprises the following steps: 3.

1. Extract fine-tuning images and corresponding text files from the data folder after manual interaction to build a new training dataset; 3.

2. Use the saved basic target detection model to build the teacher network, use the YOLOv8 original model to build the student network, input the new training set into the teacher network and the student network, and output the original prediction score logit of the unsafe behavior category respectively; 3.

3. Divide the logit of the teacher network and the student network by the temperature parameter T, and extract the soft label value of the softened probability distribution of each unsafe behavior category after softmax operation; 3.

4. Based on the soft label values ​​of the teacher network and the student network, the entropy difference of each unsafe behavior category is calculated, and the weighting factor is generated by maximum normalization processing; 3.

5. Construct an adaptive KL divergence based on the weighting factor and KL divergence to calculate the Softloss between the teacher network and the student network; 3.

6. The logit output by the student network is subjected to softmax operation to generate hard label values, and the Hardloss is calculated by combining classification and regression losses; 3.

7. Generate an adaptive weight coefficient λ according to the number of images in the new training set and the batch size, and add the weighted sum of Soft loss and Hard loss to form the total loss; 3.

8. Use the total loss to train the student network and replace the original base model with the updated target detection model.

5. The method for post-identification of unsafe behaviors of employees based on incremental learning according to claim 4 is characterized by: The calculation formula for entropy in step 3.4 is: H P,i =-p i log2p i H Q,i =-q i log2q i Among them, H P,i represents the entropy of the teacher network on each category, H Q,i represents the entropy of the student network on each category, p i and q i Represent the soft label values ​​of the teacher network and the student network for each category respectively.

6. The method for post-identification of unsafe behaviors of employees based on incremental learning according to claim 5 is characterized by: The calculation formula of the weighting factor in step 3.4 is: Among them, w i Represents the weight factor assigned to each category, and α is a hyperparameter.

7. The method for post-identification of unsafe behaviors of employees based on incremental learning according to claim 6 is characterized by: The calculation formula of the adaptive weight coefficient λ in step 3.7 is: Here, i represents the index of the current training round, and N represents the number of training batches.

8. The method for post-identification of unsafe behaviors of employees based on incremental learning according to claim 7 is characterized by: The trigger condition for the regular incremental improvement of the model is the amount of newly annotated data in the incremental data set counted every hour. When the data amount exceeds the preset threshold, the basic target detection model is incrementally improved, otherwise no processing is performed.

9. An employee unsafe behavior post-identification system based on incremental learning, characterized in that: include: Front-end module: Utilize the company's existing unsafe behavior identification system to identify employee behavior in real time, and store the identified employee unsafe behavior information in the database of each identification system; Data transmission module: connect to the multi-source database of the enterprise's existing identification system through the DBAPI low-code tool, obtain alarm information regularly and store it uniformly in the specified database table; Back-end module: Based on the method described in any one of claims 1 to 8, the constructed basic target detection model is used to perform secondary recognition of unsafe behaviors, perform manual interactive labeling and model timed incremental training, and update the target detection model.

10. The employee unsafe behavior post-identification system based on incremental learning according to claim 9, characterized in that: The system sets up a visual interface, displays the relevant information of the unsafe behavior secondary identified by the back-end module in the form of a list, and uses manual review operation to return the review result of the secondary identification to the back-end module in the form of parameters.

Citation Information

Patent Citations

  • Safety production supervision method and system

    CN110414320A

  • Employee behavior identification method and device based on video understanding network, and storage medium

    CN118470787A