Student illegal behavior detection method for vocational education training workshop

By improving the YOLOV8 model, combined with the C2f_DWR module, DySample, Deformable Attention and MPDIoU technology, the problem of low detection accuracy of students' violations in the vocational education training workshop is solved, and efficient and accurate detection of violations and strengthening teaching safety management is achieved.

CN120148104APending Publication Date: 2025-06-13ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510162479.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing technology has problems such as low accuracy, poor resource utilization, insufficient target positioning ability and poor teaching safety management in the detection of student violations in vocational education training workshops.

Method used

The improved YOLOV8 model is adopted, and the model's detection accuracy and efficiency in complex scenarios are improved by introducing the C2f_DWR module, DySample dynamic upsampling, Deformable Attention module and MPDIoU loss function, and the model's adaptability to changing environments is enhanced by combining multi-stage data acquisition and enhancement strategies.

Benefits of technology

It significantly improves the accuracy and efficiency of students' violation detection, can accurately identify violations under occlusion or light changes, and realizes the system's high customization and scene adaptability, strengthens teaching safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148104A_ABST
    Figure CN120148104A_ABST
Patent Text Reader

Abstract

The invention relates to a vocational education training workshop-oriented student illegal behavior detection method. The method comprises the steps of 1, shooting illegal behavior pictures of students in a training workshop and making a data set; step 2, classifying and labeling the illegal behavior data set after data enhancement by using labmeli; 3, dividing the data set marked in the step 2 into a training set, a verification set and a test set according to the proportion of 8: 1: 1; and 4, inputting the divided data set into an improved YOLOV8 network for training, accelerating convergence of a model network by using an improved module, and storing the trained model as a detection model. The method has remarkable advantages in the aspects of improving detection precision, optimizing resource utilization, improving target positioning capability and strengthening teaching safety management, and an innovative solution is provided for intelligent management and safety improvement of vocational education practical training teaching environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting students' violation behaviors in a vocational education training workshop, and particularly to a detection system for students' violation behaviors in a training workshop based on improved YOLOV8 model training and WEB technology, belonging to the technical field of intelligent safety auxiliary detection for vocational education. Background Art

[0002] The methods for detecting violation behaviors have been continuously evolving with the development of science and technology. In the early stage, the detection mainly relied on manual vision to inspect the violation behaviors in the workshop. This method has the problems of subjectivity and high human resource costs. With the improvement of automation, equipment detection has gradually replaced manual vision detection.

[0003] In recent years, with the rapid development of artificial intelligence technology, the detection methods for violation behaviors based on deep learning have gradually become a research hotspot. Deep learning algorithms can automatically extract features and perform efficient classification through the training of a large amount of data, so as to achieve more accurate detection in complex scenarios. For example, the application of object detection algorithms based on convolutional neural networks (CNNs) in video surveillance can accurately identify and judge in real time whether there are violation phenomena. Coupled with the high-speed development of Internet technology and the continuous improvement of software technology, combining deep learning algorithms and software can not only detect violation behaviors quickly and accurately, but also meet people's requirements for the convenience and portability of detection devices. Compared with the above-mentioned automated devices, the cost is greatly reduced, the safety can be guaranteed, and the requirements for the production environment are not too high. Summary of the Invention

[0004] The present invention provides a method for detecting students' violation behaviors in a vocational education training workshop, which shows significant advantages in improving detection accuracy, optimizing resource utilization, enhancing target positioning ability and strengthening teaching safety management, so as to solve the problems existing in the prior art.

[0005] A method for detecting students' violation behaviors in a vocational education training workshop includes the following steps: Step 1: Take pictures of students' violation behaviors in the training workshop and make a data set; the data set includes at least the first group of data, the second group of data and the third group of data, where the first group of data is the image data of students' violation behaviors obtained before the start of the training course; the second group of data is the image data of students' violation behaviors during the training course; the third group of data is the image data of students' violation behaviors after the end of the training course.

[0006] Step 2: Use labelimg to classify and label the data set of violation behaviors after data augmentation; Step 3: Divide the dataset marked in Step 2 into a training set, a validation set, and a test set according to the ratio of 8:1:1; Step 4: Input the divided dataset into the improved YOLOV8 network for training. The training epoch value is 300. Use the NVIDIA RTX 4090 GPU for accelerated training. The optimizer is AdamW, and the initial learning rate is 0.001. Use the improved module to accelerate the convergence of the model network. After the model is trained, save the model weight file best.pt as the core model for detecting violation behaviors. The student violation behavior detection system can customize and adjust the detection categories, select the detection model, whether to enable alarms, adjust the IOU threshold, and the detection location.

[0007] Introducing the combination of the DWR module and the original C2f can enhance the information fusion of the C2f module at different scales and improve the accuracy of object detection. At the same time, introducing the Deformable Attention module after the original SPPF module can improve the feature representation ability of the model and enhance the recognition ability for complex scenes and multi-scale objects. Modify the original Upsample in YOLOv8 to the Dy_Sample structure, reducing a large amount of computational burden and time. It realizes upsampling through point sampling instead of the traditional kernel method, which is simpler and more efficient. Replace the original regression loss function CIOU with MPDIOU to increase the convergence speed of the model and improve the performance of the model.

[0008] Step 5: Take pictures of the violation behaviors to be detected to obtain the pictures to be detected. Identify the categories of each violation behavior in the pictures to be detected through the saved model, and customize and adjust the detection-related parameters; Step 6: Warn and record the current violation behavior according to the number of categories of the violation behavior and the confidence score. Among them, the confidence score represents the similarity degree between this category and the standard category, and the higher the score, the closer it is.

[0009] The specific content of Step 1 includes: Use high-definition camera equipment to collect student operation scenario data in the training workshop. When collecting, select images at different time periods and positions to ensure data diversity. Perform data augmentation operations on the captured pictures, including image color inversion, image horizontal flipping, changing the brightness of the image, applying Gaussian blur to the original image, and applying affine transformation to change the position and scale of the image.

[0010] The specific content of Step 2 includes: The dataset is divided into four categories, namely the smoke class, the touch_power class, the use_phone class, and the violence class. Among them, the smoke class represents the behavior of students smoking in the training workshop; the touch_power class represents the behavior of students touching the power switch of the machine tool in the training workshop; the use_phone class represents the behavior of students using mobile phones in the training workshop; the violence class represents the behavior of students having fights and conflicts in the training workshop.

[0011] The improved YOLOV8 network in step 4 specifically includes: S4.1: DWR (Dilation-wise Residual) is an efficient multi-scale feature extraction method that decomposes the original single-step method into two steps: Region Residualization and Semantic Residualization. In the C2f module, first, a convolution is performed on the input feature map to expand the number of channels to twice the number of input channels. Then, features are gradually extracted through multiple Bottleneck modules. Each Bottleneck module contains multiple convolutional layers and can be configured to use or not use the shortcut connection (residual connection) as needed. However, when dealing with complex scenarios, the Bottleneck module may not be able to effectively capture the diversity of global and local features, thus limiting the expressive power of the model. To efficiently obtain multi-scale context information, we replace the Bottleneck module in the C2f module with the previously mentioned DWR module, as shown in the figure. This improved module is called the C2f_DWR module, which can enhance the information fusion ability of the C2f module at different scales, thereby improving the accuracy of object detection. In this way, YOLOv8 will perform better in capturing fine-grained information and global information in complex scenarios.

[0012] S4.2: DySample improves the resource utilization efficiency by bypassing dynamic convolution and adopting point sampling method for upsampling. In the initial YOLOv8 model, the UpSample method requires a large amount of computing resources and parameters, which limits the lightweight ability of the model in object detection. In practical applications, the initial images used to detect students' violations in the workshop are usually small and prone to pixel distortion, resulting in the loss of fine-grained information and posing challenges to feature learning. To solve these problems, DySample is introduced as an alternative to UpSample. DySample is a lightweight and efficient dynamic upsampler that performs upsampling by combining point sampling method with a learning-based sampling method. This method not only reduces the consumption of computing resources but also improves the image resolution without adding extra burden. Therefore, DySample improves the efficiency and performance of the model while reducing the computational cost. As shown in the figure, it demonstrates the sampling-based dynamic upsampling and module design in DySample.

[0013] S4.3: After adding Deformable Attention to the SPPF module of YOLOv8, the feature representation ability of the model can be improved, and the recognition ability for complex scenes and multi-scale targets can be enhanced. Deformable Attention can dynamically focus on key information, reduce the calculation of irrelevant regions, thereby improving the computational efficiency and detection accuracy. In addition, it can effectively capture long-range dependencies and improve the performance of the model in cases of occlusion or complex backgrounds. Therefore, this improvement helps to enhance the performance of YOLOv8 in multi-object detection tasks.

[0014] S4.4: The Minimum Partial Distance Intersection over Union (MPDIoU) is introduced to calculate the loss. The MPDIoU function aims to improve the accuracy of the model in defect localization by combining the traditional IoU evaluation framework with the concept of minimum positive distance. It not only considers the overlap degree between the predicted bounding box and the ground truth bounding box but also the distance between their centroids. This method can more finely adjust the performance of the model in the task of detecting students' violations, especially in the training workshop scenario with complex backgrounds and irregular defects. Using the MPDIoU loss function, we effectively solve the problem of inconsistent directions between the predicted bounding box and the ground truth bounding box, simplify the calculation process, and improve the efficiency and accuracy of detecting students' violations.

[0015] The specific steps of Step 5 include: In an actual workshop, a camera is used to capture real-time images of students' operations as the data to be detected. The captured images are input into an improved YOLOv8 model. The model outputs the detected violation behavior categories and their confidence levels in each image, and the confidence threshold for detection can be adjusted to reduce false alarms or missed detections.

[0016] The implementation method of the violation behavior detection system is as follows: The violation behavior detection system is built using Springboot and Vue. The real-time photos captured by the camera are input into the violation behavior detection model, and the detected categories and corresponding confidence scores are output. The confidence level and detection frame are retained on the detection result image, and then the image is encoded through base64 encoding and transmitted to the front end through Socket. The front end decodes the image through base64 encoding to obtain the image, which is displayed on the front-end interface of the violation behavior detection system, and the detection parameters such as the detected category and confidence score are saved in the database.

[0017] Step six specifically includes: Warning lights and buzzers are installed in the training workshop. When the system detects a violation behavior, a signal is sent to the warning device through the network, thereby triggering the device to give a warning. The detected violation behavior category and confidence level are used to determine the occurrence of students' violation behaviors, trigger the warning device in the training workshop to warn teachers and students, and record the results in the database.

[0018] The present invention has the following beneficial effects and advantages: By improving the YOLOV8 model structure (introducing the C2f_DWR module, Dysample dynamic upsampling, Deformable Attention module, and MPDIoU loss function), the present invention significantly improves the accuracy and efficiency of detecting students' violation behaviors in complex scenarios, and can accurately identify even in the case of occlusion or light changes. At the same time, through multi-stage data collection and enhancement strategies, the adaptability of the model to the changing environment of vocational education training workshops is enhanced. Combined with a flexible parameter adjustment function (such as detection categories, model selection, IOU threshold, etc.), the high customization and scene adaptation capabilities of the system are realized. In addition, the system comprehensively covers common violation behaviors such as smoking, touching the machine tool power switch, using mobile phones, and fighting, and has real-time warning and recording functions, providing technical support for managers to quickly respond and intervene in behaviors. Overall, the present invention shows significant advantages in improving detection accuracy, optimizing resource utilization, enhancing target positioning ability, and strengthening teaching safety management, providing an innovative solution for the intelligent management and safety improvement of vocational education training teaching environments. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0020] Figure 1 is the algorithm flow chart of the method for detecting students' illegal behaviors in the vocational education training workshop of the present invention; Figure 2 is the improved YOLOV8 algorithm framework diagram of the present invention; Figure 3 is the structure diagram of the smoke category of the present invention; Figure 4 is the structure diagram of the violence category of the present invention; Figure 5 is the structure diagram of the use_phone category of the present invention; Figure 6 is the structure diagram of the touch_power category of the present invention; Figure 7 is the structure diagram of the illegal behavior to be detected taken in Embodiment 1 of the present invention; Figure 8 is the recognition effect diagram of Embodiment 1 of the present invention.

[0021] Figure 9 is the effect diagram of the front-end page for detecting illegal behaviors of the present invention.

[0022] Figure 10 is the structure diagram of the improved 1 DWR module for YOLOV8 of the present invention.

[0023] Figure 11 is the structure diagram of the improved 2 Dysample module for YOLOV8 of the present invention.

[0024] Figure 12 is the structure diagram of the improved 3 Deformable Attention module for YOLOV8 of the present invention.

[0025] Figure 13 is the structure diagram of the improved 4 MPDIoU module for YOLOV8 of the present invention. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment Refer to Figures 1-13 , a method for detecting students' violation behaviors in a vocational education training workshop, comprising the following steps: Step 1: Take pictures of students' violation behaviors in the training workshop and make a data set; the data set includes at least the first group of data, the second group of data, and the third group of data, where the first group of data is the image data of students' violation behaviors obtained before the training course; the second group of data is the image data of students' violation behaviors during the training course; the third group of data is the image data of students' violation behaviors after the training course.

[0028] Use a high-definition camera device to collect students' operation scene data in the training workshop. For the following violation behaviors, smoking: students smoke in the workshop. Fighting conflict: students have physical conflicts while chasing and playing.

[0029] Touching the equipment power box: students privately turn on or touch the machine tool power box. Using mobile phones: students illegally use mobile phones in the training workshop. When collecting, select images at different time periods and positions to ensure data diversity and avoid excessive repetition of images.

[0030] Perform data enhancement operations on the captured pictures, including image color inversion, horizontal flipping of the image, changing the brightness of the image, applying Gaussian blur to the original image, and applying affine transformation to change the position and scale of the image for the pictures in the data set.

[0031] Step 2: Use labelimg to classify and label the data set of violation behaviors after data enhancement; Divide the data set into four categories, namely the smoke class, the touch_power class, the use_phone class, and the violence class; where the smoke class represents the behavior of students smoking in the training workshop; the touch_power class represents the behavior of students touching the machine tool power switch in the training workshop; the use_phone class represents the behavior of students using mobile phones in the training workshop; the violence class represents the behavior of students having fighting conflicts in the training workshop.

[0032] Step 3: Divide the data set labeled in Step 2 into a training set, a validation set, and a test set according to the ratio of 8:1:1; Step 4: Input the divided dataset into the improved YOLOV8 network for training. The training epoch value is 300. Accelerated training is performed using an NVIDIA RTX 4090 GPU. The optimizer is AdamW, and the initial learning rate is 0.001. The improved module is used to accelerate the convergence of the model network. After the model is trained, save the model weight file best.pt as the core model for detecting illegal behaviors. The student illegal behavior detection system can customize and adjust the detection categories, select the detection model, whether to enable alerts, adjust the IOU threshold, and the detection location.

[0033] The implementation method for improving the model is as follows: Reason for improvement 1: The complex background and occlusions in the workshop pose high requirements for object detection. DWR (Dilation-wise Residual) is an efficient multi-scale feature extraction method that decomposes the original single-step method into two steps: regional residualization and semantic residualization. Combining DWR and C2f can enhance the information fusion of the C2f module at different scales, improve the accuracy of object detection, and enable YOLOv8 to perform better in capturing fine-grained and global information in complex scenes.

[0034] Reason for improvement 2: Considering the deployment requirements of the model on mobile devices, the model needs to be lightweighted due to the limitations of device computing power and storage space. Dysample replaces the Upsample in YOLOv8. Dysample does not use complex dynamic convolution processing. It achieves upsampling through point sampling instead of the traditional kernel method, which is simpler and more efficient. Dysample requires fewer parameters and reduces the memory requirements of the GPU, which means it is more lightweight and runs faster.

[0035] Reason for improvement 3: The training workshop environment is complex and extensive, and the images of illegal actions are small. After adding DeformableAttention to the SPPF module, the feature representation ability of the model can be improved, which can dynamically focus on key information, reduce the calculation of irrelevant regions, effectively target small object detection, and thus improve the calculation efficiency and detection accuracy.

[0036] Reason for improvement 4: Illegal behaviors in the workshop may involve different postures (such as bending down to touch the machine, playing and fooling around, etc.), resulting in large changes in the shape of the target. Using MPDIoU to replace the loss function, MPDIoU minimizes the distance between the upper left and lower right points of the predicted box and the ground truth box, more accurately reflects the relative position between the boxes, and at the same time considers the similarity of the box shapes, improving the detection accuracy and the adaptability to irregular targets.

[0037] Through the above improvements, the DWR module enhances the information fusion ability of the backbone network at different scales. Enhance the complex scene detection ability: The DAT attention module can dynamically focus on key information and reduce the calculation of irrelevant regions. DySample reduces a large amount of computational burden and time, and is simpler and more efficient. MPDIoU can more accurately guide the model to optimize the position and shape of the bounding box.

[0038] Step Five: Take pictures of the violations to be detected to obtain the pictures to be detected, identify the categories of each violation in the pictures to be detected through the saved model, and customize and adjust the detection-related parameters; In an actual workshop, use a camera to take real-time pictures of students' operations as the data to be detected. Input the taken pictures into the improved YOLOv8 model. The model outputs the categories of violations detected in each picture and their confidence levels (such as "smoking, confidence level 0.92"). The user can adjust the confidence threshold for detection (such as the default 0.7) to reduce false positives or false negatives.

[0039] The implementation method of the violation detection system is as follows: Use Springboot and Vue to build a violation detection system. The real-time photos taken by the camera are input into the violation detection model, and the detection categories and corresponding confidence scores are output. The confidence level and detection box are retained on the detection result picture, and then the picture is encoded through base64 and transmitted to the front end through Socket. The front end decodes the picture through base64 encoding to obtain the picture, which is displayed on the front-end interface of the violation detection system, and the detection parameters such as the detection category and confidence score are saved in the database.

[0040] Step Six: Warn and record the current violations according to the number of violation categories and the confidence scores. Among them, the confidence score represents the similarity degree between this category and the standard category, and the higher the score, the closer it is.

[0041] The specific implementation method is as follows: Install warning lights and buzzers in the training workshop. When the system detects the occurrence of a violation, it sends a signal to the warning device through the network, thereby triggering the device to give a warning.

[0042] The improved YOLOV8 network in the fourth step specifically includes: S4.1: DWR (Dilation-wise Residual) is an efficient multi-scale feature extraction method that decomposes the original single-step method into two steps: Region Residualization and Semantic Residualization. In the C2f module, the input feature map is first convolved to expand the number of channels to twice the number of input channels, and then features are gradually extracted through multiple Bottleneck modules. Each Bottleneck module contains multiple convolutional layers, and can be configured to use shortcut connections (residual connections) as needed. However, when dealing with complex scenes, the Bottleneck module may not be able to effectively capture the diversity of global and local features, thereby limiting the expressive power of the model. In order to efficiently obtain multi-scale context information, we replace the Bottleneck module in the C2f module with the previously mentioned DWR module, as shown in the figure. This improved module is called the C2f_DWR module, which can enhance the information fusion capability of the C2f module at different scales, thereby improving the accuracy of target detection. In this way, YOLOv8 will perform better in capturing fine-grained and global information in complex scenes.

[0043] S4.2: DySample improves resource utilization efficiency by bypassing dynamic convolution and using a point sampling method for upsampling. In the initial YOLOv8 model, the UpSample method requires a lot of computing resources and parameters, which limits the model's lightweight ability in object detection. In practical applications, the initial images used to detect violations of workshop students are usually small and prone to pixel distortion, resulting in loss of fine-grained information, which poses challenges to feature learning. To address these issues, DySample is introduced as an alternative to UpSample. DySample is a lightweight and efficient dynamic upsampler that performs upsampling through a point sampling method combined with a learning sampling method. This method not only reduces the consumption of computing resources, but also improves image resolution without adding additional burden. Therefore, DySample improves the efficiency and performance of the model while reducing computing costs. The figure shows the sampling-based dynamic upsampling and module design in DySample.

[0044] S4.3: After adding Deformable Attention to the SPPF module of YOLOv8, the feature representation ability of the model can be improved, and the recognition ability for complex scenes and multi-scale targets can be enhanced. Deformable Attention can dynamically focus on key information, reduce the calculation of irrelevant regions, thereby improving the calculation efficiency and detection accuracy. In addition, it can effectively capture long-range dependencies and improve the performance of the model in cases of occlusion or complex backgrounds. Therefore, this improvement helps to enhance the performance of YOLOv8 in multi-object detection tasks.

[0045] S4.4: The Minimum Partial Distance Intersection over Union (MPDIoU) is introduced to calculate the loss. The MPDIoU function aims to improve the accuracy of the model in defect localization by combining the traditional IoU evaluation framework with the concept of the minimum positive distance. It not only considers the overlap degree between the predicted box and the ground truth box, but also the distance between their centroids. This method can more finely adjust the performance of the model in the student violation detection task, especially in the training workshop scenario with complex backgrounds and irregular defects. By using the MPDIoU loss function, we effectively solve the problem of inconsistent directions between the predicted box and the ground truth box, while simplifying the calculation process and improving the efficiency and accuracy of student violation detection.

[0046] The above has described the embodiments of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations made to these embodiments still fall within the protection scope of the present invention.

Claims

1. A method for detecting student violations in a vocational education training workshop, characterized by: The following steps are involved: Step 1: Take pictures of students’ illegal behaviors in the training workshop and create a data set; Step 2: Use labelimg to classify and label the enhanced violation dataset; Step 3: Divide the labeled data set in step 2 into training set, validation set and test set in a ratio of 8:1:1; Step 4: Input the divided data set into the improved YOLOV8 network for training, use the improved module to speed up the convergence of the model network, and save the model as a detection model after training; Step 5: Take a photo of the violation to be detected, obtain the image to be detected, identify the category of each violation in the image to be detected through the saved model, and customize the detection-related parameters; Step 6: Issue warnings and record current violations based on the number of violation categories and confidence scores.

2. A method for detecting student violations in a vocational education training workshop according to claim 1, characterized in that: The step 1 specifically includes: Use high-definition video equipment to collect student operation scene data in the training workshop. When collecting data, select images from different time periods and locations to ensure data diversity, perform data enhancement operations on the captured images, invert the image color, flip the image horizontally, change the image brightness, perform Gaussian blur processing on the original image, and apply affine transformation to change the image position and scale.

3. The method for detecting student violations in a vocational education training workshop according to claim 1 is characterized in that: The step 2 specifically includes: The data set is divided into four categories, namely smoke, touch_power, use_phone and violence. The smoke category represents the smoking behavior of students in the training workshop; the touch_power category represents the touching of the machine tool power switch by students in the training workshop; the use_phone category represents the using of mobile phones by students in the training workshop; and the violence category represents the fighting and conflicting behavior by students in the training workshop.

4. The method for detecting student violations in a vocational education training workshop according to claim 1 is characterized in that: The improved YOLOV8 network in step 4 specifically includes: S4.1: Decompose the original single-step method into two steps: regional residualization and semantic residualization; S4.2: Replace YOLOv8's Upsample with Dysample, which achieves upsampling through point sampling; S4.3: After adding the SPPF module, the feature representation capability of the model is improved, the key information is dynamically focused, the calculation of irrelevant areas is reduced, and the calculation efficiency and detection accuracy are improved for small target detection; S4.4: Use MPDIoU to replace the loss function. MPDIoU accurately reflects the relative position between the boxes by minimizing the distance between the upper left and lower right corners of the predicted box and the true box, while considering the similarity of the box shapes, thereby improving the detection accuracy and adaptability to irregular targets.

5. The method for detecting student violations in a vocational education training workshop according to claim 1 is characterized in that: The step five specifically includes: In the actual workshop, a camera is used to capture images of students’ operations in real time as the data to be detected. The captured images are input into the improved YOLOv8 model. The model outputs the category of violation detected in each image and its confidence. The confidence threshold of the detection can be adjusted to reduce false positives or negatives. The implementation method of the violation detection system is: Springboot and Vue are used to build a violation detection system. The real-time photos taken by the camera are input into the violation detection model, and the detection category and the corresponding confidence score are output. The confidence and detection box are retained on the detection result picture, and then the picture is encoded through base64 encoding and transmitted to the front end through Socket. The front end decodes the picture through base64 encoding and displays it on the front end interface of the violation detection system. The detection parameters such as the detection category and confidence score are saved in the database.

6. The method for detecting student violations in a vocational education training workshop according to claim 1 is characterized in that: The step six specifically includes: Warning lights and buzzers are installed in the training workshop. When the system detects violations, it sends signals to the early warning device through the network, which triggers the device to issue an early warning.