Student Behavior Recognition Method Based on Improved YOLOv5
The improved YOLOv5 model addresses computational and cost challenges in student behavior recognition by leveraging Ghost bottlenecks and spatial-channel attention, achieving high accuracy and reduced computational load on edge devices.
Patent Information
- Application Number
- CN202211644613.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-20
AI Technical Summary
The existing student behavior recognition methods have a large amount of computing power on edge devices, resulting in high computing resources occupancy, inability to effectively process large amounts of student video data, and high model deployment costs.
Using the improved YOLOv5 network architecture YOLOv5s-A-SG, the combination of Ghost bottleneck structure, H-Swish activation function, feature pyramid and path aggregation network reduces the computational amount and improves recognition accuracy, and uses the space-channel attention module to enhance detection performance.
Efficient student behavior recognition is achieved on edge devices, with an identification accuracy of 83.2%, reducing computing resource usage and deployment costs and improving detection speed.
Smart Images

Figure CN115830392B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to a student behavior recognition method based on object detection. Background Art
[0002] As an indispensable part of current classroom teaching, teaching evaluation is an activity that judges the value of the teaching process and results based on teaching objectives and serves teaching decision-making. For teachers, teaching evaluation can effectively discover the deficiencies in their teaching process. For students, teaching evaluation can help them better identify their weaknesses and make up for deficiencies. The classroom is a place for direct communication between teachers and students. It is of great research significance for schools to pay attention to the learning behaviors of students in the classroom. On the one hand, in daily class activities, the state of students in class reflects to a certain extent the teaching quality of teachers and the concentration of students in class. On the other hand, research shows that the classroom behaviors of students are easily affected by external factors. As the only supervisor of students during class, teachers focus more on educational teaching activities and cannot always pay attention to the class status of each student. The cameras in traditional classrooms often only play a role in security monitoring. However, there is a large amount of information that can be mined in class videos. From class videos, the class attendance and the performance of each student in class can be observed. If the classroom behavior information of students can be automatically detected through class videos, it can integrate information technology with education and teaching and provide support for teaching services.
[0003] Existing student behavior recognition methods often use deep learning technology to automatically learn features from training data and perform student behavior recognition. However, for deep learning, the larger the computational complexity of the model, the more computing resources it occupies. The model needs to be deployed on a cloud server. The huge amount of student video data generated by terminal devices is transmitted to the cloud server through the network, which will occupy a large amount of bandwidth. Therefore, it is necessary to implement edge computing processing for the student recognition model to achieve local processing of data. This can not only alleviate the latency problem but also reduce the deployment cost of the algorithm. However, current edge devices cannot effectively process a large amount of student videos because the algorithm performance of edge devices still needs to be improved. Therefore, the present invention will conduct research on a lightweight small-scale student classroom behavior recognition algorithm. Summary of the Invention
[0004] The present invention aims to break through the deficiencies of existing research and meet the needs of practical applications, and proposes a student behavior recognition algorithm based on improved YOLOv5, constructing a deep learning network architecture YOLOv5s-A-SG, which can recognize the behaviors of students in small-scale classrooms with fewer parameters.
[0005] 1. The technical solution provided by the present invention is: A student behavior recognition method based on improved YOLOv5, including the following steps:
[0006] S1. Collect data through classroom monitoring, and construct a student behavior detection dataset Stu-OA according to the defined student behavior categories.
[0007] S2. Input the prepared dataset into the improved network input module, which mainly performs a series of operations on the data, including Mosaic data augmentation, Focus processing, etc., to provide rich and convenient training data for model training.
[0008] S3. Further input the data processed by the input end into the backbone network. Considering the lightweight of the model, the Ghost module is adopted in the backbone network to greatly reduce the computational amount while ensuring the accuracy, and complete the feature extraction of the data.
[0009] S4. In the feature fusion network part, the extracted feature maps first transmit semantic information from top to bottom in the Feature Pyramid Network (FPN), and then transmit localization information from bottom to top in the Path Aggregation Network (PAN) to achieve feature fusion of different layers.
[0010] S5. Further process the feature maps obtained by the feature fusion network. In order to improve the feature extraction ability and student behavior recognition ability, the original convolutional layer is replaced with a spatial-channel attention module, and feature maps of different scales are obtained after the output of this module.
[0011] 2. In step S1, the produced student behavior detection dataset contains both the annotation information of students and the behavior information of students. The entire dataset contains a total of 84,264 samples. It includes 17,452 listening samples, 5,783 sleeping samples, 3,875 standing samples, 32,351 mobile phone playing samples, 4,327 writing samples, 14,583 reading samples and 5,893 raising hand samples. This dataset contains many objects with different proportions, ranging from 50×50 pixels to 300×300 pixels. Among them, nearly 80% of the objects in the dataset occupy less than 5% of the entire image.
[0012] 3. In step S2, use Mosaic data augmentation on the dataset to increase the number of samples so as to improve the ability to detect small targets, use adaptive anchor boxes to adapt to different datasets, adopt adaptive image scaling to reduce the computational amount, and add a Focus module to speed up the inference speed.
[0013] 4. In step S3, the Ghost bottleneck structure is used to replace the original CSP structure. The CSP structure contains a large number of convolutional modules, which can improve the learning ability of the model and effectively extract student features. However, the CSP structure uses unnecessary parameters during feature extraction, resulting in redundant feature extraction. The Ghost bottleneck structure uses Ghost modules and introduces a bottleneck structure similar to MobileNetv2 to build a new basic module. This bottleneck structure can avoid information loss caused by compression when information is converted between different dimensions. To address the problem of excessive parameters caused by redundant feature maps, the Ghost module divides each convolutional layer into two parts: the first part uses ordinary convolution to generate a part of the feature map, and the second part uses linear transformation to generate a part of the feature map, ultimately achieving the effect of generating more feature maps with fewer parameters.
[0014] To keep the number of connected channels consistent, the Ghost bottleneck structure uses two Ghost modules: the first Ghost module first performs dimension elevation to increase the depth of the network, which is beneficial for feature extraction; the second Ghost module performs dimension reduction to restore the original number of channels. When Stride (convolutional stride) = 2, depthwise separable convolution is added after the dimension elevation of the Ghost module, and depthwise separable convolution and point convolution are added to the branch connection, which can significantly reduce the computational amount.
[0015] Meanwhile, to bring an obvious performance improvement to the network model, using an activation function can solve the problem of insufficient expressive ability of linear functions. The original Ghost used the ReLU activation function. To simplify the network and better transmit information to the next layer, the H-Swish activation function improved by Swish, which has a faster inference speed, is selected. The H-Swish function suppresses the feature values less than -3 of the input value, linearly outputs the values greater than 3, and non-linearly transforms and outputs the remaining values.
[0016] 5. In step S4, the feature fusion network part adopts the structure of Feature Pyramid Network (FPN) plus Path Aggregation Network (PAN). The concept of FPN is to enhance the feature fusion of different layers and perform predictions at multiple scales. PAN adds bottom-up fusion on the basis of FPN. Deep feature maps carry stronger semantic features and weaker localization information, while shallow feature maps carry stronger position information and weaker semantic features. FPN transmits the deep semantic features to the shallow layer to enhance the semantic expression at multiple scales, while PAN, on the contrary, transmits the shallow localization information to the deep layer to enhance the localization ability at multiple scales.
[0017] 6. In step S5, for the output end, a spatial-channel attention fusion module is introduced to replace the original convolutional module. The spatial attention module fuses features of different scales according to semantic importance. The channel attention module uses deformable convolution to learn sparse features and performs fusion at the same spatial position. Finally, the outputs of the two attention modules are fused.
[0018] Input, backbone network feature map F in ∈R L×H×W . Where L is the number of feature layers of different scales in the feature pyramid feature map, and H, W, and C are the width, height, and number of channels of the feature map respectively.
[0019] Output, feature map F after attention mechanism out ∈R L×H×W .
[0020] Step1. Input the feature map into the channel attention module, and the channel attention module acts on the feature maps of different channels. Since the feature maps of different channels correspond to targets of different scales in the target, the fusion of features of different scales can obtain different semantic information. Therefore, the features extracted from the feature maps of different channels are also enhanced. The construction of the channel attention module is defined as D L .
[0021]
[0022] Where F is the input feature map, S is the area of a certain channel feature map, C is the number of channels of the entire feature map, f(·) is equivalent to a 1×1 convolution, and σ(x) is a hard-sigmod function.
[0023] Step2. Then input it into the spatial attention module. The spatial attention module learns semantic information in the spatial dimension. The spatial attention module first uses deformable convolution to make the attention learn sparse features, and then aggregates the information at the same spatial position. The construction of the spatial attention module is defined as D S .
[0024]
[0025] Where F l is the output result of the channel attention module, L is the number of channels of the output feature map of the channel attention module, K is the number of sparse sampling positions, and pk + Δpk is the offset of the region of interest. Δmk is the learning factor of region pk, and both are learned from the input features of the intermediate layer. Wl,k represents the convolution kernel of the deformable convolution at the l-th layer channel and k sparse sampling positions.
[0026] The output end includes three different detection parts for detecting the target. The output of each detection layer includes the coordinate information of the target, the object score, and the category to which it belongs. The loss function adopts GIoU (Generalized Intersection over Union) Loss, the object confidence loss function, and the binary cross-entropy classification loss function. Since the IoU Loss cannot directly optimize the situation where the targets do not overlap and cannot recognize various alignment methods when used as the loss function. To solve the problems existing in the above IoU Loss function, the loss function is improved to GIoU Loss. First, a minimum rectangle C that can contain both the true box A and the predicted box B is obtained, and then the ratio of the remaining area in C to the total area of C is calculated, and this ratio is brought into the formula to form the GIoU Loss.
[0027]
[0028] Loss giou = 1 - GIoU
[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0030] 1. The present invention innovatively combines the spatial-channel attention fusion module with the YOLOv5 model, and proposes an improved method for recognizing classroom behaviors of YOLOv5 to enhance the detection performance of the output end.
[0031] 2. The present invention utilizes the Ghost bottleneck structure, enabling the model to reduce the parameters of the model without degrading the performance and improving the detection speed.
[0032] 3. The present invention utilizes the H-Swish activation function to solve the problem of insufficient expression ability of the linear activation function and speeds up the inference speed.
[0033] 4. The method for recognizing student behaviors based on the improved YOLOv5 constructed by the present invention realizes and improves the accuracy on the self-made student behavior dataset, achieving an identification accuracy of 83.2%. Description of the Drawings
[0034] Figure 1 is the overall architecture of the YOLOv5s-A-SG network;
[0035] Figure 2 is the structure diagram of the Ghost bottleneck;
[0036] Figure 3 is the structure diagram of the spatial-channel attention module;
[0037] Figure 4 is the result display of the method for recognizing student behaviors. Detailed implementation manners
[0038] The present invention will be further described below in conjunction with specific embodiments.
[0039] The student behavior recognition method based on YOLOv5 proposed in this embodiment is a lightweight student behavior recognition network. As Figure 1 shown, the improved model of the present invention is named YOLOv5s-A-SG. YOLOv5s-A-SG mainly includes four parts, namely an input module (Input), a backbone network (Backbone), a feature fusion network (Neck) and an output end (Head). First is the input module. For the video frames obtained after preprocessing, this module uniformly scales them into images of 640×640×3; then is the backbone network, including a Focus module and a Ghost module. The Focus module slices the input image into four pieces and then splices them on the channels. The Ghost module is obtained by replacing the original CSP module, which greatly reduces the amount of calculation while ensuring the accuracy; then is the feature fusion network. This part first transmits semantic information from top to bottom in the Feature Pyramid Network (FPN) for the extracted feature maps, and then transmits location information from bottom to top in the Path Aggregation Network (PAN) to achieve feature fusion of different layers; finally is the output end. In order to improve the feature extraction ability and student behavior detection ability, the present invention replaces the original convolutional layer with a spatial-channel attention module. After being output by this module, feature maps of 20×20, 40×40, and 80×80 are obtained respectively.
[0040] The specific implementation of the student behavior recognition algorithm based on the improved YOLOv5 in this embodiment is as follows:
[0041] 1) Network architecture design
[0042] First, after preprocessing the input data at the input end, it is sent to the backbone network composed of Ghost bottleneck structures to extract features. After aggregating the features through the feature fusion network, it is input to the output end with a spatial-channel attention fusion module added to obtain the student coordinates and behavior categories.
[0043] 2) Feature extraction of the backbone network based on the Ghost bottleneck structure
[0044] The feature extraction of the backbone network based on the Ghost bottleneck structure uses the Ghost bottleneck structure to replace the original CSP structure. The CSP structure contains a large number of convolutional modules, which can improve the learning ability of the model and effectively extract student features. However, the CSP structure uses unnecessary parameters during feature extraction, resulting in redundant feature extraction. The Ghost bottleneck structure uses Ghost modules and introduces a bottleneck structure similar to MobileNetv2 to build a new basic module. This bottleneck structure can avoid information loss caused by compression when information is converted between different dimensions. To address the problem of excessive parameters caused by redundant feature maps, the Ghost module divides each convolutional layer into two parts: the first part uses ordinary convolution to generate a part of the feature map, and the second part uses linear transformation to generate a part of the feature map, ultimately achieving the effect of generating more feature maps with fewer parameters.
[0045] To keep the number of connected channels consistent, the Ghost bottleneck structure uses two Ghost modules: the first Ghost module first performs dimension elevation to increase the depth of the network, which is beneficial for feature extraction; the second Ghost module performs dimension reduction to restore the original number of channels. When Stride (convolution step) = 2, depthwise separable convolution is added after the dimension elevation of the Ghost module, and depthwise separable convolution and point convolution are added in the branch connection, which can significantly reduce the computational amount.
[0046] At the same time, to bring an obvious performance improvement to the network model, using an activation function can solve the problem of insufficient expressive ability of linear functions. The original Ghost used the ReLU activation function. To simplify the network and better transmit information to the next layer, the H-Swish activation function improved by Swish, which has a faster inference speed, is selected. The H-Swish function suppresses the feature values less than -3 of the input value, linearly outputs the values greater than 3, and non-linearly transforms and outputs the remaining values.
[0047] 3) Output end based on spatial-channel attention fusion
[0048] As Figure 3 shown, the spatial-channel attention fusion module is composed of a spatial attention module and a channel attention module.
[0049] The output end includes three different detection parts for detecting the target. The output of each detection layer includes the coordinate information of the target, the object score, and the category to which it belongs. The loss function uses GIoU (Generalized Intersection over Union) Loss, the object confidence loss function, and the binary cross-entropy classification loss function. Since the IoU Loss cannot directly optimize the case where the targets do not overlap and cannot recognize various alignment methods when used as the loss function. To solve the problems existing in the above IoU Loss function, the loss function is improved to GIoU Loss. First, a minimum rectangle C that can contain both the ground truth box A and the predicted box B is obtained, and then the ratio of the area of the remaining part in C to the total area of C is calculated, and this ratio is brought into the formula to form the GIoU Loss.
[0050]
[0051] Loss giou = 1 - GIoU
[0052] 4) Experimental configuration
[0053] Hardware configuration: All experiments in this chapter are run on a server configured with an Intel Xeon Gold 5218 CPU with a main frequency of 2.30 GHz and an NVIDIA RTX 6000 24-GB GPU, using the Ubuntu 18.04 operating system, Python 3.6, the development environment is VScode, and the deep learning framework is Pytorch version 1.9.0 and CUDA 10.2.
[0054] Software configuration: The data input size is 640 * 640. The SGD algorithm is used as the optimizer, the initial learning rate is set to 0.01, the momentum is 0.937, and the batch size is 128. During the training process, the number of epochs is set to 300.
[0055] 3) Dataset selection
[0056] The present invention selects the data collected from real classrooms to make a student behavior dataset. This dataset contains both the annotation information of students and the behavior information of students. The entire dataset contains a total of 84,264 samples. It includes 17,452 listening samples, 5,783 sleeping samples, 3,875 standing samples, 32,351 mobile phone playing samples, 4,327 writing samples, 14,583 reading samples, and 5,893 raising hand samples. This dataset contains many objects with different proportions, ranging from 50×50 pixels to 300×300 pixels. Among them, nearly 80% of the objects in the dataset occupy less than 5% of the entire image. Finally, the training set, test set, and validation set are divided according to the ratio of 8:1:1.
[0057] 4) Performance Analysis
[0058] By analyzing the results in Table 1, it is known that using the H-Swish function significantly improves the model performance. The detection speed of YOLOv5s-A-S can reach 62 FPS, the Precision is also increased by 1.3%, and the Recall is improved by 0.3%. YOLOv5s-A-SG, which replaces the convolutional module of the original model, has an mAP improvement of 1.4% compared to the original model. In addition, YOLObv5s-A-SG has a great advantage in terms of the number of parameters compared to the other two models, and its performance is also improved. Generally speaking, YOLOv5s-A-SG maintains high performance while being lightweight.
[0059] Table 1 Comparison of the improvement results of activation functions
[0060] Model Activation function P(%) R(%) mAP (%) FPS Number of parameters (M) YOLOv5s-A LeakyReLU 77.3 80.7 82.2 60.5 7.5 YOLOv5s-A-S H-Swish 78.6 81.0 82.8 62.2 7.3 YOLOv5s-A-SG H-Swish 80.3 82.1 84.2 65.7 5.1
[0061] To further verify the performance of the improved model, in this section, the improved model YOLOv5s-A-SG of YOLOv5s is compared with YOLORs, PP-YOLOEs, and YOLOXs to observe their performance on the Stu-OA dataset of students' classroom behavior. As shown in Table 2, the experimental results prove that YOLORs and PP-YOLOEs have higher recall rates, but lower accuracies and too many model parameters. The improved method in this chapter has an mAP index 2.4% higher than YOLORs and 2.9% higher than PP-YOLOE-s, and the number of model parameters is only 1 / 2 of the above two models. Through the analysis of the above experimental results, the effectiveness of the algorithm of the present invention is proved.
[0062] Table 2 Performance comparison of different algorithms
[0063] Method Parameters (M) P(%) R(%) mAP YOLORs 9.0 76.2 81.6 80.8 PP-YOLOEs 8.0 75.1 82.5 80.3 YOLOXs 9.0 74.4 76.5 76.7 YOLOv5s 7.2 82.5 80.3 78.8 YOLOv5s-A-SG 5.1 84.2 82.1 83.2
[0064] The above-described embodiments are only the preferred embodiments of the present invention, and do not limit the scope of implementation of the present invention. Therefore, any changes made according to the shape and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. An improved YOLOv5-based student behavior recognition method, characterized in that, It includes the following steps: S1. Collect image data through classroom monitoring and construct a student behavior detection dataset Stu-OA according to the defined student behavior categories. S2. Input the made dataset into an improved network input module, which mainly performs a series of operations on the data. S3. Further input the data processed by the input end into the backbone network. Considering the lightweight of the model, the Ghost module is adopted in the backbone network to complete the feature extraction of the data. S4. In the feature fusion network part, the extracted feature maps first transmit semantic information from top to bottom in the feature pyramid, and then transmit location information from bottom to top in the path aggregation network, realizing the feature fusion of different layers. S5. Further process the feature maps obtained through the feature fusion network. In order to improve the feature extraction ability and student behavior recognition ability, the original convolutional layer is replaced with a spatial-channel attention module. After the output of this module, feature maps of different scales are obtained respectively. In step S5, for the output end, a spatial-channel attention fusion module is introduced to replace the original convolutional module. The spatial attention module fuses feature maps of different scales according to semantic importance. The channel attention module uses deformable convolution to learn sparse features and performs fusion at the same spatial position. Finally, the outputs of the two attention modules are fused. Input, backbone network feature map F in R N×H×W , where N is the number of feature layers with different scales of the feature pyramid feature map, and H, W, and C are the width, height, and number of channels of the feature map, respectively; Output, the feature map F after the attention mechanism out R N×H×W ; Step1. Input the feature map into the channel attention module. The channel attention module acts on the feature maps of different channels. Since the feature maps of different channels correspond to targets of different scales in the target, the fusion of features of different scales can obtain different semantic information. Therefore, the features extracted from the feature maps of different channels are also enhanced. The construction of the channel attention module is defined as D L ; Where F is the input feature map, S is the area of a certain channel feature map respectively, C is the number of channels of the entire feature map, f(·) is equivalent to a 1×1 convolution, and σ(x) is a hard-sigmod function. Step2. Then it is input into the spatial attention module. The spatial attention module learns semantic information in the spatial dimension. First, the deformable convolution is used in the spatial attention module to enable the attention to learn sparse features, and then the information at the same spatial position is aggregated. The construction of the spatial attention module is defined as D S ; Among which F l is the result output by the channel attention module, L is the number of channels of the feature map output by the channel attention module, K is the number of sparse sampling positions, and p k +Δp k is the offset of the region of interest, and Δm k is the learning factor of region p k , and both are learned from the input features of the intermediate layer. W l,k represents the convolutional kernel of the deformable convolution in the first layer channels at k sparse sampling positions.
2. The student behavior recognition method based on the improved YOLOv5 according to claim 1, wherein: In step S1, the made student behavior detection dataset contains both the annotation information of students and the behavior information of students. The entire dataset contains a total of 84,264 samples, including 17,452 listening samples, 5,783 sleeping samples, 3,875 standing samples, 32,351 mobile phone playing samples, 4,327 writing samples, 14,583 reading samples and 5,893 raising hand samples. This dataset contains many objects with different proportions, ranging from 50×50 pixels to 300×300 pixels. Among them, nearly 80% of the objects in the dataset only occupy less than 5% of the entire image.
3. The student behavior recognition method based on the improved YOLOv5 according to claim 1, characterized in that: In step S2, Mosaic data augmentation is used on the dataset to increase the number of samples so as to improve the ability to detect small targets. Adaptive anchor boxes are used to adapt to different datasets. Adaptive image scaling is adopted to reduce the calculation amount. The Focus module is added to speed up the inference speed.
4. The student behavior recognition method based on the improved YOLOv5 according to claim 1, characterized in that: In step S3, the Ghost bottleneck structure is used to replace the original CSP structure. The CSP structure contains a large number of convolutional modules, which can improve the learning ability of the model and effectively extract student features. However, the CSP structure uses unnecessary parameters during feature extraction, resulting in redundant feature extraction. The Ghost bottleneck structure uses Ghost modules and introduces a bottleneck structure similar to MobileNetv2 to build a new basic module. This bottleneck structure can avoid information loss caused by compression when information is transformed between different dimensions. To address the problem of excessive parameters caused by redundant feature maps, the Ghost module divides each convolutional layer into two parts: the first part uses ordinary convolution to generate a part of the feature map, and the second part uses linear transformation to generate a part of the feature map, ultimately achieving the effect of generating more feature maps with fewer parameters. To keep the number of connected channels consistent, the Ghost bottleneck structure uses two Ghost modules: the first Ghost module first performs dimension elevation to increase the depth of the network, which is beneficial for feature extraction; the second Ghost module performs dimension reduction to restore the original number of channels. When Stride = 2, that is, the convolutional stride = 2, depthwise separable convolution is added after the Ghost module elevates the dimension, and depthwise separable convolution and point convolution are added in the branch connection, which can significantly reduce the computational amount. At the same time, to bring an obvious performance improvement to the network model, using an activation function can solve the problem of insufficient expressive ability of linear functions. The original Ghost used the ReLU activation function. To simplify the network and better transmit information to the next layer, the H-Swish activation function improved by Swish, which has a faster inference speed, is selected. The H-Swish function suppresses the feature values less than -3 in the input, linearly outputs the values greater than 3, and non-linearly transforms and outputs the remaining values.
5. The student behavior recognition method based on the improved YOLOv5 according to claim 1, characterized in that: In step S4, the feature fusion network part adopts the structure of Feature Pyramid Network (FPN) plus Path Aggregation Network (PAN). FPN+PAN draws on PANet from CVPR 2018, which was mainly applied in the field of image segmentation at that time, but Alexey split and applied it to Yolov5 to further improve the feature extraction ability. The concept of FPN is to enhance the feature fusion of different layers and perform predictions at multiple scales. PAN adds bottom-up fusion on the basis of FPN. Deep feature maps carry stronger semantic features and weaker localization information, while shallow feature maps carry stronger position information and weaker semantic features. FPN transmits the deep semantic features to the shallow layer to enhance the semantic expression at multiple scales, while PAN, on the contrary, transmits the shallow localization information to the deep layer to enhance the localization ability at multiple scales.
6. The method for recognizing student behavior based on improved YOLOv5 according to claim 1, characterized in that: The output end includes three different detection parts for detecting the target. The output of each detection layer includes the coordinate information of the target, the object score, and the category to which it belongs. The loss function uses GIoU (Generalized Intersection over Union) Loss, the object confidence loss function, and the binary cross-entropy classification loss function. Since the IoU Loss cannot directly optimize the case where the targets do not overlap and cannot recognize various alignment methods when used as the loss function, to solve the problems existing in the above IoU Loss function, the loss function is improved to GIoU Loss. First, a minimum bounding box R that can contain both the ground truth box A and the predicted box B is obtained, and then the ratio of the area of the remaining part in R to the total area is calculated. This ratio is substituted into the formula to form the GIoU Loss; Loss giou = 1 - GIou