Drowning person identification method based on improved YOLOv8

By improving the YOLOv8 algorithm, combining the SPPF_ECA module and small object detection layer, the SS-YOLOv8 model is built, which solves the problem of traditional drowning search and rescue methods relying on manpower and existing algorithms to deploy difficulties, and realizes efficient and accurate identification of drowning personnel, which is suitable for drowning rescue in various scenarios.

CN120260077APending Publication Date: 2025-07-04JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339634.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional drowning search and rescue methods rely on manpower and are susceptible to fatigue and distraction, and cannot achieve all-weather automated monitoring. In addition, existing algorithms such as FasterR-CNN and MaskR-CNN models have large parameters and slow inference speed, making it difficult to deploy to edge devices.

Method used

Improve the YOLOv8 algorithm, combine the SPPF_ECA module and the small object detection layer, enhance the detection accuracy and robustness of the model for complex scenarios, improve the recognition ability of small objects, and train and optimize by building the SS-YOLOv8 model.

Benefits of technology

While ensuring high detection accuracy, the detection efficiency and adaptability of the model are improved, and drowning personnel can be accurately identified in complex environments. It is suitable for various scenarios such as swimming pools and oceans, providing support for drowning rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260077A_ABST
    Figure CN120260077A_ABST
Patent Text Reader

Abstract

The invention discloses a target identification method based on improved YOLOv8, and belongs to the field of computer vision. The method comprises the steps that a data set is acquired, a plurality of images with 640 * 640 pixels and corresponding label files are included on the basis of real drowning person photos disclosed on the network and frame images in drowning videos are intercepted to serve as a sample data set, and the label files are in a txt format; data preprocessing: preprocessing the collected drowning person data set, and dividing the data set into a training set, a verification set and a test set; a YOLOv8 model is improved; the YOLOv8 model is improved by using an ECA attention mechanism and a small target detection layer; model training: training the improved model by using a training set and a verification set; and model evaluation: specifically evaluating the recognition performance of the improved model by using the test set, continuously optimizing the model performance according to an evaluation result, and finally inputting a picture needing to be detected into the model for detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target recognition, and particularly to a drowning person recognition method based on improved YOLOv8. Background Art

[0002] Traditional drowning search and rescue methods usually rely on human resources. Through real-time observation of surveillance videos or on-site personnel, they are easily affected by factors such as fatigue and distraction, with a high miss detection rate and unable to achieve all-weather automated monitoring. In recent years, with the rise of computer vision technology, significant progress has been made in object detection in the field of image recognition, providing new solutions for identifying drowning victims. By using computer vision technology to process images, autonomous detection of drowning victims can be achieved, thus enabling rapid rescue operations. However, although algorithms such as FasterR-CNN and MaskR-CNN can improve detection accuracy, they have a large number of model parameters and slow inference speed, making it difficult to deploy to edge devices. With the emergence of the YOLO (You Only Look Once) series, with its good detection accuracy and speed, it is being widely used in the field of object detection. Summary of the Invention

[0003] Aiming at the problems of insufficient sensitivity to detailed features and insufficient generalization ability in complex scenarios existing in the YOLOv8 algorithm, the present invention proposes a drowning person recognition method based on improved YOLOv8. By using the SPPF_ECA module, important channel features related to the target are highlighted, and unimportant features are ignored, which helps the network to perform channel selection on features of different scales, improving the detection accuracy and robustness of the model in complex scenarios; at the same time, a small target detection layer is added for small targets such as the head and arms of drowning people to improve the capture and recognition ability of small targets and enhance the generalization ability of the model in different scenarios to meet the application requirements in real scenarios.

[0004] Technical Solution: A drowning person recognition method based on improved YOLOv8 includes the following steps:

[0005] Step 1, obtain a dataset of drowning people;

[0006] Step 2, preprocess the collected dataset of drowning people and divide it into a training set, a validation set, and a test set;

[0007] Step 3, improve the original network model to construct an SS-YOLOv8 model;

[0008] Step 4, train the improved model using the images in the training set and the validation set;

[0009] Step 5: Specifically evaluate the recognition performance of the improved model using the test set, and continuously optimize the model performance according to the evaluation results to obtain the best recognition model;

[0010] Step 6: Input the image to be detected into the model for detection.

[0011] Furthermore, in the above Step 1, the sources of the relevant data sets collected for drowning victims include publicly available photos of drowning victims and frame images intercepted from drowning videos.

[0012] Furthermore, in the above Step 2, the preprocessing of the collected data set mainly involves enhancing the images, including operations such as horizontal and vertical flipping, shearing, and scaling, to generate a diverse data set. The preprocessed data set will be divided into a training set, a validation set, and a test set.

[0013] Furthermore, in the above Step 3, the construction of the SS-YOLOv8 model is specifically as follows:

[0014] (1) Based on the original YOLOv8 network, the Spatial Pyramid Pooling - Fast (SPPF) module and the ECA (Efficient Channel Attention) attention mechanism are combined to enhance the channel feature extraction ability, improve the robustness of the network and the recognition ability of targets in complex scenes.

[0015] (2) A small - target detection layer is added to the Head layer of the network, which is specifically used to improve the recognition accuracy of small targets.

[0016] Furthermore, the combination position of the Spatial Pyramid Pooling - Fast (SPPF) module and the ECA (Efficient Channel Attention) attention mechanism is: the ECA module is added after all the MaxPool2d operations are completed and before the second convolutional layer (Conv), which can make full use of the multi - scale features after the pooling operation, improve the weight distribution between channels, and enhance the expression of important features.

[0017] Furthermore, ECA is a lightweight channel attention mechanism. By introducing local one - dimensional convolution to capture the mutual relationship between channels, it assigns adaptive weights to each channel, highlights the channels that contribute more to feature expression, and suppresses irrelevant channels, avoiding the dependence on the fully - connected layer, while significantly reducing the computational complexity while ensuring the effectiveness of the channel attention mechanism.

[0018] Furthermore, the calculation method of the ECA module is as follows:

[0019] First, perform global average pooling operation on each channel to obtain the global description vector of each channel:

[0020]

[0021] where Xc (h, w) is the value at the position of height h and width w in the c-th channel of the input feature map; c = 1, 2, …, C; h = 1, 2, …, H; w = 1, 2, …, W. z c represents the global average information of channel c.

[0022] To avoid the high computational overhead brought by the fully connected layer, ECA uses one-dimensional convolution to capture the dependencies between channels. That is, a one-dimensional convolution with a kernel size of k is applied to the channel description vector z, and the output is the channel attention weight w.

[0023] w = Conv1D(z, k)

[0024] where Conv1D(z, k) represents the one-dimensional convolution operation with a kernel size of k. The size k of the convolution kernel is adaptively determined by the number of channels of the input features to capture appropriate channel dependencies.

[0025]

[0026] where γ and β are hyperparameters; |.| odd means taking the nearest odd number of the result to ensure that the size of the convolution kernel is odd and avoid information loss in the convolution operation.

[0027] Normalizing the channel weight vector w obtained by one-dimensional convolution through the Sigmoid activation function, the attention weight w of each channel can be obtained. c .

[0028] w c = σ(w c )

[0029] where c = 1, 2, …, C; σ(.) represents the Sigmoid activation function.

[0030]

[0031] Finally, multiply the original input feature map X by the generated channel weights to achieve adaptive weighting for each channel and obtain the output feature map result Y.

[0032] Y c (h, w) = w c X c (h, w)

[0033] where Y c (h, w) represents the value at the position of height h and width w in the c-th channel of the output feature map, c = 1, 2, …, C; h = 1, 2, …, H; w = 1, 2, …, W.

[0034] Furthermore, a small object detection layer is added to the Head layer of YOLOv8, specifically for detecting small objects, enhancing the network's ability to detect small objects.

[0035] Furthermore, in step 4, the model training steps are as follows:

[0036] (1) In terms of network environment configuration, the operating system is Ubuntu-18.04 64-bit Server, the CPU of the experimental machine is Intel Xeon, the GPU is NVIDIA 2080TI, and the video memory is 12G.

[0037] (2) The network structure is built based on Pytorch 1.8.1, the programming language is Python 3.9, and the GPU acceleration library is CUDA 11.1.

[0038] (3) Configure and train hyperparameters, and use the training set data to calculate the loss and continuously optimize the model weights during each training process.

[0039] Furthermore, in step 5, the specific method for evaluating the model performance using the test set is as follows:

[0040] (1) Calculate the precision rate, recall rate, and mean average precision (mAP@0.5) of the model as evaluation indicators for detection accuracy.

[0041] (2) Calculate the number of model parameters (params), detection speed (detection speed), and frame rate (Frame) as evaluation indicators for the size and real-time performance of the detection model.

[0042] Beneficial effects: The drowning person recognition method based on YOLOv8 of the present invention has a high detection accuracy and a fast detection efficiency, and can accurately identify drowning persons in complex environments. In addition, this method can adapt to different drowning scenarios, has strong practicability, and can be widely applied to various scenarios such as swimming pools and oceans, providing strong support for drowning rescue.

[0043] (1) Combine the SPPF module in the YOLOv8 backbone network with the ECA mechanism. By aggregating multi-scale context information and optimizing the weight distribution between channels, the model's ability to extract features of different scales is improved while ensuring the detection efficiency.

[0044] (2) A small object detection is added to the Head layer, specifically for processing small-scale objects existing in the water area scene, further improving the model's ability to detect small objects. Description of the Drawings

[0045] Figure 1 It is the flowchart of the drowning person recognition method based on the improved YOLOv8 of the present invention;

[0046] Figure 2 It is the schematic diagram of the network structure of the model SS-YOLOv8 of the present invention;

[0047] Figure 3 It is the schematic diagram of the network structure of the SPPF_ECA module of the present invention;

[0048] Figure 4 It is the schematic diagram of the network structure of the ECA module of the present invention;

[0049] Figure 5 It is the schematic diagram of the network structure of the small target detection head of the present invention. Detailed implementation manners

[0050] To make the technical solution of the present invention clearer, the following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments.

[0051] Embodiment 1

[0052] Step 1: Obtain the drowning person dataset.

[0053] Collect the surveillance videos of real drowning scenes and the images of public drowning incidents, covering scenes such as open waters, swimming pools, oceans, and turbulences.

[0054] Step 2: Preprocess the dataset.

[0055] To increase the diversity of the sample data, perform horizontal, vertical, and horizontal-vertical mixed flips, shears, zooms, translations, brightness adjustments, and Mosaic data augmentations on the original images. Divide the augmented dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0056] Step 3: Build an improved SS-YOLOv8 model.

[0057] The principle of the YOLOv8 model is as follows:

[0058] The official of YOLOv8 provides five models of different sizes for users to choose from, namely YOLOv8n, YOLOv8s, YOLOv8m, YOLOv8l, and YOLOv8x. These models have the same network structure and control the model scale through two parameters, depth_multiple and width_multiple. Among them, YOLOv8x is the largest model in the YOLOv8 series. It has deeper network layers and stronger feature extraction ability, and can better handle object detection in complex environments. Therefore, YOLOv8x is used as the main model framework.

[0059] The model architecture of YOLOv8 mainly consists of three parts: backbone, Neck, and Head.

[0060] The Backbone part is mainly responsible for extracting low-level features from the input image. The Backbone of YOLOv8 uses the C2f module, which consists of two convolutional layers and multiple bottleneck layers, and enhances the depth and breadth of feature extraction through residual connections.

[0061] In addition, the SPPF (SpatialPyramidPoolingFaster) module is integrated at the end of the Backbone. Through multi-level pooling operations and residual connections, the feature aggregation ability is improved without significantly increasing the computational cost.

[0062] The Neck part is used to further fuse the multi-scale features extracted by the Backbone and enhance the recognition ability of targets of different scales in the detection task. The Neck layer of YOLOv8 uses the structure of the Path Aggregation Network (PAN) and the Feature Pyramid Network (FPN) for multi-scale feature fusion, and fuses the features of different levels through upsampling and downsampling operations to improve the generalization ability of the model in complex scenarios.

[0063] The Head part is responsible for the final object detection task. The Decoupled-Head structure is adopted in YOLOv8 to separate the classification and regression tasks. This structure helps to capture the category information and location features of the target more accurately. In addition, the number of channels of the regression head is also adjusted accordingly.

[0064] The network structure of the improved model SS-YOLOv8 is as Figure 2As shown in the figure, Conv is a convolutional layer used to extract local features of the input feature map; C2f is a cross-stage partial network module used to enhance the feature reuse ability and balance the computational efficiency and feature expression ability; upsample is an upsampling module used to restore the low-resolution feature map to a high-resolution one for fusion with the feature map at the corresponding level; SPPF_ECA is the improved module of the present invention. The specific steps for constructing the SS-YOLOv8 model are as follows:

[0065] 1) The network structure of the SPPF_ECA module is as Figure 3 shown. Add the ECA module after the concatenation (Concat) operation of all max pooling layers (MaxPool2d) and before the second convolutional layer (Conv), make full use of the multi-scale features after the pooling operation, improve the weight distribution between channels, and enhance the expression of important features.

[0066] 2) The network structure of the ECA module is as Figure 4 shown, where the hyperparameters of the ECA module are set as: γ = 2, β = 1;

[0067] 3) The network structure of the small object detection head is as Figure 5 shown. Add a small object detection of 160×160 to the Head layer of YOLOv8 to form a 4-detection head structure of 160×160, 80×80, 40×40, and 20×20, which can respectively detect extremely small, small, medium, and large objects, and connect it with the shallow and deep feature maps to expand the detection range of the network for extremely small objects in the image and enhance the detection ability of the network for small objects. Through more detailed object division and higher attention to details, the missed detection and false detection are effectively reduced.

[0068] Step 4: Train the model using the training set and the validation set.

[0069] Set the hyperparameters. Set the training cycle to 100 rounds, the image batch processing to 64, use the SGD optimizer to update the network parameters, set the momentum to 0.937, the initial learning rate to 0.01, and the weight decay coefficient to 0.0005.

[0070] Start training the SS-YOLOv8 model, perform forward propagation and backward propagation using the training set data, calculate the classification loss, localization loss, and confidence loss, and update the model weights through the optimizer.

[0071] After each training cycle ends, evaluate the model performance using the validation set, monitor the loss value, accuracy, and recall rate on the validation set to prevent overfitting of the model training.

[0072] Dynamically adjust the learning rate according to the verification results, and adopt the cosine annealing strategy to smoothly decrease the learning rate from the initial value to the minimum value within the training cycle, improving the convergence stability of the model.

[0073] Step 5: Evaluate the model performance of SS-YOLOv8 using the test set.

[0074] 1) Calculate the accuracy, recall rate, and mean average precision of the model on the test set to evaluate the detection performance of the model.

[0075] 2) Calculate the number of parameters and frame rate of the model on the test set to evaluate the detection speed of the model, and judge whether the model meets the preset performance requirements according to the evaluation results.

[0076] 3) Finally, output the optimized model performance report, including the confusion matrix, PR curve, and detection results, to ensure the robustness of the model in complex scenarios.

[0077] Step 6: Input the pictures to be detected into the model for detection.

[0078] As described above, the above embodiments are only examples of the specific implementation manners of the present invention, and their purpose is to more clearly illustrate the technical solutions of the present invention, rather than constituting any limitation to the protection scope of the present invention. Those skilled in the art should understand that without departing from the core technical idea and scope defined by the claims of the present invention, reasonable adjustments or equivalent replacements can be made to the technical details, parameter configurations, and implementation methods in the embodiments. Such modifications and variations should be covered within the patent protection scope of the present invention. The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A drowning person recognition method based on improved YOLOv8, characterized in that, It includes the following steps: Step 1: Obtain the dataset of drowning persons; Step 2: Preprocess the collected dataset of drowning persons and divide it into a training set, a validation set, and a test set; Step 3: Improve the original network model to construct the SS-YOLOv8 model; Step 4: Use the images in the training set and the validation set to train the improved model; Step 5: Specifically evaluate the recognition performance of the improved model using the test set, and continuously optimize the model performance according to the evaluation results to obtain the best recognition model; Step 6: Input the picture to be detected into the model for detection.

2. The drowning person recognition method based on improved YOLOv8 according to claim 1, wherein, In Step 1, the obtained dataset is specifically: based on the publicly available real photos of drowning persons on the Internet and the frame images intercepted from drowning videos as the sample dataset, including multiple images of 640*640 pixels and the corresponding label files, and the label files are in txt format.

3. The drowning person recognition method based on improved YOLOv8 according to claim 1, wherein In Step 2, the steps for preprocessing the images in the dataset are specifically: perform horizontal, vertical, and horizontal-vertical mixed flipping, shearing, scaling, translation, brightness adjustment, and Mosaic data augmentation on the original images, and divide the dataset according to the ratio of 7:2:1 to obtain the training set, the validation set, and the test set.

4. The drowning person recognition method based on improved YOLOv8 according to claim 1, characterized in that, In Step 3, improving the original network model to construct the SS-YOLOv8 model is specifically: (1) Combine the Spatial Pyramid Pooling - Fast (SPPF) module with the ECA mechanism to enhance the model's ability to capture important features while ensuring computational efficiency; (2) Add a small object detection layer for small objects such as the heads and arms of drowning persons, so that the model can more sensitively capture the feature information of small objects.

5. A drowning person recognition method based on improved YOLOv8 according to claim 1, characterized in that, In Step 4, the model training environment is specifically: the training period is set to 100 epochs, the picture batch size is set to 64, the Stochastic Gradient Descent (SGD) optimizer is used to update the network parameters, the momentum is set to 0.937, the initial learning rate is set to 0.01, and the weight decay coefficient is set to 0.0005. After the training period ends, use the validation set to verify the model performance, observe the loss function and accuracy on the validation set, and prevent the model from overfitting.

6. The drowning person recognition method based on improved YOLOv8 according to claim 1, wherein, In Step 5, evaluating the model performance using the test set is specifically: calculate indicators such as the accuracy, recall rate, mean average precision, number of parameters, and frame rate of the model on the test set to determine the performance of the model in actual detection.