Mouse identification and detection method based on improved YOLOv8

By improving the network structure and loss function of YOLOv8, the accuracy and efficiency of mouse recognition detection are improved, and the problem of insufficient mouse detection accuracy in the prior art is solved, thereby achieving a more efficient mouse image recognition effect.

CN120164231APending Publication Date: 2025-06-17NANJING FORESTRY UNIV

Patent Information

Application Number
CN202510216629.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing mouse detection methods are insufficient in dealing with complex and changing backgrounds and target occlusions, making it difficult to effectively identify and locate mice.

Method used

The mouse identification detection method based on improved YOLOv8 is adopted, and the feature expression ability and generalization performance of the model are improved by improving YOLOv8's backbone network, neck network and loss function, reducing redundant calculations and optimizing feature delivery paths.

Benefits of technology

It significantly improves the expression ability of mouse image features and the generalization performance of the model, improves the computing efficiency and recognition accuracy, and can more effectively handle mouse image recognition tasks in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164231A_ABST
    Figure CN120164231A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv8-based mouse identification and detection method, which can identify mice appearing in various scenes, is convenient for subsequent control and capture, improves the expression ability of mouse image features and the generalization performance of a model, reduces redundant calculation, optimizes a feature transmission path, and remarkably improves the calculation efficiency. The SPPFSCF module is adopted for the trunk part of the YOLOv8 model, efficient processing and multi-level feature extraction of the image are achieved by integrating various convolution layers, pooling layers and feature map splicing operation, and richer feature representation is provided for a mouse image recognition task. A SimAM attention mechanism module is introduced into the neck network, and the expression and detection capability of a target is enhanced through feature similarity. A dynamic non-monotonic focusing mechanism is introduced, gradient gains can be distributed more reasonably, a model is made to pay more attention to an anchor frame with common quality, and therefore the overall performance of a detector is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mouse detection, and particularly relates to a mouse recognition method based on improved YOLOv8. Background Art

[0002] As a common pest, the harm of mice cannot be ignored. They not only cause direct economic losses through behaviors such as gnawing furniture and wires, but also may spread various diseases, such as the plague, hemorrhagic fever, etc., posing a serious threat to human health. In addition, mice have strong reproductive capabilities. Once they settle in a living or working environment, their numbers will increase rapidly, further intensifying their competition for food resources, resulting in food contamination and waste. At the same time, their feces and urine will also pollute the environment and air, triggering allergic reactions and hygiene problems. Therefore, effectively controlling the number and activities of mice is of great significance for protecting human health, maintaining environmental hygiene, and ensuring property safety.

[0003] In recent years, deep learning technology has made breakthrough progress in the field of image recognition, providing strong technical support for mouse detection. Traditional mouse detection methods mostly rely on manually designed features and rules, such as color, shape, texture, etc. These methods often struggle to handle problems such as complex and variable backgrounds and target occlusion. The introduction of deep learning technology, especially convolutional neural networks (CNNs), has greatly improved the accuracy and robustness of target detection by automatically learning hierarchical feature representations in images. By training a large number of labeled mouse image data, deep learning models can learn the unique features of mice and accurately identify the position and size of mice in new images. In addition, with the continuous development of efficient target detection algorithms such as the YOLO (You Only Look Once) series, the real-time performance and accuracy of deep learning technology in mouse detection have been further improved, providing efficient and intelligent solutions for fields such as agricultural pest management and laboratory animal monitoring. Summary of the Invention

[0004] Based on the above problems, the present invention provides a mouse recognition and detection method based on improved YOLOv8, which can identify mice appearing in various scenarios, facilitating subsequent control and capture, enhancing the expression ability of mouse image features and the generalization performance of the model, reducing redundant calculations, optimizing the feature transfer path, and significantly improving the calculation efficiency.

[0005] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:

[0006] A mouse recognition and detection method based on improved YOLOv8, the method comprising the following steps:

[0007] Step 1: Obtain a mouse image dataset;

[0008] Step 2: Construct an improved YOLOv8 network model; including improving the backbone network, neck network, and original loss function of YOLOv8 to obtain the improved YOLOv8 network model, and the loss function of the improved YOLOv8 network model is the Wise_IoUv2 loss function;

[0009] Step 6: Input the mouse image data into the improved YOLOv8 network model for training.

[0010] The above mouse recognition and detection method based on the improved YOLOv8, the specific content of step 1 includes:

[0011] Preprocess the obtained mouse image dataset, including: randomly dividing the dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0012] The above mouse recognition and detection method based on the improved YOLOv8, improving the backbone network of YOLOv8, specifically including:

[0013] The improved backbone network structure includes 5 Conv modules, 4 C2f modules, and an improved SPPFSCF module, which is responsible for feature extraction;

[0014] The convolution kernels of the 5 Conv convolution modules are all 3×3, the strides are all 2, the paddings are all 1, and the number of output channels gradually doubles from 64 until 1024;

[0015] Improve the original SPPF pyramid module to the SPPFSCF module for deep learning;

[0016] In the SPPFSCF module, it includes five 1×1 convolutional layers, two 3×3 convolutional layers, three max-pooling layers with pooling kernels of 3×3, 5×5, and 9×9 respectively, and two Concat splicing layers;

[0017] The data processing process of the SPPFSCF module:

[0018] 1) Conv1: Increase the number of channels and feature transformation

[0019] First, the input mouse image x passes through the first 1×1 convolutional layer, namely Conv1, which is used to increase the number of channels and perform preliminary feature transformation;

[0020] 2) Conv2: Process the original input in parallel

[0021] The second 1×1 convolutional layer, namely Conv2, processes the original input image x in parallel with Conv1, as an independent branch for subsequent splicing operations;

[0022] 3) Conv3: Feature Extraction

[0023] The first 3×3 convolutional layer, i.e., Conv3, further extracts features from the feature map output by Conv1 to capture the local texture and edge information of the mouse image and enhance the feature expression ability;

[0024] 4) Conv4: Feature Transformation and Channel Adjustment

[0025] After Conv3, the third 1×1 convolutional layer, i.e., Conv4, transforms the feature map and adjusts the number of channels to adapt to subsequent multi-scale feature fusion;

[0026] 5) Pooling Layer: Multi-scale Feature Fusion

[0027] The feature map x1 processed by Conv4 undergoes three max-pooling operations through three max-pooling layers (pooling kernels are 3×3, 5×5, and 9×9 respectively), and three feature maps x2, x3, and x4 with different sizes are obtained respectively;

[0028] 6) Concat: Feature Map Concatenation

[0029] The feature map x1 processed by Conv4, the feature maps x2, x3, and x4 after three poolings are concatenated in the channel dimension through the first Concat concatenation function. The concatenated feature map contains multi-scale information and can describe the features of the mouse image more comprehensively;

[0030] 7) Conv5 and Conv6: Processing of the Concatenated Feature Map

[0031] The concatenated feature map is reduced in dimension through the fourth 1×1 convolutional layer, i.e., Conv5, to reduce the number of channels and computational complexity; then, the second 3×3 convolutional layer, i.e., Conv6, further extracts the features of the concatenated feature map to enhance the feature expression ability;

[0032] 8) Feature Map Concatenation of the Conv2 Branch

[0033] The feature map obtained by Conv2 processing the original input image x is concatenated with the feature map output by Conv6 through the second Concat concatenation function; when concatenating, it is necessary to ensure that the shapes of the two feature maps are the same in the non-concatenated dimensions. The size of the concatenated feature map in the channel dimension is the sum of the two, and remains unchanged in other dimensions;

[0034] 9) Conv7: Output of the Final Feature Map

[0035] The concatenated feature map passes through the last 1×1 convolutional layer, i.e., Conv7, to increase the number of channels from the input channel number c1 to c2 and output the final feature map.

[0036] The above-mentioned mouse recognition and detection method based on improved YOLOv8 improves the neck network of YOLOv8, specifically including:

[0037] Add a SimAM attention mechanism module after the neck network. The SimAM attention mechanism module dynamically adjusts the weights of each pixel by calculating the similarity between each pixel in the feature map and its neighboring pixels, thereby enhancing important features and suppressing irrelevant features. Its working process:

[0038] 1) Input feature map: The input feature map X ∈ R(B×C×H×W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively;

[0039] 2) Calculate local self-similarity: For each pixel x in the feature map i,j , first calculate the square of the difference between it and all pixels in the neighborhood, then sum these differences and perform normalization to obtain the s of the pixel i,j :

[0040]

[0041] Among them, i and j respectively represent the position indexes of the pixel in the feature map, and Ω i,j represents the pixels within the neighborhood of pixel x i,j , and N is the number of pixels in the neighborhood.

[0042] 3) Generate attention weight w i,j :

[0043]

[0044] Among them, is the average value or standard deviation of each pixel s i,j in the entire feature map or local area, and ∈ is a very small constant used to prevent division by zero error;

[0045] Multiply the attention map by the feature map: Multiply the generated attention weight map W ∈ R(B×1×H×W) by the input feature map X to obtain the weighted feature map X′ = W⊙X, where ⊙ represents element-wise multiplication.

[0046] The above-mentioned mouse recognition and detection method based on improved YOLOv8 improves the loss function of YOLOv8, specifically including: changing the loss function of YOLOv8 to the Wise_IoUv2 loss function, which uses a dynamic non-monotonic focusing mechanism to evaluate the quality of the bounding box and uses the outlier degree as an index to replace the traditional IoU to measure the quality of the bounding box; the loss value L of the Wise_IoUv2 loss function WIoUv2 is:

[0047]

[0048] Where: L WIoUv1 is the loss value of the Wise_IoU v1 loss function: L WIoUv1 = R WIoU L IoU , the penalty term (x gt , y gt ) and (x, y) are the center points of the predicted bounding box and the ground truth bounding box respectively; W g and H g are the width and height of the smallest enclosing box of the predicted bounding box and the ground truth bounding box respectively;

[0049] L IoU represents the overlap degree between the predicted bounding box and the ground truth bounding box: W i and H i are the width and height of the intersection area of the predicted bounding box and the ground truth bounding box respectively; S u is the area of the union region of the predicted bounding box and the ground truth bounding box; represents the moving average with momentum m;

[0050] is the gradient gain, is the monotonic focusing coefficient, γ is the exponential power, * means to prevent R WIoU from generating a speed that hinders convergence, and separate W g and H g from the computational graph;

[0051] Backpropagation gradient:

[0052] The beneficial effects of the method of the present invention are:

[0053] (1) After replacing the SPPF module in the backbone of the YOLOv8 model with the improved SPPFSCF module, compared with the original SPPF pyramid module, the SPPFSCF module integrates the design concepts of spatial pyramid pooling (SPP), fully connected layer (FC), and cross-stage partial connection (CSP). Among them, the SPP part fuses the different-scale feature information of the mouse image through multi-scale pooling operations, effectively enhancing the model's ability to capture multi-scale features; the CSP part further improves the expression ability of the mouse image features and the generalization performance of the model through partial connection and cross-stage feature fusion mechanisms. In addition, the CSP network structure significantly improves the computational efficiency by reducing redundant calculations and optimizing the feature transmission path, thus accelerating the model's inference process. The SPPFSCF module realizes the efficient processing of the input mouse image and multi-level feature extraction by integrating various convolutional layers, pooling layers, and feature map splicing operations, providing a richer feature representation for the mouse image recognition task.

[0054] (2) Introduce the SimAM attention mechanism module into the neck network of YOLOv8. Its core idea is to enhance the representation and detection ability of the target through feature similarity. The addition of the SimAM module enables the model to more effectively focus on the key regions in the mouse image, thus improving the accuracy of target recognition. Specifically, SimAM calculates the feature similarity between different positions in the image, focuses on and integrates the similar target regions, and then strengthens the representation ability of the target and optimizes the recognition effect of the mouse image. In addition, the SimAM attention mechanism does not require the introduction of additional parameters, has a low computational cost, and can effectively improve the generalization performance of the model. Integrating the SimAM module in the neck network not only promotes the fusion and interaction between the features of mouse images at different scales, but also significantly enhances the model's detection and recognition ability for multi-scale targets, enabling it to more efficiently handle the mouse image recognition task in complex scenarios.

[0055] (3) Replace the loss function of YOLOv8 with the Wise-IoUv2 loss function. Compared with traditional loss functions (such as IoU, CIoU, etc.), Wise-IoUv2 shows superior performance in the mouse image recognition task. Its core improvement lies in the introduction of a dynamic non-monotonic focusing mechanism, which can more reasonably allocate gradient gains, enabling the model to pay more attention to ordinary-quality anchor boxes, thus significantly improving the overall performance of the detector. In addition, Wise-IoUv2 fully considers the problem of low-quality samples in the object detection dataset during design, such as the detection and recognition of small objects like mice or blurred objects in complex environments. By enhancing the fitting ability of the bounding box loss, Wise-IoUv2 can more accurately handle small objects and effectively improve the detection and recognition accuracy of the model for small objects. These characteristics make Wise-IoUv2 have broader application potential and practical value in the mouse image recognition task. Description of the Drawings

[0056] Figure 1 It is a schematic flow diagram of a mouse recognition method based on improved YOLOv8;

[0057] Figure 2 It is the structural diagram of the YOLOv8 network model before improvement;

[0058] Figure 3 It is the structural diagram of the improved SPPFSCF module;

[0059] Figure 4 It is the structural diagram of the SimAM attention mechanism module;

[0060] Figure 5 It is the structural diagram of the improved YOLOv8 network model;

[0061] Figure 6 It is the P-R curve trained by the improved YOLOv8 network. Detailed Implementation Manner

[0062] A mouse recognition and detection method based on improved YOLOv8, the method includes the following steps:

[0063] Step 1: Obtain a mouse image dataset;

[0064] Step 2: Construct an improved YOLOv8 network model, specifically including steps 3, 4, and 5;

[0065] Step 3: Improve the backbone network of the YOLOv8 network, and use the improved backbone network to extract features from mouse images. Through a series of convolutional and deconvolutional layers, as well as techniques such as residual connections and bottleneck structures, it can effectively extract key information in mouse images, while reducing the size of the network and improving performance.

[0066] Step 4: Improve the neck network of the YOLOv8 network, and use the improved neck network to fuse the feature maps of different levels extracted from the head network, enhance the feature representation ability, so as to provide richer feature information to the head network.

[0067] Step 5: Improve the original loss function of YOLOv8 to the Wise_IoUv2 loss function;

[0068] Step 6: Use the head network to predict information such as the category, location, and confidence of the target based on the feature maps extracted from the previous steps, using a series of components such as convolutional layers, classifiers, and detectors.

[0069] The specific steps of Step 1 include:

[0070] Step 1.1: Perform preprocessing operations on the obtained mouse image dataset;

[0071] Step 1.2: The preprocessing operations on the obtained mouse image dataset include: randomly dividing the dataset into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0072] The specific steps of Step 2 include:

[0073] The improved YOLOv8 model constructed includes a backbone network, a neck network, and a head network.

[0074] The specific steps of Step 3 include:

[0075] Step 3.1: The improved backbone network structure includes 5 Conv modules, 4 C2f modules, and an improved SPPFSCF module, which is responsible for feature extraction, using a series of convolutional and deconvolutional layers, and at the same time using residual connections and bottleneck structures, and also including some common improvement techniques, such as depthwise separable convolution and dilated convolution.

[0076] Step 3.2: The convolutional kernels of the 5 Conv convolutional modules are all 3×3, the strides are all 2, the paddings are all 1, and the number of output channels gradually doubles from 64 until it reaches 1024. When using a stride of 2, the spatial resolution will decrease. A stride of 2 means that the convolution sum moves two pixels each time, and the spatial size of the input image will be halved. Padding of 1 means that each circle of the input image is padded with a circle of zeros, which is beneficial to maintaining the spatial dimension after convolution. The formula for calculating the output size of the convolution operation is:

[0077]

[0078] where W1 and H1 are the width and height of the input feature map; W2 and H2 are the width and height of the output feature map.

[0079] Step 3.3: The C2f module consists of multiple bottleneck blocks and includes two key parameters: shortcut, a boolean value indicating whether the bottleneck block uses a shortcut connection. If the value of shortcut is True, the bottleneck block uses a shortcut connection; otherwise, if it is False, no shortcut connection is used; n: specifying the number of bottleneck blocks within the C2f module. After being processed by the C2f module, the resolution of the input feature map remains unchanged and its output channel number is the same as the input channel number.

[0080] Step 3.4: Improve the original SPPF pyramid module. The improved SPPFSCF module is a module used to construct a deep learning model.

[0081] In the SPPFSCF module, it includes five 1×1 convolutional layers, two 3×3 convolutional layers, three max-pooling layers with pooling kernels of 3×3, 5×5, and 9×9 respectively, and two Concat splicing layers.

[0082] This module combines the concepts of spatial pyramid pooling (SPP), fully connected layer (FC), and cross-stage partial connection (CSP), combines different convolutional layers and pooling layers, as well as the splicing of feature maps to process and extract features from the input feature map.

[0083] The data processing process of the SPPFSCF module:

[0084] 1) Conv1: Increase the number of channels and feature transformation

[0085] First, the input mouse image passes through the first 1×1 convolutional layer (Conv1) to increase the number of channels and perform preliminary feature transformation. This step can increase the information capacity and extract more diverse features.

[0086] 2) Conv2: Process the original input in parallel

[0087] The second 1×1 convolutional layer (Conv2) processes the original input image x in parallel with Conv1 as an independent branch for subsequent splicing operations. This branch retains the local detail information of the original image.

[0088] 3) Conv3: Feature extraction

[0089] Next, the first 3×3 convolutional layer (Conv3) further extracts features from the feature map output by Conv1. The 3×3 convolutional kernel can capture the local texture and edge information of the mouse image and enhance the feature expression ability.

[0090] 4) Conv4: Feature transformation and channel adjustment

[0091] After Conv3, the third 1×1 convolutional layer (Conv4) transforms the feature map to adapt to subsequent multi-scale feature fusion.

[0092] 5) Pooling layer: Multi-scale feature fusion

[0093] The feature map x1 processed by Conv4 undergoes three max-pooling operations through three max-pooling layers (with pooling kernels of 3×3, 5×5, and 9×9 respectively), resulting in three feature maps x2, x3, and x4 of different sizes; these feature maps can capture the key information of the mouse image at different resolutions.

[0094] 6) Concat: Feature map concatenation

[0095] The feature map x1 processed by Conv4, the processing results x2, x3, and x4 of the three pooling operations, are concatenated in the channel dimension through the first Concat concatenation function. The concatenated feature map contains multi-scale information and can describe the features of the mouse image more comprehensively.

[0096] 7) Conv5 and Conv6: Processing of the concatenated feature map

[0097] The concatenated feature map is reduced in dimension through the fourth 1×1 convolutional layer (Conv5) to reduce the number of channels and computational complexity. Then, the second 3×3 convolutional layer (Conv6) further extracts the features of the concatenated feature map to enhance the feature representation ability.

[0098] 8) Feature map concatenation of the Conv2 branch

[0099] Meanwhile, the feature map obtained by Conv2 processing the original input x is used as another branch and concatenated with the feature map output by Conv6 through the second Concat concatenation function. When concatenating, it is necessary to ensure that the shapes of the two feature maps are consistent in the non-concatenated dimensions. The size of the concatenated feature map in the channel dimension is the sum of the two, while remaining unchanged in other dimensions.

[0100] 9) Conv7: Output of the final feature map

[0101] Finally, the concatenated feature map passes through the last 1×1 convolutional layer (Conv7) to output the final feature map. This feature map combines multi-scale information and original details and can describe the key features of the mouse image more accurately.

[0102] The specific steps of step 4 include:

[0103] Step 4.1: Add the SimAM attention mechanism module after the neck network. SimAM is an attention mechanism based on the local self-similarity of feature maps. By calculating the similarity between each pixel in the feature map and its neighboring pixels, it dynamically adjusts the weights of each pixel, thereby enhancing important features and suppressing irrelevant features. Its working principle can be divided into the following steps:

[0104] 1) Input feature map: The input feature map X ∈ R(B×C×H×W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively.

[0105] 2) Calculate local self-similarity: For each pixel x i,j (where i, j represent the position indices of the pixel in the feature map respectively), SimAM calculates its similarity with the surrounding pixels. This similarity is usually measured by the distance between the feature vectors of pixels, and the commonly used one is the negative square of the Euclidean distance. Specifically, SimAM indirectly reflects the similarity by calculating the average value (after normalization) of the squares of the differences between each pixel and the pixels in its neighborhood. For each pixel, first calculate the squares of the differences between it and all the pixels in its neighborhood, and then sum and normalize these differences.

[0106]

[0107] where, Ω i,j represents the neighborhood of pixel x i,j (excluding x i,j itself, and N is the number of pixels in the neighborhood).

[0108] 3) Generate attention weights: Based on the calculated si,j above, SimAM generates attention weights wi,j through the following formula:

[0109]

[0110] where, is a certain form of normalization of s i,j (in the implementation of SimAM, it is usually approximated by the average value and standard deviation of s i,j over the entire feature map or a local region), ∈ is a very small constant (such as 1e-4) used to prevent division by zero errors. This formula is actually a variant of the sigmoid function, which maps s i,j to the interval (0,1) as the attention weight.

[0111] 4) Multiply the attention map by the feature map: Multiply the generated attention weight map \(W\in\mathbb{R}^{(B\times1\times H\times W)}\) (note that the channel dimension is ignored here because SimAM usually calculates the attention weights independently for each channel) by the original feature map \(X\) to obtain the weighted feature map \(X' = W\odot X\), where \(\odot\) represents element-wise multiplication.

[0112] Step 5 specifically includes:

[0113] Step 5.1: Improve the loss function of the YOLOv8 model to the Wise_IoUv2 loss function. Wise_IoUv2 is constructed based on Wise_IoUv1. It uses a dynamic non-monotonic focusing mechanism to evaluate the quality of the bounding boxes and uses the outlier degree as an index to replace the traditional IoU to measure the quality of the bounding boxes. The formula is:

[0114]

[0115] where: is the monotonic focusing coefficient, \(\gamma\) is the exponential power, \(L\) IoU is the base, and \(*\) means to prevent \(R\) WIoU from generating a speed that hinders convergence, and separate \(W\) g and \(H\) g from the computational graph; the loss value Wise_IoU v1 is \(L\) WIoUv1 \(=R\) WIoU \(L\) IoU ; the penalty term \(L\) IoU represents the overlapping degree between the predicted box and the ground truth box: \(W\) i and \(H\) i are the width and height of the intersection area of the two boxes respectively; \((x\) gt , \(y\) gt ) and \((x, y)\) are the center points of the two boxes respectively; \(S\) u is the union area of the two boxes; \(W\) g and \(H\) g are the width and height of the minimum enclosing box of the two boxes respectively.

[0116] Backpropagation gradient:

[0117]

[0118] During the training process, the gradient gain decreases as \(L\) IoU decreases, and the convergence speed is relatively slow in the later stage of training. Therefore, introduce the mean value of \(L\) IoU as the normalization factor. After introduction, the formula is:

[0119]

[0120] wherein is the gradient gain, is the monotonic focusing coefficient, represents the sliding average with momentum m, γ > 0.

[0121] The specific steps of step 6 include:

[0122] Step 6.1, Based on the feature map previously extracted from the image, use the head network to apply a series of designed components, including multi-layer convolutional layers, efficient classifiers, and detectors, etc., to comprehensively predict and analyze key information such as the class attribution, precise location, and confidence level of the target object. Make full use of the rich image features contained in the feature map, and through a professional combination of components, ensure the accuracy and reliability of the prediction results.

Claims

1. A mouse recognition detection method based on improved YOLOv8, the method comprising the following steps: Step 1: Obtain a mouse image dataset; Step 2: construct an improved YOLOv8 network model; including improving the backbone network, neck network, and original loss function of YOLOv8 to obtain an improved YOLOv8 network model, and the loss function of the improved YOLOv8 network model is the Wise_IoUv2 loss function; Step 6: Input the mouse image data into the improved YOLOv8 network model for training.

2. The mouse recognition and detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: The step 1 specifically includes: The preprocessing operation of the acquired mouse image dataset includes: randomly dividing the dataset into a training set, a validation set, and a test set in a ratio of 8:1:

1.

3. The mouse recognition and detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: Improvements to the YOLOv8 backbone network include: The improved backbone network structure includes 5 Conv modules, 4 C2f modules and an improved SPPFSCF module, which is responsible for feature extraction; The convolution kernel size of the five Conv convolution modules is 3×3, the stride is 2, the padding is 1, and the number of output channels is gradually doubled from 64 to 1024; Improve the original SPPF pyramid module to SPPFSCF module for deep learning; In the SPPFSCF module, there are five 1×1 convolutional layers, two 3×3 convolutional layers, three maximum pooling layers with pooling kernels of 3×3, 5×5, and 9×9, and two Concat splicing layers; The data processing process of the SPPFSCF module: 1) Conv1: Increase the number of channels and feature transformation First, the input mouse image x passes through the first 1×1 convolutional layer, Conv1, to increase the number of channels and perform preliminary feature transformation; 2) Conv2: Processing the original input in parallel The second 1×1 convolutional layer, Conv2, processes the original input image x in parallel with Conv1 as an independent branch of the subsequent splicing operation; 3) Conv3: Feature Extraction The first 3×3 convolutional layer, Conv3, further extracts features from the feature map output by Conv1 to capture the local texture and edge information of the mouse image and enhance the expressiveness of the features. 4) Conv4: Feature transformation and channel adjustment After Conv3, the third 1×1 convolutional layer, Conv4, transforms the feature map and adjusts the number of channels to accommodate subsequent multi-scale feature fusion; 5) Pooling layer: multi-scale feature fusion The feature map x1 processed by Conv4 is passed through three maximum pooling layers (the pooling kernels are 3×3, 5×5, and 9×9 respectively), and three maximum pooling operations are performed to obtain three feature maps of different sizes, x2, x3, and x4 respectively; 6) Concat: Feature map concatenation The feature map x1 processed by Conv4 and the feature maps x2, x3, and x4 after three pooling are concatenated in the channel dimension through the first Concat concatenation function. The concatenated feature map contains multi-scale information and can more comprehensively describe the characteristics of the mouse image. 7) Conv5 and Conv6: Processing of feature maps after splicing The concatenated feature map is reduced in dimension by the fourth 1×1 convolutional layer, Conv5, to reduce the number of channels and reduce the amount of calculation. Then, the second 3×3 convolutional layer, Conv6, further extracts the concatenated features to enhance the expressiveness of the features. 8) Feature map concatenation of Conv2 branch The feature map obtained by Conv2 processing the original input image x is concatenated with the feature map output by Conv6 through the second Concat concatenation function. When concatenating, it is necessary to ensure that the shapes of the two feature maps are consistent in the non-concatenated dimensions, and the size of the concatenated feature map in the channel dimension is the sum of the two, while remaining unchanged in other dimensions. 9) Conv7: Final feature map output The concatenated feature map passes through the last 1×1 convolution layer, Conv7, which increases the number of channels from the input channel number c1 to c2 and outputs the final feature map.

4. The mouse recognition and detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: Improvements to the YOLOv8 neck network include: A SimAM attention mechanism module is added after the neck network. The SimAM attention mechanism module dynamically adjusts the weight of each pixel by calculating the similarity between each pixel and its neighboring pixels in the feature map, thereby enhancing important features and suppressing irrelevant features. Its working process is as follows: 1) Input feature map: Input feature map X∈R(B×C×H×W), where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively; 2) Calculate local self-similarity: For each pixel x in the feature map i,j , first calculate the square of the difference between it and all pixels in the neighborhood, then sum and normalize these differences to get the pixel's s i,j : Among them, i, j represent the position index of the pixel in the feature map, Ω i,j Represents pixel x i,j The pixels in the neighborhood of , N is the number of pixels in the neighborhood; 3) Generate attention weight w i,j : in, is the pixel s in the entire feature map or local area i,j The mean or standard deviation of , ∈ is a small constant used to prevent division by zero errors; Multiply the attention map with the feature map: Multiply the generated attention weight map W∈R(B×1×H×W) with the input feature map X to obtain the weighted feature map X′=W⊙X, where ⊙ represents element-by-element multiplication.

5. The mouse recognition and detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: Improve the loss function of YOLOv8, including: The loss function of YOLOv8 is changed to Wise_IoUv2 loss function, which uses a dynamic non-monotonic focusing mechanism to evaluate the quality of the bounding box and uses the outlier index instead of the traditional IoU to measure the quality of the bounding box; the loss value L of the Wise_IoUv2 loss function WIoUv2 for: Where: L WIoUv1 is the loss value of Wise_IoU v1 loss function: L WIoUv1 =R WIoU L IoU , penalty term (x gt ,y gt ) and (x, y) are the center points of the predicted box and the real box respectively; W g and H g are the width and height of the minimum connected box of the predicted box and the real box respectively; L IoU Indicates the degree of overlap between the predicted box and the true box: W i and H i are the width and height of the intersection area between the predicted box and the real box; S u is the area of ​​the joint area of ​​the predicted box and the true box; represents a sliding average with momentum m; is the gradient gain, is the monotone focusing coefficient, γ is the exponential power, and * indicates that in order to prevent R WIoU Produce a speed that hinders convergence, and W g and H g Separation from the computational graph; Back propagating gradients:

Citation Information

Patent Citations

  • Low-illumination target detection method based on convolutional neural network

    CN117576540A

  • Roadblock detection and distance measurement method

    CN118351513A

  • Sheep face expression recognition method, system and device and medium

    CN119360420A

Cited By

  • Pavement structure internal disease detection method based on YOLOv8n

    CN120747086A