A method for detecting the age of squabs

By improving the YOLOv5s network model and combining the CBAM attention mechanism, a pigeon age detection method was developed, which solved the problem of inaccurate pigeon age identification in the existing technology, and realized the accurate feed delivery of an automated feeding system.

CN114972949BActive Publication Date: 2025-05-30ZHONGKAI UNIV OF AGRI & ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210520214.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-05-30
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify the age stage of a suckling pigeon, which makes it difficult for an automated feeding system to accurately calculate the feed output of each pigeon cage.

Method used

By improving the YOLOv5s network model, combined with the CBAM attention mechanism and improved NMS strategy, a pigeon age detection method is developed to accurately identify the pigeon's age stage in the case of insufficient light, overlapping pigeons and blurred pictures.

Benefits of technology

The accuracy and speed of pigeon age detection is achieved, the large amount of effort is spent on designing characteristics of traditional machine learning algorithms, and the accuracy and efficiency of automated feeding systems are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972949B_ABST
    Figure CN114972949B_ABST
Patent Text Reader

Abstract

A method for detecting the age of squabs includes the following steps: improving the YOLOv5s network model. First, improve the spatial attention in CBAM to pointwise spatial attention, and then embed the improved CBAM attention mechanism at the end of each residual branch in the backbone network of the YOLOv5s model. Improve the NMS strategy of YOLOv5s from performing NMS for each class separately to performing NMS for multiple classes simultaneously to obtain an improved YOLOv5s network model; dataset production, annotating squab pictures according to the age stage; inputting the dataset into the improved YOLOv5s network model for model training; after loading the best weight data, inputting the picture to be recognized for recognition. The present invention enhances the recognition efficiency of overlapping squabs, also avoids detecting a single pigeon at multiple different growth stages, improves the recognition accuracy, has a fast detection speed, and promotes the automation of feeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly relates to a method for detecting the age of squabs. Background Art

[0002] Pigeon breeding has great economic prospects. Pigeon meat is tender, with a strong and fresh taste, and rich in nutrition. Compared with common poultry such as chickens, ducks, and geese, pigeons have a shorter growth cycle and higher economic benefits. Traditional manual breeding requires a large number of professional breeders to push feed buckets to feed each pigeon cage point by point, and often need to return to the warehouse to replenish the feed in the feed buckets, with a large labor intensity. And when the pigeon farm reaches a certain scale, this manual feeding method is difficult to meet the requirements. Therefore, it is necessary to apply automation to the pigeon breeding industry. During the feeding process of pigeons, the pigeon members in different cages are inconsistent, and pigeons have different nutritional requirements at different age stages. The traditional manual feeding method can quickly know the amount of feed needed for this cage when the personnel pass by. How can the automated feeding method achieve this? Therefore, there is a difficult problem that urgently needs to be solved in the current automated pigeon breeding - how to quickly calculate the feed delivery amount for each pigeon cage. To solve this problem, it is urgent to develop an intelligent method for detecting the age of squabs, so that the automated machine can accurately identify the age stage of each squab when passing by, and calculate the amount of feed needed for each pigeon cage according to the nutritional requirements of different breeds at different growth stages.

[0003] The YOLOv5 algorithm is a currently popular one - stage target detection algorithm. Compared with the current two - stage target detection algorithms, its main features are a small model and fast speed, and it can also have a detection accuracy similar to that of two - stage target detection. Compared with traditional machine learning algorithms, YOLOv5 can perform end - to - end task training without manual feature design, and performs far better than traditional machine learning algorithms on large datasets. However, YOLOv5 also has deficiencies in the recognition of overlapping squabs. The space in the pigeon cage is limited, resulting in a relatively high degree of overlap of squabs in the picture, and even missed detections. The feeding environment of squabs is often insufficient in light, and the obtained pictures often have blurred boundaries, many occlusions and foreign objects. At the same time, the differences between squabs of similar ages are often small, resulting in the same squab often being recognized as two stages, corresponding to two squabs, etc. These factors all bring certain difficulties to pigeon age recognition. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above - mentioned disadvantages of the prior art, and provide a method for detecting the age of squabs with good target detection accuracy, fast speed, and close to the actual application scenario. This method can provide technical support for the automated feeding of squabs and can also be applied to the breeding of other poultry.

[0005] The present invention is realized by the following technical solutions:

[0006] A method for detecting the age of squabs, comprising the following steps:

[0007] Improve the YOLOv5s network model, improve the spatial attention in CBAM to point-wise spatial attention, and the improved CBAM becomes a new attention structure combining channel attention and point-wise spatial attention. Then embed the improved CBAM attention mechanism at the end of each residual branch in the backbone network of the YOLOv5s model. Improve the NMS strategy of YOLOv5s from performing NMS for each class separately to performing NMS for multiple classes simultaneously, and delete the allocation intervals for different classes to obtain an improved YOLOv5s network model.

[0008] Dataset production: Use LabelImg to label clear squab pictures according to the squab age stage, then randomly allocate the picture set into a training set and a validation set according to a certain ratio to obtain an initial squab dataset, and then perform mosaic data augmentation on the initial dataset to obtain a data-augmented dataset.

[0009] Model training: Input the dataset into the improved YOLOv5s network model for model training to obtain the best weight data of the improved YOLOv5s network model.

[0010] Squab age detection: Load the best weight data into the improved YOLOv5s network model, input a to-be-recognized picture containing a squab, and obtain an output picture with one or more rectangular boxes identifying the age.

[0011] Further, the method for improving the NMS strategy from performing NMS for each class separately to performing NMS for multiple classes simultaneously and deleting the allocation intervals for different classes is as follows:

[0012] Delete the allocation interval I×L in the NMS bounding box locking interval formula B′ = B + I×L. The improved NMS interval determination formula is: B′ = B, so that only one candidate box with the highest IOU score among all classes is retained for the object to be detected, and all other generated candidate boxes are eliminated; IOU refers to the intersection over union, which is the ratio of the intersection and union of the predicted box and the ground truth box.

[0013] Further, in the CBAM attention mechanism, let the input feature map be F, and the two transformations be F′, F″, where F, F′, F″ ∈ R c *w*H , the channel attention module is H ∈ R c*1*1 , the spatial attention module is S ∈ R 1*w*H , the symbol represents the multiplication operation on corresponding elements through the broadcast mechanism. Then the formula for CBAM is:

[0014]

[0015]

[0016] Further, the clear pigeon pictures are obtained by the following method: Collect the pigeon pictures at the breeding site, calculate the pixel variance of each picture through Laplace transform, use the pixel variance as the index value of the picture blurriness, discard the pictures with the index value lower than 20 as blurred pictures, and retain the rest as clear pictures.

[0017] Further, the daily age stages of the pigeons include Stage 1, Stage 2, Stage 3, Stage 4, and Stage 5. Stage 1 is from 1 to 4 days old, Stage 2 is from 5 to 8 days old, Stage 3 is from 9 to 14 days old, Stage 4 is from 15 to 20 days old, and Stage 5 is from 21 days old to slaughter and market.

[0018] Further, the calculation formula of the loss function L of the YOLOv5s model is:

[0019]

[0020] where L B , L c , L O represent the bounding box loss, class loss, and object loss respectively. β, γ, are all constants used to control the weights of the three losses, and are set to 0.05, 0.5, and 1.0 respectively;

[0021] The class loss L C is calculated in the same way as the object loss L O , both by binary cross-entropy. Let x n be the class label output by the model, y n be the true class label. In L O , x n is the object label output by the model, y n is the true object label. Taking the object loss L C as an example, its calculation formula is as follows:

[0022] L C =-∑[y n ×ln x n +(1 - y n )×ln(1 - x n )]

[0023] The bounding box loss L B adopts the CIOU calculation method. Let IOU be the intersection over union between the predicted box and the true box, and ρ 2 (b, bgt ) is the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, c is the diagonal distance of the smallest closed region enclosing the predicted bounding box and the ground truth bounding box, α is a parameter for balancing the ratio, and v is a parameter describing the aspect ratio consistency between the predicted bounding box and the ground truth bounding box. Then the bounding box loss function L B The calculation formula is as follows:

[0024] L B = 1 - CIOU

[0025]

[0026] Furthermore, during the training of the YOLOv5s network model, its initial parameters are as follows: set epochs = 300, batch-size = 16, so that the model undergoes 300 batches of training, with 16 images each time; set the parameter rect to use rectangular training, set the parameter device = cuda, so that the model is trained on the GPU; set evolve for hyperparameter evolution, set img-size = [640, 640], set the learning rate lr0 = 0.01, lrf = 0.2; set the warm-up training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.1.

[0027] Furthermore, during the production of the dataset, set the mosaic data augmentation related parameter mosaic = 1.0, the probability of horizontal flipping of the image fliplr = 0.5, hsv_h = 0.015, hsv_s = 0.7, hsv_v = 0.4.

[0028] Furthermore, mosaic data augmentation stitches together four randomly cropped images, changes the image color, and increases the diversity of the image data by horizontal flipping of the images, making the trained model more applicable.

[0029] Furthermore, the best weight data of the YOLOv5s network model is obtained by statistically analyzing the gradient information of the YOLOv5s loss function and iteratively updating the network model parameters multiple times to make the function tend to a minimum value.

[0030] The present invention improves the YOLOv5 model and uses the improved YOLOv5s to perform real-time detection of the ages of squabs, avoiding the large amount of effort consumed in designing features by traditional machine learning algorithms. At the same time, a large amount of data is used for end-to-end training tasks. The spatial attention in CBAM is improved to pointwise spatial attention and embedded as the backbone residual branch of YOLOv5, strengthening the recognition efficiency of overlapping squabs and improving the recognition accuracy. The NMS strategy of the YOLOv5 algorithm is improved to avoid detecting a single pigeon at multiple different growth stages, further improving the recognition accuracy. The pixel variance is calculated using Laplace transform to screen out clear pictures for model training, avoiding the influence of blurred pictures caused by irresistible factors such as light and jitter on the training effect of the model. Overall, the model of the present invention is small in size, fast in detection speed, and the data is collected from the real pigeon breeding environment, providing technical support for the machine to quickly identify the feed input amount of each pigeon cage, and can also be applied to the breeding of other poultry, promoting the automation of the breeding industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic flowchart of an embodiment of the present invention.

[0032] Figure 2 It is a diagram of the improved YOLOv5s network model in an embodiment of the present invention.

[0033] Figure 3 It is a diagram showing examples of five growth stages of squabs in an embodiment of the present invention.

[0034] Figure 4 It is a diagram showing the distribution of the number of squab samples in an embodiment of the present invention.

[0035] Figure 5 It is a diagram of an example after mosaic data augmentation in an embodiment of the present invention.

[0036] Figure 6 It is a schematic diagram showing the combination method of CBAM and YOLOv5s in an embodiment of the present invention.

[0037] Figure 7 (a) It is the standard CBAM structure diagram in an embodiment of the present invention; Figure 7 (b) It is the improved CBAM structure diagram in an embodiment of the present invention.

[0038] Figure 8 (a) It is the spatial attention structure diagram of the standard CBAM in an embodiment of the present invention; Figure 8 (b) It is the pointwise spatial attention structure diagram of the improved CBAM in an embodiment of the present invention.

[0039] Figure 9(a) is the heat map of the three detection layers of the standard YOLOv5s in the embodiment of the present invention; Figure 9 (b) is the heat map of the three detection layers of YOLOv5s integrated with CBAM in the embodiment of the present invention; Figure 9 (c) is the heat map of the three detection layers of YOLOv5s combined with the improved CBAM in the embodiment of the present invention.

[0040] Figure 10 (a) is the true label in the embodiment of the present invention; Figure 10 (b) is the detection result of the standard YOLOv5s in the embodiment of the present invention; Figure 10 (c) is the detection result of YOLOv5s integrated with CBAM in the embodiment of the present invention; Figure 10 (d) is the detection result of YOLOv5s combined with the improved CBAM in the embodiment of the present invention.

[0041] Figure 11 is the schematic diagram of the improved NMS strategy in the embodiment of the present invention.

[0042] Figure 12 (a) is the detection effect of the unimproved NMS strategy in the embodiment of the present invention; Figure 12 (b) is the detection effect of the improved NMS strategy in the embodiment of the present invention.

[0043] Figure 13 (a) is the true label map in the embodiment of the present invention; Figure 13 (b) is the heat map of the three detection layers of class 1 exported under the improved NMS; Figure 13 (c) is the heat map of the three detection layers of class 2 exported under the improved NMS.

[0044] Figure 14 is the partial result display diagram of target recognition using the improved YOLOv5s network model in the embodiment of the present invention.

[0045] Figure 15 is the evaluation index diagram of target recognition using the improved YOLOv5s network model in the embodiment of the present invention. Detailed implementation manners

[0046] A method for detecting the age of squabs, as Figure 1 shown, includes the following steps:

[0047] (1) Improve the YOLOv5s network model. First, improve the spatial attention in CBAM to point-wise spatial attention, and then embed the improved CBAM attention mechanism at the end of each residual branch structure in the backbone network of the YOLOv5s model. Improve the NMS strategy of YOLOv5s from performing NMS separately for each class to performing NMS for multiple classes simultaneously, and delete the allocation intervals for different classes to obtain the improved YOLOv5s network model.

[0048] (2) Dataset preparation. Use LabelImg to annotate clear pigeon pictures according to the pigeon's age stage, and then randomly allocate the picture set into a training set and a validation set according to a certain ratio to obtain the initial pigeon dataset. Then perform mosaic data augmentation on the initial dataset to obtain the dataset after data augmentation.

[0049] In this embodiment, the clear pigeon pictures are obtained by the following method: collect pigeon pictures at the breeding site, calculate the pixel variance of each picture through Laplace transform, use the pixel variance as the index value of the picture blurriness, discard the pictures with an index value lower than 20 as blurred pictures, and retain the rest as clear pictures.

[0050] The age stages of pigeons can be divided according to the actual breeding situation of each farm. For example, in this embodiment, they can be divided into five stages. Specifically, the age stages of pigeons include stage one, stage two, stage three, stage four, and stage five. Stage one is 1 - 4 days old, stage two is 5 - 8 days old, stage three is 9 - 14 days old, stage four is 15 - 20 days old, and stage five is from 21 days old to slaughter and market, corresponding to the five categories finally output by the recognition model. The sample pictures of each stage are as Figure 3 shown. In this embodiment, random allocation is performed according to the ratio of 70% for the training set and 30% for the validation set. The number of samples in each stage is as Figure 4 shown.

[0051] In this embodiment, during dataset preparation, set the relevant parameters for mosaic data augmentation: mosaic = 1.0, the probability of horizontal flipping of pictures fliplr = 0.5, hsv_h = 0.015, hsv_s = 0.7, hsv_v = 0.4. Mosaic data augmentation increases diverse picture data by randomly cropping and splicing four pictures, changing the picture color, and horizontally flipping the pictures, making the trained model have stronger applicability. The sample pictures after augmentation are as Figure 5 shown.

[0052] (3) Model training. Input the dataset into the improved YOLOv5s network model for model training to obtain the best weight data of the improved YOLOv5s network model.

[0053] (4) Squab age detection. Load the optimal weight data into the improved YOLOv5s network model, input an image to be recognized containing a squab, and obtain an output image with one or more rectangular boxes identifying the squab age.

[0054] As Figure 2 , the overall structure of YOLOv5s is divided into three parts: backbone, neck, and head. The backbone uses CSPDarknet53 that integrates SPP and Focus to extract features; the neck uses FPN and PAN to transmit the feature maps of different sizes from the backbone upward and then fuse them downward, so as to fully combine the information of feature maps of different sizes; finally, YOLO Head is used for detection.

[0055] Figure 2 In, for the backbone and neck parts, Focus is a module that slices the image before it enters the network, replaces the original convolutional layer, stacks the spatial information into the channel space, and is used to reduce FLOPS and increase speed; the SPP module extracts and fuses feature maps of different scales, which can effectively improve the receptive field of the backbone network; the Bottleneck module uses a residual structure to solve the problem of network degradation; the C3 module contains n Bottleneck modules to enhance the learning ability of the network and reduce the computational cost. In the Head part, YOLOv5s adopts a prediction mode of multi-layer feature maps and focuses on objects of different sizes according to the size of the feature maps. Figure 2 What is shown is the prediction mode of three-layer feature maps. YOLOv5s performs a 1×1 convolution on each of the three layers of feature maps. In the present invention, the output channels are 3×(4 + 1 + 5), divided into three layers of anchors. Each layer of anchor contains 4 coordinate information (x, y, w, h), 1 object score, and 5 classification scores. The highest one of the five classification scores is multiplied by the object score to obtain the probability of this classified object. According to the coordinate information and classification, NMS processing is performed, and redundant boxes are removed according to the IOU threshold to obtain the final result.

[0056] In this embodiment, the calculation formula of the loss function L of the YOLOv5s model is:

[0057]

[0058] Among them, L B , L c , L O respectively represent the bounding box loss, class loss, and object loss. β, γ, are all constants used to control the weights of the three losses, and are respectively set to 0.05, 0.5, and 1.0;

[0059] Class loss L C is consistent with the target loss L O in calculation, both are calculated by binary cross-entropy. Let x n be the class label output by the model, and y n be the true class label. In L O , x n is the object label output by the model, and y n is the true object label. Taking the target loss L C as an example, its calculation formula is as follows:

[0060] L C = -∑[y n × ln x n + (1 - y n ) × ln(1 - x n )]

[0061] Bounding box loss L B adopts the CIOU calculation method. Let IOU be the intersection over union between the predicted box and the ground truth box, ρ 2 (b, b gt ) be the Euclidean distance between the centers of the predicted box and the ground truth box, c is the diagonal distance of the smallest enclosing region of the predicted box and the ground truth box, α is the parameter for balancing the ratio, and v is the parameter for describing the aspect ratio consistency of the predicted box and the ground truth box. Then the bounding box loss function L B has the following calculation formula:

[0062] L B = 1 - CIOU

[0063]

[0064] In this embodiment, during the training of the YOLOv5s network model, its initial parameters are as follows: set epochs = 300, batch-size = 16, so that the model is trained for 300 batches, with 16 images each time; set the parameter rect to use rectangular training, set the parameter device = cuda, so that the model is trained on the GPU; set evolve for hyperparameter evolution, set img-size = [640, 640], set the learning rate lr0 = 0.01, lrf = 0.2; set the warm-up training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.1.

[0065] The best weight data of the YOLOv5s network model is obtained by statistically analyzing the gradient information of the YOLOv5s loss function and iteratively updating the network model parameters multiple times to make the function tend to the minimum value.

[0066] Let the average precision be mAP (average precision), the precision be P, and the recall be R. The calculation formulas for mAP, P, and R are as follows:

[0067]

[0068] Among them, TP is the number of positive classes determined as positive classes, FP is the number of negative classes determined as positive classes, FN is the number of positive classes determined as negative classes, and Q(P, R): a function that calculates the area of the P-R curve formed by P and R for each category.

[0069] CBAM (Convolutional Block Attention Module) is an attention mechanism based on convolutional neural networks. Its role is to make the network more inclined to pay attention to a certain place, increase the weight of useful features, and at the same time suppress the weight of invalid features, and at the same time, it does not increase additional overhead. Compared with channel attention, it adds the position information of the feature map, and compared with spatial attention, it adds the channel information of the feature map.

[0070] The present invention introduces CBAM into YOLOv5s, so that the attention of the network model is concentrated on the squab object, reducing the influence of light, foreign objects, etc. in the picture. The spatial attention in CBAM is improved to pointwise spatial attention, which can avoid the problem of the average pooling layer in the spatial attention of CBAM blurring the boundary information between overlapping squabs under limited arithmetic expressions. The improved CBAM becomes a new attention structure composed of channel attention and pointwise spatial attention. Let the input feature map be F, and the two changes be F′, F″, F, F′, F″ ∈ R c*w*H , the channel attention module is H ∈ R c*1*1 , the spatial attention module is S ∈ R 1*w*H , the symbol represents the multiplication operation of corresponding elements through the broadcast mechanism, then the formula of CBAM is:

[0071]

[0072] The formula expression of the improved CBAM is the same as that of CBAM, only the internal spatial attention structure is different. Figure 6 is the position where CBAM is embedded in each residual structure of the YOLOv5s model. Figure 7 (a) is the structure of the standard CBAM. Figure 7 (b) is the structure of the improved CBAM. Figure 8 (a) is the spatial attention structure of the standard CBAM. Figure 8(b) is the point-wise spatial attention structure of the improved CBAM, which improves the spatial attention in Figure 7 (a) to Figure 7 the point-wise attention in Figure 9 (b). The recognition results of the traditional YOLOv5s model, the YOLOv5s model embedded with the standard CBAM, and the YOLOv5s model embedded with the improved CBAM are as shown in Figure 10 . Among them, Figure 10 (a) is the ground truth label, Figure 10 (b) is the detection result of the traditional YOLOv5s, Figure 10 (c) is the detection result of the YOLOv5s embedded with the standard CBAM, Figure 10 (d) is the detection result of the YOLOv5s embedded with the improved CBAM. Figure 9 (a), Figure 9 (b), and Figure 9 (c) are the detection heatmaps of the traditional YOLOv5s, the YOLOv5s embedded with the standard CBAM, and the YOLOv5s embedded with the improved CBAM, respectively. Table 1 shows the comparison of the detection effects of the above different models.

[0073] Table 1 Comparison of the detection effects of different models

[0074]

[0075] It can be clearly observed that YOLOv5s has deficiencies in the ability to extract effective features. The information extracted by the three detection layers ( Figure 9 (a)) is relatively scattered, resulting in the omission of a pigeon in the upper left corner of Figure 10 (b). After introducing CBAM, it is found that Figure 9 (b) extracts more concentrated information. The features extracted by the last detection layer can correspond to three squabs, and the squab in the upper left corner can be detected. Therefore, the P and mAP in Table 1 increase slightly. However, due to the insufficient expression of the extracted feature information, the boundary information of the two pigeons above is adhered, and the adhered information is recognized as a pigeon, so the R decreases. By improving the spatial attention in CBAM to point-wise spatial attention, the feature information is fully expressed, making Figure 9 the heatmap information of the three detection layers in

[0076] (c) the most concentrated. The heatmap on the far right not only corresponds to the positions of the three pigeons but also distinguishes the boundaries between them, correctly detecting the three pigeons.The present invention improves the NMS strategy of YOLOv5s from single-class partition NMS to multi-class centralized NMS, thus avoiding the situation where the model misjudges a squab as two or more squabs at different stages due to the similarity of squabs in adjacent growth stages. In practical applications, this method avoids the misjudgment of the number of squabs.

[0077] The traditional NMS strategy of YOLOv5 only operates between labels of the same class. The final result may retain labels of different classes with a relatively high degree of overlap. However, the improved NMS strategy of the present invention operates on all classes, and the final result retains the label with the highest score among the overlapping labels, thus avoiding the problem of multiple-class overlap in the predicted labels. The specific method is as follows:

[0078] Let B, B′∈R 1*4 , where B is the bounding box information, which includes the center point coordinates (x, y) and the width and height (w, h) of the bounding box; B′ is the new bounding box information; I is the fixed position of the class with the highest score among all detected classes; L is the maximum value of the width and height of the bounding box, set to 4096, which only participates in the operation as a value larger than the normal calculated value. The formula for the NMS bounding box locking interval of traditional YOLOv5s is: B′ = B + I×L.

[0079] The improved YOLOv5s of the present invention removes the method of locking the interval through I×L in the formula for the NMS bounding box locking interval of traditional YOLOv5s, and improves from performing NMS processing for each class separately to performing NMS processing for multiple classes simultaneously. The improved NMS interval determination formula is: B′ = B. This makes the object to be detected only retain one candidate box with the highest IOU score among all classes, and eliminates all other generated candidate boxes, fundamentally avoiding the possibility of detecting a squab at two stages due to the similarity of adjacent growth stages.

[0080] The schematic diagrams of NMS for single-class partition and multi-class simultaneous NMS are shown in Figure 11 , where the red box is the ground truth label, the green box is class 1, and the blue box is class 2.

[0081] The recognition effects of the unimproved NMS and the improved NMS are respectively as shown in Figure 12 (a) and Figure 12 (b). Figure 12 The second picture in (a) is the stretched picture of the first picture. Figure 12 The second picture in (b) is the stretched picture of the first picture. As can be seen from Figure 12 (a), the unimproved NMS misidentifies a squab in the lower left corner as two squabs at different stages (stage2 and stage3 respectively). As can be seen from Figure 12(b) It can be seen that the improved NMS can correctly retain the class with the highest score (stage3), directly curbing the possibility of the model identifying multiple different stages. Figure 13 The heatmaps of two classes derived for one of the pictures under three-layer vision, Figure 13 (a) is the true label of the picture. The darker the red, the more attention the model devotes here, and the darker the blue, the less attention the model devotes. And Figure 13 In the heatmaps derived for the two classes, the dark red pixel areas can be concentrated at the center of the true label, indicating that the model detection ability in the training of the present invention performs well. The final detection results of YOLOv5s with the improved NMS strategy are shown in Table 2.

[0082] Table 2 Final detection results of YOLOv5s with the improved NMS strategy

[0083] Class P R mAP all 0.951 0.946 0.936 Stage1 0.967 0.942 0.941 Stage2 0.925 0.961 0.902 Stage3 0.971 0.926 0.937 Stage4 0.955 0.971 0.97 Stage5 0.935 0.93 0.93

[0084] The improved YOLOv5s network model obtained by the improved CBAM attention and the improved NMS strategy. The structure of the improved YOLOv5s network model is as Figure 2 shown. After training, the input pictures are detected. Figure 14 For the display of some recognition results, among them, Figure 14 (a) is the true label, Figure 14 (b) is the predicted label, Figure 15 are the object detection evaluation metrics. It can be seen from Figure 14 that using the improved YOLOv5s network model of the present application, the recognition result of the target after training has a high consistency with the true label, indicating that it has good object recognition ability. Figure 15 The mAP (average precision) curve in []] reaches a stable state after about 70 epoches and starts to rise slightly to 0.944. Its P (precision) and R (recall) curves also reach a stable state after about 70 epoches and start to rise slightly to 0.959 and 0.953 respectively. From the perspective of the validation set loss function, the position (box) loss reaches a stable state after about 200 epoches, while the objectness loss and classification loss show an overfitting tendency after about 100 epoches and 120 epoches respectively. The overall curve performs well. Finally, the epoch with the highest mAP in the validation set is selected as the best model weight data. The final detection results are shown in Table 3.

[0085] Table 3 Final detection results of the improved YOLOv5s

[0086] Stage P R mAP F1 Stage1 0.982 0.948 0.956 0.965 Stage2 0.925 0.961 0.903 0.943 Stage3 0.959 0.965 0.953 0.962 Stage4 0.963 0.983 0.987 0.973 Stage5 0.962 0.910 0.921 0.935

[0087] The above detailed description is a specific description of the feasible embodiments of the present invention. Such embodiments are not intended to limit the patent scope of the present invention. Any equivalent implementation or modification without departing from the present invention shall be included in the patent scope of this case.

Claims

1. A method for detecting the age of squabs, characterized in that, it includes the following steps: Improve the YOLOv5s network model, improve the spatial attention in CBAM to pointwise spatial attention, and the improved CBAM becomes a new attention structure combining channel attention and pointwise spatial attention. Then embed the improved CBAM attention mechanism at the end of each residual branch in the backbone network of the YOLOv5s model. Improve the NMS strategy of YOLOv5s from performing NMS for each class separately to performing NMS for multiple classes simultaneously, and delete the allocation intervals for different classes to obtain an improved YOLOv5s network model; Dataset production: Use LabelImg to label clear squab pictures according to the age stages of squabs, and then randomly allocate the picture set into a training set and a validation set according to a certain ratio to obtain an initial squab dataset. Then perform mosaic data augmentation on the initial dataset to obtain a data-augmented dataset; Model training: Input the dataset into the improved YOLOv5s network model for model training to obtain the best weight data of the improved YOLOv5s network model; Squab age detection: Load the best weight data into the improved YOLOv5s network model, input a to-be-recognized picture containing a squab, and obtain an output picture with one or more rectangular boxes identifying the age.

2. A method for detecting the age of squabs according to claim 1, characterized in that, The method of improving the NMS strategy from performing NMS for each class separately to performing NMS for multiple classes simultaneously and deleting the allocation intervals for different classes is: Delete the allocation interval I×L in the NMS bounding box locking interval formula B′ = B + I×L. The improved NMS interval determination formula is: B′ = B, so that only the candidate box with the highest IOU score among all classes for the object to be detected is retained, and all other generated candidate boxes are eliminated; IOU refers to the intersection over union, which is the ratio of the intersection and union of the predicted box and the ground truth box.

3. A method for detecting the age of squabs according to claim 1, characterized in that, In the CBAM, let the input feature map be F, and the two changes be F' and F″, where F, F′, F″ ∈ R c*w*H , the channel attention module is H ∈ R c*1*1 , the spatial attention module is S ∈ R 1*w*H , the symbol represents the multiplication operation on the corresponding elements through the broadcasting mechanism. Then the formula of CBAM is:

4. A method for detecting the age of squabs according to claim 1, characterized in that, The clear squab pictures are obtained by the following method: Collect squab pictures at the breeding site, calculate the pixel variance of each picture through Laplace transform, use the pixel variance as an index value of the picture blurriness, discard the pictures with an index value lower than 20 as blurred pictures, and retain the rest as clear pictures.

5. A method for detecting the age of squabs according to claim 1, characterized in that, The age stages of the squabs include stage one, stage two, stage three, stage four, and stage five. Stage one is from 1 to 4 days old, stage two is from 5 to 8 days old, stage three is from 9 to 14 days old, stage four is from 15 to 20 days old, and stage five is from 21 days old to slaughter and market.

6. A method for detecting the age of squabs according to claim 1, characterized in that, The calculation formula of the loss function L of the YOLOv5s model is: L = βL B + γL c + ζL O Among them, L B , L C , L O respectively represent the bounding box loss, the class loss, and the object loss. β, γ, and ζ are all constants used to control the weights of the three losses, which are set to 0.05, 0.5, and 1.0 respectively; Category loss L C is consistent with the target loss L O and both are calculated by binary cross-entropy. Let x n be the category label output by the model, and y n be the true category label. In L O , x n is the object label output by the model, and y n is the true object label. Taking the target loss L C as an example, its calculation formula is as follows: L C = -∑[y n ×ln x n + (1 - y n ) × ln(1 - x n )] Bounding box loss L B Using the calculation method of CIOU, let IOU be the intersection over union between the predicted box and the ground truth box, and ρ 2 (b, b gt ) is the Euclidean distance between the centers of the predicted box and the ground truth box, c is the diagonal distance of the smallest enclosing region of the predicted box and the ground truth box, α is the parameter for balancing the ratio, and v is the parameter for describing the aspect ratio consistency between the predicted box and the ground truth box. Then the bounding box loss function L B The calculation formula is as follows: L B = 1 - CIOU 7. A method for detecting the age of squabs according to claim 1, characterized in that, During the training of the YOLOv5s network model, its initial parameters are as follows: Set epochs = 300, batch-size = 16, so that the model undergoes 300 batches of training, with 16 images each time; Set the parameter rect to use rectangular training, set the parameter device = cuda, so that the model is trained on the GPU; Set evolve for hyperparameter evolution, set img-size = [640, 640], set the learning rate lr0 = 0.01, lrf = 0.2; Set the warm-up training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.

1.

8. A method for detecting the age of squabs according to claim 1, characterized in that, During the production of the dataset, set the relevant parameters of mosaic data augmentation: mosaic = 1.0, the probability of horizontal flipping of the image fliplr = 0.5, hsv_h = 0.015, hsv_s = 0.7, hsv_v = 0.

4.

9. A method for detecting the age of squabs according to claim 1, characterized in that, The mosaic data augmentation randomly crops and stitches four images, changes the image color, and increases diverse image data by horizontally flipping the images, so that the trained model has stronger applicability.

10. A method for detecting the age of squabs according to claim 1, characterized in that, The best weight data of the YOLOv5s network model is obtained by statistically calculating the gradient information of the YOLOv5s loss function and iteratively optimizing the network model parameters multiple times to make the function tend to the minimum value.

Citation Information

Patent Citations

  • Pigeon egg quality identification method

    CN113538389A

  • Mask wearing detection method based on YOLOv5 network

    CN114399799A