Dead Chicken Detection Method Based on Image Fusion and Improved YOLOv9

By using image fusion, improving YOLOv9 model, attention mechanism and adaptive convolution kernel in dead chicken detection technology, the misidentification and misidentification problems of dead chicken detection in high-density breeding environments are solved, and higher detection accuracy and robustness are achieved.

CN119251865BActive Publication Date: 2025-05-30CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411239415.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-05-30
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

The existing dead chicken detection technology has problems with misidentification and misidentification in high-density breeding environments, especially in the case of uneven light and occlusion caused by structures such as multi-layer chicken cages, baffles and sinks.

Method used

Using the dead chicken detection method based on image fusion and improved YOLOv9, thermal infrared and visible light images are fused through PIAFusion technology, combined with EMA attention mechanism and Rep-DCNv3 module, feature extraction and object detection capabilities are enhanced, and target positioning is improved using MPDIoU loss function.

Benefits of technology

It significantly improves the accuracy and robustness of dead chicken detection, reduces the misidentification and misidentification rates, adapts to complex lighting environments and occlusion conditions, and improves the efficiency and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251865B_ABST
    Figure CN119251865B_ABST
Patent Text Reader

Abstract

The present invention discloses a dead chicken detection method based on image fusion and improved YOLOv9. First, the PIAFusion technology is used to achieve the fusion of thermal infrared and visible light images, so as to improve the saliency of dead chicken target features and eliminate the influence of uneven illumination. Secondly, the EMA attention mechanism is introduced into the neck of YOLOv9 to effectively distinguish the targets in low-light environments from dead chicken targets, thereby improving the accuracy and generalization ability of the model. Then, the Rep-DCNv3 module is introduced to enhance the backbone feature extraction network and better extract the features of dead chickens in cases of partial occlusion. Finally, MPDIoU is used to replace the loss function of YOLOv9 to more accurately evaluate the degree of overlap between targets. The dead chicken detection method of the present invention solves the problems of low automation degree of dead chicken patrol inspection and time-consuming and laborious manual patrol inspection in large-scale breeding environments. It not only ensures a high detection speed but also significantly improves the accuracy of target detection, and can meet the requirements of real-time detection of dead chickens in actual production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of dead chicken detection and computer deep learning, and particularly relates to a dead chicken detection method based on image fusion and improved YOLOv9. Background Art

[0002] Chicken is the second largest meat consumer in China after pork. With the continuous improvement of people's requirements for food safety and quality, China's broiler industry is gradually transforming and upgrading towards the direction of intelligence and greenness. At present, most domestic farms adopt a multi-layer cage farming mode with a high breeding density. Most still rely on manual inspections to discover and remove dead chickens, which has problems such as high work intensity and strong subjectivity. There will also be a phenomenon of dead chickens remaining due to visual fatigue caused by long-term work of breeding personnel. If dead poultry cannot be processed in time, a large number of germs will be generated and may cause serious disease transmission risks, such as large-scale influenza epidemics.

[0003] In recent years, significant progress has been made in non-contact dead chicken recognition technology, mainly including sound monitoring methods and machine vision methods. The sounds emitted by poultry can reflect their health status and emotional state, and thus identify dead chickens. However, affected by the noise of the chicken flock, the recognition accuracy of the sound monitoring method is relatively low. The machine vision method identifies the behavior or state of chickens through image analysis, with high efficiency and accuracy. The infrared imaging technology combined with neural networks has realized a dead chicken self-recognition system based on machine vision; the research on poultry pose estimation systems and the combination of morphological features and machine learning algorithms has significantly improved the accuracy of dead chicken recognition. Nevertheless, the complexity of the breeding environment and technical limitations are still challenges for improving the current recognition accuracy.

[0004] In the monitoring of multi-layer cage chickens, the fusion of infrared and visible light images is the key to improving the accuracy of dead chicken detection. Infrared images capture the temperature distribution, while visible light images provide morphological and color details. The fused image can significantly enhance the integrity of the model input and improve the robustness of the detection. Image fusion technology is divided into traditional methods and deep learning methods. Traditional methods are mainly divided into pixel operations, transform domains, image decomposition, and energy functions. Although effective in specific situations, they are limited by artificial strategies and have limited adaptability. Deep learning methods, such as the image fusion framework of generative adversarial networks, the image fusion framework of convolutional neural networks, and the image fusion framework of autoencoders, can provide more flexible fusion strategies. The complex lighting, multi-layer structure, and obstacles in the farm, as well as the occlusion between chickens, result in uneven lighting, which affects the accuracy of dead chicken recognition. The Progressive Image Fusion based on Illumination Awareness (PIAFusion) effectively reduces the impact of uneven lighting by dynamically adjusting the fusion parameters to adapt to different lighting conditions, better adapts to the complex lighting environment, and provides a progressive and highly adaptable image fusion solution for dead chicken detection.

[0005] In the field of object detection, algorithms are mainly divided into two categories: single-stage and two-stage. The representative of two-stage object detection algorithms is the Region-based Convolutional Neural Network (R-CNN) and its derivative algorithms. Single-stage algorithms have an advantage in real-time applications due to their fast and accurate detection capabilities. Since the release of the YOLO series in 2015, its structure and performance have been continuously optimized through successive iterations. YOLOv9, released in February 2024, includes four main structures: Backbone, Neck, Head, and an auxiliary branch. Through innovations in key structures, it solves the information bottleneck problem in deep learning and significantly improves the detection speed and performance.

[0006] Current dead chicken detection technologies face multiple challenges, mainly including: 1. Existing dead chicken recognition technologies mainly rely on images from a single source, such as thermal infrared imaging or visible light images. In high-density farming scenarios, the presence of structures such as multi-layer chicken coops, baffles, and water troughs has a significant impact on the light distribution, which may cause target detection algorithms to have difficulty distinguishing actual dead chickens from visual errors caused by the environment, greatly increasing the risk of misidentifying dead chickens. 2. Traditional dead chicken detection methods have high false detection rates and missed detection rates when dead chickens are blocked by other live chickens. Traditional dead chicken detection methods usually rely on morphological features and texture information in the images. When dead chickens are blocked by other live chickens, these features cannot be effectively captured, and they are not robust enough in dealing with occlusion and complex backgrounds, resulting in a significant increase in false detection rates and missed detection rates. 3. When target detection algorithms face multiple dead chickens generating overlapping regression boxes, the false detection rates and missed detection rates are high. Especially in high-density farming environments, there is a large amount of mutual occlusion or partial overlap between target dead chickens, resulting in the detection algorithm generating overlapping regression boxes, which cannot accurately cover the entire body of each dead chicken, thus affecting the accuracy and reliability of the detection. Summary of the Invention

[0007] The problem to be solved by the present invention is to provide a dead chicken detection method based on image fusion and improved YOLOv9. Taking broiler farming in a stacked cage rearing mode as the research object, an image fusion technology is adopted to register and fuse thermal infrared images and visible light images, and the recognition of dead chickens is realized through a target detection algorithm, achieving higher efficiency and accuracy.

[0008] The present invention adopts the following technical solutions: A dead chicken detection method based on image fusion and improved YOLOv9, comprising the following steps:

[0009] S1. Collect chicken coop picture data, and adopt a progressive image fusion method based on illumination perception to fuse visible light and infrared light, improve the saliency of dead chicken target features, and eliminate the influence of uneven illumination caused by multi-layer chicken coops, baffles, and water trough structures;

[0010] S2. Build a dead chicken detection model based on the YOLOv9 object detection algorithm. Introduce the EMA attention mechanism into the neck of YOLOv9 to perform exponential smoothing on the model parameters, distinguish the low-light environment and the dead chicken target, and fuse targets of different sizes.

[0011] S3. Introduce the Rep-DCNv3 module into the dead chicken detection model to enhance the backbone feature extraction network and extract the features of dead chickens in the case of partial occlusion.

[0012] S4. Use MPDIoU to replace the loss function of YOLOv9. By evaluating the overlap degree between targets, provide more fine-grained error feedback and improve the detection accuracy of the dead chicken detection model.

[0013] Specifically, in step S1, for the progressive image fusion method based on illumination perception, the backbone network is built on an end-to-end CNN framework, including a feature extractor and an image reconstructor.

[0014] The feature extractor is used to mine the complementary and common features of infrared images and visible light images. Its structure contains five convolutional layers. The first layer uses a 1×1 convolutional layer to process the infrared image and the visible light image respectively to reduce the differences between different modality images. The second to fifth layers use four convolutional layers with shared weights to extract the depth features of infrared and visible light images respectively.

[0015] The depth features extracted from the infrared image and the visible light image are pooled and connected, and then input into the image reconstructor. The image reconstructor consists of five convolutional layers, without introducing downsampling, and the size of the fused image is the same as that of the source image. It is used to integrate common and complementary information to generate a fused image to adapt to the complex lighting environment caused by multi-layer chicken coops, baffles, and water troughs.

[0016] Furthermore, after the second, third, and fourth layers of the feature extractor, a CMDAF module is introduced respectively to comprehensively extract the features of infrared and visible light images and realize the exchange of modality complementary features.

[0017] Specifically, in step S2, the EMA attention mechanism is used to enrich the feature extraction and attention weight allocation of the dead chicken detection model at different scales. The method is as follows:

[0018] S2.1. In the case of a live chicken occluding a dead chicken, through convolutional kernels of different sizes, adapt to semantic information at different levels, perform convolutional operations, and capture multi-scale features. Smaller-sized convolutional kernels are used to extract edge and texture detail features, and larger-sized convolutional kernels are used to identify large-scale patterns and shapes in the image.

[0019] S2.2. Introduce an attention module, and through a lightweight network architecture, learn and evaluate the contribution degrees of features at each scale to achieve dynamic weighting of features;

[0020] S2.3. Integrate a subsampling method to reduce the feature dimension through downsampling, reduce the computational amount, and at the same time ensure the integrity and accuracy of feature representation;

[0021] S2.4. In the feature fusion stage, integrate the weighted multi-scale features, retain the multi-scale features, strengthen the model's recognition ability for key information through a weighting mechanism, form the final feature representation, reduce the missed detection of dead chicken targets caused by live chicken occlusion, reduce the interference of live chickens and chicken cage backgrounds on the dead chicken detection model, and improve the accuracy and robustness of detection;

[0022] S2.5. Through optimizing information processing, reduce information loss and confusion in the feature extraction process, and improve the adaptability to complex environments and the accuracy of detection results by strengthening the model's extraction and utilization of key information.

[0023] Specifically, in step S3, in the Rep-DCNv3 module, the convolutional kernel in the RepNCSPELAN module of the YOLOv9 algorithm is replaced by DCNv3 in the Backbone feature extraction network to perform feature capture when the dead chicken detection model processes heavily occluded targets;

[0024] The DCNv3 dynamically adjusts through an adaptive convolutional kernel to adapt to the target shape. In the case where the dead chicken is occluded by the live chicken, it accurately captures the target features and improves the recognition ability for occluded targets; through the offset prediction network of DCNv3, it enhances the adaptability to target position and shape changes and improves the detection accuracy; by optimizing bilinear interpolation to linear interpolation on the axis, it reduces the computational amount and avoids information distortion caused by interpolation.

[0025] Furthermore, the network structure of the Rep-DCNv3 module includes: LN, FFN, and GELU components, as well as convolutional layers and downsampling layers for obtaining hierarchical feature maps.

[0026] First, the input multi-scale feature data is successively subjected to preliminary feature extraction through convolutional layers, and then dynamically adjusted through the adaptive convolutional kernel in Rep-DCNv3 to adapt to the target shape and accurately capture the target features, especially in the case where the dead chicken is occluded by the live chicken.

[0027] Subsequently, the feature map enhances the non-linear expression ability through FFN and GELU, and performs normalization processing through LN to stabilize the training process.

[0028] Finally, the spatial dimension of the feature map is gradually reduced through the downsampling layer, while the depth of the feature map is increased to capture more abstract feature representations for visual task processing, obtaining high-dimensional feature representations for object detection, which are subsequently used for the precise recognition and localization of dead chicken targets.

[0029] Furthermore, the Rep-DCNv3 module involves an offset prediction network and deformable convolution operations. First, it predicts the offset of the input position, and then performs an adaptive convolution operation based on the offset. The principle is as follows:

[0030]

[0031] where H represents the total number of aggregation groups; k h represents the position-independent projection weight of the h-th group; q hm represents the modulation scalar of the m-th sampling point in the h-th group; x h represents the input feature map of the slice; Δr hm represents the offset corresponding to the grid sampling position r m in the m-th group; y(r 0 ) represents the output.

[0032] Specifically, in step S4, the MPDIoU is an improved IoU metric method used to accurately evaluate the overlap degree between targets in a high-density farming environment, balance the detection performance between different categories, and reduce overfitting to the majority class, including the following sub-steps;

[0033] S4.1. By considering the partial overlap and distance information between dead chicken targets, evaluate the matching degree MPDIoU of the target boxes, calculated as follows:

[0034]

[0035] where w and h are the width and height of the input image; A and B are two arbitrary convex shapes, d 1 is the Euclidean distance from the upper left corner of shape B to the upper left corner of shape A, and d 2 is the Euclidean distance from the lower right corner of shape B to the lower right corner of shape A; are the upper left coordinates and lower right coordinates of A respectively, are the upper left coordinates and lower right coordinates of B respectively;

[0036] S4.2. Based on MPDIoU, define the loss function as follows:

[0037] L MPDIoU = 1 - MPDIoU

[0038] S4.3. During the training process of the dead chicken detection model, by introducing the MPDIoU loss, learn a more accurate target box localization strategy, generate more precise detection results, and enhance the adaptability of the dead chicken detection model to complex scenarios.

[0039] The technical solution of the present invention also provides: an electronic device, including:

[0040] One or more processors;

[0041] A storage device having one or more programs stored thereon;

[0042] When the one or more programs are executed by the one or more processors, the one or more processors implement the dead chicken detection method based on image fusion and improved YOLOv9 described in any of the above.

[0043] The technical solution of the present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps in any of the above dead chicken detection methods based on image fusion and improved YOLOv9 are implemented.

[0044] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0045] 1. The dead chicken detection method based on image fusion and improved YOLOv9 of the present invention takes the broiler chicken breeding in the stacked cage culture mode as the research object. Aiming at the dead chicken detection task of caged chickens and the deficiencies such as low accuracy of dead chicken recognition in the existing technology, a dead chicken detection method based on multi-source information fusion is designed.

[0046] 2. The method of the present invention uses PIAFusion to fuse visible light images and infrared light images, effectively solving problems such as uneven illumination and complex backgrounds; at the same time, the EMA (Exponential Moving Average) attention mechanism is incorporated into the target detection algorithm to distinguish low-dark targets such as chicken cage baffles and waterers from dead chicken targets and achieve the fusion of targets of different sizes.

[0047] 3. The method of the present invention introduces the Rep-DCNv3 module in the feature extraction stage of target detection, effectively extracting the features of dead chickens blocked by live chickens. This innovative module solves the problem that traditional methods are difficult to accurately detect when dead chickens are blocked, significantly improving the detection accuracy and robustness.

[0048] 4. The method of the present invention uses the MPDIoU (Modified Partial Distance-IoU) technology as the loss function of the object detection algorithm, which can more accurately evaluate the overlapping degree between objects, balance the detection performance between different categories, reduce the overfitting of the model to the majority class, and improve the detection ability for the minority class (such as dead chickens). Description of the Drawings

[0049] Figure 1 It is the framework structure diagram of the dead chicken detection model of the present invention;

[0050] Figure 2 It is the structure diagram of the image fusion module of the present invention;

[0051] Figure 3 It is the structure diagram of the GELAN network in the YOLOv9 algorithm of the present invention;

[0052] Figure 4 It is the structure diagram of the Rep-DCNv3 network of the present invention;

[0053] Figure 5 It is the structure diagram of the improved YOLOv9 object detection module of the present invention;

[0054] Figure 6 It is the comparison result diagram of three evaluation indexes in the embodiment;

[0055] Figure 7 It is the heat map result diagram of the improved YOLOv9 model in the embodiment;

[0056] Figure 8 It is the heat map result diagram of the original YOLOv9 model in the embodiment. Detailed Embodiment

[0057] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the application will be further elaborated in detail below with reference to the drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments of other researchers in the field on this embodiment belong to the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0058] Aiming at the dead chicken detection task of caged chickens, the present invention proposes an innovative improved YOLOv9 model, which uses multi-source information fusion technology to solve the problem of uneven illumination in the farm.

[0059] First, the PIAFusion (Progressive Image Fusion based on Illumination Awareness) method is adopted to fuse visible light and infrared light, enhancing the saliency of the dead chicken target features to overcome the problem of uneven illumination caused by structures such as multi-layer chicken coops, baffles, and water troughs. This fusion enhances the overall brightness and contrast of the image by comprehensively utilizing various light source information, thereby improving the robustness and detection accuracy of the model under complex illumination conditions.

[0060] Secondly, based on the YOLOv9 object detection algorithm, a dead chicken detection model is constructed. To further enhance the model performance, an EMA attention mechanism is introduced into the neck of YOLOv9. By performing exponential smoothing on the model parameters, it can effectively distinguish low-dark targets such as chicken coop baffles and waterers from dead chicken targets and achieve the fusion of targets of different sizes. This technology not only optimizes the stability of the model but also improves the detection ability for various targets.

[0061] Then, the Rep-DCNv3 module is introduced into the feature extraction network of the dead chicken detection model to enhance the backbone feature extraction network, ensuring that the features of dead chickens even when blocked by live chickens can be effectively extracted, solving the problem that traditional methods are difficult to accurately detect dead chickens when they are blocked, and significantly improving the accuracy of dead chicken detection.

[0062] Finally, in the selection of the loss function, the MPDIoU loss function is adopted as the core loss function of the algorithm to replace the loss function of YOLOv9. The MPDIoU loss function provides more fine-grained error feedback by more accurately evaluating the overlap degree between targets, thereby improving the detection accuracy of the model.

[0063] The overall framework of the dead chicken detection model of the present invention is as Figure 1 shown, demonstrating innovation and superiority in the task of detecting dead chickens in caged chickens. By introducing the Rep-DCNv3 module and the EMA attention mechanism, it shows significant advantages in dealing with occluded targets and fusing targets of different sizes, bringing a new solution to the task of detecting dead chickens in caged chickens. It not only improves the accuracy of dead chicken detection but also provides a new technical approach for the task of detecting dead chickens in caged chickens. The specific method is as follows:

[0064] (1) The present invention adopts the PIAFusion (Progressive Image Fusion based on Illumination Awareness) method for image processing, aiming to improve the quality and usability of images in an environment with uneven illumination conditions.

[0065] In the task of detecting dead chickens in caged chickens, the PIAFusion technology realizes the efficient fusion of visible light and infrared light images by comprehensively considering the scene illumination characteristics and image content, significantly enhancing the robustness of the detection system.

[0066] The core advantage of the PIAFusion technology lies in its progressive fusion strategy. First, it comprehensively evaluates the illumination conditions of the input image, identifies areas with uneven illumination, and performs targeted optimization processing. PIAFusion adopts a multi-scale analysis method, decomposes the image into different scale levels, and implements customized fusion strategies for each layer to adapt to the complex illumination environment caused by multi-layer chicken coops, baffles, water troughs, etc.

[0067] The image fusion module of the present invention is as Figure 2 shown. The PIAFusion backbone network is constructed on an end-to-end CNN-based framework and mainly consists of two key components: a feature extractor and an image reconstructor.

[0068] The feature extractor aims to deeply mine the complementary and common features of infrared images and visible light images. Its structure includes five convolutional layers. To mitigate the differences between different modality images, a 1×1 convolutional layer is used at the initial stage to process the infrared image and the visible light image respectively. Secondly, four convolutional layers with shared weights are used to extract the deep features of the infrared and visible light images respectively.

[0069] In particular, the present invention introduces the CMDAF module after the second layer, the third layer, and the fourth layer respectively to achieve the exchange of modality complementary features.

[0070] The introduction of the CMDAF module enables the network to gradually integrate complementary information during the feature extraction stage, thereby comprehensively extracting the features of infrared and visible light images.

[0071] Subsequently, the deep features extracted from the two images are pooled and concatenated as the input of the image reconstructor. The image reconstructor consists of five convolutional layers, and its task is to fully integrate the common and complementary information to generate the final fused image. To avoid information loss during the image fusion process, no downsampling is introduced in the network, and the size of the fused image is the same as that of the source images.

[0072] (2) The present invention constructs a dead chicken detection model based on the YOLOv9 object detection algorithm, and mainly performs dead chicken detection based on the color information difference between dead chickens and live chickens in the fused image.

[0073] First, the YOLOv9 object detection algorithm, on the basis of YOLOv8 and YOLOv5, conducts in-depth innovation and optimization for the information bottleneck problem in deep learning, effectively solves the loss of inter-layer data transmission, and alleviates the problems of slow convergence speed and unstable results. On the basis of YOLOv8, YOLOv9 introduces the Generalized Efficient Layer Aggregation Network (GELAN) and the Programmable Gradient Information (PGI) control framework.

[0074] The GELAN inherits and optimizes the advantages of CSPNet and ELAN. The network structure is as Figure 3 shown. By integrating diverse computational modules, it significantly enhances the flexibility and accuracy of the model. The design of GELAN optimizes the inference time and significantly improves the overall detection accuracy of the model.

[0075] The described PGI control framework, as an auxiliary structure, optimizes the information flow path through the collaborative action of three sub-units, alleviates the impact of information bottlenecks, and thus improves the inference efficiency of the algorithm.

[0076] The YOLOv9 object detection algorithm not only solves the convergence problem in deep learning but also significantly improves the running efficiency and accuracy of the model. The decoupled head architecture and anchor-free strategy of YOLOv9 simplify the object detection process, reduce unnecessary computations, and accelerate the object detection speed. At the same time, task-specific loss functions designed for object classification and bounding box regression estimation, such as Distribution Focal Loss (DFL) and Complete IoU Loss (CIoU), are adopted to further improve the detection accuracy.

[0077] Then, the present invention makes important improvements in the Neck part of the YOLOv9 model through EMA (Efficient Multi-Scale Attention Mechanism), significantly enhancing the model's perception and recognition capabilities for multi-scale targets. The EMA attention mechanism effectively improves the model's ability to capture different-scale information in images by combining multi-scale feature extraction and attention weight allocation, while maintaining optimized computational efficiency and resource consumption.

[0078] In the actual fused images of the present invention, low-contrast targets such as chicken coop baffles and drinkers are likely to cause misidentification of dead chicken targets. At the same time, the heavy occlusion of dead chickens by live chickens and the mutual overlap between multiple dead chickens lead to a decrease in the detection accuracy of the model.

[0079] The introduction of the EMA attention mechanism not only enriches the feature extraction of the model at different scales but also provides the model with a more comprehensive and detailed understanding of image content through the fusion of multi-scale information. This multi-dimensional feature integration strategy enables the model to more sensitively capture local changes when dealing with complex scenarios such as live chickens occluding dead chickens, thus significantly improving the accuracy of target localization.

[0080] The local and global attention mechanisms in the EMA module work together to finely perceive local features and integrate global context understanding, achieving precise focusing on the target area. The introduction of this dual attention mechanism greatly reduces the problem of missed detection of dead chicken targets caused by live chicken occlusion and also reduces the interference of live chickens and the chicken coop background on the model's decision-making, further improving the detection accuracy and robustness.

[0081] Furthermore, in the testing stage, the improved YOLOv9 model, with the enhanced generalization ability of the EMA module, can effectively handle various scene and target pose changes. This improvement in generalization performance is not only reflected in the accurate recognition of known samples but also in the adaptability to unknown or rare samples, ensuring the stability and reliability of the model in practical applications.

[0082] In addition, the EMA module optimizes the information processing flow, effectively reducing information loss and confusion during feature extraction. This optimization strategy not only avoids misjudgments caused by feature confusion but also further enhances the model's adaptability to complex environments and the accuracy of detection results by strengthening the model's extraction and utilization of key information.

[0083] For the EMA attention mechanism of the present invention, first, convolution operations are performed through a series of convolutional kernels with different sizes to capture multi-scale features. These convolutional kernels are specifically designed to adapt to semantic information at different levels. Smaller-sized convolutional kernels are used to extract detailed features such as edges and textures, while larger-sized convolutional kernels are used to identify large-scale patterns and shapes in the image.

[0084] Subsequently, a series of carefully designed attention modules are introduced. These modules, through a lightweight network architecture, learn and evaluate the contribution degrees of features at each scale to achieve dynamic weighting of features. This weighting strategy not only enhances the model's sensitivity to significant features but also effectively reduces the processing requirements for redundant information, further improving the operation efficiency.

[0085] Next, to further reduce the computational complexity and memory consumption, a subsampling technique is integrated. This is a technique that reduces the feature dimension through downsampling, ensuring the integrity and accuracy of feature representation while reducing the computational amount.

[0086] In the feature fusion stage, the EMA mechanism integrates the weighted multi-scale features to form the final feature representation. This fusion not only retains the advantages of multi-scale features but also strengthens the model's ability to identify key information through the weighting mechanism.

[0087] (3) In view of the limitations of traditional convolution in dealing with heavily occluded targets, the present invention introduces an improved scheme of Rep-DCNv3. When detecting dead chickens in an automated breeding environment, the problem of heavy occlusion has always been a key bottleneck in improving the algorithm performance. When traditional convolutional neural networks handle such problems, they are often limited by their convolutional kernels with fixed shapes and are difficult to effectively capture the features of occluded targets. To solve this problem, the present invention proposes an improved scheme of Rep-DCNv3, aiming to improve the performance of the YOLOv9 algorithm in dealing with heavily occluded targets.

[0088] In the YOLOv9 algorithm, on the one hand, DCN (Deformable Convolutional Networks) brings revolutionary improvements to traditional convolutional neural networks based on the design of deformable convolutions. The flexibility of deformable convolutions enables the network to learn the shape and position of input features more adaptively, thereby better capturing the relationships between features. On the other hand, the RepNCSPELAN module plays an important role in feature processing and fusion. However, in the face of the occlusion of dead chickens by live chickens, traditional convolutional kernels have great limitations, which are mainly manifested in inaccurate feature extraction of occluded targets, resulting in a decrease in detection accuracy.

[0089] To address this problem, the present invention replaces the ordinary convolutional kernels in the RepNCSPELAN module with DCNv3. As Figure 4 shown, by integrating the Rep-DCNv3 module into the Backbone feature extraction network, advanced components such as LN, FFN, and GELU are adopted, as well as convolutional layers and downsampling layers for obtaining hierarchical feature maps, improving the network performance and providing the best solution for the processing of visual tasks. Compared with traditional central neural networks, the basic blocks are closer to ViTs and are equipped with more advanced components. This design demonstrates excellent performance and flexibility in various visual tasks, enabling the dead chicken detection model of the present invention to more effectively extract the features of occluded dead chickens to achieve more accurate feature capture of occluded targets.

[0090] Specifically, the adaptive convolutional kernel of DCNv3 is one of its core advantages. By dynamically adjusting the convolutional kernel to adapt to the shape of the target, it can more accurately capture the target features in complex environments. This adaptability significantly improves the model's ability to recognize occluded targets, especially in the case where dead chickens are occluded by live chickens.

[0091] In addition, the offset prediction network of DCNv3 further enhances the model's adaptability to changes in target position and shape, thereby improving the detection accuracy.

[0092] Furthermore, in order to improve the processing speed while maintaining high detection accuracy, the DCNv3 of the present invention introduces an innovative optimization strategy. By optimizing bilinear interpolation into linear interpolation on the axis, the network structure of the YOLOv9 algorithm is optimized, achieving a significant improvement in computational efficiency. It not only reduces the computational amount but also avoids information distortion caused by interpolation, ensuring the accuracy of the detection results and significantly enhancing the model's performance in processing heavily occluded targets.

[0093] The present invention introduces the Rep-DCNv3 module into the feature extraction network, which involves an offset prediction network and deformable convolution operations. By first predicting the offsets of the input positions and then performing adaptive convolution operations based on these offsets, more refined processing of the input features is achieved, with better performance and robustness. The principles are as follows:

[0094]

[0095] Among them, H represents the total number of aggregation groups; k h represents the position-independent projection weight of the h-th group; q hm represents the modulation scalar of the m-th sampling point in the h-th group; x h represents the input feature map of the slice; Δr hm represents the offset corresponding to the grid sampling position r m in the m-th group; y(r 0 ) represents the output.

[0096] (4) The MPDIoU loss function used in the present invention is for bounding box regression. The core idea is to optimize the accuracy and efficiency of bounding box regression by considering the distances between the top-left and bottom-right points of the predicted bounding box and the ground truth bounding box. This method comprehensively considers factors such as the position, size, and shape of the bounding box, thus more comprehensively evaluating the results of object detection.

[0097] The MPDIoU loss function independently calculates the average IoU of each category, which helps to reduce the overfitting of the model to the training data and improve the generalization ability of the model. By introducing the MPDIoU loss function, the training process becomes more stable, the model can learn features more comprehensively, and thus achieve better detection results in complex environments.

[0098] In a high-density breeding environment, dead chickens often block or partially overlap each other, which poses challenges to traditional object detection algorithms and makes it difficult to accurately distinguish and locate each dead chicken. Traditional regression boxes often cannot accurately cover the entire body of each dead chicken. Especially in the case of overlapping objects, the detection algorithm is prone to misjudgment or omission. This overlapping phenomenon not only affects the detection accuracy but also reduces the interpretability of the results, making subsequent data processing and analysis more complex.

[0099] To solve this problem, the present invention adopts the MPDIoU technique as the loss function of the YOLOv9 algorithm. MPDIoU is an improved IoU measurement method that more precisely evaluates the matching degree of the target boxes by considering the partial overlap and distance information between the targets.

[0100] The advantage of the MPDIoU technique lies in its sensitivity to overlapping objects and its ability to balance when dealing with objects of different categories. It can not only reduce the overfitting of the model to the majority class but also improve the detection ability for minority classes (such as dead chickens). This balancing mechanism makes the model more robust and reliable when handling multi-object detection tasks in complex environments.

[0101] In the YOLOv9 algorithm, replacing the traditional IoU loss function with the MPDIoU loss function significantly improves the model's detection performance for overlapping objects. By introducing the MPDIoU loss, the model can learn a more accurate object bounding box localization strategy during training, thus generating more precise detection results. This improvement not only enhances the detection accuracy but also strengthens the model's adaptability to complex scenarios.

[0102] In summary, the improved YOLOv9 object detection module structure of the present invention, as Figure 5 shown, in the feature extraction network, the Rep-DCNv3 module is introduced to ensure that the features of occluded dead chickens can also be effectively extracted. In addition, the EMA mechanism is introduced after the RepNCSPELAN4 module in the Neck part to better distinguish objects from the background. In the selection of the loss function, aiming at the limitations of the traditional IoU loss function in dealing with overlapping or occluded objects, the MPDIoU loss function is selected. The MPDIoU loss function can more accurately evaluate the overlapping degree between objects, balance the detection performance between different categories, reduce overfitting to the majority class, and improve the detection ability for minority classes (such as dead chickens).

[0103] In an embodiment of the present invention, in the broiler promotion demonstration base in Lujicun, Daxing Town, Suyu District, Suqian City, Jiangsu Province, with the white - feather cage - raised chicken industry as the test object, the dead chicken detection experiment is carried out according to the above - mentioned method. The chicken cages for cage - raised chickens have a total of four layers from top to bottom. When conducting the experiment, the second - layer chicken cage is selected to collect data, and the experimental data is collected in March 2024.

[0104] The registered and fused images are screened to eliminate images of poor quality, and finally 1348 effective images are obtained. The labeled images are randomly divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. According to the uncertainty of the position of the dead chicken target in the image, the Mosaic data augmentation method is used to perform operations such as scaling, flipping, and splicing on the images, which enriches the training data and is conducive to improving the robustness and generalization ability of the model.

[0105] In this embodiment, the hardware devices used in the experiment and their configurations are as follows: Intel(R) Xeon(R) Bronze 3106 CPU, 62GB of running memory, and a maximum frequency of 1.7GHz; RTX 2080Ti graphics cards (three pieces), 36GB of video memory. The experimental operating system is Ubuntu 20.04, the model framework is built based on PyTorch 1.8, and the Python version is 3.9.

[0106] Furthermore, three metrics, namely Precision (P), Recall (R), and mean average precision at 0.5 (mAP@0.5), are used to evaluate the model performance. In addition, FPS is used to measure the model detection speed. FPS represents how many frames of images the algorithm can process per unit of time (seconds). The higher the frame rate, the better the real-time performance of the algorithm.

[0107] In this embodiment, the detection speed is evaluated by measuring the time required to process a single image, the ability of the model to accurately identify dead chickens is evaluated by Precision (P), the coverage ability of the model to identify dead chickens is evaluated by Recall (R), and the mean average precision (mAP) comprehensively processes P and R. By considering the balance between detection precision and recall, it becomes a comprehensive indicator of model performance. mAP@0.5 refers to the average precision of each class at the intersection over union (IoU) threshold of 0.5, and the calculation formula is as follows:

[0108]

[0109] Among them, TP is the number of true samples (true positives); FP is the number of false positive samples; FN is the number of false negative samples; and C is the number of classes.

[0110] To comprehensively evaluate the performance of the algorithm proposed in the present invention, this embodiment selects a series of representative benchmark models for comparative experiments. These models include the original YOLOv9 model, YOLOv8 and YOLOv7, the classic object detection algorithms Faster-RCNN and SSD, and the EfficientDet model, an efficient convolutional neural network architecture achieved through compound scaling methods and network structure optimization. In addition, the large model RT-DETR model, which performs excellently in terms of the number of parameters, and the DAMO YOLO model, which can handle dynamically changing scenarios and provide real-time object detection and tracking capabilities, are also considered. In particular, Zhao Yiming et al. constructed the improved YOLOv5-SE model by introducing the SE attention module and applying CIoU_Loss and DIoU_NMS to the original YOLOv5 model.

[0111] All these algorithm models are comprehensively compared with the algorithm proposed in the present invention, and the experimental results are shown in Table 1 below:

[0112] Table 1: COMPARATIVE EXPERIMENTS OF DIFFERENT METHODS

[0113]

[0114] It can be seen that the dead chicken detection model of the present invention has achieved the highest performance in terms of both the mAP@0.5 value and the P value, and is only slightly lower than the YOLOv5-SE model in terms of the R value. Compared with the initial YOLOv9 model, the proposed algorithm has improved by 3.4%, 6% and 4.8% in terms of mAP@0.5, R and P values respectively.

[0115] Through the analysis of Table 1, the FPS of EfficientDet, SSD, Faster-RCNN, and RT-DETR is significantly lower than that of other algorithms, and they also fail to reach the level of other evaluation models in terms of mAP@0.5, indicating that these models are difficult to meet the requirements of real-time detection tasks for real-time detection speed and accuracy. Compared with the YOLOv5-SE model, the proposed algorithm has higher mAP@0.5 and P values than the YOLOv5-SE model. Although the R value is slightly lower by 0.3%, the FPS is slightly higher. Considering the detection accuracy and speed comprehensively, the proposed model performs more excellently. In addition, compared with YOLOv9, YOLOv8, and YOLOv7, the algorithm introduced in this study provides an improved trade-off between detection speed and detection accuracy. It should be noted that although the FPS value of YOLOv8 is higher than that of the proposed algorithm, the proposed method has increased the mAP value by 0.5% and the P value by 4.3%, thus confirming the superior performance of the proposed algorithm in object detection.

[0116] Furthermore, in order to compare the influence of different structures on the detection performance, ablation experiments were carried out on the collected fusion dataset. In the experiments, YOLOv9 was used as the basic model. Rep-DCNv3 was only added to the backbone of YOLOv9 (YOLOv9+Rep-DCNv3), EMA was only added to the NECK part of YOLOv9, MPDIoU technology was only used as the loss function, both Rep-DCNv3 and the EMA module were used simultaneously, and both Rep-DCNv3 and the MPDIoU loss function were used simultaneously. The experimental results are shown in Table 2 below.

[0117] Table 2: ABLATION RESULTS

[0118]

[0119] From the analysis of Table 2, by comparing the experimental results of Group ⑦ (the basic YOLOv9 model) and Group ⑥ (YOLOv9+Rep-DCNv3), it can be observed that simply adding the Rep-DCNv3 module significantly improves the mAP@0.5, R, and P values of the model, increasing from 96.5%, 93.1%, 94.4% to 97.4%, 96.3%, 97.8% respectively. This comparison shows that the Rep-DCNv3 module has a positive effect on improving detection accuracy. Further, by comparing the results of Group ⑥ and Group ① (Ours), it is noted that when the Rep-DCNv3 module is combined with the EMA module and the MPDIoU loss function, the performance is further improved.

[0120] Specifically, the mAP@0.5, R, and P values increase from 97.4%, 96.3%, 97.8% to 98.9%, 97.4%, 99.2% respectively. This indicates that the combined use of the Rep-DCNv3 module, the EMA module, and the MPDIoU loss function can achieve complementary effects and further improve the detection performance. In addition, by comparing the results of Group ③ (YOLOv9+RepDCNv3+MPDIoU) and Group ④ (YOLOv9+MPDIoU), it is found that although the use of only the MPDIoU loss function has a certain improvement effect on performance, it is not as significant as the Rep-DCNv3 module. The addition of the Rep-DCNv3 module increases the mAP@0.5 and P values from 96.8%, 95.4% to 97.7%, 98.1% respectively, and the R value also increases.

[0121] Finally, the comparison results of Group ② (YOLOv9+Rep-DCNv3+EMA) and Group ⑤ (YOLOv9+EMA) show that the addition of the Rep-DCNv3 module also improves the performance in the presence of the EMA module, with the mAP@0.5, R, and P values increasing from 97.1%, 94.7%, 96.9% to 98.2%, 95.7%, 98.3% respectively. The Rep-DCNv3 module shows its significant effect in improving detection performance in various combinations. Whether used alone or in combination with other improved components, the Rep-DCNv3 module can effectively improve the accuracy and robustness of object detection.

[0122] Furthermore, in this embodiment, through a strict evaluation framework, the performance differences between the EMA attention mechanism and other advanced attention mechanisms in the object detection algorithm are quantified. To this end, YOLOv9 is selected as the benchmark model, and the attention mechanisms selected include Channel Attention (CA), Squeeze-and-Excitation (SE), Convolutional Block Attention Module (CBAM), Efficient Channel Attention (ECA), and Normalization-based Attention Module (NAM). In the experimental design, all attention mechanisms are integrated into the position of the EMA mechanism to ensure the consistency and fairness of the evaluation. The relevant experimental results are detailed in Table III below:

[0123] Table III: Experimental results of different attention mechanisms

[0124]

[0125] Among them, the CA mechanism aims to optimize the channel-level representation of features by emphasizing the importance and interdependence between channels; the SE mechanism further enhances the channel-level expression of features by recalibrating channel weights; the CBAM mechanism combines spatial and channel attention to strengthen the cross-dimensional information of local features, thereby improving the model's sensitivity to local features; the ECA mechanism uses global average pooling to capture the global spatial information between channels, providing strong support for the channel-level reconstruction of features; the NAM mechanism dynamically adjusts the allocation of attention by adjusting the degree of normalization of features.

[0126] The experimental results show that although the CA module has a slight improvement in the P value, its performance on mAP@0.5 is 0.3% lower than that of the original YOLOv9 model and 0.9% lower than that of the EMA module. The CBAM and SE modules are better than the EMA module in terms of the R value, but they fail to exceed the EMA in terms of mAP and P values, further highlighting the advantages of the EMA mechanism in comprehensive performance. The ECA and NAM modules do not reach the performance level of the EMA module in multiple evaluation indicators.

[0127] Based on the comprehensive experimental data, the integration of YOLOv9 and EMA shows excellent performance under the specific settings of the method of the present invention, achieving an mAP@0.5 value of 97.1% and a P value of 96.9%.

[0128] Furthermore, in the performance study of verifying the MPDIoU loss function, this embodiment selected two variants of the YOLOv9 model: the basic YOLOv9 model and the YOLOv9 model integrated with the proven effective EMA module. These two models respectively applied the MPDIoU loss function and were trained on the fused image dataset. The results are as Figure 6 shown.

[0129] Figure 6 shows the comparison results of three key evaluation metrics of the two models and their variants on the validation set. Figure 6 (a) in Figure 6 shows the mAP@0.5 metric result, Figure 6 (b) in Figure 6 shows the P metric result,

[0130] Finally, to comprehensively evaluate the performance of this method in object detection, especially its ability to detect dead chicken targets in complex backgrounds, this embodiment experiment used a visualization display in the form of heatmaps, as Figure 7 and Figure 8 shown.

[0131] Figure 7 and Figure 8 respectively show the heatmap results of using the method of the present invention and the original YOLOv9 model, intuitively reflecting the detection effects of the two models under different occlusion conditions.

[0132] In Figure 7 and Figure 8 (a) of, the heatmap results of multiple dead chickens in the improved YOLOv9 model of the present invention and the original YOLOv9 model under the condition of not serious occlusion of dead chickens are shown. It can be seen that the heatmap of the method of the present invention, that is, Figure 7 (a) inFigure 8 In (a), there is excessive attention to the background, resulting in inaccurate detection of dead chicken targets. This indicates that although the original YOLOv9 model is still significantly interfered by background factors when occlusion is not severe.

[0133] In Figure 7 and Figure 8 Figure (b), the heatmap results in the case of severe occlusion of multiple dead chickens are shown. In this more challenging scenario, the heatmap of the method of the present invention, namely Figure 7 Figure (b) in the figure still performs excellently, accurately identifying and locating each dead chicken, and the confidence level exceeds 0.9. This result is attributed to the introduction of the Rep-DCNv3 module, which accurately captures target features through an adaptive convolution kernel, effectively solves the occlusion problem, and enhances the recognition ability of occluded targets. In contrast, the heatmap of the original YOLOv9 model in the same scenario, namely Figure 8 Figure (b) in the figure fails to reach such a high confidence level, showing deficiencies in dealing with scenarios with severe occlusion.

[0134] In addition, the MPDIoU loss function is adopted in the case of multiple dead chickens by this method, further improving the accuracy of the regression box. The design of the MPDIoU loss function takes into account the partial overlap and distance information between targets, more accurately evaluates the overlap degree between targets, and thus provides a more effective training signal for the model. This is reflected in the heatmap of this method. Even in the case of multiple dead chickens occluding each other, the regression box can still accurately cover the entire body of each dead chicken. In summary, through the synergistic effect of the EMA module and the Rep-DCNv3 module, this method significantly improves the accuracy and robustness of dead chicken detection. At the same time, the application of the MPDIoU loss function further improves the accuracy of the regression box. These improvements enable this method to show significant performance advantages compared with the original YOLOv9 model in complex backgrounds and scenarios with severe occlusion.

[0135] In summary, the present invention develops an efficient dead chicken detection method based on image fusion and improved YOLOv9 model for the automation requirement of dead chicken detection in large-scale breeding environments. By using the PIAFusion technology to fuse thermal infrared and visible light images, this method effectively improves the saliency of dead chicken target features and solves the influence of uneven illumination on the detection results. In addition, the introduction of the EMA attention mechanism significantly improves the ability of the model to distinguish dead chicken targets in low-light environments, enhancing the accuracy and generalization of the model. Further, the integration of the Rep-DCNv3 module strengthens the backbone feature extraction network, enabling the model to more accurately extract the features of dead chickens in the case of partial occlusion.

[0136] Most importantly, by adopting the MPDIoU loss function to replace the original loss function of YOLOv9, the present invention is more accurate in evaluating the degree of target overlap, further improving the detection performance. The accuracy of the experimental results reaches 99.2%, and the mAP@0.5 value is as high as 98.9%, significantly superior to other current advanced algorithms, verifying the superiority of the method of the present invention in the dead chicken detection task. Therefore, the proposed dead chicken detection method not only ensures high detection speed but also greatly meets the demand for real-time detection of dead chickens in actual production. These results indicate that the proposed method has broad application potential in large-scale breeding environments and provides a reliable solution for automated dead chicken inspection.

[0137] The dead chicken detection model proposed by the present invention has demonstrated excellent detection effects in the multi-layer cage-raised broiler environment, providing a solid theoretical basis for automated dead chicken inspection. This achievement not only promotes the practice of welfare-based breeding of caged chickens but also provides technical support for large-scale breeding management. To address the more complex situation of occlusion between chickens, subsequent work will explore more advanced image registration and fusion technologies. Considering that the growth cycle of broilers is approximately 21 days, future research will expand to chicken datasets throughout the entire life cycle to obtain a more comprehensive performance evaluation. At the same time, to improve the overall performance of the monitoring system, future research will also consider synchronous detection of multi-layer chicken cages to achieve more efficient automated inspection. This research direction will provide new perspectives and solutions for improving the method of automatic dead chicken detection.

[0138] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A dead chicken detection method based on image fusion and improved YOLOv9, characterized in that: The steps include: S1. Collect chicken cage image data, use a progressive image fusion method based on lighting perception to fuse visible light with infrared light, improve the significance of dead chicken target features, and eliminate the impact of uneven lighting caused by multi-layer chicken cages, baffles, and water tank structures; S2. Build a dead chicken detection model based on the YOLOv9 target detection algorithm, introduce the EMA attention mechanism into the YOLOv9 neck, perform exponential smoothing on the model parameters, distinguish between low-light environments and dead chicken targets, and fuse targets of different sizes; S3, introduce the Rep-DCNv3 module into the dead chicken detection model, enhance the backbone feature extraction network, and extract the features of dead chickens under partial occlusion; The Rep-DCNv3 module replaces the convolution kernel in the RepNCSPELAN module of the YOLOv9 algorithm by DCNv3 in the Backbone feature extraction network to capture the features of the dead chicken detection model when processing heavily occluded targets; The DCNv3 dynamically adjusts to the target shape through an adaptive convolution kernel. When a dead chicken is occluded by a live chicken, the target features are accurately captured, thereby improving the ability to recognize the occluded target. The DCNv3 offset prediction network is used to enhance the adaptability to changes in target position and shape, thereby improving the accuracy of detection. The bilinear interpolation is optimized to linear interpolation on the axis, thereby reducing the amount of calculation and avoiding information distortion caused by interpolation. The Rep-DCNv3 module has a network structure including LN, FFN and GELU components, as well as convolutional layers and downsampling layers, which are used to obtain hierarchical feature maps. The method is as follows: S3.1, the input multi-scale feature map is subjected to preliminary feature extraction through the convolution layer, and then dynamically adjusted through the adaptive convolution kernel in the Rep-DCNv3 module to adapt to the target shape and capture the target features; S3.2, the feature map is sequentially processed by FFN and GELU to enhance the nonlinear expression ability, and is normalized by LN to stabilize the training process; S3.3, gradually reduce the spatial dimension of the feature map through the downsampling layer, and increase the depth of the feature map to perform visual task processing, and obtain a high-dimensional feature representation for target detection, which is used for accurate identification and positioning of dead chicken targets; S4. MPDIoU is used to replace the loss function of YOLOv9. By evaluating the degree of overlap between targets, it provides more fine-grained error feedback and improves the detection accuracy of the dead chicken detection model.

2. The dead chicken detection method based on image fusion and improved YOLOv9 according to claim 1 is characterized in that: In step S1, the progressive image fusion method based on illumination perception has a backbone network constructed in an end-to-end CNN framework, including a feature extractor and an image reconstructor; The feature extractor is used to mine the complementary and common features of infrared images and visible light images. The structure includes five convolutional layers. The first layer uses a 1×1 convolutional layer to process the infrared image and the visible light image respectively to reduce the difference between images of different modalities. The second to fifth layers use four convolutional layers with shared weights to extract the deep features of the infrared and visible light images respectively. The deep features extracted from the infrared image and the visible light image are aggregated and connected and input into the image reconstructor. The image reconstructor consists of five convolutional layers without introducing downsampling. The size of the fused image is consistent with the source image. It is used to integrate common and complementary information to generate a fused image to adapt to the complex lighting environment caused by multi-layer chicken cages, baffles, and sinks.

3. The dead chicken detection method based on image fusion and improved YOLOv9 according to claim 2 is characterized in that: A CMDAF module is introduced after the second, third and fourth layers of the feature extractor to comprehensively extract the features of infrared and visible light images and realize the exchange of modal complementary features.

4. The dead chicken detection method based on image fusion and improved YOLOv9 according to claim 1, characterized in that: In step S2, the EMA attention mechanism is used to enrich the feature extraction and attention weight allocation of the dead chicken detection model at different scales. The specific method is as follows: S2.

1. When the live chicken occludes the dead chicken, convolution operations are performed by using convolution kernels of different sizes to adapt to semantic information at different levels and capture multi-scale features. Smaller convolution kernels are used to extract edge and texture detail features, while larger convolution kernels are used to identify large-scale patterns and shapes in the image. S2.2, introduce the attention module, learn and evaluate the contribution of each scale feature through a lightweight network architecture, and realize dynamic weighting of features; S2.3, integrated subsampling method, reduces feature dimensions by downsampling, reduces the amount of computation while ensuring the integrity and accuracy of feature representation; S2.

4. In the feature fusion stage, the weighted multi-scale features are integrated, the multi-scale features are retained, and the model's recognition ability of key information is strengthened through the weighting mechanism to form the final feature representation, thereby reducing the missed detection of dead chicken targets caused by live chickens, reducing the interference of live chickens and chicken cage backgrounds on the dead chicken detection model, and improving the accuracy and robustness of detection; S2.

5. Reduce information loss and confusion in the feature extraction process by optimizing information processing, and enhance the model to extract and utilize key information to improve adaptability to complex environments and the accuracy of detection results.

5. The dead chicken detection method based on image fusion and improved YOLOv9 according to claim 1 is characterized in that: The Rep-DCNv3 module involves an offset prediction network and a deformable convolution operation. It first predicts the offset of the input position and then performs an adaptive convolution operation based on the offset. The principles are as follows: Where H represents the total number of aggregation groups; k h represents the position-independent projection weight of the hth group; q hm represents the modulation scalar of the mth sampling point in the hth group; x h represents the input feature map of the slice; Δr hm Represents the grid sampling position r in the mth group m The corresponding offset; y(r0) represents the output.

6. The dead chicken detection method based on image fusion and improved YOLOv9 according to claim 1 is characterized in that: In step S4, the MPDIoU is an improved IoU measurement method, which is used to accurately evaluate the degree of overlap between targets in a high-density farming environment, balance the detection performance between different categories, and reduce overfitting of the majority class, including the following sub-steps; S4.

1. By considering the partial overlap and distance information between the target dead chickens, the matching degree MPDIoU of the target box is evaluated and calculated as follows: Where w and h are the width and height of the input image; A and B are two arbitrary convex shapes, d1 is the Euclidean distance from the upper left corner of shape B to the upper left corner of shape A, and d2 is the Euclidean distance from the lower right corner of shape B to the lower right corner of shape A; are the upper left corner coordinates and lower right corner coordinates of A respectively, They are the upper left corner coordinates and lower right corner coordinates of B respectively; S4.

2. Based on MPDIoU, the loss function is defined as follows: L MPDIoU =1-MPDIoU S4.

3. During the training of the dead chicken detection model, by introducing the MPDIoU loss, a more accurate target box positioning strategy is learned to generate more precise detection results, thereby enhancing the adaptability of the dead chicken detection model to complex scenes.

7. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the dead chicken detection method based on image fusion and improved YOLOv9 as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the steps in the dead chicken detection method based on image fusion and improved YOLOv9 described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method for identifying dead chicken in infrared image cage based on improved YOLOv5 model

    CN115527234A

  • Method for improving YOLOv8 network and application of method in strip steel surface defect detection

    CN118365599A