Cow ear tag detection method and device based on convolutional neural network
By optimizing the convolutional neural network of the YOLOv11 algorithm, a cattle ear tag detection model was constructed, which solved the problems of low efficiency, high cost and poor usability of existing cattle identification technologies. It achieved efficient and low-cost non-contact detection, adapts to complex scenarios, and promotes intelligent management of animal husbandry.
Patent Information
- Application Number
- CN202510939228.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing cattle identification technologies have shortcomings in terms of efficiency, accuracy, cost, ease of use, and automation integration. Traditional physical ear tags are inefficient and prone to misreading, while electronic ear tags are costly, easily damaged, and difficult to integrate with digital systems.
A convolutional neural network was used to optimize the YOLOv11 algorithm. A bull's ear mark detection model was constructed by using adaptive downsampling, dynamic sample attention scale sequence fusion, and a dual-kernel convolutional cross-stage partial dark image recovery module. Non-contact detection was performed using a general-purpose camera.
It improves the performance of cattle ear tag detection, reduces costs, minimizes safety risks, adapts to complex scenarios, and achieves efficient batch identification, thus contributing to refined breeding and disease prevention and control.
Smart Images

Figure CN120894801A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of computer vision, and particularly relate to a method and device for detecting cow ear tags based on a convolutional neural network. BACKGROUND
[0002] In modern and large-scale livestock management, accurate and efficient identity recognition of individual livestock such as cattle is the core basis for realizing fine feeding, disease prevention and control, pedigree tracing, production performance measurement, and product quality safety tracing. The current mainstream cattle identity recognition technology has obvious limitations: traditional physical ear tag recognition involves wearing ear tags with unique alphanumeric codes on cattle, relying on manual close-range observation and recording. Although it is low-cost and easy to implement, it is extremely inefficient in large-scale cattle herds, and is easily affected by light, angle, dirt, and personnel state, leading to misreading and missing reading. Data need to be manually entered, making it difficult to integrate with digital systems, and there is a safety risk of personnel contacting large livestock. Electronic ear tags (RFID) recognition involves non-contact reading through a radio frequency chip, which has high automation and speed. However, it requires expensive specialized readers, has high deployment and maintenance costs, has a short effective reading distance and strict location requirements, and has much higher ear tag and equipment costs than traditional ear tags. The internal ID read needs to be additionally associated with a visual number, which lacks intuitiveness, and the ear tag is easily damaged and falls off, leading to recognition failure. In summary, the existing technology has deficiencies in efficiency, accuracy, cost, ease of use, and automation integration. Therefore, there is an urgent need to explore a new type of cattle identity recognition method that is non-contact, high-efficiency, low-cost, high-precision, and easy to integrate with existing management processes. SUMMARY
[0003] To solve the above problems of the prior art, embodiments of the present application provide a method and device for detecting cow ear tags based on a convolutional neural network.
[0004] In a first aspect, embodiments of the present application provide a method for detecting cow ear tags based on a convolutional neural network, comprising: obtaining multiple cow ear tag images to form a model construction dataset; replacing the third, fifth, and seventh layers of ordinary convolution in the YOLOv11 backbone network with adaptive down-sampling Adown; introducing a dynamic sample attention scale sequence fusion at the 25th and 26th layers of the YOLOv11 neck network; replacing all double-kernel convolution cross-stage partial modules C3k2 modules in the model with double-kernel convolution cross-stage partial dark image restoration modules C3k2-Dblock to optimize the YOLOv11 target detection algorithm and construct an optimized cow ear tag detection model; training the optimized cow ear tag detection model using a training set; and detecting cow ear tag images using the trained cow ear tag detection model.
[0005] On the basis of the method embodiment content, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the application comprises: acquiring a plurality of ear tag images of cattle, and assembling the images into a model construction dataset, which comprises: collecting a plurality of original ear tag images of cattle, sharpening and denoising each ear tag image to obtain a plurality of preprocessed images, and assembling the preprocessed images into the model construction dataset.
[0006] On the basis of the method embodiment content, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the application comprises: replacing the third layer, the fifth layer and the seventh layer of ordinary convolution in the YOLOv11 backbone network with adaptive downsampling Adown, which comprises: the adaptive downsampling Adown uses a generalized efficient layer aggregation network GELAN, which comprises a cross-stage partial network CSPNet and an efficient layer aggregation network ELAN.
[0007] On the basis of the method embodiment content, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the application comprises: introducing dynamic sample attention scale sequence fusion into the 25th layer and the 26th layer of the neck network of YOLOv11, which comprises: the dynamic sample attention scale sequence fusion is obtained by combining an adaptive multi-scale feature fusion mechanism ASFF and a content-aware upsampling method DySample.
[0008] On the basis of the method embodiment content, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the application comprises: replacing all double-kernel convolution cross-stage partial modules C3k2 modules in the model with double-kernel convolution cross-stage partial dark image restoration modules C3k2-Dblock, which comprises: the double-kernel convolution cross-stage partial dark image restoration module C3k2-Dblock is obtained by improving the double-kernel convolution cross-stage partial module C3k2 module through a dense block DBlock in dark image restoration DarkIR, and the dense block DBlock fuses a dense connection DenseNet, local feature fusion LFF and residual learning ResNet.
[0009] On the basis of the method embodiment content, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the application comprises: training the optimized ear tag detection model using a training set, which comprises: training the optimized ear tag detection model using the model construction dataset to obtain a trained ear tag detection model.
[0010] On the basis of the method embodiment content, the cow ear tag detection method based on the convolutional neural network provided in the embodiment of the application adopts the trained cow ear tag detection model to detect the cow ear tag image, and includes: after obtaining the trained cow ear tag detection model, saving the model weight parameter with the extension name ".pt"; reloading the model weight parameter to obtain a loaded cow ear tag detection model, and adopting the loaded cow ear tag detection model to detect the cow ear tag image.
[0011] In a second aspect, an embodiment of the application provides a cow ear tag detection device based on a convolutional neural network, comprising: a first main module configured to obtain a plurality of cow ear tag images to form a model construction dataset; a second main module configured to replace the third layer, the fifth layer and the seventh layer of ordinary convolution in the YOLOv11 backbone network with adaptive down-sampling Adown; a third main module configured to introduce a dynamic sample attention scale sequence fusion at the 25th layer and the 26th layer of the YOLOv11 neck network; a fourth main module configured to replace all double-core convolutional cross-stage partial modules C3k2 modules in the model with double-core convolutional cross-stage partial dark image restoration modules C3k2-Dblock, optimize the YOLOv11 target detection algorithm, and construct an optimized cow ear tag detection model; a fifth main module configured to train the optimized cow ear tag detection model using a training set; and a sixth main module configured to detect the cow ear tag image using the trained cow ear tag detection model.
[0012] In a third aspect, an embodiment of the application provides an electronic device, comprising:
[0013] at least one processor, at least one memory, and a communication interface; wherein,
[0014] The processor, the memory, and the communication interface communicate with each other;
[0015] The memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the cow ear tag detection method based on the convolutional neural network provided in any one of the various implementation manners of the first aspect.
[0016] In a fourth aspect, an embodiment of the application provides a non-transitory computer readable storage medium storing computer instructions, and the computer instructions cause a computer to execute the cow ear tag detection method based on the convolutional neural network provided in any one of the various implementation manners of the first aspect.
[0017] The ear tag detection method and device based on a convolutional neural network provided by the embodiment of the present application optimize YOLOv11, improve the ear tag detection performance, and increase mAP50 by 2.3% compared with YOLOv11 and increase mAP50-95 by 2.7%. The general camera is adopted, the cost is lower than that of the electronic ear tag scheme, the non-contact detection reduces the safety risk, and the device is suitable for complex scenes such as light and pollution. The device can be integrated with a digital system to realize efficient batch recognition, help fine feeding and disease prevention and control, and promote intelligent management of the livestock industry. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 The ear tag detection method based on a convolutional neural network provided by the embodiment of the present application is shown in the flowchart.
[0020] Figure 2 The ear tag detection device based on a convolutional neural network provided by the embodiment of the present application is shown in the structural diagram.
[0021] Figure 3 The physical structure of the electronic device provided by the embodiment of the present application is shown in the diagram.
[0022] Figure 4 The improved YOLOv11 ear tag detection network structure provided by the embodiment of the present application is shown in the diagram.
[0023] Figure 5 The improved YOLOv11 ear tag detection result provided by the embodiment of the present application is shown in the diagram. DETAILED DESCRIPTION
[0024] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application. In addition, the technical features in each of the embodiments or in a single embodiment provided by the present application can be combined with each other at will to form a feasible technical solution, and such combination is not subject to the order of steps and / or structure mode, but should be based on the fact that it can be realized by those of ordinary skill in the art. When the combination of technical solutions appears to be contradictory or unfeasible, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection required by the present application. If there are step numbers in the following embodiments, they are set only to facilitate the description and explanation, and the order between the steps is not limited in any way. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those of ordinary skill in the art.
[0025] The embodiment of the present application provides a kind of ear tag detection method based on convolutional neural network, refer to Figure 1 , the method comprises: obtaining multiple ear tag images, model construction dataset is formed;The 3rd layer, the 5th layer and the 7th layer ordinary convolution in YOLOv11 main network are replaced by adaptive downsampling Adown;Dynamic sample attention scale sequence fusion is introduced in the 25th layer and the 26th layer of YOLOv11 neck network;All double-core convolution cross-stage partial module C3k2 module in model is replaced by double-core convolution cross-stage partial dark image restoration module C3k2-Dblock, YOLOv11 target detection algorithm is optimized, and the optimized ear tag detection model is constructed;The training set is used to train the optimized ear tag detection model;The trained ear tag detection model is used to detect ear tag image.
[0026] Based on the content of the above method embodiment, as an optional embodiment, the ear tag detection method based on convolutional neural network provided in the embodiment of the present application, the multiple ear tag images are obtained, and the model construction dataset is formed, comprising: collecting multiple original ear tag images, sharpening and denoising each ear tag image to obtain multiple preprocessed images, and the multiple preprocessed images form the model construction dataset.
[0027] Refer to Figure 4, acquire multiple ear tag images of cattle, and perform sharpening and denoising, replace the ordinary convolution in the YOLOv11 backbone network, i.e., the 3rd layer, the 5th layer and the 7th layer, with Adown downsampling, introduce a dynamic sample attention scale sequence fusion ASF-DySample at the 25th and 26th layers of the neck network, and replace all C3k2 modules in the model with C3k2-Dblock, finally optimize the YOLOv11 target detection algorithm, construct a cattle ear tag detection model, and train the optimized cattle ear tag target detection model based on the training set; and detect the cattle ear tag image by using the trained cattle ear tag target detection model and the image to be detected.
[0028] Based on the content of the above method embodiment, as an optional embodiment, the cattle ear tag detection method based on the convolutional neural network provided in the embodiment of the application, the replacement of the 3rd layer, the 5th layer and the 7th layer ordinary convolution in the YOLOv11 backbone network with the adaptive downsampling Adown, comprises: the adaptive downsampling Adown uses a generalized efficient layer aggregation network GELAN, which includes a cross-stage partial network CSPNet and an efficient layer aggregation network ELAN.
[0029] The ordinary convolution in the YOLOv11 backbone network, i.e., the 3rd layer, the 5th layer and the 7th layer, is replaced with the Adown downsampling. The ADOWN downsampling uses a generalized efficient layer aggregation network (GELAN) that comprehensively considers the lightweight, inference speed and accuracy. The GELAN integrates the design concepts of CSPNet and ELAN, creating a flexible network architecture that can be adjusted according to different application requirements and computing resources. CSPNet reduces redundant computation by splitting and merging feature maps, while ELAN enhances the representation ability of features using layer aggregation. Unlike traditional deep network architectures that rely on specific types of computation units, GELAN allows the use of various types of computation blocks within its framework, such as traditional convolution layers, deep separable convolutions or other types of novel computation units.
[0030] GELAN improves the convolution operation, so that the network can better retain and utilize gradient information during training. Gradient information is crucial for the optimization process in deep learning, as it indicates how the loss function changes with respect to the model parameters. In this way, GELAN achieves higher parameter efficiency than traditional convolutional neural networks, maintaining high detection accuracy even with smaller network volumes. It can solve the information loss problem of traditional convolutional neural networks when processing target detection tasks and improve the parameter utilization efficiency of the model. GELAN optimizes the network structure through a gradient path planning technique, enabling the network to retain the integrity of input information while more effectively optimizing weights, thereby improving detection accuracy.
[0031] Based on the content of the above method embodiment, as an optional embodiment, the cow ear tag detection method based on the convolutional neural network provided in the embodiment of the application introduces a dynamic sample attention scale sequence fusion at the 25th and 26th layers of the YOLOv11 neck network, which includes that the dynamic sample attention scale sequence fusion is obtained by combining an adaptive multi-scale feature fusion mechanism ASFF and a content-aware upsampling method DySample.
[0032] The dynamic sample attention scale sequence fusion ASF-DySample is introduced at the 25th and 26th layers of the neck network. The two modules of ASFF and DySample work together to significantly improve the performance of the YOLO series model (especially for small targets) in multi-scale feature fusion and upsampling.
[0033] ASFF solves the problem of information conflict and inconsistency in multi-scale feature fusion of traditional feature pyramid networks (such as FPN, PANet). In target detection, different scale feature maps (usually from different depths of the backbone network) contain information of different semantic levels and spatial details. Simply connecting or adding them together may interfere with the reliable information of the deep layer (low resolution, strong semantics but fuzzy details) by the noise of the shallow layer (high resolution, rich details but weak semantics), and vice versa, resulting in poor fusion effect, especially harmful to small target detection.
[0034] ASFF core idea: Sequential fusion path: ASFF does not fuse all scale features at once, but adopts an ordered and adaptive fusion process. Generally, it performs sequential fusion along the scale of the feature pyramid (from large to small or from small to large). Spatial and channel attention: This is the soul of ASFF. When fusing two adjacent scale feature maps (e.g., fusing P_i with higher resolution and P_i+1 with lower resolution but stronger semantics to generate a new P_i): Channel attention: Apply channel attention (such as SE module or its variants) to P_i+1 (providing semantic guidance) to generate a channel weight vector. This aims to emphasize the feature channels in P_i+1 that are more relevant to the current task.
[0035] Spatial attention: Apply spatial attention (such as convolution to generate a spatial weight map) to P_i (providing spatial details). This aims to focus on more important spatial regions in P_i (such as regions that may contain the target).
[0036] Adaptive weighted fusion: Up-sample (usually with interpolation) the channel-weighted P_i+1 to the same resolution as P_i. Then, weight-sum the spatial-weighted P_i and the up-sampled channel-weighted P_i+1. The weights here are not fixed, but learned through a small network (such as 1x1 convolution) that considers the context information of the two feature maps to dynamically decide which scale information should be trusted more at each position.
[0037] ASFF adaptively selects and fuses information through attention mechanisms, effectively reducing conflicts and inconsistencies between different scale features. Channel attention focuses on important semantic features, and spatial attention focuses on key spatial positions, making the fused features more discriminative. The fusion of shallow detail information and deep semantic information is more accurate, significantly improving the feature representation of small targets (such as the cow ear tag which may occupy a small proportion in the image). The entire fusion process (including attention modules and fusion weight generation) is differentiable and can be optimized through training data.
[0038] DySample provides an efficient, lightweight, and high-performance up-sampling method for enlarging low-resolution feature maps to high-resolution in the feature pyramid to facilitate fusion with shallow features. Traditional up-sampling methods (such as nearest neighbor interpolation, bilinear interpolation) are fixed and parameterless, unable to adapt to image content; while learnable up-sampling (such as transposed convolution) has high computational complexity and is prone to introducing checkerboard artifacts.
[0039] DySample core idea: The core of DySample is to use the local information of the input feature map itself to dynamically generate an up-sampling kernel. Dynamic kernel generation: For each position (i, j) to be filled on the target high-resolution map, DySample does not use a pre-defined fixed kernel (such as bilinear kernel), but according to the information of the corresponding position and its neighborhood of the input low-resolution feature map, it calculates a small, content-related up-sampling kernel K(i, j) in real time. The size of this kernel is usually very small (such as 4x4). Light generator: The network structure for generating this dynamic kernel K(i, j) is very simple and efficient. It is usually composed of several lightweight operations: extract features from the low-resolution input feature map in the area (e.g. 4x4 or 5x5) around the low-resolution coordinates (i / s, j / s) mapped back from the target position (i, j) (s is the scaling factor). Apply a very small network (e.g. one or more 1x1 convolution layers) to the local area features to directly predict the up-sampling kernel weights K(i, j) (a 4x4 weight matrix) for that position. Weighted summation: Use the generated dynamic kernel K(i, j) to perform weighted summation on the corresponding local neighborhood (4x4 area) of the input feature map to get the output value of the target position (i, j).
[0040] DySample up-sampling kernel is dynamically generated according to the local content of the input image, which can better adapt to edges, textures and other structures, produce clearer and more natural high-resolution features, and reduce blurring and artifacts. Compared with transposed convolution, the kernel generation network of DySample is extremely light (mainly 1x1 convolution), with greatly reduced computational complexity and parameter quantity, which can be almost ignored. Due to its working principle (each position generates a kernel independently, the kernel size is small and the calculation is overlapped), it effectively avoids the common checkerboard artifact problem of transposed convolution. DySample can very conveniently replace the up-sampling operation (interpolation or transposed convolution) in existing models (such as the FPN / PANet part of YOLO). Its performance improvement over fixed interpolation and transposed convolution is verified on multiple tasks (such as object detection, semantic segmentation).
[0041] The ASFF inevitably involves an upsampling operation (enlarging deep low-resolution features to be fused with shallow high-resolution features) when sequentially fusing different scale feature maps. At this time, the DySample serves as an ideal upsampler in the ASFF fusion step. The ASFF is responsible for adaptively selecting and fusing different scale semantic and spatial information and solving fusion conflicts. The DySample is responsible for reconstructing spatial details (enlarging feature maps) in a high quality and high efficiency in the fusion process, and providing the ASFF with more accurate high-resolution feature representations. The ASFF ensures that the information of the shallow layer (containing small target details) is effectively utilized in the fusion and is not overwhelmed by deep layer information. The DySample ensures that the deep semantic information can retain clearer structures after upsampling, and provides more accurate semantic guidance for shallow details, to jointly improve the feature representation capability and detection accuracy for small targets. The lightweight nature of the DySample makes the ASFF fusion process increase the computational overhead very limited while introducing strong adaptive capability, thus maintaining the overall efficiency of the model.
[0042] The ASFF is an adaptive multi-scale feature fusion mechanism that effectively solves the information conflict problem in feature pyramid fusion by sequentially fusing paths combined with channel and spatial attention and dynamic fusion weights, and significantly improves the feature representation capability of the model (especially for small targets). The DySample is an efficient and content-aware upsampling method that dynamically generates a tiny upsampling kernel, which can provide clearer and more natural high-resolution feature reconstruction at a very low computational cost compared to fixed interpolation and transposed convolution. The two are closely combined: the DySample provides high-quality upsampling for the ASFF, and the ASFF uses the results of the DySample for adaptive fusion. This collaborative work greatly enhances the model's ability to handle multi-scale targets (especially small targets such as cow ear tags), while maintaining the efficiency of the model.
[0043] Based on the content of the above method embodiment, as an optional embodiment, the cow ear tag detection method based on a convolutional neural network provided in the embodiment of the application replaces all double-kernel convolution cross-stage partial modules C3k2 modules in the model with double-kernel convolution cross-stage partial dark image restoration modules C3k2-Dblock, which includes: the double-kernel convolution cross-stage partial dark image restoration module C3k2-Dblock is obtained by improving the double-kernel convolution cross-stage partial module C3k2 module through a dense block DBlock in dark image restoration DarkIR, and the dense block DBlock fuses a dense connection DenseNet, a local feature fusion LFF, and a residual learning ResNet.
[0044] Replace all C3k2 modules in the model with C3k2-Dblock. The core idea of DBlock comes from DenseNet, but usually combines the idea of residual learning (ResNet) to form a residual dense block (RDB) or similar variants.
[0045] Dense connectivity is the most core feature of DBlock. The input of each layer (convolutional layer) in the block not only comes from the output of the previous layer, but also comes from the output of all previous layers in the block. Formally: the input of the l-th layer x_l = Concat([x_0, x_1,..., x_{l-1}]), where x_0 is the initial input of the DBlock.
[0046] This feature forces the network to retain and reuse low-level features (such as edges, textures) from shallow layers and abstract features from deep layers. Each layer can directly obtain the gradient from the input and all previous layers, greatly improving the backpropagation of the gradient. Each layer receives a feature concatenation containing different levels of abstract information, enhancing the richness and expressiveness of the features. Since the features are highly reused, each layer can be designed to be narrower (i.e. the out_channels of the convolution can be smaller), and the overall efficiency may be higher than simply stacking convolutional layers.
[0047] Another feature of DBlock is Local Feature Fusion (LFF): directly concatenating the features of all previous layers will cause the number of channels to increase dramatically, increasing the computational burden of subsequent layers. To solve this problem, DBlock will add a 1x1 convolutional layer (or a small bottleneck structure) after the dense connection (usually the last of the block). The high channel dimension after concatenation is compressed, reducing the computational load. Selectively fuse information from different layers to learn a compact and informative feature representation as the output of the DBlock. This 1x1 convolutional layer is usually called the Local Feature Fusion Layer (LFF Layer).
[0048] On top of the standard Dense Block, a global residual connection is usually added. That is, the final output of the whole DBlock F_out = LFF(DenseBlock(x_in)) + x_in. The advantage is to allow the network to learn identity transformation more easily, which means if no complex feature transformation is needed, the network can simply pass the input (which is important in image enhancement, because the input image itself already contains a lot of useful information). Also, it can further alleviate the vanishing gradient problem, and speed up the convergence. It ensures that the low-level information (such as the overall structure) of the original input can be passed to the subsequent layers without loss.
[0049] Intra-block structure details (typical):
[0050] Input: x_in
[0051] Layer 1: Conv(3x3, Growth Rate G) -> ReLU. Input: x_in
[0052] Layer 2: Conv(3x3, G) -> ReLU. Input: Concat(x_in, layer1_out)
[0053] Layer 3: Conv(3x3, G) -> ReLU. Input: Concat(x_in, layer1_out, layer2_out)
[0054] ...(There can be more layers, such as 4 or 5 layers)
[0055] Local feature fusion: Conv(1x1, Output Channels). Input: Concat(x_in, layer1_out, layer2_out, layer3_out,...)
[0056] Output: LFF_out + x_in (residual connection)
[0057] The dense connection and local fusion mechanism of the DBlock enables the DBlock to effectively extract and integrate multi-level features (from details to semantics) from input with severe noise and weak information, and provide a rich information base for subsequent image reconstruction (such as denoising, brightening, and contrast adjustment). The shallow layer detail information is retained and transmitted to the deep layer through the dense connection, which helps to identify the fine texture in the image. At the same time, the abstract features learned by the deep network help to distinguish the real image structure and noise. The gradient disappearance problem is particularly serious in the deep network, and the dense and residual connection design of the DBlock ensures the smooth flow of information and gradient. Although the structure is complex, by controlling the Growth Rate (G) and using 1x1 convolution dimension reduction, the DBlock maintains a relatively reasonable model size and computational complexity while providing powerful performance.
[0058] The DBlock is a powerful feature extraction module that integrates the essence of DenseNet, local feature fusion (LFF) and residual learning (ResNet). It is optimized for the core challenges of information deficiency, severe noise interference and gradient propagation difficulty in low-light image enhancement by forcing feature reuse, promoting information flow, efficiently fusing multi-scale features and stabilizing the training process. The DBlock is one of the key architectural innovations of the DarkIR model that effectively improves the quality of low-light images, restores lost details and suppresses noise.
[0059] Based on the content of the above method embodiment, as an optional embodiment, the cow ear tag detection method based on a convolutional neural network provided in the embodiment of the application is trained by using a training set on the optimized cow ear tag detection model, and includes: training the optimized cow ear tag detection model by using a data set constructed by the model to obtain a trained cow ear tag detection model.
[0060] The cow ear tag detection model is trained by using the data set constructed by the model, and a trained cow ear tag detection model is obtained. The size of the input image is set to 640x640, the batch size is 16, the number of training iterations is 100, the initial learning rate is 0.01, and the optimizer uses SGD.
[0061] The comparison between the detection results of the application and YOLOv11 is shown in Table 1:
[0062] Table 1
[0063]
[0064] As can be seen from the above table, the performance of the application is obviously improved after improvement on YOLOv11, with an increase of 2.3% in mAP50 and an increase of 2.7% in mAP50-95 compared with YOLOv11.
[0065] In order to verify the effectiveness of the three improvements, the present application carries out the following six groups of ablation experiments for the three innovations, and the ablation experiment results are shown in Table 2:
[0066] Table 2
[0067]
[0068] Through the data in the three tables in the above comparative experiments and ablation experiments, the effective improvement of the algorithm can be obviously felt.
[0069] Based on the content of the above method embodiment, as an optional embodiment, the method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the present application adopts a trained ear tag detection model to detect ear tag images, and includes: after obtaining the trained ear tag detection model, saving the model weight parameters with an extension name of ".pt"; reloading the model weight parameters to obtain a loaded ear tag detection model, and adopting the loaded ear tag detection model to detect ear tag images. The detection result can be referred to Figure 5 .
[0070] The method for detecting ear tags of cattle based on a convolutional neural network provided in the embodiment of the present application optimizes YOLOv11, improves the ear tag detection performance, and the mAP50 is improved by 2.3% compared with YOLOv11 and by 2.7% compared with mAP50-95; a general camera is adopted, the cost is lower than that of an electronic ear tag scheme, non-contact detection reduces safety risks, and the method is suitable for complex scenes such as illumination and contamination; the method can be integrated with a digital system to realize efficient batch recognition, help fine feeding and disease prevention and control, and promote intelligent management of the livestock industry.
[0071] The implementation basis of each embodiment of the present application is that the processing of the device with processor function is programmed (i.e., software). Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present application can be packaged into various modules. Based on this actual situation, on the basis of the above embodiments, the embodiment of the present application provides a device for detecting ear tags of cattle based on a convolutional neural network, which is used to execute the method for detecting ear tags of cattle based on a convolutional neural network in the above method embodiment. Referring to Figure 2The device comprises: a first main module for realizing acquisition of multiple ear tags of cattle images to constitute a model construction dataset; a second main module for realizing replacement of normal convolution in the 3rd layer, the 5th layer and the 7th layer of a YOLOv11 backbone network with adaptive down sampling Adown; a third main module for realizing introduction of a dynamic sample attention scale sequence fusion at the 25th layer and the 26th layer of a YOLOv11 neck network; a fourth main module for realizing replacement of all double-core convolution cross-stage partial modules C3k2 modules in the model with double-core convolution cross-stage partial dark image restoration modules C3k2-Dblock, optimizing a YOLOv11 target detection algorithm to construct an optimized ear tag detection model of cattle; a fifth main module for realizing training of the optimized ear tag detection model of cattle by using a training set; and a sixth main module for realizing detection of ear tag images of cattle by using the trained ear tag detection model of cattle.
[0072] The ear tag detection device based on the convolutional neural network provided in the embodiment of the application adopts Figure 2 a plurality of modules in the YOLOv11, improves the ear tag detection performance, and increases mAP50 by 2.3% and mAP50-95 by 2.7% compared with YOLOv11; the general camera is adopted, the cost is lower than that of the electronic ear tag scheme, the non-contact detection reduces the safety risk, and the device is suitable for complex scenes such as illumination and contamination; the device can be integrated with a digital system to realize efficient batch recognition, assist fine feeding and disease prevention and control, and promote intelligent management of the livestock industry.
[0073] It should be noted that the device in the device embodiment provided in the application can be used to realize the methods in the method embodiments and can be used to realize the methods in other method embodiments provided in the application, the difference is only that the corresponding function modules (i.e., software) are set, the principle is basically the same as that of the above-mentioned device embodiments, as long as the technical personnel in the art can obtain the corresponding technical means by combining the technical features on the basis of the above-mentioned device embodiments, and the technical solutions formed by these technical means, on the premise of ensuring the practicability of the technical solutions, the device in the above-mentioned device embodiment can be improved, so as to obtain the corresponding device class embodiment (i.e., software) for realizing the methods in other method class embodiments. For example:
[0074] Based on the content of the above-mentioned device embodiment, as an optional embodiment, the ear tag detection device based on the convolutional neural network provided in the embodiment of the application further comprises a first sub-module for realizing the acquisition of multiple ear tag images of cattle to constitute a model construction dataset, which comprises: collecting multiple original ear tag images of cattle, sharpening and denoising each ear tag image to obtain multiple preprocessed images, and the multiple preprocessed images constitute the model construction dataset.
[0075] Based on the content of the above device embodiment, as an optional embodiment, the cow ear tag detection device based on the convolutional neural network provided in the embodiment of the application further includes a second sub-module for replacing the third layer, the fifth layer and the seventh layer of the YOLOv11 backbone network with adaptive downsampling Adown, including: the adaptive downsampling Adown uses a generalized efficient layer aggregation network GELAN, including a cross-stage partial network CSPNet and an efficient layer aggregation network ELAN.
[0076] Based on the content of the above device embodiment, as an optional embodiment, the cow ear tag detection device based on the convolutional neural network provided in the embodiment of the application further includes a third sub-module for introducing a dynamic sample attention scale sequence fusion at the 25th layer and the 26th layer of the YOLOv11 neck network, including: the dynamic sample attention scale sequence fusion is obtained by combining an adaptive multi-scale feature fusion mechanism ASFF and a content-aware upsampling method DySample.
[0077] Based on the content of the above device embodiment, as an optional embodiment, the cow ear tag detection device based on the convolutional neural network provided in the embodiment of the application further includes a fourth sub-module for replacing all double-core convolution cross-stage partial modules C3k2 modules in the model with double-core convolution cross-stage partial dark image restoration modules C3k2-Dblock, including: the double-core convolution cross-stage partial dark image restoration module C3k2-Dblock is obtained by improving the double-core convolution cross-stage partial module C3k2 module through a dense block DBlock in dark image restoration DarkIR, and the dense block DBlock fuses a dense connection DenseNet, a local feature fusion LFF and a residual learning ResNet.
[0078] Based on the content of the above device embodiment, as an optional embodiment, the cow ear tag detection device based on the convolutional neural network provided in the embodiment of the application further includes a fifth sub-module for training the optimized cow ear tag detection model using a training set, including: training the optimized cow ear tag detection model using a data set constructed by the model to obtain a trained cow ear tag detection model.
[0079] Based on the content of the above device embodiment, as an optional embodiment, the ear tag detection device based on the convolutional neural network provided in the embodiment of the application further includes: a sixth sub-module, configured to implement detection of the ear tag image by using the trained ear tag detection model, including: after obtaining the trained ear tag detection model, saving the model weight parameter with an extension name of “.pt”; reloading the model weight parameter to obtain a loaded ear tag detection model, and detecting the ear tag image by using the loaded ear tag detection model.
[0080] The method of the embodiment of the application is implemented by relying on an electronic device, and therefore it is necessary to introduce the related electronic device. For this purpose, the embodiment of the application provides an electronic device, as shown in the accompanying drawings, which includes at least one processor, a communications interface, at least one memory and a communications bus, wherein the at least one processor, the communications interface and the at least one memory complete mutual communication through the communications bus. The at least one processor can call the logical instructions in the at least one memory to execute all or part of the steps of the method provided in the foregoing various method embodiments. Figure 3
[0081] In addition, the logical instructions in the at least one memory described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the various method embodiments of the application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program codes that can be stored in the medium.
[0082] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment of the application according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0083] Those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary universal hardware platform, and of course can also be implemented by hardware, through the description of the above embodiments. Based on such understanding, the above technical solutions can be embodied in the form of a software product in essence or in the form of a part of the prior art. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment or some parts of the embodiment.
[0084] The flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the system, method and computer program product according to the embodiments of the present application. Based on this understanding, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which includes one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order from that noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0085] It should be noted that the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without more limitations, the elements defined by the statement "comprising" do not exclude the presence of other identical elements in the process, method, article or device including the elements. Any "predetermined threshold", "preset threshold" or similar expression, if not marked with a specific numerical value, can be determined by a person skilled in the art through simple experiments or corresponding debugging.
[0086] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting cow ear tags based on convolutional neural networks, characterized in that, include: Multiple images of cow ear tags were acquired to form a dataset for model construction. The ordinary convolutional layers (layers 3, 5, and 7) of the YOLOv11 backbone network were replaced with adaptive downsampling (Adown). Dynamic sample attention scale sequence fusion was introduced in layers 25 and 26 of the YOLOv11 neck network. All C3k2 modules in the dual-kernel convolutional cross-stage were replaced with C3k2-Dblock modules for dark image recovery in the dual-kernel convolutional cross-stage, thus optimizing the YOLOv11 object detection algorithm and constructing an optimized cow ear tag detection model. The optimized cow ear tag detection model was trained using the training set. The trained cow ear tag detection model was then used to detect cow ear tag images.
2. The method for detecting cow ear tags based on convolutional neural networks according to claim 1, characterized in that, The process of acquiring multiple cow ear tag images to form a model construction dataset includes: acquiring multiple original cow ear tag images, performing data enhancement processing on each cow ear tag image such as sharpening, increasing brightness, increasing noise, and inverting phase to obtain multiple pre-processed images, which together form the model construction dataset.
3. The method for detecting cow ear tags based on convolutional neural networks according to claim 2, characterized in that, The replacement of the ordinary convolutions in layers 3, 5, and 7 of the YOLOv11 backbone network with adaptive downsampling (Adown) includes: the adaptive downsampling Adown uses a generalized efficient layer aggregation network (GELAN), which includes the cross-stage partial network CSPNet and the efficient layer aggregation network ELAN.
4. The method for detecting cow ear tags based on convolutional neural networks according to claim 3, characterized in that, The introduction of dynamic sample attention scale sequence fusion in layers 25 and 26 of the YOLOv11 neck network includes: the dynamic sample attention scale sequence fusion is obtained by combining the adaptive multi-scale feature fusion mechanism ASFF with the content-aware upsampling method DySample.
5. The method for detecting cow ear tags based on convolutional neural networks according to claim 4, characterized in that, The step of replacing all the dual-kernel convolutional cross-stage modules C3k2 modules in the model with dual-kernel convolutional cross-stage dark image restoration modules C3k2-Dblock includes: the dual-kernel convolutional cross-stage dark image restoration module C3k2-Dblock is obtained by improving the dual-kernel convolutional cross-stage modules C3k2 modules through dense blocks DBlock in Dark IR, and the dense blocks DBlock fuse dense connections DenseNet, local feature fusion LFF and residual learning ResNet.
6. The method for detecting cow ear tags based on convolutional neural networks according to claim 5, characterized in that, The step of training the optimized cow ear tag detection model using a training set includes: using the model to construct a dataset to train the optimized cow ear tag detection model, thereby obtaining a trained cow ear tag detection model.
7. The method for detecting cow ear tags based on convolutional neural networks according to claim 6, characterized in that, The step of using a trained cow ear tag detection model to detect cow ear tags in an image includes: after obtaining the trained cow ear tag detection model, saving the model weight parameters with the extension ".pt"; reloading the model weight parameters to obtain a loaded cow ear tag detection model; and using the loaded cow ear tag detection model to detect cow ear tags in an image.
8. A bovine ear tag detection device based on a convolutional neural network, characterized in that, include: The first main module is used to acquire multiple cow ear tag images and form a dataset for model construction. The second main module is used to replace the ordinary convolutions in layers 3, 5, and 7 of the YOLOv11 backbone network with adaptive downsampling (Adown). The third main module is used to introduce dynamic sample attention scale sequence fusion in layers 25 and 26 of the YOLOv11 neck network. The fourth main module is used to replace all the dual-kernel convolutional cross-stage modules (C3k2) in the model with the dual-kernel convolutional cross-stage dark image recovery module (C3k2-Dblock) to optimize the YOLOv11 object detection algorithm and build an optimized cow ear tag detection model. The fifth main module is used to train the optimized cow ear tag detection model using the training set. The sixth main module is used to detect cow ear tag images using the trained cow ear tag detection model.
9. An electronic device, characterized in that, include: At least one processor, at least one memory, and a communication interface; wherein, The processor, memory, and communication interface communicate with each other; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to perform the method described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Remote sensing target detection method based on convolutional neural network
CN119152367A
Multi-head cattle identification and tracking algorithm based on improved YOLOv10s
CN119418396A
Trademark small target detection method based on improved YOLOv8
CN120198637A
Transfer learning-based method for improved VGG16 network pig identity recognition
WO2022252272A1
Method and system for adaptive feature fusion in convolutional neural network
WO2023077998A1