Body condition scoring method for dairy cows based on attention mechanism and lightweight convolutional neural network

Through the method based on attention mechanism and lightweight convolutional neural network, the existing cow body condition assessment methods are solved, and more efficient and accurate cow body condition scores are achieved.

CN114997725BActive Publication Date: 2025-05-23ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210759970.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-05-23
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing methods for assessing the physical condition of dairy cows are subjective, labor-intensive, and limiting the number and frequency of scores.

Method used

A cow body condition scoring method based on attention mechanism and lightweight convolutional neural network was adopted. A cow tail picture was taken through a 2D visible light camera, a convolutional neural network model was constructed, and a preprocessed image was used for scoring.

Benefits of technology

It effectively eliminates subjectivity in the scoring process, reduces pressure on animals, increases the number and frequency of scoring cows, and improves the accuracy and efficiency of scoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114997725B_ABST
    Figure CN114997725B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for scoring the body condition of dairy cows based on an attention mechanism and a lightweight convolutional neural network, comprising: using a 2D visible light camera to take a picture of a dairy cow's tail to obtain a dairy cow body condition data set; constructing a convolutional neural network model, and using the dairy cow body condition data set for pre-training to obtain a trained convolutional neural network model; obtaining a dairy cow picture to be detected and pre-processing it to obtain a pre-processed image; the pre-processing includes a standardization operation and data enhancement; inputting the pre-processed image into the trained convolutional neural network model to obtain a dairy cow body condition scoring result. The present invention uses deep learning technology to eliminate the problems of strong subjectivity of manual scoring; the present invention adds an efficient channel attention mechanism to enhance the network's ability to extract dairy cow body conditions; and at the same time introduces a Dropout operation to enhance the generalization ability of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cow physical condition detection, and in particular to a cow physical condition scoring method based on an attention mechanism and a lightweight convolutional neural network. Background Art

[0002] The body condition of dairy cows is an indirect indicator of the status of subcutaneous fat reserves. Studies have shown that cows with too high or too low fat reserves at calving and a rapid loss of fat reserves in early lactation can reduce milk production, negatively affect health, and reduce reproductive performance. In addition, excessive consumption of fat reserves in early lactation is associated with reduced survival of dairy cows in the herd. Therefore, timely assessment and management of dairy cows' fat reserves during lactation enables dairy cows to achieve optimal body condition, which can maintain and improve dairy cow production and welfare.

[0003] The assessment of the body condition of dairy cows is usually performed by at least two professional operators according to a standardized BCS table. Since the difference in subcutaneous fat reserves of dairy cows is more obvious visually, operators can rely on visual observation of key areas at the rear end of the cow's body (such as the spine tail, long ribs, etc.) for assessment. However, this manual method contains factors such as operator subjectivity and the requirement that the animal be restrained during the assessment process, as well as being a labor-intensive method that limits the number or frequency of dairy cows that can be scored. Therefore, there is an urgent need to develop an automatic scoring system for dairy cow body condition. Summary of the invention

[0004] The object of the present invention is to provide a cow body condition scoring method based on an attention mechanism and a lightweight convolutional neural network, which can eliminate the subjectivity in the scoring process to a certain extent, minimize the stress on animals, and increase the number and frequency of scored cows.

[0005] To achieve the above object, the present invention adopts the following technical solution: a method for scoring the body condition of dairy cows based on an attention mechanism and a lightweight convolutional neural network, the method comprising the following steps in order:

[0006] (1) Use a 2D visible light camera to take pictures of the cow's tail and obtain a cow body condition dataset;

[0007] (2) constructing a convolutional neural network model and using the cow body condition dataset for pre-training to obtain a trained convolutional neural network model;

[0008] (3) obtaining a cow image to be detected and preprocessing it to obtain a preprocessed image; the preprocessing includes a standardization operation and data enhancement;

[0009] (4) The preprocessed image is input into the trained convolutional neural network model to obtain the cow body condition score result.

[0010] The step (1) comprises the following steps in order:

[0011] (1a) Using a 2D visible light camera, take 50,000 pictures of cow tails from a bird's-eye view, and ensure that each picture of cow tails contains at least one cow tail;

[0012] (1b) Delete unclear cow tail pictures and pictures with repeated cow tail content;

[0013] (1c) The torso of each cow’s tail in the cow’s tail image is taken as the target and labeled to obtain a cow body condition dataset. The cow body condition dataset contains 8972 cow body condition images and corresponding score values.

[0014] In step (2), the convolutional neural network model includes six layers from input to output, namely:

[0015] The first layer is the first convolutional layer, which contains 24 convolution kernels, the convolution kernel size is 3×3, and the step size is 2;

[0016] The second layer is the maximum pooling layer, which contains 24 pooling kernels, the pooling kernel is 3×3, and the stride is 2;

[0017] The third layer is the deep feature extraction layer, which contains 3 attention units;

[0018] The fourth layer is the second convolution layer, which contains 1024 convolution kernels, the convolution kernel size is 1×1, and the step size is 2;

[0019] The fifth layer is the global pooling layer, which contains one pooling kernel with a size of 7×7.

[0020] The sixth layer is a fully connected layer, containing 1000 neurons;

[0021] In the deep feature extraction layer, the attention unit includes 1 first basic unit and 3 second basic units; the first basic unit includes a first branch, a second branch, a first splicing layer and a first channel fusion layer; the second basic unit includes a third branch, a fourth branch, a second splicing layer and a second channel fusion layer; the first branch has the same structure as the third branch, the first splicing layer has the same structure as the second splicing layer, and the first channel fusion layer has the same structure as the second channel fusion layer;

[0022] The first branch is divided into five layers:

[0023] The first layer is the third convolutional layer, with a convolution kernel size of 1×1 and a step size of 2;

[0024] The second layer is the first depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1;

[0025] The third layer is the first ECA module;

[0026] The fourth layer is the fourth convolutional layer, the convolution kernel size is 1×1, and the step size is 2;

[0027] The fifth layer is the first Dropout layer;

[0028] The second branch is divided into two layers:

[0029] The first layer is the second depth convolution layer, with a convolution kernel size of 3×3 and a stride of 2;

[0030] The second layer is the fifth convolution layer, with a convolution kernel size of 1×1 and a step size of 1;

[0031] The third branch is divided into five layers:

[0032] The first layer is the sixth convolutional layer, with a convolution kernel size of 1×1 and a step size of 2;

[0033] The second layer is the third depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1;

[0034] The third layer is the second ECA module;

[0035] The fourth layer is the seventh convolution layer, with a convolution kernel size of 1×1 and a step size of 2;

[0036] The fifth layer is the second Dropout layer;

[0037] The first basic unit divides the feature map obtained by the maximum pooling layer or the second basic unit into two branches in the dimension of the number of channels: a first branch and a second branch; wherein the first branch performs convolution, depth convolution, and convolution operations in sequence, and then inputs the result into the first ECA module, the first ECA module multiplies the generated channel weights with each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the first Dropout layer, and the first Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second branch performs convolution and depth convolution in sequence; the first splicing layer splices the result obtained by the first branch and the result obtained by the second branch in the dimension of the number of channels, and the first channel fusion layer rearranges the result of the first splicing layer in the dimension of the number of channels;

[0038] The second basic unit divides the feature map obtained by the first basic unit into two branches in the dimension of the number of channels: the third branch and the fourth branch; the third branch performs convolution, depth convolution and convolution operations in sequence, and then inputs the result into the second ECA module, the second ECA module multiplies the generated channel weights by each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the second Dropout layer, and the second Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second splicing layer splices the results obtained by the third branch and the results obtained by the fourth branch in the dimension of the number of channels, and the second channel fusion layer rearranges the results of the second splicing layer in the dimension of the number of channels.

[0039] In step (3), the formula for the standardization operation is:

[0040]

[0041] Where x is the pixel value of the channel, μ is the mean of the channel pixel values, and σ is the variance of the channel pixels.

[0042] In step (3), the data enhancement includes geometric transformation and color change.

[0043] It can be seen from the above technical scheme that the beneficial effects of the present invention are: first, compared with traditional manual scoring, the present invention uses deep learning technology to eliminate the problem of strong subjectivity in manual scoring; second, the present invention adds an efficient channel attention mechanism to enhance the network's ability to extract the body condition of dairy cows; at the same time, the Dropout operation is introduced to enhance the generalization ability of the network; third, in order to ensure the lightweight requirements of practical applications, the present invention trims the network structure, and the classification accuracy is higher than that of the algorithm with the same lightweight network, the network parameters are lower, and the model size is lower; fourth, the present invention can be applied to lightweight scenarios, which promotes the commercial development of dairy cow body condition scoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flow chart of the method of the present invention;

[0045] Figure 2 This is a diagram showing the enhanced effect of the cow image data in the present invention;

[0046] Figure 3 It is a workflow diagram of the convolutional neural network in the present invention;

[0047] Figure 4 It is a working flow chart of the first basic unit and the second basic unit in the present invention;

[0048] Figure 5 It is a working schematic diagram of the ECA module in the present invention. DETAILED DESCRIPTION

[0049] like Figure 1 As shown, a method for scoring the body condition of dairy cows based on an attention mechanism and a lightweight convolutional neural network comprises the following steps in order:

[0050] (1) Use a 2D visible light camera to take pictures of the cow's tail and obtain a cow body condition dataset;

[0051] (2) constructing a convolutional neural network model and using the cow body condition dataset for pre-training to obtain a trained convolutional neural network model;

[0052] (3) obtaining a cow image to be detected and preprocessing it to obtain a preprocessed image; the preprocessing includes a standardization operation and data enhancement;

[0053] (4) The preprocessed image is input into the trained convolutional neural network model to obtain the cow body condition score result.

[0054] The step (1) comprises the following steps in order:

[0055] (1a) Using a 2D visible light camera, take 50,000 pictures of cow tails from a bird's-eye view, and ensure that each picture of cow tails contains at least one cow tail;

[0056] (1b) Delete unclear cow tail pictures and pictures with repeated cow tail content;

[0057] (1c) The torso of each cow’s tail in the cow’s tail image is taken as the target and labeled to obtain a cow body condition dataset. The cow body condition dataset contains 8972 cow body condition images and corresponding score values.

[0058] In step (2), if Figure 3 As shown, the convolutional neural network model includes six layers from input to output, namely:

[0059] The first layer is the first convolutional layer, which contains 24 convolution kernels, the convolution kernel size is 3×3, and the step size is 2;

[0060] The second layer is the maximum pooling layer, which contains 24 pooling kernels, the pooling kernel is 3×3, and the stride is 2;

[0061] The third layer is the deep feature extraction layer, which contains 3 attention units;

[0062] The fourth layer is the second convolution layer, which contains 1024 convolution kernels, the convolution kernel size is 1×1, and the step size is 2;

[0063] The fifth layer is the global pooling layer, which contains one pooling kernel with a size of 7×7.

[0064] The sixth layer is a fully connected layer, containing 1000 neurons;

[0065] In the deep feature extraction layer, the attention unit includes 1 first basic unit and 3 second basic units; the first basic unit includes a first branch, a second branch, a first splicing layer and a first channel fusion layer; the second basic unit includes a third branch, a fourth branch, a second splicing layer and a second channel fusion layer; the first branch has the same structure as the third branch, the first splicing layer has the same structure as the second splicing layer, and the first channel fusion layer has the same structure as the second channel fusion layer;

[0066] The first branch is divided into five layers:

[0067] The first layer is the third convolutional layer, with a convolution kernel size of 1×1 and a step size of 2;

[0068] The second layer is the first depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1;

[0069] The third layer is the first ECA module;

[0070] The fourth layer is the fourth convolutional layer, the convolution kernel size is 1×1, and the step size is 2;

[0071] The fifth layer is the first Dropout layer;

[0072] The second branch is divided into two layers:

[0073] The first layer is the second depth convolution layer, with a convolution kernel size of 3×3 and a stride of 2;

[0074] The second layer is the fifth convolution layer, with a convolution kernel size of 1×1 and a step size of 1;

[0075] The third branch is divided into five layers:

[0076] The first layer is the sixth convolutional layer, with a convolution kernel size of 1×1 and a step size of 2;

[0077] The second layer is the third depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1;

[0078] The third layer is the second ECA module;

[0079] The fourth layer is the seventh convolution layer, with a convolution kernel size of 1×1 and a step size of 2;

[0080] The fifth layer is the second Dropout layer;

[0081] The fourth branch does not include any layers and does not process the image. For example, if there are six feature maps, one branch processes three feature maps, the third branch performs some operations on the first three feature maps, and the fourth branch retains the last three feature maps and does not operate on the last three feature maps.

[0082] like Figure 4 As shown, the first basic unit divides the feature map obtained by the maximum pooling layer or the second basic unit into two branches in the dimension of the number of channels: a first branch and a second branch; wherein the first branch performs convolution, depth convolution, and convolution operations in sequence, and then inputs the result into the first ECA module, the first ECA module multiplies the generated channel weights by each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the first Dropout layer, and the first Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second branch performs convolution and depth convolution in sequence; the first splicing layer splices the result obtained by the first branch and the result obtained by the second branch in the dimension of the number of channels, and the first channel fusion layer rearranges the result of the first splicing layer in the dimension of the number of channels;

[0083] like Figure 4 As shown, the second basic unit divides the feature map obtained by the first basic unit into two branches in the dimension of the number of channels: the third branch and the fourth branch; the third branch performs convolution, depth convolution and convolution operations in sequence, and then inputs the result into the second ECA module, the second ECA module multiplies the generated channel weights by each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the second Dropout layer, and the second Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second splicing layer splices the results obtained by the third branch and the results obtained by the fourth branch in the dimension of the number of channels, and the second channel fusion layer rearranges the results of the second splicing layer in the dimension of the number of channels.

[0084] like Figure 5 As shown in the figure, the ECA module is an efficient channel attention module, and its functions are as follows: First, for the input feature tensor (H×W×C), the average value of all pixels is calculated on the feature map of each channel through global average pooling, thereby reducing the number of parameters and calculations, and then a one-dimensional convolution with a convolution kernel of 3×3 is used to achieve the purpose of cross-channel interaction. The feature tensor output by the one-dimensional convolution is mapped by the Sigmoid (σ) function to generate channel weights between 0 and 1. Finally, the generated channel weights are multiplied element by element with the input feature tensor to generate a refined feature map.

[0085] In step (3), the formula for the standardization operation is:

[0086]

[0087] Where x is the pixel value of the channel, μ is the mean of the channel pixel values, and σ is the variance of the channel pixels.

[0088] In step (3), if Figure 2 As shown, the data enhancement includes geometric transformation and color change.

[0089] The geometric transformation includes:

[0090] Random translation: randomly move the image in the X or Y direction (or both) by a random distance;

[0091] Flip horizontally, flip the image horizontally;

[0092] Random rotation, rotate the image, the rotation degree is a random number;

[0093] Color changes refer to saturation enhancement, brightness enhancement, contrast enhancement, and sharpness enhancement.

[0094] In summary, the present invention adopts a low-cost 2D camera to collect the top-view image of the cow's tail, combined with the target monitoring and recognition capabilities of the deep learning algorithm, firstly pre-processes the acquired cow's tail image, and then classifies the feature image by adding the attention mechanism and using Dropout, etc., in order to achieve accurate and efficient scoring of the body condition of dairy cows in the natural environment of the breeding pen.

Claims

1. A method for scoring dairy cow body condition based on attention mechanism and lightweight convolutional neural network. Features: The method comprises the following steps in order: (1) Use a 2D visible light camera to take pictures of the cow's tail and obtain a cow body condition dataset; (2) constructing a convolutional neural network model and using the cow body condition dataset for pre-training to obtain a trained convolutional neural network model; (3) obtaining a cow image to be detected and preprocessing it to obtain a preprocessed image; the preprocessing includes a standardization operation and data enhancement; (4) Inputting the preprocessed image into the trained convolutional neural network model to obtain the body condition score result of the dairy cow; In step (2), the convolutional neural network model includes six layers from input to output, namely: The first layer is the first convolutional layer, which contains 24 convolution kernels, the convolution kernel size is 3×3, and the step size is 2; The second layer is the maximum pooling layer, which contains 24 pooling kernels, the pooling kernel is 3×3, and the stride is 2; The third layer is the deep feature extraction layer, which contains 3 attention units; The fourth layer is the second convolution layer, which contains 1024 convolution kernels, the convolution kernel size is 1×1, and the step size is 2; The fifth layer is the global pooling layer, which contains one pooling kernel with a size of 7×7. The sixth layer is a fully connected layer, containing 1000 neurons; In the deep feature extraction layer, the attention unit includes 1 first basic unit and 3 second basic units; the first basic unit includes a first branch, a second branch, a first splicing layer and a first channel fusion layer; the second basic unit includes a third branch, a fourth branch, a second splicing layer and a second channel fusion layer; the first branch has the same structure as the third branch, the first splicing layer has the same structure as the second splicing layer, and the first channel fusion layer has the same structure as the second channel fusion layer; The first branch is divided into five layers: The first layer is the third convolutional layer, with a convolution kernel size of 1×1 and a step size of 2; The second layer is the first depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1; The third layer is the first ECA module; The fourth layer is the fourth convolutional layer, the convolution kernel size is 1×1, and the step size is 2; The fifth layer is the first Dropout layer; The second branch is divided into two layers: The first layer is the second depth convolution layer, with a convolution kernel size of 3×3 and a stride of 2; The second layer is the fifth convolution layer, with a convolution kernel size of 1×1 and a step size of 1; The third branch is divided into five layers: The first layer is the sixth convolutional layer, with a convolution kernel size of 1×1 and a step size of 2; The second layer is the third depth convolution layer, with a convolution kernel size of 3×3 and a step size of 1; The third layer is the second ECA module; The fourth layer is the seventh convolution layer, with a convolution kernel size of 1×1 and a step size of 2; The fifth layer is the second Dropout layer; The first basic unit divides the feature map obtained by the maximum pooling layer or the second basic unit into two branches in the dimension of the number of channels: a first branch and a second branch; wherein the first branch performs convolution, depth convolution, and convolution operations in sequence, and then inputs the result into the first ECA module, the first ECA module multiplies the generated channel weights with each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the first Dropout layer, and the first Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second branch performs convolution and depth convolution in sequence; the first splicing layer splices the result obtained by the first branch and the result obtained by the second branch in the dimension of the number of channels, and the first channel fusion layer rearranges the result of the first splicing layer in the dimension of the number of channels; The second basic unit divides the feature map obtained by the first basic unit into two branches in the dimension of the number of channels: the third branch and the fourth branch; the third branch performs convolution, depth convolution and convolution operations in sequence, and then inputs the result into the second ECA module, the second ECA module multiplies the generated channel weights by each element in the basic feature map to generate a refined feature map, and inputs the refined feature map into the second Dropout layer, and the second Dropout layer randomly assigns the elements in the refined feature map to 0 with a probability of 20%; the second splicing layer splices the results obtained by the third branch and the results obtained by the fourth branch in the dimension of the number of channels, and the second channel fusion layer rearranges the results of the second splicing layer in the dimension of the number of channels.

2. The method for scoring the body condition of dairy cows based on the attention mechanism and lightweight convolutional neural network according to claim 1, Features: The step (1) comprises the following steps in order: (1a) Using a 2D visible light camera, take 50,000 pictures of cow tails from a bird's-eye view, and ensure that each picture of cow tails contains at least one cow tail; (1b) Delete unclear cow tail pictures and pictures with repeated cow tail content; (1c) The torso of each cow’s tail in the cow’s tail image is taken as the target and labeled to obtain a cow body condition dataset. The cow body condition dataset contains 8972 cow body condition images and corresponding score values.

3. The method for scoring the body condition of dairy cows based on the attention mechanism and lightweight convolutional neural network according to claim 1, Features: In step (3), the formula for the standardization operation is: In the formula, x is the pixel value of the channel, μ is the mean value of the channel pixel value, σ is the variance of channel pixels.

4. The method for scoring the body condition of dairy cows based on the attention mechanism and lightweight convolutional neural network according to claim 1, Features: In step (3), the data enhancement includes geometric transformation and color change.