Dairy cow feeding area feed intake estimation system based on computer vision
By designing a computer vision-based feeding intake estimation system for dairy cow feeding areas, using U2-Net model and deep learning technology, the problem of inaccurate estimation of dairy cow feeding intake in the existing technology is solved, and high-precision feeding intake evaluation and improvement of ranch production efficiency is achieved.
Patent Information
- Application Number
- CN202510091863.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing feed intake estimation technology for dairy cattle feeding areas cannot meet the actual needs of the pastures, and the estimates are inaccurate.
A computer vision-based feeding intake estimation system for dairy cows is designed, including a data acquisition module, a feed area segmentation module, a data processing module and an overall feeding intake estimation module. The system uses the U2-Net model to perform image segmentation, and uses feature extraction and feed intake estimation subnet for different images through feature extraction subnet and feed intake estimation subnet.
A high-precision assessment of the overall feed intake of dairy cows from a height of 2.95m is achieved, which meets the actual needs of the pasture, reduces interference to the cows and improves production efficiency.
Smart Images

Figure CN119942458A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a technical system for estimating feed intake in a feeding area of a dairy cow. Background Art
[0002] Accurate estimation of feed intake is a key part of modern and intelligent ranch management. The feed intake of dairy cows has a direct impact on the health status, genetic evaluation and breeding of dairy cows. Reasonable feed supply can promote the healthy growth of animals and improve production performance such as milk production or meat quality. On the contrary, excessive or insufficient feed intake may have a negative impact on the health of dairy cows, such as obesity, malnutrition and other problems. Therefore, accurate estimation of feed intake is crucial to maintaining the good health of dairy cows and improving the overall production efficiency of the ranch.
[0003] At present, the methods for estimating feed intake can be roughly divided into two categories: one is the behavioral and physiological model method, and the other is the visual and image analysis method. The behavioral and physiological model method is mainly based on indicators such as the cow's eating behavior, chewing time, and milk production efficiency, and infers its feed intake by establishing a regression model. For example, by monitoring the feeding behavior of dairy cows, the feed intake can be estimated based on the chewing sound; the feed intake can also be estimated by using pressure sensors to measure the feeding time and chewing time of dairy cows. This type of method focuses on analyzing the behavioral characteristics and physiological responses of individual dairy cows to indirectly evaluate their feeding situation.
[0004] The vision and image analysis principle estimates the change in the feed pile before and after feeding the cows through physical shooting equipment, algorithms, and deep learning technologies to determine the overall feed intake within the field of view. For example, an RGB-D (Red Green Blue-Depth) camera is used to collect data at a height of less than 1.5m from the ground. The feed trough is used for loading, and the depth map before and after the simulated feeding is put into the deep network model for training to estimate the corresponding feed intake.
[0005] The behavioral and physiological model method is relatively rough in estimation, with poor robustness and accuracy. Compared with the behavioral and physiological model method, the visual and image analysis method focuses more on the precise measurement of physical changes in feed, and can provide more intuitive and accurate feed intake data. However, the field of view presented by the shooting distance of the current existing technology can only capture a single cow. Because when the shooting height exceeds 2.5 meters, a large number of holes and noise will appear in the depth map taken, resulting in unreliable depth map data, and thus the estimated feed intake is inaccurate. In summary, it is difficult for existing technologies to meet the actual needs of ranches. Summary of the invention
[0006] The invention aims to solve the problem that the existing estimation of the feed intake of dairy cows in the feeding area cannot meet the actual needs of the pasture and the estimation is inaccurate.
[0007] A system for estimating the feed intake of dairy cows in a feeding area based on computer vision, comprising:
[0008] Data acquisition module: used to obtain RGB images of the feed area before and after feeding the cows;
[0009] Feed area segmentation module: Use the image segmentation model to process the RGB image, identify and extract the mask image of the feed area;
[0010] Data processing module: color matching is performed on the mask images corresponding to before and after feeding the cows to obtain the RGB image of the feed area corresponding to the mask image; a difference image is obtained based on the RGB image of the feed area in the mask image corresponding to before and after feeding the cows;
[0011] The overall feed intake estimation module: uses the overall feed intake estimation model to estimate the overall feed intake of the acquired difference image; the overall feed intake estimation model includes a feature extraction subnetwork and a feed intake estimation subnetwork;
[0012] Feature extraction subnetwork: used to extract features from the difference image obtained after processing by the data processing module;
[0013] Feed intake estimation subnetwork: The feature map obtained by the feature extraction subnetwork is passed through a multi-scale pooling layer and spliced, passed through a fully connected layer, and then processed by an activation function, and then passed through a fully connected layer to map the feature vector to a scalar value, that is, the estimated value of feed intake.
[0014] Furthermore, the image segmentation model image adopts U 2 -Net model.
[0015] Furthermore, the U 2 -Net model is an encoder-decoder structure; the encoder includes 4 layers of RSU submodules and 2 layers of RSU4F submodules, respectively recorded as encode_1 to encode_6; the decoder includes 1 layer of RSU4F submodule and 4 layers of RSU submodules, respectively recorded as decode_5 to decode_1; starting from the salient image of the last layer of encode_6, it is fused layer by layer with the salient image of each layer of the decoder, and finally the final salient mask image is generated through 1×1Conv and a Sigmoid function.
[0016] Furthermore, the process of color matching the mask images corresponding to the cows before and after feeding includes:
[0017] For the generated mask image, the RGB image of the corresponding feed area is implanted into the mask image area to obtain the RGB image of the feed area in the mask image.
[0018] Furthermore, the processing process of the feature extraction subnetwork includes:
[0019] The difference image obtained after processing by the data processing module is recorded as the original difference image, and the original difference image is sent to the feature extraction sub-network. The processing is divided into two paths. One path performs a maximum pooling operation on the original difference image; the other path sends the original difference image to the backbone network of the feature extraction sub-network for deeper feature extraction;
[0020] The backbone network of the feature extraction subnetwork is based on the RseNet network architecture, including 4 layers of Residual submodules. Each Residual submodule consists of several Bottleneck Blocks. Each Bottleneck Block includes three convolutional layers, namely 1×1 convolution, 3×3 convolution, and 1×1 convolution.
[0021] Furthermore, the specific structure of the 4-layer Residual sub-module of the backbone network of the feature extraction sub-network is as follows: the first layer contains 3 Bottleneck Blocks; the second layer contains 4 Bottleneck Blocks; the third layer contains 23 Bottleneck Blocks; the fourth layer contains 3 Bottleneck Blocks; each Bottleneck Block contains a skip connection.
[0022] Furthermore, the original difference image is fed into the feature extraction subnetwork, and the processing is divided into two paths for processing as follows:
[0023] First, the input difference image I undergoes the first 2×2 Max Pooling to obtain the feature map F1; I then undergoes 1×1 convolution and the first layer of Residual submodule to obtain the feature map F1′; F1′ is downsampled to obtain F1' downsample Then concatenate it with F1 to get the feature map F1″;
[0024] After F1 is subjected to a second 2×2 Max Pooling, feature map F2 is obtained; after F1″ is subjected to a 1×1 convolution and a second layer of Residual submodule, feature map F2′ is obtained; F2′ is downsampled to obtain F2' downsample Then concatenate it with F2 to get the feature map F2″;
[0025] Then F2 undergoes a third 2×2 MaxPooling to obtain feature map F3; F2″ undergoes a 1×1 convolution and the third layer of Residual submodule to obtain feature map F3′; F3′ is downsampled to obtain F3' downsample Then concatenate it with F3 to get the feature map F3″;
[0026] Then, F3″ is subjected to 1×1 convolution and the fourth layer of Residual submodule to obtain the feature map F4′;
[0027] The feature map F4′ is first normalized, and then the Query, Key, and Value are calculated; then the Query and Key are matrix multiplied to obtain the corresponding energy map A1; then the softmax function is passed to obtain the attention map A2; then the Value is multiplied by A2 to obtain the weighted feature map A3; finally, A3 is connected to the residual of the feature map F4′ and output.
[0028] Furthermore, in the feed intake estimation subnetwork, the Swish activation function is used when processing the activation function.
[0029] Furthermore, the overall feed intake estimation model is pre-trained. During the training process of the overall feed intake estimation model, different difference images need to be used to train the overall feed intake estimation model; in the process of obtaining different difference images, it is necessary to use a permutation and combination method to subtract the same pile of feed images pairwise according to the relationship before and after the feed to obtain difference images.
[0030] Furthermore, the loss function used in the training process of the overall feed intake estimation model is as follows:
[0031]
[0032] Where n is the number of samples, y i is the true label of the total feed intake of the i-th sample, is the model's predicted value for the overall feed intake of the ith sample.
[0033] Beneficial effects brought by the technical solution of the present invention:
[0034] According to actual production needs, the present invention designs and implements a system for estimating the total feed intake of dairy cows at a height of 2.95m. The system not only has a wide field of view coverage, which can ensure high-precision evaluation of the total feed intake of dairy cows and meet the actual needs of the ranch, but also enhances the adaptability to different environments.
[0035] The non-contact feed intake estimation method of the present invention effectively reduces interference to dairy cows and reduces the stress response caused by sensor equipment estimating feed intake, thereby helping to maintain the health of dairy cows and improve productivity. In addition, the present invention significantly saves time and labor costs through visual data processing and feature extraction. Compared with traditional processing methods, visual image analysis reduces operational complexity and makes the overall feed intake estimation of dairy cows more efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a schematic diagram of the processing flow of the feed intake estimation system for dairy cows in the feeding area based on computer vision;
[0037] Figure 2 It is the overall deployment diagram for data collection;
[0038] Figure 3 is the camera field of view;
[0039] Figure 4 This is a schematic diagram of the U2-Net model structure;
[0040] Figure 5 This is the preprocessing effect diagram of a single image;
[0041] Figure 6 This is the difference diagram before and after simulated feeding;
[0042] Figure 7 Model diagram for estimating overall feed intake;
[0043] Figure 8 Figure 2 is a diagram of the multi-scale feature fusion process. DETAILED DESCRIPTION
[0044] Based on actual production needs, this application designs and implements a system for estimating the total feed intake of dairy cows at a height of 2.95m. The system not only has a wide field of view coverage, ensuring high-precision evaluation of the total feed intake of dairy cows, but also enhances the ability to adapt to different environments. This provides more practical technology for optimizing feed costs, genetic evaluation, breeding of breeding cattle, and ensuring the health of dairy cows.
[0045] Specific implementation method 1: Combination Figure 1 To explain this embodiment,
[0046] This embodiment is a system for estimating the feed intake of dairy cows in a feeding area based on computer vision, comprising:
[0047] Data acquisition module: Use a binocular camera to obtain RGB images before and after feeding, which serve as the basic data source for subsequent analysis.
[0048] In this embodiment, a horizontal aluminum crossbar is set at a height of 2.95 meters from the ground. The binocular camera is installed at a suitable position of the crossbar, such as Figure 2 As shown, to ensure that the shooting range of multiple cows can be covered from this perspective, such as Figure 3 As shown. The data acquisition module is one of the core components of this application, and is intended to provide high-quality basic data for subsequent image processing and analysis. In the early stage of developing the model of the present invention, the module designed two different data set acquisition schemes to meet the research needs at different stages. The first type of data set is designed specifically for training segmentation models, with the aim of allowing the model to learn and understand the visual characteristics of feed under various conditions, thereby improving its recognition accuracy and adaptability. This type of data set is crucial to ensure that the model can accurately distinguish feed from other background elements under different lighting conditions. The second type of data set focuses more on exploring the impact of sunlight on the color perception of feed, and combines the simulation of the natural eating behavior of dairy cows to gain a deeper understanding of how environmental factors affect the visual recognition of feed, while intuitively capturing the changes in the amount of feed intake before and after feeding of dairy cows. In the actual use process after the model of the present invention is developed, the images of dairy cows before and after eating are directly collected (without collecting the two types of data in the training model), and the feed area segmentation module is used for segmentation processing.
[0049] Feed area segmentation module: Use the trained image segmentation model to process the RGB image to accurately identify and extract the mask image of the feed area, ensuring that the feed and background elements can be clearly distinguished.
[0050] Since the original RGB image obtained from the cattle farm is a snapshot of the real scene, it contains feed and various objects in the surrounding environment. The feed area segmentation module can accurately extract the area where the feed is located from the complex background. A new image is output, called a mask image. In this new image, only the feed area will be marked (usually white), while the background and other objects are blocked (usually black). This method is particularly suitable for processing images within a large field of view, and can effectively eliminate the interference of non-feed objects, thereby laying the foundation for more accurate estimation of feed changes in the future. Compared with traditional methods, this method is more in line with the needs of practical applications.
[0051] The image segmentation model in this embodiment adopts U 2 -Net model, which is an encoder-decoder structure, such as Figure 4As shown. The encoder includes 4 layers of RSU (Residual U-block) submodules and 2 layers of RSU4F (Factorized ResidualU-block) submodules, respectively recorded as encode_1 to encode_6; the decoder includes 1 layer of RSU4F submodules and 4 layers of RSU submodules, respectively recorded as decode_5 to decode_1. The salient image of the last layer of encode_6 is fused with the salient image of each layer of the decoder, and finally the final salient mask image is generated through 1×1Conv and a Sigmoid function. This design enables the model to capture detailed features while maintaining high resolution, which is very suitable for segmentation tasks of objects with complex textures such as feed.
[0052] It should be noted that each decoder generates a saliency map based on the feature map of the current decoder layer. Since the size of the feature maps generated by the network at different decoder layers is different, the size of the saliency map of each layer is also different, and is usually smaller than the input image size. Because the saliency maps of each layer are of different sizes, in order to be able to splice these saliency maps (that is, to fuse them together), upsampling operations are required so that they can be spliced with the saliency maps of other layers. Finally, all the spliced saliency maps (restored to the original image size by upsampling) are further processed by a 1×1 convolutional layer. The role of the 1×1 convolutional layer is to reduce the number of channels, and the spliced saliency maps are usually compressed to a single channel output. Then, through a Sigmoid function, the output is mapped to the [0,1] range, and finally a binary saliency map (i.e., saliency mask) is generated. This mask represents the salient areas in the image. Figure 4 The encoder-decoder shown in does not show 1×1Conv and a Sigmoid function.
[0053] Data processing module: Receive the above mask image and use color matching algorithm to further optimize the feed area to generate an image that reflects the actual feed distribution. Pair these processed feed images according to the time series (i.e. before and after feeding), and compare and calculate the difference between the two at the pixel level to form a difference image. This process assigns a label value representing the quality of feed intake before and after feeding to each sample, which intuitively shows the feed consumption. Finally, the difference image and its corresponding label value are sent to the feed intake estimation module.
[0054] The data processing module is mainly responsible for color matching of the mask image (black and white binary feed area mask image), optimizing image quality and generating a difference image. On this basis, the white feed area in the mask image is matched with the color of the feed area in the original color image. A new image is generated with a black background and the feed area retaining the original color. Since each image after color matching can only express a state of the feed at the current moment, it cannot express the amount of change in the feed. Therefore, this module generates a difference image by comparing two color-matched feed images and calculating the difference between them. That is, the visualization result of the amount of change in the feed.
[0055] The specific implementation method is as follows:
[0056] First, color matching of the target area is performed on all RGB images before and after simulated feeding with the mask image generated by segmentation technology. Color matching of the target area is performed on the mask image: the feed detail color of the original RGB image is transplanted to the white part of the mask image. In this way, a feed area segmentation image with the original feed color details can be obtained, rather than a mask image with a white feed area. In the actual processing process, color matching is to extract the corresponding pixel value from the original RGB image for each white pixel in the mask image, and put it on the corresponding point of the white area of the mask image.
[0057] Next, these color-matched images are cropped to a standard size of 360 × 1080 pixels, such as Figure 5 shown.
[0058] In order to evaluate the change in the amount of feed intake of dairy cows, the permutation and combination method is used to subtract the same pile of feed images according to the relationship before and after the feed to obtain difference images. For example, the first picture of the same pile of feed is taken before feeding a1, the cow goes in and eats and then exits to take a2, the cow goes in and eats and then exits to take a3, the cow goes in and eats and then exits to take a4, and the cow goes in and eats and then exits to take a5. The difference images representing the amount of feed intake are a2-a1, a3-a2, a4-a3, a5-a4, a3-a1, a4-a1, a4-a2, a5-a1, a5-a2, a5-a3, so there are C(5,2)=10 matching methods. Considering that the single feed intake of dairy cows under the conditions of this application will not exceed 10kg, the weight change reflected by the difference image is controlled within this range. In this process, in order to avoid negative values affecting the integrity of feature extraction, the absolute value operation is used to retain the complete difference information.
[0059] Finally, 8,000 difference images were selected from the processed images for further analysis. Given the large size of the images, which may consume too much time and computing resources in the training phase, this application downsampled these difference images and converted them into a smaller 120×360 pixel format, such as Figure 6 As shown. In addition, data enhancement methods such as random horizontal flipping, random vertical flipping and random rotation are also applied to increase data diversity and prevent overfitting. A total of 16,000 difference images and their corresponding weight differences are obtained, and divided into training sets and test sets in a ratio of 8:2. The training set and test set are used to train and test the overall feed intake estimation model in the overall feed intake estimation module. Through continuous feature extraction and model training, the model gradually fits the true label value, thereby obtaining a model that can estimate the overall feed intake within the field of view. In actual use (actual estimation), the difference image of this module can be directly sent to the feed intake estimation module for estimation.
[0060] Overall feed intake estimation module: The overall feed intake estimation model is used to estimate the overall feed intake of the acquired difference images.
[0061] The overall feed intake estimation model (DMS-RGB-Net) proposed in this application mainly includes image feature extraction and feed intake estimation. Figure 7 As shown. In the feature extraction part, the difference image is processed through a series of multi-scale fusion and self-attention mechanisms to reduce the influence of noise in the feature extraction process. And while extracting subtle features, the global connection and the relationship between each part of the subtle features are also taken into account, which also avoids the influence of the color change of the feed surface caused by light to a greater extent. In the feed intake estimation part, the extracted feature vector is linearly transformed through the fully connected layer, and nonlinearity is introduced through the activation function to finally estimate the feed intake.
[0062] Feature extraction part:
[0063] First, the difference image obtained after processing by the data processing module is recorded as the original difference image, and the original difference image is sent to the feature extraction part. Then, the processing of the feature map is divided into two paths, one of which performs a maximum pooling operation on the original difference image; the other path sends the original difference image to the backbone network of the feature extraction part for deeper feature extraction.
[0064] The backbone network of the feature extraction part is based on the RseNet network architecture. Each Residual submodule consists of several BottleneckBlocks, and each BottleneckBlock includes three convolutional layers, namely 1×1 convolution, 3×3 convolution, and 1×1 convolution. In this embodiment, the entire feature extraction backbone network is composed of 4 layers of Residual submodules, and the specific structure is as follows: the first layer contains 3 BottleneckBlocks; the second layer contains 4 BottleneckBlocks; the third layer contains 23 BottleneckBlocks; the fourth layer contains 3 BottleneckBlocks. Each BottleneckBlock contains a jump connection, which allows the gradient to be directly passed to the previous layer, thereby effectively alleviating the common gradient vanishing and gradient exploding problems during deep network training.
[0065] At the same time, it is also considered that when the shooting angle range is expanded, the feeding area under the overall angle range is also expanding. And according to the behavior of a single cow when eating, the feeding distribution law of feed is funnel-shaped. There will be a pattern of large feeding range in the middle and small feeding range around. Through downsampling and corresponding residual block processing, combined with the splicing and convolution operations of feature maps, the network can effectively fuse multi-scale feature information. This structure not only captures local details, but also captures global structures. This rich feature helps the model make more accurate predictions in complex tasks.
[0066] More specifically, the difference image goes through the detailed steps of the two paths, such as Figure 8 As shown:
[0067] First, the input difference image I undergoes the first 2×2 MaxPooling to obtain the feature map F1, as shown in formula (1). I then undergoes 1×1 convolution and the first layer of Residual submodule to obtain a 360×180×64 feature map F1′. F1′ is obtained by downsampling to obtain F1' downsample Then concatenate it with F1 to obtain the feature map F1″, as shown in formula (2).
[0068]
[0069] After F1 is subjected to a second 2×2 Max Pooling, the feature map F2 is obtained, as shown in formula (3). After F1″ is subjected to a 1×1 convolution and the second layer of Residual submodule, a 180×60×128 feature map F2′ is obtained. F2′ is downsampled to obtain F2' downsample Then concatenate it with F2 to obtain the feature map F2″, as shown in formula (4).
[0070]
[0071] Then F2 is subjected to the third 2×2 MaxPooling to obtain the feature map F3, as shown in formula (5). F2″ is subjected to the 1×1 convolution and the third layer of Residual submodule to obtain the 90×30×256 feature map F3′. F3′ is obtained by downsampling operation to obtain F3' downsample Then concatenate it with F3 to obtain the feature map F3″, as shown in formula (6).
[0072]
[0073] Finally, F3″ is subjected to a 1×1 convolution and the fourth layer of Residual submodule to obtain a 45×15×512 feature map F4′.
[0074] After multi-scale feature fusion, the relationship between the subtle features obtained by the model is relatively complex. In order to allow the features at each position to interact with the features at other positions, this application introduces a self-attention mechanism, thereby enhancing the robustness of the model to input changes.
[0075] The specific process is as follows: In order to improve numerical stability, the feature map F4′ is first normalized, and then Query, Key, and Value are calculated; then Query and Key are matrix multiplied to obtain the corresponding energy map A1, as shown in formula (7); then the softmax function is passed to obtain the attention map A2, as shown in formula (8); then Value is multiplied by A2 to obtain the weighted feature map A3, as shown in formula (9); finally, A3 is connected to the feature map F4′ residual and output.
[0076] A1=Query×Key T (7)
[0077] A2=softmax(A1) (8)
[0078] A3=A2×Value (9)
[0079] Feed intake estimation part: In the feed intake estimation part, the feature map obtained after the self-attention mechanism is passed through a multi-scale pooling layer. Specifically, three adaptive average pooling layers of different scales are used:
[0080] (1) Global average pooling, capturing the global information of the entire feature map;
[0081] (2) Medium-scale pooling to capture information in local areas;
[0082] (3) Small-scale pooling to capture finer-grained local information.
[0083] The specific process is as follows: After being processed by the self-attention mechanism, the feature map of each channel is simultaneously subjected to three different scale pooling operations, namely global average pooling, medium scale pooling, and small scale pooling. Then, 1x1 output, 2x2 output, and 4x4 output are obtained respectively; finally, the three different scale pooling results are spliced together.
[0084] The concatenated high-dimensional features are mapped to a lower-dimensional space through a fully connected layer. Nonlinearity is introduced through the Swish activation function (SiLU). Finally, another fully connected layer maps the feature vector to a scalar value, which is the estimated value of feed intake.
[0085] When the output of the fully connected layer passes through the Swish activation function (SiLU), nonlinearity is introduced to enhance the expressiveness of the model. The calculation method of Swish is shown in formula (10):
[0086] Swish(x)=x i ·σ(x) (10)
[0087] Where σ(x) is the Sigmoid function, and its calculation method is shown in formula (11):
[0088]
[0089] Among them, x is the feature vector, x i is the element value in the feature vector. The output range of the Sigmoid function is between (0,1). It acts as a "gate" in the Swish function to control the input x i The degree of passing. i When x is positive, the output of the Sigmoid function is close to 1, which means that x i Almost completely passed; when x i When x is negative, the output of the Sigmoid function is close to 0, which means that x i is suppressed, but not completely truncated to 0. This can adaptively adjust its output to better capture complex patterns in the data.
[0090] The loss function used in the training of the overall feed intake estimation model using the training set is as follows:
[0091] In order to update the parameters of the model, this application uses MSE-LOSS as the loss function, which is used to measure the mean square difference between the model prediction value and the true value. The calculation method is shown in formula (12):
[0092]
[0093] Where n is the number of samples, yi is the true label (target value) of the total feed intake of the ith sample, is the model's predicted value for the overall feed intake of the ith sample.
[0094] After obtaining the trained overall feed intake estimation model, in actual use, the difference image obtained by the data processing module can be directly sent to the feed intake estimation module for estimation.
[0095] For the first time, the segmentation model was used to segment the feed area in the image, solving the impact of other interference factors in the large field of view. In the feed intake estimation module, the ResNet architecture, self-attention mechanism and multi-scale fusion technology were innovatively combined to design and implement a deep learning model for estimating the overall feed intake within the camera field of view; at the same time, this application was experimented in a real environment, and the overall process only required RGB images to participate, which significantly reduced costs.
[0096] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A system for estimating the feed intake of dairy cows in feeding areas based on computer vision, characterized in that: include: Data acquisition module: used to obtain RGB images of the feed area before and after feeding dairy cows; Feed area segmentation module: Use the image segmentation model to process the RGB image, identify and extract the mask image of the feed area; Data processing module: color matching is performed on the mask images corresponding to before and after feeding the cows to obtain the RGB image of the feed area corresponding to the mask image; a difference image is obtained based on the RGB image of the feed area in the mask image corresponding to before and after feeding the cows; The overall feed intake estimation module: uses the overall feed intake estimation model to estimate the overall feed intake of the acquired difference image; the overall feed intake estimation model includes a feature extraction subnetwork and a feed intake estimation subnetwork; Feature extraction subnetwork: used to extract features from the difference image obtained after processing by the data processing module; Feed intake estimation subnetwork: The feature map obtained by the feature extraction subnetwork is passed through a multi-scale pooling layer and spliced, passed through a fully connected layer, and then processed by an activation function, and then passed through a fully connected layer to map the feature vector to a scalar value, that is, the estimated value of feed intake.
2. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 1, characterized in that: The image segmentation model image adopts U 2 -Net model.
3. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 2, characterized in that: The U 2 -Net model is an encoder-decoder structure; the encoder includes 4 layers of RSU submodules and 2 layers of RSU4F submodules, respectively recorded as encode_1 to encode_6; the decoder includes 1 layer of RSU4F submodule and 4 layers of RSU submodules, respectively recorded as decode_5 to decode_1; starting from the salient image of the last layer of encode_6, it is fused layer by layer with the salient image of each layer of the decoder, and finally the final salient mask image is generated through 1×1Conv and a Sigmoid function.
4. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 1, characterized in that: The process of color matching the mask images corresponding to the cows before and after feeding includes: For the generated mask image, the RGB image of the corresponding feed area is implanted into the mask image area to obtain the RGB image of the feed area in the mask image.
5. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 1, characterized in that: The processing process of the feature extraction subnetwork includes: The difference image obtained after processing by the data processing module is recorded as the original difference image, and the original difference image is sent to the feature extraction sub-network. The processing is divided into two paths. One path performs a maximum pooling operation on the original difference image; the other path sends the original difference image to the backbone network of the feature extraction sub-network for deeper feature extraction; The backbone network of the feature extraction subnetwork is based on the RseNet network architecture, including 4 layers of Residual submodules. Each Residual submodule consists of several Bottleneck Blocks. Each Bottleneck Block includes three convolutional layers, namely 1×1 convolution, 3×3 convolution, and 1×1 convolution.
6. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 5, characterized in that: The specific structure of the 4-layer Residual sub-module of the backbone network of the feature extraction sub-network is as follows: the first layer contains 3 Bottleneck Blocks; the second layer contains 4 Bottleneck Blocks; the third layer contains 23 Bottleneck Blocks; the fourth layer contains 3 Bottleneck Blocks; each Bottleneck Block contains a skip connection.
7. The computer vision-based system for estimating the feed intake of dairy cows in feeding areas according to claim 5, characterized in that: The original difference image is sent to the feature extraction subnetwork, and the processing is divided into two paths for processing as follows: First, the input difference image I undergoes the first 2×2 Max Pooling to obtain the feature map F1; I then undergoes 1×1 convolution and the first layer of Residual submodule to obtain the feature map F1′; F1′ is downsampled to obtain F1' downsample Then concatenate it with F1 to get the feature map F1″; After F1 is subjected to a second 2×2 Max Pooling, feature map F2 is obtained; after F1″ is subjected to a 1×1 convolution and a second layer of Residual submodule, feature map F2′ is obtained; F2′ is downsampled to obtain F2' downsample Then concatenate it with F2 to get the feature map F2″; Then F2 undergoes a third 2×2 Max Pooling to obtain feature map F3; F2″ undergoes a 1×1 convolution and the third layer of Residual submodule to obtain feature map F3′; F3′ is downsampled to obtain F3' downsample Then concatenate it with F3 to get the feature map F3″; Then, F3″ is subjected to 1×1 convolution and the fourth layer of Residual submodule to obtain the feature map F4′; The feature map F4′ is first normalized, and then the Query, Key, and Value are calculated; then the Query and Key are matrix multiplied to obtain the corresponding energy map A1; then the softmax function is passed to obtain the attention map A2; then the Value is multiplied by A2 to obtain the weighted feature map A3; finally, A3 is connected to the residual of the feature map F4′ and output.
8. The computer vision-based system for estimating the feed intake of dairy cows in feeding areas according to claim 1, characterized in that: In the feed intake estimation subnetwork, the Swish activation function is used when processing the activation function.
9. A computer vision-based system for estimating feed intake in a dairy cow feeding area according to any one of claims 5 to 8, characterized in that: The overall feed intake estimation model is pre-trained. During the training process of the overall feed intake estimation model, different difference images need to be used to train the overall feed intake estimation model; in the process of obtaining different difference images, a permutation and combination method is needed to subtract the same pile of feed images from each other according to the relationship before and after the feed to obtain difference images.
10. The computer vision-based dairy cow feeding area feed intake estimation system according to claim 9, characterized in that: The loss function used in the training process of the overall feed intake estimation model is as follows: Where n is the number of samples, y i is the true label of the total feed intake of the i-th sample, is the model's predicted value for the total feed intake of the ith sample.
Citation Information
Patent Citations
Cervical image processing method and device based on regional proposal
CN108090906A
Feed intake monitoring method and system based on twin network and depth data
CN114358163A
Virtual food box defect generation method and system based on neural network
CN114359269A
Palm print recognition method fusing attention mechanism and residual network
CN116978075A
Low-power-consumption monitoring awakening method, device and equipment based on human shape recognition
CN118982847A
Cited By
Video stream-based beef cattle feed intake intelligent monitoring and analysis method
CN121095200A
Intelligent monitoring and analysis method for beef cattle intake based on video stream
CN121095200B