A method and system for estimating the weight of meat poultry based on 2D images

By building an image classification and semantic segmentation network, combining it with a convolutional neural network, and using a consumer-grade camera to capture three-view images, the problems of time-consuming and low-precision weight measurement of meat poultry were solved, achieving efficient and stable weight estimation results.

CN119380375BActive Publication Date: 2025-10-14CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411483188.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-10-14
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The existing technology for measuring the weight of meat poultry has the problems of being time-consuming, low-precision and affecting animal behavior. Especially for active meat poultry animals, traditional methods are difficult to achieve efficient and stable weight estimation.

Method used

A 2D image-based poultry weight estimation method is adopted. By constructing an image classification network and a semantic segmentation network, combined with a convolutional neural network, and using a consumer-grade camera to capture three-view images, it automatically identifies images that meet the shooting conditions and estimates weight, reducing computing power requirements and environmental adaptability.

Benefits of technology

It achieves efficient and stable weight estimation of meat poultry in their natural eating state, reduces power consumption and computing power requirements of the recognition process, improves prediction accuracy and adaptability, and is suitable for weight estimation in meat poultry breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380375B_ABST
    Figure CN119380375B_ABST
Patent Text Reader

Abstract

The application discloses to the technical field of computer vision, and particularly relates to a kind of meat poultry weight estimation method and system based on 2D image, the method includes: in response to sensor perception meat poultry enters shooting place, by shooting device collection meat poultry three view image set;Image classification network is constructed, three view image set is filtered, and three view image satisfying shootable condition is obtained;Image segmentation network is constructed, and three view image satisfying shootable condition is handled with semantic segmentation, and three view image after segmentation is obtained;Based on three view image after segmentation, input convolutional neural network estimates meat poultry weight.The system includes: meat poultry image acquisition device and operation module.2D image is collected using consumer camera, and efficient, stable meat poultry weight estimation is realized, to assist breeding work, and it has great significance to meat poultry breeding industry.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer vision, and particularly relates to a meat poultry weight estimation method and system based on a 2D image. BACKGROUND

[0002] Growth traits are one of the important breeding directions of meat poultry, and the accuracy of body weight measurement directly affects the accuracy of offspring selection. In the prior art, the duck body size and weight are measured by traditional manual methods using a tape measure, a yardstick, an electronic scale and other instruments, the entire measurement process is time-consuming, labor-intensive, and the measured duck body size index data is easily affected by the subjectivity of the measurer. Meanwhile, the sensors in the precision instruments such as the weighing module of the intelligent measurement device are greatly affected by animal behavior in the naked state, resulting in inaccurate measurement results. Manual measurement also causes serious stress reactions to the ducks, affecting the normal behavior of the ducks, and does not meet the requirements of animal welfare. With the rapid development of machine vision technology, the contactless weight measurement method based on machine vision for pigs, cattle, sheep and other animals has changed the traditional contact weight measurement method, and can better solve the problems in the traditional measurement of animal weight data, such as the inability of measurement personnel to continuously and non-contact measure animal weight data. Related machine vision technologies mainly include:

[0003] Image classification is an important task in computer vision, which involves dividing input images into different pre-defined categories. This process involves analyzing and extracting features from the content of the image and comparing them with a pre-trained model to identify the category to which the image belongs. Image classification has a wide range of applications in many fields, such as object detection, face recognition, medical image analysis, etc.

[0004] Semantic segmentation is a part of neural network applications in the field of computer vision, which involves taking some simple filtered and processed raw data as input and highlighting the part of interest. Image semantic segmentation is to take an image as input and highlight the region of interest.

[0005] Convolutional neural networks are a type of deep neural network with convolutional structure, which can reduce the amount of memory occupied by deep networks. Its three key operations are: first, local receptive field, second, weight sharing, and third, pooling layer, which effectively reduces the number of network parameters and alleviates the overfitting problem of the model. It is commonly used in the field of computer vision.

[0006] Existing technologies present significant difficulties in estimating the weight of broiler poultry. Considering that broiler poultry are relatively active and have a relatively light weight, the use of scales for measurement is highly unstable. Existing technologies that use computer vision technology to estimate animal weight rely on high-definition 2D or 3D cameras to capture animal images, extract key parameters such as the animal's body shape through image processing techniques, and then use machine learning or deep learning regression methods to predict weight based on the extracted features. These methods are primarily used for species such as pigs and cattle and are not suitable for the naturally active broiler poultry. For example, broiler poultry constantly moves their bodies while their wings move, causing their wings to obscure the body, significantly increasing the difficulty of capturing images of the poultry. Therefore, there is an urgent need for a 2D image-based broiler poultry weight estimation method and system. Using consumer-grade cameras to capture 2D images allows for efficient and stable broiler poultry weight estimation to assist in breeding work, which is of great significance to the broiler poultry farming industry. Summary of the Invention

[0007] The present invention aims to provide a method for estimating the weight of meat poultry based on 2D images, which is characterized by comprising the following steps:

[0008] Step S1: In response to a sensor sensing that a poultry enters a shooting location, a shooting device collects a set of three-view images of the poultry;

[0009] Step S2: construct an image classification network to screen the three-view image set and obtain three-view images that meet the conditions for being photographed;

[0010] Step S3: construct an image segmentation network, perform semantic segmentation processing on the three-view images that meet the shooting conditions, and obtain the segmented three-view images;

[0011] Step S4: Based on the segmented three-view image, input the convolutional neural network to estimate the weight of the poultry.

[0012] The three-view images in step S1 include: a front view, a top view, and a side view; and the shooting equipment includes: a front view shooting equipment, a top view shooting equipment, and a side view shooting equipment.

[0013] The sensor in step S1 is a radio frequency identification sensor, and the poultry is equipped with a radio frequency identification tag to sense the poultry entering and leaving the shooting location and sense the poultry's identity information.

[0014] The image classification network in step S2 is a DeformAttn-ShuffleNetV2 network, which constructs a deformable attention module based on ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution; the input of the DeformAttn-ShuffleNetV2 network is the side view in the three-view image; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the meat poultry status; the classification result includes: zhanli and qita; when the classification result is zhanli, it is determined that the three-view image corresponding to the input side view meets the shooting condition.

[0015] The network structure of the deformable attention module is as follows: the input feature channel is divided into a first branch and a second branch; the first branch includes: 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, and linear rectification function in sequence; the second branch includes: 1×1 convolution, batch normalization, linear rectification function, 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, linear rectification function, 1×1 convolution, batch normalization, linear rectification function, DA convolution, batch normalization, 1×1 convolution, batch normalization, and linear rectification function; the first branch performs channel shuffling after the second branch is merged; the DA convolution includes: depthwise separable convolution with input and output channels both c, batch normalization, linear rectification function, deformable convolution with input and output channels both c, batch normalization, linear rectification function, the first fully connected layer, the second fully connected layer, and the Softmax function.

[0016] The conditions for shooting in step S2 are: the poultry enters the cage and puts its head through the partition to start eating; the poultry is in a standing position; and the wings of the poultry are in a tightened state.

[0017] The image segmentation network in step S3 is a Segment Anything segmentation model.

[0018] The convolutional neural network in step S4 receives a composite image composed of the segmented three-view images and outputs an estimated weight of meat poultry; the composite image passes through the first convolution layer, the first downsampling layer, the second convolution layer, the second downsampling layer, the first fully connected layer, the second fully connected layer and the output layer in sequence to obtain the estimated weight of meat poultry; the kernel size of the first convolution layer is 11×11, the convolution step is 1, and there is no padding; the kernel size of the second convolution layer is 12×12, the convolution step is 1, and there is no padding; the first downsampling layer includes: a first ReLu activation function and an initial pooling layer, the initial pooling layer has a 6×6 pooling window; the second downsampling layer includes: a second ReLu activation function and a second pooling layer, the second pooling layer has a 5×5 pooling window; the first fully connected layer contains 80 nodes; the second fully connected layer contains 16 nodes; the output layer has no activation function.

[0019] The parameter adjustment process of the convolutional neural network is as follows:

[0020] First, adjust the learning rate from large to small, from 0.1 to 0.0001;

[0021] Secondly, adjust the batch size according to the size of the dataset and the computing resources;

[0022] Adjust the optimizer again, the optimizer is SGD or Adam;

[0023] Finally, adjust the regularization parameters, including: L1 regularization, L2 regularization and Dropout.

[0024] Another object of the present invention is to disclose a meat poultry weight estimation system according to the meat poultry weight estimation method based on 2D images of the present invention, characterized in that it comprises: a meat poultry image acquisition device and a calculation module;

[0025] The meat and poultry image acquisition device includes: a shell, a front view camera, a top view camera and a side view camera; a partition is provided in the shell, which divides the interior of the shell into a feed storage area and a meat and poultry standing area; a feeding port and a feeding area are provided in the feed storage area; the meat and poultry standing area can accommodate a single meat and poultry standing; a feeding hole is provided on the partition, which restricts the meat and poultry to only pass its head through the feeding hole to complete feeding; the front view camera is located on the inner side of the shell side wall near the feeding port along the axial direction of the feeding hole, and is used to capture meat and poultry images facing the head of the meat and poultry; the top view camera is located on the inner side of the shell top wall, and is used to capture meat and poultry images from the top; the side view camera is located on the inner side of the shell side wall along the axial direction perpendicular to the feeding hole, and is used to capture meat and poultry images facing the torso of the meat and poultry;

[0026] The operation module includes: an image classification network module, an image segmentation network module and a convolutional neural network module; the image classification network module is used to screen the three-view image set, input the side view in the three-view image into the DeformAttn-ShuffleNetV2 network, and output the classification result of the meat poultry state to obtain the screened three-view image; the image segmentation network module is used to perform semantic segmentation processing on the three-view images that meet the shooting conditions to obtain the segmented three-view images; the convolutional neural network module is used to use the segmented three-view images to estimate the weight of meat poultry.

[0027] The beneficial effects of the present invention are:

[0028] The present invention discloses a method for estimating the weight of meat poultry based on 2D images. Based on existing intelligent weighing equipment or intelligent breeding poultry equipment, three consumer-grade infrared cameras are installed in front, on the side and on the top of the breeding cage respectively to collect three-view images of a single meat poultry in a natural eating state from three perspectives: front view, side view and top view. On the basis of reusing existing breeding equipment, it is possible to collect complete body images of meat poultry at different ages throughout the growth cycle. By integrating image information from three perspectives, it is possible to better adapt to the active characteristics of meat poultry. Taking advantage of the feeding nature of meat poultry, the meat poultry is likely to actively eat after entering the shooting area. After sensing that the meat poultry has entered the collection area, multiple sets of three-view images are collected within 1-3 seconds, and the camera is automatically set to standby mode, which greatly reduces the power consumption of the camera.

[0029] Existing computer vision-based weight estimation requires detailed identification of subject features, such as chest width and body length, which requires high computing power and power consumption. This drawback is particularly pronounced for broiler animals, such as ducks and chickens, which are naturally active and have wings that significantly impact body contour recognition, and are found in relatively large numbers. The present invention, through extensive observation and experimentation, has discovered that broiler animals are likely to feed immediately after entering the collection area. Furthermore, given the positioning of the partitions and feeding bowls, broilers are likely to briefly maintain a standing posture with their wings folded while feeding, extending their necks to pass through the openings in the partitions. In this state, multiple sets of three-view images are captured within 1-3 seconds of sensing the broiler entering the collection area. This eliminates the need for detailed identification of specific body parts, especially for side views, and avoids differences in the degree to which different broiler individuals extend their necks while feeding. By constructing an image classification network and defining conditions for capturing three-view images, the present invention automatically identifies three-view images that meet these conditions, effectively reducing the computing power required during the recognition process.

[0030] The image classification network is a DeformAttn-ShuffleNetV2 network. Based on ShuffleNetV2, a deformable attention module (DeformAttn Blocks) is built to introduce the SKnet attention mechanism and deformable convolution. This deformable attention module allows for better dynamic adjustment to various environmental changes within the duck cage, such as lighting conditions and random clutter. As a result, the DeformAttn-ShuffleNetV2 network disclosed in this invention achieves superior performance in terms of accuracy, number of parameters, and number of FLOPs.

[0031] Based on the DeformAttn-ShuffleNetV2 network's accurate identification of three-view images that meet the requirements for capture, the Segment Anything segmentation model is used as the image segmentation network to achieve high segmentation accuracy in complex environments. This approach eliminates the need for data annotation, requiring only the pre-defined pixel values ​​of the segmented area and a few label points. Because the feeding area of ​​meat poultry is relatively fixed, selecting a few fixed label points for each of the three-view images yields highly accurate segmented three-view images.

[0032] Based on the segmented three-view images, a convolutional neural network estimates the weight of meat poultry. By adjusting the parameters of the convolutional neural network, a more accurate estimation result is obtained. Experimental verification shows that the predicted value obtained by applying the meat poultry weight estimation method based on 2D images disclosed in the present invention has a high degree of fit between the actual value, and can achieve better prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of a flow chart of a method for estimating the weight of meat poultry based on 2D images according to the present invention;

[0034] Figure 2 The three-view segmentation images in the embodiment of the present invention, where (a) is the front view - RGB image; (b) is the front view - segmentation image; (c) is the side view - RGB image; (d) is the side view - segmentation image; (e) is the top view - RGB image; (f) is the top view - segmentation image;

[0035] Figure 3 Schematic diagram of the camera installation position in an embodiment of the present invention, wherein (a) is a front view; (b) is a side view; (c) is a top view;

[0036] Figure 4 A schematic diagram of a group of RGB images of an object to be estimated in a natural eating state according to an embodiment of the present invention;

[0037] Figure 5A schematic diagram of the structure of the DeformAttn-ShuffleNetV2 network provided by the present invention;

[0038] Figure 6 for Figure 5 A magnified image of the network structure of the deformable attention module;

[0039] Figure 7 for Figure 6 A magnified image of the network structure of the DA convolution part;

[0040] Figure 8 This is a schematic diagram of the structure of the CNN network provided by the present invention. DETAILED DESCRIPTION

[0041] The present invention provides a method and system for estimating the weight of meat poultry based on 2D images, which will be further described in detail below with reference to the accompanying drawings.

[0042] like Figure 1 The embodiment of the present invention shown in FIG. 1 discloses a method for estimating the weight of meat poultry based on 2D images, comprising the following steps:

[0043] Step S1: In response to a sensor sensing that a poultry enters a shooting location, a shooting device collects a set of three-view images of the poultry;

[0044] Step S2: construct an image classification network to screen the three-view image set and obtain three-view images that meet the conditions for being photographed;

[0045] Step S3: construct an image segmentation network, perform semantic segmentation processing on the three-view images that meet the shooting conditions, and obtain the segmented three-view images;

[0046] Step S4: Based on the segmented three-view image, input the convolutional neural network to estimate the weight of the poultry.

[0047] In this embodiment, the present invention discloses a 2D image-based poultry weight estimation method suitable for use in scenarios involving weight estimation of subjects, such as poultry breeding. In this embodiment, weight estimation of ducks is used as an example. The following describes the specific implementation of each step in conjunction with a specific embodiment.

[0048] Step S1: In response to a sensor sensing that a poultry enters a shooting location, a shooting device collects a set of three-view images of the poultry;

[0049] The three-view images in step S1 include: a front view, a top view, and a side view; and the shooting equipment includes: a front view shooting equipment, a top view shooting equipment, and a side view shooting equipment.

[0050] In this embodiment, the shooting equipment is three consumer-grade infrared cameras, which are respectively installed in front, on the side and above the weight estimation device to capture three-view images of a single meat poultry in its natural eating state from three perspectives: front view, top view and side view. Figure 4 The shooting time is from the time the poultry stands up and enters the weight estimation device to the time it leaves the device after eating. Multiple sets of three-view images of the poultry in a natural eating state are obtained to form a three-view image set of the poultry.

[0051] Those skilled in the art should know that in response to the sensor sensing that the poultry meets the conditions for filming, the camera is triggered to start filming. The filming conditions are:

[0052] The poultry enters the cage and puts its head through the partition to start eating; the poultry is in a standing position; and its wings are in a tightened state;

[0053] The sensors include but are not limited to infrared sensors, optical sensors, bioelectric sensors, and weight scales, etc., which are not specifically limited in this embodiment.

[0054] According to the inventors' extensive experimental observations, in most cases, driven by their natural instincts, poultry will immediately rush to eat when entering a cage, prompting them to stretch their necks, pass their heads through the cage partitions, maintain a standing posture, and keep their wings folded. Therefore, in most cases, when the sensor senses that the poultry has entered the cage within 1 to 3 seconds, it is likely that the poultry is in the photographable condition. At this time, the shooting timing is triggered and the camera is called to capture 3 to 5 sets of three-view images. In this embodiment, even if the poultry does not enter the photographable condition within 1 to 3 seconds, unqualified images can be manually excluded during the model training phase. In actual production environments, unqualified images can be excluded by using an image classification network in subsequent steps. Therefore, in this embodiment, the timing for capturing multiple sets of three-view images is within 1-3 seconds of the poultry entering the cage via the sensor. Therefore, the method for sensing the shooting timing used in this embodiment has the advantages of simple structure, rapid judgment, and low cost.

[0055] The sensor in step S1 is a radio frequency identification sensor, and the poultry is equipped with a radio frequency identification tag to sense the poultry entering and leaving the shooting location and sense the poultry's identity information.

[0056] In a preferred embodiment, the poultry is equipped with an electronic tag, and the sensor is an electronic tag sensor. The electronic tag sensor, in conjunction with the electronic tag, not only detects when the poultry enters a designated area within the cage but also obtains the poultry's identity information, facilitating weight tracking throughout the poultry's growth cycle. The electronic tag includes, but is not limited to, RFID tags, NFC tags, Bluetooth tags, and ZigBee tags, and is not specifically limited in this embodiment.

[0057] In this embodiment, multiple groups of three-view images of meat and poultry in a natural eating state are captured, and then the multiple groups of images in a natural eating state obtained are selected to obtain multiple groups of images in a natural eating state that meet the shooting conditions, and the camera is controlled to enter a standby state.

[0058] In an optional embodiment, in response to the sensor sensing that the poultry leaves the photographing location, the photographing device enters a standby state.

[0059] In this embodiment, the camera is installed at a position such as Figure 3 As shown, (a) is the front view; (b) is the side view; (c) is the top view; by installing consumer-grade 2D cameras in front, on the side and above the weight estimation device, the images of meat ducks eating in their natural state were captured manually. A total of 275 ducks were collected, and the three-view Figure 1 A total of about 15,000 images are used as the data set for subsequent training models.

[0060] In an optional embodiment, multiple groups of three-view images of meat poultry in a natural eating state are captured, and then the multiple groups of images in a natural eating state are selected to obtain multiple groups of images in a natural eating state that meet the conditions for being photographed. When the sensor senses that the meat poultry has left the breeding cage, a weighing scale is used to measure the actual weight of the poultry, and the actual weight value and the multiple groups of images in a natural eating state that meet the conditions for being photographed are added to the data set at the same time for model training and testing.

[0061] Those skilled in the art will appreciate that, given the significant weight fluctuations of broiler poultry throughout their growth, the camera should be positioned to ensure that it can capture a complete image of the body of the older broiler poultry. These positions will be adjusted based on factors such as the camera's performance parameters and the size of the breeding cage, and are not specifically limited in this embodiment.

[0062] Step S2: construct an image classification network to screen the three-view image set and obtain three-view images that meet the conditions for being photographed;

[0063] The image classification network in step S2 is a DeformAttn-ShuffleNetV2 network, and the DeformAttn-ShuffleNetV2 network builds a deformable attention module on the basis of ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution; the input of the DeformAttn-ShuffleNetV2 network is the side view in the three-view image; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the meat poultry state; the classification result includes: zhanli and qita; when the classification result is zhanli, it is determined that the three-view image corresponding to the input side view meets the shooting condition;

[0064] like Figure 6 As shown, the network structure of the deformable attention module is as follows: the input feature channel is divided into a first branch and a second branch; the first branch includes: 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, linear rectification function; the second branch includes: 1×1 convolution, batch normalization, linear rectification function, 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, linear rectification function, 1×1 convolution, batch normalization, linear rectification function, DA convolution, batch normalization, 1×1 convolution, batch normalization, linear rectification function; the first branch performs channel shuffling after the second branch is merged. Figure 7 As shown, the DA convolution includes: depth-wise separable convolution with c input and output channels, batch normalization, linear rectification function, deformable convolution with c input and output channels, batch normalization, linear rectification function, the first fully connected layer, the second fully connected layer and the Softmax function.

[0065] In this embodiment, the second depth-wise separable convolution in the DA convolution is replaced by a deformable convolution to improve the adaptability to geometric deformation. The deformable convolution can dynamically adjust the position of the convolution kernel according to the content of the input feature map, so as to better capture the changes in the geometric shape and posture of the object. In contrast, the traditional depth-wise separable convolution has a fixed sampling position and cannot handle spatial deformation well; the feature expression ability is enhanced, and the deformable convolution can adaptively adjust the sampling point position by introducing additional offset learning to better capture the detailed information and complex background of the target object; the detection and segmentation performance of the model are improved, and the deformable convolution can enhance the model's ability to detect edges and shapes, thereby improving the overall detection and segmentation performance; at the same time, the advantages of multi-scale features and spatial deformation can be combined, so that the model can not only better utilize multi-scale information, but also better adapt to the deformation characteristics of the target object.

[0066] The conditions for shooting in step S2 are: the poultry enters the cage and puts its head through the partition to start eating; the poultry is in a standing position; and the wings of the poultry are in a tightened state.

[0067] Considering that meat poultry have different feeding postures in their natural state, such as fully extended legs, half-squatting legs, and lying down, etc. By observing the changes in posture of meat poultry during feeding, it was found that the length of the neck extension of meat poultry when eating was very different compared to the natural standing posture; when eating naturally, the naturally spread wings of meat poultry cover about half of the body outline; in addition, it was found that meat poultry will most likely change from a standing posture to a prone posture after eating for a few seconds. Therefore, the posture of meat poultry has a significant impact on the accuracy of the subsequent weight estimation model. Image classification technology is used to identify and extract the posture images most suitable for weight estimation, and an image classification network model is introduced to automatically screen images to select meat poultry images that meet the aforementioned photographic conditions.

[0068] In this embodiment, multiple sets of three-view images of animals in natural feeding positions are selected and classified to obtain a set of three-view images that meet the capture conditions. The selected sets of images of animals in natural feeding positions are classified into the "zhanli" category, based on the side view. Images of animals in natural feeding positions with legs fully upright, wings fully folded, and head lowered are classified into the "qita" category. These other feeding positions include, but are not limited to, lying down and half-squatting.

[0069] In this embodiment, the DeformAttn-ShuffleNetV2 network performs image classification based on the side view in the three-view image as input, and divides the side view into two categories: standing posture and other postures, where other postures include half-squatting and lying down to eat. The SKnet attention mechanism is introduced based on the ShuffleNetV2 network, named SK-ShuffleNetV2, and deformable convolution is further introduced based on the SK-ShuffleNetV2 network, named DeformAttn-ShuffleNetV2. As shown in Table 1, the DeformAttn-ShuffleNetV2 network achieves good performance in terms of accuracy, number of parameters, and FLOPs.

[0070] Table 1 Image classification results

[0071] Model Top-1(%) Params(M) FLOPs(G) ShuffleNetV2 97.20 1.26 1.45 SK-ShuffleNetV2 97.66 1.6 1.46 DeformAttn-ShuffleNetV2 98.13 6.15 6.87

[0072] ShuffleNetV2 is a novel neural network architecture in the prior art, which aims to effectively reduce the parameter and computational complexity by implementing channel-wise group convolution and channel shuffle operation, thereby improving the computational efficiency. The model introduces a channel-wise group convolution method, which divides the input channels into multiple groups and applies convolution operation to each group independently. Subsequently, the results of these operations are merged in the channel dimension, which can effectively improve the network capacity and information flow, thereby optimizing the performance. By effectively reducing the number of parameters and computational load, ShuffleNetV2 achieves high classification accuracy while minimizing resource requirements. To further improve performance, ShuffleNetV2 uses repeated modules. Each module consists of three operations: channel-wise group convolution, channel mixing, and again channel-wise group convolution. By stacking multiple modules, the overall depth and width of the network are increased, thereby improving the model's expressive ability.

[0073] In this embodiment, in order to improve the accuracy of detecting broiler image in complex environment, the picture data to the best posture is selected, and two technologies are introduced on the basis of ShuffleNetV2: SKnet attention mechanism (Selective Kernel Networks) and deformable convolution (Deformable Convolution).

[0074] The SKnet attention mechanism enables each neuron to adaptively adjust the size of its receptive field (i.e., convolution kernel) according to the different scales of input information, effectively capturing multi-scale features in complex image space without consuming excessive computational resources like traditional CNNs. In addition, the SKnet attention mechanism can integrate deep features, thereby improving the interpretability and understandability of the captured features.

[0075] Deformable convolution can perform non-uniform and adaptive sampling and perception on input features. Deformable convolution learns offsets to adaptively adjust the size and shape of the receptive field according to the context information, thereby more effectively capturing the spatial structure and details of the target. By adding a learnable offset to deformable convolution, the adaptability of traditional convolution operations is enhanced, thereby better adjusting the spatial relationship and deformation between features and improving the expressive ability of the model.

[0076] The structure of the DeformAttn-ShuffleNetV2 network is as shown in Figure 5 .

[0077] The input of the DeformAttn-ShuffleNetV2 network is a side view in a three-view image of a meat bird, and the output is a classification result of the meat bird state; on the basis of the standard ShuffleNetV2 network structure, Stage2, Stage3 and Stage4 are composed of deformable attention modules, realizing the fusion of SKnet attention mechanism (Selective Kernel Networks) and deformable convolution (Deformable Convolution); the network structure of the deformable attention module part is as shown in Figure 6

[0078] In this embodiment, for the 13320 RGB three-view images of 275 meat birds obtained, each meat bird corresponds to multiple groups of natural feeding state images, the multiple groups of natural feeding state images corresponding to each meat bird are screened, and after the screened multiple groups of natural feeding state images are classified, 423 side view RGB images of the "zhanli" category are classified, 652 side view RGB images of the "qita" category are classified, and 1075 sample natural feeding state images are obtained.

[0079] In this embodiment, after using 275 meat duck side view images, the images are divided into standing posture and other postures, and an image classification model DeformAttn-ShuffleNetV2 is trained.

[0080] The model classification results and related data are as follows, wherein ShuffleNetV2 is the basic model, SK-ShuffleNetV2 and DeformAttn-ShuffleNetV2 are improved models, and finally the DeformAttn-ShuffleNetV2 model is selected as the actual use:

[0081] Step S3: constructing an image segmentation network, performing semantic segmentation processing on the three-view images that meet the shootable conditions, and obtaining segmented three-view images;

[0082] The image segmentation network in the step S3 is a Segment Anything segmentation model.

[0083] In this embodiment, the meat bird feeding area is relatively fixed, that is, feeding into the feed port of the body weight estimation device. The Segment Anything segmentation model can obtain a high segmentation precision image in a complex environment, and does not need to be labeled, only needs to pre-set the region to be segmented and a few label point pixel values. Since the meat bird feeding area is relatively fixed, a few fixed label points are selected for the three-view images, and a high segmentation precision image can be obtained.

[0084] ​In this embodiment, based on the Segment Anything segmentation model, there is no need to perform image annotation, which greatly reduces the workload while ensuring the segmentation accuracy. By setting different segmentation rectangular areas, labels of target objects and backgrounds for the three-view images respectively, the final segmented image is obtained. The mIoU indicators obtained for the front view, top view and side view in the dataset after image classification are shown in Table 3, which are: 0.9697 for the front view, 0.9871 for the top view, and 0.9769 for the side view. According to the obtained multiple groups of sample images of natural eating states, after obtaining the image classification results, they are sequentially input into the SegmentAnything segmentation model, and rectangular areas and labels are set for the three view images respectively, to obtain the recognition results output by the initial image segmentation model as shown in Table 3. Figure 2 shown.

[0085] Table 3 Segment Anything segmentation model mIoU results statistics

[0086]

[0087] Step S4: Based on the segmented three-view image, input the convolutional neural network to estimate the weight of the poultry.

[0088] The convolutional neural network in step S4 receives a composite image composed of the segmented three-view images and outputs an estimated weight of meat poultry; the composite image passes through the first convolution layer, the first downsampling layer, the second convolution layer, the second downsampling layer, the first fully connected layer, the second fully connected layer and the output layer in sequence to obtain the estimated weight of meat poultry; the kernel size of the first convolution layer is 11×11, the convolution step is 1, and there is no padding; the kernel size of the second convolution layer is 12×12, the convolution step is 1, and there is no padding; the first downsampling layer includes: a first ReLu activation function and an initial pooling layer, the initial pooling layer has a 6×6 pooling window; the second downsampling layer includes: a second ReLu activation function and a second pooling layer, the second pooling layer has a 5×5 pooling window; the first fully connected layer contains 80 nodes; the second fully connected layer contains 16 nodes; the output layer has no activation function.

[0089] In this example, a convolutional neural network was used to estimate duck weight. The segmented three-view images obtained by the image segmentation model were used as input to the convolutional neural network model using channel stitching technology. The weight values ​​corresponding to the individuals were used as the true values ​​to train a duck weight estimation model. For specific data, see Table 2.

[0090] The convolutional neural network receives a composite image composed of segmented three-view images and outputs an estimated weight of the poultry. Figure 8As shown, the synthetic image passes through the first convolutional layer, the first downsampling layer, the second convolutional layer, the second downsampling layer, the first fully connected layer, the second fully connected layer and the output layer in sequence to obtain the estimated weight of meat poultry; the kernel size of the first convolutional layer is 11×11, the convolution step is 1, and there is no padding; the kernel size of the second convolutional layer is 12×12, the convolution step is 1, and there is no padding; the first downsampling layer includes: a first ReLu activation function and an initial pooling layer, and the initial pooling layer has a 6×6 pooling window; the second downsampling layer includes: a second ReLu activation function and a second pooling layer, and the second pooling layer has a 5×5 pooling window; the first fully connected layer contains 80 nodes; the second fully connected layer contains 16 nodes; the output layer has no activation function.

[0091] In this embodiment, the process of obtaining the data set required for training the convolutional neural network is as follows:

[0092] Three cameras were installed on the weight estimation equipment in the poultry feeding area. By capturing three-view images of the poultry eating naturally, a comprehensive dataset encompassing all three perspectives of the feeding process was generated. To accurately estimate weight, these three-view images were then run through a Segment Anything segmentation model to obtain segmented images that met the unified input requirements of the weight estimation model. Immediately after capturing the images of the ducks, staff weighed the ducks and recorded the weight as the ground truth. Notably, the model's input is a composite image composed of the segmented three-view images obtained from these three perspectives.

[0093] In this embodiment, a total of 275 ducks were collected, and three Figure 1 A total of about 15,000 images are used as the data set for subsequent training models, with the ratio of training set to test set being 8:2.

[0094] The segmented three-view images in the training set are input into the convolutional neural network to obtain an estimated weight of the meat poultry. The estimated weight of the meat poultry is compared with the actual weight value, and the parameters in the convolutional neural network are continuously changed until the training stop condition is met to obtain a trained convolutional neural network.

[0095] In this embodiment, the specific process of changing the parameters in the convolutional neural network is as follows:

[0096] First, adjust the learning rate from large to small, from 0.1 to 0.0001;

[0097] In an optional embodiment, cosine annealing and adaptive learning rate are used to adjust the learning rate;

[0098] Secondly, adjust the batch size according to the size of the dataset and the computing resources;

[0099] Adjust the optimizer again, which is SGD or Adam;

[0100] Finally, adjust the regularization parameters, including L1 regularization, L2 regularization, and Dropout;

[0101] In this embodiment, when the training data is less, L1 and L2 regularization help to prevent overfitting. Dropout reduces the complexity of the model by randomly discarding a portion of neurons.

[0102] In this embodiment, after obtaining the image segmentation result, the corresponding three-view image is uniformly used as the input of the convolutional neural network, and the live meat duck weight estimation model is trained. The weight estimation result and related data are shown in Table 2:

[0103] Table 2 Weight estimation result

[0104] Metric MAE <![CDATA[R 2 ]]> MRE MAD RMSE CNN 40g 0.9631 1.64% 208.6g 48.8g

[0105] The related numerical values in Table 2 are defined as follows:

[0106] MAE (Mean Absolute Error)

[0107] Definition: MAE is the average of the absolute difference between the predicted value and the actual value. The formula is:

[0108]

[0109] Where y i is the actual value, is the predicted value, and n is the sample size.

[0110] Effect: MAE is used to measure the average absolute difference between the predicted value and the actual value of the model.

[0111] R 2

[0112] Definition: R 2 measures the model's ability to explain the variance of the data. The formula is:

[0113]

[0114] Where, is the average of the actual value.

[0115] Effect: R 2 is used in regression tasks to represent the degree of fitting between the predicted value and the actual value. The closer the R 2 value is to 1, the better the model explains the variance of the data.

[0116] MRE (Mean Relative Error)

[0117] Definition: MRE is the average relative error between the actual value and the predicted value, usually expressed as a percentage. The formula is:

[0118]

[0119] Function: MRE is used to evaluate the relative difference between the predicted value and the actual value. It is suitable for situations where the actual value is large and has obvious changes.

[0120] MAD (Median Absolute Deviation)

[0121] Definition: MAD is the median of the absolute differences between the actual and predicted values, and the formula is:

[0122]

[0123] Role: MAD is used as a robust alternative to MAE in regression tasks because the median is more robust to outliers than the mean.

[0124] RMSE (Root Mean Squared Error)

[0125] Definition: RMSE is the square root of the mean of the squares of the errors between the predicted values ​​and the actual values. The formula is:

[0126]

[0127] Function: RMSE is used in regression tasks to emphasize larger errors because the square term amplifies larger errors.

[0128] The convolutional neural network disclosed in this embodiment differs from existing techniques that rely on 2D or 3D cameras to capture animal images. Instead, it uses image processing techniques to extract key parameters such as animal size. Based on these extracted features, machine learning or deep learning regression methods are used to predict poultry weight. The disclosed method for estimating poultry weight based on 2D images eliminates the need for high-definition 2D or 3D cameras to capture animal images. Instead, it uses consumer-grade cameras to capture 2D images, achieving efficient and stable poultry weight estimation.

[0129] In a specific embodiment, the method for estimating the weight of poultry based on 2D images disclosed in the present invention is applied to automatically estimate the weight of ducks based on the trained convolutional neural network. The specific process is as follows:

[0130] 2D three-view images of several meat ducks are collected to form a three-view image set of meat birds; the three-view image set of meat birds is screened by using an image classification network to obtain screened three-view images meeting preset meat duck posture requirements; the three-view images meeting the shootable conditions are subjected to semantic segmentation processing by using an image segmentation network to obtain segmented three-view images; and the segmented three-view images are uniformly taken as inputs of a convolutional neural network to estimate the body weight of the meat duck.

[0131] In the embodiment, a consumer-grade 2D camera is installed at three positions of front, side and top of the body weight estimation device, images of meat ducks in a natural feeding state are automatically collected by using an image classification model DeformAttn-ShuffleNetV2, three-view images after classification are segmented by using an image segmentation model Segment Anything Model, then the group of images with the best segmentation effect of the duck are taken as inputs of a body weight estimation model, and the body weight of the meat duck is automatically estimated.

[0132] Another embodiment of the application discloses a meat bird body weight estimation system based on 2D images, comprising a meat bird image collection device and an operation module.

[0133] The meat bird image collection device comprises a shell, an orthographic view camera, a top view camera and a side view camera; a partition is arranged in the shell, and the partition divides the interior of the shell into a feed storage area and a meat bird standing area; a feeding area with a feeding port is arranged in the feed storage area; the meat bird standing area can accommodate a single meat bird to stand; a feeding hole is arranged on the partition, and the feeding hole limits the meat bird to only pass the head through the feeding hole to complete feeding; the orthographic view camera is arranged on the inner side of the shell side wall on the side close to the feeding port in the axial direction of the feeding hole, and is used for collecting meat bird images facing the head of the meat bird; the top view camera is arranged on the inner side of the top wall of the shell, and is used for collecting meat bird images from the top; and the side view camera is arranged on the inner side of the shell side wall in the direction perpendicular to the axial direction of the feeding hole, and is used for collecting meat bird images facing the trunk of the meat bird.

[0134] The input of the DeformAttn-ShuffleNetV2 network is a side view in the three-view images; the output of the DeformAttn-ShuffleNetV2 network is a classification result of the meat bird state; and the classification result comprises zhanli and qita; when the classification result is zhanli, it is determined that the three-view image corresponding to the input side view meets the shootable condition.

[0135] The operation module comprises: an image classification network module, an image segmentation network module, and a convolutional neural network module; the image classification network module is used for screening the three-view image set, inputting a side view in the three-view image into a DeformAttn-ShuffleNetV2 network, outputting a classification result of a meat bird state, and obtaining the screened three-view image; the image segmentation network module is used for performing semantic segmentation processing on the three-view image meeting the shootable condition, and obtaining the segmented three-view image; and the convolutional neural network module is used for estimating the meat bird weight by using the segmented three-view image.

[0136] In view of the relatively stable posture of the top view and the front view under the shooting standard, in the embodiment, the image classification network module screens the side view, and the top view and the front view take corresponding results.

[0137] For example, the object to be estimated is a meat bird, three consumer-grade 2D infrared cameras are installed in front, side and top of the weight estimation device, three-view images of the meat bird in a natural feeding state standing on the weight estimation device are captured, a Python program is used to parse an RTSP protocol stream to obtain the three-view images of the meat bird, and the first image of at least one object to be estimated can be obtained. The 2D infrared camera complies with the RTSP protocol.

[0138] In the embodiment, a space rectangular system coordinate is established to accurately obtain the meat bird natural feeding state image, as shown in Figure 2 . Figure 3 is a mounting position diagram of the three cameras in the weight estimation device provided by the application, wherein (a) is a front view; (b) is a side view; (c) is a top view; and Figure 3 the lower left corner of the front view shown in (a) is the coordinate origin, which extends to Figure 3 the right side direction in (a) is the positive direction of the x axis, and the length is 46 cm; which extends to Figure 3 the direction inside the breeding cage in (a) is the positive direction of the y axis, and the length is 58 cm; which extends to Figure 3 the top direction in (a) is the positive direction of the z axis, and the length is 53 cm; the coordinates of the three cameras are measured and obtained using a ruler, and the unit is centimeter. The front view shooting device coordinates are x: 19 cm, y: 0 cm, and z: 43.5 cm. The side view shooting device coordinates are x: 0 cm, y: 41 cm, and z: 27.5 cm. The top view shooting device coordinates are x: 20 cm, y: 45.5 cm, and z: 53 cm.

[0139] In this embodiment, those skilled in the art should know that, considering that the weight of meat poultry changes greatly during the entire growth process, the installation position of the camera should ensure that the camera can capture a complete image of the body of the large-day-old meat poultry. Those skilled in the art will make corresponding adjustments based on the performance parameters of the selected camera, the breeding cage ruler and other factors, and no specific limitation is made in this embodiment. Compared with the existing technology that mostly only collects 2D images from one or two perspectives for weight estimation, the present invention uses the front view, top view and side view to obtain the weight of the meat poultry. Figure 3 We collected images from three directions, taking into full consideration the fact that meat poultry are more active than pigs, cattle, sheep and other animals, and their body shape changes greatly during the breeding cycle. The standing posture and wing state have a great impact on weight estimation. We used the image data collected from three perspectives to train the model and obtain more accurate estimation results. The weight estimation results and related data are shown in Table 2. For example, R 2 It is 0.9631, which proves that the fit between the predicted value and the actual value obtained by applying the meat poultry weight estimation method based on 2D images disclosed in the present invention is high, and a good prediction effect can be obtained.

Claims

1. A method for estimating the weight of meat poultry based on 2D images, characterized in that: The steps include: Step S1: In response to a sensor sensing that a poultry enters a shooting location, a shooting device collects a set of three-view images of the poultry; Step S2: construct an image classification network to screen the three-view image set and obtain three-view images that meet the conditions for being photographed; Step S3: construct an image segmentation network, perform semantic segmentation processing on the three-view images that meet the shooting conditions, and obtain the segmented three-view images; Step S4: Based on the segmented three-view image, input the convolutional neural network to estimate the weight of the meat poultry; The image classification network in step S2 is a DeformAttn-ShuffleNetV2 network, which builds a deformable attention module on the basis of ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution; the input of the DeformAttn-ShuffleNetV2 network is the side view in the three-view image; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the meat poultry state; the classification result includes: zhanli and qita; when the classification result is zhanli, it is determined that the three-view image corresponding to the input side view meets the shooting condition; The network structure of the deformable attention module is as follows: the input feature channel is divided into a first branch and a second branch; the first branch includes: 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, and linear rectification function in sequence; the second branch includes: 1×1 convolution, batch normalization, linear rectification function, 3×3 convolution, batch normalization, 1×1 convolution, batch normalization, linear rectification function, 1×1 convolution, batch normalization, linear rectification function, DA convolution, batch normalization, 1×1 convolution, batch normalization, and linear rectification function; the first branch performs channel shuffling after the second branch is merged; the DA convolution includes: depthwise separable convolution with input and output channels both c, batch normalization, linear rectification function, deformable convolution with input and output channels both c, batch normalization, linear rectification function, the first fully connected layer, the second fully connected layer, and the Softmax function.

2. The method for estimating the weight of poultry based on 2D images according to claim 1, characterized in that: The three-view images in step S1 include: a front view, a top view, and a side view; and the shooting equipment includes: a front view shooting equipment, a top view shooting equipment, and a side view shooting equipment.

3. The method for estimating the weight of poultry based on 2D images according to claim 1, characterized in that: The sensor in step S1 is a radio frequency identification sensor, and the poultry is equipped with a radio frequency identification tag to sense the poultry entering and leaving the shooting location and sense the poultry's identity information.

4. The method for estimating the weight of meat poultry based on 2D images according to claim 1, characterized in that: The conditions for shooting in step S2 are: the poultry enters the cage and puts its head through the partition to start eating; the poultry is in a standing position; and the wings of the poultry are in a tightened state.

5. The method for estimating poultry weight based on 2D images according to claim 1, characterized in that: The image segmentation network in step S3 is a Segment Anything segmentation model.

6. The method for estimating poultry weight based on 2D images according to claim 1, characterized in that: The convolutional neural network in step S4 receives a composite image composed of the segmented three-view images and outputs an estimated weight of meat poultry; the composite image passes through the first convolution layer, the first downsampling layer, the second convolution layer, the second downsampling layer, the first fully connected layer, the second fully connected layer and the output layer in sequence to obtain the estimated weight of meat poultry; the kernel size of the first convolution layer is 11×11, the convolution step is 1, and there is no padding; the kernel size of the second convolution layer is 12×12, the convolution step is 1, and there is no padding; the first downsampling layer includes: a first ReLu activation function and an initial pooling layer, the initial pooling layer has a 6×6 pooling window; the second downsampling layer includes: a second ReLu activation function and a second pooling layer, the second pooling layer has a 5×5 pooling window; the first fully connected layer contains 80 nodes; the second fully connected layer contains 16 nodes; the output layer has no activation function.

7. The method for estimating poultry weight based on 2D images according to claim 6, characterized in that: The parameter adjustment process of the convolutional neural network is as follows: First, adjust the learning rate from large to small, from 0.1 to 0.0001; Secondly, adjust the batch size according to the size of the dataset and the computing resources; Adjust the optimizer again, the optimizer is SGD or Adam; Finally, adjust the regularization parameters, including: L1 regularization, L2 regularization and Dropout.

8. A poultry weight estimation system according to the 2D image-based poultry weight estimation method according to any one of claims 1 to 7, characterized in that: include: Meat and poultry image acquisition device and computing module; The meat and poultry image acquisition device includes: a shell, a front view camera, a top view camera and a side view camera; a partition is provided in the shell, which divides the interior of the shell into a feed storage area and a meat and poultry standing area; a feeding port and a feeding area are provided in the feed storage area; the meat and poultry standing area can accommodate a single meat and poultry standing; a feeding hole is provided on the partition, which restricts the meat and poultry to only pass its head through the feeding hole to complete feeding; the front view camera is located on the inner side of the shell side wall near the feeding port along the axial direction of the feeding hole, and is used to capture meat and poultry images facing the head of the meat and poultry; the top view camera is located on the inner side of the shell top wall, and is used to capture meat and poultry images from the top; the side view camera is located on the inner side of the shell side wall along the axial direction perpendicular to the feeding hole, and is used to capture meat and poultry images facing the torso of the meat and poultry; The operation module includes: an image classification network module, an image segmentation network module and a convolutional neural network module; the image classification network module is used to screen the three-view image set, input the side view in the three-view image into the DeformAttn-ShuffleNetV2 network, and output the classification result of the meat poultry state to obtain the screened three-view image; the image segmentation network module is used to perform semantic segmentation processing on the three-view images that meet the shooting conditions to obtain the segmented three-view images; the convolutional neural network module is used to use the segmented three-view images to estimate the weight of meat poultry.

Citation Information

Patent Citations

  • Sound wave and image fused food volume measurement method and system in smartphone

    CN112150535A

  • Remote sensing scene image classification method based on grouping mixed attention

    CN115546654A