Broiler weight estimation method and system based on two-view image
Through the broiler weight estimation method of two-view images, the use of radio frequency identification sensors and image processing networks, the shortcomings of contact equipment in broiler weight measurement are solved, and efficient and accurate weight estimation is achieved to adapt to the activity characteristics and environmental changes of broiler chickens.
Patent Information
- Application Number
- CN202510405580.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
The existing broiler weight measurement methods mainly rely on contact equipment, which have problems such as time-consuming and labor-intensive operation, affecting animal health, unstable measurement accuracy and poor data consistency. The non-contact measurement technology is not effective in broiler weight estimation, especially the influence of small body size and frequent activities.
The weight estimation method of broiler chickens based on two-view images is adopted, and the top view and side rear image shooting equipment is used, combined with radio frequency identification sensors, DeformAttn-ShuffleNetV2 network and CoRSeaFormer segmentation model is used to estimate the weight of broiler chickens through the RepVGGVan network, reduce computing power requirements and power consumption, and adapt to the activity characteristics of broiler chickens.
It realizes high-precision broiler weight estimation under low computing cost and low memory requirements, reduces interference to animals, improves measurement efficiency and accuracy, adapts to complex environment changes, and reduces the impact on wing occlusion.
Smart Images

Figure CN120279547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method and system for estimating the weight of broiler chickens based on two-view images. Background Art
[0002] With the continuous development of the poultry breeding industry, weight, as an important indicator to measure the growth status of broiler chickens, occupies a crucial position in production management and genetic breeding. Weight data can not only reflect the health status and nutritional level of broiler chickens, but also provide a scientific basis for individual selection in genetic breeding. At the same time, it plays a key role in the production management of farms, helping to determine the optimal slaughter time and avoid economic losses caused by overfeeding. Therefore, how to accurately and quickly obtain the weight information of broiler chickens is of great significance for the efficient management of farms and breeding selection.
[0003] In the prior art, the measurement of broiler chicken weight mainly relies on two types of methods: contact and non-contact. Contact measurement usually uses devices such as electronic scales and measuring rulers, and manual collection of weight or body measurement data is carried out. Although this type of method has a certain measurement accuracy, there are many limitations in practical applications. First of all, manual operation is time-consuming and laborious, and it is difficult to meet the production needs of large-scale farms. Secondly, contact measurement is likely to cause stress reactions in animals, affecting their normal growth and health status. In addition, the accuracy of measurement data is easily affected by the experience and subjective judgment of operators, and the data consistency is poor, thus affecting the subsequent decision-making accuracy and breeding accuracy. Finally, under the interference of animal behavior, the weighing device may have measurement errors, further affecting the reliability of the data.
[0004] In order to address the above problems, in recent years, with the rapid development of machine vision technology, non-contact measurement methods based on 2D images have gradually been applied to the field of livestock weight estimation. Compared with contact measurement, non-contact measurement has multiple advantages, including reducing interference with animals, reducing stress reactions, and meeting the requirements of animal welfare. At the same time, non-contact measurement technology can achieve automated operation, reduce manual intervention, and improve the efficiency and accuracy of data collection. In addition, non-contact measurement can estimate the weights of multiple animals in a short time, greatly improving production efficiency. However, the application of existing non-contact measurement technology in broiler chicken weight estimation still faces many challenges. Since broiler chickens are light in weight and highly active, the wing movements are likely to block key body parts, resulting in insufficient accuracy of traditional 2D images for capturing whole-body features and affecting the accuracy of weight estimation.
[0005] At present, a large number of machine vision-based weight estimation technologies have been applied to large livestock such as pigs, cows, and sheep. These technologies usually rely on 2D or 3D cameras to capture the side or top images of animals, use image processing algorithms to extract the body shape characteristics of animals, and combine regression algorithms for weight estimation, achieving good results in large livestock. However, for broiler chickens, due to their small size and frequent activities, the existing technologies have unsatisfactory effects when applied to broiler chicken weight estimation.
[0006] In the Chinese patent document "A Method and System for Estimating the Weight of Poultry Based on 2D Images" with the application number 202411483188.X, a method and system for estimating the weight of poultry based on three-view images are disclosed, which require the use of three image acquisition devices; and through subsequent research by the inventors, it is found that there is still room for optimization in the image segmentation network and convolutional neural network. Therefore, there is an urgent need to develop a method and system for estimating the weight of broiler chickens based on two-view images, which are optimized for the characteristics of broiler chickens, use fewer consumer-grade camera devices to obtain broiler chicken images, and combine more advanced image processing and deep learning algorithms to achieve efficient and stable weight estimation. Summary of the Invention
[0007] The object of the present invention is to provide a method for estimating the weight of broiler chickens based on two-view images, which is characterized by including the following steps:
[0008] Step S1: In response to the sensor sensing that the broiler chicken enters the shooting location, the shooting device collects a set of two-view images of the broiler chicken;
[0009] Step S2: Construct an image classification network, screen the set of two-view images, and obtain two-view images that meet the shootable conditions;
[0010] Step S3: Construct an image segmentation network, perform semantic segmentation processing on the two-view images that meet the shootable conditions, and obtain the segmented two-view images;
[0011] Step S4: Based on the segmented two-view images, input them into the RepVGGVan network to estimate the weight of the broiler chicken.
[0012] The two-view images in the step S1 include: a top view and a side rear view; the shooting device includes: a top view shooting device and a side rear view shooting device.
[0013] The sensor in the step S1 is a radio frequency identification sensor, and the broiler chicken wears a radio frequency identification tag to realize sensing the entry and exit of the broiler chicken from the shooting location and sensing the identity identification information of the broiler chicken.
[0014] The image classification network in step S2 is the DeformAttn-ShuffleNetV2 network. The DeformAttn-ShuffleNetV2 network constructs a deformable attention module based on ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution. The input of the DeformAttn-ShuffleNetV2 network is the side rear view in the two-view images. The output of the DeformAttn-ShuffleNetV2 network is the classification result of the broiler state. The classification results include: zhanli and qita. When the classification result is zhanli, it is determined that the two-view image corresponding to the input side rear view satisfies the shootable condition.
[0015] The network structure of the deformable attention module is as follows: the input feature channels are divided into a first branch and a second branch. The first branch sequentially includes: a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit. The second branch sequentially includes: a 1×1 convolution, batch normalization, a rectified linear unit, a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit, a 1×1 convolution, batch normalization, a rectified linear unit, a DA convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit. After the first branch and the second branch go through a merging operation, channel shuffle is performed. The DA convolution includes: a depthwise separable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a deformable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a first fully connected layer, a second fully connected layer, and a Softmax function.
[0016] The shootable condition in step S2 is as follows: the broiler enters the cage, puts its head through the partition and starts eating. And the broiler is in a standing posture. And the wings of the broiler are in a tightened state.
[0017] The image segmentation network in step S3 is the CoRSeaFormer segmentation model. The CoRSeaFormer network constructs a compression-enhanced axial attention module under coordinate convolution and a receptive field attention convolution fusion module based on SeaFormer. The compression-enhanced axial attention module under coordinate convolution is used to introduce coordinate convolution, and the receptive field attention convolution fusion module is used to introduce receptive field attention convolution. The input of the CoRSeaFormer network is the two-view images. The output of the CoRSeaFormer network is the segmentation result of the broiler two-view images.
[0018] The network structure of the compression-enhanced axial attention module is as follows: the input feature channels calculate queries, keys, and values through coordinate convolution; the queries, keys, and values are divided into a detail-enhanced kernel branch and an axial attention compression branch; the detail-enhanced kernel branch sequentially includes: concatenating the query, key, and value channels, 3×3 depthwise separable convolution, batch normalization, ReLU function, 1×1 convolution, batch normalization; the axial attention compression branch is divided into a horizontal sub-branch and a vertical sub-branch, the horizontal sub-branch sequentially includes: horizontally compressing the query, key, and value, multi-head attention; the vertical sub-branch sequentially includes: vertically compressing the query, key, and value, multi-head attention; the horizontal sub-branch and the vertical sub-branch perform pointwise addition using tensor broadcasting operation and then perform 1×1 convolution; the detail-enhanced kernel branch and the axial attention compression branch are combined through multiplication calculation; the coordinate convolution includes: coordinate generation, normalization, adding coordinate channels.
[0019] The network structure of the receptive field attention convolution fusion module is as follows: the input feature channels are composed of a third branch and a fourth branch; the third branch sequentially includes: a receptive field attention convolution module; the fourth branch sequentially includes: 3×3 convolution, batch normalization, Sigmoid activation function, upsampling operation; the third branch and the fourth branch are combined after multiplication; the receptive field attention convolution module includes: a regional attention mechanism, a convolution layer, and a residual connection.
[0020] The RepVGGVan network in step S4 receives a synthetic image composed of two segmented view images and outputs an estimated broiler weight; the synthetic image sequentially passes through a first 3×3 convolution layer, a first 1×1 convolution layer, a first downsampling layer, a second 3×3 convolution layer, a second 1×1 convolution layer, a first residual connection layer, a third 3×3 convolution layer, a third 1×1 convolution layer, a second residual connection layer, a fourth 3×3 convolution layer, a fourth 1×1 convolution layer, a third residual connection layer, and a visual attention module to obtain the estimated broiler weight; the convolution stride of the first 3×3 convolution layer is 2, and the padding is 1; the convolution stride of the other 3×3 convolution layers is 1, and the padding is 1; the convolution stride of the 1×1 convolution layer is 1, without padding; the first downsampling layer is a convolution function with a convolution stride of 2; the residual connection layer is a residual operation in Resnet; the visual attention module includes: channel attention, spatial attention, and their weighted fusion.
[0021] The parameter adjustment process of the RepVGGVan network is as follows:
[0022] First, adjust the learning rate, adjust the learning rate from large to small, from 0.1 to 0.0001;
[0023] Second, adjust the batch size, adjust it according to the size of the dataset and different computing resources;
[0024] Adjust the optimizer again, where the optimizer is SGD or Adam;
[0025] Finally, adjust the regularization parameters, including: L1 regularization, L2 regularization, and Dropout.
[0026] Another object of the present invention is to disclose a broiler weight estimation system according to the broiler weight estimation method based on two-view images described in the present invention, which is characterized by including: a broiler image acquisition device and a weight estimation device;
[0027] The broiler image acquisition device includes: a cage, a top-view camera, a side-rear view camera, and an acquisition trigger device; there is a cage opening on one side of the cage, and the broiler to be measured enters the inside of the cage through the cage opening; there is a partition inside the cage, and the partition divides the inside of the cage into a feed storage area and a broiler standing area; there is a feed container in the feed storage area; the broiler standing area can accommodate a single broiler to be measured standing; there is a feeding hole on the partition; the feeding hole restricts the broiler to be measured to only pass its head through the feeding hole to complete feeding; the top-view camera is fixed on the top of the cage and is used to collect the top-view image of the broiler to be measured during feeding; the side-rear view camera is fixed outside the cage, close to the side of the cage opening, and is used to collect the side-rear view image of the broiler to be measured during feeding; the acquisition trigger device is used to trigger the acquisition of pictures and perform weight estimation;
[0028] The two-view images include: the top-view image and the side-rear view image of the broiler to be measured during feeding.
[0029] The weight estimation device includes: an image classification network module, an image segmentation network module, and a RepVGGVan network module; the image classification network module is used to screen the two-view image set, input the side-rear view in the two-view images into the DeformAttn-ShuffleNetV2 network, and output the classification result of the broiler state to obtain the screened two-view images; the image segmentation network module is used to perform semantic segmentation processing on the two-view images that meet the shooting conditions to obtain the segmented two-view images; the RepVGGVan network module is used to estimate the broiler weight by using the segmented two-view images.
[0030] The beneficial effects of the present invention are as follows:
[0031] The present invention discloses a method for estimating broiler weight based on two-view images. Based on the existing intelligent weighing equipment or intelligent poultry breeding equipment, two consumer-grade infrared cameras are installed on the side and top of the breeding cage respectively to collect two-view images of a single broiler in a natural eating state from two perspectives: side and rear and top view. On the basis of reusing the existing breeding equipment, it is possible to collect complete body images of broilers of different ages throughout the growth cycle. By integrating the image information from two perspectives, it is possible to better adapt to the active characteristics of broilers. Taking advantage of the eating nature of broilers, there is a high probability that broilers will actively eat after entering the shooting area. After sensing that the broilers have entered the collection area, multiple groups of two-view images are collected within 1-3 seconds, and the camera is automatically set to standby mode, which greatly reduces the power consumption of the camera.
[0032] In the prior art, when weight estimation is performed based on computer vision technology, it is necessary to perform fine recognition of the subject features, such as chest width and body length, which requires high computing power and power consumption. This defect is particularly obvious for broiler animals such as broilers, which are naturally active, have a large influence on the recognition of body contours due to wings, and are quite large in number. The present invention has found through a large number of observations and experiments that broiler animals are likely to eat immediately after entering the collection area, and in conjunction with the position of the partition and the feeding bowl, broilers are likely to maintain a standing posture with wings tightened for a short period of time during the eating process, and extend their necks to pass their heads through the openings of the partitions to eat. In this state, multiple groups of two-view images are collected within 1-3 seconds after the broiler enters the collection area. There is no need to perform fine recognition of specific body parts of the broiler, especially for the side and rear images, while avoiding differences in the degree of neck extension of different broiler individuals when eating; the present invention automatically identifies two-view images that meet the conditions for shooting by constructing an image classification network and defining the conditions for shooting two-view images, effectively reducing the computing power requirements in the recognition process.
[0033] The image classification network is a DeformAttn-ShuffleNetV2 network. A deformable attention module (DeformAttn Blocks) is constructed on the basis of ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution. By constructing a deformable attention module, dynamic adjustment can be better performed to adapt to various environmental changes in the chicken coop, such as light brightness, random debris changes, etc. Therefore, the DeformAttn-ShuffleNetV2 network disclosed in the present invention has achieved good performance in accuracy, parameter quantity and FLOPs.
[0034] Based on the accurate recognition of two-view images that meet the shootable conditions by the DeformAttn-ShuffleNetV2 network, using the CoRSeaFormer segmentation model as the image segmentation network, high-precision segmented images can be obtained in complex environments, and only a small amount of data annotation is required to obtain segmented two-view images with high segmentation accuracy under the premise of low computational cost and low memory requirements.
[0035] Based on the segmented two-view images, the RepVGGVan network estimates the weight of broilers. By adjusting the parameters through the RepVGGVan network, a relatively accurate estimation result can be obtained. Through experiments, it is verified that the fitting degree between the predicted value and the actual value obtained by applying the broiler weight estimation method based on two-view images disclosed in the present invention is relatively high, and a good prediction effect can be obtained. Brief Description of the Drawings
[0036] Figure 1 It is a schematic flowchart of a method for estimating the weight of broilers based on two-view images according to the present invention;
[0037] Figure 2 It is a two-view segmented image in an embodiment of the present invention, where (a) is a side-back view - RGB image; (b) is a side-back view - segmented image; (c) is a top view - RGB image; (d) is a top view - segmented image;
[0038] Figure 3 It is a schematic diagram of the installation positions of the side-back view and top view cameras in an embodiment of the present invention;
[0039] Figure 4 It is a schematic diagram of a group of RGB images of the object to be estimated in the natural feeding state in an embodiment of the present invention;
[0040] Figure 5 It is a schematic diagram of the structure of the DeformAttn-ShuffleNetV2 network provided by the present invention;
[0041] Figure 6 For Figure 5 an enlarged view of the network structure of the deformable attention module part in
[0042] Figure 7 For Figure 6 an enlarged view of the network structure of the DA convolution part in
[0043] Figure 8 It is a schematic diagram of the structure of the CoRSeaFormer provided by the present invention;
[0044] Figure 9 It is a schematic diagram of the structure of the compression-enhanced axial attention module under coordinate convolution provided by the present invention;
[0045] Figure 10 is Figure 8 the enlarged structure diagram of the receptive field attention convolution fusion module;
[0046] Figure 11 is Figure 9 the enlarged structure diagram of the coordinate convolution in
[0047] Figure 12 the structural schematic diagram of the RepVGGVan network provided by the present invention.
[0048] Figure 13 is Figure 12 the enlarged structure diagram of the visual attention module in
[0049] Figure 14 the top view structural schematic diagram of a broiler chicken image acquisition device provided by the present invention;
[0050] Figure 15 the rear view structural schematic diagram of a broiler chicken image acquisition device provided by the present invention;
[0051] Figure 16 the structural schematic diagram of a broiler chicken weight estimation system based on two-view images provided by the present invention;
[0052] Among them, the broiler chicken image acquisition device - 100, the weight estimation device - 200, the cage body - 101, the partition - 102, the feed storage area - 103, the broiler chicken standing area - 104, the feed container - 105, the feeding hole - 106, the top view camera - 107, the side and rear view camera - 108, the cage body opening - 109. Detailed implementation manners
[0053] The present invention provides a broiler chicken weight estimation method and system based on two-view images. The following further describes the present invention in detail with reference to the accompanying drawings.
[0054] As Figure 1 shown in the embodiment of the present invention, a broiler chicken weight estimation method based on two-view images is disclosed, including the following steps:
[0055] Step S1: In response to the sensor detecting that the broiler chicken enters the shooting location, the shooting device collects a set of two-view images of the broiler chicken;
[0056] Step S2: Construct an image classification network to screen the set of two-view images and obtain two-view images that meet the shootable conditions;
[0057] Step S3: Construct an image segmentation network to perform semantic segmentation processing on the two-view images that meet the shootable conditions and obtain the segmented two-view images;
[0058] Step S4: Based on the two segmented view images, input them into the RepVGGVan network to estimate the weight of the broiler chicken.
[0059] In this embodiment, a method for estimating the weight of a broiler chicken based on two view images disclosed by the present invention is applicable to the scenario of estimating the weight of an object to be estimated, such as weight estimation in broiler breeding. In this embodiment, the weight estimation of broiler chickens is taken as an example. The following will introduce the specific implementation manners of each step in combination with specific embodiments.
[0060] Step S1: In response to the sensor detecting that the broiler chicken enters the shooting location, the shooting device collects a set of two view images of the broiler chicken;
[0061] The two view images in Step S1 include: a top view and a side rear view; the shooting device includes: a top view shooting device and a side rear view shooting device.
[0062] In this embodiment, the shooting device is two consumer-grade infrared cameras. The two consumer-grade infrared cameras are respectively installed on the side and above the weight estimation device to collect two view images of a single broiler chicken in a natural feeding state from two perspectives of side rear and top views. The shooting time is from when the broiler chicken stands on the weight estimation device until it finishes eating and leaves the device, obtaining multiple sets of two view images in the natural feeding state as Figure 4 shown, forming a set of two view images of the broiler chicken.
[0063] Those skilled in the art should know that in response to the sensor detecting that the broiler chicken meets the shootable conditions, the camera is triggered to start shooting. The shootable conditions are:
[0064] The broiler chicken enters the cage, puts its head through the partition and starts eating; and the broiler chicken is in a standing posture; and the wings of the broiler chicken are in a tightened state;
[0065] The sensor includes but is not limited to infrared sensors, optical sensors, bioelectric induction sensors, and weighing scales, etc., which are not specifically limited in this embodiment.
[0066] According to the inventor's extensive experimental observations, in most cases, driven by the nature of broiler chickens, when broiler chickens enter the cage, they will immediately rush towards the food to eat, prompting the broiler chickens to stretch their necks and pass their heads through the partition in the cage, maintaining a standing posture and keeping their wings in a tightened state. Therefore, in most cases, when the sensor senses that the broiler chicken is in the cage within 1 to 3 seconds, the broiler chicken is most likely in the shootable condition. At this time, the shooting opportunity is triggered to call the camera to take 3 to 5 groups of two-view images. In this embodiment, even if the broiler chicken does not enter the shootable condition within 1 to 3 seconds, unqualified pictures can be manually excluded during the model training stage. In the actual production environment, unqualified pictures are excluded by using an image classification network in the subsequent steps. Therefore, in this embodiment, the shooting opportunity for multiple groups of two-view images is within 1 - 3 seconds when the sensor senses that the broiler chicken enters the cage. Therefore, the method for sensing the shooting opportunity used in this embodiment has the advantages of simple structure, rapid judgment, and low cost.
[0067] In step S1, the sensor is a radio frequency identification sensor, and the broiler chicken wears a radio frequency identification tag to sense the entry and exit of the broiler chicken from the shooting location and sense the identity identification information of the broiler chicken.
[0068] In a preferred embodiment, the broiler chicken wears an electronic tag, and the sensor is an electronic tag sensor. By cooperating the electronic tag sensor with the electronic tag, not only can it sense that the broiler chicken enters the designated area in the cage, but also the identity information of the broiler chicken can be obtained, which is convenient for weight tracking throughout the growth cycle of the broiler chicken. The electronic tag includes but is not limited to RFID tags, NFC tags, Bluetooth tags, and ZigBee tags, etc., and is not specifically limited in this embodiment.
[0069] In this embodiment, multiple groups of two-view images of the broiler chicken in the natural feeding state are captured, and then the multiple groups of images obtained in the natural feeding state are selected to obtain multiple groups of natural feeding state images that meet the shootable conditions, and the camera is controlled to enter the standby state.
[0070] In an alternative embodiment, in response to the sensor sensing that the broiler chicken leaves the shooting location, the shooting device enters the standby state.
[0071] In this embodiment, the installation position of the camera is as Figure 3 shown, where the green rectangular area is the installation position of the top-view shooting device, and the red rectangular area is the installation position of the side-back view shooting device. By installing consumer-grade 2D cameras on the side and above the weight estimation device, images of the broiler chicken eating in the natural state are captured manually. A total of 1122 broiler chickens are collected, and two views Figure 1 A total of about 74,932 images are used as the dataset for subsequent model training.
[0072] In an optional embodiment, multiple groups of two-view images of broilers in their natural feeding state are captured, and then the obtained multiple groups of images in the natural feeding state are selected to obtain multiple groups of natural feeding state images that meet the shootable conditions. When the sensor senses that the broiler leaves the breeding cage, a weighing scale is used to measure the true weight value of the broiler, and the true weight value and the multiple groups of natural feeding state images that meet the shootable conditions are simultaneously added to the dataset for model training and testing.
[0073] Those skilled in the art should know that considering the large weight change of broilers during the entire growth process, the installation position of the camera should ensure that the camera can capture the complete image of the broiler with a large age. Those skilled in the art make corresponding adjustments according to factors such as the performance parameters of the selected camera and the breeding cage ruler, which are not specifically limited in this embodiment.
[0074] Step S2: Construct an image classification network to screen the two-view image set and obtain two-view images that meet the shootable conditions;
[0075] The image classification network in the step S2 is the DeformAttn-ShuffleNetV2 network. The DeformAttn-ShuffleNetV2 network constructs a deformable attention module on the basis of ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution; the input of the DeformAttn-ShuffleNetV2 network is the side-back view in the two-view images; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the broiler state; the classification results include: zhanli and qita; when the classification result is zhanli, it is determined that the two-view image corresponding to the input side-back view meets the shootable conditions.
[0076] As Figure 6 shown, the network structure of the deformable attention module is: the input feature channels are divided into a first branch and a second branch; the first branch sequentially includes: a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit; the second branch sequentially includes: a 1×1 convolution, batch normalization, a rectified linear unit, a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit, a 1×1 convolution, batch normalization, a rectified linear unit, a DA convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit; after the first branch and the second branch go through a merging operation, channel shuffling is performed. As Figure 7As shown in the figure, the DA convolution includes: a depthwise separable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a deformable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a first fully connected layer, a second fully connected layer, and a Softmax function.
[0077] In this embodiment, replacing the second depthwise separable convolution in the DA convolution with a deformable convolution improves the adaptability to geometric deformations. The deformable convolution can dynamically adjust the position of the convolution kernel according to the content of the input feature map, thus better capturing the changes in the geometric shape and pose of the object. In contrast, the traditional depthwise separable convolution has fixed sampling positions and cannot handle spatial deformation well; it enhances the feature expression ability. The deformable convolution can adaptively adjust the sampling point positions by introducing additional offset learning, better capturing the detailed information of the target object and the complex background; it improves the detection and segmentation performance of the model. The deformable convolution can enhance the model's ability to detect edges and shapes, thus improving the overall detection and segmentation performance; at the same time, it can combine the advantages of multi-scale features and spatial deformations, enabling the model to not only better utilize multi-scale information but also better adapt to the deformation characteristics of the target object.
[0078] The shootable conditions in step S2 are as follows: the broiler enters the cage, puts its head through the partition and starts eating; and the broiler is in a standing posture; and the wings of the broiler are in a tightened state.
[0079] Considering that the foraging postures of broilers in the natural state are different, such as fully extended legs, semi-squat legs, and reclining broilers, etc. By observing the posture changes during the broiler's eating process, compared with the natural standing posture, there are significant differences in the neck extension length of broilers when eating; when eating naturally, the wings naturally spread out by the broilers cover about half of the body contour; in addition, it is found that after eating for a few seconds, broilers are likely to change from a standing posture to a lying posture. Therefore, the posture of the broiler has an important impact on the accuracy of the subsequent body weight estimation model. Image classification technology is used to identify and extract the most suitable posture images for body weight estimation, and an image classification network model is introduced to automatically screen images to select broiler images that meet the shootable conditions.
[0080] In this embodiment, multiple groups of two-view images in the natural eating state are selected and classified to obtain a set of two-view images that meet the shootable conditions. The filtered multiple groups of natural eating state images are classified. From the perspective of the side-rear view, natural eating state images with fully upright legs, fully tightened wings, and eating with the head down are classified into the "zhanli" category; other eating postures are classified into the "qita" category, and other eating postures include, but are not limited to, lying down to eat, semi-squatting to eat, etc.
[0081] In this embodiment, the DeformAttn-ShuffleNetV2 network performs image classification based on the rear-side view image among the two-view images, and classifies the rear-side view image into two categories: standing posture and other postures, where other postures include situations such as squatting and lying down to eat. Based on the ShuffleNetV2 network, the SKnet attention mechanism is introduced, named SK-ShuffleNetV2, and further, the deformable convolution is introduced on the basis of the SK-ShuffleNetV2 network, named DeformAttn-ShuffleNetV2, as the classification network actually used in this embodiment. As shown in Table 1, the DeformAttn-ShuffleNetV2 network has achieved good performance in terms of accuracy, number of parameters, and FLOPs.
[0082] Table 1 Image Classification Result Table
[0083] Model Top-1(%) Params(M) FLOPs(G) ShuffleNetV2 97.20 1.26 1.45 SK-ShuffleNetV2 97.66 1.6 1.46 CoRSeaFormer 98.13 6.15 6.87
[0084] ShuffleNetV2 is a novel neural network architecture in the prior art, aiming to effectively reduce the parameters and computational complexity by implementing per-channel group convolution and channel shuffle operations, thereby improving the computational efficiency. This model introduces the per-channel group convolution method, divides the input channels into multiple groups, and independently applies convolution operations to each group. Subsequently, the results of these operations are merged in the channel dimension, which can effectively improve the network capacity and information mobility, thereby optimizing the performance. By effectively reducing the number of parameters and computational load, ShuffleNetV2 achieves high classification accuracy while minimizing resource requirements. To further improve the performance, ShuffleNetV2 adopts repeated modules. Each module consists of three operations: per-channel group convolution, channel mixing, and per-channel group convolution again. By stacking multiple modules, the overall depth and width of the network are increased, thereby improving the expressive ability of the model.
[0085] In this embodiment, in order to improve the accuracy of detecting broiler images in complex environments and select the picture data of the best posture, two technologies are introduced on the basis of ShuffleNetV2: the SKnet attention mechanism (SelectiveKernel Networks) and the deformable convolution (Deformable Convolution).
[0086] The SKnet attention mechanism enables each neuron to adaptively adjust the size of its receptive field (i.e., the convolution kernel) according to different scales of the input information, can effectively capture multi-scale features in the complex image space, and does not consume as much computational resources as traditional CNNs. In addition, the SKnet attention mechanism can also integrate deep features, thereby improving the interpretability and comprehensibility of the captured features.
[0087] Deformable convolution can perform non-uniform and adaptive sampling and perception on the input features. By learning the offsets, deformable convolution adaptively adjusts the size and shape of the receptive field according to the context information, so as to more effectively capture the spatial structure and details of the target. Adding learnable offsets to the deformable convolution enhances the adaptability of the traditional convolution operation, thus better adjusting the spatial relationship and deformation between features and improving the expression ability of the model.
[0088] The structure of the DeformAttn-ShuffleNetV2 network is as Figure 5 shown.
[0089] The input of the DeformAttn-ShuffleNetV2 network is the side-back view in the two-view images of broilers, and the output is the classification result of the broiler state; based on the standard ShuffleNetV2 network structure, Stage2, Stage3, and Stage4 are composed of deformable attention modules to achieve the fusion of the SKnet attention mechanism (Selective KernelNetworks) and deformable convolution (Deformable Convolution); the network structure of the deformable attention module part is as Figure 6 shown.
[0090] In this embodiment, for the 74,932 RGB two-view images of 1,122 broilers obtained, each broiler corresponds to multiple groups of natural feeding state images. The multiple groups of natural feeding state images corresponding to each broiler are screened, and the screened multiple groups of natural feeding state images are classified. Respectively, 7,036 side-back view RGB images of the "zhanli" category and 4,615 side-back view RGB images of the "qita" category are classified, and 11,651 sample natural feeding state images are obtained.
[0091] In this embodiment, after using the side-back images of 1,122 broilers, the images are divided into two categories: standing posture and other postures, and the image classification model DeformAttn-ShuffleNetV2 is trained.
[0092] The model classification results and related data are as follows, where ShuffleNetV2 is the basic model, SK-ShuffleNetV2 and DeformAttn-ShuffleNetV2 are the improved models, and finally the DeformAttn-ShuffleNetV2 model is selected for actual use.
[0093] Step S3: Construct an image segmentation network, perform semantic segmentation processing on the two-view images that meet the photographable conditions, and obtain the segmented two-view images;
[0094] In step S3, the image segmentation network is the CoRSeaFormer segmentation model. The CoRSeaFormer network constructs a compression-enhanced axial attention module and a receptive field attention convolution fusion module under coordinate convolution based on SeaFormer. The compression-enhanced axial attention module under coordinate convolution is used to introduce coordinate convolution, and the receptive field attention convolution fusion module is used to introduce receptive field attention convolution. The input of the CoRSeaFormer network is a two-view image. The output of the CoRSeaFormer network is the segmentation result of the two-view image of the broiler, and the recognition result output by the initial image segmentation model is as follows Figure 2 shown, where (a) is the side-rear view - RGB image; (b) is the side-rear view - segmentation image; (c) is the top view - RGB image; (d) is the top view - segmentation image.
[0095] As Figure 9 shown, the network structure of the compression-enhanced axial attention module is as follows: the input feature channels calculate queries, keys, and values after coordinate convolution. The queries, keys, and values are divided into a detail-enhanced kernel branch and an axial attention compression branch. The detail-enhanced kernel branch successively includes: concatenating the query, key, and value channels, 3×3 depthwise separable convolution, batch normalization, ReLU function, 1×1 convolution, batch normalization. The axial attention compression branch is divided into a horizontal sub-branch and a vertical sub-branch. The horizontal sub-branch successively includes: horizontally compressing the query, key, and value, multi-head attention. The vertical sub-branch successively includes: vertically compressing the query, key, and value, multi-head attention. The horizontal sub-branch and the vertical sub-branch perform pointwise addition using tensor broadcast operation and then perform 1×1 convolution. The detail-enhanced kernel branch and the axial attention compression branch are combined through multiplication calculation. The coordinate convolution includes: coordinate generation, normalization, and adding coordinate channels.
[0096] In this embodiment, the coordinate convolution is introduced before calculating the queries, keys, and values, and the position information (coordinates) is directly input into the convolutional layer, making it easier for the network to learn features related to spatial positions and enhancing the spatial perception ability. The coordinate convolution explicitly provides position information, enabling the network to directly understand the spatial structure in the image. Since the network no longer needs to implicitly learn spatial information through complex structures, CoordConv can accelerate the convergence speed of the network and improve the learning efficiency of the model. By introducing coordinate channels, the network can obtain stronger expressive power without adding too many parameters. With explicit position information, the network can exhibit better robustness and generalization ability when processing objects in different scenarios, positions, or scales.
[0097] As Figure 10As shown, the network structure of the receptive field attention convolution fusion module is as follows: the input feature channels are composed of a third branch and a fourth branch; the third branch sequentially includes: a receptive field attention convolution module; the fourth branch sequentially includes: a 3×3 convolution, batch normalization, a Sigmoid activation function, and an upsampling operation; the third branch and the fourth branch are merged after multiplication; the receptive field attention convolution module includes: a regional attention mechanism, a convolutional layer, and a residual connection.
[0098] In this embodiment, replacing the convolution in the receptive field attention convolution fusion module with a receptive field attention convolution module can improve the feature extraction ability of the convolutional neural network in a specific area. Since the receptive field of RFAConv can be dynamically adjusted, it can better process global context information and local detail information simultaneously. Compared with traditional convolution, it can extract features more effectively at multiple scales; RFAConv can automatically adjust the size of the receptive field according to the content of the input image, thereby enhancing the adaptability of the convolutional layer to different types of images; the dynamic adjustment of the receptive field can help the model more accurately identify the details in the image.
[0099] In this embodiment, the CoRSeaFormer network performs image segmentation based on the side rear view and the top view in the two views. Based on the SeaFormer network, coordinate convolution and receptive field attention are introduced, named CoRSeaFormer, as the segmentation network actually used in this embodiment. As shown in Table 2, the CoRSeaFormer network has achieved good performance in mIoU, the number of parameters, and FLOPs.
[0100] Table 2 Image Segmentation Result Table
[0101]
[0102] SeaFormer is an efficient Transformer network designed for optimizing computational complexity. By introducing a sparse attention mechanism, SeaFormer significantly reduces the global computational requirements in the traditional self-attention mechanism, restricting the attention scope to a local window, thereby reducing the computational complexity from O(n2) to O(n). This sparsification strategy not only improves the efficiency of the model in processing large-scale image tasks but also retains the ability to capture global context information. In addition, SeaFormer combines multi-scale feature fusion technology, enabling it to effectively process image details at different scales while maintaining efficient global information extraction. The network also features a lightweight design, significantly reducing the number of parameters and computational requirements, making it particularly suitable for deployment on resource-constrained devices (such as mobile devices or embedded systems).
[0103] In this embodiment, the image segmentation network in step S3 is the CoRSeaFormer segmentation model. The CoRSeaFormer network constructs a compression-enhanced axial attention module and a receptive field attention convolution fusion module under coordinate convolution based on the SeaFormer. The compression-enhanced axial attention module under coordinate convolution is used to introduce coordinate convolution, and the receptive field attention convolution fusion module is used to introduce receptive field attention convolution. The structure of the CoRSeaFormer network is as Figure 8 shown. Compared with the SeaFormer network in the prior art, in the CoRSeaFormer network disclosed in the present invention, CoSeaFormerLayer uses CoSeaFormer to replace SeaFormer in the SeaFormer network; and uses RFAConv to replace Conv+BN in the SeaFormer network.
[0104] In this embodiment, in order to improve the segmentation accuracy of broiler chicken images in complex environments and obtain the best picture data, two technologies are introduced based on the SeaFormer: Coordinate Convolution and Receptive Field Attention Convolution.
[0105] Coordinate Convolution is a convolution method that enhances the spatial perception ability of convolutional neural networks. It directly introduces the input coordinate information (such as the x and y coordinates of each pixel) into the convolution operation, so that the network can better understand the spatial relationship.
[0106] Receptive Field Attention Convolution is a convolution operation that enhances the ability of convolutional neural networks by dynamically adjusting the receptive field. RFAConv introduces the receptive field attention mechanism, enabling the network to adaptively adjust the receptive field size according to the content of the input image, thereby improving the network's feature extraction ability.
[0107] The input of the CoRSeaFormer network is the two-view image of the broiler chicken, and the output is the two-view segmentation result of the broiler chicken. Based on the standard SeaFormer network structure, the fusion of Coordinate Convolution and Receptive Field Attention Convolution is realized. The network structure of the coordinate convolution module part is as Figure 11 shown, and the structure of the receptive field attention convolution fusion module is as Figure 10 shown.
[0108] In this embodiment, for the 45,258 RGB two-view segmentation images of 1,122 broilers obtained after passing through the image classification network, each broiler corresponds to multiple groups of natural feeding state images. The multiple groups of natural feeding state images corresponding to each broiler are screened, and after screening, the multiple groups of natural feeding state images are labeled. 893 side-rear view RGB images and 536 top-view RGB images are respectively labeled, and 1,429 sample natural feeding state images are obtained.
[0109] In this embodiment, after using the side-rear and top-view images of 1,122 broilers, an image segmentation model CoRSeaFormer is trained.
[0110] The model segmentation results and related data are as follows. Among them, SeaFormer is the basic model, and CoRSeaFormer is the improved model. Finally, the CoRSeaFormer model is selected for actual use.
[0111] Step S4: Based on the segmented two-view images, input them into the RepVGGVan network to estimate the weight of the broilers.
[0112] In the step S4, the weight estimation network is the RepVGGVan segmentation model. The RepVGGVan network constructs a visual attention module based on RepVGG. The input of the RepVGGVan network is the two-view segmentation images. The output of the RepVGGVan network is the weight value of the broilers.
[0113] As Figure 12 shown, the synthetic image sequentially passes through the first 3×3 convolutional layer, the first 1×1 convolutional layer, the first downsampling layer, the second 3×3 convolutional layer, the second 1×1 convolutional layer, the first residual connection layer, the third 3×3 convolutional layer, the third 1×1 convolutional layer, the second residual connection layer, the fourth 3×3 convolutional layer, the fourth 1×1 convolutional layer, the third residual connection layer, and the visual attention module to obtain the estimated weight of the broilers. The convolutional stride of the first 3×3 convolutional layer is 2, and the padding is 1. The convolutional stride of the other 3×3 convolutional layers is 1, and the padding is 1. The convolutional stride of the 1×1 convolutional layer is 1, and there is no padding. The first downsampling layer is a convolutional function with a convolutional stride of 2. The residual connection layer is the residual operation in Resnet. The visual attention module includes: channel attention, spatial attention, and their weighted fusion.
[0114] In this embodiment, a visual attention module is introduced based on the RepVGG network. Using its efficient visual attention mechanism, it can adaptively focus on key regions in the image, achieve precise feature extraction, and perform excellently in the fusion of global and local information. Through lightweight design, VAN significantly reduces the computational complexity and the number of parameters, making it suitable for resource-constrained devices and real-time applications. At the same time, the sparse attention mechanism of VAN can greatly reduce unnecessary computations and improve the overall performance.
[0115] RepVGG is an innovative convolutional neural network architecture whose design goal is to improve the model's expressive ability while maintaining efficient inference. This network adopts a multi-branch design during the training phase, including complex structures such as 3×3 convolutions, 1×1 convolutions, and skip connections, which can enhance the model's feature extraction ability. However, during the inference phase, RepVGG uses the structural reparameterization technique to fuse all branches into a simple VGG-like structure, retaining only single-path convolution operations. This design greatly simplifies the model's inference process, significantly improves the computational efficiency, and maintains the high performance during training. RepVGG not only has obvious advantages in the inference speed of the model, but also due to the simplicity of its structure, it is very suitable for large-scale deployment, especially in resource-constrained environments (such as mobile devices or edge computing scenarios).
[0116] In this embodiment, in order to improve the accuracy of broiler weight estimation in complex environments and obtain the best weight estimation value, a visual attention (Visual Attention Network) is introduced based on RepVGG.
[0117] Visual Attention Network is a neural network architecture based on the visual attention mechanism, aiming to improve the feature extraction efficiency in image processing tasks. Different from traditional convolutional neural networks, VAN can adaptively focus on important local regions and ignore irrelevant information when processing images by introducing a visual attention module.
[0118] The structure of the RepVGGVan network is as Figure 12 shown, Figure 12 where VAN represents the Visual Attention Network;
[0119] The input of the RepVGGVan network is the two-view images of broilers, and the output is the weight estimation result of broilers. Based on the standard RepVGG network structure, the fusion of visual attention is achieved. The network structure of the VAN part of the visual attention module is as Figure 13 shown.
[0120] In this embodiment, the process of obtaining the dataset required for training the RepVGGVan network is as follows:
[0121] Install 2 cameras on the weight estimation device in the broiler feeding area. By capturing two-view images of broilers eating in a natural state, a comprehensive dataset containing two perspectives of the eating process is obtained. To facilitate accurate weight estimation, the two-view images will pass through the CoRSeaFormer segmentation model to obtain segmented images that meet the unified input requirements of the weight estimation model. Immediately after taking the broiler images, the staff weighs the chickens and records the weight values as the true values. It should be noted that the input of this model is a synthetic image formed by splicing the segmented two-view images obtained from the above two perspectives.
[0122] In this embodiment, a total of 1122 broilers are collected, and about Figure 1 a total of 45,258 images are used as the dataset for subsequent model training, and the ratio of the training set to the test set is 8:2.
[0123] Input the segmented two-view images in the training set into the RepVGGVan network to obtain the estimated weight of the broilers. Compare the estimated weight of the broilers with the true weight value, and continuously change the parameters in the RepVGGVan network until the training stop condition is met, and a trained RepVGGVan network is obtained.
[0124] In this embodiment, the specific process of changing the parameters in the RepVGGVan network is as follows:
[0125] First, adjust the learning rate, which is adjusted from large to small, from 0.1 to 0.0001;
[0126] In an optional embodiment, the cosine annealing and adaptive learning rate are used to adjust the learning rate;
[0127] Second, adjust the batch size, which is adjusted according to the size of the dataset and different computing resources;
[0128] Third, adjust the optimizer, and the optimizer is SGD or Adam;
[0129] Finally, adjust the regularization parameters, including: L1 regularization, L2 regularization, and Dropout;
[0130] In this embodiment, when the training data is less, L1 and L2 regularizations help prevent overfitting. Dropout reduces the complexity of the model by randomly discarding a part of the neurons.
[0131] In this embodiment, after obtaining the image segmentation result, the corresponding two-view images are uniformly used as the input of the RepVGGVan network to train a weight estimation model for live broiler chickens. The weight estimation results and related data are shown in Table 3:
[0132] Table 3 Weight Estimation Results Table
[0133] Model MAE Params FLOPs CNN 17.180g 3.27M 5.15G RepVGG 15.335g 7.83M 62.95G RepVGGVan 14.675g 12.85M 73.09G
[0134] MAE (Mean Absolute Error) in Table 3
[0135] Definition: MAE is the average of the absolute differences between the predicted value and the actual value. The formula is:
[0136]
[0137] where y i is the actual value, is the predicted value, and n is the number of samples.
[0138] Function: MAE is used to measure the average absolute gap between the predicted value and the actual value of the model.
[0139] Considering that traditional image segmentation models need to rely on data annotation and need to label fixed regions, while in the broiler weight assessment scenario, broiler chickens have strong hyperactivity, resulting in the inability of the existing technology to accurately estimate the fixed regions of broiler chickens. Applying the image segmentation model CoRSeaFormer disclosed in this embodiment does not require relying on the labels of fixed regions and can obtain more accurate segmentation images.
[0140] The RepVGGVan network disclosed in this embodiment is different from the prior art that relies on 2D or 3D cameras to capture animal images, extracts key parameters such as the animal's body shape through image processing technology, and uses machine learning or deep learning regression methods to predict the weight of broiler chickens according to the extracted features. A method for estimating the weight of broiler chickens based on two-view images disclosed in the present invention does not require relying on high-definition 2D or 3D cameras to capture animal images, uses a consumer-grade camera to collect 2D images, and realizes efficient and stable weight estimation of broiler chickens.
[0141] In a specific embodiment, applying a method for estimating the weight of broiler chickens based on two-view images disclosed in the present invention, based on the trained RepVGGVan network, realizes automatic estimation of the weight of broiler chickens. The specific process is as follows:
[0142] Collect 2D two-view images of several broiler chickens to form a set of two-view images of broiler chickens; use an image classification network to screen the set of two-view images of broiler chickens to obtain the screened two-view images that meet the preset posture requirements of broiler chickens; use an image segmentation network to perform semantic segmentation on the two-view images that meet the shooting conditions to obtain the segmented two-view images; use the segmented two-view images as the input of the RepVGGVan network to estimate the weight of broiler chickens.
[0143] In this embodiment, a consumer-grade 2D camera is installed at two positions, the side and the top, of the weight estimation device. Images of broiler chickens in their natural feeding state are automatically collected through the image classification model DeformAttn-ShuffleNetV2. The image segmentation model Segment Anything Model respectively segments the two classified two-view images. Then, the group of images with the best segmentation effect of this chicken is used as the input of the weight estimation model to automatically estimate the weight of broiler chickens.
[0144] As Figure 16 shown, another embodiment of the present invention discloses a broiler chicken weight estimation system according to the broiler chicken weight estimation method based on two-view images described in the present invention, including: a broiler chicken image acquisition device 100 and a weight estimation device 200;
[0145] As Figure 14 and Figure 15 shown, the broiler chicken image acquisition device 100 includes: a cage body 101, a top-view camera 107, a side-rear-view camera 108, and an acquisition trigger device; a cage body opening 109 is provided on one side of the cage body 101, and the broiler chicken to be measured enters the inside of the cage body 101 through the cage body opening 109; a partition 102 is provided inside the cage body 101, and the partition 102 divides the inside of the cage body 101 into a feed storage area 103 and a broiler chicken standing area 104; a feed container 105 is provided in the feed storage area 103; the broiler chicken standing area 104 can accommodate a single broiler chicken to stand; a feeding hole 106 is provided on the partition 102; the feeding hole 106 restricts the broiler chicken to be measured to only pass its head through the feeding hole 106 to complete feeding; the top-view camera 107 is fixed on the top of the cage body 101 for collecting top-view images of the broiler chicken to be measured during feeding; the side-rear-view camera 107 is fixed outside the cage body 101, close to the side of the cage body opening 109, for collecting side-rear-view images of the broiler chicken to be measured during feeding; the acquisition trigger device is used to trigger the acquisition of pictures and perform weight estimation;
[0146] The two-view images include: top-view images and side-rear-view images of the broiler chicken to be measured during feeding
[0147] The input of the DeformAttn-ShuffleNetV2 network is the side-rear view image in the two-view images; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the broiler state; the classification results include: zhanli and qita; when the classification result is zhanli, it is determined that the two-view image corresponding to the input side-rear view image meets the shootable conditions.
[0148] The weight estimation device 200 includes: an image classification network module, an image segmentation network module, and a RepVGGVan network module; the image classification network module is used to screen the two-view image set, input the side-rear view image in the two-view images into the CoRSeaFormer network, and the output is the classification result of the broiler state, obtaining the screened two-view images; the image segmentation network module is used to perform semantic segmentation processing on the two-view images that meet the shootable conditions to obtain the segmented two-view images; the RepVGGVan network module is used to estimate the weight of the broiler using the segmented two-view images.
[0149] Considering that the posture of the top view is relatively stable under the shooting standard, in this embodiment, the image classification network module screens the side-rear view, and the top view takes the corresponding result.
[0150] For example, the object to be estimated is a broiler. By installing 2 consumer-grade 2D infrared cameras behind and above the weight estimation device, the two-view images of the broiler in the natural feeding state standing on the weight estimation device are captured. Using a Python program to parse the RTSP protocol stream to obtain the two-view images of the broiler, at least one first image of the object to be estimated can be obtained. The 2D infrared camera follows the RTSP protocol. As Figure 16 shown, the system structure of the broiler weight estimation system based on two-view images realizes the data interaction between the broiler image acquisition device 100 and the weight estimation device 200.
[0151] In this embodiment, by establishing a spatial rectangular coordinate system, the image of the broiler in the natural feeding state is accurately obtained, as Figure 2 shown, where (a) is the side-rear view - RGB image; (b) is the side-rear view - segmented image; (c) is the top view - RGB image; (d) is the top view - segmented image. Figure 3 is the installation position diagram of the 2 cameras in the weight estimation device provided by the present invention. The green rectangular area is the installation position of the top view shooting device, and the red rectangular area is the installation position of the side-rear view shooting device; the specific coordinates of the cameras are as follows: the position of the top view camera is 8.7 cm away from the left pivot of the iron plate, the side-rear iron plate is 18.8 cm away from the top of the iron gate, and the camera is 15.3 cm away from the edge of the iron plate.
[0152] In this embodiment, those skilled in the art should know that considering the large weight change of broiler chickens during the entire growth process, the installation position of the camera should ensure that the camera can capture a complete image of the body of broiler chickens at a large age. Those skilled in the art make corresponding adjustments according to factors such as the performance parameters of the selected camera and the dimensions of the breeding cages, which are not specifically limited in this embodiment. Compared with the prior art where most weight estimations are made by collecting 2D images from only one or two perspectives, the present invention collects pictures from two directions, namely the side rear view and the top view, fully considering the problems that broiler chickens are more active than animals such as pigs, cows, and sheep, have large body shape changes during the breeding cycle, and the standing posture and wing state have a great impact on weight prediction. The model is trained using the image data collected from the two perspectives to obtain a relatively accurate prediction result. The weight estimation results and related data are shown in Table 3. For example, the MAE is 14.675 g, which proves that the fitting degree between the predicted value and the actual value obtained by applying the broiler chicken weight estimation method based on two-view images disclosed in the present invention is relatively high, and a good prediction effect can be obtained.
Claims
1. A method for estimating the weight of broiler chickens based on two-view images, characterized in that, Including the following steps: Step S1: In response to the sensor detecting that the broiler enters the shooting location, a two-view image set of the broiler is collected by the shooting device; Step S2: Construct an image classification network to screen the two-view image set and obtain two-view images that meet the shootable conditions; Step S3: Construct an image segmentation network to perform semantic segmentation processing on the two-view images that meet the shootable conditions and obtain the segmented two-view images; Step S4: Based on the segmented two-view images, input them into the RepVGGVan network to estimate the weight of the broiler.
2. The method for estimating the body weight of broiler chickens based on two-view images according to claim 1, wherein In the two-view images in Step S1, they include: a top view and a side rear view; the shooting device includes: a top view shooting device and a side rear view shooting device; In Step S1, the sensor is a radio frequency identification sensor, and the broiler wears a radio frequency identification tag to realize detecting the entry and exit of the broiler from the shooting location and detecting the identity identification information of the broiler.
3. The method for estimating the body weight of broilers based on two-view images according to claim 1, characterized in that, The image classification network in Step S2 is the DeformAttn-ShuffleNetV2 network. The DeformAttn-ShuffleNetV2 network constructs a deformable attention module based on ShuffleNetV2 to introduce the SKnet attention mechanism and deformable convolution; the input of the DeformAttn-ShuffleNetV2 network is the side rear view in the two-view images; the output of the DeformAttn-ShuffleNetV2 network is the classification result of the broiler state; the classification result includes: zhanli and qita; when the classification result is zhanli, it is determined that the two-view image corresponding to the input side rear view meets the shootable conditions; The network structure of the deformable attention module is: the input feature channels are divided into a first branch and a second branch; the first branch sequentially includes: a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit; the second branch sequentially includes: a 1×1 convolution, batch normalization, a rectified linear unit, a 3×3 convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit, a 1×1 convolution, batch normalization, a rectified linear unit, a DA convolution, batch normalization, a 1×1 convolution, batch normalization, a rectified linear unit; after the first branch and the second branch are merged, channel shuffling is performed; the DA convolution includes: a depthwise separable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a deformable convolution with both input and output channels being c, batch normalization, a rectified linear unit, a first fully connected layer, a second fully connected layer, and a Softmax function.
4. The method for estimating the body weight of broiler chickens based on two-view images according to claim 1, characterized in that, The shootable conditions in Step S2 are: the broiler enters the cage, puts its head through the partition and starts eating; and the broiler is in a standing posture; and the wings of the broiler are in a tightened state.
5. The method for estimating the weight of broilers based on two-view images according to claim 1, wherein In step S3, the image segmentation network is the CoRSeaFormer segmentation model. The CoRSeaFormer network constructs a compression-enhanced axial attention module and a receptive field attention convolution fusion module under coordinate convolution based on the SeaFormer. The compression-enhanced axial attention module under coordinate convolution is used to introduce coordinate convolution, and the receptive field attention convolution fusion module is used to introduce receptive field attention convolution. The input of the CoRSeaFormer network is a two-view image. The output of the CoRSeaFormer network is the segmentation result of the two-view image of the broiler chicken.
6. The method for estimating the body weight of broilers based on two-view images according to claim 5, wherein, The network structure of the compression-enhanced axial attention module is as follows: the input feature channels calculate queries, keys, and values after coordinate convolution; The queries, keys, and values are divided into a detail enhancement kernel branch and an axial attention compression branch; The detail enhancement kernel branch sequentially includes: concatenating the query, key, and value channels, 3×3 depthwise separable convolution, batch normalization, ReLU function, 1×1 convolution, and batch normalization. The axial attention compression branch is divided into a horizontal sub-branch and a vertical sub-branch. The horizontal sub-branch sequentially includes: horizontally compressing the query, key, and value, and multi-head attention. The vertical sub-branch sequentially includes: vertically compressing the query, key, and value, and multi-head attention. The horizontal sub-branch and the vertical sub-branch perform pointwise addition using tensor broadcast operation and then perform 1×1 convolution. The detail enhancement kernel branch and the axial attention compression branch are combined through multiplication calculation. The coordinate convolution includes: coordinate generation, normalization, and adding coordinate channels.
7. The method for estimating the body weight of broilers based on two-view images according to claim 5, wherein, The network structure of the receptive field attention convolution fusion module is as follows: the input feature channels consist of a third branch and a fourth branch. The third branch sequentially includes: a receptive field attention convolution module. The fourth branch sequentially includes: 3×3 convolution, batch normalization, Sigmoid activation function, and upsampling operation. The third branch and the fourth branch are combined after multiplication. The receptive field attention convolution module includes: a regional attention mechanism, a convolutional layer, and a residual connection.
8. The method for estimating the weight of broiler chickens based on two-view images according to claim 1, characterized in that In step S4, the RepVGGVan network receives the synthetic image composed of the segmented two-view images and outputs the estimated weight of the broiler chicken. The synthetic image sequentially passes through a first 3×3 convolutional layer, a first 1×1 convolutional layer, a first downsampling layer, a second 3×3 convolutional layer, a second 1×1 convolutional layer, a first residual connection layer, a third 3×3 convolutional layer, a third 1×1 convolutional layer, a second residual connection layer, a fourth 3×3 convolutional layer, a fourth 1×1 convolutional layer, a third residual connection layer, and a visual attention module to obtain the estimated weight of the broiler chicken. The convolutional stride of the first 3×3 convolutional layer is 2, and the padding is 1. The convolutional stride of the other 3×3 convolutional layers is 1, and the padding is 1. The convolutional stride of the 1×1 convolutional layer is 1, and there is no padding. The first downsampling layer is a convolutional function with a convolutional stride of 2. The residual connection layer is the residual operation in Resnet; The visual attention module includes: channel attention, spatial attention, and their weighted fusion.
9. The method for estimating the weight of broilers based on two-view images according to claim 8, wherein, The parameter adjustment process of the RepVGGVan network is as follows: First, adjust the learning rate, which is adjusted from large to small, from 0.1 to 0.0001; Second, adjust the batch size, which is adjusted according to the size of the dataset and different computing resources; Third, adjust the optimizer, and the optimizer is SGD or Adam; Finally, adjust the regularization parameters, including: L1 regularization, L2 regularization, and Dropout.
10. A broiler weight estimation system for the broiler weight estimation method based on two-view images according to any one of claims 1-9, characterized in that, Including: A broiler chicken image acquisition device (100) and a weight estimation device (200); The broiler chicken image acquisition device (100) includes: a cage body (101), a top view camera (107), a side and rear view camera (108), and an acquisition trigger device; there is a cage body opening (109) on one side of the cage body (101), and the broiler chicken to be measured enters the inside of the cage body (101) through the cage body opening (109); a partition (102) is provided inside the cage body (101), and the partition (102) divides the inside of the cage body (101) into a feed storage area (103) and a broiler chicken standing area (104); a feed container (105) is provided in the feed storage area (103); the broiler chicken standing area (104) can accommodate a single broiler chicken to stand; an eating hole (106) is provided on the partition (102); the eating hole (106) restricts the broiler chicken to be measured to only pass its head through the eating hole (106) to complete eating; the top view camera (107) is fixed on the top of the cage body (101) and is used to collect the top view image of the broiler chicken to be measured when eating; the side and rear view camera (107) is fixed outside the cage body (101), close to the side of the cage body opening (109), and is used to collect the side and rear view image of the broiler chicken to be measured when eating; the acquisition trigger device is used to trigger image acquisition and perform weight estimation; The two-view images include: the top view image and the side and rear view image of the broiler chicken to be measured when eating; The weight estimation device (200) includes: an image classification network module, an image segmentation network module, and a RepVGGVan network module; the image classification network module is used to screen the two-view image set, input the side and rear view in the two-view images into the image classification network, and output the classification result of the broiler chicken state to obtain the screened two-view images; the image segmentation network module is used to perform semantic segmentation processing on the two-view images that meet the shooting conditions to obtain the segmented two-view images; the RepVGGVan network module is used to estimate the weight of the broiler chicken by using the segmented two-view images.
Citation Information
Patent Citations
Method and system for estimating weight of meat poultry based on 2D image
CN119380375A
Cited By
Laying hen body type and weight evaluation method and system based on image recognition
CN120783396A
Broiler chicken online grading and intelligent segmentation system based on machine vision
CN121605991A