Beef cattle tachypnea behavior identification method based on inspection robot and YOWOv2-Enhanced
By embedding the improved YOLOv2 model on the inspection robot, combining the SE attention mechanism, feature fusion module and Tversky Loss loss function, the problem that traditional monitoring is difficult to identify shortness of beef cattle in real time is solved, and efficient and accurate health detection is achieved.
Patent Information
- Application Number
- CN202411969576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional manual monitoring is difficult to effectively monitor the shortness of breathing in beef cattle in real time, especially in complex breeding environments. It is inefficient and prone to missed inspections, making it difficult to meet the needs of large-scale breeding.
The beef cattle shortness of breath behavior recognition method based on patrol robots and improved YOLOv2 model (YOWOv2-Enhanced) is adopted to improve the detection accuracy and adaptability of the model through video acquisition, frame extraction processing, production of training data sets, introduction of SE attention mechanism module, feature fusion module and Tversky Loss loss function.
It realizes rapid and accurate identification of shortness of breathing behaviors of beef cattle, improves the efficiency and reliability of health testing, adapts to complex cattle farm environments, and improves detection accuracy and adaptability.
Smart Images

Figure CN120014699A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular, relates to a method for identifying the rapid breathing behavior of beef cattle based on a patrol robot and YOWOv2-Enhanced. Background Art
[0002] In modern animal husbandry, health management of beef cattle is crucial to improving production efficiency and economic benefits. However, due to the complex breeding environment, traditional manual monitoring methods often have problems such as low efficiency and easy to miss detection, which is difficult to meet the needs of large-scale breeding. Therefore, it is particularly important to develop an intelligent system that can monitor the health status of beef cattle in real time, especially for the important health indicator of beef cattle shortness of breath. Summary of the invention
[0003] The purpose of the present invention is to provide a method for identifying the rapid breathing behavior of beef cattle based on a patrol robot and YOWOv2-Enhanced. The method of the present invention can realize rapid and accurate identification of the rapid breathing behavior of beef cattle, facilitate rapid detection of the health of beef cattle, can adapt to complex cattle farm environments, and has high reliability and accuracy.
[0004] The present invention is achieved by adopting the following technical solutions: A method for identifying the rapid breathing behavior of beef cattle based on a patrol robot and YOWOv2-Enhanced comprises the following steps: S1. Video acquisition, including using a patrol robot equipped with a camera to shoot, capture beef cattle, extract abdominal features, and obtain beef cattle's rapid breathing behavior video; S2. Select a video and perform frame processing, including selecting a number of valid videos from the collected videos and extracting the videos into images; S3. Create a two-dimensional training data set to obtain training weights, including randomly selecting a number of images obtained after each video frame is extracted, annotating them, annotating a rectangular box on the abdomen of the beef cattle, and then dividing the annotated images into a data set, a validation set, and a test set in proportion, and finally using YOLOv8 training to obtain training weights; S4. Use VIA to annotate the action, including using the VIA tool to annotate the action category, using the two-dimensional training weight to obtain a rectangular box for action annotation, and finally extracting and uploading the annotated file; S5. Create a 3D training dataset, including using some python scripts to get the AVA dataset format, annotating the beef cattle ID in the VIA annotated folder, and then correcting the annotations; S6. Improve the model and train it, including introducing the SE attention mechanism module into the two-dimensional and three-dimensional backbone networks of the YOWOv2 model, adding a feature fusion module after the backbone network, and finally introducing Tversky Loss into the loss function. Use the improved YOWOv2-Enhanced model to train the beef cattle tachypnea behavior dataset, and save the weight file of the trained model. S7. Test the improved YOWOv2-Enhanced model, including embedding the trained model into the camera and using the inspection robot to test it in the cattle farm to verify the detection effect. If the detection effect does not meet the requirements, continue to improve the model.
[0005] Furthermore, in step S1, the collected videos include videos shot at different time periods, lighting conditions and angles during a day, and the videos include complex backgrounds with and without occlusions.
[0006] Furthermore, in step S4, the action categories are divided into rapid breathing and normal, and the file contains the name of the video, the number of the video frame, the coordinate value of the abdomen of the beef cattle, and the action category number.
[0007] Furthermore, in step S6, the SE attention mechanism module consists of three parts: Squeeze compression, Excitation excitation, and Scale reweighting. The principle can be explained by the following formula: , , Where: Z c Output of Squeeze compression module; F sq It is a Squeeze compression operation; H and W are the height and width of the feature image respectively; is the input data channel feature map information; s is the channel attention feature information; z is the feature map information of all channels; W net is the network parameter; W 1 is the parameter of the first layer network; W 2 is the parameter of the second layer network; F ex (z, W net ) is the Excitation operation; σ[g(z, W net )] is the feature information processed by Sigmoid function; δ( W 1z ) is the feature information processed by the ReLU function; The final output is expressed as follows: After the Scale reweighting part obtains a 1×1×C vector, it performs a Scale operation on the original feature map.
[0008] Furthermore, in step S6, the principle of the feature fusion module can be explained by the following formula: Polymerization process: , , , Output: , Where: Sc represents the aggregated spatial features, T1 and T2 represent the dual-temporal features, Avg (·) and Max (·) represent the global average pooling and global maximum pooling across spatial dimensions, respectively; W c1 , W c2 represents the bi-phase channel weight, Conv1(·) and Conv2(·) represent one-dimensional convolution; W c1 ', W c2 ' represents the output bi-phase channel weight.
[0009] Further, in step S6, the Tversky Loss principle is explained as follows: , , Where: p The label probability predicted by the model. Usually, p>0.5 is considered a positive sample, otherwise it is a negative sample. γ is the adjustment factor.
[0010] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention introduces the SE attention mechanism module into the 2D and 3D backbone networks of the YOWOv2 model. The module can adaptively recalibrate the channel feature response and enhance the model's ability to extract key features, thereby improving the accuracy of target detection. (2) The present invention adds a feature fusion module after the backbone network, which can effectively fuse features at different levels, make full use of multi-scale information, and improve the model's detection ability for complex scenes and occluded targets in beef cattle farms; (3) The present invention optimizes the loss function of the YOWOv2 model and introduces Tversky Loss as the optimized loss function, which can more flexibly balance the relationship between positive samples, negative samples and false positives, and is particularly suitable for dealing with problems of sample imbalance and overlap between classes, thereby improving the detection accuracy of the model for the rapid breathing behavior of beef cattle; (4) The present invention applies the improved network to the identification of rapid breathing behavior of beef cattle, which can further determine the health status of beef cattle and ultimately realize intelligent animal husbandry. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a flow chart of the method of the present invention; Figure 2 A labeling diagram for extracting an image from a certain frame in the present invention; Figure 3 It is a structural schematic diagram of the inspection robot in the present invention; In the picture: 1. Inspection vehicle; 2. Navigation radar; 3. ZED2 binocular depth camera; 4. Server. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the examples of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] like Figures 1 to 3 As shown, a method for identifying rapid breathing behavior of beef cattle based on a patrol robot and YOWOv2-Enhanced includes the following specific steps: S1. Video acquisition.
[0014] The inspection robot moves forward in the aisle of the beef cattle shed and uses the ZED2 binocular depth camera to capture the beef cattle and extract abdominal features to obtain a video of the beef cattle's rapid breathing behavior. The video includes videos of beef cattle shot at different times of the day, lighting, and angles. The video includes complex backgrounds with and without obstructions. All the captured videos are stored in MP4 format with a size of 1280*720 pixels.
[0015] S2. Select the video and process the frame.
[0016] Several valid videos of different scenes were selected, including videos of rapid breathing and normal behavior, as well as videos with and without occlusion. The length of each video was cut to 11 seconds. A python script was used to extract 30 frames per second from the video into images with a size of 1280*720 pixels.
[0017] S3. Create a two-dimensional training data set and obtain training weights.
[0018] We randomly selected 15 to 20 images from the images obtained after extracting frames from videos of different scenes, annotated them using the labelimg tool, marked a rectangular box on the abdomen of the beef cattle with the label "beef cattle", and then divided them into data set, validation set, and test set in a ratio of 7:2:1. We converted them into VOC format and finally used YOLOv8 training to obtain the training weights. Figure 2 Shown is a labeled diagram of a certain frame of extracted image in the present invention.
[0019] S4. Use VIA to annotate actions.
[0020] Use the VIA tool to annotate the action categories, use the two-dimensional training weights to get the rectangular box for action annotation, the action categories are rapid breathing, normal, and finally extract and upload the annotated json file. The json file contains: the name of the video, the number of the video frame, the coordinate value of the beef cattle's abdomen, and the action category number. This information is required for the annotation file, and the information in the json file needs to be integrated.
[0021] S5. Create a three-dimensional training dataset.
[0022] Use some python scripts to get the ava dataset format. Name each annotated file in the VIA annotated folder: video name_finish.json. Deepsort needs to feed 2 frames of pictures in advance before it can start annotating the person's ID from the third frame. Dense_proposals_train.pkl starts from the third frame (that is, 0 and 1 are missing), so 0 and 1 need to be added. Next, use deep sort to associate the ID of the beef cattle, then put train_personID.csv and train_without_personID.csv together, make corrections in the next step, and then run the train.py file to get train.csv, and use the improved network for training.
[0023] S6. Improve the model and train it.
[0024] (1) Introduce the SE (Squeeze-and-Excitation) attention mechanism module into the 2D and 3D backbone networks of the YOWOv2 model.
[0025] YOWOv2 is an efficient multi-level framework for real-time spatiotemporal action detection. The model uses an efficient multi-level feature fusion strategy to effectively improve the recognition efficiency and accuracy of action targets. The model combines 2D and 3D backbone networks for accurate action detection, and designs multi-level detection channels to detect action instances of different scales. In addition, a dynamic label allocation strategy and an anchor-free mechanism are introduced to further enhance the generalization and adaptability of the model.
[0026] The SE attention mechanism module is a method of introducing a channel attention mechanism. It redistributes the information between channels by learning an attention weight vector that represents the relationship between channels, thereby enhancing the network's attention to important feature channels.
[0027] The SE attention mechanism module consists of three parts: Squeeze compression, Excitation excitation, and Scale reweighting. The principle can be explained by the following formula: , , Where: Z c Output of Squeeze compression module; F sq It is a Squeeze compression operation; H and W are the height and width of the feature image respectively; is the input data channel feature map information; s is the channel attention feature information; z is the feature map information of all channels; W net is the network parameter; W 1 is the parameter of the first layer network; W 2 is the parameter of the second layer network; F ex (z, W net ) is the Excitation operation; σ[g(z, W net )] is the feature information processed by Sigmoid function; δ( W 1z ) is the feature information processed by the ReLU function.
[0028] The final output is expressed as follows: After the Scale reweighting part obtains a 1×1×C vector, it performs a Scale operation on the original feature map.
[0029] (2) A feature fusion module is added after the backbone network. The feature fusion module is an added component that is used to effectively fuse features at different levels, make full use of multi-scale information, and improve the model's ability to detect complex scenes and occluded targets in beef farms. The feature fusion module can be implemented through convolution operations, upsampling operations, and element addition.
[0030] The specific method is: create a python file in the folder with train.py, write the feature fusion module, and call it after the backbone network in train.py.
[0031] The principle of feature fusion module can be explained by the following formula: Polymerization process: , , , Output: , Where: Sc represents the aggregated spatial features, T1 and T2 represent the dual-temporal features, Avg (·) and Max (·) represent the global average pooling and global maximum pooling across spatial dimensions, respectively; W c1 , W c2 represents the bi-phase channel weight, Conv1(·) and Conv2(·) represent one-dimensional convolution; W c1 ', W c2 ' represents the output bi-phase channel weight.
[0032] (3) The loss function of the YOWOv2 model is optimized and Tversky Loss is introduced as the optimization loss function. Tversky Loss is a loss function designed to solve the problems of imbalance between positive and negative samples and imbalance between difficult and easy samples in target detection tasks. It can more flexibly balance the relationship between positive samples, negative samples and false positives. It is particularly suitable for dealing with problems of sample imbalance and inter-class overlap, thereby improving the model's detection accuracy for rapid breathing behavior of beef cattle. It adjusts the weights of easy-to-classify and difficult-to-classify samples so that the model pays more attention to difficult-to-classify samples, thereby improving the model's prediction ability.
[0033] The specific method is: create a new Python file tversky_loss.py, write the code into this file, then modify train.py, call the tversky_loss.py file, and combine it with the original loss function.
[0034] The Tversky Loss principle is explained as follows: , , Where: p The label probability predicted by the model. Usually, p>0.5 is considered a positive sample, otherwise it is a negative sample. γ is the adjustment factor.
[0035] S7. Test the improved YOWOv2-Enhanced model.
[0036] The trained RT-DETR-DMSE target detection network model is embedded in the camera and tested in the cattle farm using a patrol robot. First, the sorted test images are put into a folder named beef cattle; then the detect.py program is modified, and the best.bt file in the weights folder after training is set as the weight file of the detect.py program; finally, the detect.py program and the python script are embedded in the camera, and the patrol robot is used to move forward in the aisle of the cowshed to test the model effect. If the training results do not meet the requirements, the model will continue to be improved.
[0037] like Figure 3 As shown in the figure, the inspection robot includes an inspection vehicle 1, a navigation radar 2, a ZED2 binocular depth camera and a server 4. As an automated and intelligent monitoring tool, the inspection robot has been widely used in many fields. They have functions such as autonomous navigation, environmental perception and data collection, and can replace manual work for efficient and accurate monitoring. In the field of animal husbandry, inspection robots have been used to monitor the growth status, behavioral characteristics and disease warning of animals.
[0038] YOWOv2-Enhanced is an improved algorithm of the YOWOv2 model. As an advanced computer vision algorithm, it has high-precision and high-efficiency target detection and recognition capabilities. It can achieve real-time monitoring and early warning of beef cattle's shortness of breath behavior by extracting and classifying features of targets in the video. The application of this algorithm can not only improve the accuracy and reliability of monitoring, but also reduce the risk of missed detection and false alarms.
[0039] The present invention combines the inspection robot with the YOWOv2-Enhanced algorithm to realize intelligent monitoring of the shortness of breath behavior of beef cattle. The inspection robot can autonomously navigate to the breeding area and use sensors such as cameras to collect image or video data of beef cattle. Then, the YOWOv2-Enhanced algorithm processes and analyzes these data, identifies beef cattle with shortness of breath, and issues early warning information in real time. This method can not only improve monitoring efficiency, but also reduce the workload of breeders and improve breeding efficiency.
[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying shortness of breath in beef cattle based on inspection robots and YOWOv2-Enhanced, characterized in that: The following steps are involved: S1. Video acquisition, including using a patrol robot equipped with a camera to shoot, capture beef cattle, extract abdominal features, and obtain beef cattle's rapid breathing behavior video; S2. Select a video and perform frame processing, including selecting a number of valid videos from the collected videos and extracting the videos into images; S3. Create a two-dimensional training data set to obtain training weights, including randomly selecting a number of images obtained after each video frame is extracted, annotating them, annotating a rectangular box on the abdomen of the beef cattle, and then dividing the annotated images into a data set, a validation set, and a test set in proportion, and finally using YOLOv8 training to obtain training weights; S4. Use VIA to annotate the action, including using the VIA tool to annotate the action category, using the two-dimensional training weight to obtain a rectangular box for action annotation, and finally extracting and uploading the annotated file; S5. Create a three-dimensional training data set, including using some python scripts to get the ava data set format, annotate the beef cattle ID in the VIA annotated folder, and then correct the annotation; S6. Improve the model and train it, including introducing the SE attention mechanism module into the two-dimensional and three-dimensional backbone networks of the YOWOv2 model, adding a feature fusion module after the backbone network, and finally introducing Tversky Loss into the loss function. Use the improved YOWOv2-Enhanced model to train the beef cattle tachypnea behavior dataset, and save the weight file of the trained model. S7. Test the improved YOWOv2-Enhanced model, including embedding the trained model into the camera and using the inspection robot to test it in the cattle farm to verify the detection effect. If the detection effect does not meet the requirements, continue to improve the model.
2. The method for identifying shortness of breath behavior of beef cattle based on inspection robot and YOWOv2-Enhanced according to claim 1, characterized in that: In step S1, the collected videos include videos shot at different time periods, lighting conditions and angles during a day, and the videos include complex backgrounds with and without occlusions.
3. The method for identifying shortness of breath behavior of beef cattle based on inspection robot and YOWOv2-Enhanced according to claim 1, characterized in that: In step S4, the action categories are classified into rapid breathing and normal, and the file contains the name of the video, the number of the video frame, the coordinate value of the abdomen of the beef cattle, and the action category number.
4. The method for identifying shortness of breath behavior of beef cattle based on inspection robot and YOWOv2-Enhanced according to claim 1, characterized in that: In step S6, the SE attention mechanism module consists of three parts: Squeeze compression, Excitation excitation and Scale reweighting. The principle can be explained by the following formula: , , Where: Z c Output of Squeeze compression module; F sq It is a Squeeze compression operation; H and W are the height and width of the feature image respectively; is the input data channel feature map information; s is the channel attention feature information; z is the feature map information of all channels; W net is the network parameter; W 1 is the parameter of the first layer network; W 2 is the parameter of the second layer network; F ex (z, W net ) is the Excitation operation; σ[g(z, W net )] is the feature information processed by Sigmoid function; δ( W 1z ) is the feature information processed by the ReLU function; The final output is expressed as follows: After the Scale reweighting part obtains a 1×1×C vector, it performs a Scale operation on the original feature map.
5. The method for identifying shortness of breath behavior of beef cattle based on inspection robot and YOWOv2-Enhanced according to claim 1, characterized in that: In step S6, the principle of the feature fusion module can be explained by the following formula: Polymerization process: , , , Output: , Where: Sc represents the aggregated spatial features, T1 and T2 represent the dual-temporal features, Avg (·) and Max (·) represent the global average pooling and global maximum pooling across spatial dimensions, respectively; W c1 , W c2 represents the bi-phase channel weight, Conv1(·) and Conv2(·) represent one-dimensional convolution; W c1 ', W c2 ' represents the output bi-phase channel weight.
6. The method for identifying shortness of breath behavior of beef cattle based on a patrol robot and YOWOv2-Enhanced according to claim 1, characterized in that , In step S6, the Tversky Loss principle is explained as follows: , , Where: p The label probability predicted by the model. Usually, p>0.5 is considered a positive sample, otherwise it is a negative sample. γ is the adjustment factor.