Beef cattle salivation behavior identification method based on inspection robot and RT-DETR-DMSE
By applying the improved RT-DETR-DMSE model on the inspection robot, the salivation behavior of beef cattle is identified, and the problems of low efficiency and insufficient adaptability of beef cattle salivation behavior recognition in the prior art are solved, and rapid and accurate identification and health detection are achieved in complex environments.
Patent Information
- Application Number
- CN202411966810.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to quickly and accurately identify beef cattle salivation behavior, resulting in low efficiency in health detection of beef cattle and insufficient ability to adapt to complex cattle farm environments.
The beef cattle salivation behavior recognition method based on patrol robots and improved RT-DETR-DMSE model is adopted. Through video acquisition, frame lifting, manual annotation, data set division, model improvement and training, the rapid and accurate identification of beef cattle salivation behavior is achieved.
It realizes rapid and accurate identification of beef cattle salivation behavior in complex cattle environments, improves the efficiency and accuracy of beef cattle health testing, and has high adaptability and reliability.
Smart Images

Figure CN120014698A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of target detection, and in particular to a method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE. Background Art
[0002] In the modern agricultural system, beef cattle breeding occupies an important position. As an important part of the agricultural economy, beef cattle breeding not only provides people with a rich source of meat food and meets people's demand for high-quality animal protein, but also drives the development of related industries and promotes farmers' income. However, in the process of beef cattle breeding, salivation not only seriously affects the daily life and health of beef cattle, but also may cause serious complications such as choking, dyspnea, and aspiration pneumonia when eating and calling. For beef cattle, long-term salivation will lead to a decline in their quality of life and even affect their growth and development and slaughter performance. In addition, salivation may also cause environmental pollution, increase breeding costs, and bring economic losses to the beef cattle breeding industry. At present, the discovery of salivation behavior requires manual naked eye observation, which consumes a lot of manpower and material resources and is inefficient.
[0003] With the development of intelligent and automated animal husbandry, computer vision technology is gradually combined with animal husbandry. Target recognition in the field of computer vision is also further applied in beef cattle breeding, which can provide favorable assistance for disease diagnosis and growth monitoring of beef cattle. At present, cattle farms need to be able to accurately identify targets raised in complex environments to adapt to the complex environment of farms, quickly identify the salivation behavior of beef cattle, and ensure the health of beef cattle. To meet the needs, the present invention proposes a method for identifying salivation behavior of beef cattle based on inspection robots and RT-DETR-DMSE. Summary of the invention
[0004] The purpose of the present invention is to provide a method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE. The method can realize rapid and accurate identification of salivation behavior of beef cattle, facilitate rapid detection of health of beef cattle, adapt to complex cattle farm environments, and have high reliability and accuracy.
[0005] In order to solve the above problems, the technical solution of the present invention is: A method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE comprises the following steps: S1. Video collection: using patrol robots to collect videos of cattle behavior at different times of the day, lighting conditions, and angles; S2. Frame extraction, extracting the video into images; S3. Manual labeling: Manually label the images of beef cattle salivation behavior. According to the different behaviors of beef cattle, the labeling labels are divided into two categories: abnormal and normal. Abnormal means salivation behavior, and normal means normal behavior. S4. Create a data set by using the annotated files and the images of beef cattle salivation behavior; S5. Dataset division: divide the dataset into training set, validation set and test set; S6. Improve the model and train it, introduce the DBRA structure of the DAT module and the BRA module in the shallow layer of the RT-DETR model, integrate the EMA module in the encoder, use the improved RT-DETR-DMSE model to train the dataset of beef cattle salivation behavior, and save the weight file of the trained model; S7. Test the improved model, embed the trained model into the camera, use the inspection robot to test it in the cattle farm, and verify the detection effect. If the detection effect does not meet the requirements, continue to improve the model.
[0006] Furthermore, step S3 also includes: Normalize the annotation box: , , , , In the formula , , , They represent the normalized center coordinates of the labeled rectangle and the width and height of the labeled rectangle. , , , are the coordinates of the upper left corner and lower right corner of the manually marked rectangular box, , are the width and height of the image respectively.
[0007] Furthermore, in step S6, the RT-DETR model includes a backbone network, a neck and a decoder, wherein the backbone network adopts a ResNet convolutional neural network to extract feature representations of the input image; the neck network includes intra-scale feature interaction (AIFI) and a cross-scale feature fusion module (CCFM) to further process and fuse the features extracted by the backbone network; the decoder is constructed using a multi-layer Transformer decoder layer, each layer of which includes a self-attention mechanism and a feedforward neural network system component to generate the final detection result based on the features provided by the neck network.
[0008] Furthermore, in step S6, the principle of the DAT module can be expressed by the following formula: , , , Where: q represents a vector, x represents a scalar, , denotes key and value embeddings respectively, represents a specific linear transformation, W k represents the output vector of keys, W v An output vector representing the values, P Indicates the parameter quantity, ∆p Indicates the parameter increment, θ offset (q) represents a function that accepts an input q and outputs ∆p, ϕ (⋅;⋅) represents bilinear interpolation, making it differentiable, z represents input data, r represents coordinates; The final output is expressed as: , Where: z (m) represents the mth data feature, q (m) represents the mth vector, represents the mth key embedding, represents the mth value embedding, d represents the dimension of the key vector, B Indicates location information. R Represents features that control position encoding.
[0009] Further, in step S6, the principle of the BRA module is explained as follows: first, the input feature map is converted into a feature vector, the feature vector is divided into small feature map areas through reshape deformation, and then these feature map areas are linearly transformed to obtain three small feature vectors, and the features to be paid attention to are further obtained.
[0010] Further, in step S6, the DBRA structure consists of two branches, one branch uses DWConv to generate feature maps X 1. In another branch, DBRA aims to reduce feature maps through DWConv and DAT modules to improve computational efficiency while maintaining feature details, thereby obtaining X2. Then, DAT is used to further convert the spatial information into channel information to facilitate the fusion of spatial and depth features. After that, BRA is used to identify key feature areas and supplement global information to obtain X 3. Finally, the DBRA fusion component will X 1. X 2 and X 3 Generate feature maps using connected components (CAT) along the channel dimension X '.
[0011] Furthermore, in step S6, the principle of the EMA module can be expressed by the following formula: , Among them: Q represents the query vector, which is used to specify the content to focus on, K represents the key vector, which is used to calculate the similarity between the query vector and other vectors, and V represents the value vector, which is used to weight according to the attention weight.
[0012] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention introduces a deep fusion module (DBRA) in the shallow layer of the backbone to replace the downsampling layer DWConv of the original model, which can convert the spatial features of the image into deep features and realize the deep interaction between the local features and the global features of the image; (2) By integrating the EMA module into the encoder, the present invention can further enhance the sub-attention learning of high-level semantic features and fully connect the context to achieve the representation of deep feature mapping; (3) The present invention applies the improved network to the identification of salivation behavior of beef cattle, and accurately identifies the breeding targets in complex environments to adapt to the complex environment of the farm. It can quickly and accurately identify the salivation behavior of beef cattle, further judge the health status of beef cattle, ensure the health of beef cattle, and ultimately realize intelligent livestock breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 Flow chart of the method of the present invention.
[0014] Figure 2 This is a network framework diagram of the present invention.
[0015] Figure 3 It is a structural schematic diagram of the inspection robot in the present invention.
[0016] In the picture: 1. Inspection vehicle; 2. Navigation radar; 3. ZED2 binocular depth camera; 4. Server. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the examples of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] like Figures 1 to 3 As shown in FIG. 1 , a method for identifying salivation behavior of beef cattle based on RT-DETR-DMSE is implemented according to the following specific steps: The first step is video acquisition.
[0019] A patrol robot was used to move forward in the aisle of the beef cattle shed. The ZED2 binocular depth camera was used to shoot in the aisle of the beef cattle shed at different times of the day, lighting and angles, capture the beef cattle, extract the mouth features, and obtain the salivation behavior video of the beef cattle. The video contained complex backgrounds with and without obstructions. All the video formats were stored in MP4 format with a size of 1280*720 pixels.
[0020] The second step is frame processing.
[0021] The captured video is processed by frame extraction; a python script is used to extract 30 frames per second from the video into images; the size is 1280*720 pixels.
[0022] The third step is manual labeling.
[0023] The frame-edited images were manually annotated using Labelimg. According to the different behaviors of beef cattle, the annotation labels were divided into two categories: abnormal and normal. Abnormal means salivation behavior, and normal means normal behavior. A total of 1,000 images were annotated.
[0024] In order to improve training efficiency and speed up model training, normalize the annotation boxes: , , , , In the formula , , , They represent the normalized center coordinates of the labeled rectangle and the width and height of the labeled rectangle. , , , are the coordinates of the upper left corner and lower right corner of the manually marked rectangular box, , are the width and height of the image respectively.
[0025] The fourth step is to create a data set.
[0026] Create folders according to the format of the VOC dataset, and put the beef cattle salivation behavior images and the corresponding annotation xml files into folders named images_beef cattle and Annotation_beef cattle, respectively.
[0027] The fifth step is data set division.
[0028] Use a python script to divide the dataset into training set, test set and validation set in a ratio of 7:2:1, save the divided dataset image names in the ImageSets_beef cattle folder in the format of a txt file, and store the paths corresponding to the images of the training set, test set and validation set in the data directory in the form of a yaml file named beef cattle.yaml, so that the image and the image target box can be one-to-one correspondence. Finally, use another Python script to convert the xml file into a txt file that can be recognized by the RT-DETR-DMSE model, and save the txt file in the Labels folder.
[0029] In steps 6 and 7, the model is improved and trained and tested.
[0030] (1) Create a new yaml file named RT-DETR-DMSE.yaml, copy the code in the original network to RT-DETR-DMSE.yaml, and introduce a deep fusion module (DBRA) in the shallow layer of the RT-DETR model to replace the downsampling layer DWConv of the original model. The DBRA module includes DAT and BRA modules.
[0031] The RT-DETR model consists of a backbone network, a neck, and a decoder. The backbone network uses a ResNet convolutional neural network to extract feature representations of the input image. The neck network contains intra-scale feature interaction (AIFI) and a cross-scale feature fusion module (CCFM) to further process and fuse the features extracted by the backbone network. The decoder is constructed using a multi-layer Transformer decoder layer, each of which contains a self-attention mechanism and a feedforward neural network system component to generate the final detection result based on the features provided by the neck network.
[0032] The principle of the DAT module can be expressed by the following formula: , , , Where: q represents a vector, x represents a scalar, , denotes key and value embeddings respectively, represents a specific linear transformation, W k represents the output vector of keys, W v An output vector representing the values, P Indicates the parameter quantity, ∆p Indicates the parameter increment, θ offset (q) represents a function that accepts an input q and outputs ∆p, ϕ (⋅;⋅) represents bilinear interpolation, making it differentiable, z represents input data, r represents coordinates; The final output is expressed as: , Where: z (m) represents the mth data feature, q (m) represents the mth vector, represents the mth key embedding, represents the mth value embedding, d represents the dimension of the key vector, B Indicates location information. R Represents features that control position encoding.
[0033] The principle of the BRA module is explained as follows: first, the input feature map is converted into a feature vector, the feature vector is divided into small feature map areas through reshape deformation, and then these feature map areas are linearly transformed to obtain three small feature vectors, and further obtain the features to be paid attention to.
[0034] The DBRA structure principle is explained as follows: The DBRA fusion module consists of two branches, one branch uses DWConv to generate feature maps X 1. In another branch, DBRA aims to reduce feature maps through DWConv and DAT modules to improve computational efficiency while maintaining feature details, thereby obtaining X 2. Then, DAT is used to further convert the spatial information into channel information to facilitate the fusion of spatial and depth features. After that, BRA is used to identify key feature areas and supplement global information to obtain X 3. Finally, the DBRA fusion component will X 1. X 2 and X3 Generate feature maps using connected components (CAT) along the channel dimension X '.
[0035] (2) Integrate the EMA module into the encoder of the network. Create a new file named EMA.py in the models file, copy the model code into the EMA.py file and save it. Import EMA.py at the beginning of the YOLO.py file, add the EMA channel parameter in parse_model.py, and finally add the EMA attention mechanism module to the encoder of RT-DETR-DMSE.yaml. EMA can retain the information of each channel while reducing the amount of calculation, fully connect the context information, make the spatial semantic features well distributed in each feature group, reduce other interference, and aggregate the output features of parallel branches through cross-dimensional interactions, fuse the context information of different scales, and then capture the pairwise relationship at the pixel level.
[0036] The principle of the attention mechanism EMA module can be expressed by the following formula: , Where: Q represents the query vector, which is used to specify what to focus on, K represents the key vector, which is used to calculate the similarity between the query vector and other vectors, and V represents the value vector, which is used to weight according to the attention weight.
[0037] (3) Replace the yaml file in the cfg module in train with the RT-DETR-DMSE.yaml file, and replace the yaml file in the data module with beef cattle.yaml. Set eppochs to 300 and batch-size to 16, and train the improved network. Save the weight file and global average precision (map) after the model training is completed. Compared with the training results of the original model, the size of the training weight file is reduced from 15.1mb to 11.2mb, and the map value is increased from 0.856 to 0.923.
[0038] (4) The trained RT-DETR-DMSE target detection network model was embedded in the camera and tested in the cattle farm using a patrol robot. First, the sorted test images were placed in a folder named beef cattle. Then, the detect.py program was modified, and the best.bt file in the weights folder after training was set as the weight file of the detect.py program. Finally, the detect.py program and the python script were embedded in the camera, and the patrol robot was used to move forward in the aisle of the cattle house to test the model effect. If the training results did not meet the requirements, the model was improved.
[0039] like Figure 3 As shown in the figure, the inspection robot includes an inspection vehicle 1, a navigation radar 2, a ZED2 binocular depth camera and a server 4. As an automated and intelligent monitoring tool, the inspection robot has been widely used in many fields. They have functions such as autonomous navigation, environmental perception and data collection, and can replace manual work for efficient and accurate monitoring. In the field of animal husbandry, inspection robots have been used to monitor the growth status, behavioral characteristics and disease warning of animals.
[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying salivation behavior of beef cattle based on inspection robots and RT-DETR-DMSE, characterized in that: The following steps are involved: S1. Video collection, including using patrol robots to collect videos of cattle behavior at different times of the day, lighting conditions, and angles; S2. Frame extraction, including extracting the video into images; S3. Manual labeling, including manual labeling of images of salivation behavior of beef cattle. According to different behaviors of beef cattle, the labeling labels are divided into two categories: abnormal and normal. Abnormal means salivation behavior, and normal means normal behavior. S4. Create a data set, including using the annotated files and the beef cattle salivation behavior images to create a data set; S5. Dataset division, including dividing the data set into training set, validation set and test set; S6. Improve the model and train it, including introducing a DBRA structure fused with a DAT module and a BRA module in the shallow layer of the RT-DETR model, fusing an EMA module in the encoder, using the improved RT-DETR-DMSE model to train a dataset of salivation behavior of beef cattle, and saving a weight file of the trained model; S7. Test the improved model, embed the trained model into the camera, use the inspection robot to test it in the cattle farm, and verify the detection effect. If the detection effect does not meet the requirements, continue to improve the model.
2. The method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE according to claim 1, characterized in that: Step S3 also includes: Normalize the annotation box: , , , , In the formula , , , They represent the normalized center coordinates of the labeled rectangle and the width and height of the labeled rectangle. , , , are the coordinates of the upper left corner and lower right corner of the manually marked rectangular box, , are the width and height of the image respectively.
3. The method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE according to claim 1, characterized in that: In step S6, the RT-DETR model includes a backbone network, a neck and a decoder, wherein the backbone network uses a ResNet convolutional neural network to extract feature representations of the input image; the neck network contains intra-scale feature interaction (AIFI) and a cross-scale feature fusion module (CCFM) to further process and fuse the features extracted by the backbone network; the decoder is constructed using a multi-layer Transformer decoder layer, each layer of which contains a self-attention mechanism and a feedforward neural network system component to generate the final detection result based on the features provided by the neck network.
4. The method for identifying salivation behavior of beef cattle based on a patrol robot and RT-DETR-DMSE according to claim 1, characterized in that: In step S6, the principle of the DAT module can be expressed by the following formula: , , , Where: q represents a vector, x represents a scalar, , denotes key and value embeddings respectively, represents a specific linear transformation, W k represents the output vector of keys, W v An output vector representing the values, P Indicates the parameter quantity, ∆p Indicates the parameter increment, θ offset (q) represents a function that accepts an input q and outputs ∆p, ϕ (⋅;⋅) represents bilinear interpolation, making it differentiable, z represents input data, r represents coordinates; The final output is expressed as: , Where: z (m) represents the mth data feature, q (m) represents the mth vector, represents the mth key embedding, represents the mth value embedding, d represents the dimension of the key vector, B Indicates location information. R Represents features that control position encoding.
5. The method for identifying salivation behavior of beef cattle based on inspection robot and RT-DETR-DMSE according to claim 1, characterized in that: In step S6, the principle of the BRA module is explained as follows: first, the input feature map is converted into a feature vector, the feature vector is divided into small feature map areas through reshape deformation, and then these feature map areas are linearly transformed to obtain three small feature vectors, and further obtain the features to be paid attention to.
6. The method for identifying salivation behavior of beef cattle based on inspection robot and RT-DETR-DMSE according to claim 1, characterized in that: In step S6, the DBRA structure consists of two branches, one of which uses DWConv to generate feature maps X 1. In another branch, DBRA aims to reduce feature maps through DWConv and DAT modules to improve computational efficiency while maintaining feature details, thereby obtaining X 2. Then, DAT is used to further convert the spatial information into channel information to facilitate the fusion of spatial and depth features. After that, BRA is used to identify key feature areas and supplement global information to obtain X 3. Finally, the DBRA fusion component will X 1. X 2 and X 3 Generate feature maps using connected components (CAT) along the channel dimension X '.
7. The method for identifying salivation behavior of beef cattle based on inspection robot and RT-DETR-DMSE according to claim 1, characterized in that: In step S6, the principle of the EMA module can be expressed by the following formula: , Among them: Q represents the query vector, which is used to specify the content to focus on, K represents the key vector, which is used to calculate the similarity between the query vector and other vectors, and V represents the value vector, which is used to weight according to the attention weight.
Citation Information
Cited By
Method and device for identifying ground layer crack image of synthetic material playground and medium
CN121639678A
Method, device and medium for identifying image of surface layer crack of synthetic material sports ground
CN121639678B