Multi-mode beef cattle behavior identification method and system
By improving the YOLOv8 algorithm and combining multimodal analysis, the problem of identifying similar posture behaviors of beef cattle in complex environments is solved, and higher recognition accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510656882.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The prior art is difficult to accurately identify similar posture behaviors of beef cattle in complex environments, and computer vision technology relies on a single visual information, making it easy to cause misjudgment or misjudgment.
The multimodal beef cattle behavior recognition method is adopted, and by improving the YOLOv8 algorithm, combining the channel reduction uneven grouping method, the C-NBottleneck module and the conical multi-scale dimensionality reduction feature extraction algorithm, the visual recognition ability is enhanced, and motion sensing and position positioning analysis are combined.
It significantly improves the recognition accuracy of similar posture behaviors of cattle in complex environments, ensuring that stable recognition performance can be maintained under conditions such as light changes and partial occlusion.
Smart Images

Figure CN120220249A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a multi-modal beef cattle behavior recognition method and system. Background Art
[0002] Traditional beef cattle breeding management methods rely on breeders to regularly patrol the cattle shed and visually observe the appearance, feeding situation, behavior performance, etc. of beef cattle to judge the health status and growth state of beef cattle. This method has many drawbacks. It is inefficient and consumes a large amount of human and time costs. Moreover, it is difficult to achieve real-time and continuous monitoring through manual observation, and it is easy to miss something, unable to promptly capture the abnormal behaviors or sudden health problems of beef cattle and take corresponding measures.
[0003] With the development of computer vision technology, YOLOv8, as an advanced object detection algorithm, can efficiently process image or video data in the beef cattle breeding scenario, identify individual beef cattle, their postures, and various behavior actions. This provides an intuitive and accurate basis for breeders to timely master the daily activity rules and health status of beef cattle, and helps to achieve refined breeding management. However, YOLOv8 introduces more technical means and network structures, such as feature pyramid networks, attention mechanisms, etc., resulting in a relatively large model size. This means that more storage space is required to save model parameters, and during model loading and inference, it will occupy more memory resources. For resource-constrained devices, it may bring certain operating pressure, and the detection effect for details is relatively poor, prone to missed detection or false detection. Moreover, computer vision technology only relies on visual information, and in some complex situations, it may make misjudgments or be unable to comprehensively and accurately evaluate the true state of beef cattle due to the singularity of information.
[0004] Therefore, currently, in the beef cattle breeding environment, the internal scene of the cattle shed is complex with many interference factors. The lighting conditions in the cattle shed are variable, ranging from bright lighting areas to dark corners, and the light intensity varies greatly at different times of the day. There are a large number of facilities and equipment in the cattle shed, such as railings, feeding troughs, water troughs, etc. These objects may form occlusions or confusions with beef cattle visually, posing a significant challenge to the recognition accuracy of YOLOv8, increasing the difficulty of object detection, and prone to misjudgments or missed detections, resulting in a series of wrong measures and causing greater cost losses.
[0005] Chinese Patent Application, Application No. CN202411190892.6, Publication Date December 6, 2024, discloses a new method and system for bovine digital twin behavior perception modeling based on the fusion of sensor data and video data. Step 1: Obtain the sensor data of the bovine and the video data of the bovine; Step 2: Based on the data of the bovine obtained in Step 1, establish a data set, and unify the time lengths of the sensor data of the bovine and the video data of the bovine in Step 1 when the start times are the same; Step 3: Perform data fusion based on the sensor data of the bovine and the video data of the bovine with unified time lengths in Step 2; Step 4: Train the data fused in Step 3; Step 5: Evaluate the indicators of the data trained in Step 4. However, this solution only uses video data for visual analysis, and the recognition accuracy is limited under complex lighting and occlusion conditions. Summary of the Invention
[0006] Aiming at the problem that a single visual model in the prior art is difficult to adapt to the recognition of similar posture behaviors of bovines in complex environments, this application provides a multi-modal beef cattle behavior recognition method and system, improves the YOLOv8 algorithm, reconstructs the backbone network using the channel reduction unequal grouping method, combines the C-NBottleneck module and the conical multi-scale dimensionality reduction feature extraction algorithm, strengthens the visual recognition ability of bovine behavior characteristics, and combines motion sensing analysis and position location analysis, etc., to improve the recognition accuracy of bovine behavior types.
[0007] One aspect of the embodiments of this specification provides a method for multi-modal management of bovines, including: using an improved YOLOv8 object detection algorithm to perform visual analysis on bovines to obtain behaviors of bovines including eating, lying down, standing, moving, and estrus behaviors as the first behavior; collecting bovine motion data and identifying walking or running behaviors through feature analysis as the second behavior; obtaining bovine position data and identifying behaviors of bovines such as forage, drinking water, or standing still by comparing the bovine positions with the coordinates of preset functional areas as the third behavior; fusing the first behavior, the second behavior, and the third behavior to obtain the final behavior type of the bovine; adjusting the cattle shed environmental parameters according to the final behavior type of the bovine; the cattle shed environmental parameters include lighting brightness, fan operation parameters, and ultraviolet disinfection parameters.
[0008] Furthermore, the improved YOLOv8 object detection algorithm is used for visual analysis of cattle, including: optimizing the backbone network of the YOLOv8 object detection algorithm by using the method of unequal grouping of channel reduction; constructing a C-NBottleneck network according to the optimized backbone network, and using the C-NBottleneck network to replace the Bottleneck component in the C2F module of the backbone network of the YOLOv8 object detection algorithm; adding a conical multi-scale dimensionality reduction feature extraction algorithm to the replaced YOLOv8 object detection algorithm, and obtaining an enhanced feature map through a channel attention mechanism and a spatial attention mechanism; obtaining the final improved YOLOv8 object detection algorithm; using the final improved YOLOv8 object detection algorithm to analyze cattle images and identify cattle behaviors.
[0009] Furthermore, the backbone network of the YOLOv8 object detection algorithm is optimized by using the method of unequal grouping of channel reduction, including: defining an unequal grouping convolution operation for channel reduction; using the defined unequal grouping convolution operation for channel reduction to replace all convolution operations in the backbone network except the first layer of convolution, and obtaining the optimized backbone network; Furthermore, defining an unequal grouping convolution operation for channel reduction includes: performing a preliminary convolution operation on the input feature map to obtain an intermediate feature map; dividing the channel dimension of the intermediate feature map into n groups. When the channel dimension cannot be divided evenly by n, the number of channels in the first n - 1 groups is set to be equal, and the last group contains the remaining channels; dividing the convolution kernels into n groups according to the same grouping strategy as the intermediate feature map; performing convolution operations on each group of feature maps using the corresponding group of convolution kernels; splicing the convolution results of each group with the intermediate feature map along the channel dimension to obtain a convolution output feature map.
[0010] Furthermore, according to the optimized backbone network, a C-NBottleneck network is constructed, including: obtaining the feature map output by the optimized backbone network as the first feature map; performing a non-linear transformation on the first feature map using the SiLU activation function; performing a convolution process on the non-linearly transformed first feature map using the optimized backbone network to obtain a second feature map; performing a normalization process on the second feature map; performing a residual connection between the normalized second feature map and the first feature map to generate a C-NBottleneck network.
[0011] Furthermore, a conical multi-scale dimensionality reduction feature extraction algorithm is added to the replaced YOLOv8 object detection algorithm. Through the channel attention mechanism and the spatial attention mechanism, an enhanced feature map is obtained, including: obtaining the feature map F output by the replaced YOLOv8 object detection algorithm; according to the feature map F, through channel reduction unequal grouped convolution based on the channel attention mechanism, obtaining a channel weighted feature map; according to the channel weighted feature map, through channel reduction unequal grouped convolution based on the spatial attention mechanism, obtaining a spatial weighted feature map; according to the spatial weighted feature map and the feature map F, through a depthwise separable convolution operation, obtaining an enhanced feature map.
[0012] Furthermore, according to the feature map F, through channel reduction unequal grouped convolution based on the channel attention mechanism, obtaining a channel weighted feature map, including: performing convolution processing on the feature map F using a channel reduction unequal grouping method with a convolution kernel size of N1*N1 to obtain an intermediate feature map; performing global average pooling operation on the intermediate feature map to obtain a channel descriptor; performing a non-linear transformation on the channel descriptor through a fully connected layer to obtain a channel weight; according to the channel weight and the feature map F, obtaining a channel weighted feature map.
[0013] Furthermore, according to the channel weighted feature map, through channel reduction unequal grouped convolution based on the spatial attention mechanism, obtaining a spatial weighted feature map, including: performing convolution processing on the channel weighted feature map using a channel reduction unequal grouping method with a convolution kernel size of N2*N2 to obtain a spatial attention feature map; performing global average pooling operation on the spatial attention feature map to obtain a spatial feature descriptor; performing a non-linear transformation on the spatial feature descriptor through a fully connected layer to obtain a spatial attention weight; According to the spatial attention weight and the channel weighted feature map, obtaining a spatial weighted feature map.
[0014] Another aspect of the embodiments of this specification also provides a multi-modal beef cattle behavior recognition system, which collects cattle movement data and identifies walking or running behaviors through feature analysis as the second behavior, including: obtaining the acceleration data of cattle using an accelerometer; collecting the magnetic field intensity data of cattle using a magnetometer; processing the acceleration data using the fast Fourier transform to obtain frequency domain features; performing time domain analysis and spectrum analysis on the magnetic field intensity data to obtain magnetic field features; using the locally linear embedding algorithm to perform non-linear dimensionality reduction and fusion on the frequency domain features and the magnetic field features to obtain a fused feature vector; using a support vector machine (SVM) to classify the fused feature vector to obtain the second behavior.
[0015] Further, obtain the cattle position data. By comparing the cattle position with the coordinates of the preset functional areas, identify the cattle's forage, drinking water, or stationary behaviors as the third type of behavior, including: arranging multiple UWB base stations in the cattle shed, equipping each cow with a UWB tag; using the two-way time-of-flight algorithm strategy to measure the distance from the UWB tag to the UWB base station; calculating the two-dimensional coordinates of the cattle using triangulation based on the distance data from at least three base stations as the cattle position data; and determining whether the cattle are in the preset functional areas according to the cattle position data and the division of the cattle shed functional areas, where the preset functional areas include the forage area and the drinking water area.
[0016] Compared with the prior art, the advantages of this application are as follows: The improved YOLOv8 algorithm enhances the ability to distinguish between similar posture behaviors such as eating and standing, lying and resting, etc. through uneven channel reduction grouping and the C-NBottleneck module. At the same time, the conical multi-scale dimensionality reduction feature extraction algorithm combined with the attention mechanism strengthens the model's ability to extract the core behavior features of cattle, enabling the system to maintain stable recognition performance under complex environmental conditions such as light changes and partial occlusion. In addition, visual analysis provides morphological features but is greatly affected by the environment, motion sensing provides dynamic features but it is difficult to distinguish static behavior details, and position positioning provides spatial semantics but lacks posture information. The three modalities work together to achieve information complementarity and cross-verification. Therefore, this application significantly improves the recognition accuracy of similar posture behaviors of cattle in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the overall structure of the multi-modal cattle management system of this application; Figure 2 It is a flowchart of implementing the uneven channel reduction grouping algorithm of this application; Figure 3 It is a flowchart of implementing the C-NBottleneck of this application; Figure 4 It is a flowchart of implementing the C-N-C2f of this application; Figure 5 It is a flowchart of implementing the conical multi-scale dimensionality reduction feature extraction of this application; Figure 6 IoU is used to measure the overlap degree between the predicted box and the ground truth box in the object detection task Figure 7 It is a structural diagram of implementing YOLOv8 of this application; Figure 8 It is the five-level bulb brightness of this application; Figure 9 It is a flowchart of generating measures according to behaviors of this application; Figure 10Model diagram for generating measures based on behavior in this application. Detailed implementation
[0018] The following describes this application in detail in combination with the accompanying drawings of the specification and specific embodiments.
[0019] Example 1 As Figure 1 shown, collect images of cattle, label the behaviors of the collected cattle, and the labeled behaviors include eating, lying down, standing, moving, and estrus. Also label the data set, validation set, and test set. Distinguish and label the identity of the cattle through ear tags, train using the improved YOLOv8 algorithm, and identify the behaviors.
[0020] Place an accelerometer on the leg of the cattle and a magnetometer on the back, and through manual observation, repeatedly record the parameters of the accelerometer and magnetometer of the cattle under the behaviors of eating, lying down, standing, walking, running, and estrus respectively. After normalization, perform data processing to obtain the non-linear feature vectors of the accelerometer and magnetometer respectively, and use a non-linear dimensionality reduction method to perform feature-level fusion and training, and identify the corresponding behaviors.
[0021] Use ultra-wideband technology to locate the position of the cattle, and analyze the behaviors of the cattle such as walking, running, drinking water, eating forage, and standing still according to the location.
[0022] Perform decision-level fusion on the obtained behaviors using the majority voting decision fusion method, and adjust the brightness of the light bulb, the fan, the ultraviolet ray, and send a warning to the farmer according to the accurately identified behaviors of the cattle.
[0023] In order to extract and fuse the spatial and channel information of the input original feature map, a standard convolution operation of 3x3 is usually used to obtain the output feature map The formula: , where F' represents the output feature map; represents the input feature map, is the set of real numbers, indicating that the elements of the feature map F are all real numbers, is the number of input channels, represents the width and height of the input feature map before the convolution operation, represents the width and height of the output feature map after the convolution operation, is the number of output channels, , represents the set of learned filter kernels, where represents the filter kernels corresponding to the 1st to the c-th channel numbers, , is the size of the convolution kernel, is the bias term, represents performing the convolution operation. And the number of floating-point operations required by F ( ) The calculation formula is: ; ; where, is the number of input channels, represents the width and height of the input feature map before the convolution operation, is the size of the convolution kernel, is the padding, is the stride, represents the width and height of the output feature map after the convolution operation, is the number of output channels.
[0024] It can be seen that especially when the image pixels are greater than 256x256, the number of floating-point operations is large, which means consuming a large amount of computing power and time. It can also be seen that the number of optimized parameters is obviously determined by the input dimension and the output feature map. The number of output feature maps of the convolutional layer often contains a lot of redundancy, and some are very similar. Therefore, it is unnecessary to use a large number of FLOPs and parameters to generate redundant feature maps one by one.
[0025] Therefore, as Figure 2 shown, the method of using channel dimension reduction and non-uniform grouping (CDR–NUG) is used to optimize the ordinary convolution operation, and the formula is derived: , where, , represents the set of learned filter kernels, where, represents the filter kernels corresponding to the 1st to the channel numbers, , , by default, n = 4 is taken, and it is necessary to ensure that n can divide . And in order to simplify the notation and improve the model speed, the bias term is omitted, and is obtained. To solve the problem of fewer output feature maps, then channels are divided into groups, and the feature maps of the gth group are denoted as . If the divided groups cannot be divided evenly, non-uniform grouping is used in the last group. The method is: the number of channels in each of the first n - 1 groups is , and the number of channels in the last group is , where / / represents integer division and % represents taking the remainder. Where is the feature map of the mth group, , with a size of , and its corresponding convolution kernel , and the number of input channels at this time is , the number of output channels is , is also divided into n groups. The last group uses unequal grouping, and the unequal grouping method is the same as above, denoted as , is the convolution kernel of the m-th group, , with a size of . Perform convolution operations separately, with a stride of , and padding , then it can be ensured that: , , to ensure that the width w and height h obtained before and after the convolution operation are equal.
[0026] In the YOLOv8 structure, a 3x3 convolution kernel is used, so p = 1. At this time, the output feature map of each group , , is the feature map of the m-th group, is the feature map of the m-th group after pooling; then, taking the channel dimension of the feature map obtained from the first-step standard convolution as the starting part, and then stacking the feature maps after the convolution operation of each group after grouping along the channel dimension in sequence to complete the splicing operation. In this way, the final output feature map obtained after splicing has a width and height of and , respectively, and the number of channels is , which is the same as the original number of output channels.
[0027] The computational complexity ratio can be verified using the formula . It can be seen that, except in the case where the number of output channels >> the number of input channels, the optimized computational complexity is significantly more efficient than ordinary convolution operations: .
[0028] As Figure 3 shows, the feature map is processed using the CDR–NUG improved convolution operation to extract features. To reduce the problem of gradient vanishing or explosion and to make the network more robust to different initializations, BN normalization is performed on it. To enable the network to learn and fit more complex data patterns, it passes through the SiLU activation function, with the expression , where is the Sigmoid activation function, x represents the input feature tensor, and then the CDR–NUG improved convolution operation, as well as BN normalization, is used and added to the feature map to obtain C-NBottleneck.
[0029] As Figure 4As shown, all Bottleneck components in the original network's C2f module can be replaced by the C-N-C2f module, and this structure adopts the method of cross-stage local connection.
[0030] As Figure 5 shown, the feature information of the detection images of cattle behavior is enhanced through the Pyramid-shaped Multi-scale Dimensionality Reduction Feature Extraction (PMDR) algorithm.
[0031] First, a convolution operation that processes the feature map F using the method of unequal grouped channel reduction is performed. A 3x3 convolution kernel is used, and the number of channels remains unchanged for feature extraction, reducing the number of parameters. It is denoted as . Among them, CN is the CDR–NUG method, 3x3 is the convolution kernel size, and C is the number of channels.
[0032] And in order to stabilize the data distribution and accelerate the training convergence, the Batch Normalization (BN) operation is added. The calculation formula of BN is: ; , where is the scaling factor of the c-th channel, which can restore the feature expression ability. is the offset factor of the c-th channel, which can adjust the feature offset. is the value at the n-th sample, c-th channel, (i, j) position of is the process of normalizing the value at the n-th sample, c-th channel, (i, j) position of is the mean of the c-th channel, is the variance of the c-th channel, is a very small constant to prevent the denominator from being zero. So the overall formula is: , where is the scaling factor, is the offset factor. is the result of the normalization step. And the activation function is used to obtain , and its formula is: , where x represents the elements in
[0033] makes the calculation efficient and enhances the sparsity and interpretability of the model. Its expression is: , in order to indicate the C of this layer The numerical distribution of weights of C feature maps in each channel of are respectively subjected to global average pooling operations and become , and its formula is: , where is a 1×1×c feature vector, is the output feature map in the two-dimensional matrix at the i-th row, j-th column, and c-th channel, is each channel of are respectively subjected to global average pooling operations.
[0034] Then, the spatial information is compressed into channel descriptors to obtain a global receptive field. At the same time, the computational complexity is reduced. A fully connected layer (referred to as the FC layer) is used to compress the vector to . After experimental verification, r = 16 is the best. Then, the activation function is used to perform a non-linear transformation on the dimension-reduced result and avoid the problem of gradient disappearance, enabling the model to converge faster. Then, a layer Linear(C / r, C) is used to restore the features to C channels. Since the channels of the restored features need to be multiplied channel by channel with the channels of the original feature map to achieve recalibration of the feature map, the obtained channel weights are in the range of 0-1. Therefore, the activation function Sigmoid is used, and its formula is , where x represents the value of each element on the feature map. The formula for this step is: , where is an intermediate function, representing the operations of applying the fully connected layer and the activation function (ReLU). is the weight matrix of the fully connected layer, is used to map from the c-dimensional space to the c / r-dimensional space, is used to map it back from the c / r-dimensional space to the c-dimensional space.
[0035] Then, starting from the first channel (channel index c = 0) in the feature map F, it is traversed in order to the last channel (channel index ). For the current channel, all elements of channel c in the original feature map F are considered. This is a two-dimensional sub-tensor with a shape of , denoted as (where ranges from 0 to , ranges from 0 to ). Meanwhile, obtain the channel weight vector of the element of channel c in , since is shaped like , so is . Multiply each element of channel c in the original feature map F by the channel weight to obtain a new two-dimensional sub-tensor , and its calculation formula is: . The overall formula is denoted as: , this operation can enhance or weaken the features of specific channels. If the value of an element in the channel weight tensor is large (close to 1), then the corresponding channel features will be enhanced; if the element value is small (close to 0), the corresponding channel features will be weakened. This allows the network to pay more attention to the channel features that are more important for the task, thereby improving the performance of the model. To reduce the computational amount and speed up the operation efficiency, perform a dimensionality reduction operation on . As can be seen from the above CDR–NUG improved convolution, when the number of output channels is less than the number of channels, the computational amount of this convolution operation drops significantly. Therefore, perform a convolution operation on the obtained result using the channel reduction unequal grouping method, and perform the operation with a convolution kernel of 7x7, and reduce its number of channels to to obtain . Among them is an integer and can divide the number of channels. By default is 4. And use BN batch normalization and ReLU activation function. Then, perform an average pooling operation on the channels to generate a tensor of size .
[0036] Meanwhile, pass the feature map obtained by convolution through the Sigmoid activation function to limit its value between 0 and 1 to obtain the spatial attention feature map , and its specific expression is: , where is the average pooling of the channels.
[0037] At each channel c of , if the eigenvalue at the spatial position is , and the spatial position weight in is , then the eigenvalue of the multiplied result at this channel and position becomes: .
[0038] The overall formula is: , if the weight of a certain spatial position is close to 1, then the features at this position remain basically unchanged after multiplication because it is considered an important spatial position; if the weight is close to 0, then the features at this position will be greatly weakened after multiplication, thus achieving the effect of suppressing unimportant spatial positions.
[0039] Finally, for , perform a dimensionality increase operation. Using the depthwise separable convolution operation can reduce the computational amount. The convolution kernel size is 7x7, the number of output channels is C, and batch normalization is used. Use the Sigmoid activation function again to obtain the weight of this module, and its expression is: .
[0040] Finally, for each element at position , multiply it with the corresponding element in : , where is the position of the height, width, channel, and batch of the feature map, and the final output result completes the spatial attention operation, and its expression is: .
[0041] Its complete expression is: ; where DSC is the depthwise separable convolution operation method.
[0042] As Figure 6 shown, this is used to measure the overlap degree between the predicted box and the ground truth box in the object detection task, and its calculation formula is: ; its symbols are as shown in the figure. However, there is a fatal defect. When there is no overlap between the predicted box and the ground truth box, that is, or when, , then the partial derivatives of and are . It can be concluded that the gradient of backpropagation disappears, which will cause or not to be updated during training.
[0043] Low-quality examples are inevitably included in the training data. Therefore, in this study, Wise-IoU is used. In the first step, the penalty term is defined as the normalized length of the connection of the center points , and its formula is . Since may produce gradients that hinder convergence, in order to effectively eliminate the factors that hinder convergence, we separate and from the computational graph and represent them with , obtaining: .
[0044] Based on the distance metric, distance attention is constructed, and a two-layer attention mechanism is obtained. Let the attention function , which can better amplify the of ordinary-quality anchor boxes, and based on the distance metric, is constructed, and its formula is: .
[0045] In the second step, to solve the problem of unbalanced sample quality, the Focal-EIOU loss function develops a monotonic focusing strategy specifically for the cross-entropy loss, significantly reducing the impact of easy-to-classify samples on the overall loss. This improvement allows the model to focus more on solving those challenging samples, thereby improving the accuracy of the classification task. Similarly, we can construct the monotonic focusing coefficient , where is the exponential power, is the base, represents separation from the computational graph, then ; during the training process, the monotonic focusing coefficient decreases as decreases, which will cause the convergence speed to gradually slow down in the later stage of training. Therefore, we first set a momentum m and introduce the moving average as the normalization factor: , ; This type of method can ensure that the overall performance always remains in a better state, effectively addressing the problem of the slowdown in the convergence speed that occurs in the later stage of model training.
[0046] In the third step: We define an outlier degree to describe the quality of the anchor boxes: , a smaller outlier degree indicates high-quality anchor boxes. Therefore, we assign a lower gradient gain to it to promote the bounding box regression to focus on anchor boxes of general quality. For anchor boxes with a larger outlier degree, by assigning a smaller gradient gain, we can prevent them from having an overly negative impact on the model, thus avoiding excessive error propagation caused by low-quality samples. We set the hyperparameters and , and obtain the formula: ; when the parameter , such that . When the outlier degree of the anchor box reaches a specific threshold where is a predefined constant), this anchor box obtains the highest gradient gain. Since is dynamically changing, the quality evaluation criteria for anchor boxes are also adjusted dynamically. This dynamic nature enables to continuously optimize the allocation of gradient gains to adapt to the current training state.
[0047] As Figure 7 shows, the backbone network is the basic feature extraction part of YOLOv8. As can be seen from the figure, it starts with a simple convolutional layer (Conv), which initially captures the features of the input image. Subsequently, there are convolutional operations processed by the unequal channel reduction grouping method, followed by the C2f module improved based on the convolutional operation of the unequal channel reduction grouping, denoted as C - N - C2f. These modules can efficiently extract features at different levels through specific grouping and convolutional methods. Among them, C - N - C2f plays an important role in the backbone network. It can mine features at different scales, enabling the backbone network to obtain more representative and discriminative features. At the end of the backbone network, there is an SPPF (Spatial Pyramid Pooling - Fast) module, which performs pooling operations on features at different scales, helping to fuse local features and ensuring that the model can obtain sufficient information when processing targets of different sizes.
[0048] The neck network is responsible for feature fusion and transmission between the backbone network and the detection head. The figure shows that the neck network contains multiple Concatenate and Upsample operations. Through the Upsample operation, the low-resolution feature map can be restored to a higher resolution so that it can be concatenated with other high-resolution feature maps. These concatenation operations can effectively fuse features at different levels. For example, fuse feature maps at different depths in the backbone network to provide more comprehensive and rich feature information for the detection head. And incorporate the conical multi-scale dimensionality reduction feature extraction algorithm to enrich feature extraction. In addition, there are also convolution operations and C - N -C2f processed by the unequal grouping method of channel reduction in the neck network, which continue to process the fused features to further optimize the feature representation and ensure higher-quality features are passed to the detection head.
[0049] Combine the accelerometer and magnetometer at the data level and perform LLE feature-level fusion. This method is named AMLM (AccMag–LLE’s Method). The accelerometer and magnetometer can be installed on the object to be monitored (such as cattle) at the same time, and data is collected at the same time. The advantage compared with YOLOv8 is that it can clearly distinguish whether the cattle is walking or running, rather than only being able to recognize movement like YOLOv8.
[0050] The accelerometer works based on Newton's second law (F = ma). It usually consists of a mass block and a sensor that can detect the force on the mass block. When the accelerometer is subjected to acceleration, the mass block will generate a corresponding displacement due to inertia, and the sensor measures the magnitude and direction of the acceleration by detecting this displacement.
[0051] Wear the accelerometer on the leg of the cattle. Assume that the accelerometer can measure the acceleration of the cattle in three-dimensional space (x, y, z), and the sampling frequency is f (unit: Hz). The sampling interval time . At time Collect acceleration data.
[0052] For each sampling moment The accelerometer records the acceleration of the cattle on the x, y, and z axes
[0053] Check whether there are outliers in the collected data. Assume the reasonable range of acceleration is , for each sampling point n, if ( Similarly), then consider this point as an outlier, and linear interpolation can be used for correction. The formula is , where is The two nearest non-outliers before and after, and by the same token .
[0054] Remove the noise in the data. Moving average filtering can be used. Let the moving average window size be M, then the x-axis acceleration value after filtering , and by the same token and .
[0055] Since the data units and ranges of the accelerometer and magnetometer may be different, the data , , needs to be normalized. For the acceleration data, the formula can be used, and by the same token , .
[0056] According to the characteristics of cattle behavior, the collected data is segmented into several segments, each corresponding to a behavior. Let the behavior category be . For each segment of data, the behavior is labeled through on-site observation or other auxiliary means (such as video recording) to obtain a labeled dataset consisting of L data pairs where is the acceleration data of the i-th segment, represents the corresponding behavior.
[0057] To reflect the overall movement amplitude and stability, the time-domain characteristics of the accelerometer are calculated: The mean value of the x-axis acceleration is , where is the number of sampling points of the i-th segment of data. By the same token, the mean value of the y-axis acceleration and the mean value of the z-axis acceleration can be obtained.
[0058] Since acceleration itself is related to force and energy (according to Newton's second law F = ma), variance can well describe the degree of energy dispersion of acceleration over a period of time. The variance of the x-axis acceleration is , and by the same token, the variance of the y-axis acceleration and the variance of the z-axis acceleration can be obtained.
[0059] The peak value of the x-axis acceleration , and by the same token, the peak value of the y-axis acceleration and the peak value of the z-axis acceleration can be obtained.
[0060] To highlight the periodic and rhythmic behavior characteristics, the Fast Fourier Transform (FFT) is used to perform the fast Fourier transform on the x-axis to obtain the frequency-domain sequence , , where \(i\) is the imaginary unit. Similarly, we get , .
[0061] Calculate the energy in the frequency domain , where represents the modulus of the complex number , and find the frequency with the maximum energy , To find the value of \(k\) that maximizes . Similarly, the frequency domain characteristics of the \(y\)-axis and \(z\)-axis can be obtained .
[0062] Thus, the feature extraction is completed. Use the obtained data to extract the feature vector as: .
[0063] A magnetometer based on the principle of electromagnetic induction. When a magnetic field passes through the coil, according to Faraday's law of electromagnetic induction ( ), an induced electromotive force will be generated in the coil. By measuring this electromotive force, the magnetic field strength can be deduced. Where \(E\) is the induced electromotive force, \(N\) is the number of turns of the coil, is the rate of change of magnetic flux.
[0064] Similar to the data acquisition of cattle using an accelerometer, assume that the data accuracy of the magnetic field strength measured by the magnetometer in the three axes of \(x\), \(y\), and \(z\) is (unit: Tesla, T), and the number of samples \(N\) collected is , where . Check whether there are outliers in the collected data and use moving average filtering to obtain .
[0065] Perform feature extraction on it. Since the computing power required by the magnetometer is more complex, in order to save costs, we choose to calculate the amplitude of the magnetic field strength rather than using the \(x\), \(y\), and \(z\) axes separately as the feature vector: .
[0066] Normalize the amplitude, that is, we get ; Calculate the amplitude change rate: , where is the sampling interval time (unit: second).
[0067] Extract the time domain features: Similar to the method of the accelerometer, calculate the mean value , the maximum and minimum values For the magnetic field strength measured by the magnetometer, its unit is Tesla (T), and the unit of the standard deviation is also Tesla. This makes the standard deviation more intuitively represent the fluctuation range of the magnetic field strength relative to the average value in terms of physical meaning. The standard deviation , calculate it, and perform fast Fourier transform FFT to obtain the frequency domain sequence . Calculate the power spectral density: , and find the frequency k among all frequencies k that makes the largest. This frequency is the main frequency, denoted as . Set 6 frequency intervals , and the band energy .
[0068] Thus, the feature extraction is completed, and the feature vector is extracted using the obtained data as .
[0069] In feature fusion, locally linear embedding (LLE) is selected to obtain . LLE is a non-linear dimensionality reduction method. Its basic idea is that each data point can be approximately reconstructed by a linear combination of its neighboring points. First, for each data point , use the Euclidean distance , where m is the original dimension of the feature vector, to find its k nearest neighbor points . Then, solve a weight matrix W such that each data point can be approximately represented by a linear combination of its neighboring points, that is and satisfy the constraint condition , and solve the weight matrix by minimizing the reconstruction error . Using the obtained weight matrix W, map the data points to a low-dimensional space. Let the data points in the low-dimensional space be , and obtain the low-dimensional embedding by solving the eigenvalue problem of a low-rank matrix M, that is, minimize the objective to solve the low-dimensional data points , where is the corresponding neighboring point in the low-dimensional space. Select the eigenvector corresponding to the smallest non-zero eigenvalue as the low-dimensional embedding result, thus removing the redundant information in the data and being able to handle non-linear features well.
[0070] Use the minimization objective to solve the low-dimensional data points , where is the corresponding neighboring point in the low-dimensional space; Let , where I is the identity matrix, and solve the eigenvalue problem of matrix M , take the eigenvectors corresponding to the smallest d + 1 non-zero eigenvalues (excluding the smallest eigenvalue corresponding to the all-ones vector) to obtain the low-dimensional embedding result , where are data points in the low-dimensional space.
[0071] Use the support vector machine (SVM) for model selection and construction: Six SVM classifiers need to be constructed. For the i-th SVM classifier (i = 1, 2, 3, 4, 5, 6), mark the samples belonging to class as the positive class (y = +1), and mark the samples of the remaining five classes as the negative class (y = -1). Select the radial basis function (RBF) as the kernel function, and the kernel , where is the kernel parameter, which determines the width of the kernel function. Use k-fold cross-validation to optimize the penalty parameter C and . and are eigenvectors (eigenvectors extracted from the acceleration data ).
[0072] For the i-th SVM classifier, use the labeled data for training. The training data set is , where when , , , . The optimization objective function for each SVM classifier is , and the constraint condition is , where C is the penalty coefficient, is the slack variable.
[0073] Solve the optimization problem of each SVM classifier through the sequential minimal optimization algorithm (SMO) to obtain the optimal weight vector and the bias .
[0074] For a new sample , input it into these six SVM classifiers respectively to obtain six decision function values: . Compare these six decision function values, and select the class corresponding to the classifier with the largest decision function value as 's predicted class, that is: if , then predict belongs to class.
[0075] UWB: A radio technology in indoor positioning systems, used to identify and locate cattle groups, and combined with other sensor data to enhance cattle behavior detection. The detected behaviors are: foraging, drinking, running, walking, and standing still. It differs from YOLOv8 and AMLM in that it effectively reduces the confusion between two similar patterns of behavior (such as eating feed and drinking water).
[0076] Deploy multiple UWB base stations in the cowshed. The positions of these base stations are known, and they will form the infrastructure of the positioning system for receiving signals sent from UWB tags worn on cattle. Equip each cow with a UWB tag, which will regularly send UWB signals. The UWB tag sends signals in the form of non-sinusoidal narrow pulses at the nanosecond to picosecond level, and the signals propagate in the cowshed. The pre-arranged UWB base stations in the cowshed receive these signals. Due to the ultra-wideband characteristics of UWB signals, the signals can effectively propagate in complex indoor environments and have strong anti-interference capabilities.
[0077] Distance measurement: When the base station receives the signal sent by the tag, use the two-way time-of-flight method (TW-TOF) to measure the round-trip time-of-flight of the signal from the tag to the base station ( ). According to the speed of light (c) and the measured round-trip time-of-flight ( ), calculate the distance (d) between the base station and the tag. The calculation formula is: . Divide by 2 here because the measured is the round-trip distance and the one-way distance is needed.
[0078] After measuring the distance, perform positioning calculations. At least three base stations need to receive the signals of the same tag in order to calculate the position of the cattle through triangulation or other positioning algorithms. Assume the coordinates of the three base stations are , and their distances to the tag are , then the coordinates of the tag (cattle) can be solved through the following equations: . By solving this system of equations, the two-dimensional coordinate position of the cattle in the cowshed can be obtained, thus realizing the positioning of the cattle.
[0079] For the last step of multimedia feature fusion, perform decision-level fusion. Since the behaviors that can be studied by the three sensors are all different and most algorithms on the market are not applicable, and in order to reduce computing power, the majority voting decision fusion method is used.
[0080] The recognized results are eating, foraging, drinking, lying down, standing, moving, walking, running, estrus, and abnormal.
[0081] Feeding: When YOLOv8 and AMLM jointly recognize feeding, and UWB does not recognize forage or drinking water, the displayed behavior is feeding.
[0082] Forage: When at least one of YOLOv8 and AMLM recognizes feeding, and UWB recognizes forage, the displayed behavior is forage.
[0083] Drinking water: When at least one of YOLOv8 and AMLM recognizes feeding, and UWB recognizes drinking water, the displayed behavior is drinking water.
[0084] Lying prone: When YOLOv8 and AMLM jointly recognize lying prone, or UWB recognizes stillness, and at least one of YOLOv8 and AMLM recognizes lying prone, the displayed behavior is lying prone.
[0085] Standing: When YOLOv8 and AMLM jointly recognize standing, or UWB recognizes stillness, and at least one of YOLOv8 and AMLM recognizes standing, the displayed behavior is standing.
[0086] Moving: When YOLOv8 and UWB jointly recognize movement, and AMLM does not recognize walking or running, the displayed behavior is moving.
[0087] Walking: When at least one of YOLOv8 and UWB recognizes movement, and AMLM recognizes walking, the displayed behavior is walking.
[0088] Running: When at least one of YOLOv8 and UWB recognizes movement, and AMLM recognizes running, the displayed behavior is running.
[0089] In heat: When at least two of YOLOv8, AMLM, and UWB recognize being in heat, the displayed behavior is being in heat.
[0090] Abnormal: When the above situations do not occur, it is abnormal.
[0091] Such as Figure 9 , Figure 10 As shown, when the system does not detect any behavior or abnormality of the cows in the cowshed, that is, when there are no cows, turn on the ultraviolet lamp for 20 minutes for disinfection, turn on the speaker at an appropriate volume to drive away the cows and make them leave the cowshed, and at the same time turn on the fan to make the air circulate and prevent ozone accumulation from causing physical discomfort to the farmers and cows. Do this only once a day. When the disinfection is over or cows are detected to enter the cowshed, immediately stop the above measures.
[0092] Use a digital light sensor BH1750 and connect it to the computer through the I2C interface. The sensor's VCC is connected to the power supply, GND is connected to the ground, SCL and SDA are connected to the corresponding I2C clock and data pins of the development board or adapter respectively, and use the OpenCV library and Python - BH1750 library to write code to set the light brightness range to 0-100.
[0093] like Figure 8 As shown, according to the recognized light intensity, PWM is used to control the brightness of the bulb, which is divided into five levels, and the brightness is increased or decreased according to the corresponding measures taken according to the recognized cow behavior.
[0094] When detected When the light bulb is set to the current brightness, the light bulb will automatically adjust the brightness to 20 minutes.
[0095] in , To detect the amount of cattle eating or grazing behavior, is the sum of all the behaviors detected by the cow at the same time. , is the detected brightness, Indicates integer divisibility.
[0096] When detected If the time exceeds 5 minutes, the bulb will automatically adjust the brightness to The gear is turned off 2 minutes after the end of the behavior, and the fan is turned on and turned off 2 minutes after the end of the behavior. At the same time, the situation of frequent cattle movement is reported to the farmer.
[0097] in , To detect the amount of movement or walking behavior of cattle, it is determined that agitation occurs. , and when , the power is turned off.
[0098] When the running behavior of cattle is detected, it is judged that there is agitation, and the light bulb automatically adjusts the brightness to The light is turned on and off 2 minutes after the restless behavior ends, and the fan is turned on and off 2 minutes after the restless behavior ends. At the same time, the identity of the cow is recorded through the ear tag information, and the fact that the cow is running is reported to the farmer.
[0099] When detected , and the time exceeds 5 minutes, the bulb automatically adjusts the brightness to The file will be closed 2 minutes after the behavior ends.
[0100] Among them, , is the number of detected lying-down behaviors of cows. Moreover, when a cow has a long-term lying-down behavior, the identity of the cow is recorded through the ear tag information, and the situation of long-term lying-down is sent back to the farmer.
[0101] When a cow's estrus behavior is detected, the video of this period of behavior is automatically saved and the automatic saving is terminated 1 minute after the end of the estrus behavior. At the same time, the identity of the cow is recorded through the ear tag information, and the situation of the cow running is sent back to the farmer.
[0102] Ultraviolet rays can damage the DNA or RNA structure of microorganisms such as bacteria, viruses, and fungi. When microorganisms are exposed to ultraviolet rays of a certain intensity and duration, their nucleic acids will absorb the energy of the ultraviolet rays, resulting in the breakage of molecular chains or the formation of pyrimidine dimers, so that the microorganisms cannot carry out normal reproduction and metabolism and finally die. For example, common pathogens in cowsheds such as Escherichia coli, Salmonella, and cowpox virus can all be effectively killed by ultraviolet rays. This broad-spectrum bactericidal property can greatly reduce the number of pathogenic microorganisms in the cowshed and reduce the risk of cows being infected with diseases.
[0103] Fine light management of the cattle's feeding environment can significantly affect their feeding behavior, digestion efficiency, and overall health. When feeding cattle, appropriately increasing the light intensity has been proven to increase the feed intake of cattle. This may be because moderate light mimics the conditions of daytime in the natural environment, thus stimulating the foraging behavior of cattle. In addition, good light conditions help cattle better observe and select their feed, thereby promoting digestion and nutrient absorption.
[0104] When cattle move frequently or show restless behaviors such as running, it may be a reaction to environmental changes or some form of stress. This may include reactions to noise, extreme temperatures, uncomfortable housing conditions, or interference from other animals. In this case, appropriately dimming the light and improving air circulation can be used as an effective management strategy to reduce the stress response of cattle. A darker environment can help cattle feel safer, reduce their tension, thereby encouraging them to stay calm, reduce unnecessary energy consumption, and turn on the fan for ventilation.
[0105] The present invention and its implementation manners are schematically described above. This description is not restrictive. Without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Any reference signs in the claims should not limit the claims involved. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of this creation, design similar structural manners and embodiments to this technical solution without creative efforts, they should all fall within the protection scope of this application. In addition, the term "including" does not exclude other elements or steps, and the word "a" before an element does not exclude including "a plurality of" such elements. The plurality of elements stated in the product claims can also be implemented by one element through software or hardware. Words such as first and second are used to represent names and do not indicate any specific order.
Claims
1. A multimodal beef cattle behavior recognition method, characterized in that: include: Improve the YOLOv8 target detection algorithm, and use the improved YOLOv8 target detection algorithm to perform visual analysis on cattle to obtain the cattle's behaviors including eating, lying, standing, moving, and estrus as the first behavior; Collect cattle movement data and identify, through feature analysis, behaviors including but not limited to walking or running as a second behavior; Obtain cattle location data, and identify cattle feeding, drinking or resting behaviors as the third behavior by comparing cattle locations with preset functional area coordinates; The first behavior, the second behavior and the third behavior are integrated to obtain the final behavior type of the cattle; Adjust the cowshed environmental parameters according to the final behavior type of the cattle; the cowshed environmental parameters include light brightness, fan operation parameters and ultraviolet disinfection parameters.
2. The multimodal beef cattle behavior recognition method according to claim 1, characterized in that: Improved YOLOv8 target detection algorithm, including: The channel reduction unequal grouping method is used to optimize the backbone network of the YOLOv8 target detection algorithm; Based on the optimized backbone network, a C-NBottleneck network is constructed, and the C-NBottleneck network is used to replace the Bottleneck component in the C2F module in the backbone network of the YOLOv8 target detection algorithm; A cone multi-scale dimensionality reduction feature extraction algorithm is added to the replaced YOLOv8 target detection algorithm, and an enhanced feature map is obtained through the channel attention mechanism and the spatial attention mechanism to obtain an improved YOLOv8 target detection algorithm.
3. The multimodal beef cattle behavior recognition method according to claim 2, characterized in that: The channel reduction unequal grouping method is used to optimize the backbone network of the YOLOv8 target detection algorithm, including: Define channel reduction unequal grouped convolution operation; All convolution operations except the first layer of convolution in the backbone network are replaced by the defined channel reduction unequal grouped convolution operation to obtain the optimized backbone network.
4. The multimodal beef cattle behavior recognition method according to claim 3, characterized in that: Define channel reduction unequal grouped convolution operation, including: Perform a preliminary convolution operation on the input feature map to obtain an intermediate feature map; Divide the channel dimension of the intermediate feature map into n groups. When the channel dimension cannot be divided by n, set the number of channels in each of the first n-1 groups to be equal, and the last group contains the remaining channels. Divide the convolution kernels into n groups according to the same grouping strategy as the intermediate feature map; For each group of feature maps, use the convolution kernel of the corresponding group to perform convolution operation; Concatenate each group of convolution results with the intermediate feature map along the channel dimension to obtain the convolution output feature map.
5. The multimodal beef cattle behavior recognition method according to claim 2, characterized in that: Build a C-NBottleneck network, including: Obtain a feature map output by the optimized backbone network as the first feature map; Use SiLU activation function to perform nonlinear transformation on the first feature map; Using the optimized backbone network to perform convolution processing on the first feature map after nonlinear transformation, a second feature map is obtained; Normalizing the second feature map; The normalized second feature map is residually connected to the first feature map to generate a C-NBottleneck network.
6. The multimodal beef cattle behavior recognition method according to claim 2, characterized in that: The enhanced feature map includes: Get the feature map F output by the replaced YOLOv8 target detection algorithm; According to the feature map F, the channel weighted feature map is obtained through channel reduction unequal grouping convolution based on the channel attention mechanism; According to the channel weighted feature map, the spatial weighted feature map is obtained by channel reduction unequal grouping convolution based on the spatial attention mechanism; According to the spatial weighted feature map and the feature map F, an enhanced feature map is obtained through a depth-separated convolution operation.
7. The multimodal beef cattle behavior recognition method according to claim 6, characterized in that: Get the channel weighted feature map, including: The feature map F is convolved using the channel reduction unequal grouping method with a convolution kernel size of N1*N1 to obtain an intermediate feature map; Perform global average pooling on the intermediate feature map to obtain the channel descriptor; The channel descriptor is transformed nonlinearly through the fully connected layer to obtain the channel weight; According to the channel weight and feature map F, the channel weighted feature map is obtained.
8. The multimodal beef cattle behavior recognition method according to claim 6, characterized in that: Get the spatial weighted feature map, including: The channel weighted feature map is convolved using the channel reduction unequal grouping method with a convolution kernel size of N2*N2 to obtain the spatial attention feature map; Perform global average pooling on the spatial attention feature map to obtain the spatial feature descriptor; The spatial feature descriptor is transformed nonlinearly through the fully connected layer to obtain the spatial attention weight; According to the spatial attention weight and the channel weighted feature map, a spatial weighted feature map is obtained.
9. The multimodal beef cattle behavior recognition method according to any one of claims 2 to 8, characterized in that: Collect cattle movement data and identify walking or running behaviors through feature analysis as the second behavior, including: The acceleration data of cattle are collected by accelerometers; the magnetic field strength data of cattle are collected by magnetometers; Fast Fourier transform is used to process the acceleration data to obtain frequency domain features; Perform time domain analysis and spectrum analysis on magnetic field intensity data to obtain magnetic field characteristics; The local linear embedding algorithm is used to perform nonlinear dimensionality reduction fusion on the frequency domain features and magnetic field features to obtain the fused feature vector; The fused feature vector is classified using support vector machine (SVM) to obtain the second behavior.
10. The multimodal beef cattle behavior recognition method according to claim 9, characterized in that: Identify cattle feeding, drinking or resting behaviors as the third behavior, including: Multiple UWB base stations are deployed in the cowshed, and each cow is equipped with a UWB tag; The distance from the UWB tag to the UWB base station is determined using a two-way time-of-flight algorithm. The two-dimensional coordinates of the cattle are calculated by using the distance data of at least three base stations using the triangulation method as the cattle position data; According to the cattle location data and the functional area division of the cattle shed, it is determined whether the cattle are in a preset functional area, and the preset functional area includes a forage area and a drinking water area.
11. A multimodal beef cattle behavior recognition system, characterized in that: include: The visual analysis module uses the improved YOLOv8 target detection network to perform visual analysis on cattle and obtain the first behavior of cattle; The improved YOLOv8 target detection network includes a channel reduction unequal grouping optimized backbone network, a C-NBottleneck network and a conical multi-scale dimension reduction feature extraction network; the first behavior includes eating, lying down, standing and moving; The motion analysis module collects the motion data of the cattle through the accelerometer and the magnetometer, and obtains the second behavior of the cattle through time-frequency domain analysis and local linear embedding calculation, wherein the second behavior includes walking or running behavior; The positioning analysis module uses a two-way time-of-flight algorithm and a triangulation method to obtain cattle location data through a UWB base station and a UWB tag, and identifies the third behavior of the cattle by comparing the cattle location data with the coordinates of a preset functional area, wherein the third behavior includes feeding grass, drinking water or resting behavior; The data fusion module performs fusion analysis on the first behavior, the second behavior and the third behavior to obtain the final behavior type of the cattle.
Citation Information
Patent Citations
Novel cow digital twin behavior perception modeling method and system based on fusion of sensor data and video data
CN119091211A
Problem behavior recognition system and method for disabled people
CN114091596A
Henhouse environment control method and system
CN116740805A
Action recognition method and device and computer equipment
CN117292435A
Gait camouflage method for resisting biological recognition
CN119818832A