Cow bone joint point extraction and posture recognition method and system based on visual sensing
Through a deep learning-based method, combined with the improved ResNet network and the object detection algorithm GCMT, the problem of low accuracy of cattle bone joint node recognition and pose estimation is solved, and efficient recognition is achieved under different environments and image quality conditions is achieved, reducing the dependence on expensive equipment.
Patent Information
- Application Number
- CN202510227017.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-24
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art has low accuracy in cattle bone joint node recognition and pose estimation, high environmental conditions and image quality requirements, making it difficult to achieve effective recognition under multiple targets and occlusion conditions.
Using a deep learning-based method, combined with the improved ResNet network and the object detection algorithm GCMT, the extraction and pose recognition of the bovine bone joint nodes are realized through a multi-step process of feature extraction, object detection and pose recognition.
It improves the accuracy and robustness of cattle bone joint node recognition, can be effectively identified under different environments and image quality conditions, reduces dependence on expensive equipment, and realizes automated and efficient health assessment.
Smart Images

Figure CN120071398A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image vision, and relates to the posture recognition of animals. It is a method and system for extracting bovine skeletal joint points and posture recognition based on visual sensing. Background Art
[0002] In modern animal husbandry, it is crucial to accurately evaluate the health and growth status of cattle. However, traditional health assessment methods usually rely on manual observation or the use of expensive equipment, and require a large amount of time and resources. Therefore, a method and system that can automatically, quickly, and accurately evaluate the health status of cattle are needed.
[0003] In the fields of computer vision and image processing, significant progress has been made in joint point detection and pose estimation technologies. These technologies are mainly used in fields such as human pose estimation and gesture recognition, and achieve pose detection and tracking by identifying key points of the human body or object. However, in terms of the recognition and pose estimation of skeletal joint points of large mammals such as cattle, there are still obvious deficiencies, and there are also requirements for the quality of the detected video or image. For example, when there are many targets in the image or occlusion occurs, or in the case of insufficient light, the processing effect of existing algorithms still has room for improvement.
[0004] Therefore, a new method and system for identifying and extracting bovine skeletal joint points based on image processing are needed, which can overcome the limitations of traditional methods and achieve automatic, rapid, and accurate identification and extraction of bovine skeletal joint points. Summary of the Invention
[0005] In order to solve problems such as low accuracy, environmental condition limitations, and image quality requirements of traditional technologies, the present invention provides an algorithm and system for extracting bovine skeletal joint points and posture recognition based on image processing. On the one hand, a method for extracting bovine skeletal joint points and posture recognition based on deep learning is provided; on the other hand, a complete bovine detection application system is established for actual bovine action recognition.
[0006] The technical solution of the present invention is: a method for extracting bovine skeletal joint points and posture recognition based on visual sensing, including the following steps:
[0007] Step 1, collect video images of cattle, normalize the images, and perform image preprocessing;
[0008] Step 2, extract features from the preprocessed image I. The feature extraction network uses an improved ResNet network. The improved ResNet network combines the SE mechanism and spatial pyramid pooling SPP, embeds an SE block after the residual block, and adds a spatial pyramid pooling layer to the last fully connected layer;
[0009] Step 3: According to the extracted features, use the target detection algorithm GCMT to detect the cows as targets, identify the positions of the target cows, obtain the detection frames of the cows, and crop the cow pictures according to the detection frames for later use;
[0010] After the target detection algorithm GCMT decomposes the features extracted in Step 2 through two-dimensional feature vector decomposition, it adds position information and outputs the prediction results of the targets through the GM Transformer module. The GM Transformer module is based on the DETR target detection model. It encapsulates the generation of query objects Object Queries into a separate module GCOQ, replaces the first-level multi-head self-attention module in the encoder and decoder structures with a dynamic resolution enhanced attention module DETA, and replaces the second-level self-attention module in the encoder structure with a dynamic multi-scale attention module DMSA; among them
[0011] GCOQ updates the query objects Object Queries by introducing a graph structure among the query objects Object Queries and using the graph attention network GAT;
[0012] DETA adds a dynamic filtering layer in the multi-head self-attention, and dynamically enhances or weakens some attention connections based on the position distance in the feature map;
[0013] DMSA assigns weights to the feature maps of different scales of the image according to global and local information, fuses the feature maps of different scales by combining the weights, and then performs self-attention calculation;
[0014] Step 4: Extract the joint points of the cow skeleton within the detection frame;
[0015] Step 5: Convert the detected key points into the poses of the cows.
[0016] Furthermore, Step 1 is specifically as follows:
[0017] Scale the collected original image I o to a unified pixel size. If the image pixels are lower than the unified pixels, use the enhanced Laplacian sharpening ELS operator to sharpen the image and enhance the image features:
[0018] The ELS operator is a linear combination of the two-dimensional Laplacian operator L and an edge enhancement high-pass filter H to simultaneously emphasize the boundaries and texture details:
[0019] ELS = αL + βH
[0020] where α and β are adjustable weights, and apply the ELS operator to the image I o to obtain the sharpened image I sharp as follows:
[0021] I sharp = I o + c·(ELS * I o )
[0022] Where c is a coefficient for adjusting the sharpening intensity, and the value of c ranges from 0.3 to 0.5.
[0023] Furthermore, in step 2, the improved ResNet network is called SE-SPP ResNet, which includes an initial convolutional layer, a max pooling layer, two residual block groups + SE module, one residual block group + grouped convolution, one residual block group + SE module, and a spatial pyramid pooling layer + fully connected layer configured in sequence.
[0024] Furthermore, the GCOQ module is specifically as follows:
[0025] First, define the query object Object Queries as the nodes of the graph. Let the initial Object Queries matrix be:
[0026]
[0027] Where N is the number of query objects, D is the feature dimension of each query object, and these query objects are initialized based on the features extracted in step 2 to construct a graph:
[0028] G(V, E)
[0029] Where V is the set of nodes corresponding to the query objects; E is the set of edges representing the potential associations between the query objects, and these edges are calculated based on similarities such as Euclidean distance or cosine similarity.
[0030] Use the graph attention network GAT to update the features of the query object Object Queries. For each node m in the graph, its updated feature h ′ m is calculated as follows:
[0031]
[0032] Where h n is the original or previous iteration feature of node m, N(m) represents the set of neighbor nodes of node m, σ is the activation function, and α mn is the attention coefficient between nodes m and n, which is calculated by the following formula:
[0033]
[0034] W is a learnable weight matrix applied to the features of each node, and a is a learnable parameter vector for calculating the attention coefficient.
[0035] Furthermore, the dynamic resolution enhancement attention module DETA includes a multi-head self-attention module and dynamic content-aware filtering, which are implemented as follows:
[0036]
[0037] Among them, F l (u, v) is the weight function of position u relative to v at level l, and Q l , K l , V l represent the query, key, and value matrices at level l respectively, is used for normalization, and F l (u, v) is a distance-based filter, which is implemented by a Gaussian function:
[0038]
[0039] Among them, P u and P v are the coordinates of positions u and v on the feature map respectively, and σ is a parameter that controls the attenuation speed.
[0040] Furthermore, in the dynamic multi-scale attention module DMSA, for features of different scales Figure X l use a neural network to calculate the fusion weight α according to global and local information l , and these weights determine the contributions of feature maps of each scale in the final fusion. The feature fusion formula is as follows:
[0041]
[0042] Among them, X′ is the fused feature map, and Transform(X l ) represents a set of transformation operations to make the dimensions of feature maps of all scales consistent, and generate the values and keys in the attention calculation with the fused feature map for self-attention calculation.
[0043] The present invention also provides a bovine skeletal joint point extraction and pose recognition system based on visual sensing, which is used to implement the above-mentioned bovine skeletal joint point extraction and pose recognition method based on visual sensing, and includes:
[0044] An image acquisition module, which is used to acquire video images of cows. The image acquisition device requires at least a resolution of 1080p to ensure image clarity;
[0045] An image preprocessing module, which is used to implement the image preprocessing in Step 1. It is used to scale the picture to a picture with a size of 720 pixels * 720 pixels, ensuring that the input image size meets the input of the network. At the same time, the ELS operator is introduced to sharpen the image, making the image features input into the network more obvious;
[0046] A target detection module, which is used to implement the feature extraction in Step 2 and the target detection of cows in Step 3. It is used for later bone joint point extraction. A computer program is loaded in the target detection module to implement the improved ResNet network and the target detection algorithm GCMT in the method described in Claim 1. The improved ResNet network is used to extract the feature map, and GCMT is used to implement target detection based on the feature map to obtain the cow bone joint points;
[0047] A pose recognition module, which is used to implement Step 4 and Step 5. It connects the bone joint points output from the target detection module into the pose output of the cow's movement. In the part of bone joint point extraction, according to the cow bone joint point characteristics, the joint points on the cow's bone are divided into 17 points, and a new type of network is established to extract 17 bone joint points. After identifying 17 bone joint points, they are connected according to prior knowledge to output the final pose.
[0048] The beneficial effects of the present invention are:
[0049] 1. A method of sharpening the image by establishing an enhanced Laplacian sharpening operator is adopted. It can adapt to target objects of different scales and sizes, making the image features input into the network more obvious. At the same time, the model has better generalization ability when processing multi-scale images.
[0050] 2. For the target detection of cows, a target detection algorithm GCMT based on the improvement of the DETR algorithm is proposed. The advantages of GCMT are as follows: GCMT improves the Object Queries structure, enabling the model to better utilize target information and improving the detection performance; GCMT introduces a spatial multi-head attention mechanism, which can better capture the spatial relationship between targets, thereby improving the accuracy and robustness of target detection; GCMT adopts a dynamic multi-scale attention mechanism, which can adaptively adjust the attention weights, effectively process targets of different scales, and improve the generalization ability of the algorithm.
[0051] 3. In bovine skeletal joint point detection, a new method for extracting basic image features was established. By integrating the SE block and the SPP layer into the traditional ResNet architecture, the model not only improved the channel-level feature adjustment ability but also increased the adaptability to spatial scales. The combination of these two enables the model to more accurately capture useful information when dealing with complex image tasks. The introduction of the SE block helps the model better generalize to new and unseen data because it enhances the effectiveness of features by learning the dependencies between channels. At the same time, the multi-scale feature extraction strategy of the SPP layer further enhances the model's adaptability to different sizes and scale changes.
[0052] 4. The present invention establishes a new network for extracting bovine skeletal joint points, named Cattle-PoseNet. This network is trained with bovine pictures, making its performance more prominent in extracting bovine skeletal joint points; Cattle-PoseNet combines deep learning and traditional computer vision methods, fully leveraging the advantages of deep learning in feature learning while integrating traditional pose estimation methods, making the algorithm more robust and reliable; Cattle-PoseNet adopts a multi-scale feature fusion module, which can effectively integrate feature information at different scales, improving the adaptability to bovines of different sizes and proportions; The model structure of Cattle-PoseNet is relatively simple, and both training and deployment are relatively easy, making the algorithm more easily promoted and applied in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flowchart of the algorithm for bovine skeletal joint point extraction and pose recognition in the present invention.
[0054] Figure 2 It is a flowchart of the basic feature extraction of the image in the present invention.
[0055] Figure 3 It is a flowchart of the residual block group + grouped convolution in the present invention.
[0056] Figure 4 It is a flowchart of the residual block group + squeeze and excitation block in the present invention.
[0057] Figure 5 It is a flowchart of the object detection algorithm in the present invention.
[0058] Figure 6 It is a flowchart of the GM Transformer network in the present invention.
[0059] Figure 7 It is a flowchart of the GCOQ module in the present invention.
[0060] Figure 8 It is a flowchart of the DETA module in the present invention.
[0061] Figure 9 This is the flowchart of the DMSA module of the present invention.
[0062] Figure 10 This is the flowchart of the Cattle - PoseNet network of the present invention.
[0063] Figure 11 This is the flowchart of the application system for extracting cattle bone joint points and recognizing postures in the present invention. Detailed implementation manners
[0064] The following will illustrate the detailed implementation manners of the present invention. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0065] As Figure 1 shown, the present invention proposes a method for extracting cattle bone joint points and recognizing postures based on visual sensing, including the following steps:
[0066] S1. Collect the color video images of cattle and perform normalization and image pre - processing.
[0067] Image pre - processing plays a crucial role in the task of extracting cattle bone joint points and recognizing postures. It can help improve the image quality, reduce noise, enhance features, etc., so as to provide better input for the subsequent processing steps.
[0068] Since the video recording devices used in the actual production process are different and the sizes of the output pictures are different, in the present invention, first, the original picture I o is scaled to a picture with a size of 720 pixels * 720 pixels to ensure that the input image size meets the input of the network. When the picture pixels are higher than 720 pixels * 720 pixels, the picture can be directly reduced. If the picture pixels are lower than 720 pixels * 720 pixels, since the image will be enlarged and the picture will be distorted, etc., the present invention uses the Enhanced Laplacian Sharpener (ELS) operator to sharpen the image, making the image features input into the network more obvious. The specific method of sharpening is described through embodiments.
[0069] The ELS operator is defined as a linear combination of the Laplacian operator L and an edge - enhancing high - pass filter H. As an embodiment, first, define the standard two - dimensional Laplacian operator:
[0070]
[0071] Then, a high-pass filter is defined, aiming to enhance the high-frequency parts of the image, such as edge details:
[0072]
[0073] The ELS operator is a combination of these two filters, aiming to simultaneously emphasize larger boundaries and finer texture details:
[0074] ELS = αL + βH (3)
[0075] Where α and β are adjustable weights. In the present invention, the commonly used settings are α = 1 and β = 0.5. With such values, it can ensure that the edges are moderately enhanced while increasing the overall contrast and clarity of the image.
[0076] Apply the ELS operator to the image I o to obtain the sharpened image I sharp as follows:
[0077] I sharp = I o + c·(ELS * I o ) (4)
[0078] Where c is a coefficient for adjusting the sharpening intensity. Usually, c takes values between 0.3 and 0.5 to ensure that the image details are significantly enhanced without introducing excessive noise.
[0079] S2. Extraction of basic image features.
[0080] In order to achieve efficient extraction of basic image features in the algorithm for bovine bone joint point extraction and pose recognition, the present invention proposes an improved ResNet network, called SE-SPP ResNet. Based on the improvement of the traditional ResNet, the specific structure is as Figure 2 shown, consisting of an initial convolutional layer, a max-pooling layer, a residual block group, a squeeze-and-excitation block, grouped convolution, and a spatial pyramid pooling layer. The specific application and implementation process of this model in the present invention are described in detail below.
[0081] The input of the model is a color bovine body image with a resolution of 720x720 pixels. The image is first subjected to normalization processing, including color normalization and adjustment of lighting conditions, to reduce the impact of environmental variations on the model performance, and the preprocessed image I is obtained.
[0082] Since the 720*720-sized image output by step S1 is relatively large, the image is further reduced first to facilitate feature extraction and reduce the computational dimension. The present invention uses an initial convolutional layer to capture the preliminary features of the image and reduce the dimension. The initial convolutional layer uses a 7x7 convolutional kernel, 128 output channels, a stride of 2, and a padding of 3 at the edges. The purpose of this layer is to quickly reduce the image size while extracting preliminary spatial features. After the convolutional operation, batch normalization and the ReLU activation function are applied to enhance the network's non-linear processing ability and stabilize the training process. The max pooling layer uses a 3x3 max pooling kernel with a stride of 2 to further reduce the spatial dimension of the feature map and increase the model's invariance to small displacements. The specific method is as follows:
[0083] First, the image is further reduced to facilitate feature extraction and reduce the computational dimension. The present invention uses an initial convolutional layer to capture the preliminary features of the image and reduce the dimension. The convolutional kernel has a size of 7*7 and a stride of 2. The specific formula is as follows:
[0084] Conv(I) = ReLU(BN(W * I + b)) (5)
[0085] Where W and b are the weights and biases of the convolutional kernel respectively, BN is batch normalization, and ReLU is the activation function. After convolution, the image size becomes approximately 360x360 pixels. To further reduce the feature dimension, the present invention uses a max pooling layer after the initial convolution. A 3*3 pooling kernel is selected with a stride of 2. After pooling, the image size becomes approximately 180x180 pixels.
[0086] The specific design of the residual block is as follows: The residual blocks of the standard ResNet are used to construct the deep feature extraction network, and each residual block is followed by an SE block. The SE block compresses the spatial information of each channel into a single number through global average pooling; through a two-layer fully connected network, it first reduces the dimension and then increases the dimension, and outputs the adjustment coefficient of each channel through the Sigmoid activation function; the original channel features are multiplied by the adjustment coefficient obtained by excitation to achieve feature recalibration.
[0087] The specific structure is as Figure 3 , in the residual convolution, let x be the input feature map. The operations in the residual block usually include two convolutional layers, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function. The formula is expressed as:
[0088]
[0089] Where represents the residual function, including two weights W 1 、W 2 and biases b 1 、b2 The convolution operation, where x is the input feature map. BN is the batch normalization operation. ReLU is the activation function.
[0090] The Squeeze-and-Excitation (SE) block explicitly models the importance of each channel. The SE block enables the network to learn how to effectively utilize global information, optimize the feature dependencies between channels, thereby enhancing the richness and accuracy of feature representation. The Squeeze-and-Excitation block first performs global average pooling on the feature map after residual convolution to achieve Squeeze, then realizes Excitation through a fully connected layer to recalibrate the importance of each channel, and finally multiplies the recalibrated weights back to the original feature map for the Scale operation. The formulas for each part are as follows:
[0091] Squeeze:
[0092]
[0093] Excitation:
[0094]
[0095] Scale:
[0096]
[0097] Among them, is the compressed feature of the c-th channel obtained through global average pooling, and H×W represents the size of the feature map. s is the excitation weight for all channels obtained through two fully connected layers, W 3 and W 4 are the weights of the fully connected layers, δ is the ReLU activation function, σ is the Sigmoid function, represents the set of compressed features for all channels, expressed as a vector, where each element corresponds to the compressed feature of one channel. s c is the excitation weight of the c-th channel, is the feature map of the c-th channel adjusted by the excitation weight. Combining the residual connection and the output feature map adjusted by SE, the formula is:
[0098]
[0099] where x is the original input feature map, is the feature map adjusted by the SE block.
[0100] After the combination of two-layer residual modules and compression and activation blocks is a combined module of residual blocks and grouped convolutions. This module introduces grouped convolutions into the traditional residual block to reduce the consumption of computing resources while maintaining the model's expressive power. By dividing the feature map into groups and performing convolutions on each independent group, grouped convolutions greatly reduce the number of parameters and the amount of computation. This not only improves the computational efficiency but also helps the model to be deployed in resource-constrained environments such as mobile devices and embedded systems commonly used in production environments. In the present invention, we divide the input feature map into multiple groups, perform convolution operations on each group of features independently, and finally combine these features. The specific module is as Figure 4 , which divides the input features Figure X into G groups, performs convolution operations on each group independently, and applies the convolution kernel W g to each group of feature maps, where g represents the g-th group.
[0101] X g = W g * X g for each group g = 1, 2, ……, G (11)
[0102] where W g is the g-th group convolution kernel, and X g is the subset of features assigned to the g-th group. By dividing the feature map into groups and processing each group of data independently, grouped convolutions can effectively reduce the number of parameters, thereby reducing the risk of overfitting and improving the computational efficiency.
[0103] Then, the convolution results of all groups are combined into a single feature map, and the formula is as follows:
[0104] Y = Concat(X 1 , X 2 , ……, X G ) (12)
[0105] Apply batch normalization and ReLU activation function to the combined feature map Y:
[0106] Z = ReLU(BN(Y)) (13)
[0107] This step is a standard operation to normalize the data distribution and introduce non-linearity, which helps to improve the training process and the generalization ability of the model.
[0108] Finally, directly add the activated feature map Z to the original input X to form the final output feature map O
[0109] O = Z + X (14)
[0110] Subsequently, there is another combination of residual block groups and compression and activation blocks. After that, there is a combination of a spatial pyramid pooling module and a fully connected layer to adjust the size of the image and output the final feature image. The spatial pyramid pooling layer integrates the spatial pyramid pooling layer at the end of the residual network and performs pooling operations of 1x1, 2x2, and 4x4 to capture context information at different scales. All pooling outputs are concatenated into a fixed-length feature vector to ensure that the dimension of the output feature remains unchanged regardless of the change in the input image. The fully connected layer converts the fixed-length feature vector output by the SPP layer into a format suitable for subsequent processing, such as predicting joint positions or pose classification labels.
[0111] The training details are as follows: The model is trained using cross-entropy loss (for classification) or mean squared error loss (for regression). The Adam optimizer is used, and the initial learning rate is set to 0.001, which is adapted during the training process through a learning rate decay strategy. The network architecture and hyperparameters, such as the learning rate and batch size, are adjusted based on the performance on the validation set to obtain the best performance.
[0112] S3. Identify the position of the target cow and crop the cow's picture
[0113] Use the Graph-based Convolutional Multitask Tracker (GCMT) algorithm to perform object detection on the input image to identify the position of the cow in the image, such as Figure 5 . The GCMT algorithm adopts a spatial attention mechanism and a dynamic multi-scale attention mechanism, which can effectively capture the spatial relationship between objects and adaptively adjust the attention weights, improving the accuracy and robustness of object detection. In the GCMT algorithm, after the features extracted in step 2 are decomposed into two-dimensional feature vectors, position information is added, and the prediction results of the objects are output through the GM Transformer module. The GM Transformer module is similar to the DETR structure of the object detection model. The improvement lies in changing the generation method of the query object Object Queries, changing the structure of the multi-head self-attention mechanism module and the multi-head attention mechanism module to make it more suitable for the task of extracting cow skeletal joint points in the present invention. Since the generation method of the query object Object Queries in DETR is changed, the present invention encapsulates it into a separate module, named the Graph-Contextual Object Queries Module (GCOQ), and the structure is as Figure 7 .
[0114] GCOQ is a method that uses graph neural networks (GNNs) to enhance object queries, especially suitable for object detection tasks in complex scenarios. This module introduces a graph structure among object queries, allowing queries to communicate information with each other and enhancing their expressive and judgment capabilities.
[0115] We first define object queries as the nodes of the graph. Let the initial object query matrix be:
[0116]
[0117] where N is the number of queries and D is the feature dimension of each query. These queries are initialized based on the features initially extracted by the neural network. On this basis, a graph is constructed:
[0118] G(V, E)(16)
[0119] where V is the set of nodes, corresponding to object queries; E is the set of edges, representing the potential associations between queries. These edges are calculated based on similarities such as Euclidean distance or cosine similarity.
[0120] The present invention uses a graph attention network (GAT) to update the features of object queries. For each node i in the graph, its updated feature h i ′ is calculated as follows:
[0121] h i ′ = σ (∑ j∈N(i) α ij Wh j ) (17)
[0122] where h j is the original or previous iteration feature of node i. N(i) represents the set of neighbor nodes of node i. α ij is the attention coefficient between nodes i and j, calculated by the following formula:
[0123]
[0124] W is a learnable weight matrix applied to the features of each node, a is a learnable parameter vector for calculating the attention coefficient, and σ is an activation function. In the present invention, a non-linear function such as ReLU is selected.
[0125] After being processed by the GAT layer, each Query not only retains its original target information but also integrates the information of its neighboring Queries, achieving context enhancement of information. This enables each Query to more comprehensively consider the surrounding environment and the states of other targets during object detection, thereby improving the accuracy and robustness of detection.
[0126] Through the above steps, the GCOQ module effectively utilizes the graph structure to enhance the functionality of Object Queries, enabling the object detection model to have better performance in complex scenarios. This method is particularly suitable for production environments where the interactions between objects are significant, or there are a large number of occlusions and intersections.
[0127] After generating the three vectors q, k, and v through the GCOQ module, it will enter the Dynamic Resolution Enhanced Attention (DETA) module. Figure 8 , adding a dynamic filtering layer. The present invention improves the multi-head attention of DETR to a multi-head self-attention module plus a dynamic content-aware filtering module, which can more effectively process features of various scales and focus on more critical information regions.
[0128] The present invention introduces a dynamic content-aware filter. The purpose of this filter is to enhance or weaken certain attention connections based on the content of the features, so that the attention mechanism can pay more attention to the important or information-rich parts of the image. This method is particularly suitable for processing complex visual scenes, where some regions may contain more information about the objects to be detected than other regions. The specific formula is as follows:
[0129]
[0130] Among them, F l (u, v) is the weight function of position u relative to v at level l, and Q l , K l , V l represent the query, key, and value matrices at level l respectively. is used for normalization. F l (u, v) realizes dynamic content-aware filtering through weights, using a distance-based filter. This filter considers the spatial distance between two positions in the feature map, used to emphasize the relationship between neighboring positions while reducing the influence of distant positions. It can be implemented through a Gaussian function:
[0131]
[0132] where P u and P v are the coordinates of positions u and v on the feature map respectively, and σ is a parameter controlling the attenuation rate.
[0133] After passing through the Add&Norm module, the output value v and the query q and key k of the encoder part are input into the Dynamic Multi-Scale Attention (DMSA) module, and the structure is as Figure 9 . This module aims to enhance the model's detection ability for targets of different sizes by adaptively combining features from different scales, especially in complex scenarios where the target sizes and scales vary significantly.
[0134] First, use the Feature Pyramid Network (FPN) with multi-resolution output to generate a series of feature maps {X 1 , X 2 , ……, X L} from the input image, where X 1 is the feature map with the highest resolution and X L is the feature map with the lowest resolution. Then, for the features of each scale Figure X l use a small neural network to calculate a fusion weight α l based on global and local information. These weights determine the contribution of each scale feature map in the final fusion. The feature fusion formula is as follows:
[0135]
[0136] where Transform(X l ), a set of transformation operations, makes the dimensions of feature maps of all scales consistent for easy summation. The processed fused features are the same as those of the traditional attention module. It should be noted that the Q l , K l , and V l used in the attention calculation in the Dynamic Multi-Scale Attention module DMSA are obtained from the query, key, and value matrices generated by X ′ instead of directly operating on the directly input q, k, and v.
[0137] After passing through the GM Transformer, the target detection of cows can be achieved, and the detection boxes classified as cows are recorded for later use.
[0138] S4. Implement the extraction of joint points of cow skeletons inside the detection box
[0139] The Cattle-PoseNet network is used to extract the joint points of cattle bones. Cattle-PoseNet is a network specifically designed for cattle pose estimation tasks, which can accurately identify the positions of key points on cattle, including key parts such as the head, tail, and limbs. The image of the target cattle is input into the Cattle-PoseNet network, and through the forward propagation process of the network, the extraction of cattle bone joint points is achieved. The Cattle-PoseNet network internally contains multiple convolutional neural networks and key point regression layers, which can effectively learn image features and predict the positions of key points on cattle. The output results of the Cattle-PoseNet network include the coordinate information of each key point on cattle, such as the positions of key points like the head, tail, and limbs. These coordinate information can be obtained through the output layer of the network, that is, the predicted key point coordinates. Mapping the predicted key point coordinate information to the original image realizes the visualization of cattle bone joint points. This can be achieved by drawing position marks of key points on the image, thus intuitively displaying the recognition results of cattle poses.
[0140] In the present invention, the joint points on the cattle bones are divided into 17 points, which are respectively named as head, left eye, right eye, chest, back, left shoulder, left arm, left hand, right shoulder, right arm, right hand, left hip, left leg, left foot, right hip, right leg, and right foot. According to the above characteristics of cattle bone joint points, the present invention establishes a new network to extract 17 bone joint points, which is named Cattle-PoseNet, as Figure 10 , and this network is trained with cattle pictures to make its performance more prominent in extracting cattle bone joint points.
[0141] First, the cattle detection pictures generated by S3 are scaled to a unified size, and the SE-SPP ResNet network is used for feature extraction. These two processes are similar to the processes described in S2 and will not be elaborated here. After feature extraction, it enters the depthwise separable convolution module. Depthwise Separable Convolution is an efficient convolution operation method that decomposes the standard convolution operation into two smaller operations: depthwise convolution and pointwise convolution. This structure not only reduces the number of parameters of the model but also reduces the computational complexity while maintaining or enhancing the performance of the network. Assuming the input feature Figure XThe dimension is C×H×W, where C is the number of channels, and H and W are the height and width of the feature map. During depth convolution, the convolutional kernel is independently applied to each input channel, which mainly processes spatial features without cross-channel information. The size of the convolutional kernel for each channel is 3×3. Pointwise convolution: Use a 1×1 convolutional kernel to aggregate the output channels of depth convolution. This step is essentially a fully connected layer processing for feature points at each position, thus achieving feature integration and dimensional transformation. The formula is as follows:
[0142] Y = (X * K depthwise ) * K 1 ×1 (22)
[0143] Among them, X * K depthwise represents applying the corresponding depth convolutional kernel to each channel of X, which is 3×3 in the present invention, and K 1×1 is the pointwise convolutional kernel with a dimension of C′×C×1×1, used to convert C channels to C′ channels.
[0144] After that is the spatial attention module. The main function of this module is to dynamically adjust the focus of the feature map it processes, strengthen the features most relevant to the current task (i.e., the key point positions of the cow), and at the same time ignore background noise and irrelevant information. This mechanism helps the network accurately identify key points in various scenarios, especially in cases where the background is complex or the lighting conditions are not ideal. The spatial attention module usually consists of the following steps: First, compress the feature map along the channel direction through average pooling and max pooling layers, simplifying each H×W feature map into a 1-dimensional response. This step aims to aggregate global spatial information. Then, combine the outputs of average pooling and max pooling and process them through one or more convolutional layers (the present invention uses 1×1 convolution) to generate a spatial attention map. Next, through an activation function, normalize the values of the attention map between 0 and 1 to generate the attention weights for each spatial position. The overall calculation formula is as follows:
[0145] X out = (σ(Conv 1×1 ([AvgPool(X); MaxPool(X)]))) ⊙ X (23)
[0146] Among them, the input feature map AvgPool and MaxPool represent average pooling and max pooling, and both will have dimensions Conv 1×1 is a 1×1 convolutional kernel, σ is the Sigmoid function, used to normalize the attention weights between [0,1], ⊙ represents element-wise multiplication, multiply the generated spatial attention map with the original feature map element-wise (dot product) to strengthen the features in the key region and suppress the features in other regions, Xout Denote the output feature map, where the weighted features are more prominent in the key regions of the original feature map. By introducing this spatial attention mechanism, Cattle-PoseNet can accurately identify and predict the poses of cows in various environments, especially showing efficient performance in automated pasture management and veterinary diagnosis.
[0147] Then there is the multi-scale feature fusion module. The main function of this module is to effectively fuse the feature maps from different scales, so as to obtain a richer and more representative feature representation. This helps to improve the adaptability of the pose detection model to targets of different scales, enabling it to accurately detect key points on cows of different sizes and proportions. Let the feature maps from different levels be F 1 , F 2 , ……, F n , the calculation of the multi-scale feature fusion module can be expressed as:
[0148]
[0149] where F fusion is the fused feature map, and α i are the weights of each feature map, which are obtained through training.
[0150] The output convolutional layer of the Cattle-PoseNet algorithm is responsible for mapping the features learned by the network to the output space of the target detection task, that is, predicting the positions of the key points of the cow. This convolutional layer is usually a 1×1 convolutional kernel, which is used to map the feature map of the last layer of the network to the final key point coordinates.
[0151] Finally, an activation function is applied to limit the output range, ensure that the output is within a reasonable range, and make the output have appropriate non-linear characteristics. In Cattle-PoseNet, Softmax normalization is required for each pixel point to ensure that the sum of the probabilities of all pixel points of each key point is 1:
[0152]
[0153] where Y k represents the heat map of the k-th key point, X 4k is the fused feature map obtained by multi-scale feature fusion, and j is used to traverse all pixel positions. By doing so, the heat map of the cow's skeletal joint points can be output, and the position with the maximum joint point probability is extracted as the output of the cow's skeletal joint points.
[0154] S5. Convert the key points to poses
[0155] Based on the 17 bovine skeletal joints output by S4, connect them according to prior knowledge. Specifically: the head, left eye, and right eye are connected to form a triangle; the left shoulder is connected to the left arm, and the left arm is connected to the left hand to form a line. The same is done for the right shoulder, right arm, and right hand, the left hip, left leg, and left foot, and the right hip, right leg, and right foot; then the left shoulder and right shoulder are respectively connected to the chest, and the left hip and right hip are connected to the back; finally, connect the head to the chest and the chest to the back. In this way, the posture of the cow is completely extracted.
[0156] Compared with the prior art, the ELS operator proposed by the present invention can perform image denoising and enhancement processing on the noise caused by poor light in the image, so that the algorithm has better anti-noise ability.
[0157] The GM-Transformer proposed by the present invention introduces the GCOQ module, in which the graph attention network strengthens the context enhancement of information, so as to obtain better results in the case of occlusion between the cow and the background. At the same time, the proposed DETA module, compared with the self-attention module in the traditional Transformer, introduces a dynamic content-aware filtering module, which can make the attention mechanism pay more attention to the important or information-rich parts in the image, especially providing better results in the case of occlusion. The DMSA module of the present invention, compared with the traditional Transformer, introduces the network architecture FPN with multi-resolution output, which improves the output effect of the model for the case of more targets with different sizes.
[0158] Based on object detection, the present invention proposes a network called Cattle-PoseNet to extract bovine skeletal joints, introduces a depthwise separable convolution mechanism to reduce the computational amount, thereby reducing the resources consumed by the method and system of the present invention. At the same time, a spatial attention mechanism is introduced, which can pay more attention to the key region information to improve the extraction accuracy of skeletal key points. In addition, when extracting bovine skeletal joints, the present invention also introduces a multi-scale feature fusion module, which improves the applicability of the model to targets of different sizes and enables it to have good robustness in different environments.
[0159] As Figure 11 shown, the present invention also provides an application system for bovine skeletal joint extraction and posture recognition based on image processing, including:
[0160] Image acquisition module. The image acquisition module is a fundamental module in the bovine skeletal joint point extraction and pose recognition system. Its main responsibility is to obtain high-quality image or video data for subsequent image processing, joint point detection, and pose analysis. Regarding the image acquisition device, a resolution of at least 1080p (1920x1080) is preferred to ensure image clarity sufficient to distinguish fine joint points; the camera should have good low-light performance, especially when used in indoor environments with insufficient lighting or at night. Regarding the image acquisition environment, ensure sufficient and uniform lighting, and avoid strong backlighting or shadows, which may affect image quality and the accuracy of joint point detection; simplifying the background can reduce interference during image processing and improve the accuracy of joint point detection. Using a background wall with a unified color is a good choice.
[0161] The image preprocessing module is a crucial part of the bovine skeletal joint point extraction and pose recognition system, directly affecting the efficiency and accuracy of subsequent modules. The main purpose of this module is to improve image quality, reduce noise during processing, and convert the image size and format to adapt to the subsequent processing flow. Since different video recording devices are used in actual production processes and the output picture sizes are different, in this invention, the picture is first scaled to a picture of 720 pixels * 720 pixels in size to ensure that the input image size conforms to the input of the network. At the same time, the ELS operator is introduced to sharpen the image, making the image features input into the network more obvious.
[0162] The target detection module is responsible for extracting the feature map from the preprocessed image and separating the cow from the complex background for subsequent skeletal joint point extraction. In the part of extracting the feature map, this invention introduces an improved ResNet network based on the traditional ResNet network, named SE-SPP ResNet, which consists of an initial convolutional layer, a max pooling layer, a residual block group, a squeeze and excitation block, grouped convolution, a spatial pyramid pooling layer, etc. Through the combination of these technologies, the SE-SPP ResNet model not only has a significant improvement in the accuracy and efficiency of feature extraction but also shows excellent performance and adaptability when processing high-resolution images, which is crucial for the bovine skeletal joint point extraction and pose recognition tasks; in the part of separating the cow from the background, this invention proposes an improved object detection algorithm based on the DETR algorithm, named GCMT algorithm, which includes an improved transformer structure, named GM-Transformer, changing the generation method of Object Queries and the structures of the multi-head self-attention mechanism module and the multi-head attention mechanism module to make it more suitable for the bovine skeletal joint point extraction task of this invention.
[0163] The pose recognition module is responsible for extracting skeletal joint points from the region of interest output by the target detection module and connecting them into the pose of the cattle's movement. In the part of skeletal joint point extraction, according to the characteristics of cattle skeletal joint points, the invention divides the joint points on the cattle skeleton into 17 points and establishes a new type of network to extract 17 skeletal joint points, which is named Cattle-PoseNet. This network is trained with cattle pictures to make its performance more prominent in extracting cattle skeletal joint points. After identifying 17 skeletal joint points, they are connected according to prior knowledge to output the final pose.
[0164] The algorithm and application system for cattle skeletal joint point extraction and pose recognition based on image processing of the present invention solve the problems of low accuracy, environmental condition limitations and image quality requirements of traditional technologies, do not require expensive equipment, have low usage costs, do not require manual monitoring, and are convenient to apply.
[0165] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for extracting joint points and recognizing postures of cattle bones based on visual sensing, characterized by: The following steps are involved: Step 1, collecting video images of cattle, normalizing the images, and performing image preprocessing; Step 2, feature extraction is performed on the preprocessed image I, and the feature extraction network adopts an improved ResNet network. The improved ResNet network combines the SE mechanism and the spatial pyramid pooling SPP, embeds the SE block after the residual block, and adds the spatial pyramid pooling layer to the last fully connected layer; Step 3: Based on the extracted features, the target detection algorithm GCMT is used to detect the target cow, identify the position of the target cow, obtain the detection frame of the cow, and crop the cow image according to the detection frame for later use; The target detection algorithm GCMT decomposes the features extracted in step 2 into a two-dimensional feature vector, adds the position information, and outputs the prediction result of the target through the GM Transformer module. The GM Transformer module is based on the DETR target detection model, encapsulates the generation of query objects into a separate module GCOQ, and replaces the first-level multi-head self-attention module in the encoder and decoder structure with a dynamic resolution enhanced attention module DETA, and replaces the second-level self-attention module in the encoder structure with a dynamic multi-scale attention module DMSA; in GCOQ introduces a graph structure between query objects and uses the graph attention network GAT to update query objects. DETA adds a dynamic filter layer to the multi-head self-attention, dynamically strengthening or weakening certain attention connections based on the position distance in the feature map; DMSA assigns weights to feature maps of different scales of the image based on global and local information, combines the weights to fuse feature maps of different scales, and then performs self-attention calculations; Step 4, extracting the joint points of the cow bones within the detection frame; Step 5: Convert the detected key points into the cow’s posture.
2. The method for extracting and recognizing the joint points of cattle bones based on visual sensing according to claim 1 is characterized in that Step 1 is as follows: The collected original image I o Scale to a uniform pixel size. If the image pixel is lower than the uniform pixel, use the enhanced Laplace sharpening ELS operator to sharpen the image and enhance the image features: The ELS operator is a linear combination of a two-dimensional Laplacian operator L and an edge enhancement high-pass filter H to emphasize both boundary and texture details: ELS=αL+βH Where α and β are adjustable weights, applying the ELS operator to the image I o The sharpened image I is obtained sharp as follows: I sharp =I o +c·(ELS*I o ) Where c is the coefficient for adjusting the sharpening strength, and the value of c is between 0.3 and 0.
5.
3. The method for extracting and recognizing the joint points of cattle bones based on visual sensing according to claim 1 is characterized in that In step 2, the improved ResNet network is called SE-SPP ResNet, which includes an initial convolution layer, a maximum pooling layer, two layers of residual block groups + SE modules, one layer of residual block groups + grouped convolution, one layer of residual block groups + SE modules, and a spatial pyramid pooling layer + fully connected layer configured in sequence.
4. The method for extracting and recognizing the joint points of cattle bones based on visual sensing according to claim 3, characterized in that SE -SPP ResNet is specifically: First, use the initial convolution layer to reduce the image size and extract preliminary spatial features. After the initial convolution, use the maximum pooling layer to further reduce the image size and complete the secondary dimensionality reduction of the image. In the residual block group + SE module, the operations in the residual block group include two convolutional layers, each of which is followed by a batch normalization layer and a ReLU activation function. The SE module first performs global average pooling on the feature map output by the residual block group to achieve compression, and then activates it through a fully connected layer, recalibrates the weight of each channel, and then multiplies the recalibrated weight back to the original feature map for Scale operation. The feature map adjusted by the SE block is output and superimposed with the feature map of the original input as the output feature map of the residual block group + SE module; In the residual block group + group convolution, the input feature map X is grouped, each group is convolved independently, and then the convolution results of all groups are merged into a single feature map. Batch normalization and ReLU activation function are applied to the merged feature map, and the activated feature map is directly added to the input feature map X to form the output; It then passes through a layer of residual block group + SE module, followed by a combination of a spatial pyramid pooling module and a fully connected layer to adjust the size of the image and output the final feature image, that is, the feature map extracted from the preprocessed image I.
5. The method for extracting joint points and recognizing postures of cattle bones based on visual sensing according to claim 1, characterized in that G The COQ modules are: First, define the query object Object Queries as the node of the graph, and set the initial Object Queries matrix to: Where N is the number of query objects, D is the feature dimension of each query object, and these query objects are initialized based on the features extracted in step 2 to construct a graph: G(V,E) Among them, V is a set of nodes, corresponding to the query objects; E is a set of edges, representing the potential associations between the query objects. These edges are based on the similarity calculated by metrics such as Euclidean distance or cosine similarity. Use the graph attention network GAT to update the features of the query object Object Queries. For each node m in the graph, its updated feature h ′ m It is calculated as follows: Among them, h n is the original or previous iteration feature of node m, N(m) represents the set of neighbor nodes of node m, σ is the activation function, α mn is the attention coefficient between nodes m and n, calculated as follows: W is the learnable weight matrix applied to each node feature and a is the learnable parameter vector used to calculate the attention coefficient.
6. The method for extracting and recognizing bovine bone joints based on visual sensing according to claim 1 is characterized in that The dynamic resolution enhanced attention module DETA includes a multi-head self-attention module and dynamic content-aware filtering, which is implemented as follows: Among them, F l (u,v) is the weight function of level l position u relative to v, Q l , K l 、V l Represent the query, key, and value matrices of level l respectively. For normalization, F l (u,v) is a distance-based filter implemented by a Gaussian function: Where P u and P v are the coordinates of positions u and v on the feature map, and σ is a parameter that controls the decay speed.
7. The method for extracting and recognizing the joint points of cattle bones based on visual sensing according to claim 1 is characterized in that: In the multi-scale attention module DMSA, for feature maps X of different scales l Use a neural network to calculate the fusion weight α based on global and local information l , these weights determine the contribution of each scale feature map in the final fusion, and the feature fusion formula is as follows: Where X′ is the fused feature map, Transform(X l ) represents a set of transformation operations to make the dimensions of feature maps of all scales consistent, and use the fused feature maps to generate the values and keys in the attention calculation for self-attention calculation.
8. The method for extracting and recognizing the joint points of cattle bones based on visual sensing according to claim 1 is characterized by: Step 4 is as follows: Divide the joints on the cow skeleton into 17 points, and establish a Cattle-PoseNet network to extract the 17 bone joints. The Cattle-PoseNet network includes basic feature extraction, deep separable convolution, spatial attention operation, and multi-scale feature fusion, and then outputs a heat map after convolution and activation function: First, the image in the cow detection box generated in step 3 is scaled to a uniform size, and the improved ResNet network is used for feature extraction. Then, the image is fed into the deep separable convolution module to realize feature integration and dimension transformation. After the deep separable convolution module, the spatial attention module is input. The feature map is first compressed along the channel direction through the average pooling and maximum pooling layers, and each feature map is simplified into a one-dimensional response to aggregate global spatial information. Then, the outputs of the average pooling and maximum pooling are merged and processed through one or more convolutional layers to generate a spatial attention map. Then, the value of the attention map is normalized to between 0 and 1 through the activation function to generate the attention weight of each spatial position. The multi-scale features are fused by combining the attention weights, and the fused features are convolutionally mapped to the output space of the target detection task, that is, the key point positions of the cow are predicted. Finally, an activation function is applied to limit the output range and Softmax normalization is performed on each pixel to ensure that the sum of all pixel probabilities for each key point is 1: where Y k represents the heat map of the kth key point, X 4k is the fusion feature, j is used to traverse all pixel positions.
9. The method for extracting and recognizing bovine bone joints based on visual sensing according to claim 1, characterized in that In step 5, the key points of the cow skeleton are connected according to prior knowledge to obtain the posture of the cow.
10. A cattle bone joint point extraction and posture recognition system based on visual sensing, characterized by The method for extracting and recognizing bovine bone joints based on visual sensing as claimed in claim 1 comprises: An image acquisition module is used to acquire video images of cattle. The image acquisition device needs to have a resolution of at least 1080p to ensure image clarity. The image preprocessing module is used to implement the image preprocessing in step 1, and is used to scale the image to a size of 720 pixels * 720 pixels to ensure that the input image size meets the network input. At the same time, the ELS operator is introduced to sharpen the image to make the image features of the input network more obvious; A target detection module, used to implement the feature extraction of step 2 and the target detection of cattle in step 3, and used for the subsequent extraction of bone joints. The target detection module is loaded with a computer program to implement the improved ResNet network and the target detection algorithm GCMT in the method of claim 1. The improved ResNet network is used to extract feature maps, and GCMT is used to implement target detection based on the feature maps to obtain the bone joints of cattle; The posture recognition module is used to implement steps 4 and 5, connecting the bone joints output from the target detection module into the posture output of the cow's movement. In the bone joint extraction part, according to the characteristics of the cow's bone joints, the joints on the cow's skeleton are divided into 17 points, and a new type of network is established to extract the 17 bone joints. After identifying the 17 bone joints, they are connected according to prior knowledge to output the final posture.
Citation Information
Patent Citations
Animal behavior identification method based on Pose-Transform network
CN115830713A
Animal skeleton key point detection and animal posture recognition method and system
CN116469128A
Attitude reconstruction interaction behavior understanding method based on skeleton and image features
CN117238026A
Cited By
PLFuNet-based end-to-end adaptive visual feature extraction method
CN120563956A
Transform lightweight model-based cattle health assessment method
CN120998498A
Cow posture recognition method based on triaxial angular velocity signal of gyroscope
CN121542858A
A method for recognizing posture of a cow based on three-axis angular velocity signals of a gyroscope
CN121542858B
Image recognition method of robot for home environment
CN121661622A