Cow face recognition method based on dense point cloud three-dimensional reconstruction
Through the three-dimensional reconstruction of dense point clouds, the problems of external pathogen invasion, susceptibility to replacement and inefficiency in existing cattle identity recognition technology are solved, efficient and accurate cattle identity recognition is achieved, and the smooth development of insurance claims and loans is promoted.
Patent Information
- Application Number
- CN202411942368.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing cattle identity identification technology has problems such as external pathogen invasion, ease of replacement, and inefficiency, making it difficult to achieve accurate individual cattle identity identification, affecting the smooth development of insurance claims and loans.
The cow face recognition method based on three-dimensional reconstruction of dense point clouds is adopted. By receiving cow pictures from different perspectives, the improved ALIKED lightweight network is used to extract key points and descriptors, combined with the attention-enhanced feature point matching network and 3D convolutional neural network, dense point cloud coordinates are generated, and identity recognition is performed through cosine similarity calculation.
It improves the accuracy and efficiency of cow face recognition, can better pay attention to significant features, simplify feature extraction steps, improve model operation speed, and ensure the accuracy and consistency of identity recognition.
Smart Images

Figure CN119992588A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional vision, and in particular to a cow face recognition method based on dense point cloud three-dimensional reconstruction. Background Art
[0002] With the development of large-scale integrated dairy and beef cattle farming, live cattle have become the largest asset component of farming enterprises and herders. However, due to factors such as epidemic diseases and cattle identity confirmation, it is difficult to purchase insurance or asset mortgage loans for live cattle as assets, and inclusive finance is difficult to reach. Currently, the commonly used cattle identification technologies include ① embedded RFID glass tube tags, which have problems such as external pathogen invasion, easy migration of glass tubes in the cattle body, difficulty in finding and removing them during slaughter, and easy replacement during insurance claims; ② liquid nitrogen numbering technology, which is only applicable to shelf cattle after 10 months of age, and manual branding and numbering are time-consuming and labor-intensive, and inefficient; ③ hanging ear tags, which are easy to be replaced and cannot be used as a basis for confirming the identity of cattle during insurance claims. In order to effectively control diseases and achieve precise feeding during cattle breeding, identifying individual cattle and connecting insurance and loans has become the focus and difficulty of future intelligent and integrated ranch research. Accurate identification of individual identities will promote the smooth development of beef cattle insurance, solve the problem that the claiming beef cattle and the insured beef cattle cannot be fully matched, and improve the accuracy of underwriting claims. It can be seen that accurate identification of individual cattle identities has become the focus and difficulty of future intelligent and integrated ranch research. It is also a key technology for finding support for beef cattle disease prevention and safety insurance applications. It has extremely high theoretical research significance and application value. Summary of the invention
[0003] Purpose of the invention: In order to overcome the shortcomings of the prior art, the present invention provides a cow face recognition method based on dense point cloud three-dimensional stereo reconstruction.
[0004] The technical solution adopted by the present invention is:
[0005] A cow face recognition method based on dense point cloud three-dimensional reconstruction specifically comprises the following steps:
[0006] S1, receives pictures of cows from different perspectives to capture the 2D feature points of the cows;
[0007] S2, using 3D reconstruction technology, converts pictures of cows from different perspectives into 3D dense point cloud coordinates to ensure the efficiency and accuracy of facial recognition;
[0008] S3, extract the constructed 3D dense point cloud features and perform cow face comparison.
[0009] In step S2, 2D features of cow images from different perspectives are extracted and converted into 3D feature volumes, and then dense 3D point cloud coordinates are generated, including the following processes:
[0010] S21, extracting cow image key points and descriptors through the improved ALIKED lightweight network;
[0011] The use of an improved lightweight feature point extraction network has the following advantages: First, using feature points for 3D reconstruction helps to focus on local significant features; second, using a lightweight feature extraction network can greatly speed up the processing while maintaining accuracy; third, improving the ALIKED network and using dynamic activation and enhanced linear transformation activation functions can improve the nonlinearity of the network and capture complex feature relationships.
[0012] First, the primary features of the input image are extracted through a series of convolutional layers and pooling layers, and then passed to a differentiable key point detection convolutional layer to detect the key features of cow face recognition. The calculation process is as follows:
[0013]
[0014] where w(i,j) is the weight, x is the input feature map, and Δx(i,j) and Δy(i,j) are offsets relative to the standard grid.
[0015] Then, for each detected key point, a sparse deformable descriptor head SDDH is used to learn the deformable sampling position and construct a deformable descriptor. The feature points are used for processing in order to better focus on the local significant features of cow face recognition, namely the nose pattern, the swirl on the top of the cow's head, and the patterns on both sides of the cow. In addition, the activation function of ALIKED is replaced by the following A(x) to enhance the nonlinear expression ability of the network and effectively reduce the inference time.
[0016] A(x)=(1-λ)ReLU(x)+λx
[0017] Where ReLU(x) is the original activation function A(x), and λ is a hyperparameter used to balance the nonlinearity of the modified activation function A(x).
[0018] The working process includes the following steps:
[0019] S211, extracts the primary features of the image through a series of convolutional layers and pooling layers;
[0020] S212, using a differentiable keypoint detection convolutional layer to determine the keypoint locations in the image;
[0021] S213, for each detected key point, a deformable sampling position is learned through a sparse deformable descriptor head SDDH, and a deformable descriptor is constructed.
[0022] Finally, the detected feature points and their descriptors are output for feature point matching.
[0023] S22, feature point matching based on attention-enhanced feature point matching network (AGMNet) to provide accurate correspondence for 3D feature body construction;
[0024] First, a convolutional neural network combined with an attention mechanism is used to enable the feature points to communicate with each other, so as to calculate the matching descriptors. The attention mechanism is added here to enhance the matching process, which can deepen the weight of the significant features of cow face recognition, highlight the cow's nose pattern, the whorl on the top of the cow's head, and the patterns on both sides of the cow. Then, the optimal transmission theory is used to deal with the feature point matching problem, and a score matrix is calculated to represent the similarity between different feature points. Then, the multi-scale sparse Sinkhorn algorithm is used to obtain the optimal match.
[0025] S23, using homography transformation to convert 2D feature points into points in 3D space, and constructing a 3D feature body;
[0026] The RANSAC algorithm is used to iteratively estimate the homography matrix. The number of iterations can be estimated by the following formula to convert the 2D feature points into 3D feature volumes.
[0027]
[0028] Among them, p is the probability that RANSAC will get the correct model, w is the proportion of inliers in the data, and n is the number of randomly selected points in each iteration.
[0029] S24, uses 3D convolutional neural network (3D CNN) to generate depth map directly from 3D feature volume;
[0030] First, the 3D convolution kernel slides in the feature volume to extract the spatial and depth features of the 3D feature volume, and then an upsampling layer predicts the depth information and maps it to the output space together with the extracted features to generate a depth map.
[0031] S25, generates a dense point cloud based on the depth map, each point containing its coordinates in the 3D space.
[0032] For each pixel (u, v) in the depth map, the corresponding 3D coordinate (X, Y, Z) is calculated as follows:
[0033]
[0034]
[0035]
[0036] Where d is the depth value in the depth map, depth scale factor is the scaling factor of the depth value, which is usually used to convert the value of the depth map into actual distance units, (fx, fy) is the camera focal length, and (cx, cy) is the camera optical center.
[0037] In step S3, the constructed 3D dense point cloud features are extracted, and the cow face recognition is performed by calculating the cosine similarity between vectors, including the following process:
[0038] S31, using the improved RandLA-Net network to extract 3D high-level features from dense point clouds
[0039] Using an end-to-end network structure, tedious feature engineering and modular design are avoided in the 3D feature extraction process, the entire extraction process is simplified, and the global relationship between input and output can be better captured, thereby improving system performance.
[0040] The improved RandLA-Net network is mainly divided into two stages: encoding and decoding. Point features are extracted and enhanced through four encoding layers. Each encoding layer consists of a local feature aggregation module and an attention pooling layer. 25% of the point cloud is retained by downsampling, the number of point clouds is gradually reduced, and the feature dimension of each point is increased. Then, through 4 decoding layers, the nearest neighbor interpolation and MLP are used to upsample the features, restore the number of point clouds, and fuse the feature information from the encoding layer. Finally, a fully connected layer and Dropout operation are used to fuse the global context to obtain the final semantic label of each point.
[0041] S32, the features extracted by the improved RandLA-Net are further processed through a fully connected layer to map the features to a low-dimensional space;
[0042] A fully connected layer is used to map the high-dimensional feature vector from S31 into a 256-dimensional space, which facilitates the comparison of the feature vectors of two Niuniuniu.
[0043] S33, using the output of the fully connected layer to calculate the cosine similarity, to identify whether the two cows have the same identity.
[0044] After obtaining the 3D feature vector of the cow, it is compared with other feature vectors in the library through cosine similarity to determine whether the two cows are the same. The cosine similarity calculation formula is as follows.
[0045]
[0046] Where a and b are two eigenvectors, and γ is the value of cosine similarity, ranging from -1 to 1, where 1 means that the two vectors are exactly the same, -1 means they are completely opposite, and 0 means that the two vectors are orthogonal (no relationship).
[0047] Beneficial effects:
[0048] 1. The present invention performs three-dimensional reconstruction based on 2D feature points. Compared with 2D feature maps, it is more conducive to focusing on the significant distinguishing features of cow face recognition, namely, the nose pattern of the cow, the whorl on the top of the cow's head, and the patterns on both sides of the cow. The feature map emphasizes global features more.
[0049] 2. The 2D feature point extraction network used in the present invention is improved based on the ALIKED lightweight network. The activation function is replaced by a nonlinear activation function. The expression power of the activation function is fully utilized during training, and the network reasoning speed is improved by merging convolutional layers in the reasoning stage.
[0050] 3. The present invention adopts a feature point matching network based on attention enhancement, which not only ensures the matching accuracy, but also pays more attention to the significant distinguishing features of cow face recognition.
[0051] 4. The present invention improves the RandLA-Net network to extract 3D point cloud features, simplifies the extraction steps while ensuring accuracy, helps to improve the model running speed and accelerate reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Shown is a flow chart of the present invention;
[0053] Figure 2 The figure shows the network architecture of the cow face recognition method based on dense point cloud three-dimensional reconstruction of the present invention. Specific implementation methods
[0055] The present invention will be further described below in conjunction with the accompanying drawings.
[0056] The present invention will be further described below in conjunction with examples.
[0057] like Figure 1 As shown in the figure, a cow face recognition method based on dense point cloud 3D reconstruction first receives 4 cow pictures from different perspectives, uses 3D reconstruction technology to extract feature points from them respectively, then performs feature point matching to facilitate the construction of 3D point cloud coordinates, and then extracts 3D features from the dense point cloud to obtain the most comprehensive and representative feature vector, and then performs cow face recognition. Specifically, it includes the following steps:
[0058] S1,receives four images of the cow from different perspectives to capture the detailed features of the cow.
[0059] Input four pictures of cows from different perspectives into the network, taken from four directions: front, top, left, and right.
[0060] S2, using 3D reconstruction technology, extracts and matches feature points from pictures of Niu Niu from different perspectives, and reconstructs them into 3D dense point cloud coordinates, which specifically includes the following steps:
[0061] S21, extracting key points and descriptors of Niu Niu images through the improved ALIKED lightweight network;
[0062] First, the local features of the input image are extracted through a series of convolutional layers and pooling layers. In this process, the dimension of the extracted feature map gradually decreases, while the feature dimension gradually increases to retain more information. Then it is passed to a differentiable key point detection convolutional layer to detect the key feature points for cow face recognition. The key point detection calculation process is as follows:
[0063]
[0064] where w(i,j) is the weight, x is the input feature map, and Δx(i,j) and Δy(i,j) are offsets relative to the standard grid.
[0065] For each detected key point, a sparse deformable descriptor head SDDH is used to learn the deformable positions of the surrounding features. This process involves calculating the local features around the key points and adjusting the sampling positions based on these features to capture more accurate local structural information. After determining the deformable positions, the SDDH module extracts descriptors at these positions. These descriptors can capture the local features around the key points. The feature points are used for processing in order to better focus on the local significant features of cow face recognition, namely the nose print, the whorl on the top of the cow's head, and the patterns on both sides of the cow. Finally, the detected feature points and their descriptors are output for feature point matching.
[0066] S22, feature point matching based on attention-enhanced feature point matching network (AGMNet) to provide accurate correspondence for 3D feature body construction;
[0067] First, the feature points and descriptors obtained by S21 are sent to a convolutional neural network based on the attention mechanism enhancement, so that the feature points can communicate with each other and calculate the matching descriptors. The core of the communication between feature points is to aggregate the information of feature points through the attention mechanism, calculate the similarity between feature points using the attention mechanism, and aggregate features based on this. The calculation process can be expressed as:
[0068]
[0069] Where m ε→i represents the aggregated features of feature point i, α ij Represents the similarity weight between feature points i and j.
[0070] α ij =Softmax j (scaled_scores ij )
[0071] where scaled_scores ij Represents the scaled dot product.
[0072]
[0073] Among them scores ij is the dot product of the query and the key, expressed as:
[0074] scores ij =Q i K j T
[0075] Among them, Q i is the query, which represents the feature vector of the feature point to be queried. K is the key, which represents the feature vector of the source feature point. k is the dimension of the key vector, Used to scale the dot product to prevent the vanishing gradient problem.
[0076] The update of feature points is achieved through convolutional neural network (CNN). During the update process, the weight w of the convolution layer will be updated through the back propagation algorithm, and the updated gradient can be expressed as:
[0077]
[0078] Where L is the loss function.
[0079] Then use the gradient descent algorithm to update the feature points:
[0080] w k+1 =w k -αΔw
[0081] Where α is the learning rate.
[0082] The update formula of feature points is as follows:
[0083] x k+1 =x k +w k+1 ·D
[0084] Among them, x k+1 is the updated feature point state, x k is the state before the feature point is updated, and D is the feature point descriptor.
[0085] Then, the optimal transmission theory is used to deal with the feature point matching problem, a score matrix is calculated to represent the similarity between different feature points, and then the multi-scale sparse Sinkhorn algorithm is used to obtain the optimal match. The iterative formula of the Sinkhorn algorithm is as follows:
[0086] Z=log_sinkhorn_iterations(Z,log_mu,log_nu,iters)
[0087] Where log_mu and log_nu are the logarithmic distributions of rows and columns, and iters is the number of iterations.
[0088] S23, using homography transformation to convert 2D feature points into points in 3D space, and constructing a 3D feature body;
[0089] The RANSAC algorithm is used to iteratively estimate the homography matrix. The number of iterations can be estimated by the following formula to convert the 2D feature points into 3D feature volumes.
[0090]
[0091] Among them, p is the probability that RANSAC will get the correct model, w is the proportion of inliers in the data, and n is the number of randomly selected points in each iteration.
[0092] S24, uses 3D convolutional neural network (3D CNN) to generate depth map directly from 3D feature volume;
[0093] Based on the 3D feature volume obtained above, the spatial and depth features of the 3D feature volume are first extracted by sliding the 3D convolution kernel in the feature volume, and then the depth information is predicted through an upsampling layer and mapped to the output space together with the extracted features to generate a depth map.
[0094] S25, generates a dense point cloud based on the depth map, each point containing its coordinates in the 3D space.
[0095] For each pixel (u, v) in the depth map, the corresponding 3D coordinate (X, Y, Z) is calculated as follows:
[0096]
[0097]
[0098]
[0099] Where d is the depth value in the depth map, depth scale factor is the scaling factor of the depth value, which is usually used to convert the value of the depth map into actual distance units, (fx, fy) is the camera focal length, and (cx, cy) is the camera optical center.
[0100] S3, extracting the constructed 3D dense point cloud features and performing facial comparison, specifically including the following steps:
[0101] S31, using the improved RandLA-Net network to extract 3D high-level features from dense point clouds
[0102] It is mainly divided into two stages, encoding and decoding. Point features are extracted and enhanced through four encoding layers. Each encoding layer consists of a local feature aggregation module and an attention pooling layer. 25% of the point cloud is retained by downsampling, the number of point clouds is gradually reduced, and the feature dimension of each point is increased. Then, through 4 decoding layers, the nearest neighbor interpolation and MLP are used to upsample the features, restore the number of point clouds, and fuse the feature information from the encoding layer. Finally, a fully connected layer and Dropout operation are used to fuse the global context to obtain the final semantic label of each point.
[0103] S32, further processes the features extracted by the improved RandLA-Net through a fully connected layer to map the features to a low-dimensional space;
[0104] A fully connected layer is used to map the high-dimensional feature vector from S31 into a 256-dimensional space, which facilitates the comparison of the feature vectors of two Niuniuniu.
[0105] S33, using the output of the fully connected layer to calculate the cosine similarity, to identify whether the two cows have the same identity.
[0106] After obtaining the 3D feature vector of the cow, it is compared with other feature vectors in the library through cosine similarity to determine whether the two cows are the same. The cosine similarity calculation formula is as follows.
[0107]
[0108] Where a and b are two eigenvectors, and γ is the value of cosine similarity, ranging from -1 to 1, where 1 means that the two vectors are exactly the same, -1 means they are completely opposite, and 0 means that the two vectors are orthogonal (no relationship).
[0109] If the cosine similarity of the 3D feature vectors of two cows reaches 0.95 or above, the two cows should be judged to be the same identity, and the name of the cow to be identified in the system library should be output; if there is no cow in the system library whose cosine similarity with the cow to be identified reaches 0.95, then unknow is output.
[0110] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A cow face recognition method based on dense point cloud 3D reconstruction, characterized in that The following steps are involved: S1, receives pictures of cows from different perspectives to capture 2D feature points of cows; S2, using 3D reconstruction technology, converts pictures of cows from different perspectives into 3D dense point cloud coordinates to ensure the efficiency and accuracy of facial recognition; S3, extract the constructed 3D dense point cloud features and perform cow face comparison.
2. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 1, characterized in that: S2 extracts the 2D feature points of the cow image input in S1 and completes the construction of the 3D feature body, thereby generating a 3D dense point cloud, including the following processes: S21, extracting cow image key points and descriptors through the improved ALIKED lightweight network; S22, feature point matching based on attention-enhanced feature point matching network layer, providing accurate correspondence for 3D feature body construction; S23, using homography transformation to convert 2D feature points into points in 3D space, using RANSAC algorithm to estimate the homography matrix, and then converting the 2D feature points into 3D feature bodies; S24, uses a 3D convolutional neural network to generate a depth map directly from a 3D feature volume; S25, generates a dense point cloud based on the depth map, each point containing its coordinates in the 3D space.
3. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 2, characterized in that In step S21, the activation function of ALIKED is replaced by the following A(x) in the improved ALIKED lightweight network: A(x)=(1-λ)ReLU(x)+λx Where ReLU(x) is the original activation function of ALIKED, and λ is a hyperparameter used to balance the nonlinearity of the modified activation function A(x). The working process includes the following steps: S211, extracts the primary features of the image through a series of convolutional layers and pooling layers; S212, using a differentiable keypoint detection convolutional layer to determine the keypoint locations in the image; S213, for each detected key point, a deformable sampling position is learned through a sparse deformable descriptor head SDDH, and a deformable descriptor is constructed.
4. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 2, characterized in that The specific steps of step S22 are as follows: S221, a convolutional neural network combined with an attention mechanism enables feature points to communicate with each other, thereby calculating matching descriptors; S222, uses the optimal transmission theory to deal with the feature point matching problem, calculates a score matrix to represent the similarity between different feature points, and then obtains the optimal match through the multi-scale sparse Sinkhorn algorithm.
5. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 2, characterized in that In step S24, a 3D convolutional neural network is used to directly generate a depth map from the 3D feature volume to supplement the lost depth information in the 2D image. The working process includes the following steps: S241, extracts spatial and depth features by sliding the 3D convolution kernel in the feature volume; S242, predict the depth information through an upsampling layer and map it to the output space together with the extracted features to generate a depth map.
6. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 2, characterized in that In step S25, the process of generating dense point cloud coordinates according to the depth map includes the following steps: For each pixel (u, v) in the depth map, the corresponding 3D coordinate (X, Y, Z) is calculated as follows: Where d is the depth value in the depth map, depth scale factor is the scaling factor of the depth value, which is usually used to convert the value of the depth map into actual distance units, (fx, fy) is the camera focal length, and (cx, cy) is the camera optical center.
7. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 1, characterized in that In step S3, high-level feature extraction is performed on the 3D dense point cloud constructed in S2, and it is converted into a 256-dimensional feature vector to facilitate the comparison of cow facial features. The specific steps are as follows: S31, using the improved RandLA-Net network to extract 3D high-level features from dense point clouds S32, the features extracted by the improved RandLA-Net are further processed through a fully connected layer to map the features into a 256-dimensional space; S33, using the output of the fully connected layer to calculate the cosine similarity, to identify whether the two cows are of the same identity.
8. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 7, characterized in that In step S31, an improved RandLA-Net network is used to extract 3D high-level features from dense point clouds, and point features are enhanced to highlight important features that are beneficial to cow face recognition. The working process includes the following steps: S311, extracts and enhances point features through four encoding layers, each of which consists of a local feature aggregation module and an attention pooling layer, gradually reducing the number of point clouds and increasing the feature dimension of each point; S312, through 4 decoding layers, uses nearest neighbor interpolation and MLP to upsample features and fuse feature information from the encoding layer.
9. The method for cow face recognition based on dense point cloud three-dimensional reconstruction according to claim 7, characterized in that The specific steps of step S33 are as follows: After obtaining the 3D feature vector of the cow, it is compared with other feature vectors in the library through cosine similarity to determine whether the two cows are of the same identity. The cosine similarity calculation formula is as follows: Where a and b are two eigenvectors, and γ is the value of cosine similarity, ranging from -1 to 1, where 1 means that the two vectors are exactly the same, -1 means they are completely opposite, and 0 means that the two vectors are orthogonal.
Citation Information
Cited By
A multi-modal information fusion-based cattle identity recognition method
CN122715212A