Method, device and system for individual identity recognition of farmed animals by fusing local topological invariance and metric learning

CN120808386BActive Publication Date: 2026-09-25SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510664381.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2026-09-25
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

[0004]现有方法多使用传统人工特征提取算法或者简单的卷积神经网络进行特征提取,这些方法对奶牛图像中的关键特征(如花纹、轮廓、纹理等)提取和表达能力有限,尤其是在复杂环境下,如多视角、光照变化或背景干扰时,难以保证识别的准确性和鲁棒性

Benefits of technology

[0058]1.本发明不仅采用融合注意力机制(CSA)的改进型卷积神经网络提取奶牛个体的精细视觉外观特征,还创新性地并行引入了局部模式拓扑不变性(Local PatternTopology Invariance,LPTI)提取路径,专门用于捕捉对姿态和视角变化引起的非刚性形变具有鲁棒性的结构化拓扑特征。通过有效融合这两种具有互补优势的特征,生成了更全面的增强个体表征,显著提高了模型对于姿态、视角变化剧烈情况下的特征提取鲁棒性和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808386B_ABST
    Figure CN120808386B_ABST
Patent Text Reader

Abstract

The application discloses a kind of fusion local topological invariance and metric learning's breeding animal individual identity recognition method, equipment and system, method includes: (1) obtaining animal image, constructs sample data set;(2) construct feature extraction deep network model, the feature extraction deep network model includes basic feature extraction path, local mode topological invariance extraction path, feature fusion module;(3) using sample data set to train the model;(4) using the model trained, extract the enhanced embedding vector of several known identity's animal image, and store in feature database;(5) the animal image to be identified is input into the trained model to extract its enhanced embedding vector, calculate the similarity between its and the enhanced embedding vector stored, select the highest similarity enhanced embedding vector corresponding identity as the animal individual identity to be identified.The application is more accurate and robust.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision technology, and more particularly to a method, device, and system for identifying individual farmed animals by integrating local topological invariance and metric learning. Background Technology

[0002] Animal identification technology is crucial for the daily management of farms, such as monitoring the health status of dairy cows, recording milk production, and controlling diseases. Traditional contact-based methods, such as ear tags and radio frequency identification (RFID), are prone to damage or loss due to cow activity or equipment malfunctions, and their installation and maintenance costs are high.

[0003] In recent years, computer vision-based methods for identifying individual dairy cows have been extensively studied. Compared with traditional contact methods, computer vision technology has the characteristics of non-contact, remote identification, high efficiency and automation, which can significantly reduce costs and improve the intelligence level of farms.

[0004] Existing methods mostly use traditional manual feature extraction algorithms or simple convolutional neural networks for feature extraction. These methods have limited ability to extract and represent key features (such as patterns, contours, and textures) in dairy cow images, especially in complex environments such as multi-viewpoints, varying lighting, or background interference, making it difficult to guarantee the accuracy and robustness of recognition. Therefore, there is a need to improve existing computer vision-based methods for individual identification of farmed animals (such as dairy cows). Summary of the Invention

[0005] To address the problems existing in the prior art, the purpose of this invention is to provide a method, device, and system for identifying individual farmed animals that integrates local topological invariance and metric learning, which has high accuracy and robustness.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0007] A method for identifying individual farmed animals that integrates local topological invariance and metric learning includes the following steps:

[0008] (1) Obtain animal images from different perspectives and postures, and preprocess them to construct a sample dataset;

[0009] (2) Construct a feature extraction deep network model, the feature extraction deep network model including:

[0010] The basic feature extraction path is used to extract convolutional feature vectors from animal images using a ResNet 50 network with a fusion attention mechanism.

[0011] A local pattern topology invariant extraction path is used to identify key local pattern points of the convolutional feature vector, construct a local graph between key local pattern points, and process the local graph using a graph neural network to extract topological feature vectors.

[0012] The feature fusion module is used to fuse the convolutional feature vector with the topological feature vector to generate an enhanced embedding vector;

[0013] (3) The feature extraction deep network model is trained using the sample dataset;

[0014] (4) Using the trained feature extraction deep network model, extract the enhanced embedding vectors of several animal images with known identities, and store the enhanced embedding vectors and corresponding identities in the feature database;

[0015] (5) Input the image of the animal to be identified into the trained feature extraction deep network model to extract its enhanced embedding vector, calculate the similarity between the enhanced embedding vector and the enhanced embedding vector stored in the feature database, and select the identity corresponding to the enhanced embedding vector with the highest similarity in the feature database as the identity of the individual animal to be identified.

[0016] Furthermore, the ResNet 50 network with the fused attention mechanism specifically adds a channel-spatial attention module to the end of each residual module in the ResNet 50 network. The residual module is used to perform the following calculations:

[0017]

[0018] Where X″ represents the output of the residual module, X represents the input of the residual module, and Conv 1×1 Represents 1×1 convolution, BN represents batch normalization, Conv 3×3 This indicates a 3×3 convolution, PReLU represents the Parametric ReLU activation function, CSA represents the channel-spatial attention module, and + indicates a residual connection. This indicates the composition of functions.

[0019] Furthermore, the channel-spatial attention module includes a channel attention module and a spatial attention module, wherein the channel attention module is used to perform the following calculations:

[0020]

[0021] Among them, X CA F represents the output of the channel attention module. in This represents the input to the channel attention module, GAP represents the global average pooling operation, and Conv... 1×k This represents a 1×k convolution operation, where k represents the adaptive selection of the convolution kernel stride. C represents F in The dimension, γ, b are adjustment parameters, | | odd This indicates finding the nearest odd number, sigmoid represents the sigmoid activation function, and ⊙ represents element-wise multiplication. This indicates the composition of functions;

[0022] The spatial attention module is connected in series after the output of the channel attention module and is used to perform the following calculations:

[0023]

[0024] Among them, X CSA Represents the output of the spatial attention module, Conv 3×3 This represents a 3×3 convolution operation, Concat represents feature concatenation, and AvgPool and MaxPool represent average pooling and max pooling operations, respectively.

[0025] Furthermore, the local pattern topology invariance extraction path includes:

[0026] The key feature point detection unit is used to find the points with the highest non-maximum suppression after the convolutional feature vectors are aggregated as key local pattern points.

[0027] The local graph construction unit is used to construct a local graph with each key local pattern point as a vertex and the Euclidean distance between vertices as edges.

[0028] The graph neural network unit is used to take all key local pattern points and local graphs as input, and learn the node feature matrix through a K-layer graph attention network (GAT).

[0029] The graph readout unit is used to aggregate the node feature matrices output by the graph neural network unit into graph-level topological feature vectors.

[0030] Furthermore, the key feature point detection unit is specifically used to perform the following operations:

[0031]

[0032] In the formula, H (0) =(h1,h2,...,h N ) represents the key local pattern point vector, h1, h2, ..., h N X represents the 1st, 2nd, ..., Nth key local pattern points, where N is the number of key local pattern points. cnn denoted as convolutional feature vector, Agg represents channel aggregation, NMS represents non-maximum suppression, and TopN represents finding the N points with the highest values ​​as key local pattern points.

[0033] Furthermore, the calculation formula performed by the feature fusion module is as follows:

[0034] X out =W proj *Concat(X cnn ,X topo )+b proj

[0035] Among them, X out This represents an augmented embedding vector, Concat represents concatenation, and W... proj Let b represent the projection weight matrix. proj The bias matrices, X, are all obtained through training. cnn X represents the convolutional feature vector. topo This represents the topological feature vector.

[0036] Furthermore, in step (3), when training the feature extraction deep network model, an improved metric learning loss function is used, and the specific calculation formula is as follows:

[0037] L Total =L PC +ξ·L SR +β·L TC

[0038] Where: L Total For the total loss, L PC The comparison loss for class proxies is calculated using the following formula:

[0039]

[0040] In the formula, N batch N represents the batch training sample size. c Let X represent the number of categories, λ represent the radius of the projected hypersphere, δ represent the distance between the similarity between the augmented embedding vector and the positive class agent and the similarity between the augmented embedding vector and the negative class agent, and X represent the distance between the similarity between the augmented embedding vector and the positive class agent. out Represents the augmented embedding vector, y i y j Let P(y) represent the i-th and j-th categories respectively. i ), P(y j ) represent y respectively i y j Class proxy with multi-center representation within the class, and M represents y i The number of multicenters within a class, τ represents the temperature scaling factor, W FC (y i ,m),W FC (y i (n) represent y i The m-th and n-th intra-class centers in the class;

[0041] L SR The total sparsification regularization constraint term for intra-class multicenters is calculated using the following formula:

[0042]

[0043] In the formula, Indicates y i The intra-class multicenter sparse regularization of a class is calculated using the following formula: W FC (y i ,1) W FC (y i ,s), W FC (y i ,t) represent y i The 1st, s, and t-th intra-class centers in the class, where ξ represents the sparsity regularization coefficient;

[0044] L TC The loss term is the intra-class consistency term for topological features, and its calculation formula is as follows:

[0045]

[0046] in, For training samples belonging to y in a batch i The set of training sample indices for class X topo (j) is the topological feature vector of training sample j, μ topo (y i ) is y i The mean of the topological feature vectors of the class within the training sample batch, where β is the weight coefficient of the topological consistency loss term. This represents the square of the second norm.

[0047] Furthermore, the formula for calculating the similarity between the enhanced embedding vector described in step (5) and the enhanced embedding vector stored in the feature database is as follows:

[0048]

[0049] In the formula, X′ out It is the enhanced embedding vector of the image of the animal to be identified. Represents y in the feature database i The augmented embedding vector of the k-th training sample of class, with size X′ out The same, || || represents the norm.

[0050] A device for identifying individual farmed animals that integrates local topological invariance and metric learning includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0051] A system for identifying farmed animals that integrates local topological invariance and metric learning includes:

[0052] The data acquisition module is used to acquire images of farmed animals;

[0053] The edge reasoning module is used to execute the above methods;

[0054] A cloud storage module is used to store and analyze the recognition results of the edge inference module;

[0055] The model optimization module is used to optimize the model in the edge inference module based on the data from the cloud storage module;

[0056] The front-end display module is used to display the recognition results of the edge inference module and the analysis results of the cloud storage module.

[0057] Compared with the prior art, the beneficial effects of this invention are:

[0058] 1. This invention not only employs an improved convolutional neural network with fused attention mechanism (CSA) to extract fine visual appearance features of individual cows, but also innovatively introduces a Local Pattern Topology Invariance (LPTI) extraction path in parallel. This path is specifically designed to capture structured topological features robust to non-rigid deformations caused by changes in posture and viewpoint. By effectively fusing these two complementary features, a more comprehensive enhanced individual representation is generated, significantly improving the model's robustness and accuracy in feature extraction under conditions of drastic changes in posture and viewpoint.

[0059] 2. This invention employs an improved metric learning loss function to train the model. A multi-center class proxy contrastive loss, combined with sparsity regularization constraints, is applied to the final fused features, effectively enhancing the intra-class compactness and inter-class separability of the overall features. Simultaneously, a specially introduced intra-class consistency loss term for topological features directly optimizes the topological features output by the LPTI module, forcing them to remain stable and consistent among samples of the same class. The combination of these loss functions makes the final learned enhanced embedding vectors more discriminative, further improving the accuracy and generalization ability of individual cow identification. Attached Figure Description

[0060] Figure 1This is a flowchart illustrating the method for identifying individual farmed animals that integrates local topological invariance and metric learning provided by the present invention.

[0061] Figure 2 This is a structural diagram of the feature extraction deep network model provided by the present invention;

[0062] Figure 3 This is a structural diagram of the basic feature extraction path provided by the present invention;

[0063] Figure 4 This is a structural diagram of the residual module in the basic feature extraction path provided by this invention;

[0064] Figure 5 This is a schematic diagram of the channel-attention module provided by the present invention;

[0065] Figure 6 This is a schematic diagram of the LPTI extraction path provided by the present invention;

[0066] Figure 7 This is a schematic diagram of the feature fusion module provided by the present invention;

[0067] Figure 8 This is a schematic diagram of the structure of an embodiment of the farmed animal individual identification system that integrates local topological invariance and metric learning provided by the present invention;

[0068] Figure 9 This is a schematic diagram of another embodiment of the livestock individual identification system that integrates local topological invariance and metric learning provided by the present invention. Detailed Implementation

[0069] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0070] Example 1

[0071] This invention provides a method for identifying individual farmed animals by integrating local topological invariance and metric learning, such as... Figure 1 As shown, it includes the following steps:

[0072] S1. Obtain animal images from different perspectives and postures, perform preprocessing, and construct a sample dataset.

[0073] Data acquisition can be divided into the following four operations: image acquisition, image enhancement, normalization processing, and dataset creation.

[0074] The following is a detailed explanation of these four operations:

[0075] Image Acquisition: Using professional image acquisition equipment, dairy cows are photographed from multiple angles and in various poses within the dairy farming environment. This ensures that the acquired images cover all common appearances of dairy cows in actual farming conditions. During the shooting process, attention must be paid to the lighting conditions of the shooting environment, avoiding shadows, direct sunlight, and other factors as much as possible. Simultaneously, the stability of the shooting equipment must be maintained to prevent image blurring.

[0076] Image enhancement: Perform a series of enhancement operations on the acquired images, such as randomly rotating the image within ±30 degrees, randomly scaling the image within a scale range of 0.8 to 1.2 times, and horizontally flipping the image with a 50% probability, to increase the diversity of data and the generalization ability of the model.

[0077] Normalization: First, divide each pixel value of the image by the maximum pixel value of the image to normalize it to the range of 0-1. Then, normalize the image by mean and standard deviation to make the image data have a consistent distribution in the feature space, thereby improving the training efficiency and stability of the model.

[0078] Dataset creation: The enhanced and normalized images are divided into training, validation and test sets according to a certain ratio, ensuring that each subset contains images from different cows and from various perspectives and poses, and each image is labeled with its corresponding cow individual identifier.

[0079] S2. Construct a deep network model for feature extraction. For example... Figure 2 As shown, the feature extraction deep network model includes a basic feature extraction path, a Local Pattern Topology Invariance (LPTI) extraction path, and a feature fusion module.

[0080] The following section details the specific structure, parameters, and mathematical expressions of the above modules.

[0081] (1) The basic feature extraction path uses a ResNet 50 network with a fusion attention mechanism to extract convolutional feature vectors from animal images. The ResNet 50 network with a fusion attention mechanism is an improved version of the ResNet 50 network, such as... Figure 3 As shown, it contains an input convolutional layer, 16 bottleneck blocks (residual modules), and a global average pooling layer.

[0082] The input convolutional layer contains the following three modules: 7×7 convolution (Conv... 7×7 Batch normalization (BN) and activation function (ReLU) are used to represent the data processed by the input convolutional layer in a composite form as follows:

[0083]

[0084] In the above expression, I represents the animal image, which is a three-dimensional matrix of size D×D×3, corresponding to the input image data, with a side length of D; X is a three-dimensional matrix of size D / 2×D / 2×C1, representing the output feature map of the input convolutional layer, where C1 is the number of output channels of the convolutional layer. This indicates that the function performs a conformance operation, specifically executing the last operation first, and then proceeding from the right to the left. For example, in the above expression, Conv is executed first. 7×7 (I), then Conv 7×7 (I) Perform batch normalization, and then apply the ReLU activation function to the result of batch normalization.

[0085] The following is a detailed explanation of the data processing flow of the functions corresponding to the three modules in the above formula.

[0086] First, we introduce the k×k convolution module function. The 7×7 convolution is a special case where k=7. The formula for convolution with a kernel size of k×k is:

[0087]

[0088] Among them, X in It is the three-dimensional input matrix of the convolution module function, with a size of H×W×C, where H, W, and C represent the length, width, and number of channels of the input image, respectively, and X... in (h+m-1,w+n-1,c) is X in The value at channel c and position (h+m-1, w+n-1); X′ is the three-dimensional output matrix of the convolution module function, with a size of H′×W′×C′; the convolution kernel W k×k It is a four-dimensional matrix of size k×k×C×C′; bias b k×k It is a vector of size C′. In the 7×7 convolution module of this invention, the input is I; the output is X. (1) The size is D / 2×D / 2×C1; the parameters to be trained include the convolution kernel W. 7×7 and bias b 7×7 Their sizes are 7×7×3×C1 and C1, respectively.

[0089] Secondly, batch normalization is performed after each convolutional layer to accelerate the training and convergence speed of the model. For an input 3D matrix X′ of size H′×W′×C′, the mean and variance are calculated for each channel. Then, the feature values ​​of each channel are normalized so that the mean is close to 0 and the variance is close to 1. The specific calculation process is as follows:

[0090]

[0091]

[0092]

[0093]

[0094] Where X(h,w,c) is the value of the input feature map X′ at channel c and position (h,w), and μ(c) and σ 2 (c) represents the mean and variance of X′ on channel c, respectively. ε is a very small constant used to prevent the denominator from being zero. γ and β are the scaling and offset parameters to be trained. X BN This is the output value after batch normalization, and its magnitude is the same as X. In the batch normalization module of this invention, the input is X. (1) The output is The parameters to be trained have a scaling parameter γ (1) and offset parameter β (1) All are C1

[0095] Finally, there is the ReLU activation function module, whose calculation formula is as follows:

[0096]

[0097] Among them, X BN (h,w) represents the feature map X after batch normalization. BN At position (h, w), X ReLU It is the output value processed by the ReLU activation function, and its magnitude is related to X. BN The same. In the ReLU activation function module of this invention, the input is... The output is No training parameters are needed.

[0098] The following section provides a detailed introduction to the improved bottleneck block of the ResNet 50 network. The bottleneck block structure is as follows: Figure 4 As shown, it contains the following modules: 1×1 convolution (Conv 1×1 Batch normalization (BN), activation function (PReLU), 3×3 convolution (Conv) 3×3 ), Channel-Spatial Attention (CSA) module. In the above modules, "Conv 1×1 The three modules "BN" and "PReLU" appear twice, with the same structure but independent parameters. The processing of the bottleneck block can be represented using a composite function as follows:

[0099]

[0100] In the above expression, + indicates a residual connection, X corresponds to the input feature map of the bottleneck block, X″ is a three-dimensional matrix of size D / 2×D / 2×C2, which is the output feature map of the bottleneck block, and C2 is the number of output channels of the second 1×1 convolutional module in the bottleneck block.

[0101] 1×1 convolution and 3×3 convolution are special cases of k×k convolution with k=1 and k=3, respectively. Both convolution and batch normalization have been detailed above and will not be repeated here. The following section provides a detailed explanation of the data processing flow for the other modules involved in the above formulas.

[0102] First, let's look at the PReLU activation function module. The calculation formula for the PReLU activation function module is:

[0103]

[0104] Compared to the ReLU activation function, the PReLU activation function alleviates the gradient vanishing problem of ReLU in the negative input region by introducing a learnable parameter α in the negative domain.

[0105] CSA module such as Figure 5 As shown, it includes a channel attention module (CA) and a spatial attention module (SA). The CA module includes global average pooling (GAP) and 1×k convolution (Conv... 1×k The activation function (Sigmoid) can be used to represent data processed by the CA module in a composite form as follows:

[0106]

[0107] In the above expression, X CA The output of the channel attention module is a three-dimensional matrix of size D / 2 × D / 2 × C3, where ⊙ represents element-wise multiplication, and F... in The input to the channel attention module is a three-dimensional matrix of size D / 2×D / 2×C3, which also corresponds to the output of the second PReLU activation function in the bottleneck block. C3 is the number of output channels of the 3×3 convolutional module in the bottleneck block.

[0108] The SA module includes average pooling (AvgPool), max pooling (MaxPool), feature concatenation (Concat), and 3×3 convolution (Conv...). 3×3 The activation function (Sigmoid) can be used to represent the data processed by the SA module in a composite form as follows:

[0109]

[0110] In the above expression, X CSA It is a three-dimensional matrix of size D / 2×D / 2×C3, which is the output of the CSA module and the SA module.

[0111] The following is a detailed explanation of the data processing flow of the corresponding functions in the above formula.

[0112] The Global Average Pooling (GAP) module in the CA module averages the pixel values ​​across the entire feature map space. The calculation process is as follows:

[0113]

[0114] in, yes In channel c, at position (h, w), X GAP It is a three-dimensional matrix of size 1×1×C3, corresponding to the output after global average pooling.

[0115] The max pooling module in the SA module takes the maximum value on each channel of the feature map. The calculation process is as follows:

[0116]

[0117] Among them, X CA (c,h,w) is X CA In channel c, at position (h, w), X Maxpool It is a three-dimensional matrix of size D / 2×D / 2×, corresponding to the output of the max pooling module.

[0118] The average pooling module in the SA module takes the average value across the channels of each feature point in the feature map. The calculation process is as follows:

[0119]

[0120] Among them, X Avgpool It is a three-dimensional matrix of size D / 2×D / 2×1, corresponding to the output of the average pooling module. The feature concatenation module in the SA module concatenates the two input feature maps along the channel dimension, as shown below:

[0121]

[0122] Among them, X Concat X is a three-dimensional matrix of size D / 2×D / 2×2, corresponding to the output of the feature stitching module. Concat (h,w,c) is X Concat The value at position (h, w) in channel c.

[0123] The global average pooling layer mainly consists of one module: global average pooling. The data processing calculation formula is as follows:

[0124]

[0125] Where X″′ represents the input of the global average pooling layer, corresponding to the output feature map after the 16th bottleneck block, with a size of H″′×W″′×C4, where C4 is the number of output channels of the 16th bottleneck block; X qavg This represents the output of the global average pooling layer, i.e., the convolutional feature vector, with a size of 1×1×C4; there are no parameters to be trained.

[0126] (2) The LPTI extraction path is used to identify the key local pattern points of the convolutional feature vector, construct a local graph between the key local pattern points, and use a graph neural network to process the local graph to extract the topological feature vector.

[0127] like Figure 6 As shown, the LPTI extraction path includes:

[0128] The key feature point detection unit is used to find the points with the highest non-maximum suppression after the convolutional feature vector aggregation is used as key local pattern points; specifically, it performs the following operations:

[0129]

[0130] In the formula, H (0) =(h1,h2,...,h N ) represents the key local pattern point vector, h1, h2, ..., h N X represents the 1st, 2nd, ..., Nth key local pattern points, where N is the number of key local pattern points. cnn represents the convolutional feature vector, Agg represents channel aggregation, NMS represents non-maximum suppression, and TopN represents finding the N points with the highest values ​​as key local pattern points.

[0131] The local graph construction unit is used to calculate the Euclidean distance between all key local pattern points as the vertex set V, and to determine the adjacency information by using the Euclidean distance between vertices as the edge set E, and to construct the local graph G = (V, E).

[0132] The graph neural network unit is used to take all key local pattern points and local graphs as input, and learn the node feature matrix through a K-layer graph attention network (GAT). Specifically, for the k-th (k from 0 to K-1) layer of the GAT, the attention value e is first calculated. ij The calculation formula is as follows:

[0133]

[0134] in, H represents the input feature matrix of the k-th layer GAT. (k) The row vector, W (k) Let || denote the linear transformation matrix that can be learned by this layer, || denotes vector concatenation, and a (k) This represents the learnable attention parameter vector of this layer, and LeakyReLU represents the activation function;

[0135] Then, the normalized attention coefficient is calculated using the Softmax activation function. The calculation formula is as follows:

[0136]

[0137] Finally, attention coefficients are used to weight and aggregate the features of neighboring nodes to obtain the features of the next layer nodes. The calculation formula is as follows:

[0138]

[0139] Where σ represents the activation function. This continues until the calculation yields... The updated node feature matrix

[0140] The graph readout unit is used to aggregate the node feature matrices output by the graph neural network units into a graph-level topological feature vector X. topo .

[0141] (3) The feature fusion module is used to fuse the convolutional feature vector with the topological feature vector to generate an enhanced embedding vector. For example... Figure 7 As shown, the feature fusion module includes feature stitching and feature projection, and the calculation formula is as follows:

[0142] X out =W proj *Concat(X cnn ,X topo )+b proj

[0143] Among them, X out This represents an augmented embedding vector, Concat represents concatenation, and W... proj Let b represent the projection weight matrix. proj The bias matrices, X, are all obtained through training. cnn X represents the convolutional feature vector. topo This represents the topological feature vector.

[0144] X outThe output of the model is the sum of its parts, but during training, a loss function needs to be calculated. Therefore, a multi-center embedding layer is required after feature projection during training. The multi-center embedding layer mainly consists of three modules: fully connected (FC), an activation function (Softmax), and a weighted summation (WS). The data processed by the multi-center embedding layer can be represented by a composite function as follows:

[0145]

[0146] In the above expression, P is a variable of size 1×1×C4×N. c The four-dimensional matrix, corresponding to the output class proxy of the multi-center embedding layer, N c The number of categories to be identified.

[0147] The role of the fully connected module in the multi-center embedding layer is to calculate the similarity between the input features and the intra-class multicenters. The data processing calculation formula is as follows:

[0148] X FC (y i ,m)=X out ·W FC (y i ,m)

[0149] Among them, W FC It is a 1×1×C4×M·N c A four-dimensional matrix, where M is the number of in-class centers for each class, and W... FC (y i (m) represents the y-th i The m-th intra-class center in the class has a size of 1×1×C4; X FC It is a size of 1×1×M·N c The three-dimensional matrix, X FC (y i (m) represents X out With the yth i Output the similarity of the m-th intra-class center in the class.

[0150] The Softmax activation function normalizes the output of the fully connected layer, transforming it into a probability distribution. The data processing calculation formula is as follows:

[0151]

[0152] Among them, M p It is a size of 1×1×M·N c The matrix, M p (y i (m) represents the y-th i The m-th intra-class center in the class is at the y-th position iThe weight of the class, where τ represents the temperature scaling factor.

[0153] The weighted summation module calculates a weighted sum based on the probabilities output by Softmax, generating a class proxy for each class. The data processing calculation formula is as follows:

[0154]

[0155] S3. Train the feature extraction deep network model using the sample dataset.

[0156] During training, the network loss is calculated based on the improved metric learning loss function, and the network parameters are optimized and adjusted through backpropagation algorithm, and the process is continuously iterated.

[0157] Before network training begins, the network needs to be initialized. Based on the above explanation of the feature extraction deep network training model and parameters, the network layers and parameters that need to be initialized are as follows:

[0158] Input convolutional layer: 7×7 convolution kernel W 7×7 and bias b 7×7 ; Scaling parameter γ for batch normalization (1) and offset parameter β (1) ;

[0159] bottleneck block: 1×1 convolution kernel W 1×1 、W′ 1×1 and bias b 1×1 b′ 1×1 ; Scaling parameter γ for batch normalization (2) γ (3) and offset parameter β (2) β (3) The learnable parameters α and α′ of the PReLU activation function; the convolution kernel W of the 3×3 convolution. 3×3 and bias b 3×3 ;CSA's 1×k convolution kernel W 1×k and bias b 1×k Convolution kernel W′ 3×3 and bias b′ 3×3 ;

[0160] K-layer graphical attention network: linear transformation matrix W (k) and attention parameter vector a (k) ;

[0161] Feature fusion module: Projection weight matrix W proj and bias matrix b proj The weight matrix W of the multi-center embedding layer FC .

[0162] The initialization method is as follows:

[0163] Different layer modules have different parameter initialization methods:

[0164] Convolutional Modules: For the weights of each convolutional module, if the following activation function is an asymmetric activation function such as ReLU, the He initialization method is used, including the convolutional kernel W of the input convolutional layer. 7×7 A convolutional kernel W with 16 bottleneck blocks 1×1 、W′ 1×1 W 3×3 The linear transformation matrix W of the K-layer graph attention network (k) If the activation function following it is a symmetric activation function such as Sigmoid, the Xavier initialization method is used, including the W of CSA in the 16 bottleneck blocks. 1×k and W′ 3×3 .

[0165] He initializes the convolution weight calculation process as follows: Assume the number of input channels of the convolution module is C. in The number of output channels is C out The kernel size is k×k. The convolution weights W k×k The initial values ​​of are 0 and the standard deviation is . The normal distribution, i.e., W k×k With a mean of 0 and a standard deviation of Random sampling initialization in a normal distribution.

[0166] The Xavier initialization convolution weight calculation process is as follows: Convolution weight W k×k The initial values ​​follow a mean of 0 and a standard deviation of 2 / (C). in +C out The normal distribution of W k×k With a mean of 0 and a standard deviation of 2 / (C) in +C out Random sampling initialization is performed on a normal distribution.

[0167] The biases of each convolutional module are typically initialized to 0, including the bias b of the input convolutional layer. 7×7 The bias b of 16 bottleneck blocks 1×1 b′ 1×1 b 3×3 b 1×k b′ 3×3 The bias b of the feature fusion module proj The attention parameter vector a of the K-layer graph attention network (k) .

[0168] Fully connected module: Weight W for the fully connected module FCand W proj Using the Xavier initialization method, W FC and W proj From uniform distribution Random sampling initialization, where N in N represents the number of input neurons. out This represents the number of output neurons.

[0169] Batch normalization module: Typically initializes the scaling factor to 1, including the γ of the input convolutional layer. (1) and γ of 16 bottleneck blocks (2) γ (3) The offset factor is initialized to 0, including the β of the input convolutional layer. (1) and β of 16 bottleneck blocks (2) β (3) .

[0170] PReLU activation function module: The slope parameters α and α′ in the PReLU activation function module are usually initialized to 0.25.

[0171] After initialization, network training is performed, which can be divided into the following four operations: feature acquisition, loss calculation, parameter optimization, and iterative training. These four operations will be described in detail below:

[0172] Feature extraction: An improved ResNet 50 network with parameter initialization is used as the basic feature extraction path, and a K-layer graph attention network is used as part of the LPTI extraction path. The preprocessed cow image I is input and passed through the basic feature extraction path, the LPTI extraction path, and the feature fusion module, finally obtaining a feature vector X of size 1×1×C4. out .

[0173] Loss calculation: The obtained feature vector X out Through the multi-center embedding layer after parameter initialization, the class proxy P(y) of each class is obtained. i Then, calculate the class proxy comparison loss, using the following formula:

[0174]

[0175] Class proxy contrastive loss improves SoftMax loss by introducing multiple centers for each class, which helps reduce intra-class variance and does not require triple sampling as in traditional triple loss.

[0176] To enable adaptive adjustment of the number of class centers, a sparse regularization constraint is applied to the intra-class multicenters, and its calculation formula is as follows:

[0177]

[0178] In the formula, Indicates y i The intra-class multicenter sparse regularization of a class is calculated using the following formula: W FC (y i ,1) W FC (y i ,s), W FC (y i ,t) represent y i The 1st, s, and t-th intra-class centers in the class, where ξ represents the sparsity regularization coefficient;

[0179] Add a topological feature intra-class consistency loss term L TC The calculation formula is as follows:

[0180]

[0181] in, For training samples belonging to y in a batch i The set of training sample indices for class X topo (j) is the topological feature vector of training sample j, μ topo (y i ) is y i The mean of the topological feature vectors of the class within the training sample batch, where β is the weight coefficient of the topological consistency loss term. This represents the square of the second norm.

[0182] The improved metric learning loss function is calculated using the following formula:

[0183] L Total =L PC +ξ·L SR +β·L TC

[0184] Parameter optimization: Parameter optimization uses the backpropagation algorithm to adjust network parameters based on the loss value. The backpropagation algorithm is based on the chain rule, calculating the gradient of the loss function with respect to each parameter in the network. Its core idea is to start with the loss function and propagate the gradient backward step by step, calculating the contribution of each parameter to the loss function, thereby determining how to adjust the parameters to reduce the loss. Parameter optimization can be divided into two operations: gradient calculation and parameter update. These two operations are described in detail below:

[0185] Calculate the gradient: First, calculate the gradient of the loss function with respect to the output feature vector. For the joint loss function consisting of the class proxy contrastive loss, its associated sparse regularization constraints, and the topological feature intra-class consistency loss, derive the gradient with respect to the output feature X based on its expression and differentiation rules. outgradient And regarding intraclass multicenter W FC gradient

[0186] According to the chain rule, the gradient is backpropagated to each module of the network step by step, and finally the gradient of the loss function with respect to all the network parameters θ to be trained is calculated.

[0187] Parameter Update: The network parameters θ are updated using the optimizer based on the calculated gradients. This step uses the Adam optimizer, and its parameter update process is as follows:

[0188] The initialization parameters are defined as follows: the first-order moment estimation vector *m0* and the second-order moment estimation vector *v0* are all-zero vectors with the same dimension as the parameter to be optimized, *θ*. In each iteration:

[0189] The first-order moment estimate is calculated using the following formula:

[0190] m t =β1m t-1 +(1-β1)g t

[0191] Among them, g t θ represents the gradient of the loss function with respect to the parameter θ, t represents the current training iteration step, and β1 is a hyperparameter that usually takes a value in the interval (0,1), with a common value of 0.9.

[0192] The second-order moment estimate is calculated using the following formula:

[0193]

[0194] β2 is another hyperparameter, which usually takes a value in the interval (0,1), with a common value of 0.999.

[0195] The bias correction for the first-order moment estimate is calculated using the following formula:

[0196]

[0197] As the number of iterations t increases, It will gradually approach 0. As training progresses to a certain point, the impact of this bias correction will gradually decrease, making... It can more accurately reflect the actual average gradient.

[0198] The bias correction for the second-order moment estimate is calculated using the following formula:

[0199]

[0200] Similarly, as the number of iterations t increases, this bias correction allows... It more accurately reflects the actual gradient variance.

[0201] Finally, the parameters are updated based on the bias-corrected first-order moment estimates and second-order moment estimates, calculated as follows:

[0202]

[0203] Here, η is the learning rate, a hyperparameter that controls the step size of each parameter update, and ε is a very small constant added to prevent numerical instability when the denominator is zero.

[0204] Based on the above calculation formula, and according to the current parameter value θ t The learning rate η and the corrected average gradient and gradient variance This is used to update the parameters, adjusting them in the direction that the loss function decreases.

[0205] Iterative Training: Repeat the above process of feature acquisition, loss calculation, and parameter optimization, continuously iterating until the loss converges to a set threshold. In each iteration, a batch of data is randomly selected from the training dataset for training, enabling the model to learn more general feature representations and avoiding overfitting. As the number of iterations increases, the network parameters are gradually adjusted according to the selected optimizer, causing the value of the loss function to continuously decrease. When the value of the loss function changes less than the set threshold in multiple consecutive iterations, the model is considered to have converged, and training ends.

[0206] S4. Using the trained feature extraction deep network model, extract the enhanced embedding vectors of several animal images with known identities, and store the enhanced embedding vectors and their corresponding identities in the feature database.

[0207] In the database, a record is created for each individual cow, and its corresponding feature vector and related identification information (such as cow number, collection time, etc.) are stored in the corresponding fields for quick querying and retrieval later.

[0208] S5. Input the image of the animal to be identified into the trained feature extraction deep network model to extract its enhanced embedding vector, calculate the similarity between the enhanced embedding vector and the enhanced embedding vector stored in the feature database, and select the identity corresponding to the enhanced embedding vector with the highest similarity in the feature database as the identity of the individual animal to be identified.

[0209] The formula for calculating similarity is:

[0210]

[0211] In the formula, X′out It is the enhanced embedding vector of the image of the animal to be identified. Represents y in the feature database i The augmented embedding vector of the k-th training sample of class, with size X′ out The similarity is defined by ||, where || represents the norm. The similarity value ranges from -1 to 1; the closer the value is to 1, the more similar the two vectors are.

[0212] Example 2

[0213] This invention provides a device for identifying individual farmed animals that integrates local topological invariance and metric learning. This invention provides services for implementing the method described in Embodiment 1 above. The device may include: a memory storing a computer-executable program; a processor coupled to the memory; and the processor calling the computer-executable program stored in the memory to execute the steps of the method described in Embodiment 1.

[0214] Example 3

[0215] This invention provides a system for identifying individual farmed animals that integrates local topological invariance and metric learning. This system is implemented on a standalone terminal device, belonging to an edge computing scenario. Figure 8 As shown, the system includes a data acquisition module, a model inference module, and a front-end display module. These modules are integrated into a single device, such as a smart camera with certain computing capabilities. These end devices have image acquisition functions, enabling them to collect image data of farmed animals such as dairy cows in real time and run the method described in Embodiment 1 locally. The end device inputs the acquired image data into a locally trained model for real-time inference, quickly obtaining the identification results of individual farmed animals such as dairy cows.

[0216] Data Acquisition Module: High-resolution, low-light intelligent cameras with autofocus and automatic aperture adjustment are carefully deployed in key areas of the ranch, such as feeding areas, rest areas, and activity areas. These cameras use advanced optical sensors to clearly capture images of cows under different lighting conditions, ensuring that the image clarity and detail richness meet the needs of subsequent model inference. The camera installation angles are precisely calculated and field-tested to ensure comprehensive coverage of the cows' possible positions and postures, reducing blind spots and avoiding image distortion or occlusion caused by improper installation. The image acquisition strategy involves setting reasonable acquisition time intervals and trigger conditions based on the cows' activity patterns and the ranch's daily operational procedures. The acquired image data is transmitted to the model inference module of the end device.

[0217] Model Inference Module: A pre-trained, mature cow feature extraction model is stored on the edge device. During model loading, the hardware acceleration resources of the edge device (such as a dedicated AI accelerator chip) are utilized for acceleration. Through rapid decompression of model parameters and optimized memory allocation, the model can be loaded and enter inference mode in a short time. Received image data is input into the loaded model, initiating the real-time inference process. The model rapidly extracts and analyzes cow features from the image, accurately identifies the individual cow, and outputs the corresponding recognition results. The inference process fully utilizes the hardware acceleration capabilities of the edge device to ensure the inference task is completed quickly, meeting the real-time requirements of the field.

[0218] Front-end display module: The front-end display module is connected to the model inference module, enabling it to acquire the latest recognition results and related data in real time and dynamically update the display on the interface, adapting the interface to the screen size and resolution of the device. Once the model completes inference for an image, the recognition result is immediately displayed on the corresponding interface. Furthermore, based on the individual information and historical data of the cows, additional functions such as monitoring their health status and abnormal behavior can be developed.

[0219] Example 4

[0220] This invention provides a system for identifying individual farmed animals that integrates local topological invariance and metric learning. This system is implemented in an edge-cloud scenario. Figure 9 As shown, it includes a data acquisition module, an edge inference module, a front-end display module, a cloud storage module, and a model optimization module.

[0221] The following will provide a detailed introduction to these five modules.

[0222] Data acquisition module: Specifically, it consists of acquisition devices such as cameras. The acquisition devices deployed in edge-cloud and edge scenarios are the same. However, in edge-cloud scenarios, the acquired image data is transmitted to the edge inference module in real time through high-speed wireless networks (such as 5G or high-performance Wi-Fi).

[0223] Edge Inference Module: An edge inference module is deployed on edge computing nodes in the ranch, such as small servers or gateway devices near the breeding area, to execute the individual identification method described in Example 1. Edge computing nodes have more powerful computing capabilities and storage resources than end devices, but offer lower network latency compared to cloud servers. The edge inference module is responsible for receiving animal image data of dairy cows and other farmed animals collected from multiple end devices and centrally processing and analyzing this data. At the edge node, the preprocessed image data is input into the model for rapid inference calculations, outputting preliminary individual identification results for dairy cows and other farmed animals, which are then pushed to the front-end display module. The edge inference module, connected to the cloud storage module, can further upload the collected data and inference results to the cloud storage module.

[0224] Cloud Storage Module: The cloud server constructs a large-scale distributed storage system based on an advanced distributed file system and distributed database, providing reliable and efficient storage services for data uploaded from edge nodes across various ranches. Upon receiving data uploaded from edge computing devices, the cloud storage module categorizes and stores cow image data, recognition results, feature vectors, and related breeding information according to multiple dimensions such as ranch, time, and individual cow, and establishes a comprehensive indexing mechanism for rapid querying and retrieval. Furthermore, the cloud storage module is responsible for regularly backing up and archiving the stored data, backing up important data to off-site data centers or storage media to ensure data security and reliability.

[0225] Model Optimization Module: Compared to edge-cloud models which are fixed to hardware resources (such as chips), edge-cloud models can be dynamically optimized and updated. The cloud server retrieves a large amount of cow image data and related information from the cloud storage module, performs further preprocessing and annotation on this data to meet the requirements of model training. Then, based on the previous network training steps, the parameters are further optimized and adjusted. The trained model needs to be comprehensively evaluated. Based on the evaluation results, if the model's performance meets or exceeds the expected target, the model is marked as a usable version and released and pushed through the version management system. Model update notifications are sent to each edge node to guide the edge nodes in updating and deploying the model, ensuring that the edge inference module uses the optimal model version. If the model performance does not meet the requirements, the reasons are analyzed in depth, targeted improvements and optimizations are made, and then retraining and evaluation are performed until a model version that meets the performance requirements is obtained.

[0226] Front-end display module: The front-end display module in the edge-cloud scenario has richer functions than that in the terminal scenario. In addition to obtaining the latest recognition results in real time, it can also interact with the cloud storage module and model optimization module, receive data analysis reports, model update notifications and other information from the cloud, and present them in a visual way, providing strong support for making scientific and reasonable breeding decisions, thereby realizing a complete closed loop from data collection to decision execution.

[0227] It should be understood that the embodiments and descriptions above are only the principles, main features and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope of the invention, and all such changes and modifications fall within the protection scope of the present invention.

Claims

1. A method for identifying individual farmed animals that integrates local topological invariance and metric learning, characterized in that, Includes the following steps: (1) Obtain animal images from different perspectives and postures, and preprocess them to construct a sample dataset; (2) Construct a feature extraction deep network model, wherein the feature extraction deep network model includes: The basic feature extraction path is used to extract convolutional feature vectors from animal images using a ResNet 50 network with a fusion attention mechanism. A local pattern topology invariant extraction path is used to identify key local pattern points of the convolutional feature vector, construct a local graph between key local pattern points, and process the local graph using a graph neural network to extract topological feature vectors. The feature fusion module is used to fuse the convolutional feature vector with the topological feature vector to generate an enhanced embedding vector; (3) The feature extraction deep network model is trained using a sample dataset; (4) Using the trained feature extraction deep network model, extract the enhanced embedding vectors of several animal images with known identities, and store the enhanced embedding vectors and corresponding identities in the feature database; (5) Input the image of the animal to be identified into the trained feature extraction deep network model to extract its enhanced embedding vector, calculate the similarity between the enhanced embedding vector and the enhanced embedding vector stored in the feature database, and select the identity corresponding to the enhanced embedding vector with the highest similarity in the feature database as the identity of the individual animal to be identified. The local pattern topology invariance extraction path includes: The key feature point detection unit is used to find the points with the highest non-maximum suppression after the convolutional feature vectors are aggregated as key local pattern points. The local graph construction unit is used to construct a local graph with each key local pattern point as a vertex and the Euclidean distance between vertices as edges. Graph neural network units are used to take all key local pattern points and local graphs as input, and then... The layered graph attention network (GAT) learns the node feature matrix. The graph readout unit is used to aggregate the node feature matrix output by the graph neural network unit into a graph-level topological feature vector; The key feature point detection unit is specifically used to perform the following operations: , In the formula, Represents the key local pattern point vector. This represents the 1st, 2nd, ..., Nth key local pattern points, where N is the number of key local pattern points. Represents the convolutional feature vector. Indicates channel aggregation, Indicates nonmaximum suppression. This indicates searching for the highest value. These points are used as key local pattern points. This indicates the composition of functions.

2. The method for identifying individual identities of farmed animals by integrating local topological invariance and metric learning according to claim 1, characterized in that, The ResNet 50 network with the fusion attention mechanism specifically adds a channel-spatial attention module to the end of each residual module in the ResNet 50 network. The residual module is used to perform the following calculations: , in, X represents the output of the residual module, and X represents the input of the residual module. express Convolution, BN stands for Batch Normalization. express Convolution, PReLU represents the Parametric ReLU activation function, CSA represents the channel-spatial attention module, and + represents the residual connection.

3. The method for identifying individual identities of farmed animals by integrating local topological invariance and metric learning according to claim 2, characterized in that, The channel-space attention module includes a channel attention module and a spatial attention module. The channel attention module is used to perform the following calculations: , in, This represents the output of the channel attention module. This represents the input to the channel attention module. This indicates a global average pooling operation. express Convolution operation, This indicates adaptive selection of the convolution kernel stride. C represents Dimensions , To adjust the parameters, This means finding the nearest odd number. This represents the Sigmoid activation function. Represents element-wise multiplication; The spatial attention module is connected in series after the output of the channel attention module and is used to perform the following calculations: , in, This represents the output of the spatial attention module. express Convolution operation, Indicates feature splicing, and These represent average pooling and max pooling operations, respectively.

4. The method for identifying individual identities of farmed animals by integrating local topological invariance and metric learning according to claim 1, characterized in that, The calculation formula executed by the feature fusion module is as follows: , in, Represents the augmented embedding vector, Indicates splicing, Represents the projection weight matrix. The bias matrices are all obtained through training. Represents the convolutional feature vector. This represents the topological feature vector.

5. The method for identifying individual identities of farmed animals by integrating local topological invariance and metric learning according to claim 1, characterized in that, In step (3), when training the feature extraction deep network model, an improved metric learning loss function is used, and the specific calculation formula is as follows: , in: For the total loss, The comparison loss for class proxies is calculated using the following formula: , In the formula, Indicates the batch training sample size. Indicates the number of categories. Indicates the radius of the projected hypersphere. This represents the distance between the similarity between the augmented embedding vector and the positive proxy and the similarity between the augmented embedding vector and the negative proxy. Represents the augmented embedding vector, , Representing the i-th and j-th categories respectively, , They represent , Class proxy with multicentric representation within the class, and , express Number of intraclass multicenters Indicates the temperature scaling factor. , They represent The m-th and n-th intra-class centers in the class; The total sparsification regularization constraint term for intra-class multicenters is calculated using the following formula: , In the formula, express The intra-class multicenter sparse regularization of a class is calculated using the following formula: , , , They represent The 1st, sth, and tth class centers within the class, Represents the sparse regularization coefficient; The loss term is the intra-class consistency term for topological features, and its calculation formula is as follows: , in, For training samples within a batch The set of training sample indices for the class. For training samples The topological feature vector, for The mean of the topological feature vectors of a class within a batch of training samples. These are the weighting coefficients for the topology consistency loss term. This represents the square of the second norm.

6. The method for identifying individual farmed animals by integrating local topological invariance and metric learning according to claim 1, characterized in that, The formula for calculating the similarity between the enhanced embedding vector and the enhanced embedding vector stored in the feature database in step (5) is as follows: , In the formula, It is the enhanced embedding vector of the image of the animal to be identified. Representation feature database Class 1 The augmented embedding vectors of each training sample, with a size equal to... same, Represents the norm.

7. A device for identifying individual farmed animals by integrating local topological invariance and metric learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the method as described in any one of claims 1-6.

8. A system for identifying individual farmed animals that integrates local topological invariance and metric learning, characterized in that, include: The data acquisition module is used to acquire images of farmed animals; Edge reasoning module, used to execute the method of any one of claims 1-6; A cloud storage module is used to store and analyze the recognition results of the edge inference module; The model optimization module is used to optimize the model in the edge inference module based on the data from the cloud storage module; The front-end display module is used to display the recognition results of the edge inference module and the analysis results of the cloud storage module.

Citation Information

Patent Citations

  • Deep convolutional network target identification method based on dual-channel attention mechanism

    CN115601583A

  • Sheep individual identity recognition method and system based on deep metric learning

    CN116798066A