Long-term identity recognition method and system for multi-cat individuals

By employing multimodal feature fusion and temporal evolution prediction methods, the problem of declining recognition rates in cat identification during long-term growth has been solved, achieving accurate and robust identification of cats and adapting to gradual changes in biometrics.

CN121074993AActive Publication Date: 2025-12-05HANGZHOU DIANZI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511138865.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-05
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing cat identification technologies have seen a decline in recognition rates over the long term and lack a mechanism to proactively adapt to feature drift, resulting in insufficient recognition accuracy and robustness.

Method used

By employing a multimodal feature fusion and temporal evolution prediction method, and combining computer vision, near-infrared imaging, and bioimpedance detection with a self-attention mechanism and recurrent neural network, we can achieve active prediction and updating of features and solve the feature drift problem.

Benefits of technology

It achieves accurate identification of cats' identities over a long period of growth, improving the accuracy and robustness of identification, and can proactively adapt to changes in biological characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074993A_ABST
    Figure CN121074993A_ABST
Patent Text Reader

Abstract

The invention discloses a long-term identity recognition method and system for multi-cat individuals. Face three-dimensional point cloud data, cat eye iris data and palm print data of a cat are collected and converted into the same three-dimensional joint coordinate system, and alignment of a physical space is carried out. Then, three parallel feature extraction networks are used for performing feature extraction on the aligned data, and preliminary fusion features are obtained after splicing and dimension reduction; and carrying out attention enhancement and residual connection on the preliminary fusion feature to obtain a fusion feature vector. And the fusion feature vectors of the individuals in different periods are stored in a historical feature sequence according to a time sequence, and a feature library of the multiple individuals is constructed. And through cosine similarity comparison, finding an individual most similar to the individual to be identified in the feature library as a candidate cat. And predicting the time sequence evolution characteristics of the candidate cats, judging the identity of the to-be-identified individual according to the similarity among the time sequence evolution characteristics, the historical characteristics and the to-be-identified individual, and performing rolling updating on the characteristic library.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and relates to a biological feature recognition technology, in particular to a long-term identity recognition method and system for multiple cat individuals. BACKGROUND

[0002] In pet medical management, pet shelters, and intelligent devices such as outdoor cat houses, the recognition of pet identity is an important task, and the core purpose is to uniquely identify different animal individuals through biological features to achieve personalized services and management.

[0003] The existing solution realizes cat identity recognition based on computer vision and AI classification algorithm, and mainly focuses on single or simple multi-modal data processing. Compared with common non-biological targets, the individual changes of biological targets are particularly obvious, especially for young cats. After more than six months, the face shape will change greatly, resulting in a significant decrease in recognition rate after more than six months for single modal recognition methods. In addition, the cross-modal features are inconsistent in physical space, and the deviation is large, which easily introduces noise in the fusion process, thereby seriously affecting the accuracy and robustness of recognition. Moreover, the feature library of the traditional solution is static, or only passive and simple feature replacement is performed when recognition fails. Such a mechanism cannot learn and predict the internal law of the evolution of biological features over time, and thus cannot actively adapt to the long-term and gradual drift of features.

[0004] In summary, the existing technology lacks a high-robustness solution for long-term identity recognition of animal targets. SUMMARY

[0005] In view of the deficiencies of the prior art, the application provides a long-term identity recognition method and system for multiple cat individuals, which solves the feature drift failure problem caused by growth in existing cat identity recognition through deep fusion of multi-modal features such as computer vision, near-infrared imaging, and biological impedance detection, learning of high-discrimination feature space, and active time-series evolution prediction.

[0006] A long-term identity recognition method for multiple cat individuals, comprising the following steps: Step 1, collecting multi-modal data of the cat, including face three-dimensional point cloud data, cat eye iris data, and palm print data.

[0007] Step 2, taking the cat's nose tip position as the origin of the three-dimensional joint coordinate system, and converting the iris data and the palm print data to the three-dimensional joint coordinate system through matrix transformation, so as to ensure the alignment of the multi-modal data in the three-dimensional space.

[0008] Step 3, using three parallel feature extraction networks to extract features from the data aligned in step 2 respectively, then concatenating the feature vectors output by the three feature extraction networks in dimension, and projecting back to the original input dimension to obtain the preliminary fusion feature F'.

[0009] Step 4, feature enhancement of the preliminary fusion feature F' through self-attention mechanism, then residual connection with the preliminary fusion feature F' to obtain the fusion feature vector .

[0010] Step 5, saving the fusion feature vector of each individual in different periods in the historical feature sequence in chronological order to construct a feature library.

[0011] Finding the closest feature vector to the fusion feature vector of the individual to be identified in the feature library as a candidate cat. Input the historical feature sequence of the candidate cat into the recurrent neural network to generate a predicted feature vector. Calculate the similarity between the fusion feature vector and the nearest historical feature vector and the predicted feature vector of the candidate cat, and obtain the identity confirmation confidence Confidence after weighted summation. If the identity confirmation confidence Confidence is lower than the set threshold, the current cat is judged as an unknown individual, and its feature is added to the feature library. Otherwise, the identity of the current cat is judged as the candidate cat, and the fusion feature vector is added to the historical feature sequence of the candidate cat, and the earliest historical feature is removed to realize intelligent rolling update of the feature library.

[0012] A long-term identity recognition system for multiple cat individuals, comprising a data acquisition module, a spatial alignment module, a feature extraction module, a nonlinear fusion module, and a feature library update module.

[0013] The data acquisition module includes a ToF depth sensor for acquiring three-dimensional point cloud data of cat faces, a multispectral camera working in the near-infrared band for capturing cat eye iris texture, and a 4x4 bioimpedance sensor array for acquiring impedance distribution grid data of cat paw pads when the cat touches.

[0014] The spatial alignment module takes the cat's nose tip position as the coordinate origin, and converts the multi-modal data collected by the data acquisition module into the same three-dimensional joint coordinate system.

[0015] The feature extraction module uses three parallel feature extraction networks to extract features from the aligned data respectively.

[0016] The nonlinear fusion module performs splicing and projection on the features output by the feature extraction network in the dimension to obtain preliminary fusion features, and then applies a self-attention mechanism to enhance the preliminary fusion features and perform residual connection with the preliminary fusion features to output a fusion feature vector.

[0017] The recognition and feature updating module uses a recurrent neural network to perform time series evolution prediction from the historical feature sequence, confirms the identity of the cat individual by the similarity between the predicted features, historical features and fusion features, and updates the historical feature sequence using the fusion features.

[0018] The present application has the following beneficial effects: 1. The time series evolution prediction is introduced, which changes from passive compensation of biological feature changes to active prediction of feature evolution trend, solves the feature drift problem, and accurately identifies individuals in the long-term growth change process.

[0019] 2. The nonlinear fusion network based on the attention mechanism is used to replace the traditional linear weighting method, which can deeply mine and utilize the complex correlation between different biological features, generate fusion features with higher information density and stronger expression ability, and fully exert the advantages of multi-modal collaboration.

[0020] 3. Through strict multi-modal space alignment steps, the cross-modal data deviation caused by different physical positions of sensors is eliminated, providing high-quality and unbiased input for all subsequent processing procedures, ensuring the physical meaning and accuracy of the fusion results. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 It is a multi-modal data acquisition and space alignment schematic diagram; Figure 2 It is a feature extraction and preliminary fusion schematic diagram; Figure 3 It is a nonlinear fusion schematic diagram. DETAILED DESCRIPTION

[0022] The present application will be further explained in conjunction with the accompanying drawings; A long-term identity recognition system for multiple cat individuals includes a data acquisition module, a space alignment module, a feature extraction module, a nonlinear fusion module and a feature library updating module.

[0023] The data acquisition module includes a ToF depth sensor for acquiring three-dimensional point cloud data of a cat face, a multispectral camera working in the near-infrared band for capturing cat eye iris texture, and a 4x4 bioimpedance sensor array for acquiring impedance distribution grid data of the cat's paw pad when the cat touches it.

[0024] The spatial alignment module takes the position of the cat's nose tip as the coordinate origin, and converts the multi-modal data collected by the data collection module into the same three-dimensional joint coordinate system.

[0025] The feature extraction module uses three parallel feature extraction networks to extract features from the aligned data.

[0026] The nonlinear fusion module performs dimension concatenation and projection on the features output by the feature extraction network, obtains preliminary fusion features, and then applies a self-attention mechanism to enhance the preliminary fusion features and perform residual connection with the preliminary fusion features to output a fusion feature vector.

[0027] The recognition and feature update module uses a recurrent neural network to perform time series evolution prediction from the historical feature sequence, confirms the identity of the cat individual by comparing the similarity between the predicted features, historical features and fusion features, and updates the historical feature sequence using the fusion features.

[0028] A long-term identity recognition method for multiple cat individuals, the specific steps are as follows: Step 1, as shown in Figure 1 , collect the cat's face three-dimensional point cloud data, cat eye iris data and palm print data.

[0029] Step 2, spatially align the multi-modal features collected in step 1 to solve the problem of cross-modal feature fragmentation.

[0030] First, a lightweight pose estimation model MediaPipe Animal Pose is used to locate the position of the cat's nose tip, which is taken as the origin of the three-dimensional joint coordinate system. Then, through a pre-calibrated rotation and translation matrix, the coordinates of the iris data and palm print data are converted to the three-dimensional joint coordinate system, obtaining palm print data , iris data and face data in the same three-dimensional joint coordinate system, ensuring that the multi-modal data is aligned in three-dimensional space.

[0031] Step 3, as shown in Figure 2 , three parallel feature extraction networks are used to extract features from the palm print data , iris data and face data aligned in step 2, then the feature vectors output by the three feature extraction networks are concatenated in dimension, and a fully connected layer and a ReLU activation function are used for nonlinear transformation, followed by convolution operation, projecting back to the original input dimension, obtaining a 256-dimensional preliminary fusion feature F': wherein, and are learnable weight matrix and bias vector. represents the spliced features. The fully connected layer can extract abstract features, and the ReLU activation function introduces nonlinearity, enabling the network to learn complex relationships.

[0032] Specifically, a 3D sparse convolutional network is used to extract features from unstructured facial data , obtaining a 256-dimensional three-dimensional geometric feature . A ResNet-18 residual network is used to extract features from single-channel near-infrared iris data , obtaining a 256-dimensional iris feature . A lightweight deep separable CNN network is used to extract features from palmprint data , obtaining a 256-dimensional bioimpedance feature .

[0033] Step 4, as shown in Figure 3 , apply self-attention mechanism to the preliminary fusion feature F', calculate the correlation between different elements inside, adaptively identify the most critical feature dimension in the fusion task, increase its weight, and suppress those relatively secondary or redundant features. Then use residual connection to connect the attention-enhanced feature and the preliminary fusion feature F' with residual connection, to obtain the final fusion feature vector : where α is a learnable scalar parameter initialized to 0, used to control the initial influence of the self-attention mechanism and ensure training stability.

[0034] The features before and after attention enhancement are added through residual connection, which not only increases the weight of more useful features, but also preserves the original features.

[0035] Step 5, collect the cat's facial 3D point cloud data, eye iris data and palmprint data at different times, and obtain the fusion feature vector at different times as the historical feature vector according to steps 2-4, and save it in the historical feature sequence in chronological order. Save the historical feature sequence of different cats in the feature library.

[0036] For the individual to be identified, obtain the fusion feature vector by steps 1-4, and traverse the historical feature sequence in the feature library to obtain the fusion feature vector Cosine similarity is calculated with the last historical feature vector in the historical feature sequence, and the individual corresponding to the highest similarity historical feature vector is selected as the candidate cat i.

[0037] The historical feature sequence of the candidate cat i in the last n times is selected . The historical feature sequence is input into the long short-term memory network LSTM to obtain the predicted feature vector of the time evolution .

[0038] The similarity between the fusion feature vector of the individual to be identified and the historical feature vector of the candidate cat i in the last time , the predicted feature vector is calculated respectively and , and then the final identity confirmation confidence is calculated: wherein, represents the historical similarity, represents the evolution similarity. , are the weights of the historical similarity and the evolution similarity respectively, which are artificial set hyperparameters, in the embodiment =0.4, =0.6.

[0039] The confidence threshold τ=0.85 is set, when Confidence>τ, the identity of the individual to be identified is confirmed as the candidate cat i, and is added to the latest historical feature , while the earliest historical feature in is removed , so that the length of the historical feature sequence is kept as n. If Confidence≤τ, it is determined as an unknown individual, a new identity is registered in the feature library, and is added to the historical feature sequence of the new individual.

Claims

1. A method for long-term identity recognition for multiple cat individuals, characterized in that: The specific steps are as follows: Step 1, collecting multi-modal data of cats, including facial three-dimensional point cloud data, cat eye iris data and palm print data; Step 2, converting the facial three-dimensional point cloud data, cat eye iris data and palm print data to a three-dimensional joint coordinate system through matrix transformation; Step 3, using three parallel feature extraction networks to extract features from the data aligned in step 2, then splicing in the dimension, and projecting back to the original input dimension to obtain the preliminary fusion feature F'; Step 4, feature enhancement is performed on the preliminary fusion feature F' through a self-attention mechanism, and then residual connection is performed with the preliminary fusion feature F' to obtain a fusion feature vector ; Step 5, fusion feature vector of each cat individual in different period The historical feature sequence is saved in chronological order, and historical feature sequences of different cats constitute a feature library; Finding the fusion feature vector in the feature library that is closest to the individual to be identified The closest individual feature vector as a candidate cat; input the historical feature sequence of the candidate cat into the recurrent neural network to generate a predicted feature vector; calculate the fusion feature vector The similarity between the closest historical feature vector and the predicted feature vector of the candidate cat, and the identity confirmation confidence Confidence obtained by weighted summation; if the identity confirmation confidence Confidence is lower than the set threshold, the current cat is judged as an unknown individual, and its features are added to the feature library to register the historical feature sequence of the new individual; otherwise, the identity of the current cat is judged as the candidate cat, and the fusion feature vector is added to the historical feature sequence of the candidate cat, and the earliest historical feature is removed to realize intelligent rolling update of the feature library.

2. The method of claim 1, wherein the method is a multi-cat individual oriented long-term identity recognition method. The nose tip position of the cat in the facial three-dimensional point cloud data is taken as the origin of the three-dimensional joint coordinate system, and then the coordinates of the iris data and the palm print data are converted to the three-dimensional joint coordinate system through the pre-calibrated rotation and translation matrix.

3. The method of claim 2, wherein the method is for long-term identification of multiple cats. The nose tip position of the cat is located by the pose estimation model MediaPipe Animal Pose.

4. The method of claim 1, wherein the method is for long-term identification of multiple cats. The aligned facial three-dimensional point cloud data, cat eye iris data and palm print data are extracted using a 3D sparse convolution network, a ResNet-18 residual network and a lightweight depth separable CNN network, respectively.

5. The method of claim 4, wherein the method is for long-term identification of multiple cats. The feature vectors output by the three feature extraction networks are spliced in the dimension, nonlinearly transformed using a fully connected layer and a ReLU activation function, and then projected back to the original input dimension through convolution operation to obtain the preliminary fusion feature F'.

6. The method of claim 1, wherein the method is a multi-cat individual oriented long-term identity recognition method. Selecting a sequence of historical features of a candidate cat In an input long short-term memory network (LSTM), obtaining a predicted feature vector for the time series evolution .

7. The method of claim 6, wherein the method is for long-term identification of multiple cats. The identity confirmation confidence Is: wherein, represents a historical similarity, represents an evolutionary similarity; , are weights for the historical similarity and the evolutionary similarity, respectively; represents a recent historical feature vector of the candidate cat.

8. The method of claim 7, wherein the method is for long-term identification of multiple cats. Setting weights = 0.4, = 0.6, identity confirmation confidence threshold τ = 0.

85.

9. A long-term identity recognition system for multiple cat individuals, characterized by: The method for realizing the identity recognition method as claimed in any of claims 1-8 comprises a data acquisition module, a spatial alignment module, a feature extraction module, a nonlinear fusion module and a feature library updating module; The data acquisition module is used to acquire facial three-dimensional point cloud data, cat eye iris data and palm print data of cats; The spatial alignment module is used to convert the multi-modal data collected by the data acquisition module into the same three-dimensional joint coordinate system; The feature extraction module uses three parallel feature extraction networks to extract features from the aligned data, respectively; The nonlinear fusion module splices and projects the features output by the feature extraction network in the dimension to obtain the preliminary fusion feature, applies a self-attention mechanism to enhance the preliminary fusion feature, and performs residual connection with the preliminary fusion feature to output a fusion feature vector; The recognition and feature updating module uses a recurrent neural network to perform time series evolution prediction from the historical feature sequence, confirms the identity of the cat individual through the similarity between the predicted feature, the historical feature and the fusion feature, and updates the historical feature sequence using the fusion feature.

10. The multi-cat individual-oriented long-term identity recognition system of claim 9, wherein: The facial three-dimensional point cloud data is obtained by a ToF depth sensor, the cat eye iris data is obtained by a multispectral camera working in the near-infrared band, and the palm print data is obtained by a bioimpedance sensor array.

Citation Information

Patent Citations

  • Judicial scene-oriented multi-modal fusion identity authentication method, medium and equipment

    CN116797895A

  • Wind power plant personnel identity verification system and method based on face recognition and biological characteristics

    CN118940242A

  • Photovoltaic power station fire early warning method fusing visual search and multi-mode large model

    CN119475239A

  • Face authentication apparatus

    US20210056289A1

  • Methods and systems for face recognition

    WO2019100436A1