Pet identification method and device, storage medium and computer equipment

Through the multi-task model architecture and comparative learning method, the accuracy of pet identification is improved, and the problem of insufficient pet identification accuracy in the prior art is solved, which is suitable for pet identification tasks of edge devices.

CN120472503AActive Publication Date: 2025-08-12SHENZHEN LIBRO TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510913169.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The existing pet identification methods have low accuracy, especially in complex scenarios, which cannot fully utilize feature sharing between pet detection and individual recognition tasks, resulting in insufficient accuracy of pet identification.

Method used

The multi-task model architecture is adopted, including image encoder, pet detection branch and pet identification branch. The unsupervised pre-trained image encoder is used, and the model parameter update is performed by combining contrast learning and multi-task label data to realize collaborative training of pet detection and identification tasks.

Benefits of technology

It improves the accuracy of pet identification, especially the ability to identify different pets in a multi-pet environment, reduces the dependence on labeled data, and is suitable for resource-constrained edge device deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472503A_ABST
    Figure CN120472503A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pet identification method and device, a storage medium and computer equipment. The method comprises the following steps: acquiring a pet image collected by the edge pet equipment; calling a preset multi-task model to carry out pet recognition on the pet image to obtain a pet recognition result; the multi-task model comprises an image encoder, a pet detection branch and a pet identification branch, and the training process of the multi-task model comprises the following steps: carrying out unsupervised pre-training on the image encoder based on a historical pet image collected by an edge pet device; screening a positive sample image set and a negative sample image set from the historical pet images, and updating model parameters of the multi-task model based on the positive sample image set and the negative sample image set; and obtaining training sample data with a multi-task label, and updating parameters of the multi-task model based on the training sample data. The method can improve the accuracy of pet recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of pet intelligent technology, and in particular to a pet identification method, device, electronic device and storage medium. Background Art

[0002] In recent years, as the pace of life continues to accelerate, people's life pressures are also increasing. The company of pets can greatly relieve people's mental stress and bring them physical and mental pleasure, which has led to the rapid development of the pet products market in recent years.

[0003] Among them, some pet management applications are provided in the related art, which analyze the pet's behavior by collecting videos or images of the pet's behavior to analyze the pet's behavioral habits, health status, etc. Then, the pet is given corresponding education or treatment based on the analysis results.

[0004] Capturing pet behavior videos or images requires good pet recognition capabilities. However, current pet recognition methods have low accuracy. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a pet identification method, device, electronic device and storage medium, aiming to improve the accuracy of pet identification.

[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a pet identification method, the method comprising: Obtaining a pet image captured by the edge pet device; Calling a preset multi-task model to perform pet recognition on the pet image to obtain a pet recognition result; The multi-task model includes an image encoder, a pet detection branch, and a pet recognition branch. The training process of the multi-task model includes: performing unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; Filtering a positive sample image set and a negative sample image set from the historical pet images, the positive sample image set including multiple different images of the same pet, and the negative sample image set including multiple images of different pets, and updating model parameters of the multi-task model based on the positive sample image set and the negative sample image set; Acquire training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and update the parameters of the multi-task model based on the training sample data.

[0007] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a pet identification device, the device comprising: an acquisition unit, configured to acquire the pet image captured by the edge pet device; an identification unit, configured to call a preset multi-task model to perform pet identification on the pet image to obtain a pet identification result; The multi-task model includes an image encoder, a pet detection branch, and a pet recognition branch. The training process of the multi-task model includes: performing unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; Filtering a positive sample image set and a negative sample image set from the historical pet images, the positive sample image set including multiple different images of the same pet, and the negative sample image set including multiple images of different pets, and updating model parameters of the multi-task model based on the positive sample image set and the negative sample image set; Acquire training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and update the parameters of the multi-task model based on the training sample data.

[0008] In some embodiments, the present application further provides a model training device, comprising: a pre-training unit, configured to perform unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; a first updating unit, configured to screen a positive sample image set and a negative sample image set from the historical pet images, wherein the positive sample image set includes multiple different images of the same pet, and the negative sample image set includes multiple images of different pets, and update model parameters of the multi-task model based on the positive sample image set and the negative sample image set; The second updating unit is used to obtain training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and to update the parameters of the multi-task model based on the training sample data.

[0009] Optionally, in some embodiments, the first updating unit includes: a first screening subunit, configured to screen out a plurality of different images of the same pet captured at different shooting angles from the historical pet images to obtain a positive sample image set; a second screening subunit, configured to screen out images of pets of the same breed captured from multiple angles from the historical pet images, to obtain a negative sample image set; A first updating subunit is configured to calculate a contrastive learning loss based on the positive sample image set and the negative sample image set, and update model parameters of the multi-task model based on the contrastive learning loss.

[0010] Optionally, in some embodiments, the first updating subunit includes: A first calculation module is used to calculate the contrastive learning loss based on the positive sample image set and the negative sample image set; A first updating module is used to update the parameters of the image encoder and the pet recognition branch according to the contrastive learning loss.

[0011] Optionally, in some embodiments, the second updating unit includes: A first acquisition subunit is configured to acquire training sample data with multi-task labels, wherein the training sample data includes a plurality of sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image; A prediction subunit, configured to input the sample pet image into the multi-task model to obtain a detection result output by the pet detection branch and a recognition result output by the pet recognition branch; a first calculation subunit, configured to calculate a pet detection loss based on the detection result and the corresponding pet detection tag; a second calculation subunit, configured to calculate a pet identification loss based on the identification result and the pet identification tag; The second updating subunit is configured to calculate a target loss based on the pet detection loss and the pet identification loss, and to update parameters of the pet detection branch and the pet identification branch based on the target loss.

[0012] Optionally, in some embodiments, the second updating subunit includes: a determination module, configured to determine a first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet identification loss; a second calculation module, configured to perform weighted calculation on the pet detection loss and the pet identification loss based on the first weight coefficient and the second weight coefficient to obtain a target loss; The second updating module is used to update the parameters of the pet detection branch and the pet identification branch based on the target loss.

[0013] Optionally, in some embodiments, the determination module includes: An acquisition submodule, configured to acquire the task difficulty and convergence status corresponding to the pet detection branch and the pet identification branch; A determination submodule is used to determine a first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet recognition loss according to the task difficulty and the convergence state.

[0014] Optionally, in some embodiments, the prediction subunit includes: a first prediction module, configured to input the sample pet image into the image encoder to obtain output image features; a second prediction module, configured to process the image features based on the first gating network to obtain pet detection features, and input the pet detection features into the pet detection branch to obtain a detection result; The third prediction module is used to process the image features based on the second gating network to obtain pet identification features, and input the pet identification features into the pet identification branch to obtain an identification result.

[0015] Optionally, in some embodiments, the identification unit includes: A training subunit, configured to perform knowledge distillation based on the pet recognition branch of the preset multi-task model to obtain a student model; a deployment subunit, configured to deploy the student model in the edge pet device; The recognition subunit is used to call the student model to perform pet recognition on the pet image to obtain a pet recognition result.

[0016] Optionally, in some embodiments, the model training device provided by the present application further includes: A second acquisition subunit is configured to acquire recognition correction data uploaded by a user and generate a fine-tuning sample based on the recognition correction data; A fine-tuning subunit is configured to periodically fine-tune the student model based on the fine-tuning samples.

[0017] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the pet identification method described in the first aspect when executing the computer program.

[0018] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the pet identification method described in the first aspect.

[0019] To achieve the above-mentioned purpose, the fifth aspect of the embodiment of the present application proposes a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the pet identification method described in the first aspect.

[0020] The pet identification method proposed in the embodiment of the present application obtains pet images collected by an edge pet device; calls a preset multi-task model to perform pet identification on the pet images to obtain a pet identification result; the multi-task model includes an image encoder, a pet detection branch, and a pet identification branch, and the training process of the multi-task model includes: unsupervised pre-training of the image encoder based on historical pet images collected by the edge pet device; screening out a positive sample image set and a negative sample image set from the historical pet images, the positive sample image set including multiple different images of the same pet, and the negative sample image set including multiple images of different pets, and updating the model parameters of the multi-task model based on the positive sample image set and the negative sample image set; obtaining training sample data with multi-task labels, the training sample data including multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and updating the parameters of the multi-task model based on the training sample data.

[0021] It can be seen from this that the pet recognition method provided in the embodiment of the present application trains the pet recognition model through a three-stage learning architecture. On the one hand, a large number of historical pet images are used to pre-train the image encoder in the model, so that the image encoder can learn to obtain good image feature extraction capabilities. The model is then compared and learned through positive and negative samples to improve the generalization ability of the model. Furthermore, the model in this case adopts a multi-task model that includes pet detection tasks and recognition tasks, so that the detection tasks and recognition tasks promote each other, thereby improving the pet recognition ability of the model. Therefore, the pet recognition method provided in this case can greatly improve the accuracy of pet recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0023] Figure 1 A flowchart of the pet identification method provided in this application; Figure 2 A schematic diagram of the model structure of the multi-task model provided in this application; Figure 3 A flowchart of the model training method provided in this application; Figure 4 A schematic diagram of the structure of the student model provided for this application; Figure 5 A schematic diagram of the structure of the pet identification device provided in this application; Figure 6 A schematic diagram of the structure of the model training device provided in this application; Figure 7 This is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0025] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations: Contrastive learning: Contrastive learning aims to teach a model to distinguish between similar and dissimilar data samples. During training, the model maximizes the similarity of pairs of positive samples while minimizing the similarity of pairs of negative samples. A similarity function (such as cosine similarity) is typically used to measure the similarity between pairs of samples, and model parameters are updated by optimizing an objective function (such as contrastive loss). Contrastive learning can utilize unlabeled data for pre-training, reducing reliance on labeled data and saving manpower and material resources. By comparing positive and negative samples, the model learns the intrinsic structure and semantic information of the data, deriving feature representations with strong generalization capabilities, and improving the model's performance on downstream tasks.

[0026] Knowledge Distillation (KD): A model compression technique whose core idea is to transfer the knowledge from a large, complex teacher model to a small, simple student model, so that the student model can keep the performance as close as possible to the teacher model while maintaining a small scale.

[0027] In related technologies, various edge pet devices, such as pet feeders, pet smart toilets, and pet toys, only have simple command execution functions, such as feeding food at regular intervals and in fixed quantities according to the pet feeding plan, and processing pet excrement after detecting that the pet has gone to the toilet. In addition to achieving the above functions, the edge pet device provided in this case can also be equipped with an image / video acquisition device and a data processing device. The edge pet device can capture pet images or pet videos of the pet in real time, and then use a data processing device to detect the pet's behavioral habits or health status based on the pet images or pet videos, and then provide corresponding guidance or medical advice based on the detection results.

[0028] In some multi-pet scenarios, such as in multi-pet households or pet stores, the edge pet device needs to perform pet identification after collecting pet images, and then filter out the accurate pet image of each pet based on the pet identification results, and then detect the behavioral habits or health status based on the accurate pet image to obtain accurate detection results.

[0029] However, current pet detection methods typically use a pet detection model to identify the presence of a pet in an image and, if so, its location. A pet identification model is then used to individually identify each detected pet. This independent or cascaded model architecture handles pet detection and individual identification tasks separately, failing to fully leverage feature sharing between these tasks. This results in insufficient pet identification accuracy in complex scenarios.

[0030] In summary, the current methods for pet recognition based on pet images are not very accurate.

[0031] Based on this, the embodiments of the present application provide a pet identification method, device, electronic device and storage medium, in order to improve the accuracy of pet identification to a certain extent.

[0032] Reference Figure 1 In some embodiments, the pet identification method provided in the embodiments of the present application includes but is not limited to steps S101 to S102.

[0033] Step S101, obtaining a pet image captured by an edge pet device; Step S102: Calling a preset multi-task model to perform pet recognition on the pet image to obtain a pet recognition result.

[0034] In steps S101 to S102 shown in the embodiment of the present application, the edge pet device may be the pet feeder, pet smart toilet, etc. in the aforementioned examples. The edge pet device may be equipped with an image acquisition device, such as a camera, to acquire pet images of pets in the environment. The image acquisition device may acquire images in real time, or may acquire images when a pet is detected in the field of view, or may acquire images based on specific instruction triggers. The specific instructions here may be pet behavior analysis instructions, pet health analysis instructions, or pet food intake assessment instructions, etc. These instructions may trigger image acquisition of the pet, and then perform tasks such as pet behavior analysis, health analysis, and food intake assessment based on the acquired pet images.

[0035] After capturing pet images, the edge pet device can first perform pet recognition on the pet images to identify the individual pet information in each pet image. Based on the identified individual pet information, the pet images captured by the edge pet device are then classified into different pet individuals, generating a set of pet images corresponding to each individual pet. This allows for further downstream tasks such as pet behavior analysis, pet health analysis, or food intake assessment to be performed on each pet based on the pet image set corresponding to each individual pet.

[0036] Among them, in the embodiment of the present application, the specific process of pet recognition on pet images is that the edge pet device calls a preset multi-task model to perform pet recognition on pet images. The preset multi-task model can be a multi-task model deployed on the edge pet device or a multi-task model deployed on a cloud server. The multi-task model provided in this application includes an image encoder, a pet detection branch, and a pet recognition branch. Figure 2 As shown in FIG, a schematic diagram of the model structure of the multi-task model provided by the present application. As shown in the figure, the multi-task model includes an image encoder 210, a pet detection branch 220, and a pet recognition branch 230. The pet detection branch 220 and the pet recognition branch 230 share the underlying image encoder 210 to extract image features. Specifically, in order to improve the ability of pet recognition, the present application also provides a method for training the above multi-task model, such as Figure 3 , which is a flow chart of the model training method provided in this application, and the method includes but is not limited to steps S301 to S303.

[0037] Step S301, performing unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; Step S302: Filtering a positive sample image set and a negative sample image set from historical pet images, where the positive sample image set includes multiple different images of the same pet, and the negative sample image set includes multiple images of different pets, and updating the model parameters of the multi-task model based on the positive sample image set and the negative sample image set; Step S303: Acquire training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and update the parameters of the multi-task model based on the training sample data.

[0038] The model training method provided in this application adopts an innovative three-stage training process. First, the image encoder used for shared feature extraction is unsupervisedly pre-trained to enable the image encoder to learn the ability to extract rich image features from the input image. Then, a contrastive learning method is used to train the model's ability to distinguish different pets. Finally, a certain amount of labeled samples is used to train the model's ability to identify individual pets. This method does not require a large amount of annotation information, which can reduce the difficulty of obtaining training sample data. In addition, this method addresses the problem of high similarity between pets of the same breed in pet identification scenarios. By combining contrastive learning tasks with pet identification tasks, it can improve the recognition accuracy of similar pets of the same breed. In addition, the model provided in this application adopts a multi-task head structure. When using labeled sample data for end-to-end training, since the parameters of the two task branches are updated simultaneously, the tasks can promote each other's improvement, thereby improving the model's accuracy in pet identification.

[0039] Specifically, the image encoder can be pre-trained in an unsupervised manner, and historical pet images collected by edge pet devices can be used as samples for training. The edge pet device can be deployed in the pet's living environment, and can collect images of the pet to obtain a large number of pet images. The image encoder can then be trained based on the historical pet images collected in the environment. When the image encoder is unsupervisedly trained based on the collected historical pet images, the image encoder can be pre-trained through self-supervised tasks such as image reconstruction and contrast prediction. In the related art, training a pet recognition model generally uses a large amount of labeled data to perform supervised training on the pet recognition model, but supervised training is difficult to label and samples are difficult to obtain, resulting in insufficient sample size. The present application uses pet images collected in real time by edge pet devices to perform regular unsupervised training on the model's image encoder, so that the image encoder has the ability to learn to obtain richer feature representations, and the image encoder's ability to capture pet features can continue to evolve as the pet grows and changes, thereby greatly improving the model's ability to capture pet features.

[0040] In some embodiments, a positive sample image set and a negative sample image set are screened from historical pet images, where the positive sample image set includes multiple different images of the same pet, and the negative sample image set includes multiple images of different pets. Model parameters of the multi-task model are updated based on the positive sample image set and the negative sample image set, including: Filter multiple different images of the same pet taken from different shooting angles from historical pet images to obtain a set of positive sample images; Filter out images of pets of the same breed from different angles from historical pet images to obtain a negative sample image set; The contrastive learning loss is calculated based on the positive sample image set and the negative sample image set, and the model parameters of the multi-task model are updated based on the contrastive learning loss.

[0041] In this embodiment of the present application, to improve the model's recognition capabilities in a multi-pet environment, that is, to improve the model's ability to distinguish between different pets, a contrastive learning method is used to train the model's pet differentiation capabilities. In pet identification scenarios, there is a difficult task, namely, the task of identifying different pets of the same breed. Therefore, in this embodiment of the present application, when performing contrastive learning training on the model, some difficult samples can be selected for training to improve the model's ability to distinguish between different pets.

[0042] Specifically, multiple different images of the same pet taken from different shooting angles can be screened out from historical pet images collected by edge pet devices to form a positive sample image set.

[0043] Then, images of different pets of the same breed taken from multiple angles are filtered out from historical pet images collected by edge pet devices to form a negative sample image set. For example, negative sample image sets can be obtained by images of different pets of the same breed taken from the face angle, or by images of different pets of the same breed taken from the body angle.

[0044] After collecting the positive sample image set and the negative sample image set, the positive sample image set and the negative sample image set can be used to calculate the contrastive learning loss, and then the contrastive learning loss can be used to train the multi-task model.

[0045] Positive pairs are not constrained to specific parts of the face or body. Instead, they intentionally include images of the same pet from multiple angles and poses, allowing the model to learn the invariant characteristics of the individual in various situations. Similarly, negative pairs do not require specific constraints to maximize the model's generalization ability in complex real-world scenarios.

[0046] In some embodiments, calculating a contrastive learning loss based on a set of positive sample images and a set of negative sample images, and updating model parameters of a multi-task model based on the contrastive learning loss includes: Compute contrastive learning loss based on a set of positive and negative images. The parameters of the image encoder and pet recognition branches are updated according to the contrastive learning loss.

[0047] In an embodiment of the present application, after obtaining a positive sample image set and a negative sample image set, the multi-task model provided by the present application can be unsupervisedly trained based on the positive sample image set and the negative sample image set. Specifically, the corresponding contrastive learning loss can be calculated for the pet recognition task based on the positive sample image set and the negative sample image set. Specifically, positive sample pairs can be collected from the positive sample image set and negative sample pairs can be collected from the negative sample image set respectively, and then the contrastive learning loss can be calculated by calculating the InfoNCE loss. The goal of this loss is to bring the positive sample pairs closer and push the negative sample pairs further away.

[0048] After calculating the contrastive learning loss, the model parameters of the pet recognition branch can be further updated based on the contrastive learning loss, thereby achieving pre-training of the pet recognition branch so that it has the ability to distinguish different pets.

[0049] The contrastive learning training process in the second phase of this application must be based on the unsupervised pre-training of the image encoder in the first phase, so that the model has already learned universal representations of pets at different angles, postures, and parts (such as the face and body). Based on the model's ability to extract accurate universal representations of pets, positive and negative samples are used to further update the parameters of the image encoder and pet recognition branches, allowing the model to further learn the ability to accurately distinguish pets based on its ability to accurately extract pet image features.

[0050] In some embodiments, obtaining training sample data with multi-task labels, the training sample data including multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and updating the parameters of the multi-task model based on the training sample data includes: Obtaining training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image; Input the sample pet image into the multi-task model to obtain the detection result output by the pet detection branch and the recognition result output by the pet recognition branch; Calculate pet detection loss based on the detection results and the corresponding pet detection labels; Calculate pet identification loss based on the identification result and the pet identification tag; The target loss is calculated based on the pet detection loss and the pet recognition loss, and the parameters of the pet detection branch and the pet recognition branch are updated according to the target loss.

[0051] After pre-training the image encoder in the multi-task model using a large number of unlabeled pet images collected by edge pet devices, and then using positive and negative sample pairs screened from the images collected by edge pet devices to perform comparative learning training on the image encoder and pet recognition branch network, training sample data with multi-task labels can be further obtained. The training sample data can include multiple sample pet images and the pet detection label and pet identification label corresponding to each sample pet image. The pet detection label can be the location box of the pet in the image, and the pet identification label can be the pet ID corresponding to the pet. A sample pet image can include multiple pet detection labels and multiple pet identification labels.

[0052] Specifically, after obtaining the training sample data, the sample pet images included in each training sample data can be first input into an image encoder for image feature extraction, and then the image features can be input into the pet detection branch and the pet recognition branch for pet detection and pet recognition, respectively, to obtain output detection results and recognition results. Furthermore, the pet detection loss can be calculated based on the detection results and the corresponding pet detection labels, and the pet recognition loss can be calculated based on the recognition results and the pet recognition labels. Furthermore, the target loss is calculated based on the pet detection loss and the pet recognition loss, and the parameters of the pet detection branch and the pet recognition branch are further updated based on the target loss to obtain a trained multi-task model.

[0053] In related technologies, independent or cascaded model architectures are usually used to handle pet detection and individual recognition tasks respectively, and features between tasks cannot be shared. Moreover, the separate model design increases the redundancy of parameters, resulting in high computational overhead during the inference process, making it difficult to deploy on resource-constrained edge devices. In addition, the isolated feature learning method prevents the model from benefiting from multi-task collaboration, thereby limiting the improvement of the overall performance of the model. The model architecture and corresponding training method provided in this application integrate pet detection and pet recognition tasks into a single model, and then train the model end-to-end, thereby achieving sharing and complementarity between features within the model, which can significantly improve the accuracy of pet recognition.

[0054] The multi-task model proposed in this application differs from the dual-tower multi-gate hybrid expert network architecture in that both the pet detection branch and the pet recognition branch rely on shared features output by a pre-trained image encoder. These shared features have stronger generalization properties and are incorporated into subsequent contrastive learning and multi-task fine-tuning stages, enabling knowledge transfer and complementarity between tasks.

[0055] In some embodiments, calculating the target loss based on the pet detection loss and the pet identification loss, and updating the parameters of the pet detection branch and the pet identification branch based on the target loss include: Determining a first weight coefficient corresponding to a pet detection loss and a second weight coefficient corresponding to a pet identification loss; Perform weighted calculation on the pet detection loss and the pet recognition loss based on the first weight coefficient and the second weight coefficient to obtain the target loss; The parameters of the pet detection branch and the pet recognition branch are updated based on the target loss.

[0056] In the disclosed embodiments, the model's effectiveness can be adjusted by assigning different weights to multiple task branches. Specifically, a first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet recognition loss can be determined. Then, based on the first and second weight coefficients, a weighted calculation can be performed on the pet detection loss and the pet recognition loss to obtain the target loss.

[0057] The weight coefficients corresponding to the pet detection loss and the pet recognition loss can be set to fixed weight coefficients, such as 0.5 and 0.5, according to the importance of the task; or, in order to enhance the importance of the pet recognition branch, the weight coefficients can be set to 0.4 and 0.6. In some embodiments, the weight coefficients corresponding to the pet detection loss and the pet recognition loss can also be dynamic weights, that is, in the early stages of training, the pet detection loss and the pet recognition loss are kept consistent, and then as the number of training iterations increases, the weight of the pet detection loss is continuously reduced and the weight of the pet recognition loss is increased, thereby continuously enhancing the importance of the pet recognition branch.

[0058] In some embodiments, determining a first weight coefficient corresponding to a pet detection loss and a second weight coefficient corresponding to a pet identification loss includes: Obtain the task difficulty and convergence status corresponding to the pet detection branch and the pet recognition branch; A first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet recognition loss are determined according to the task difficulty and the convergence state.

[0059] In an embodiment of the present application, a multi-task loss weighting method based on uncertainty is provided to allow the model to automatically learn the relative weights between different tasks. This method uses the inherent uncertainty of each task as the basis for determining the weights, allowing the model to dynamically adjust the respective contributions of the pet detection task and the pet recognition task based on the difficulty and convergence status of the corresponding task branches. This adaptive weighting strategy avoids the subjectivity caused by manually set fixed weights, allowing the model to naturally balance the various tasks during training, thereby achieving better training results.

[0060] Specifically, when supervised training a multi-task model based on training sample data, the task difficulty and convergence status of the pet detection and pet identification branches can be obtained in real time. A first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet identification loss are then determined based on the task difficulty and convergence status. A target loss is then calculated based on the first and second weight coefficients, and the model parameters are updated based on the target loss.

[0061] In some embodiments, when performing supervised fine-tuning based on training sample data, a progressive training strategy may be adopted to avoid interference between tasks, thereby ensuring recognition accuracy.

[0062] In some embodiments, a sample pet image is input into a multi-task model to obtain a detection result output by a pet detection branch and a recognition result output by a pet recognition branch, including: Input the sample pet image into the image encoder and obtain the output image features; Processing the image features based on the first gating network to obtain pet detection features, inputting the pet detection features into the pet detection branch to obtain a detection result; The image features are processed based on the second gating network to obtain pet identification features, which are then input into the pet identification branch to obtain the identification results.

[0063] In an embodiment of the present application, a gating network can also be set up for each task branch in the model structure. This gating network is used to adjust the adaptability between the output of the image encoder and each task branch. When there is no significant conflict between the pet detection task and the pet recognition task, the output of the image encoder can be directly used as the input feature of the pet detection branch and the pet recognition branch, that is, the gating network can be ineffective at this time. When significant conflict is identified between multiple task branches during the multi-task fine-tuning process, the gating network corresponding to each task branch can be used to separate the feature learning paths. In this case, a sample pet image can be first input into the image encoder to obtain output image features. Then, based on the first gating network corresponding to the pet detection branch, the image features are processed to obtain pet detection features, which are then input into the pet detection branch to obtain a detection result. Finally, based on the second gating network corresponding to the pet recognition branch, the image features are processed to obtain pet recognition features, which are then input into the pet recognition branch to obtain a recognition result.

[0064] Furthermore, corresponding loss values are calculated based on the detection results and recognition results and the corresponding pet detection tags and pet recognition tags, and the model parameters are updated based on the loss values.

[0065] In some embodiments, calling a preset multi-task model to perform pet recognition on a pet image to obtain a pet recognition result includes: The student model is obtained by performing knowledge distillation on the pet recognition branch of the preset multi-task model; Deploy student models in edge PET devices; Call the student model to perform pet recognition on the pet image and obtain the pet recognition result.

[0066] As mentioned above, when an edge pet device calls a preset multi-task model to perform pet recognition on pet images, it can call a locally deployed multi-task model or a multi-task model deployed on a cloud server. However, pet recognition on images collected by edge pet devices is generally a subtask performed in tasks such as pet behavior analysis and pet health analysis based on pet images. At this time, the pet images collected are generally daily images of pet owners. Uploading these images to a cloud server for pet recognition can easily lead to privacy leaks. Uploading a large number of pet images to the cloud also consumes a lot of bandwidth and transmission time. Therefore, it is necessary to be able to deploy a model locally on the edge pet device to complete pet recognition. However, due to insufficient device hardware conditions of edge pet devices, their computing resources are limited and cannot meet the deployment conditions of large-structure models.

[0067] In an embodiment of the present application, by adopting the knowledge distillation method, a complete multi-task model is first trained in the cloud, and then the image encoder and the corresponding pet recognition branch capabilities in the multi-task model are transferred to the student model through the knowledge distillation method, thereby training the student model, and then the student model can be deployed on the edge pet device. In particular, since the pet detection task branch can promote and complement the pet recognition task branch during the model training process, the recognition ability of the pet recognition branch is improved. During model inference, the features output by the image encoder can be directly input into the pet recognition branch to obtain the pet recognition result, that is, the pet detection branch of the model does not need to be used during the inference process. In this way, in the distillation learning stage, only the image encoder and pet recognition branches can be distilled and learned to obtain the student model. As Figure 4 Figure 2 shows a schematic diagram of the student model provided in this application. The student model includes a feature extraction layer 410 and a pet recognition layer 420. The feature extraction layer is used to extract image features from pet images, and the pet recognition layer is used to predict pet recognition results based on the image features. Feature extraction layer 410 and pet recognition layer 420 are obtained through distillation learning based on the image encoder 210 and the pet recognition branch 230.

[0068] After deploying the student model obtained through distillation learning on the edge pet device, when a pet image is obtained, the student model deployed on the edge pet device can be directly called to perform pet recognition on the obtained pet image, thereby obtaining accurate pet recognition results.

[0069] In this embodiment of the present application, the teacher model deployed in the cloud can be continuously trained using desensitized data uploaded by users, and the student model deployed on the edge pet device can be regularly updated from the cloud using an incremental update mechanism to adapt to new pet individuals and feature changes. Specifically, during the distillation learning phase, the student model can learn from the intermediate feature representations of the teacher model and imitate the teacher model's predictive distribution to learn.

[0070] In some embodiments, the model training method provided by this application further includes: Obtain recognition correction data uploaded by users and generate fine-tuning samples based on the recognition correction data; The student model is periodically fine-tuned based on the fine-tuning samples.

[0071] Among them, in the pet handling application, some tasks will provide users with feedback on the results of pet identification. For example, in a pet behavior analysis task, when a pet is identified to have abnormal behavior, the pet handling application can use the identified pet information (such as pet ID), pet abnormality information (such as the presence of housebreaking behavior), and provide the corresponding pet image as evidence of the pet abnormality information. At this time, the pet owner can use the provided pet image to determine whether the pet in the pet image corresponds to the pet ID. If they do, the pet identification result is accurate; if they do not, the pet identification result is inaccurate. At this time, the pet owner can provide feedback through a pre-set feedback interface, such as providing the accurate pet ID corresponding to the pet image as identification correction data.

[0072] In an embodiment of the present application, the recognition correction data uploaded by the user can be collected and summarized, and then fine-tuning samples can be generated based on the recognition correction data, and finally the student model can be regularly fine-tuned based on the fine-tuning samples.

[0073] Among them, in the related art, training of pet recognition models generally involves collecting training sample data for training, but the importance and quality of the training samples are not distinguished between the samples, making it impossible to make targeted improvements to the capabilities of the pet recognition model. In the embodiments of the present application, the correlation between the pet recognition task and other tasks is utilized, and the interactive function between the task and the user is provided to collect verification data of the user's recognition results of the model. From the verification data, high-quality fine-tuning samples that can be used to make targeted improvements to some of the model's capabilities can be extracted. These fine-tuning samples are then used to make targeted fine-tuning of the student model, thereby continuously improving the model's pet recognition capabilities.

[0074] Reference Figure 5 In some embodiments, the present application further provides a pet identification device 500, which includes: An acquisition unit 510 is configured to acquire pet images captured by an edge pet device; The recognition unit 520 is used to call a preset multi-task model to perform pet recognition on the pet image to obtain a pet recognition result.

[0075] Reference Figure 6 In some embodiments, the present application further provides a model training device 600, which includes: A pre-training unit 610 is configured to perform unsupervised pre-training on an image encoder based on historical pet images collected by an edge pet device; A first updating unit 620 is configured to filter out a positive sample image set and a negative sample image set from historical pet images, where the positive sample image set includes multiple different images of the same pet, and the negative sample image set includes multiple images of different pets, and update model parameters of the multi-task model based on the positive sample image set and the negative sample image set; The second updating unit 630 is used to obtain training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and to update the parameters of the multi-task model based on the training sample data.

[0076] Optionally, in some embodiments, the first updating unit includes: The first screening subunit is configured to screen multiple different images of the same pet taken from different shooting angles from historical pet images to obtain a positive sample image set; The second screening subunit is used to screen out images of pets of the same breed captured from multiple angles from historical pet images to obtain a negative sample image set; The first updating subunit is used to calculate the contrastive learning loss based on the positive sample image set and the negative sample image set, and update the model parameters of the multi-task model based on the contrastive learning loss.

[0077] Optionally, in some embodiments, the first updating subunit includes: A first calculation module is used to calculate the contrastive learning loss based on the positive sample image set and the negative sample image set; The first update module is used to update the parameters of the image encoder and pet recognition branch based on the contrastive learning loss.

[0078] Optionally, in some embodiments, the second updating unit includes: A first acquisition subunit is configured to acquire training sample data with multi-task labels, where the training sample data includes a plurality of sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image; The prediction subunit is used to input the sample pet image into the multi-task model to obtain the detection result output by the pet detection branch and the recognition result output by the pet recognition branch; a first calculation subunit, configured to calculate a pet detection loss based on the detection result and the corresponding pet detection tag; a second calculation subunit, configured to calculate a pet identification loss based on the identification result and the pet identification tag; The second updating subunit is used to calculate the target loss based on the pet detection loss and the pet recognition loss, and to update the parameters of the pet detection branch and the pet recognition branch based on the target loss.

[0079] Optionally, in some embodiments, the second updating subunit includes: a determination module, configured to determine a first weight coefficient corresponding to a pet detection loss and a second weight coefficient corresponding to a pet identification loss; A second calculation module is used to perform weighted calculation on the pet detection loss and the pet recognition loss based on the first weight coefficient and the second weight coefficient to obtain a target loss; The second updating module is used to update the parameters of the pet detection branch and the pet recognition branch based on the target loss.

[0080] Optionally, in some embodiments, the determination module includes: The acquisition submodule is used to obtain the task difficulty and convergence status corresponding to the pet detection branch and the pet recognition branch; The determination submodule is used to determine a first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet recognition loss according to the task difficulty and the convergence state.

[0081] Optionally, in some embodiments, the prediction subunit includes: A first prediction module is configured to input a sample pet image into an image encoder to obtain output image features; A second prediction module is configured to process the image features based on the first gating network to obtain pet detection features, and input the pet detection features into the pet detection branch to obtain a detection result; The third prediction module is used to process the image features based on the second gating network to obtain pet identification features, and input the pet identification features into the pet identification branch to obtain a recognition result.

[0082] Optionally, in some embodiments, the identification unit includes: The training subunit is used to perform knowledge distillation based on the pet recognition branch of the preset multi-task model to obtain a student model; a deployment subunit for deploying the student model in an edge PET device; The recognition subunit is used to call the student model to perform pet recognition on the pet image and obtain the pet recognition result.

[0083] Optionally, in some embodiments, the model training device provided by the present application further includes: A second acquisition subunit is used to acquire the recognition correction data uploaded by the user and generate a fine-tuning sample based on the recognition correction data; The fine-tuning subunit is used to periodically fine-tune the student model based on the fine-tuning samples.

[0084] Reference Figure 7 , Figure 7 The hardware structure of a computer device according to another embodiment is shown. The computer device includes: The processor 701 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 702 and are called by the processor 701 to execute the pet identification method of the embodiments of this application. Input / output interface 703, used to implement information input and output; Communication interface 704, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 705 , which transmits information between various components of the device (e.g., processor 701 , memory 702 , input / output interface 703 , and communication interface 704 ); The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .

[0085] The embodiment of the present application further provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device executes and implements the above-mentioned pet identification method.

[0086] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0087] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0088] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0089] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0090] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0091] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0093] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0094] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A pet identification method, characterized in that: The method is applied to an edge pet device, and the method includes: Obtaining a pet image captured by the edge pet device; Calling a preset multi-task model to perform pet recognition on the pet image to obtain a pet recognition result; The multi-task model includes an image encoder, a pet detection branch, and a pet recognition branch. The training process of the multi-task model includes: performing unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; Filtering a positive sample image set and a negative sample image set from the historical pet images, the positive sample image set including multiple different images of the same pet, and the negative sample image set including multiple images of different pets, and updating model parameters of the multi-task model based on the positive sample image set and the negative sample image set; Acquire training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and update the parameters of the multi-task model based on the training sample data.

2. The method according to claim 1, characterized in that The method further comprises: screening a positive sample image set and a negative sample image set from the historical pet images, wherein the positive sample image set includes multiple different images of the same pet, and the negative sample image set includes multiple images of different pets; and updating the model parameters of the multi-task model based on the positive sample image set and the negative sample image set, including: Filtering multiple different images of the same pet captured from different shooting angles from the historical pet images to obtain a positive sample image set; Filtering images of pets of the same breed captured from multiple angles from the historical pet images to obtain a negative sample image set; A contrastive learning loss is calculated based on the positive sample image set and the negative sample image set, and model parameters of the multi-task model are updated based on the contrastive learning loss.

3. The method according to claim 2, characterized in that The calculating the contrastive learning loss based on the positive sample image set and the negative sample image set, and updating the model parameters of the multi-task model based on the contrastive learning loss, includes: Calculating contrastive learning loss based on the positive sample image set and the negative sample image set; Parameters of the image encoder and the pet recognition branch are updated according to the contrastive learning loss.

4. The method according to claim 1, wherein The method includes obtaining training sample data with multi-task labels, wherein the training sample data includes a plurality of sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and updating the parameters of the multi-task model based on the training sample data, including: Acquire training sample data with multi-task labels, wherein the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image; Inputting the sample pet image into the multi-task model to obtain a detection result output by the pet detection branch and a recognition result output by the pet recognition branch; Calculating a pet detection loss based on the detection result and the corresponding pet detection tag; Calculating a pet identification loss based on the identification result and the pet identification tag; A target loss is calculated based on the pet detection loss and the pet identification loss, and parameters of the pet detection branch and the pet identification branch are updated based on the target loss.

5. The method according to claim 4, characterized in that The calculating the target loss according to the pet detection loss and the pet identification loss, and updating the parameters of the pet detection branch and the pet identification branch according to the target loss, includes: Determining a first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet identification loss; Performing weighted calculation on the pet detection loss and the pet identification loss based on the first weight coefficient and the second weight coefficient to obtain a target loss; Parameters of the pet detection branch and the pet recognition branch are updated based on the target loss.

6. The method according to claim 5, characterized in that The determining of the first weight coefficient corresponding to the pet detection loss and the second weight coefficient corresponding to the pet identification loss includes: Obtaining the task difficulty and convergence status corresponding to the pet detection branch and the pet identification branch; A first weight coefficient corresponding to the pet detection loss and a second weight coefficient corresponding to the pet recognition loss are determined according to the task difficulty and the convergence state.

7. The method according to claim 4, characterized in that Inputting the sample pet image into the multi-task model to obtain a detection result output by the pet detection branch and a recognition result output by the pet recognition branch includes: Inputting the sample pet image into the image encoder to obtain output image features; Processing the image features based on the first gating network to obtain pet detection features, and inputting the pet detection features into the pet detection branch to obtain a detection result; The image features are processed based on the second gating network to obtain pet identification features, and the pet identification features are input into the pet identification branch to obtain an identification result.

8. The method according to claim 1, characterized in that Calling a preset multi-task model to perform pet recognition on the pet image to obtain a pet recognition result, including: Perform knowledge distillation based on the pet recognition branch of the preset multi-task model to obtain a student model; deploying the student model in the edge PET device; The student model is called to perform pet recognition on the pet image to obtain a pet recognition result.

9. The method according to claim 8, characterized in that The method further comprises: Obtaining recognition correction data uploaded by the user, and generating fine-tuning samples based on the recognition correction data; The student model is periodically fine-tuned based on the fine-tuning samples.

10. A pet identification device, characterized in that: The device is applied to edge pet equipment, and the device includes: an acquisition unit, configured to acquire the pet image captured by the edge pet device; an identification unit, configured to call a preset multi-task model to perform pet identification on the pet image to obtain a pet identification result; The multi-task model includes an image encoder, a pet detection branch, and a pet recognition branch. The training process of the multi-task model includes: performing unsupervised pre-training on the image encoder based on historical pet images collected by the edge pet device; Filtering a positive sample image set and a negative sample image set from the historical pet images, the positive sample image set including multiple different images of the same pet, and the negative sample image set including multiple images of different pets, and updating model parameters of the multi-task model based on the positive sample image set and the negative sample image set; Acquire training sample data with multi-task labels, where the training sample data includes multiple sample pet images and a pet detection label and a pet identification label corresponding to each sample pet image, and update the parameters of the multi-task model based on the training sample data.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the pet identification method according to any one of claims 1 to 9 is implemented.

12. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the pet identification method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program, characterized in that The computer program is read and executed by a processor of a computer device, so that the computer device executes the pet identification method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Mammary gland molybdenum target image segmentation method and system based on multi-view self-supervised deep learning

    CN115170505A

  • Multi-task data real-time detection method and device for table tennis match video

    CN116958861A

  • Pulmonary nodule detection and semantic attribute rating method based on multi-task learning

    CN117274198A

  • Systems and methods for contrastive learning of visual representations

    US20210319266A1

  • Self-Supervised Privacy Preservation Action Recognition System

    US20240331389A1