Pet nose print feature extraction method, device, electronic device and storage medium
By calling the matching object detection model based on the pet species and feeding time, the problem of low identity recognition accuracy caused by the differences in the distribution of different pet facial features is solved, and a higher accuracy of pet identity recognition is achieved.
Patent Information
- Application Number
- CN202210597619.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, due to the different distributions of different pet facial features, using the same model for object detection leads to low compatibility and segmentation accuracy of pet identity recognition, which affects the accuracy of identity recognition.
According to the type of pet and the length of feeding, a matching object detection model is called to conduct targeted object detection and feature extraction on pet images to improve the segmentation accuracy of the nose.
Through the use of targeted detection models, the segmentation accuracy of pet noses is improved, thus making the extracted nose pattern features more accurate and improving the accuracy of pet identity recognition.
Smart Images

Figure CN115171179B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image recognition, and particularly relates to a method, device, electronic device and storage medium for extracting pet nose print features. Background Art
[0002] With the development of the economy, more and more people raise pets, and the pet-related service industry is also developing rapidly. In order to provide better customized services for pets, it is necessary to identify pets accurately. At present, there are mainly two ways to identify pet identities. The first way is to implant a biochip in the pet's body and identify the pet's identity information by reading the information recorded in the biochip. However, this way has a relatively high cost and causes relatively great harm to the pet. The second way is to develop a pet identity recognition method based on pet nose prints because pet nose prints are unique and do not change with age. Specifically, first, the area where the pet's nose is located in the image is segmented through border detection technology, then features are extracted from the area where the nose is located to obtain the pet's nose print features; finally, the pet's identity information is identified based on the pet's nose print features. Therefore, this way does not require any operation on the pet, has a relatively low cost, and is harmless to the pet. The pet identity recognition method based on pet nose prints is gradually becoming popular.
[0003] However, when currently performing pet identity recognition, due to the different distributions of the facial features of different pets, using the same model for object detection for all pets results in poor compatibility and low segmentation accuracy of the nose, thus leading to low pet identity recognition accuracy. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, electronic device and storage medium for extracting pet nose print features, which call a matching model for nose segmentation for different pets to improve the accuracy of pet identity recognition.
[0005] In a first aspect, an embodiment of the present application provides a method for extracting pet nose print features, including:
[0006] Obtaining the species and feeding duration of the pet to be processed;
[0007] According to the species and the feeding duration, calling a target detection model corresponding to the pet to be processed;
[0008] Based on the target detection model, the species, and the feeding duration, performing object detection on the image of the pet to be processed to obtain the target area of the nose of the pet to be processed in the image of the pet to be processed;
[0009] Cropping out a target image corresponding to the target area from the image of the pet to be processed;
[0010] Extract features from the target image to obtain the nose print features of the pet to be processed.
[0011] In a second aspect, an embodiment of the present application provides a pet nose print feature extraction device, which includes an acquisition unit and a processing unit;
[0012] The acquisition unit is used to acquire the species and breeding duration of the pet to be processed;
[0013] The processing unit is used to call the target detection model corresponding to the pet to be processed according to the species and the breeding duration;
[0014] Based on the target detection model, the species, and the breeding duration, perform target detection on the pet image of the pet to be processed to obtain the target area of the nose of the pet to be processed in the pet image to be processed;
[0015] Crop out the target image corresponding to the target area from the pet image to be processed;
[0016] Extract features from the target image to obtain the nose print features of the pet to be processed.
[0017] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, the processor is connected to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program enables a computer to execute the method described in the first aspect.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer is operable to enable a computer to execute the method described in the first aspect.
[0020] Implementing the embodiments of the present application has the following beneficial effects:
[0021] It can be seen that in the embodiments of the present application, before performing object detection on the pet image to be processed, the type and feeding duration of the pet to be processed are first obtained, and then the object detection model corresponding to the type and feeding duration is called. In this way, for different pets, the object detection model matching them can be called to perform targeted object detection on the pet image to be processed, rather than using the same model for object detection on all pets. This improves the segmentation accuracy of the pet's nose, and thus the extracted nose print features can be more accurate, improving the accuracy of pet identity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 Schematic diagram of a pet nose print feature extraction system provided by an embodiment of the present application;
[0024] Figure 2 Schematic diagram of an application scenario of pet nose print feature extraction provided by an embodiment of the present application;
[0025] Figure 3 Schematic flowchart of a pet nose print feature extraction method provided by an embodiment of the present application;
[0026] Figure 4 Schematic flowchart of object detection based on a first object detection model provided by an embodiment of the present application;
[0027] Figure 5 Schematic diagram of constructing an association relationship provided by an embodiment of the present application;
[0028] Figure 6 Schematic flowchart of object detection based on a second object detection model provided by an embodiment of the present application;
[0029] Figure 7 Block diagram of the functional units of a pet nose print feature extraction device provided by an embodiment of the present application;
[0030] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0032] The terms "first", "second", "third", "fourth", etc. in the description and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0033] Referring to the embodiments herein means that the specific features, results, or characteristics described in conjunction with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0034] Refer to Figure 1 , Figure 1 which is a schematic diagram of a pet nose print feature extraction system provided by an embodiment of the present application. The pet nose print feature extraction system includes an image acquisition device 10, a pet nose print feature extraction device 20, and a cloud server 30. Among them, the image acquisition device 10, the pet nose print feature extraction device 20, and the cloud server 30 maintain a communication connection, and the image acquisition device 10 can be integrated on the pet nose print feature extraction device 20, and the cloud server 30 is an optional device.
[0035] Exemplarily, the image acquisition device 10 can take a picture of the pet to be processed to obtain the pet image to be processed; then, the pet nose print feature extraction device 20 can obtain the pet image to be processed from the image acquisition device 10. Further, the pet nose print feature extraction device 20 obtains the type and breeding duration of the pet to be processed. For example, if an input-output device (such as a keyboard) is provided on the pet nose print feature extraction device 20, the user can input the type and breeding duration of the pet to be processed to the pet nose print feature extraction device 20 through the input-output device; according to the type and the breeding duration, a target detection model corresponding to the pet to be processed is called; based on the target detection model, the type, and the breeding duration, target detection is performed on the pet image to be processed of the pet to be processed to obtain a target area of the nose of the pet to be processed in the pet image to be processed; a target image corresponding to the target area is cropped from the pet image to be processed; and feature extraction is performed on the target image to obtain the nose print feature of the pet to be processed.
[0036] Optionally, after the pet nose print feature extraction device 20 obtains the pet image to be processed, the type of the pet to be processed, and the breeding duration, it can send the pet image to be processed, the type of the pet to be processed, and the breeding duration to the cloud server 30, and the cloud server 30 extracts the nose print feature of the pet. And after the cloud server 30 extracts the nose print feature of the pet, it sends it to the pet nose print feature extraction device 20.
[0037] It can be seen that in the embodiment of the present application, before performing target detection on the pet image to be processed, the pet nose print feature extraction device 20 first obtains the type and breeding duration of the pet to be processed, and then calls a target detection model corresponding to the type and the breeding duration. In this way, for different pets, a target detection model matching them can be called to perform targeted target detection on the pet image to be processed, rather than using the same model for target detection of all pets, thereby improving the segmentation accuracy of the pet's nose, and further enabling the extracted nose print feature to be relatively accurate, and improving the accuracy of pet identity recognition.
[0038] Refer to Figure 2 , Figure 2 which is a schematic diagram of an application scenario for pet nose print feature extraction provided by the embodiment of the present application. Among them, Figure 2 the shown application scenario is the pet admission management in an amusement park. It should be noted that in actual applications, the pet nose print feature extraction of the present application can also have various other application scenarios, such as pet beauty appointment, medical appointment, and so on. The present application does not limit the application scenarios.
[0039] Such as Figure 2As shown, the owner of the pet to be processed (i.e., the user of this application) reserved an amusement park online, noted that the pet to be processed would be brought, and sent the nose print information of the pet to be processed to the server of the amusement park. In this way, the server of the amusement park locally stores the nose print information of the pet to be processed in advance. When the user enters the amusement park, the identity of the pet to be processed is verified based on the local nose print information to determine whether the pet to be processed has made an advance reservation.
[0040] Specifically, the above image acquisition device 10 is set at the entrance of the amusement park, so that the image of the pet to be processed can be acquired through the image acquisition device 10; then, the image of the pet to be processed is sent to the server of the amusement park, that is, the pet nose print feature extraction device 20. Finally, the pet nose print feature extraction device 20 executes the pet nose print feature extraction method of this application to obtain the nose print feature of the pet to be processed, and compares the nose print feature with the locally stored nose print information. If the comparison of the nose print feature is successful, it means that the above pet to be processed has made a reservation, and the pet to be processed can be released; if the comparison fails, it indicates that the verification fails and the pet to be processed is not released.
[0041] Refer to Figure 3 , Figure 3 is a schematic flowchart of a pet nose print feature extraction method provided by an embodiment of this application. This method is applied to the above pet nose print feature extraction device. This method includes but is not limited to the following steps:
[0042] 301: The pet nose print feature extraction device obtains the species and feeding duration of the pet to be processed.
[0043] Exemplarily, an input-output device is set on the pet nose print feature extraction device, and the user can input the species and feeding duration of the pet to be processed to the pet nose print feature extraction device through the input-output device. Of course, the pet nose print feature extraction device can also obtain the species and feeding duration in other ways. For example, the user inputs through voice. This application does not limit the way of obtaining the species and feeding duration.
[0044] 302: The pet nose print feature extraction device calls the target detection model corresponding to the pet to be processed according to the species and the feeding duration.
[0045] Exemplarily, the pet nose print feature extraction device determines the relative proportion of the nose of the pet to be processed in the image of the pet to be processed according to the species and the feeding duration. Exemplarily, the pet nose print feature extraction device obtains the growth rule of the pet to be processed according to the species of the pet to be processed. According to the growth rule of the pet to be processed and the feeding duration, the growth model of the pet to be processed is determined. That is, the pet to be processed grows according to the above growth rule, grows under normal conditions, and the growth model when it grows to the feeding duration. Obtain the proportion of each organ on the head of the pet to be processed relative to the entire head under this growth model; take the proportion of the nose relative to the entire head as the relative proportion of the nose of the pet to be processed in the image of the pet to be processed. Finally, according to the above relative proportion, the target detection model corresponding to the pet to be processed is called.
[0046] 303: The pet nose print feature extraction device performs target detection on the image of the pet to be processed based on the target detection model, the species, and the feeding duration, and obtains the target area of the nose of the pet to be processed in the image of the pet to be processed.
[0047] Optionally, when the relative proportion is less than the first threshold, that is, when the nose of the pet to be processed is a small target in the image of the pet to be processed, the target detection model corresponding to this relative proportion can be used to perform target detection on the image of the pet to be processed. In this application, the target detection model called when the relative proportion is less than the first threshold is called the first target detection model, that is, target detection is performed through the first target detection model.
[0048] Exemplarily, a plurality of query vectors corresponding to the above species are obtained, where each query vector is used to represent a facial feature corresponding to the pet of the above species. For example, a certain query vector represents the nose feature of the pet of the above species, and another query vector represents the eye feature of the pet of the above species. It should be noted that the above plurality of query vectors are obtained through pre-training. In this application, for each species of pet, query vectors corresponding to each species of pet are set, rather than using the same query vectors for all pets. In this way, when performing target detection on each species of pet, targeted detection can be performed, improving the efficiency and accuracy of target detection. Then, according to the species and the feeding duration, the association relationship between the facial organs of the pet to be processed is determined, where the association relationship mainly refers to the spatial position relationship between the facial organs.
[0049] Exemplarily, according to the type of the pet to be processed, determine the initial relative distances and initial relative directions between the respective facial organs of the pet to be processed, that is, the relative distances and relative directions between the respective facial organs when this kind of pet is newly born. Exemplarily, the initial relative distances and initial relative directions between the respective facial organs of various pets are pre-stored in the pet nose print feature extraction device. Then, according to the feeding duration, as well as the initial relative distances and initial relative directions, and the growth rules of the pets of the above type, determine the relative distances and relative directions between the respective facial organs of the pet to be processed at the moment corresponding to the feeding duration; determine the above-mentioned association relationship according to the relative distances and relative directions between the respective facial organs at the moment corresponding to the feeding duration.
[0050] Optionally, the above-mentioned association relationship can be represented by a three-dimensional matrix. Exemplarily, take the center of the nose of the pet to be processed as the coordinate origin of the three-dimensional space coordinate system, and obtain the relative directions and relative distances between the center points of the other respective facial organs and the center point of the nose. Based on the relative directions and relative distances between the center points of the other respective facial organs and the center point of the nose, determine the three-dimensional space coordinates of the center points of the other respective facial organs in the above three-dimensional space coordinate system. Finally, construct the above three-dimensional matrix based on the three-dimensional space coordinates of the other respective facial organs.
[0051] Specifically, obtain the maximum value of the X-axis, the maximum value of the Y-axis, and the maximum value of the Z-axis in the three-dimensional space coordinates of the center points of the other respective facial organs. As Figure 5 shown, the maximum values of the X-axis, the Y-axis, and the Z-axis obtained are x1, y1, and z1 respectively. Then, take the maximum value of the X-axis as the length of the three-dimensional space coordinate as the length of the above three-dimensional matrix, the maximum value of the Y-axis as the width of the above three-dimensional matrix, and the maximum value of the Z-axis as the height of the above three-dimensional matrix, to obtain a three-dimensional matrix with lengths, widths, and heights of x1, y1, and z1 respectively. Finally, divide the three-dimensional matrix by taking values at intervals of 1, obtain the spatial coordinates of each element in the three-dimensional matrix, and take the distance between the spatial coordinates of each element and the coordinate origin as the value of each element, so as to obtain the three-dimensional matrix used to represent the above-mentioned association relationship. For example, Figure 5 the element in the upper right corner of the three-dimensional matrix shown (that is, Figure 5 the gray part in
[0052] Finally, according to the association relationship, multiple query vectors, and the target detection model, perform target detection on the image of the pet to be processed of the pet to be processed, and obtain the target area of the nose of the pet to be processed in the image of the pet to be processed.
[0053] Specifically, as Figure 4As shown, the first object detection model includes an encoding network, a feature extraction network, an encoder, and a decoder. Among them, the encoding network can be an embedding layer, the feature extraction network can be a Feature Pyramid Networks (FPN), and both the encoder and the decoder can be the encoder and decoder of the transform structure.
[0054] Exemplarily, based on Figure 4 the object detection model, according to the association relationship, feature extraction is performed on the pet image to be processed, and a plurality of first feature vectors are obtained. Specifically, feature extraction is performed on the pet image to be processed to obtain a first feature map, that is, the pet image to be processed is input into the feature extraction network for feature extraction to obtain a first feature map; the association relationship is encoded to obtain a second feature map. Exemplarily, the association relationship is input into the encoding network for encoding, and a second feature map can be obtained, that is, the above three-dimensional matrix is mapped so that the dimension of the obtained second feature map is the same as the size (i.e., length and width) of the first feature map, so as to facilitate subsequent fusion with the first feature map.
[0055] Furthermore, the first feature map and the second feature map are fused to obtain a third feature map. It should be understood that both the first feature map and the second feature map are three-dimensional matrices and have the same size, so the first feature map and the second feature map can be concatenated vertically (height) to obtain a third feature map. For example, if the height of the first feature map is h1 and the height of the second feature map is h2, then the height of the third feature map obtained by concatenating the first feature map and the second feature map is (h1 + h2).
[0056] Further, the third feature map is tiled to obtain a plurality of first feature vectors. Exemplarily, first, the third feature map is tiled by height (layer by layer) to obtain a plurality of two-dimensional matrices. For example, if the height of the third feature map is (h1 + h2), then (h1 + h2) two-dimensional matrices can be tiled, and the size of each two-dimensional matrix is the same as that of the third feature map. Then, for each two-dimensional matrix, it is tiled by row or by column (in this application, tiling by row is taken as an example for illustration) to obtain a plurality of one-dimensional sequences, where the number of elements in each one-dimensional sequence is the number of columns of the two-dimensional matrix, and the number of the plurality of one-dimensional sequences is the number of rows of each two-dimensional matrix. For example, if the size of each two-dimensional matrix is w * l, then after tiling by row, w one-dimensional sequences can be obtained, and the number of elements in each one-dimensional sequence is l. Finally, the plurality of one-dimensional sequences corresponding to each two-dimensional matrix are concatenated to obtain a feature vector corresponding to each two-dimensional matrix, and the number of elements in the feature vector is w * l. Finally, the feature vector corresponding to each two-dimensional matrix is used as a first feature vector to obtain a plurality of first feature vectors. Therefore, for a third feature map with a size of w * l * (h1 + h2), (h1 + h2) first feature vectors with the number of elements being w * l can be tiled.
[0057] Further, the plurality of first feature vectors are input into an encoder for encoding to obtain a plurality of second feature vectors, where the encoding process can refer to the encoding process of the transform encoder and will not be described again; then, the plurality of query vectors and the plurality of second feature vectors are input into a decoder for decoding to obtain a plurality of third feature vectors, where the decoding process can refer to the decoding process of the transform decoder and will not be described again. Finally, object detection is performed based on each third feature vector to obtain a candidate box and a classification category corresponding to each third feature vector, that is, based on each third feature vector, class prediction and candidate box prediction are performed, and a candidate box and a classification category corresponding to each third feature vector can be obtained.
[0058] Finally, based on the candidate box and the classification category corresponding to each third feature vector, a target candidate box is determined. That is, the classification category belonging to the nose is determined from the classification categories corresponding to each third feature vector, and then the candidate box corresponding to the classification category is used as the target candidate box. And the area framed by the target candidate box in the to-be-processed pet image is used as the target area of the nose of the to-be-processed pet in the to-be-processed pet image.
[0059] Optionally, when the relative proportion is greater than or equal to the first threshold, that is, when the nose of the pet to be processed is a small target in the pet image to be processed, the pet image to be processed can be subjected to object detection by an object detection model corresponding to the relative proportion. In this application, the object detection model called by the relative proportion greater than or equal to the first threshold is called the second object detection model, that is, object detection is performed through the second object detection model.
[0060] Refer to Figure 6 , Figure 6 shows a schematic diagram of the second object detection model. The second object detection model includes a feature extraction network, a mapping layer, and a convolutional layer. Among them, the feature extraction network and the mapping layer are similar to those Figure 4 shown and will not be described again. It should be noted that, different from the first object detection model, before using the second object detection model, multiple reference frames and multiple reference features corresponding to each type of pet need to be trained. Among them, the multiple reference frames can roughly frame the objects included in the pet, and the multiple reference features correspond to the multiple reference frames and are used to characterize the features that the objects framed by each reference frame probably contain.
[0061] Exemplarily, feature extraction is performed on the pet image to be processed to obtain a first feature map, that is, it is input into the feature extraction network for feature extraction to obtain a first feature map. Then, multiple reference frames and multiple reference features corresponding to the above types are obtained. It should be noted that in this application, reference frames corresponding to each type of pet are trained, rather than using the same reference frame for all pets, so as to achieve targeted object detection for each type of pet and improve the accuracy of object detection. Further, candidate images corresponding to each reference frame are cropped from the first feature map.
[0062] It should be understood that the first feature map is a three-dimensional matrix stitched together by the feature maps extracted from each channel in the feature extraction network. Therefore, when using the reference frame to crop the image from the first feature map, in essence, the image is cropped from the feature map corresponding to each channel, and the images cropped from multiple channels are stitched together to obtain a candidate image corresponding to each candidate box. Therefore, the candidate image corresponding to each candidate box is also a three-dimensional matrix.
[0063] Further, the candidate image corresponding to each candidate box is input into the mapping layer for mapping to obtain a fourth feature map corresponding to each reference frame, that is, the candidate image corresponding to each candidate box is input into the mapping layer to perform ROI align operation, and the candidate image corresponding to each candidate box is mapped into a fifth feature map to meet the data requirements of the convolutional layer and input into the convolutional layer for calculation.
[0064] Further, perform convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain a fifth feature map corresponding to each reference box. Exemplarily, as Figure 6 shown, map the reference feature corresponding to each reference box to obtain a two-dimensional matrix corresponding to each reference box, where the number of columns of the two-dimensional matrix is the same as the height of the fourth feature map corresponding to each reference box to ensure that the dimension of the fifth feature vector is the same as the height of the fourth feature map, and the number of rows of the two-dimensional matrix depends on the number of convolution kernels set in the convolutional layer. Then, tile the two-dimensional matrix corresponding to each reference box to obtain multiple fifth feature vectors, where the dimension of each fifth feature vector is the same as the number of columns of the above two-dimensional matrix; then, use the multiple fifth feature vectors as the convolution parameters of multiple convolutional kernels, so that each fifth feature vector in the multiple fifth feature vectors can be used to perform convolution processing with the fourth feature map corresponding to each reference box respectively to obtain a sub-feature map corresponding to each fifth feature vector; splice the multiple sub-feature maps corresponding to the multiple fifth feature vectors to obtain a fifth feature map corresponding to each reference box. Therefore, the fifth feature map corresponding to each reference box is also a three-dimensional matrix.
[0065] Finally, as Figure 6 shown, tile the fifth feature map corresponding to each reference box to obtain a fourth feature vector corresponding to each reference box, where the process of tiling the fifth feature map is similar to the process of tiling the third feature map above, and a two-dimensional matrix corresponding to the fifth feature map can be tiled, and then the two-dimensional matrix is tiled into multiple one-dimensional sequences. Different from tiling the third feature map, after tiling the fifth feature map into multiple one-dimensional sequences, the multiple one-dimensional sequences are combined to obtain a long one-dimensional sequence, and this long one-dimensional sequence is used as the fourth feature vector corresponding to each reference box.
[0066] Similarly, perform object detection according to the fourth feature vector corresponding to each reference box to obtain a candidate box and a classification category corresponding to each reference box; determine a target candidate box according to the candidate box and the classification category corresponding to each reference box; use the area selected by the target candidate box in the to-be-processed pet image as the target area of the nose of the to-be-processed pet in the to-be-processed pet image.
[0067] 304: The pet nose print feature extraction device extracts a target image corresponding to the target area from the to-be-processed pet image.
[0068] Exemplarily, perform cropping on the to-be-processed pet image to extract a target image corresponding to the target area from the to-be-processed pet image.
[0069] 305: The pet nose print feature extraction device extracts features from the target image to obtain the nose print feature of the to-be-processed pet.
[0070] Exemplarily, feature extraction can be performed on the target image to obtain nose print features. For example, a trained neural network can be used to perform feature extraction on the target image to obtain nose print features.
[0071] It can be seen that in the embodiments of the present application, before performing object detection on the pet image to be processed, the type and breeding duration of the pet to be processed are first obtained, and then the object detection model corresponding to the type and breeding duration is called. In this way, for different pets, the object detection model that matches them can be called to perform targeted object detection on the pet image to be processed, rather than using the same model for object detection for all pets, thereby improving the segmentation accuracy of the pet's nose, and further enabling the extracted nose print features to be relatively accurate, improving the accuracy of pet identity recognition.
[0072] Optionally, after obtaining the nose print features of the pet to be processed, pet identity processing can be performed based on the nose print features. Exemplarily, when performing pet identity verification, the nose print features can be matched with each template pet feature in the template library, and the maximum matching value can be obtained. When the maximum matching value is greater than the threshold, it is determined that the identity verification of the pet to be processed passes; otherwise, it is determined that the verification fails. Exemplarily, when performing pet identity recognition, the nose print features can also be matched with each template pet feature in the template library, and the maximum matching value can be obtained. The identity information corresponding to the template pet feature corresponding to the maximum matching value is used as the identity information of the pet to be processed. When performing pet identity entry, the obtained pet nose print features can be entered into the database and bound to the identity information of the pet to be processed, and a new pet identity is entered into the database.
[0073] Refer to Figure 7 , Figure 7 The functional unit composition block diagram of a pet nose print feature extraction device provided by the embodiments of the present application. The pet nose print feature extraction device 700 includes: an acquisition unit 701 and a processing unit 702.
[0074] The acquisition unit 701 is configured to acquire the type and breeding duration of the pet to be processed;
[0075] The processing unit 702 is configured to call the object detection model corresponding to the pet to be processed according to the type and the breeding duration;
[0076] Based on the object detection model, the type, and the breeding duration, perform object detection on the pet image to be processed of the pet to be processed, and obtain the target area of the nose of the pet to be processed in the pet image to be processed;
[0077] Extract a target image corresponding to the target area from the pet image to be processed;
[0078] Extract features from the target image to obtain the nose print feature of the pet to be processed.
[0079] In some possible implementation manners, in terms of calling a target detection model corresponding to the pet to be processed according to the species and the breeding duration, the processing unit 702 is specifically configured to:
[0080] Determine the relative proportion of the nose of the pet to be processed in the pet image to be processed according to the species and the breeding duration;
[0081] Call a target detection model corresponding to the pet to be processed according to the relative proportion.
[0082] When the relative proportion is less than a first threshold, in terms of performing target detection on the pet image to be processed of the pet to be processed based on the target detection model, the species, and the breeding duration to obtain a target area of the nose of the pet to be processed in the pet image to be processed, the processing unit 702 is specifically configured to:
[0083] Obtain a plurality of query vectors corresponding to the species, where each query vector is used to represent a pet facial feature corresponding to the species;
[0084] Determine the association relationship between the facial organs of the pet to be processed according to the species and the breeding duration;
[0085] Perform target detection on the pet image to be processed of the pet to be processed according to the association relationship, the plurality of query vectors, and the target detection model to obtain a target area of the nose of the pet to be processed in the pet image to be processed.
[0086] In some possible implementation manners, in terms of performing target detection on the pet image to be processed of the pet to be processed according to the association relationship, the plurality of query vectors, and the target detection model to obtain a target area of the nose of the pet to be processed in the pet image to be processed, the processing unit 702 is specifically configured to:
[0087] Extract features from the pet image to be processed according to the association relationship to obtain a plurality of first feature vectors;
[0088] Input the plurality of first feature vectors into an encoder for encoding to obtain a plurality of second feature vectors;
[0089] Input the plurality of query vectors and the plurality of second feature vectors into a decoder for decoding to obtain a plurality of third feature vectors;
[0090] Perform object detection based on each third eigenvector to obtain candidate bounding boxes and classification categories corresponding to each third eigenvector;
[0091] Determine target candidate bounding boxes based on the candidate bounding boxes and classification categories corresponding to each third eigenvector;
[0092] Take the area outlined by the target candidate bounding boxes in the pet image to be processed as the target area of the nose of the pet to be processed in the pet image to be processed.
[0093] In some possible implementation manners, in terms of extracting features from the pet image to be processed according to the association relationship to obtain a plurality of first eigenvectors, the processing unit 702 is specifically configured to:
[0094] Extract features from the pet image to be processed to obtain a first feature map;
[0095] Encode the association relationship to obtain a second feature map;
[0096] Fuse the first feature map and the second feature map to obtain a third feature map;
[0097] Tile the third feature map to obtain the plurality of first eigenvectors.
[0098] In some possible implementation manners, when the relative proportion is greater than or equal to the first threshold, in terms of performing object detection on the pet image to be processed based on the object detection model, the species, and the breeding duration to obtain the target area of the nose of the pet to be processed in the pet image to be processed, the processing unit 702 is specifically configured to:
[0099] Extract features from the pet image to be processed to obtain a first feature map;
[0100] Obtain a plurality of reference bounding boxes and a plurality of reference features corresponding to the species, where the plurality of reference bounding boxes and the plurality of reference features are in one-to-one correspondence;
[0101] Extract candidate images corresponding to each reference bounding box from the first feature map;
[0102] Map the candidate image corresponding to each reference bounding box to obtain a fourth feature map corresponding to each reference bounding box;
[0103] Perform convolution processing on the fourth feature map corresponding to each reference bounding box and the reference feature corresponding to each reference bounding box to obtain a fifth feature map corresponding to each reference bounding box;
[0104] Tile the fifth feature map corresponding to each reference box to obtain a fourth feature vector corresponding to each reference box;
[0105] Perform object detection based on the fourth feature vector corresponding to each reference box to obtain a candidate box and a classification category corresponding to each reference box;
[0106] Determine a target candidate box according to the candidate box and the classification category corresponding to each reference box;
[0107] Use the area outlined by the target candidate box in the to-be-processed pet image as the target area of the nose of the to-be-processed pet in the to-be-processed pet image.
[0108] In some possible implementation manners, in terms of performing convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain the fifth feature map corresponding to each reference box, the processing unit 702 is specifically configured to:
[0109] Map the reference feature corresponding to each reference box to obtain a two-dimensional matrix corresponding to each reference box, where the number of columns of the two-dimensional matrix is the same as the height of the fourth feature map corresponding to each reference box;
[0110] Tile the two-dimensional matrix corresponding to each reference box to obtain a plurality of fifth feature vectors, where the dimension of each feature vector is the same as the number of columns of the two-dimensional matrix;
[0111] Perform convolution processing on each fifth feature vector and the fourth feature map corresponding to each reference box to obtain a sub-feature map corresponding to each fifth feature vector;
[0112] Concatenate the plurality of sub-feature maps corresponding to the plurality of fifth feature vectors to obtain the fifth feature map corresponding to each reference box.
[0113] Refer to Figure 8 , Figure 8 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 800 includes a transceiver 801, a processor 802, and a memory 803. They are connected through a bus 804. The memory 803 is used to store computer programs and data, and can transmit the data stored in the memory 803 to the processor 802.
[0114] The processor 802 is configured to read the computer program in the memory 803 and perform the following operations:
[0115] Control the transceiver 801 to obtain the type and feeding duration of the to-be-processed pet;
[0116] Call a target detection model corresponding to the pet to be processed according to the species and the feeding duration;
[0117] Based on the target detection model, the species, and the feeding duration, perform target detection on the image of the pet to be processed to obtain the target area of the nose of the pet to be processed in the image of the pet to be processed;
[0118] Crop a target image corresponding to the target area from the image of the pet to be processed;
[0119] Extract features from the target image to obtain the nose print feature of the pet to be processed.
[0120] In some possible implementation manners, in terms of calling a target detection model corresponding to the pet to be processed according to the species and the feeding duration, the processor 802 is specifically configured to perform the following operations:
[0121] Determine the relative proportion of the nose of the pet to be processed in the image of the pet to be processed according to the species and the feeding duration;
[0122] Call a target detection model corresponding to the pet to be processed according to the relative proportion.
[0123] When the relative proportion is less than a first threshold, in terms of performing target detection on the image of the pet to be processed based on the target detection model, the species, and the feeding duration to obtain the target area of the nose of the pet to be processed in the image of the pet to be processed, the processor 802 is specifically configured to perform the following operations:
[0124] Obtain a plurality of query vectors corresponding to the species, where each query vector is used to represent a pet facial feature corresponding to the species;
[0125] Determine the association relationship between the facial organs of the pet to be processed according to the species and the feeding duration;
[0126] Perform target detection on the image of the pet to be processed according to the association relationship, the plurality of query vectors, and the target detection model to obtain the target area of the nose of the pet to be processed in the image of the pet to be processed.
[0127] In some possible implementation manners, in terms of performing target detection on the image of the pet to be processed according to the association relationship, the plurality of query vectors, and the target detection model to obtain the target area of the nose of the pet to be processed in the image of the pet to be processed, the processor 802 is specifically configured to perform the following operations:
[0128] According to the association relationship, perform feature extraction on the to-be-processed pet image to obtain a plurality of first feature vectors;
[0129] Input the plurality of first feature vectors into an encoder for encoding to obtain a plurality of second feature vectors;
[0130] Input the plurality of query vectors and the plurality of second feature vectors into a decoder for decoding to obtain a plurality of third feature vectors;
[0131] Perform object detection according to each third feature vector to obtain a candidate box and a classification category corresponding to each third feature vector;
[0132] Determine a target candidate box according to the candidate box and the classification category corresponding to each third feature vector;
[0133] Take the area outlined by the target candidate box in the to-be-processed pet image as the target area of the nose of the to-be-processed pet in the to-be-processed pet image.
[0134] In some possible implementation manners, in terms of performing feature extraction on the to-be-processed pet image according to the association relationship to obtain a plurality of first feature vectors, the processor 802 is specifically configured to perform the following operations:
[0135] Perform feature extraction on the to-be-processed pet image to obtain a first feature map;
[0136] Encode the association relationship to obtain a second feature map;
[0137] Fuse the first feature map and the second feature map to obtain a third feature map;
[0138] Tile the third feature map to obtain the plurality of first feature vectors.
[0139] In some possible implementation manners, when the relative proportion is greater than or equal to a first threshold, in terms of performing object detection on the to-be-processed pet image of the to-be-processed pet based on the object detection model, the species, and the feeding duration to obtain the target area of the nose of the to-be-processed pet in the to-be-processed pet image, the processor 802 is specifically configured to perform the following operations:
[0140] Perform feature extraction on the to-be-processed pet image to obtain a first feature map;
[0141] Obtain a plurality of reference boxes and a plurality of reference features corresponding to the species, wherein the plurality of reference boxes and the plurality of reference features correspond one by one;
[0142] Extract candidate images corresponding to each reference box from the first feature map;
[0143] Map the candidate images corresponding to each reference box to obtain a fourth feature map corresponding to each reference box;
[0144] Perform convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain a fifth feature map corresponding to each reference box;
[0145] Tile the fifth feature map corresponding to each reference box to obtain a fourth feature vector corresponding to each reference box;
[0146] Perform object detection based on the fourth feature vector corresponding to each reference box to obtain a candidate box and a classification category corresponding to each reference box;
[0147] Determine a target candidate box according to the candidate box and the classification category corresponding to each reference box;
[0148] Use the area framed by the target candidate box in the pet image to be processed as the target area of the nose of the pet to be processed in the pet image to be processed.
[0149] In some possible implementation manners, in terms of performing convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain a fifth feature map corresponding to each reference box, the processor 802 is specifically configured to perform the following operations:
[0150] Map the reference feature corresponding to each reference box to obtain a two-dimensional matrix corresponding to each reference box, where the number of columns of the two-dimensional matrix is the same as the height of the fourth feature map corresponding to each reference box;
[0151] Tile the two-dimensional matrix corresponding to each reference box to obtain a plurality of fifth feature vectors, where the dimension of each feature vector is the same as the number of columns of the two-dimensional matrix;
[0152] Perform convolution processing on each fifth feature vector and the fourth feature map corresponding to each reference box to obtain a sub-feature map corresponding to each fifth feature vector;
[0153] Stitch the plurality of sub-feature maps corresponding to the plurality of fifth feature vectors to obtain a fifth feature map corresponding to each reference box.
[0154] Specifically, the transceiver 801 may be Figure 7 the acquisition unit 701 of the pet nose print feature extraction device 700 in the above embodiment, and the processor 802 may be Figure 7 the processing unit 702 of the pet nose print feature extraction device 700 in the above embodiment.
[0155] It should be understood that the electronic devices in this application may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, handheld computers, laptop computers, mobile Internet devices MID (Mobile Internet Devices, abbreviated as MID), or wearable devices, etc. The above-mentioned electronic devices are only examples, not exhaustive, including but not limited to the above-mentioned electronic devices. In practical applications, the above-mentioned electronic devices may also include: intelligent vehicle terminals, computer devices, and so on.
[0156] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any one of the pet nose print feature extraction methods described in the above method embodiments.
[0157] The embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the pet nose print feature extraction methods described in the above method embodiments.
[0158] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0159] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0160] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0161] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software program module.
[0163] If the above integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0164] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memories (abbreviation: ROM), random access memories (abbreviation: RAM), magnetic disks, or optical discs, etc.
[0165] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for extracting pet nose print features, characterized in that Including: Obtain the type and raising duration of the pet to be processed; According to the type and the raising duration, call the target detection model corresponding to the pet to be processed; Based on the target detection model, the type, and the raising duration, perform target detection on the pet image of the pet to be processed, and obtain the target area of the nose of the pet to be processed in the pet image; Extract the target image corresponding to the target area from the pet image; Extract features from the target image to obtain the nose print feature of the pet to be processed.
2. The method according to claim 1, characterized in that The step of calling the target detection model corresponding to the pet to be processed according to the type and the raising duration includes: According to the type and the raising duration, determine the relative proportion of the nose of the pet to be processed in the pet image; According to the relative proportion, call the target detection model corresponding to the pet to be processed.
3. The method according to claim 2, wherein When the relative proportion is less than the first threshold, the step of performing target detection on the pet image of the pet to be processed based on the target detection model, the type, and the raising duration, and obtaining the target area of the nose of the pet to be processed in the pet image includes: Obtain a plurality of query vectors corresponding to the type, where each query vector is used to represent a pet facial feature corresponding to the type; According to the type and the raising duration, determine the association relationship between the facial organs of the pet to be processed; According to the association relationship, the plurality of query vectors, and the target detection model, perform target detection on the pet image of the pet to be processed, and obtain the target area of the nose of the pet to be processed in the pet image.
4. The method according to claim 3, characterized in that The step of performing target detection on the pet image of the pet to be processed according to the association relationship, the plurality of query vectors, and the target detection model, and obtaining the target area of the nose of the pet to be processed in the pet image includes: According to the association relationship, extract features from the pet image to obtain a plurality of first feature vectors; Input the plurality of first feature vectors into an encoder for encoding to obtain a plurality of second feature vectors; Input the plurality of query vectors and the plurality of second feature vectors into a decoder for decoding to obtain a plurality of third feature vectors; Perform target detection according to each third feature vector to obtain a candidate box and a classification category corresponding to each third feature vector; Determine the target candidate box according to the candidate box and the classification category corresponding to each third feature vector; Use the area outlined by the target candidate box in the pet image as the target area of the nose of the pet to be processed in the pet image.
5. The method according to claim 4, wherein The step of extracting features from the pet image according to the association relationship to obtain a plurality of first feature vectors includes: Extract features from the pet image to obtain a first feature map; Encode the association relationship to obtain a second feature map; Fuse the first feature map and the second feature map to obtain a third feature map; Tile the third feature map to obtain the multiple first feature vectors.
6. The method according to claim 2, characterized in that, When the relative proportion is greater than or equal to the first threshold, performing object detection on the to-be-processed pet image of the to-be-processed pet based on the object detection model, the category, and the feeding duration, to obtain a target region of the nose of the to-be-processed pet in the to-be-processed pet image, including: Performing feature extraction on the to-be-processed pet image to obtain a first feature map; Obtaining a plurality of reference boxes and a plurality of reference features corresponding to the category, wherein the plurality of reference boxes and the plurality of reference features are in one-to-one correspondence; Cropping candidate images corresponding to each reference box from the first feature map; Performing mapping on the candidate image corresponding to each reference box to obtain a fourth feature map corresponding to each reference box; Performing convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain a fifth feature map corresponding to each reference box; Tiling the fifth feature map corresponding to each reference box to obtain a fourth feature vector corresponding to each reference box; Performing object detection according to the fourth feature vector corresponding to each reference box to obtain a candidate box and a classification category corresponding to each reference box; Determining a target candidate box according to the candidate box and the classification category corresponding to each reference box; Taking the area framed by the target candidate box in the to-be-processed pet image as the target region of the nose of the to-be-processed pet in the to-be-processed pet image.
7. The method according to claim 6, wherein The performing convolution processing on the fourth feature map corresponding to each reference box and the reference feature corresponding to each reference box to obtain a fifth feature map corresponding to each reference box includes: Mapping the reference feature corresponding to each reference box to obtain a two-dimensional matrix corresponding to each reference box, wherein the number of columns of the two-dimensional matrix is the same as the height of the fourth feature map corresponding to each reference box; Tiling the two-dimensional matrix corresponding to each reference box to obtain a plurality of fifth feature vectors, wherein the dimension of each feature vector is the same as the number of columns of the two-dimensional matrix; Performing convolution processing on each fifth feature vector and the fourth feature map corresponding to each reference box to obtain a sub-feature map corresponding to each fifth feature vector; Concatenating the plurality of sub-feature maps corresponding to the plurality of fifth feature vectors to obtain a fifth feature map corresponding to each reference box.
8. A pet nose print feature extraction device, characterized in that, The device includes an acquisition unit and a processing unit; The acquisition unit is configured to acquire the category and the feeding duration of the to-be-processed pet; The processing unit is configured to call an object detection model corresponding to the to-be-processed pet according to the category and the feeding duration; Performing object detection on the to-be-processed pet image of the to-be-processed pet based on the object detection model, the category, and the feeding duration, to obtain a target region of the nose of the to-be-processed pet in the to-be-processed pet image; Cropping a target image corresponding to the target region from the to-be-processed pet image; Performing feature extraction on the target image to obtain the nose pattern feature of the to-be-processed pet.
9. An electronic device, characterized in that, including: A processor and a memory, the processor being connected to the memory, the memory being configured to store a computer program, and the processor being configured to execute the computer program stored in the memory so that the electronic device performs the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Nose print feature extraction method and device and nonvolatile storage medium
CN112784742A
Auxiliary shooting method and device for pet
CN113132632A