Method and electronic device for photographing an object for pet identification

By detecting and focusing on the pet's nose region within images and adjusting focus dynamically, the method improves the quality and suitability of pet nose print images for AI-based identification, addressing the challenges of low recognition rates and image clarity in pet identification systems.

JP7710531B2Active Publication Date: 2025-07-18PETNOW
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023568646
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2022-06-27
Publication Date
2025-07-18
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The challenge in pet identification systems is the difficulty in obtaining high-quality images of pet nose prints due to factors like shooting angle, focus, distance, and environment, leading to low recognition rates, especially with insufficient data for machine learning, and the inability of pets to stay still for clear image capture.

Method used

A method and electronic device that detect and focus on the pet's nose region within a captured image, assess image quality, and adjust focus accordingly to ensure clarity, discarding unsuitable images and providing feedback for improved image capture.

Benefits of technology

This approach enhances the quality of pet nose print images suitable for artificial intelligence-based identification, reducing computational complexity and improving recognition rates by ensuring clear, focused images are captured and stored for learning or identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710531000004
    Figure 0007710531000004
  • Figure 0007710531000005
    Figure 0007710531000005
  • Figure 0007710531000006
    Figure 0007710531000006
Patent Text Reader

Abstract

The present invention provides a method and electronic device capable of improving image quality of an object for identifying a pet. The method for photographing an object for identifying a pet according to the present invention includes the steps of acquiring an image including the pet, detecting an object for identifying the pet from the image, and photographing a next image with the focus set on the position of the detected object. According to the present invention, an object for identifying a pet is detected from a photographed image and the focus is continuously set on the position of the object, thereby obtaining an image for identifying an object with higher quality, and thus storing an image for identifying an object of a pet that can be used for learning or identification based on artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and an electronic device for photographing an object for pet identification, and more particularly, to a method and an electronic device for obtaining an image of an object for pet identification suitable for learning or identification based on artificial intelligence.

Background Art

[0002] In modern society, the demand for pets that can be emotionally relied on while living with people is increasing. As a result, the need to database and manage information about various pets for health management and the like of pets is increasing. To manage pets, identification information of pets such as human fingerprints is required, and objects that can be used according to the pets can be defined respectively. For example, in the case of dogs, since the nose prints (the shape of the wrinkles on the nose) are different for each dog, the nose print can be used as identification information for each dog.

[0003] As shown in FIG. 1(a), the method of registering a nose print is performed by photographing a face including the pet's nose (S110) and storing and registering an image including the nose print in a database (S120), similar to registering a human fingerprint or face. Also, as shown in FIG. 1(b), the method of querying a nose print can be performed by photographing the pet's nose print (S130), searching for a nose print that matches the photographed nose print and information related thereto (S140), and outputting information that matches the photographed nose print (S150). As shown in FIG. 1, each pet can be identified by the process of registering and querying the pet's nose print, and the information of the pet can be managed. The nose print information of the pet is stored in a database and can be used as data for learning or identification based on AI.

[0004] However, there are some problems when obtaining and storing the nose print of a pet.

[0005] First, photos can be difficult to recognize depending on factors such as the shooting angle, focus, distance, size, and environment. There have been attempts to apply human face recognition technology to nose print recognition. However, while sufficient data has been accumulated for human face information, insufficient data has been secured for pet nose print information, resulting in a problem of low recognition rates. Specifically, for recognition based on AI, learning data processed in a form that allows the machine to learn is required. However, since insufficient data has been accumulated for pet nose prints, it is difficult to recognize nose prints.

[0006] In addition, for pet nose print recognition, an image with distinct nose wrinkles is required. However, unlike humans, pets cannot perform actions that would stop their movement for a while, so it is not easy to obtain an image of distinct nose wrinkles. For example, dogs keep moving their faces or licking their tongues, making it very difficult to obtain a nose print image of the desired quality. For example, an image with distinct nose wrinkles is required for nose print recognition, but in actual captured images, the nose wrinkles are often not clearly captured due to shaking or the like. To solve such problems, a method of taking a picture with the dog's nose forcibly fixed has been considered, but it is evaluated as inappropriate because it involves performing a forced action on the pet.

Summary of the Invention

Problems to be Solved by the Invention

[0007] The present invention provides a method and an electronic device capable of improving the image quality of an object for pet identification.

[0008] The problems to be solved by the present invention are not limited to those described above, and other problems not described above will be clearly understood by those skilled in the art from the following description.

Means for Solving the Problems

[0009] A method for photographing an object for pet identification according to the present invention includes the steps of obtaining an image including the pet, detecting an object for pet identification from the image, and photographing a next image with focus set at the position of the detected object.

[0010] According to the present invention, the step of obtaining an image including the pet may include the step of photographing the pet with focus set at the position of the object detected from an image photographed in a previous frame to generate the image.

[0011] According to the present invention, the step of detecting an object for pet identification from the image may include the step of setting a first feature region for determining the species of the pet from the image, and the step of setting a second feature region including an object for pet identification according to the species of the pet within the first feature region.

[0012] According to the present invention, the step of setting the second feature region may include the step of determining whether an image of the object for pet identification is suitable for learning or identification based on artificial intelligence.

[0013] According to the present invention, the step of determining whether an image of the object is suitable for learning or identification based on artificial intelligence may include the step of determining whether the quality of the image of the object satisfies reference conditions, the step of transmitting the image of the object to a server when the quality satisfies the reference conditions, and the step of discarding the image of the object and performing photographing for the next image when the quality does not satisfy the reference conditions.

[0014] A method for photographing an object for pet identification according to the present invention includes the steps of acquiring an image including the pet, detecting an object for pet identification from the image, setting a focus on the position of the object, and outputting the image with a graphic element indicating an object detection state overlaid on the position of the object where the focus is set.

[0015] According to the present invention, the step of outputting the image may include determining whether the quality of the image of the object satisfies a reference condition, overlaying a first graphic element indicating a good quality state on the object when the quality of the image of the object satisfies the reference condition, and overlaying a second graphic element indicating a poor quality state on the object when the quality of the object does not satisfy the reference condition.

[0016] According to the present invention, the step of outputting the image may include outputting score information indicating a photographing quality state of the object in the image.

[0017] According to the present invention, in the method for photographing an object for pet identification, the step of outputting the image may further include providing feedback to the user so that a good object image is photographed according to the quality state of the image of the object.

[0018] According to the present invention, the step of acquiring an image including the pet may include photographing the pet to generate the image with the focus set on the position of the object detected from the image photographed in the previous frame.

[0019] An electronic device for photographing an object for pet identification according to the present invention includes a camera that generates an original image including the pet, and a processor that detects an object for pet identification from the image and controls the camera to photograph the next image with the focus set at the position of the detected object.

[0020] According to the present invention, in the step where the processor acquires an image including a pet, the camera can be controlled to photograph the pet with the focus set at the position of the object detected from the image captured in the previous frame to generate the image.

[0021] According to the present invention, the processor can set a first feature region for determining the species of the pet from the image, and set a second feature region including an object for pet identification within the first feature region.

[0022] According to the present invention, the processor can determine whether the image of the object for pet identification is suitable for learning or identification based on artificial intelligence.

[0023] According to the present invention, the processor determines whether the quality of the image of the object satisfies the reference conditions. When the quality satisfies the reference conditions, the image of the object is transmitted to the server. When the quality does not satisfy the reference conditions, the image of the object is discarded, and the camera can be controlled to perform photographing for the next image.

[0024] The electronic device for photographing an object for pet identification according to the present invention can further include a display that outputs an image of the pet during photographing.

[0025] According to the present invention, the processor can output to the display an image in which a graphic element indicating the object detection state is overlaid at the position of the detected object.

[0026] According to the present invention, the processor determines whether the quality of the image of the object satisfies the reference conditions. When the quality of the image of the object satisfies the reference conditions, a first graphic element indicating a good quality state is overlaid on the object. When the quality of the object does not satisfy the reference conditions, a second graphic element indicating a poor quality state can be overlaid on the object.

[0027] According to the present invention, the processor can output score information indicating the shooting quality state of the object in the image via the display.

[0028] According to the present invention, the processor can provide feedback to the user so that a good object image is captured according to the quality state of the image of the object.

Advantages of the Invention

[0029] According to the present invention, an object for identifying a pet is detected from a captured image, and by continuously positioning the focus at the position of the object, a clearer quality object identification image is obtained, thereby accumulating an object identification image (nose print image) of the pet that can be used for learning or identification based on artificial intelligence.

[0030] The effects of the present invention are not limited to those described above, and other effects not described above will be clearly understood by those skilled in the art from the following description.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0032] Hereinafter, with reference to the accompanying drawings, embodiments of the present invention will be described in detail so that those having ordinary knowledge in the technical field to which the present invention pertains can easily implement them. The present invention can be realized in various different forms and is not limited to the embodiments described here.

[0033] To clearly explain the present invention, parts not related to the explanation are omitted, and the same reference numerals are given to the same or similar components throughout the specification.

[0034] In addition, in some embodiments, for components having the same configuration, only the representative embodiments are described using the same reference numerals, and in other embodiments, only the configurations different from the representative embodiments are described.

[0035] Throughout the specification, when a part is "connected (or coupled)" to another part, this includes not only the case where it is "directly connected (or coupled)", but also the case where it is "indirectly connected (or coupled)" with another member interposed therebetween. Also, when a part "includes" a certain component, this means that, unless otherwise stated to the contrary, it can further include other components, rather than excluding other components.

[0036] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention pertains. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with the meaning in the context of the related art, and should not be interpreted in an idealized or overly formal sense unless clearly defined in this application.

[0037] In this document, the content mainly focuses on extracting identification information by utilizing the wrinkle shape (nose pattern) of a dog's nose. However, in the present invention, the scope of pets is not limited to dogs, and the features used as identification information are not limited to nose patterns, and various physical characteristics of pets can be used.

[0038] As described above, the nose pattern images of pets suitable for learning or identification based on AI are not sufficient, and the quality of the nose pattern images of pets is likely to be low. Therefore, it is necessary to selectively store the nose pattern images in a database for learning or identification based on AI.

[0039] Figure 2 shows a procedure for pet nose pattern management based on AI to which the determination of the fitness of nose pattern images for learning or identification according to the present invention is applied. In the present invention, after photographing the nose pattern of a pet, it is first determined whether the photographed nose pattern image is suitable as data for learning or identification based on AI. If it is determined to be suitable, it is transmitted and stored in a server for learning or recognition based on AI and used as data for subsequent learning or identification.

[0040] As shown in FIG. 2, the nose print management procedure according to the present invention includes a nose print acquisition procedure and a nose print recognition procedure.

[0041] According to the present invention, when newly registering a pet's nose print, after taking an image including the pet, a nose print image is extracted from the pet's face area. In particular, it is first determined whether the nose print image is suitable for identifying or learning the pet. If the captured image is determined to be suitable for identification or learning, the image is transmitted to a server (artificial intelligence neural network) and stored in a database.

[0042] When querying the identification information of a pet via the nose print, similarly, after taking an image including the pet, a nose print image is extracted from the pet's face area. In particular, it is first determined whether the nose print image is suitable for identifying or learning the pet. If the captured image is determined to be suitable for identification or learning, the image is transmitted to the server, and the identification information of the pet is extracted through matching with the previously stored nose print image.

[0043] In the case of the nose print registration procedure, as shown in FIG. 2(a), the pet is photographed (S205), the face area (hereinafter described as the first feature area) is first detected from the photographed pet image (S210), the area occupied by the nose within the face area (hereinafter described as the second feature area) is detected, and a nose print image is output through a quality inspection for whether the captured image is suitable for learning or identification (S215). The output image is transmitted to a server constituting an artificial neural network and stored and registered (S220).

[0044] In the case of the nose print inquiry procedure, as shown in Fig. 2(b), the pet is photographed (S230), the face area is detected from the pet's image (S235), the area occupied by the nose within the face area is detected, and a nose print image is output through a quality inspection for whether the photographed image is suitable for learning or identification (S240), which is similar to the nose print registration procedure. In the subsequent procedures, a process of comparing the output nose print image with the previously stored and learned nose print images to search for matching information (S245) and an output process for the search result (S250) are performed.

[0045] Fig. 3 shows the procedure for detecting an object corresponding to the nose of a pet in the pet nose print management system according to the present invention.

[0046] Referring to Fig. 3, first, the pet is photographed to generate an initial image (S305), and the step of detecting the face area from the initial image is performed first (S310). Then, the step of detecting the nose area is performed considering the breed of the pet within the face area (S315). Detecting the face area first and then the nose area can reduce the computational complexity compared to detecting the nose area considering all breeds by cascaded detection, and can also improve the detection accuracy. Then, a quality inspection is performed to check whether the image of the detected nose area is suitable for future nose print identification or learning (S320). If, as a result of the quality inspection, it is determined that the image is suitable, the image can be transmitted to the server and used for nose print identification, or stored for future learning or identification (S325).

[0047] Also, according to the present invention, the camera can be controlled so that the detected nose area is focused, so that an image of an object for pet identification, such as the wrinkles (nose print) of a dog's nose, is not blurred when photographed (S330). This is to prevent the quality of the image from deteriorating due to the focus of the nose being off by making the focus of the camera match the nose area.

[0048] FIG. 4 shows an example of a UI (User Interface) screen for acquiring a nose print image of a pet to which the present invention is applied. FIG. 4 shows a case for acquiring the nose print of a dog among various pets.

[0049] Referring to FIG. 4, the species of the pet in the image being taken is identified to determine whether the pet currently being photographed is a dog. If the pet being photographed is not a dog, a message such as "Can't find a dog" is output as shown in FIG. 4(a). If the pet being photographed is a dog, the procedure for acquiring the nose print of the dog is performed. To determine whether the pet being photographed is a dog, the face region of the pet included in the image is first extracted, and the image included in the face region is compared with existing learned data to determine the species of the pet.

[0050] Thereafter, as shown in FIGS. 4(b) to (e), after setting the region corresponding to the dog's nose on the dog's face, shooting can be performed with the focus on the region corresponding to the nose. That is, the camera can be controlled so that the focus is set on the position (center point) of the region corresponding to the object for pet identification. Also, in order to feedback to the user that the object currently being tracked (e.g., the nose) is being photographed with the focus, a graphic element can be overlaid on the position of the object being tracked. By displaying a graphic element indicating the detection state of the object at the position of the object being tracked, the user can recognize that object recognition is being performed on the pet currently being photographed.

[0051] As shown in FIGS. 4(b) to 4(e), when the image quality of the object being currently photographed is good (when the quality of the object image satisfies the reference conditions), a first graphic element 410a indicating a good quality state (e.g., a smiling icon or a green icon) can be overlaid on the object and output. When the image quality of the object being currently photographed is poor (when the quality of the object image does not satisfy the reference conditions), a second graphic element 410b indicating a poor quality state (e.g., a crying icon or a red icon) can be overlaid on the object and output.

[0052] As shown in FIG. 4, even when the dog moves continuously, shooting can be performed while tracking the dog's nose and focusing on the nose. At this time, in each captured image, it may be determined whether the nose pattern image of the dog is suitable for pet identification or learning, and the degree of suitability may be output.

[0053] For example, the degree to which the captured nose pattern image of the dog is suitable for pet identification or learning can be calculated as a numerical value, and according to the numerical value of the degree of fitness, score information 420 in a form where the gauge is filled in the "Bad" direction as the degree of fitness is lower and in the "Good" direction as the degree of fitness is higher can be output. That is, score information 420 indicating the shooting quality of the object in the image can be output.

[0054] In addition, quality evaluation (such as size, brightness, sharpness, etc.) can be performed on the image of the nose pattern being currently captured, and feedback can be provided to the user so that a nose pattern image suitable for identification or learning based on artificial intelligence is captured. As user feedback, a message 430 can be output to guide the user to appropriately capture the object of the pet. For example, when the size of the image of the dog's nose pattern is smaller than the reference value, a message such as "Please adjust the distance between the dog's noses" can be output as shown in (c) of FIG. 4 so that a nose pattern image of a larger size is captured. During the process of capturing an object (e.g., a dog's nose pattern) for pet identification, as feedback to guide the user to appropriately capture the image, an audio message (e.g., "Please adjust the distance between the dog's noses") can be output via the speaker. Alternatively, the haptic module can provide feedback to the user by generating vibrations so that the object for pet identification is appropriately captured. For example, when the distance of the pet from the camera 1010 is too far or too close, the haptic module can generate vibrations to allow the user to adjust to an appropriate distance.

[0055] In addition, progress rate information 440 indicating the degree of progress in obtaining an image of an object having a quality suitable for pet identification can be output. For example, if 5 nose pattern images with suitable quality are required and 1 suitable image has been obtained so far, progress rate information 440 indicating that the progress rate is 25% can be output as shown in FIG. 4.

[0056] When the image of the dog's nose pattern has been sufficiently obtained, the photographing is terminated, and the identification information can be stored in the database together with the image of the dog's nose pattern, or the identification information of the dog can be output.

[0057] In the present invention, after first detecting the face region of a pet, the nose region is detected within the face region. This is to reduce the object detection difficulty while reducing the computational complexity. During the process of taking an image, an object other than the object to be detected, or unnecessary or incorrect information can be included in the image. Therefore, the present invention first determines whether a desired object (the nose of the pet) exists in the image being taken.

[0058] Also, in order to identify the nose pattern of a pet, an image having a resolution above a certain level is required. However, there is a problem that the amount of computation for image processing increases as the resolution of the image becomes higher. Also, as the types of pets increase, since the learning methods differ for each type of pet, there is a problem that the computational difficulty of artificial intelligence further increases. In particular, similar types of animals have similar shapes (e.g., the nose of a dog and the nose of a wolf are similar), so classifying the nose together with the type of animal for similar animals can have a very high computational difficulty.

[0059] Therefore, in order to reduce such computational complexity, the present invention uses a cascaded object detection method. For example, after first detecting the face region of a pet while taking a picture of the pet, the type of the pet is identified, and based on the detected face region of the pet and the identified type of the pet, the nose region of the pet is detected. This first performs the process of identifying the type of pet at a relatively low resolution with a relatively low computational complexity, applies the object detection method determined according to the type of pet to maintain a high resolution in the face region of the pet, and performs nose region detection. Therefore, the present invention can effectively detect the nose region of a pet while relatively reducing the computational complexity.

[0060] Figure 5 shows the general image processing process for pet identification according to the present invention. As shown in Figure 5, the method for processing an input image according to the present invention includes a step of receiving an input image from a camera (S505), a first preprocessing step of adjusting the size of the input image to generate a primary processed image (S510), a first feature region detection step of detecting the position and species of an animal from the processed image generated in the first preprocessing step (S515), a first postprocessing step of extracting a first feature value of the animal image from the result of the first feature region detection step (S520), a step of determining a detector for detecting an object (e.g., nose) for pet identification according to the species of the pet from the image processed through the first postprocessing step (S525), a second preprocessing step of adjusting the size of the image for image processing for pet identification (S530), at least one second feature region detection step corresponding to each species of animal that can be detected in the first feature detection step (S535), and a second postprocessing step of extracting a second feature value of the animal image corresponding to each second feature region detection step (S540).

[0061] The first preprocessing step

[0062] The step of applying the first preprocessing to the original image (S510) is a step of adjusting the size, ratio, direction, etc. of the original image to convert the image into a form suitable for object detection.

[0063] With the development of camera technology, most input images are composed of millions to tens of millions of pixels, and it is not preferable to directly process such large images. In order for object detection to operate efficiently, a preprocessing process must be performed to make the input image suitable for processing. Such a process is mathematically performed by coordinate system transformation.

[0064] It is obvious that an arbitrary processed image can be generated by corresponding any four points in the input image to the four vertices of the processed image and passing through an arbitrary coordinate system transformation process. However, when using an arbitrary non-linear transformation function in the coordinate system transformation process, the inverse transformation for obtaining the feature region of the input image from the bounding box acquired as the result of the feature region detector must be possible. For example, when using an Affine Transformation that linearly transforms by corresponding any four points in the input image to the four vertices of the processed image, the inverse transformation process can be easily obtained, so it is preferable to use this.

[0065] As an example of a method for determining any four points in the input image, a method of directly using the four vertices of the input image can be considered. Alternatively, a method of adding margins to the input image or cropping a part of the input image can be used so that the horizontal length and the vertical length are transformed at the same ratio. Alternatively, various interpolation methods can be applied to reduce the size of the input image.

[0066] First Feature Region Detection Step

[0067] In this step, by first detecting the region where the pet exists and the species of the animal in the pre-processed image, the first feature region that can be used in the second feature region detection step described later is set, and at the same time, by selecting a second feature region detector optimized for each pet species, the purpose is to improve the final feature point detection performance.

[0068] In this process, any object detection and classification method can be easily combined by those with ordinary knowledge in the relevant field. However, compared with conventional methods, it is known that the method based on artificial neural network is superior in performance. Therefore, it is preferable to use the feature detection technique based on artificial neural network as much as possible. For example, a feature detector of SSD (Single Vision-Shot Multibox Detection) method, which is an algorithm for detecting objects of various sizes for a single image by an artificial neural network, can be used.

[0069] The input image normalized by the above-mentioned pre-processor constitutes the first feature image to the nth feature image hierarchically by an artificial neural network. At this time, the method of extracting feature images for each layer can be mechanically learned in the learning step of the artificial neural network.

[0070] The hierarchically extracted feature images are combined with a corresponding predefined box (Priori Box) list for each layer to generate a bounding box, an individual type, and a list of confidence values. Such an operation process can also be mechanically learned in the learning step of the artificial neural network. For example, the result value is returned in the form as shown in Table 1 below. At this time, the number of types that can be determined by the neural network is determined in the neural network design step, and when it is implicitly determined that there is no object, that is, "background" is defined.

[0071]

Table 1

[0072] Such result boxes are merged with overlapping result boxes by the NMS (Non-Maximum Suppression) step and finally returned as the object detection results present in the image. NMS is a process of deriving the final feature region from a plurality of feature region candidates, and the feature region candidates can be generated by considering the probability values in Table 1 according to the procedure as shown in FIG. 6.

[0073] Next, this process will be described in detail.

[0074] 1. For each species excluding the background, perform the following process respectively.

[0075] A. In the bounding box list, exclude the boxes whose probability of being the said species is lower than a specific threshold. If there are no remaining boxes, end without results

[0076] B. In the said bounding box list, designate the box with the highest probability of being the said species as the first box (the first boundary region) and exclude it from the bounding box list.

[0077] C. For the remaining bounding box list, perform the following process respectively in the order of decreasing probability.

[0078] i. Calculate the IoU (Intersection over Union, the area ratio of the intersection to the union) with the first box.

[0079] ii. If the IoU is higher than a specific threshold, this box is a box that overlaps with the first box. Merge it with the first box.

[0080] D. Add the first box to the result box list.

[0081] E. If there are boxes remaining in the bounding box list, repeat from step C for the remaining boxes.

[0082] For two boxes A and B, the IoU can be effectively calculated as shown in Equation 1 below.

[0083]

Number

[0084] Next, a method for merging the boxes that overlap with the first box in the above process will be described. For example, the first box can be maintained as it is, and the second box can be merged by deleting it from the bounding box list (Hard NMS). Or, the first box can be maintained as it is, and the probability that the second box is a specific type can be weighted and decreased by a value between (0, 1). If the attenuated result value is smaller than a specific threshold, it can be merged by deleting it from the bounding box list for the first time (Soft NMS).

[0085] Finally, it is natural that for one or more bounding boxes determined in this way, the feature region in the original image can be obtained by passing through the inverse transformation process for any transformation process used in the preprocessing step. Depending on the configuration, it can be adjusted so that the second detection step described later can be performed well by adding a certain amount of margin to the feature region in the original image.

[0086] The First Post-Processing Step

[0087] For each feature region of the input image obtained in the first feature region setting step (S515) described above, by performing an additional post-processing step, the first feature value can be generated. For example, in order to obtain the brightness information (the first feature value) for the first feature region of the input image, an operation as shown in Equation 2 below can be performed.

[0088]

Number

[0089] At this time, L is the Luma value according to the BT.601 standard, and V is the brightness value defined in the HSV color space. M and N are the horizontal width and vertical height of the target feature region.

[0090] Using the first feature value additionally generated in this way, it is possible to predict whether the first feature region obtained in the first feature region detection step (S515) is suitable for use in the application field combined with this patent. It is obvious that the first feature value to be additionally generated must be appropriately designed according to the application field. If the conditions of the first feature value defined in the application field are not met, the system can be configured to selectively omit the second feature region setting and object detection steps described later.

[0091] Second Feature Region Detection Step

[0092] This step aims to extract the specific feature region required in the application field from the region where the animal exists. For example, in the application field of detecting the positions of eyes, nose, mouth, and ears from the face region of an animal, in the first feature region detection step, first distinguish the face region of the animal and the species information of the animal, and in the second feature region detection step, aim to detect the positions of eyes, nose, mouth, and ears according to the species of the animal.

[0093] In this process, the second feature region detection step can be composed of a plurality of feature region detectors independent of each other specialized for each species of animal. For example, if it is possible to distinguish dogs, cats, and hamsters in the first feature region detection step, it is preferable to have three second feature region detectors, each designed to be specialized for dogs, cats, and hamsters. By doing so, it is possible to reduce the types of features to be learned by individual feature region detectors and reduce the learning complexity, and it is obvious that neural network learning is possible with only a smaller number of data also in terms of learning data collection.

[0094] Since each second feature region detector is configured independently of the others, a person with ordinary knowledge can easily configure independent individual detectors. Each feature region detector is preferably configured individually according to the feature information to be detected for each species. Alternatively, in order to reduce the complexity of the system configuration, some or all of the second feature region detectors share feature region detectors with the same structure, but a method of configuring the system to be suitable for each species by alternating the learning parameter values can be used. Furthermore, a method of further reducing the system complexity can be considered by using a feature region detector with the same structure as the first feature region detection step as the second feature region detector, but only alternating the learning parameter values and the NMS method.

[0095] For one or more feature regions set through the first feature region detection step and the first post-processing step, determine which second feature region detector to use using the species information detected in the first feature region detection step, and perform the second feature region detection step using the determined second feature region detector.

[0096] First, perform a preprocessing process. At this time, it is obvious that a conversion process that allows inverse conversion must be used in the process of converting coordinates. In the second preprocessing process, since the first feature region detected in the input image must be converted into the input image of the second feature region detector, it is preferable to define the four points required for the design of the conversion function as the four vertices of the first feature region.

[0097] Since the second feature region obtained through the second feature region detector is a value detected using the first feature region, the first feature region must be considered when calculating the second feature region within the entire input image.

[0098] For the second feature area obtained via the second feature area detector, by performing additional post-processing steps in the same manner as the first post-processing step, second feature values can be generated. For example, in order to obtain the sharpness of an image, a Sobel filter can be applied, or information such as the posture of the animal to be detected can be obtained using the presence or absence of detection and the relative positional relationship between the feature areas.

[0099] Using the second feature values additionally generated in this way, it is possible to predict whether the feature areas obtained in the second object detection step are suitable for use in the application fields where they are combined with this patent. It is obvious that the second feature values additionally generated must be appropriately designed according to the application field. If the conditions of the second feature values defined in the application field are not met, it is preferable to design so that data can be obtained to suit the application field, such as excluding not only the second detection area but also the first detection area from the detection results.

[0100] System Expansion

[0101] In the present invention, by constituting a two-step detection step, an example of a system and a configuration method are given in which the position and species of an animal are detected in the first feature position detection step, and based on the result, a detector used in the second feature position detection step is selected.

[0102] Such a cascade configuration can be easily extended to a multi-layer cascade configuration. For example, in the first feature position detection step, the entire body of the animal is detected, in the second feature position detection step, the position of the face and the position of the limbs of the animal are detected, and in the third feature position detection step, the positions of the eyes, nose, mouth, and ears are detected from the face, etc. Application configurations are possible.

[0103] By using such a multi-level stepped structure, a system capable of simultaneously acquiring characteristic positions of multiple levels can be easily designed. When designing a multi-level stepped system, in order to determine the number of levels, it is obvious that the optimal hierarchical structure can be designed by considering the hierarchical domain of the characteristic positions to be acquired, the operation time and complexity of the entire system, and the resources required to configure each individual characteristic area detector, etc.

[0104] The process of variably adjusting the focal position to obtain a high-quality object image (e.g., nose pattern image) for pet identification will be described. Many techniques for adjusting the focus according to the position of an object are applied to general smartphones, but such techniques mainly target objects that do not move during the shooting time like humans. When shooting a pet, the identification object (e.g., nose) often moves continuously, making it difficult to focus. When the focus cannot be adjusted to the identification object, data not suitable for learning or identification based on artificial intelligence is likely to be generated. Therefore, it is preferable to adjust the focus so that it can be adjusted to the identification object of the pet.

[0105] Therefore, the present invention provides a method of continuously setting the focal position to the position of the object according to the position of the detected object after detecting the pet object.

[0106] According to the present invention, an image including a pet is acquired, an area (first characteristic area, e.g., face area) for determining the species of the pet is set, and the image in the first characteristic area is extracted. Then, an area (second characteristic area, e.g., nose area) for identifying the pet according to the species of the pet is set in the first characteristic area, and it is inspected whether the image (e.g., nose pattern image) in the second characteristic area is suitable for learning or identification based on artificial intelligence. When the image of the object in the second characteristic area has a quality suitable for learning or identification, it is transmitted to and stored in a neural network (e.g., server) for learning or identification. When it is not suitable, the image is discarded.

[0107] After that, when taking an image of the (N + 1)-th frame, the camera is adjusted so that the focus is positioned at the position of the object detected in the N-th frame, whereby an object image (nose print image) for identifying a pet can be obtained so as to be suitable for learning or identification based on artificial intelligence.

[0108] FIG. 7 shows a process of photographing an object for identifying a pet by adjusting a focus position according to the present invention. Referring to FIG. 7, while photographing a pet, an object (e.g., nose) for identifying the pet is continuously tracked, and the position of the focus for photographing the pet in a specific frame is set to the position of the object set in the previous frame. According to the present invention, detection of a stepped object is performed for each frame. As described with reference to FIG. 5, a first feature region is set primarily to determine the species of an animal, and secondarily, a second feature region where an object for identification (e.g., nose) is located within the first feature region is derived in consideration of the species of the animal. After the detection of the object is completed, post-processing (second post-processing) is performed to determine whether the image of the frame is suitable for learning or identification. When a suitable image is obtained, the image of the frame is stored for subsequent learning or identification.

[0109] The photographing of the next frame is continued to obtain an object image (e.g., nose print image) for identifying a pet. Here, the position of the focus at the time of photographing is the position of the object determined in the previous frame. By changing the position of the focus in accordance with the position of the object detected in the previous frame, even when the pet moves without staying still, the focus can be continuously moved following the object for identification. Such a process can reduce the possibility of the focus being shifted and obtain a high-quality object image. Such a process is repeatedly performed, and the possibility of obtaining an object image more suitable for learning or identification can be increased through the movement of the camera focus.

[0110] FIG. 8 is a flowchart of a method for photographing an object corresponding to a pet's nose in a pet nose print management system according to the present invention.

[0111] A method for photographing an object for pet identification according to the present invention includes a step of obtaining an image including a pet (S810), a step of detecting an object for pet identification from the image (S820), and a step of photographing a next image with the focus set at the position of the detected object (S830). Thereafter, object detection and focus position change for the next image (frame) can be repeatedly performed.

[0112] According to the present invention, the step of obtaining an image including a pet (S810) can include a step of photographing the pet with the focus set at the position of the object detected from the image photographed in the previous frame to generate an image.

[0113] According to the present invention, the step of photographing a next image (S830) can include a step of detecting a changed position of the object from the next image, a step of determining whether an image of the object whose position has changed in the next image is suitable for learning or identification based on artificial intelligence, and a step of performing the next photographing with the focus set at the changed position.

[0114] As shown in FIG. 7, an identification object (e.g., a pet's nose) is detected for the image of the current frame, and it is determined whether the image of the object (nose print image) is suitable for learning or identification based on artificial intelligence. The position of the object detected from the image of the current frame can be set as the focus position when photographing the image of the next frame. By adjusting the focus of the current frame with reference to the object position of the previous frame as in the present invention, the focus can be maintained on the object for pet identification, and an image (nose print image) of the identification object having a quality that can be learned or identified can be obtained.

[0115] According to the present invention, the step of detecting an object for pet identification from an image may include a step of setting a first feature region for determining the species of the pet from the image, and a step of setting a second feature region including an object for pet identification according to the species of the pet within the first feature region. Instead of simultaneously detecting the species of the pet and the object for pet identification from the image, the present invention first performs the detection of the species of the pet that can be easily processed with a relatively low resolution preferentially, and then detects the object for pet identification in consideration of the species of the pet, thereby effectively obtaining an object identification image (nose print image) while reducing the computational complexity.

[0116] According to the present invention, the step of setting the second feature region includes a step of determining whether an image of an object for pet identification (e.g., nose print image) is suitable for learning or identification based on artificial intelligence. After photographing a pet, instead of storing all the images, it is first determined whether an image of an object for pet identification is suitable for learning or identification based on artificial intelligence, and only the suitable object image is stored in a server or a neural network for learning or identification, thereby preventing unnecessary data from being stored.

[0117] According to the present invention, the step of determining whether the image of the object is suitable for learning or identification based on artificial intelligence may include a step of determining whether the quality of the image of the object satisfies a reference condition, a step of transmitting the image of the object to a learning or identification server when the quality satisfies the reference condition, and a step of discarding the image of the object and performing photographing for the next image when the quality does not satisfy the reference condition. The quality of the image of the object can be expressed by the aforementioned first quality value and second quality value. If the first quality value and the second quality value are greater than the reference value, it can be determined that the object image has a quality suitable for learning or identification.

[0118] FIG. 9 shows the process of photographing an object for identifying a pet by object tracking according to the present invention. The method of photographing an object for identifying a pet according to the present invention includes a step of acquiring an image including the pet (S910), a step of detecting an object for identifying the pet from the image (S920), a step of setting a focus on the position of the object (S930), and a step of outputting an image in which a graphic element indicating an object detection state is overlaid on the position of the object where the focus is set (S940). For example, as shown in FIG. 4, an object (nose) for identifying a pet is detected from the image being photographed, the focus is set on the position of the object and photographed, and graphic elements 410A and 410B indicating the detection state of the object are overlaid and output. By causing the graphic element to track the object detected for each frame, the user can recognize whether the object is correctly tracked and can confirm whether the image quality of the object currently being photographed is appropriate.

[0119] According to an embodiment of the present invention, the step of outputting the image (S940) can include a step of determining whether the quality of the object image satisfies a reference condition, a step of overlaying a first graphic element 410A indicating a good quality state on the object when the quality of the object image satisfies the reference condition, and a step of overlaying a second graphic element 410B indicating a bad quality state on the object when the quality of the object does not satisfy the reference condition.

[0120] According to an embodiment of the present invention, the step of outputting an image (S940) may include a step of outputting score information 420 indicating the shooting quality state of the object in the image. As shown in FIG. 4, the degree to which the photographed nose pattern image of a dog is suitable for pet identification or learning can be calculated numerically, and based on the numerical value for the fitness, the score information 420 in a form where the gauge is filled in the "BAD" direction as the fitness is lower and in the "GOOD" direction as the fitness is higher can be output.

[0121] According to an embodiment of the present invention, a method of photographing an object for identifying a pet may further include a step of providing feedback to the user so that a good object image is photographed according to the quality state of the image of the object. The feedback can include one of a message output together with the image, an audio message output via a speaker, or haptic feedback output via vibration. For example, when the size of the nose pattern image of a dog is smaller than a reference value, a message such as "Please adjust the distance of the dog's nose" can be output as shown in FIG. 4(c) so that a nose pattern image of a larger size is photographed. For example, during the process of photographing an object (e.g., the pattern of a dog's nose) for pet identification, an audio message (e.g., "Please adjust the distance of the dog's nose") can be output via a speaker as feedback to guide the user so that the image is properly photographed. Alternatively, the haptic module can provide feedback to the user by generating vibration so that the object for pet identification is properly photographed. For example, when the distance of the pet from the camera 1010 is too far or too close, the haptic module can generate vibration to allow the user to adjust to an appropriate distance.

[0122] According to an embodiment of the present invention, the step of obtaining an image including a pet (S910) may include the step of photographing the pet with the focus set on the position of the object detected from the image captured in the previous frame to generate the image.

[0123] FIG. 10 is a block diagram of an electronic device 1000 according to the present invention. The electronic device 1000 according to the present invention includes a camera 1010 that generates an original image including a pet, and a processor 1020 that detects an object for identifying the pet from the image and controls the camera 1010 to capture the next image with the focus set on the position of the detected object.

[0124] The camera 1010 may include an optical module such as a lens and a CCD (Charge-Coupled Device) or a CMOS (Complementary Metal-Oxide Semiconductor) that generates an image signal from the input light, and may generate image data through image capture and provide it to the processor 1020.

[0125] The processor 1020 controls each module of the electronic device 1000 and performs operations necessary for image processing. The processor 1020 may be composed of a plurality of microprocessors (processing circuits) according to its function. As described above, the processor 1020 can detect an object (e.g., nose) for identifying a pet (e.g., dog) and determine the validity of the image for the object. For example, the processor 1020 can detect an object (e.g., nose) for identifying a pet object from the captured image and adjust the lens distance of the camera 1010 so that the object is in focus. In addition, the processor 1020 can overlay and output a graphic element indicating that the position of the detected object is being tracked, or output score information indicating the quality state of the object image and a message providing feedback to the user to the display 1050.

[0126] The communication module 1030 can transmit or receive data with an external entity via a wired / wireless network. In particular, the communication module 1030 can exchange data for artificial intelligence-based processing via communication with a server for learning or identification.

[0127] Furthermore, the electronic device 1000 can include a memory 1040 for storing image data and information necessary for image processing, and a display 1050 for outputting a screen to the user, and can include various modules according to the application. The display 1050 can output an image of the pet during shooting and a pet interface screen. Also, although not shown, the electronic device 1000 can include a speaker for outputting an acoustic signal and a haptic module for providing haptic feedback to the user via vibration. For example, during the process of shooting an object (e.g., a dog's nose print) for pet identification, as feedback to guide the user so that the image is properly captured, an audio message (e.g., "Please adjust the distance of the dog's nose") can be output via the speaker. Or, the haptic module can provide feedback to the user by generating vibration so that an object for pet identification is properly captured. For example, when the distance of the pet from the camera 1010 is too far or too close, the haptic module can generate vibration to allow the user to adjust to an appropriate distance.

[0128] According to the present invention, the processor 1020 can control the camera 1010 to capture the pet to generate the image with the focus set at the position of the object detected from the image captured in the previous frame.

[0129] According to the present invention, the processor 1020 can detect the changed position of the object in the next image, determine whether the image of the object whose position has changed in the next image is suitable for learning or identification based on artificial intelligence, and control the camera 1010 to perform the next shooting with the focus set at the changed position.

[0130] According to the present invention, the processor 1020 can set a first feature region for determining the species of the pet from the image, and set a second feature region including an object for identifying the pet within the first feature region.

[0131] According to the present invention, the processor 1020 can determine whether the image of the object for identifying the pet is suitable for learning or identification based on artificial intelligence.

[0132] According to the present invention, the processor 1020 can determine whether the quality of the image of the object satisfies the reference conditions. When the quality satisfies the reference conditions, the image of the object is transmitted to the learning or identification server. When the quality does not satisfy the reference conditions, the image of the object is discarded, and the camera 1010 can be controlled to perform shooting for the next image.

[0133] According to the present invention, the processor 1020 can output to the display 1050 an image in which a graphic element indicating the object detection state is overlaid at the detected position of the object.

[0134] According to the present invention, the processor 1020 can determine whether the quality of the image of the object satisfies the reference conditions. When the quality of the image of the object satisfies the reference conditions, a first graphic element 410A indicating a good quality state is overlaid on the object and output via the display 1050. When the quality of the object does not satisfy the reference conditions, a second graphic element 410b indicating a bad quality state is overlaid on the object and output via the display 1050.

[0135] According to the present invention, the processor 1020 can output score information 420 indicating the shooting quality state of an object in an image to the display 1050.

[0136] According to the present invention, the processor 1020 can output a message 430 for providing feedback to the user to the display 1050 so that a good object image is captured according to the quality state of the object image.

[0137] This embodiment and the drawings attached to this specification only clearly show a part of the technical idea included in the present invention, and all modifications and specific embodiments that can be easily analogized by those skilled in the art within the scope of the technical idea included in the specification and drawings of the present invention are self-evidently included in the scope of the rights of the present invention.

[0138] Therefore, the idea of the present invention should not be defined as being limited to the described embodiments, and not only the scope of the following claims, but also all those having an equivalent or equivalent transformation to this scope of claims belong to the category of the idea of the present invention.

Claims

1. A method for photographing an object for pet identification, comprising: obtaining an image including the pet; detecting an object for pet identification from the image; photographing a next image with the focus set at the position of the detected object; The step of obtaining an image including the pet includes photographing the pet with the focus set at the position of the object detected from the image photographed in the previous frame to generate the image. The position of the focus when photographing the image of the (N + 1)-th frame is changed according to the position of the object for pet identification detected in the N-th frame. The method, wherein the changing of the focus position and the image photographing are repeated until a target number of images that satisfy the set quality conditions are obtained.

2. The step of detecting an object for pet identification from the image includes: setting a first feature region for determining the species of the pet from the image; setting a second feature region including an object for pet identification according to the species of the pet within the first feature region. The method according to claim 1.

3. The step of setting the second feature region includes determining whether the image of the object for pet identification is suitable for learning or identification based on artificial intelligence. The method according to claim 2.

4. The step of determining whether the image of the object is suitable for learning or identification based on artificial intelligence includes: determining whether the quality of the image of the object satisfies reference conditions; when the quality satisfies the reference conditions, transmitting the image of the object to a server; when the quality does not satisfy the reference conditions, discarding the image of the object and photographing the next image. The method according to claim 3.

5. A method for photographing an object for pet identification, comprising: obtaining an image including the pet; detecting an object for pet identification from the image; setting the focus at the position of the object; Outputting the image with a graphic element indicating the object detection state overlaid on the position of the set object of the focus; The step of obtaining the image including the pet includes the step of photographing the pet with the focus set at the position of the object detected from the image photographed in the previous frame to generate the image; The position of the focus when photographing the image of the (N + 1)-th frame is changed according to the position of the object for identifying the pet detected in the N-th frame; The method in which the focus position change and image photographing are repeated until the target number of nose pattern images satisfying the set quality conditions are obtained. [

6. ] The step of outputting the image includes: Determining whether the quality of the image of the object satisfies the reference conditions; When the quality of the image of the object satisfies the reference conditions, overlaying a first graphic element indicating a good quality state on the object; When the quality of the object does not satisfy the reference conditions, overlaying a second graphic element indicating a bad quality state on the object, the method according to claim 5. [

7. ] The step of outputting the image includes the step of outputting score information indicating the photographing quality state of the object in the image, the method according to claim 5. [

8. ] The method according to claim 5, further including the step of providing feedback to the user so that a good object image is photographed according to the quality state of the image of the object. [

9. ] An electronic device for photographing an object for identifying a pet, including: A camera for generating an image including the pet; A processor for detecting an object for identifying the pet from the image and controlling the camera to photograph the next image with the focus set at the position of the detected object; The processor: Controls the camera to photograph the pet with the focus set at the position of the object detected from the image photographed in the previous frame to generate the image; The position of the focus when photographing the image of the (N + 1)-th frame is changed according to the position of the object for identifying the pet detected in the N-th frame. An electronic device in which the position change of the focus and image capturing are repeated until a target number of nose print images that satisfy set quality conditions are acquired.

10. The processor sets a first feature region for determining the species of the pet from the image, and sets a second feature region including an object for identifying the pet within the first feature region. The electronic device according to claim 9.

11. The processor determines whether an image of an object for identifying the pet is suitable for learning or identification based on artificial intelligence. The electronic device according to claim 10.

12. The processor determines whether the quality of the image of the object satisfies reference conditions, when the quality satisfies the reference conditions, transmits the image of the object to the server, when the quality does not satisfy the reference conditions, discards the image of the object, and controls the camera to capture the next image. The electronic device according to claim 11.

13. The electronic device according to claim 9, further including a display that outputs an image of the pet during shooting.

14. The processor outputs to the display an image in which a graphic element indicating an object detection state is overlaid at the position of the detected object. The electronic device according to claim 13.

15. The processor determines whether the quality of the image of the object satisfies reference conditions, when the quality of the image of the object satisfies the reference conditions, overlays a first graphic element indicating a good quality state on the object, when the quality of the object does not satisfy the reference conditions, overlays a second graphic element indicating a poor quality state on the object. The electronic device according to claim 14.

16. The processor outputs score information indicating the shooting quality state of the object in the image via the display. The electronic device according to claim 13.

17. The processor provides feedback to the user so that a good object image is captured according to the quality state of the image of the object. The electronic device according to claim 9.

Citation Information

Patent Citations

  • Method and system for identifying dogs in community and readable storage medium

    CN110414390A

  • Photographic device

    JP2009055448A

  • Detection information registration device, electronic device, method for controlling detection information registration device, method for controlling electronic device, program for controlling detection information device, and program for controlling electronic device

    JP2010044516A

  • Electronic camera

    JP2011130043A

  • Muzzle pattern collation system, muzzle pattern collation method and muzzle pattern collation program

    JP2019113959A