Method and electronic device for photographing objects for pet identification - Patents.com
Through image processing technology, including identifying pet species and feature areas, detecting and extracting nose print features, the problems of low recognition rate of pet nose prints and poor image quality are solved, and efficient and accurate pet recognition and image quality improvement are achieved.
Patent Information
- Application Number
- JP2023569731
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-28
- Filing Date
- 2022-06-27
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The prior art faces problems that factors such as angle, focus, distance, size, environment and other factors affect the quality of photos when obtaining and storing pet nose prints, resulting in low recognition rate of artificial intelligence and pets cannot stay like humans to obtain clear nose print photos.
Image processing methods are used to obtain the original pet image, determine the pet species and characteristic areas, and then detect and extract specific objects for identification (such as nose prints), and improve the recognition rate and image quality through neural networks and preprocessing technologies.
Effectively reduces computational complexity, improves the accuracy and efficiency of pet nose print recognition, and ensures that only high-quality images suitable for artificial intelligence learning or recognition are stored.
Smart Images

Figure 0007675214000009 
Figure 0007675214000010 
Figure 0007675214000011
Abstract
Description
[Technical field]
[0001] The present invention relates to a method and electronic device for photographing an object for pet identification, and more particularly to a method and electronic device for acquiring an image of an object for pet identification suitable for artificial intelligence based learning or identification. [Background technology]
[0002] In modern society, there is an increasing demand for pets, which people can rely on emotionally while living with them. This has led to an increasing need to manage a variety of pet information in a database for purposes such as pet health management. In order to manage pets, pet identification information, like a person's fingerprint, is necessary, and objects that can be used for each pet can be defined. For example, in the case of dogs, each dog has a different nose print (the shape of the wrinkles on the nose), so the nose print can be used as identification information for each dog.
[0003] As shown in FIG. 1(a), a method for registering a nose print is performed by photographing a face including a pet's nose (S110) in the same manner as registering a human fingerprint or face, and storing and registering the image including the nose print in a database (S120). Also, as shown in FIG. 1(b), a method for querying a nose print can be performed by photographing a pet's nose print (S130), searching for a nose print and related information that matches the photographed nose print (S140), and outputting information that matches the photographed nose print (S150). As shown in FIG. 1, each pet can be identified and information about the pet can be managed by the process of registering and querying the pet's nose print. The pet's nose print information is stored in a database and can be used as data for AI-based learning or identification.
[0004] However, there are several problems associated with capturing and storing a pet's nose print.
[0005] First, photos can be difficult to recognize due to factors such as the shooting angle, focus, distance, size, and environment. There have been attempts to apply human face recognition technology to nose print recognition, but while sufficient data has been accumulated for human face information, there is not enough data on pet nose print information, resulting in a low recognition rate. Specifically, AI-based recognition requires learning data that has been processed into a form that machines can learn from, but there is not enough data on pet nose prints, making it difficult to recognize nose prints.
[0006] Furthermore, for the purpose of recognizing a pet's nose print, an image with clear nose wrinkles is required. However, unlike humans, pets cannot stop moving for a while, so it is not easy to obtain a clear image of nose wrinkles. For example, it is very difficult to obtain a nose print image of a desired quality for a dog because it keeps moving its face and licking its tongue. For example, for the purpose of recognizing a nose print, an image with clear nose wrinkles is required, but in most cases, the nose wrinkles are not clearly captured in the image actually taken due to shaking or the like. In order to solve this problem, a method of photographing a dog's nose while forcibly fixing it has been considered, but this method is evaluated as inappropriate because it forces the pet to perform an action. Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention provides an image processing method and electronic device capable of effectively detecting objects for pet identification while reducing computational complexity.
[0008] The present invention provides an image processing method and electronic device capable of effectively detecting objects for pet identification while reducing computational complexity.
[0009] The present invention provides a method and electronic device capable of effectively filtering low quality images during the process of acquiring images of objects for pet identification.
[0010] The problems to be solved by the present invention are not limited to those described above, and other problems to be solved not described above will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0011] A method for detecting an object for identifying a pet according to the present invention includes the steps of acquiring an original image including the pet, determining a first feature area and the species of the pet through image processing on the original image, and detecting an object for identifying the pet within the first feature area based on the determined species of the pet.
[0012] According to the present invention, the step of determining the pet species may include a step of applying a first pre-processing to the original image, a step of determining the pet species in the pre-processed image to set the first feature region, and a step of extracting a first feature value through a first post-processing to the first feature region.
[0013] According to the present invention, the step of setting the first feature region includes the steps of generating a plurality of feature images from the preprocessed image using a learning neural network, applying a predefined bounding box to each of the plurality of feature images, calculating a probability value for each pet type within the bounding box, and configuring the first feature region to include the bounding box if the calculated probability value for a particular animal species is equal to or greater than a reference value.
[0014] According to the present invention, if the first feature value is greater than a reference value, object detection for identifying the pet is performed, and if the first feature value is less than a reference value, additional processing may be omitted.
[0015] According to the present invention, the step of applying a first pre-processing to the original image may include a step of converting the original image into an image of a first resolution lower than the original resolution, and a step of applying the first pre-processing to the image converted to the first resolution.
[0016] According to the present invention, the step of detecting an object for identifying the pet may include a step of applying a second pre-processing to a first feature region for identifying the species of the pet, a step of setting a second feature region for identifying the pet based on the species of the pet in the second pre-processed first feature region, and a step of applying a second post-processing to the second feature region to extract a second feature value.
[0017] According to the present invention, the second pre-processing on the first feature region may be performed at a second resolution higher than the first resolution at which the first pre-processing for setting the first feature region is applied.
[0018] According to the present invention, the step of setting the second feature area may include a step of setting the second feature area based on the probability that an object for identifying the pet is located in the first feature area depending on the species of the pet.
[0019] According to the present invention, if the second feature value is greater than a reference value, an image including the second feature region may be transmitted to a server.
[0020] According to the present invention, the step of generating the first feature region may include a step of generating feature region candidates for determining the species of the pet in the image, and a step of generating a first feature region whose position and size are determined based on a reliability value of each of the feature region candidates.
[0021] The electronic device according to the present invention includes a camera that generates an original image including the pet, a processor that determines a first feature area and the species of the pet through image processing on the original image and detects an object for identifying the pet within the first feature area based on the determined species of the pet, and a communication module that transmits an image of the object to a server if the object for identifying the pet is valid.
[0022] According to the present invention, the processor can apply a first pre-processing to the original image, determine the species of the pet in the pre-processed image to set the first feature region, and extract a first feature value through a first post-processing to the first feature region.
[0023] According to the present invention, the processor can use a learning neural network to generate a plurality of feature images from the preprocessed image, apply predefined bounding boxes to each of the plurality of feature images, calculate a probability value for each pet type within the bounding boxes, and configure the first feature region to include the bounding box if the calculated probability value for a specific animal species is greater than or equal to a reference value.
[0024] According to the present invention, if the first feature value is greater than a reference value, object detection for identifying the pet is performed, and if the first feature value is less than a reference value, additional processing may be omitted.
[0025] According to the present invention, the processor can convert the original image into an image with a first resolution lower than the original resolution, and apply the first pre-processing to the image converted to the first resolution.
[0026] According to the present invention, the processor can apply a second pre-processing to a first feature region for identifying the species of the pet, set a second feature region for identifying the pet based on the species of the pet in the second pre-processed first feature region, and apply a second post-processing to the second feature region to extract a second feature value.
[0027] According to the present invention, the second pre-processing on the first feature region may be performed at a second resolution higher than the first resolution at which the first pre-processing for setting the first feature region is applied.
[0028] According to the present invention, the processor can set the second feature area based on the probability that an object for identifying the pet is located in the first feature area according to the species of the pet.
[0029] According to the present invention, if the second feature value is greater than a reference value, an image including the second feature region may be transmitted to the server.
[0030] According to the present invention, the processor can generate candidate feature regions for determining the species of the pet in the image, and generate a first feature region whose position and size are determined based on the reliability values of each of the candidate feature regions. Effect of the Invention
[0031] The method and electronic device for detecting an object for identifying a pet according to the present invention can effectively obtain an image of an object corresponding to the pet's nose for learning or identification by immediately selecting an image for learning or identifying a nose print after photographing the pet and storing the image in a server database.
[0032] In addition, the method and electronic device for detecting objects for pet identification according to the present invention can reduce computational complexity by first determining the species of the pet and then extracting the pet's nose print image.
[0033] According to the present invention, in the process of determining a feature region for determining the species of a pet, a final feature region of a larger area is generated by taking into account the reliability values of each of a number of feature region candidates, thereby enabling more accurate detection by subsequently detecting an object for identifying the pet within the final feature region.
[0034] According to the present invention, by inspecting the quality of an object image for identifying a pet, such as a dog's nose, in a captured image, it is possible to determine whether the image is suitable for artificial intelligence-based learning or identification, and only suitable images can be saved to optimize a neural network for learning or identification.
[0035] The effects of the present invention are not limited to those described above, and other effects not described above will be clearly understood by those skilled in the art from the following description. [Brief description of the drawings]
[0036] [Figure 1] A schematic procedure for AI-based pet management is shown. [Diagram 2] 1 shows a procedure for AI-based pet nose print management to which the suitability judgment of object images for learning or identification according to the present invention is applied. [Diagram 3] FIG. 2 is a diagram showing a procedure for detecting objects for pet identification in the pet management system according to the present invention. [Figure 4] 1 shows an example of a UI (User Interface) screen for detecting an identification object of a pet to which the present invention is applied. [Diagram 5] 4 illustrates a process of detecting objects for pet identification according to the present invention. [Figure 6] 4 shows a process for setting feature regions according to the present invention. [Figure 7] 4 is a flow chart illustrating a process for detecting an object for pet identification according to the present invention. [Figure 8]4 shows a process of deriving feature regions for determining the species of a pet according to the present invention. [Figure 9] 4 is a flow chart illustrating a process for processing an image of an object for pet identification according to the present invention. [Figure 10] Here is an example of the resulting image after applying the Canny edge detector to the input image. [Figure 11] 1 shows an example of a pattern shape of pixel blocks in which boundaries are located, which is used to determine whether or not there is blur in an image as a result of applying a Canny boundary detector. [Figure 12] 1 is a flow chart of a method for filtering images of objects for pet identification. [Figure 13] 1 is a block diagram of an electronic device according to the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0037] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will now be described in detail with reference to the accompanying drawings, in which: FIG. 1 is a block diagram of a semiconductor device according to an embodiment of the present invention;
[0038] In order to clearly describe the present invention, parts that are not relevant to the description will be omitted, and the same reference numerals will be used throughout the specification to refer to the same or similar components.
[0039] Furthermore, in some embodiments, components having the same configuration will be described only in the representative embodiment using the same symbols, and in other embodiments, only configurations that differ from the representative embodiment will be described.
[0040] Throughout the specification, when a part is described as being "connected (or coupled)" to another part, this includes not only the case where the part is "directly connected (or coupled)" to another part, but also the case where the part is "indirectly connected (or coupled)" via another member. Furthermore, when a part is described as "comprising" a certain component, this means that the part can further include the other component, not excluding the other component, unless otherwise specified.
[0041] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention pertains. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with the contextual meaning of the relevant art, and should not be interpreted in an ideal or overly formal sense unless expressly defined in this application.
[0042] This document will mainly explain how to extract identification information by utilizing the shape of wrinkles on a dog's nose (nose print), but in this invention, the range of pets is not limited to dogs, and the features used as identification information are not limited to nose prints, and various physical features of pets can be used.
[0043] As mentioned above, there are not enough pet nose print images suitable for AI-based learning or identification, and the quality of pet nose print images is likely to be low, so nose print images need to be selectively stored in a database for AI-based learning or identification.
[0044] 2 shows the procedure for AI-based pet nose print management to which the suitability judgment of the nose print image for learning or identification according to the present invention is applied. After photographing the nose print of the pet, the present invention first judges whether the photographed nose print image is suitable as data for AI-based learning or identification, and if it is judged to be suitable, it transmits and stores it in a server for AI-based learning or recognition and uses it as data for subsequent learning or identification.
[0045] As shown in FIG. 2, the nose print management procedure according to the present invention includes a nose print acquisition procedure and a nose print recognition procedure.
[0046] According to the present invention, when registering a new nose print of a pet, an image including the pet is photographed, and then a nose print image is extracted from the face area of the pet, and in particular, it is first determined whether the nose print image is suitable for identifying or learning the pet. If the photographed image is determined to be suitable for identification or learning, the image is transmitted to a server (artificial intelligence neural network) and stored in a database.
[0047] In the case of searching for identification information of a pet through a nose print, an image including the pet is photographed, and then a nose print image is extracted from the face area of the pet, and it is first determined whether the nose print image is suitable for identifying or learning the pet. If it is determined that the photographed image is suitable for identification or learning, the image is transmitted to a server, and the identification information of the pet is extracted through matching with a pre-stored nose print image.
[0048] In the case of the nose print registration procedure, as shown in FIG. 2(a), a pet is photographed (S205), a face region (hereinafter referred to as a first feature region) is detected from the photographed image of the pet (S210), and an area occupied by the nose within the face region (hereinafter referred to as a second feature region) is detected. A nose print image is output through a quality inspection to determine whether the photographed image is suitable for learning or identification (S215), and the output image is transmitted to a server constituting an artificial neural network for storage and registration (S220).
[0049] In the case of the nose print inquiry procedure, as shown in Fig. 2(b), a pet is photographed (S230), a face region is detected from the pet image (S235), the area occupied by the nose is detected within the face region, and a nose print image is output through a quality inspection to see if the photographed image is suitable for learning or identification (S240), which is similar to the nose print registration procedure. Subsequent procedures include a process of comparing the output nose print image with previously stored and learned nose print images to search for matching information (S245), and a process of outputting the search results (S250).
[0050] FIG. 3 shows a procedure for detecting an object corresponding to a pet's nose in the pet's nose print management system according to the present invention.
[0051] Referring to FIG. 3, first, an initial image is generated by photographing a pet (S305), and a step of detecting a face area from the initial image is performed first (S310). Then, a step of detecting a nose area in the face area by considering the species of the pet is performed (S315). The first detection of the face area and the second detection of the nose area can reduce the computational complexity and improve the detection accuracy compared to detecting the nose area by considering all species by cascaded detection. Then, a quality check is performed to check whether the image of the detected nose area is suitable for future identification or learning of a nose print (S320), and if it is determined to be a suitable image as a result of the quality check, the image can be transmitted to a server and used for identifying a nose print, or can be stored for future learning or identification (S325).
[0052] In addition, according to the present invention, the camera can be controlled to focus on the detected nose area so that the image of the object for identifying the pet, such as the wrinkles on a dog's nose (nose print), is not blurred (S330). This is to prevent the quality of the image from deteriorating due to the nose being out of focus by focusing the camera on the nose area.
[0053] Fig. 4 shows an example of a UI (User Interface) screen for acquiring a nose print image of a pet to which the present invention is applied. Fig. 4 shows a case in which the nose print of a dog, one of various pets, is acquired.
[0054] Referring to FIG. 4, the pet species is identified from the image being captured to determine whether the pet currently being captured is a dog. If the pet being captured is not a dog, a message such as "Dog not found" is output as shown in FIG. 4(a), and if the pet being captured is a dog, a procedure is performed to obtain the dog's nose print. In order to determine whether the pet being captured is a dog, the face area of the pet contained in the image is first extracted, and the image contained in the face area is compared with existing learned data to determine the species of the pet.
[0055] Thereafter, as shown in (b) to (e) of FIG. 4, an area corresponding to the dog's nose is set on the dog's face, and then the area corresponding to the nose is focused on for photographing. That is, the camera can be controlled to focus on the position (center point) of the area corresponding to the object for identifying the pet. Also, in order to provide feedback to the user that the image is being photographed with the focus on the object (e.g., nose) currently being tracked, a graphic element can be overlayed on the position of the object being tracked. By displaying a graphic element indicating the detection state of the object at the position of the object being tracked, the user can recognize that object recognition is being performed on the pet currently being photographed.
[0056] 4(b) to (e), if the image quality of the currently captured object is good (if the image quality of the object satisfies the reference condition), a first graphic element 410a (e.g., a smiling icon or a green icon) indicating a good quality state may be overlaid on the object and output. If the image quality of the currently captured object is bad (if the image quality of the object does not satisfy the reference condition), a second graphic element 410b (e.g., a crying icon or a red icon) indicating a bad quality state may be overlaid on the object and output.
[0057] As shown in Fig. 4, even if the dog continues to move, the nose of the dog can be tracked and photographed while focusing on the nose. In this case, in each photographed image, it may be determined whether the dog's nose print image is suitable for identifying or learning the pet, and the degree of suitability may be output.
[0058] For example, the degree to which a photographed dog's nose print image is suitable for identifying or learning about a pet can be calculated as a numerical value, and score information 420 can be output in a form in which a gauge is filled toward "Bad" when the suitability is low and toward "Good" when the suitability is high according to the numerical value for the suitability. In other words, score information 420 indicating the photographing quality of an object in an image can be output.
[0059] Also, a quality evaluation (size, brightness, clarity, etc.) of the currently captured nose print image can be performed, and a message 430 can be output to provide feedback to the user so that a nose print image suitable for recognition or learning based on artificial intelligence can be captured. For example, if the size of the dog's nose print image is smaller than a reference value, a message such as "Please adjust the distance of the dog's nose" can be output so that a larger size nose print image can be captured, as shown in (c) of FIG. 4. Also, progress rate information 440 indicating the progress of acquiring an image of an object having a quality suitable for identifying a pet can be output. For example, if four nose print images having suitable quality are required and one suitable image has been acquired so far, progress rate information 440 indicating that the progress rate is 25% can be output as shown in FIG. 4.
[0060] When a sufficient number of dog's muzzle images have been acquired, the shooting is terminated and the dog's identification information can be stored in a database together with the dog's muzzle image, or the dog's identification information can be output.
[0061] In the present invention, the face area of the pet is first detected, and then the nose area is detected within the face area. This is to reduce the difficulty of object detection while reducing the computational complexity. In the process of capturing an image, an object other than the object to be detected, or unnecessary or erroneous information may be included in the image. Therefore, the present invention first determines whether a desired object (pet's nose) is present in the image being captured.
[0062] In addition, to identify a pet's nose print, an image with a certain level of resolution is required, but the higher the image resolution, the greater the amount of calculation required for image processing.In addition, as the number of pet types increases, the learning method differs for each pet type, which increases the difficulty of the AI calculations.In particular, since animals of similar types have similar shapes (e.g., a dog's nose and a wolf's nose are similar), it can be very difficult to classify the noses of similar animals along with their animal types.
[0063] Therefore, the present invention uses a cascaded object detection method to reduce such computational complexity. For example, while photographing a pet, the pet's face area is first detected, the pet's type is identified, and the pet's nose area is detected based on the detected pet's face area and the identified pet's type. This is done by first identifying the pet's type at a low resolution with relatively low computational complexity, and then applying an object detection method determined according to the pet's type to maintain high resolution in the pet's face area and detect the nose area. Therefore, the present invention can effectively detect the pet's nose area while relatively reducing computational complexity.
[0064] FIG. 5 shows the overall image processing procedure for pet identification according to the present invention. As shown in FIG. 5, the method for processing an input image according to the present invention includes a step of receiving an input image from a camera (S505), a first pre-processing step of adjusting the size of the input image to generate a first processed image (S510), a first feature region detection step of detecting the position and species of an animal from the processed image generated in the first pre-processing step (S515), a first post-processing step of extracting a first feature value of an animal image from a result of the first feature region detection step (S520), a step of determining a detector for detecting an object (e.g., nose) for identifying a pet according to the species of the pet from the image processed through the first post-processing step (S525), a second pre-processing step of adjusting the size of the image for image processing for identifying a pet, at least one second feature region detection step (S535) corresponding to each species of animal that can be detected in the first feature detection step, and a second post-processing step (S540) of extracting a second feature value of an animal image corresponding to each second feature region detection step.
[0065] First pretreatment step
[0066] The step of applying the first pre-processing to the original image (S510) is a step of adjusting the size, ratio, direction, etc. of the original image to convert the image into a form suitable for object detection.
[0067] With the development of camera technology, most input images are composed of millions to tens of millions of pixels, and it is not desirable to directly process such large images. In order for object detection to work efficiently, a pre-processing process must be performed to make the input image suitable for processing. Mathematically, this process is performed by coordinate system transformation.
[0068] It is clear that any processed image can be generated by associating any four points in the input image with four vertices of the processed image and going through any coordinate system transformation process. However, when any nonlinear transformation function is used in the coordinate system transformation process, it is necessary to be able to perform an inverse transformation to obtain the feature region of the input image from the bounding box obtained as a result of the feature region detector. For example, it is preferable to use an affine transformation, which linearly transforms any four points of the input image by associating them with four vertices of the processed image, since the inverse transformation process can be easily obtained.
[0069] As an example of a method for determining any four points in the input image, a method of using the four vertices of the input image as they are can be used. Alternatively, a method of adding a margin to the input image or cutting out a part of the input image can be used so that the horizontal length and vertical length are converted at the same ratio. Alternatively, various interpolation methods can be applied to reduce the size of the input image.
[0070] First feature region detection step
[0071] The purpose of this step is to first detect the area where a pet is present and the species of that animal in the preprocessed image, thereby setting a first feature region that can be used in the second feature region detection step described below, and also to improve the final feature point detection performance by selecting a second feature region detector optimized for each pet species.
[0072] In this process, any object detection and classification method can be easily combined by anyone with ordinary knowledge in the related field. However, since it is known that a method based on an artificial neural network has superior performance compared to conventional methods, it is preferable to use a feature detection technique based on an artificial neural network as much as possible. For example, a feature detector of the SSD (Single Vision-Shot Multibox Detection) type, which is an algorithm for detecting objects of various sizes in one image, can be used in the artificial neural network.
[0073] The input image normalized by the above-mentioned preprocessor is hierarchically constructed into the first feature image through the n-th feature image by the artificial neural network. At this time, the method of extracting the feature image for each layer can be mechanically learned in the learning step of the artificial neural network.
[0074] The hierarchical feature image extracted in this way is combined with a predefined box (Priori Box) list corresponding to each layer to generate a bounding box, individual type, and confidence value list. This calculation process can also be mechanically learned in the learning step of the artificial neural network. For example, the result value is returned in the format shown in Table 1 below. At this time, the number of types that the neural network can determine is determined in the neural network design step, and the case where no object exists, i.e., "background", is implicitly defined.
[0075] [Table 1]
[0076] These result boxes are merged with overlapping result boxes by the NMS (Non-Maximum Suppression) step, and are finally returned as the object detection result present in the image. NMS is a process of deriving the final feature region from multiple feature region candidates, and the feature region candidates can be generated by taking into account the probability values in Table 1 according to the procedure shown in Figure 6.
[0077] This process will now be described in detail.
[0078] 1. For each species, excluding the background, perform the following steps:
[0079] A. In the bounding box list, filter out boxes whose probability of being the species is below a certain threshold. If no boxes remain, exit with no results.
[0080] B. In the bounding box list, the box with the highest probability of being the species is designated as the first box (first bounding region) and removed from the bounding box list.
[0081] C. For the remaining bounding box list, perform the following process in descending order of probability.
[0082] i. Calculate the IoU (Intersection over Union, the area ratio of the intersection set to the union set) with the first box.
[0083] ii. If the IoU is higher than a certain threshold, this box is a box that overlaps with the first box. Merge it with the first box.
[0084] D. Add the first box to the results box list.
[0085] E. If any boxes remain in the bounding box list, repeat from step C for the remaining boxes.
[0086] For two boxes A and B, the IoU can be effectively calculated as follows:
[0087] [Formula 1]
number
[0088] That is, according to the present invention, the step of generating candidate feature regions includes the steps of selecting a first boundary region (first box) from the characteristic image that has the highest probability of corresponding to a specific animal species, and calculating the area ratio of intersection to union (IoU) with the first boundary region for the remaining boundary regions excluding the selected boundary region (first box) from the characteristic image in accordance with the order of probability values, and including in the candidate feature regions of the characteristic image boundary regions whose area of intersection to union is greater than a standard area ratio.
[0089] Next, a method of merging boxes overlapping with the first box in the above process will be described. For example, the first box can be merged by keeping it as it is and deleting the second box from the bounding box list (Hard NMS). Alternatively, the first box can be merged as it is, and the probability that the second box is a specific species is reduced by weighting it by a value between (0, 1), and only if the reduced result value is smaller than a specific threshold, it is removed from the bounding box list (Soft NMS).
[0090] As one embodiment proposed by the present invention, a new method (Expansion NMS) can be used to merge the first box (first feature region candidate) and the second box (first feature region candidate) according to a probability value, as shown in the following Equation 2.
[0091] [Formula 2]
number
[0092] In this case, p 1 , p 2are the probability values of the first box (first feature region candidate) and the second box (first feature region candidate), respectively, and C (x、y) 1 , C (x、y) 2 , C (x、y) n are the (x, y) coordinates of the center points of the first box, the second box, and the merged box, respectively. In a similar way, W1, W2, and W n are the widths of the first box, the second box, and the merged box, respectively, and H 1 , H 2 , H n represents the vertical height. The probability value of the merged box can use the probability value of the first box. The first feature region derived by the extended NMS according to the present invention is determined by considering the confidence value of a specific seed being located in each feature region candidate.
[0093] That is, the center point C where the first feature region is located (x、y) n is the center point C of the feature region candidate as shown in Equation 2. (x、y) 1 , C (x、y) 2 Confidence value p for 1 , p 2 It can be determined by a weighted sum of
[0094] In addition, the width W of the first feature region n is the reliability value p for the widths W1 and W2 of the feature region candidates as shown in Equation 2. 1 , p 2 The height of the first feature region H is determined by the weighted sum of n is the height H of the feature region candidate 1 , H 2 Confidence value p for 1 , p 2 It can be determined by a weighted sum of
[0095] By generating a new box according to the above embodiment, a box with a larger width is obtained as compared with the existing Hard-NMS or Soft-NMS methods. According to the present embodiment, in the pre-processing detection for performing the multi-stage detector, a configuration is possible in which a certain margin is added, but by using the expansion NMS according to the present invention, such a margin can be adaptively determined.
[0096] FIG. 8 shows an example of detecting a pet's feature region by applying the extended NMS according to the present invention in comparison with an existing NMS. FIG. 8(a) shows a plurality of feature region candidates generated in an original image, FIG. 8(b) shows an example of a first feature region derived by the existing NMS, and FIG. 8(c) shows an example of a second feature region derived by applying the extended NMS according to the present invention. As shown in FIG. 8(b), the existing NMS (Hard NMS, Soft NMS) selects one box (feature region candidate) with the highest reliability from among a plurality of boxes (feature region candidates), so there is a possibility that a region necessary for acquiring a muzzle pattern, such as a nose region, may be missed in the subsequent second feature region detection process.
[0097] Therefore, the present invention applies a weighted average based on the reliability value to multiple boxes (candidate feature regions), sets one box with a large width and height as the first feature region (pet's face region) as shown in (c) of Fig. 8, and detects a second feature region (nose region) for identifying the pet within the first feature region. By setting a more expanded first feature region as in the present invention, it is possible to reduce the occurrence of errors in which the second feature region is not detected in the subsequent steps.
[0098] Finally, it is obvious that the feature region in the original image can be obtained by performing an inverse transformation process for any transformation process used in the preprocessing step on one or more bounding boxes determined in this way. Depending on the configuration, a certain amount of margin can be added to the feature region in the original image to adjust it so that the second detection step described later can be performed well.
[0099] First post-processing step
[0100] A first feature value can be generated by performing an additional post-processing step for each feature region of the input image acquired in the above-mentioned first feature region setting step (S515). For example, in order to acquire brightness information (first feature value) for the first feature region of the input image, a calculation such as the following Equation 3 can be performed.
[0101] [Formula 3]
number
[0102] In this case, L is the luma value according to the BT.601 standard, V is the brightness value defined in the HSV color space, and M and N are the width and height of the target feature region.
[0103] Using the additionally generated first feature value in this manner, it is possible to predict whether the first feature region obtained in the first feature region detection step (S515) is suitable for use in an application field combined with this patent. It is obvious that the additionally generated first feature value must be appropriately designed according to the application field. If the condition of the first feature value defined in the application field is not satisfied, the system can be configured to selectively omit the second feature region setting and object detection steps described below.
[0104] Second feature region detection step
[0105] The purpose of this step is to extract feature regions specifically required for the application field from the area where the animal exists. For example, in the application field of detecting the positions of the eyes, nose, mouth, and ears from the face area of an animal, the first feature area detection step first separates the face area of the animal from the species information of the animal, and the second feature area detection step detects the positions of the eyes, nose, mouth, and ears according to the species of the animal.
[0106] In this process, the second feature region detection step can be composed of a plurality of independent feature region detectors specialized for each animal species. For example, if dogs, cats, and hamsters can be distinguished in the first feature region detection step, it is preferable to provide three second feature region detectors, each designed to be specialized for dogs, cats, and hamsters. This reduces the number of types of features to be learned by the individual feature region detectors, thereby reducing the learning complexity, and it is self-evident that neural network learning is possible with only a smaller number of numerical data in terms of learning data collection.
[0107] Since each second feature region detector is configured independently of the others, a person having ordinary knowledge can easily configure an independent individual detector. It is preferable that each feature region detector is configured individually according to the feature information to be detected for each species. Alternatively, in order to reduce the complexity of the system configuration, a method can be used in which some or all of the second feature region detectors share a feature region detector with the same structure, but the learning parameter values are changed to configure the system to be suitable for each species. Furthermore, a method can be considered in which the feature region detector with the same structure as the first feature region detection step is used as the second feature region detector, but only the learning parameter values and the NMS method are changed to further reduce the system complexity.
[0108] For one or more feature regions set via the first feature region detection step and the first post-processing step, a determination is made as to which second feature region detector to use using seed information detected in the first feature region detection step, and the second feature region detection step is performed using the determined second feature region detector.
[0109] First, a pre-processing step is performed. At this time, it is obvious that a conversion step capable of inversion must be used in the process of converting coordinates. In the second pre-processing step, the first feature region detected in the input image must be converted into an input image for the second feature region detector, so that the four points required for designing the conversion function are preferably defined as the four vertices of the first feature region.
[0110] Since the second feature region obtained through the second feature region detector is a value detected using the first feature region, the first feature region must be taken into consideration when calculating the second feature region within the entire input image.
[0111] A second feature value can be generated by performing an additional post-processing step similar to the first post-processing step on the second feature region obtained through the second feature region detector. For example, a Sobel filter can be applied to obtain image sharpness, or information such as the posture of the animal to be detected can be obtained using the presence or absence of detection and the relative positional relationship between feature regions. Furthermore, image quality inspections (e.g., focus blur, motion blur) can be performed as described below.
[0112] Using the second feature value additionally generated in this manner, it is possible to predict whether the feature region acquired in the second object detection step is suitable for use in the application field to which this patent is applied. It is obvious that the second feature value additionally generated must be appropriately designed according to the application field. If the condition of the second feature value defined in the application field is not satisfied, it is preferable to design the second feature value so that data suitable for the application field can be acquired, for example, by excluding not only the second detection area but also the first detection area from the detection result.
[0113] System Expansion
[0114] In the present invention, an example of a system and configuration method is given in which a two-step detection step is configured to detect the location and species of an animal in a first feature location detection step, and a detector to be used in a second feature location detection step is selected based on the results.
[0115] Such a cascade configuration can be easily expanded to a multi-layer cascade configuration, for example, where the first feature position detection step detects the entire body of an animal, the second feature position detection step detects the positions of the face and arms and legs of the animal, and the third feature position detection step detects the positions of the eyes, nose, mouth, and ears from the face.
[0116] By using such a multi-layered hierarchical structure, it is easy to design a system capable of simultaneously acquiring feature positions of multiple layers. When designing a multi-layered hierarchical system, it is obvious that the number of layers can be determined by considering the hierarchical domain of the feature positions to be acquired, the operation time and complexity of the entire system, and the resources required to configure each individual feature region detector, and thus an optimal hierarchical structure can be designed.
[0117] FIG. 7 is a flowchart of a method for detecting an object corresponding to a pet's nose in a pet's nose print management system according to the present invention.
[0118] A method for detecting an object for identifying a pet according to the present invention includes a step of acquiring an original image including a pet (e.g., a dog) (S710), a step of determining a first feature region and the species of the pet through image processing on the original image (S720), and a step of detecting an object for identifying the pet (e.g., a nose) within the first feature region based on the determined species of the pet (S730).
[0119] In step S710, an original image including a pet is acquired through an activated camera while an application for pet object recognition is being executed. Here, illumination and focus can be adjusted so that the pet can be photographed smoothly. The acquired image may be provided as an input image for Figs. 5 and 6. Thereafter, as described above, a step of determining the species of the pet (S720) and a step of detecting the pet object (S730) can be performed for cascaded object detection.
[0120] In step S720, a procedure for identifying the pet species is performed. According to the present invention, the step of determining the pet species (S720) can include the steps of applying a first pre-processing to the original image, identifying the pet species in the pre-processed image and setting the first feature region, and extracting a first feature value through a first post-processing to the first feature region.
[0121] The step of applying the first pre-processing to the original image is a step of adjusting the size, ratio, orientation, etc. of the original image to convert the image into a form suitable for object detection, as described above with reference to step S510 of FIG. 5.
[0122] The step of setting a first feature region is a step of detecting the area in the image where a pet is present and the species of the pet, and is intended to set a first feature region that can be used in the second feature region detection step described below, as well as to improve the final feature point detection performance by selecting a second feature region detector optimized for each pet species.
[0123] According to the present invention, the step of setting the first feature region includes the steps of dividing the preprocessed image into a plurality of feature images using a learning neural network, applying a predefined bounding box to each of the plurality of feature images, calculating a probability value for each pet type within the bounding box, and configuring the first feature region to include the bounding box if the calculated probability value for a specific animal species is equal to or greater than a reference value.
[0124] As described above, the input image normalized by the preprocessor is hierarchically constructed from the first feature image to the nth feature image by the artificial neural network. At this time, the method of extracting feature images for each layer can be mechanically learned in the learning step of the artificial neural network.
[0125] The hierarchical feature image extracted in this way is combined with a predefined bounding box (Priori Box) list corresponding to each layer to generate a list of bounding boxes, object types, and confidence values (probability values), and the result can be output in the format shown in Table 1.
[0126] Thereafter, if the probability value of the specific animal type in a specific bounding box is equal to or greater than a reference value, the first feature region is set so that the bounding box is included in the first feature region.
[0127] Meanwhile, the process for determining the pet's face region (first feature region) as described above can be performed at a relatively low resolution since a high resolution is not required. That is, the step of applying the first pre-processing to the original image can include a step of converting the original image into an image with a first resolution lower than the original resolution, and a step of applying the first pre-processing to the image converted to the first resolution.
[0128] Meanwhile, when the first feature region for identifying the species of the pet is set, a first feature value is extracted through a first post-processing for the first feature region in order to determine in advance whether the dog's nose print image extracted from the acquired image is suitable as data to be used for learning or identification in the future.
[0129] That is, if the first feature value is greater than the reference value, object detection for identifying the pet is performed, and if the first feature value is less than the reference value, no additional processing is performed and processing for another image is performed. The first feature value may vary depending on the embodiment, and may be, for example, brightness information of the image to be processed.
[0130] In step S730, detection of an object for identifying the pet is performed. Although various parts such as eyes, nose, mouth, and ears can be used as objects for identifying the pet, the nose will be described as a representative example for using a nose print. This step is performed in consideration of the species of the pet previously identified. If the pet is a dog, object detection for identification optimized for dogs can be performed. The optimized object detection can differ depending on the type of animal. Furthermore, if the captured image contains multiple types of pets, object detection for identification can be performed for each animal.
[0131] The step of detecting an object for pet identification may include the steps of applying a second pre-processing to a first feature region for identifying a species of pet, setting a second feature region for identifying the pet based on the species of pet in the second pre-processed first feature region, and applying a second post-processing to the second feature region.
[0132] The second pre-processing for detecting an object for identifying a pet is a process of adjusting the size of an image, similar to the first pre-processing. The second pre-processing for the first feature region may be performed at a second resolution higher than the first resolution to which the first pre-processing is applied. This is because, unlike the process of determining the type of animal, the process of detecting an object (e.g., a nose) for identifying a pet and inspecting the identification data (nose print image) requires a relatively high quality image. Thereafter, a second feature region is set as an object for identifying a pet in the pre-processed image.
[0133] The step of setting the second feature region includes a step of setting the second feature region (e.g., nose region) based on a probability that an object for identifying the pet (e.g., nose) is located in the first feature region (e.g., face region) according to the species of the pet. When the species of the pet is previously determined in step S720, an individual feature region detector and parameters optimized according to the species are selected, and the selected detector and parameters can be used to detect the object for identifying the pet (e.g., nose region) with smaller computational complexity.
[0134] Post-processing can be performed to check whether the image of the pet identification object detected as the second feature region is suitable for use in subsequent learning or identification. As a result of the post-processing, a second feature value indicating the suitability of the image can be derived. If the second feature value is greater than a reference value, the image including the second feature region is transmitted to a server.
[0135] 13 is a block diagram of an electronic device 1300 in accordance with the present invention. The electronic device 1300 in accordance with the present invention may include a camera 1310, a processor 1320, a communication module 1330, a memory 1340, and a display 1350.
[0136] The camera 1010 may include an optical module such as a lens and a Charge-Coupled Device (CCD) or Complementary Metal-Oxide Semiconductor (CMOS) that generates an image signal from input light, and may generate image data through image capture and provide it to the processor 1320.
[0137] The processor 1320 controls each module of the electronic device 1300 and performs calculations necessary for image processing. The processor 1320 can be composed of multiple microprocessors (processing circuits) depending on the function. As described above, the processor 1330 can detect an object (e.g., a nose) for identifying a pet (e.g., a dog) and determine the validity of an image for the object.
[0138] The communication module 1330 can transmit or receive data to or from an external entity via a wired / wireless network. In particular, the communication module 1330 can exchange data for artificial intelligence-based processing via communication with a server for learning or identification.
[0139] Furthermore, the electronic device 1300 can include various modules depending on the application, including a memory 1340 for storing image data and information required for image processing, and a display 1350 for outputting images to a user.
[0140] The electronic device 1300 according to the present invention includes a camera 1310 for generating an original image including a pet, a processor 1320 for determining a first feature area and the species of the pet through image processing of the original image and detecting an object for identifying the pet within the first feature area based on the determined species of the pet, and a communication module 1330 for transmitting an image of the object to a server if the object for identifying the pet is valid.
[0141] According to the present invention, the processor 1320 can apply a first pre-processing to the original image, determine the species of the pet in the pre-processed image to set a first feature region, and extract a first feature value by a first post-processing to the first feature region.
[0142] According to the present invention, the processor 1320 can use a learning neural network to generate a plurality of feature images from the preprocessed image, apply predefined bounding boxes to each of the plurality of feature images, calculate a probability value for each pet type within the bounding boxes, and configure a first feature region to include the bounding box if the calculated probability value for a particular animal species is equal to or greater than a reference value.
[0143] According to the present invention, if the first feature value is greater than the reference value, object detection for identifying the pet may be performed, and if the first feature value is less than the reference value, additional processing may be omitted.
[0144] According to the present invention, the processor 1320 may convert an original image into an image of a first resolution lower than the original resolution, and apply a first pre-processing to the image converted into the first resolution.
[0145] According to the present invention, the processor 1320 can apply a second pre-processing to a first feature region for identifying the species of the pet, set a second feature region for identifying the pet based on the species of the pet in the second pre-processed first feature region, and apply a second post-processing to the second feature region to extract a second feature value.
[0146] According to the present invention, the second pre-processing on the first feature region can be performed at a second resolution higher than the first resolution at which the first pre-processing for setting the first feature region is applied.
[0147] According to the present invention, the processor 1320 can set the second feature region based on the probability that an object for identifying the pet is located in the first feature region according to the species of the pet.
[0148] According to the present invention, if the second feature value is greater than the reference value, an image including the second feature region can be transmitted to a server.
[0149] According to the present invention, the processor 1320 can generate candidate feature regions for determining the species of a pet in an image, and generate a first feature region whose position and size are determined based on the reliability value of each of the candidate feature regions.
[0150] FIG. 9 is a flow chart of a method for processing images of an object for pet identification.
[0151] A method for processing an image of an object for identifying a pet according to the present invention may include the steps of acquiring an image including a pet (S910), generating candidate feature regions for determining the species of the pet in the image (S920), setting a first feature region whose position and size are determined based on a reliability value of each of the candidate feature regions (S930), setting a second feature region including an object for identifying the pet in the first feature region (S940), and acquiring an image of the object in the second feature region (S950).
[0152] According to the present invention, the step of generating candidate feature regions may include the steps of hierarchically generating feature images using an artificial neural network, applying predefined boundary regions to each of the feature images to calculate a probability value that a particular type of pet is located in each boundary region, and generating candidate feature regions taking the probability values into consideration.
[0153] The input image normalized by the preprocessor is used to hierarchically generate the first to nth feature images by the artificial neural network, and the method of extracting feature images for each layer can be mechanically learned in the learning step of the artificial neural network.
[0154] The extracted hierarchical feature image is combined with a list of predefined boundary regions (boundary boxes) corresponding to each layer, and a list of probability values that a specific animal type is located in each boundary region is generated as shown in Table 1. Here, if it is not possible to determine whether an animal type is a specific animal type, it can be defined as "background."
[0155] Next, the extended NMS according to the present invention is applied to generate first feature region candidates (feature region candidates) for species identification, such as pet faces, in each feature image. Each feature region candidate can be derived using the probability values for specific animal species derived earlier.
[0156] According to the present invention, the step of generating candidate feature regions can include the steps of: selecting a first boundary region from the feature image that has the highest probability of corresponding to a specific animal species; calculating the overlapping degree with the first boundary region according to the order of probability values for the remaining boundary regions except for the first boundary region selected from the feature image; and including the boundary regions having the overlapping degree greater than a reference overlapping degree in the candidate feature regions of the feature image. In this case, for example, the area ratio of the intersection to the union between the two boundary regions can be used to evaluate the overlapping degree.
[0157] That is, feature region candidates like those shown in FIG. 8(a) can be generated by the following procedure.
[0158] 1. For each species, excluding the background, carry out the following steps:
[0159] A. From the bounding box list, remove any boxes whose probability of being the species is below a certain threshold. If no boxes remain, exit with no results.
[0160] B. In the above bounding box list, the box with the highest probability of being the relevant species is designated as the first box (first bounding region) and removed from the bounding box list.
[0161] C. For the remaining bounding box list, perform the following process in order of decreasing probability.
[0162] i. Calculate the overlap with the first box. For example, the intersection over union area ratio can be used.
[0163] ii. If the overlap value is higher than a certain threshold, the box is a box that overlaps with the first box and is merged with the first box.
[0164] D. Add the first box to the results box list.
[0165] E. If any boxes remain in the bounding box list, repeat from step C for the remaining boxes.
[0166] For two boxes A and B, for example, the intersection to union area ratio can be efficiently calculated as in Equation 1 above.
[0167] As described with reference to FIG. 8, a first feature region (eg, a pet's face region) can be derived from each of the feature region candidates derived according to the present invention based on the reliability value of each feature region candidate.
[0168] According to the present invention, the center point (C (x、y) n ) is the center point C of the feature region candidate as shown in Equation 2. (x、y) 1 , C (x、y) 2 Confidence value p for 1 , p 2 It can be determined by a weighted sum of
[0169] According to the present invention, the width W of the first feature region n is the reliability value p for the widths W1 and W2 of the feature region candidates as shown in Equation 2. 1 , p 2 The height of the first feature region H is determined by the weighted sum of n is the height H of the feature region candidate 1 , H 2 Confidence value p for1 , p 2 It can be determined by a weighted sum of
[0170] The present invention applies a weighted average based on the reliability value to multiple boxes (candidate feature regions) to set one box with a large width and height as the first feature region (pet's face region), and can detect a second feature region (nose region) for identifying the pet within the first feature region. By setting a more expanded first feature region as in the present invention, it is possible to reduce the occurrence of errors in which the second feature region is not detected in the subsequent steps.
[0171] Thereafter, detection is performed for a second feature region (e.g., nose region) for identifying the pet within the first feature region (e.g., face region of the dog). This step is performed taking into consideration the species of the pet as described above. If the pet is a dog, object detection for identification optimized for dogs can be performed. The optimized object detection may differ depending on the type of animal.
[0172] The step of setting the second feature region includes a step of setting the second feature region (e.g., nose region) based on a probability that an object for identifying the pet (e.g., nose) is located in the first feature region (e.g., face region) depending on the species of the pet.
[0173] In addition, post-processing can be performed to check whether the image of the pet identification object detected as the second feature region is suitable for use in future learning or identification. As a result of the post-processing, a second feature value representing the suitability of the image can be derived. If the second feature value is greater than a reference value, the image including the second feature region is transmitted to a server, and if the second feature value is less than the reference value, the image including the second feature region is discarded.
[0174] The electronic device 1300 according to the present invention includes a camera 1310 for generating an image including a pet, and a processor 1330 for processing the image provided by the camera 1320 to generate an image of an object for identifying the pet. The processor 1330 generates feature region candidates for determining the species of the pet in the image, sets a first feature region whose position and size are determined based on the reliability value of each of the feature region candidates, sets a second feature region including an object for identifying the pet in the first feature region, and obtains an image of the object from the second feature region.
[0175] According to the present invention, the processor 1320 can hierarchically generate a plurality of feature images from the image using an artificial neural network, apply predefined boundary regions to each of the feature images to calculate a probability value that a particular type of pet is located in each boundary region, and generate the feature region candidates taking into account the probability values.
[0176] According to the present invention, the processor 1320 can select a first boundary region from the feature image that has the highest probability of corresponding to a specific animal species, calculate the overlapping degree with the first boundary region for the remaining boundary regions from the feature image excluding the selected boundary region in accordance with the order of the probability values, and include the boundary regions having the overlapping degree greater than a reference overlapping degree as candidate feature regions of the feature image. In this case, the area ratio of the intersection to the union between two boundary regions can be used to calculate the overlapping degree.
[0177] According to the present invention, the center point at which the first feature region is located can be determined by a weighted sum of the reliability values for the center points of the candidate feature regions.
[0178] According to the present invention, the width of the first feature region is determined by a weighted sum of the reliability values for the widths of the candidate feature regions, and the height of the first feature region can be determined by a weighted sum of the reliability values for the heights of the candidate feature regions.
[0179] In accordance with the present invention, the processor 1320 can detect the changed position of the object from the next image, determine whether the image of the object with the changed position in the next image is suitable for artificial intelligence based learning or identification, and control the camera 1310 to take the next capture with the focus set at the changed position.
[0180] According to the present invention, the processor 1320 can set a first feature region for determining the species of the pet from the image, and set a second feature region within the first feature region that includes an object for identifying the pet.
[0181] According to the present invention, the processor 1320 can determine whether an image of an object for pet identification is suitable for artificial intelligence based learning or identification.
[0182] According to the present invention, the processor 1320 can control the camera 1310 to determine whether the quality of the image of the object satisfies a criterion condition, transmit the image of the object to a server if the quality satisfies the criterion condition, and discard the image of the object and take a picture of the next image if the quality does not satisfy the criterion condition.
[0183] The image of the object (e.g., a muzzle print image) derived by the above process is inspected for suitability for artificial intelligence-based learning or identification. The quality inspection of the image can be performed according to various quality conditions, which can be defined by the neural network designer. For example, the conditions can include that it is a photo of an actual dog, that the muzzle print is clearly visible, that there is no foreign object, that the image is taken from the front, and that the peripheral margin is less than a certain percentage. It is preferable that such conditions can be quantified and objective. If an image of poor quality is stored in the neural network, it may cause a general degradation of the performance of the neural network, so it is preferable to pre-filter images having a quality below the standard. Such filtering can be performed in the first or second post-processing step described above.
[0184] As an embodiment for inspecting the quality of an image of an object, a method for detecting quality degradation due to defocus and quality degradation due to camera or object shaking will be described.
[0185] A method for filtering an image of an object for identifying a pet according to the present invention includes the steps of acquiring an image including a pet (S1210), determining the species of the pet in the image and setting a first feature region (S1220), setting a second feature region including an object for identifying the pet within the first feature region taking into account the determined species of the pet (S1230), and inspecting the quality of the image of the object in the second feature region to determine whether the image of the object is suitable for artificial intelligence based learning or identification (S1240).
[0186] Also, according to the present invention, post-processing (quality inspection) is performed on the first feature region, and only when the image of the first feature region has appropriate quality, the second feature region detection and suitability determination can be performed. That is, the step of setting the first feature region includes a step of inspecting the quality of the image of the object in the first feature region and determining whether the image of the object is suitable for artificial intelligence-based learning or identification, and the second feature region can be set when the image of the object in the first feature region is determined to be suitable for artificial intelligence-based learning or identification. If it is determined that the image of the object in the first feature region is not suitable for artificial intelligence-based learning or identification, the image of the current frame can be discarded and an image of the next frame can be captured.
[0187] In some embodiments, the quality inspection (first post-processing) for the first feature region may be omitted, that is, the first post-processing step may be omitted and detection of the second feature region may be performed directly.
[0188] A quality check on the image of the object can be performed by applying different weights to different positions of the first or second feature regions.
[0189] In the first post-processing step, the brightness evaluation can be performed by a method for inspecting the image quality of an object. For example, the above-mentioned Equation 2 is calculated for the first feature region, and brightness values according to the BT.601 standard and brightness information in the HSV color space are extracted in units of pixels. If the average value is smaller than the first brightness reference value, the image is determined to be too dark, and if it is larger than the second brightness reference value, the image is determined to be too bright. If the image is too dark or too bright, subsequent steps such as detection of the second feature region can be omitted and the process can be terminated. Also, a weight can be assigned to a region determined to be important among the first feature regions for the judgment.
[0190] According to the present invention, determining whether an image of an object is suitable for artificial intelligence based learning or identification may include determining a degree of defocus blur for the object in the image of the object.
[0191] The method for detecting quality degradation due to defocus (defocus blur) is as follows. Defocus blur refers to the phenomenon in which a target area (e.g., a nose print area) becomes blurred because the camera is not in focus. An example of when defocus blur occurs is a photo taken while the autofocus adjustment is being performed on a mobile phone camera.
[0192] To determine whether an image has defocus blur, high frequency components (components having a frequency higher than a certain value) can be extracted and processed from the image. In an image, high frequency components are mainly located in areas where brightness and color change rapidly, i.e., object boundaries in the image, and low frequency components are mainly located in areas with similar brightness and color to the surroundings. Therefore, the more in-focus and clear the image is, the more high frequency components are distributed in the image. To determine this, for example, the Laplacian operator can be used. The Laplacian operator performs a second derivative on the input signal, and can effectively remove the low frequency components while leaving the high frequency components of the input signal. Therefore, using the Laplacian operator can effectively find object boundaries in an image and obtain a numerical value of how clear the boundary is.
[0193] For example, by performing a convolution operation on an input photograph using a 5×5 LoG (Laplacian of Gaussian) kernel as shown in Equation 4 below, it is possible to obtain boundary position and boundary definition information within an image.
[0194] [Formula 4]
number
[0195] For photos with little defocus blur and clear boundaries, the results of applying the Laplace operator are distributed in the range from 0 to relatively large values, and conversely, for photos with a lot of defocus blur and blurred boundaries, the results of applying the Laplace operator are distributed in the range from 0 to relatively small values. Therefore, by modeling the distribution of the results of applying the Laplace operator, it is possible to grasp the clarity.
[0196] As an example of such a method, the sharpness can be grasped using the variance value of the image as a result of applying the Laplace operator. Alternatively, various statistical methods can be adopted, such as obtaining a decile distribution diagram of the Laplacian value distribution through distribution map (histogram) analysis and calculating the distribution ratio of the highest-lowest interval. Such methods can be selectively applied depending on the application field to be used.
[0197] That is, according to the present invention, the step of determining the degree of defocus of the object may include the steps of applying a Laplacian operator, which performs second-order differentiation, to the image of the second feature region to extract an image indicating a distribution map of high-frequency components, and calculating a value indicating the defocus of the image of the second feature region from the distribution map of high-frequency components.
[0198] The importance of the sharpness of the nose region also varies depending on the position. That is, if it is in the center of the image, it is highly likely to be the center of the nose, and as it moves toward the edge of the image, it is more likely to be the outer periphery of the nose or the hair region around the nose. In order to reflect such spatial characteristics, a method of dividing the image into certain regions and assigning different weights to each region to determine the sharpness may be considered. For example, a method of setting a region of interest by dividing the image into 9 or drawing an ellipse based on the center of the image, and then multiplying the region by a weight w greater than 1 may be considered.
[0199] That is, according to the present invention, a weight applied to the center of the second feature region can be set to be greater than a weight applied to the peripheral portion of the second feature region. By applying a greater weight to the center than to the peripheral portion, image quality inspection can be performed in a focused manner on an object to be identified, such as a dog's nose print.
[0200] The closer the defocus blur score determined using the Laplace operator is to 0, the weaker the boundary line is in the image, and the larger the value, the stronger the boundary line is in the image. Therefore, if the defocus blur score is greater than a threshold value, the image is classified as clear, and if not, the image is determined as blurry. This threshold value can be empirically determined using data collected in advance, or adaptively determined by accumulating and observing multiple input images each time using a camera.
[0201] According to the present invention, determining whether an image of an object is suitable for artificial intelligence based learning or identification may include determining the degree of blurring of said object in said image of said object.
[0202] A method for detecting motion blur caused by blurring will be described below. Motion blur refers to a phenomenon in which the relative position of the subject and the camera shakes during the exposure time of the camera, causing the subject to appear blurred. When taking a photo in a low-light environment with a long exposure time setting on a mobile phone camera, this can occur when the dog moves or the user's hand shakes during the exposure time to take a photo.
[0203] Various edge detectors can be used for such image feature analysis, for example, the Canny edge detector is known as a boundary detector that efficiently detects continuous boundaries.
[0204] This is an example of the result image when the Canny boundary detector is applied to the image with diagonal upward blur in Figure 10. As shown in Figure 10, by applying the Canny boundary detector to the image, it can be confirmed that the boundary of the muzzle pattern region occurs consistently in the diagonal ( / ) direction.
[0205] By analyzing the directionality of the boundary, it is possible to effectively determine whether or not there is blur. As an example of a directional analysis method, a boundary detected by a Canny boundary detector is characterized in that it is always connected to surrounding pixels. Therefore, the directionality can be analyzed by analyzing the connection relationship with the surrounding pixels. According to an embodiment of the present invention, the overall direction and degree of blur can be calculated by analyzing the pattern distribution in a pixel block of a certain size in which a boundary is located in an image to which the Canny boundary detector is applied.
[0206] FIG. 11 shows an example of a pattern of pixel blocks with boundaries that can be used to determine the presence or absence of blur in an image after applying the Canny boundary detector.
[0207] For example, a detailed description will be given of a case where a 3×3 pixel boundary is detected as shown in FIG 11(a) as follows: For convenience of explanation, the nine pixels are numbered according to their positions as shown in FIG 11.
[0208] In the case of the 5th pixel in the middle, we can assume that it will always be found to be a border. If the 5th pixel is not a border, then this 3x3 array is not a border array and we can either skip it or count it as a non-border pixel.
[0209] If the 5th pixel in the center is a border pixel, the remaining 8 surrounding pixels are divided into 2 based on whether they are border pixels or not. 8 = 256 patterns can be defined. For example, in the case of (a) of FIG. 11, the pattern is (01000100) based on whether the {1, 2, 3, 4, 6, 7, 8, 9} pixels are on the boundary line or not, and this can be converted to the decimal system and named as the 68th pattern. This naming method can be changed to facilitate implementation.
[0210]
number
[0211]
number
[0212] In this way, a lookup table can be created for 256 different patterns. The combinations that can be entered at this time can be defined in eight directions, for example, as follows:
number
[0213] Based on this method, directional statistical information of boundary pixels can be created from the result image of the Canny edge detector. Based on this statistical information, it can be effectively determined whether motion blur has occurred in the image. It is obvious that such a determination criterion can be determined based on a large amount of data by empirically designing a classification method or by using a machine learning method. As such a method, for example, a method such as a decision tree or a random forest can be used, or a classifier using a deep neural network can be designed.
[0214] That is, according to the present invention, the step of determining the degree of blurring of the object may include the steps of applying a Canny edge detector to the image of the second feature region to construct a boundary image consisting of consecutive boundaries in the image of the object, as shown in FIG. 10, analyzing the distribution of directional patterns of blocks including the boundary in the boundary image as shown in FIG. 10, and calculating a value indicating the degree of blurring of the object from the distribution of directional patterns.
[0215] When creating the above statistical information, it is obvious that the nose region has more important information than the surrounding regions. Therefore, a method of separately collecting statistical information from a certain region in an image and weighting it can be used. As an example of such a method, the method used in the above-mentioned Defocus blur discrimination using the Laplace operator can be used. That is, the step of calculating a value indicating the degree of blur of an object from the distribution of directional patterns includes a step of calculating the distribution degree of directional patterns by applying a weight to each block of the second feature region, and the weight of a block located at the center of the second feature region can be set to be larger than the weight of a block located at the periphery of the second feature region.
[0216] 12 is a flowchart of a method for filtering an image of an object for identifying a pet. The method for filtering an image of an object for identifying a pet according to the present invention includes the steps of acquiring an image including a pet (S1210), determining the species of the pet from the image and setting a first feature region (S1220), setting a second feature region including an object for identifying the pet in the first feature region in consideration of the determined species of the pet (S1230), and inspecting the quality of the image of the object in the second feature region to determine whether the image of the object is suitable for artificial intelligence-based learning or identification (S1240). The quality inspection of the image of the object can be performed by applying different weights to each position of the first feature region or the second feature region.
[0217] Meanwhile, after the step of setting the first feature region (S1220), a step of inspecting the quality of the image of the object in the first feature region and determining whether the image of the object is suitable for artificial intelligence-based learning or identification (S1230) may be performed. In this case, a second feature region may be set when it is determined that the image of the object in the first feature region is suitable for artificial intelligence-based learning or identification. The quality inspection (first post-processing) for the first feature region may be omitted depending on the embodiment.
[0218] The step of inspecting the quality of the image of the object in the first feature region to determine whether the image of the object is suitable for artificial intelligence-based learning or identification may include a step of determining whether the brightness in the first feature region belongs to a reference range. This step may include a step of extracting luma information according to the BT.601 standard and brightness information in the HSV color space from the first feature region, and determining whether the average value is between a first threshold and a second threshold. In this step, different weights may be applied depending on the position in the image when calculating the average value.
[0219] According to the present invention, determining whether an image of an object is suitable for artificial intelligence based learning or identification may include determining a degree of defocus blur for the object from the image of the object.
[0220] According to the present invention, the step of determining the degree of defocus of the object may include the steps of extracting an image indicating a distribution map of high frequency components by applying a Laplacian operator that performs second-order differentiation to the image of the second feature region, and calculating a value indicating the defocus of the image of the second feature region from the distribution map of high frequency components.
[0221] According to the present invention, the weight applied to the central portion of the first feature region or the second feature region can be set to be greater than the weight applied to the peripheral portion of the first feature region or the second feature region.
[0222] According to the present invention, determining whether an image of an object is suitable for artificial intelligence based learning or identification may include determining the degree of blurring of said object in said image of said object.
[0223] According to the present invention, the step of determining the degree of blurring of the object may include the steps of applying a Canny edge detector to the image of the second feature region to construct a border image consisting of consecutive borders in the image of the object, analyzing the distribution of directional patterns of blocks containing the borders in the same border image, and calculating a value indicating the degree of blurring of the object from the distribution of directional patterns.
[0224] According to the present invention, the step of calculating a value indicating the degree of blurring of the object from the distribution of the directional pattern includes a step of calculating the distribution degree of the directional pattern by applying a weight to each block of the second feature region, and the weight of the block located in the center of the second feature region can be set to be greater than the weight of the block located in the periphery of the second feature region.
[0225] The electronic device 1300 according to the present invention includes a camera 1310 for generating an image including a pet, and a processor 1320 for processing the image provided by the camera 1310 to generate an image of an object for identifying the pet. The processor 1320 is configured to set a first feature region for determining the species of the pet in the image, set a second feature region including an object for identifying the pet within the first feature region taking into account the determined species of the pet, and inspect the quality of the image of the object in the second feature region to determine whether the image of the object is suitable for artificial intelligence based learning or identification.
[0226] According to the present invention, the processor 1310 can check the quality of the image of the object in the first feature region to determine whether the image of the object is suitable for artificial intelligence based learning or identification. Here, the second feature region detection and quality check can be performed only if the image of the object in the first feature region is suitable for artificial intelligence based learning or identification.
[0227] According to the present invention, the processor 1310 can determine whether the brightness in the first feature region belongs to a reference range. The quality check (first post-processing) for the first feature region may be omitted depending on the embodiment.
[0228] Here, the quality check for the image of the object can be performed by applying different weights to each position of the first feature region or the second feature region.
[0229] In accordance with the present invention, the processor 1310 can determine the degree of defocus for the object in the image of the object.
[0230] According to the present invention, the processor 1310 can extract an image representing a distribution map of high frequency components from the image of the second feature region, and calculate a value representing the defocus of the image of the second feature region from the distribution map of high frequency components.
[0231] According to the present invention, the weight applied to the central portion of the first feature region or the second feature region can be set to be greater than the weight applied to the peripheral portion of the first feature region or the second feature region.
[0232] In accordance with the present invention, the processor 1310 can determine the degree of blurring of an object in an image of the object.
[0233] According to the present invention, the processor 1310 can construct a border image consisting of the border of the image of the second feature region, analyze the distribution of directional patterns of blocks containing the border in the border image, and calculate a value representing the degree of blurring of the object from the distribution of the directional patterns.
[0234] According to the present invention, the processor 1310 calculates the distribution degree of the directional pattern by applying a weight to each block of the second feature region, and the weight of the block located in the center of the second feature region can be set to be greater than the weight of the block located on the periphery of the second feature region.
[0235] The present embodiment and the drawings attached to this specification merely clearly show a part of the technical ideas contained in the present invention, and it is self-evident that all modifications and specific embodiments that can be easily inferred by a person skilled in the art within the scope of the technical ideas contained in the specification and drawings of the present invention are included in the scope of the present invention.
[0236] Therefore, the spirit of the present invention should not be limited to the described embodiments, and all things that are equivalent or equivalent modifications to the scope of the claims, as well as the scope of the claims described below, should be considered to fall within the scope of the spirit of the present invention.
Claims
1. 1. A method for detecting objects for pet identification, comprising: acquiring an original image including the pet; determining a first feature region and a species of the pet through image processing on the original image; and detecting an object for identifying the pet within the first feature region based on the determined pet species; The step of determining the species of the pet comprises: applying a first pre-processing step to the original image; determining the species of the pet in the pre-processed image and setting the first feature region; and extracting a first feature value via a first post-processing step on the first feature region.
2. The step of setting the first feature region includes: generating a plurality of feature images from the preprocessed image using a training neural network; applying a predefined bounding box to each of the plurality of feature images; calculating a probability value for each pet type within the bounding box; and configuring the first feature region to include the bounding box if the calculated probability value for a particular animal species is greater than or equal to a reference value.
3. If the first feature value is greater than a reference value, object detection is performed to identify the pet; The method of claim 1 , wherein if the first feature value is less than a reference value, further processing is omitted.
4. The step of applying a first pre-processing to the original image includes: converting the original image into an image with a first resolution lower than the original resolution; and applying the first pre-processing to the image converted to the first resolution.
5. A method for detecting objects for pet identification, comprising: acquiring an original image including the pet; determining a first feature region and a species of the pet through image processing on the original image; and detecting an object for identifying the pet within the first feature region based on the determined pet species; The step of detecting an object for identifying a pet includes: applying a second pre-processing step to the first feature region for identifying the species of the pet; setting a second feature region for identifying the pet based on the species of the pet in the second pre-processed first feature region; and applying a second post-processing step to the second feature region to extract second feature values.
6. The method according to claim 5 , wherein the second pre-processing for the first feature region is performed at a second resolution higher than a first resolution at which the first pre-processing for setting the first feature region is applied.
7. The method of claim 5 , wherein the step of setting the second feature area includes a step of setting the second feature area based on a probability that an object for identifying the pet is located in the first feature area depending on the species of the pet.
8. The method of claim 5 , wherein if the second feature value is greater than a reference value, an image including the second feature region is transmitted to a server.
9. A method for detecting objects for pet identification, comprising: acquiring an original image including the pet; determining a first feature region and a species of the pet through image processing on the original image; and detecting an object for identifying the pet within the first feature region based on the determined pet species; The step of generating the first feature region includes: generating candidate feature regions for determining the species of the pet in the original image; generating a first feature region having a position and a size determined based on the confidence value of each of the feature region candidates.
10. 1. An electronic device for detecting objects for pet identification, comprising: A camera for generating an original image including the pet; A processor for determining a first feature region and the species of the pet through image processing of the original image, and detecting an object for identifying the pet within the first feature region based on the determined species of the pet; a communication module for transmitting an image of the object for identifying the pet to a server if the object is valid; The processor, applying a first pre-processing to the original image; determining the species of the pet in the preprocessed image to set the first feature region; An electronic device extracts a first feature value through a first post-processing on the first feature region.
11. The processor, generating a plurality of feature images from the preprocessed image using a training neural network; applying a predefined bounding box to each of the plurality of feature images; Calculate a probability value for each pet type within the bounding box; The electronic device of claim 10 , further comprising: configuring the first feature region to include the bounding box if the calculated probability value for a particular animal species is greater than or equal to a reference value.
12. If the first feature value is greater than a reference value, object detection is performed to identify the pet; The electronic device of claim 10 , wherein if the first characteristic value is less than a reference value, further processing is omitted.
13. The processor, converting the original image into an image having a first resolution lower than the original resolution; The electronic device of claim 10 , further comprising: applying the first pre-processing to an image converted to the first resolution.
14. An electronic device for detecting objects for pet identification, comprising: A camera for generating an original image including the pet; A processor for determining a first feature region and the species of the pet through image processing of the original image, and detecting an object for identifying the pet within the first feature region based on the determined species of the pet; a communication module for transmitting an image of the object for identifying the pet to a server if the object is valid; The processor, applying a second pre-processing step to the first feature region for identifying the species of the pet; setting a second feature region for identifying the pet based on the species of the pet in the second pre-processed first feature region; The electronic device applies a second post-processing to the second feature region to extract second feature values.
15. The electronic device according to claim 14 , wherein the second pre-processing on the first feature region is performed at a second resolution higher than a first resolution at which the first pre-processing for setting the first feature region is applied.
16. The electronic device of claim 14 , wherein the processor sets the second feature area based on a probability that an object for identifying the pet is located in the first feature area depending on the species of the pet.
17. The electronic device of claim 14 , wherein if the second feature value is greater than a reference value, an image including the second feature region is transmitted to the server.
18. An electronic device for detecting objects for pet identification, comprising: A camera for generating an original image including the pet; A processor for determining a first feature region and the species of the pet through image processing of the original image, and detecting an object for identifying the pet within the first feature region based on the determined species of the pet; a communication module for transmitting an image of the object for identifying the pet to a server if the object is valid; The processor, generating candidate feature regions for determining the species of the pet in the original image; The electronic device generates a first feature region whose position and size are determined based on the confidence value of each of the feature region candidates.
Citation Information
Patent Citations
Device for individual identification
JP2001202516A
Information processing system and program
JP2019010004A
Computer program and theminal for providing individual animal information based on the facial and nose pattern imanges of the animal
KR1020200044209A
System and method for matching an animal to existing animal profiles
US20150131868A1