Methods and electronic devices for photographing objects used to identify pets.

CN117296083BActive Publication Date: 2026-08-14PETNOW
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]但是,在获得并存储宠物的鼻纹方面存在一些问题

Benefits of technology

[0031]根据本发明的用于检测用于识别宠物的客体的方法以及电子装置以使得在拍摄宠物之后立即选择用于鼻纹的学习或者识别的图像,并存储在服务器的数据库,从而可以有效地获取相对应于用于学习或者识别的宠物的鼻子的客体的图像。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117296083B_ABST
    Figure CN117296083B_ABST
Patent Text Reader

Abstract

This invention provides an image processing method and electronic device for effectively detecting objects used to identify pets while reducing computational complexity. The method for detecting objects used to identify pets according to the invention includes the steps of: acquiring an original image of the pet; determining a first feature region and the species of the pet through image processing of the original image; and detecting an object used to identify the pet within the first feature region based on the determined pet species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and electronic device for photographing an object for identifying a pet, and more specifically, to a method and electronic device for acquiring images of an object for identifying a pet suitable for artificial intelligence-based learning or recognition. Background Technology

[0002] In modern society, the demand for emotionally supportive pets is increasing as people live with them. Therefore, it is necessary to improve the database management of various pet information for purposes such as pet health management. To manage pets, identification information is needed, similar to fingerprints for humans, and specific identification objects can be defined for each pet. For example, in the case of puppies, since each puppy has a unique nose print (the shape of the wrinkles on its nose), each puppy can use its nose print as identification information.

[0003] like Figure 1 As shown in (a), the method for registering a nose print is to photograph a face including the nose of a pet, similar to registering a person's fingerprint or face (S110), and then perform the process of storing the image including the nose print and registering it in a database (S120). Alternatively, the method for querying a nose print can be as follows: Figure 1 As shown in (b), the process involves photographing the pet's nose print (S130), exploring nose prints consistent with the photographed print and related information (S140), and outputting information consistent with the photographed nose print (S150). Figure 1 This would allow for the identification and management of individual pets by logging in and querying their nose prints. The pets' nose print information can be stored in a database and used as data for AI-based learning or recognition.

[0004] However, there are some issues with obtaining and storing a pet's nose print.

[0005] First, photographs are difficult to identify due to factors such as shooting angle, focus, distance, size, and environment. While attempts have been made to apply human facial recognition technology to nose print recognition, human facial information has accumulated sufficient data. Conversely, pet nose print information suffers from low recognition rates due to the lack of sufficient data. Specifically, to perform AI-based recognition, learning data needs to be processed into a form that machines can learn from, but pet nose prints lack sufficient accumulated data, making nose print recognition difficult.

[0006] Furthermore, for nose print recognition in pets, clear images of nasal wrinkles are required. However, unlike humans, pets do not exhibit behaviors such as temporarily stopping their actions, making it difficult to obtain clear images of nasal wrinkles. For example, puppies constantly move their faces and stick out their tongues, making it difficult to obtain nose print images of the desired quality. For instance, for nose print recognition, clear images of nasal wrinkles are needed, but in reality, most images are not clear due to factors such as camera shake. To address these issues, a method of forcibly fixing the puppy's nose during the photograph is being considered, but this has been deemed inappropriate because it involves forcing the pet into a certain behavior. Summary of the Invention

[0007] This invention provides an image processing method and electronic device that can effectively detect objects used for pet identification while reducing computational complexity.

[0008] This invention provides an image processing method and electronic device that can effectively detect objects used for pet identification while reducing computational complexity.

[0009] The present invention provides a method and electronic device for effectively filtering low-quality images during the acquisition of images of objects used for pet identification.

[0010] The problems solved by the present invention are not limited to those mentioned above, and those skilled in the art can clearly understand other problems not mentioned from the following description.

[0011] The method for detecting an object for identifying a pet according to the present invention includes the steps of: acquiring an original image including the pet; determining a first feature region and the species of the pet by image processing of the original image; and detecting an object for identifying the pet within the first feature region based on the determined species of the pet.

[0012] According to the present invention, the step of determining the species of the pet may include: applying a first preprocessing step to the original image; determining the species of the pet in the preprocessed image and setting a first feature region; and extracting a first feature value by a first postprocessing step to the first feature region.

[0013] According to the present invention, the step of defining the first feature region may include: generating a plurality of feature images from the preprocessed image using a learning neural network; applying a predefined bounding box to each of the plurality of feature images; calculating a probability value within the bounding box according to the species of each pet; and when the calculated probability value is above a standard value for a specific animal species, the first feature region is defined as including the bounding box.

[0014] According to the present invention, when the first feature value is greater than the standard value, object detection for identifying the pet is performed, and when the first feature value is less than the standard value, additional processing is omitted.

[0015] According to the present invention, the step of applying the first preprocessing to the original image may include: transforming the original image into an image with a first resolution lower than the original resolution; and applying the first preprocessing to the image transformed to the first resolution.

[0016] According to the present invention, the step of detecting an object for identifying the pet may include: applying a second preprocessing step to a first feature region for identifying the species of the pet; setting a second feature region for identifying the pet based on the species of the pet in the first feature region of the second preprocessing; and applying a second postprocessing step to extract a second feature value for the second feature region.

[0017] According to the present invention, the second preprocessing for the first feature region may be performed at a second resolution higher than the first resolution of the first preprocessing applied to the first feature region.

[0018] According to the present invention, the step of setting the second feature region may include: setting the second feature region based on the probability that an object for identifying the pet exists in the first feature region according to the species of the pet.

[0019] According to the present invention, when the second feature value is greater than a standard value, the image including the second feature region may be transmitted to the server.

[0020] According to the present invention, the step of generating the first feature region may include: generating feature region candidates for determining the species of the pet in the image; and generating a first feature region with determined position and size based on the confidence value of each of the feature region candidates.

[0021] The electronic device according to the invention includes: a camera for generating an original image including the pet; a processor for determining a first feature region and the species of the pet through image processing of the original image, and detecting an object for identifying the pet within the first feature region based on the determined pet species; and a communication module for transmitting an image of the object to a server when the object for identifying the pet is valid.

[0022] According to the present invention, the processor may apply a first preprocessing to the original image, and in the preprocessed image, determine the species of the pet to set the first feature region, and extract a first feature value by a first postprocessing to the first feature region.

[0023] According to the present invention, the processor may use a learning neural network to generate multiple feature images from the preprocessed image, apply a predefined bounding box to each of the multiple feature images, calculate a probability value within the bounding box according to the species of each pet, and for a specific animal species, when the calculated probability value is above a standard value, constitute the first feature region including the bounding box.

[0024] According to the present invention, when the first feature value is greater than the standard value, object detection for identifying the pet is performed, and when the first feature value is less than the standard value, additional processing is omitted.

[0025] According to the present invention, the processor may transform the original image into a first resolution image lower than the original resolution, and apply the first preprocessing to the image transformed to the first resolution.

[0026] According to the present invention, the processor may apply a second preprocessing to a first feature region for identifying the species of the pet, and set a second feature region for identifying the pet based on the species of the pet in the first feature region of the second preprocessing, and apply a second postprocessing to the second feature region to extract a second feature value.

[0027] According to the present invention, the second preprocessing for the first feature region may be performed at a second resolution higher than the first resolution of the first preprocessing applied to the first feature region.

[0028] According to the present invention, the processor may set the second feature region based on the probability that an object for identifying the pet exists in the first feature region according to the species of the pet.

[0029] According to the present invention, when the second feature value is greater than a standard value, the image including the second feature region may be transmitted to the server.

[0030] According to the present invention, the processor may generate candidate feature regions for determining the species of the pet in the image, and generate a first feature region with determined position and size based on the confidence value of each of the candidate feature regions.

[0031] The method and electronic device for detecting objects for identifying pets according to the present invention enable the immediate selection of images for learning or recognizing nose prints after photographing the pet and storage in a database on a server, thereby enabling the efficient acquisition of images of objects corresponding to the noses of the pets used for learning or recognition.

[0032] Furthermore, the method and electronic device for detecting objects used to identify pets according to the present invention first determine the species of the pet and then extract the pet's nose print image, thereby reducing computational complexity.

[0033] According to the present invention, in the process of determining the feature region for identifying the species of the pet, a final feature region with a further wide region is generated by considering the confidence value of each of a plurality of feature region candidates, so that the identification object of the pet can be detected in the final feature region, thereby enabling further accurate detection.

[0034] According to the present invention, the quality of an object image used to identify a pet, such as a dog's nose, can be checked in the captured image to determine whether the image is suitable for artificial intelligence-based learning or recognition, and only suitable images can be stored to optimize the neural network used for learning or recognition.

[0035] The effects of the present invention are not limited to those mentioned above, and those skilled in the art can clearly understand from the following description other effects not mentioned. Attached Figure Description

[0036] Figure 1 This demonstrates a basic AI-based program for pet management.

[0037] Figure 2 A procedure for AI-based pet nose print management, applicable to the learning or recognition of the fitness judgment of object images according to the present invention, is shown.

[0038] Figure 3 A procedure for detecting objects used to identify pets is shown in a pet management system according to the present invention.

[0039] Figure 4 An example of a UI (User Interface) screen for detecting the identification object of the pet to which this invention is applicable is shown.

[0040] Figure 5 The process for detecting objects used to identify pets according to the present invention is illustrated.

[0041] Figure 6 The process for defining a feature region according to the present invention is shown.

[0042] Figure 7 A flowchart of a process for detecting an object used to identify a pet, according to the present invention, is shown.

[0043] Figure 8 The process of deriving characteristic regions for determining the species of pets according to the present invention is illustrated.

[0044] Figure 9 A flowchart of the process for processing an image of an object for identifying a pet, according to the present invention, is shown.

[0045] Figure 10 This example shows the result image after applying the Canny edge detector to the input image.

[0046] Figure 11 An example is shown of a pattern of pixel blocks used to determine whether a boundary line is wobbly in a result image of an image obtained by applying the Canny boundary line detector.

[0047] Figure 12 A flowchart of a method for filtering images of objects used to identify pets.

[0048] Figure 13 This is a block diagram of an electronic device according to the present invention. Detailed Implementation

[0049] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily implement the invention. The present invention can be implemented in various different forms and is not limited to the embodiments described herein.

[0050] To clearly illustrate the present invention, parts unrelated to the description have been omitted, and the same or similar constituent elements are referred to by the same reference numerals throughout the specification.

[0051] Furthermore, in multiple embodiments, the same symbols are used for constituent elements having the same structure, and the description is only made in the representative embodiment. In other embodiments, the description is only made in the representative embodiment and other structures.

[0052] Throughout this specification, when a part is "connected (or combined)" with other parts, this includes not only "direct connection (or combination)" but also "indirect connection (or combination)" through other components. Furthermore, when a part "includes" a constituent element, this means that, unless specifically stated otherwise, it may include other constituent elements, rather than excluding them.

[0053] Unless otherwise defined, all terms used herein, including technical and scientific terms, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms such as those defined in commonly used dictionaries shall be interpreted as having the same meaning in the context of the relevant art, and shall not be construed as having an idealized or overly formal meaning unless explicitly defined in this application.

[0054] This paper focuses on extracting identification information using the shape of the wrinkles on a dog's nose (nose print). However, the scope of pets in this invention is not limited to dogs. Furthermore, the features used for identification information are not limited to nose prints and can be applied to various pet body characteristics.

[0055] As explained earlier, since there are insufficient pet nose print images suitable for AI-based learning or recognition, and the quality of pet nose print images is likely to be low, it is necessary to selectively store nose print images in a database for AI-based learning or recognition.

[0056] Figure 2 This invention illustrates an AI-based pet nose print management procedure for judging the suitability of object images for learning or recognition according to the present invention. After photographing the pet's nose print, the invention first determines whether the photographed nose print image is suitable as data for AI-based learning or recognition. If suitable, it transmits and stores the image to a server for AI-based learning or recognition, where it is used as data for learning or recognition.

[0057] like Figure 2 As shown, the nose print management program according to the present invention generally includes a nose print acquisition program and a nose print recognition program.

[0058] According to the present invention, when registering a new pet's nose print, after capturing an image including the pet, the nose print image is extracted from the pet's facial region, specifically by first determining whether the nose print image is suitable for the pet's recognition or learning. When the captured image is determined to be suitable for recognition or learning, the image is transmitted to a server (artificial intelligence neural network) and stored in a database.

[0059] When querying a pet's identification information via nose prints, the process is similar: after capturing an image of the pet, a nose print image is extracted from the pet's facial area. Specifically, it's first determined whether the nose print image is suitable for the pet's identification or learning. If the captured image is deemed suitable for identification or learning, it is transmitted to a server and matched with previously stored nose print images to extract the pet's identification information.

[0060] In the nose print registration process, such as Figure 2The pet is photographed as in (a) (S205). In the photographed pet image, the facial region (hereinafter referred to as the first feature region) is first detected (S210). The area occupied by the nose in the facial region (hereinafter referred to as the second feature region) is detected. The nose print image is output by performing a quality check on whether the photographed image is suitable for learning or recognition (S215). The output image is transmitted to the server that constitutes the artificial neural network and stored and logged (S220).

[0061] In the nose print lookup program, such as Figure 2 The process involves photographing the pet as described in (b) (S230), detecting the facial region in the pet's image (S235), detecting the area occupied by the nose within the facial region, and outputting a nasal print image by performing a quality check on the photographed image to determine if it is suitable for learning or recognition (S240), similar to the nasal print registration procedure. Subsequent procedures include comparing the output nasal print image with previously stored and learned nasal print images to explore consistency information (S245) and outputting the exploration results (S250).

[0062] Figure 3 A procedure for detecting objects corresponding to a pet's nose is shown in a pet nose print management system according to the present invention.

[0063] Reference Figure 3 First, an initial image of the pet is generated by taking a picture (S305). In this initial image, a step of detecting the facial region is performed first (S310). Then, within the facial region, a step of detecting the nose region is performed, taking into account the pet's species (S315). Because detecting the facial region first and then the nose region reduces computational complexity and improves detection accuracy compared to detecting the nose region by considering all species through cascaded detection, this approach reduces computational complexity. Next, a quality check is performed to determine whether the image of the detected nose region is suitable for subsequent nasal print recognition or learning (S320). If the quality check result indicates that the image is suitable, it is transmitted to the server for nasal print recognition or stored for future learning or recognition (S325).

[0064] Furthermore, according to the present invention, the camera can be controlled to focus on the detected nose area so that the image of an object used to identify a pet, such as the wrinkles (nose prints) on a dog's nose, is not captured blurry (S330). This is to prevent image quality degradation due to focus misalignment of the nose by focusing the camera on the nose area.

[0065] Figure 4 An example of a UI (User Interface) screen for obtaining a nose print image of a pet to which this invention is applicable is shown. Figure 4This illustrates a scenario used to obtain nose prints from a variety of pets for puppies.

[0066] Reference Figure 4 The system identifies the species of the pet in the captured image to determine if the pet in the current photo is a puppy. If the pet in the photo is not a puppy, such as... Figure 4 The algorithm outputs messages like "puppy not found," and when the pet in the photo is a puppy, it executes a program to obtain the puppy's nose print. To determine if the pet in the photo is a puppy, the facial region of the pet included in the image can be extracted first, and the image including the facial region can be compared with existing learned data to determine the pet's species.

[0067] After that, as Figure 4 After defining the area on the puppy's face corresponding to its nose as shown in (b) to (e), the camera can be focused on that area for taking a picture. That is, the camera can be controlled to focus on the area (center point) corresponding to the object being identified for the pet. Furthermore, to provide feedback to the user that the focus is on the currently tracked object (e.g., the nose) for taking a picture, graphic elements can be overlaid on the location of the tracked object. By displaying graphic elements indicating the detection status of the tracked object at its location, the user can understand that the pet being photographed is performing object recognition.

[0068] like Figure 4 As shown in (b) to (e), when the image quality of the currently captured object is good (when the image quality of the object meets the baseline conditions), the first graphic element 410A, which is shown as being in a good quality state, can be overlaid on the object and output (e.g., a smiling icon or a green icon). When the image quality of the currently captured object is poor (when the image quality of the object does not meet the baseline conditions), the second graphic element 410B, which is shown as being in a poor quality state (e.g., a crying icon or a red icon), can be overlaid on the object and output.

[0069] like Figure 4 This allows the system to track the puppy's nose and focus on it while taking pictures, even when the puppy is moving continuously. Then, it can determine whether the puppy's nose print image is suitable for pet recognition or learning in each captured image and output the degree of suitability.

[0070] For example, the suitability of a captured image of a dog's nose print for pet recognition or learning can be calculated numerically. A lower suitability score indicates a "bad" result, while a higher suitability score indicates a "good" result. In other words, the image can output a score 420 indicating the quality of the captured object.

[0071] Additionally, the system can perform a quality evaluation (size, brightness, sharpness, etc.) on the currently captured nose print image and output feedback information 430 to the user to enable the capture of nose print images suitable for AI-based recognition or learning. For example, when the size of the dog's nose print image is smaller than a standard value, the system can output information such as... Figure 4 (c) shows information such as "Please adjust the distance between the puppy's nose and the image" to capture a larger nose print image. Additionally, for pet recognition, progress information 440 can be output indicating the progress of acquiring images of the object with appropriate quality. For example, if four nose print images of appropriate quality are needed, and one appropriate image has been obtained so far, then... Figure 4 That would output progress information 440, indicating that the progress is 25%.

[0072] Once the puppy's nose print image has been fully acquired, the shooting can be stopped and the recognition information can be stored in the database along with the puppy's nose print image, or the puppy's recognition information can be output.

[0073] In this invention, the pet's facial region is detected first, followed by the nose region within the facial region. This is to reduce computational complexity and simplify object detection. During image capture, objects other than the desired object, or unnecessary or erroneous information, may be included in the image. Therefore, this invention first determines whether the desired object (the pet's nose) exists in the captured image.

[0074] Furthermore, recognizing a pet's nose print requires images with a certain level of resolution, but higher resolution images require significantly more computational resources for processing. Additionally, as the variety of pet species increases, the learning methods for each species also differ, further increasing the computational difficulty of artificial intelligence. In particular, similar species of animals often share similar shapes (e.g., a dog's nose is similar to a wolf's nose), making it computationally challenging to classify noses along with animal species for similar species.

[0075] Therefore, to reduce computational complexity, this invention employs a cascaded object detection method. For example, while photographing a pet, the facial region of the pet is first detected, then the pet's species is identified, and finally, the pet's nose region is detected based on the detected facial region and the identified pet species. This process of identifying the pet's species is performed first at a relatively low resolution with relatively low computational complexity, while a pet species-specific object detection method is applied to maintain high resolution in the pet's facial region and perform nose region detection. Thus, this invention can relatively reduce computational complexity while effectively detecting the pet's nose region.

[0076] Figure 5 The entire image processing procedure for pet identification according to the present invention is illustrated. Figure 5 As shown, the method for processing an input image according to the present invention includes a step of receiving an input image in a camera (S505), a first preprocessing step of generating a processed image by adjusting the size of the input image (S510), a first feature region detection step of detecting the location and species of an animal from the processed image generated in the first preprocessing step (S515), a first postprocessing step of extracting a first feature value of the animal image from the result of the first feature region detection step (S520), a step of determining a detector for detecting an object (e.g., nose) for identifying the pet in the image processed by the first postprocessing step (S525), a second preprocessing step of adjusting the size of the image for image processing for identifying the pet (S530), at least one second feature region detection step corresponding to the species of the animal that can be detected in the first feature detection step (S535), and a second postprocessing step of extracting a second feature value of the animal image corresponding to each second feature region detection step (S540).

[0077] First preprocessing step

[0078] The first preprocessing step (S510) for the original image is a step of adjusting the size, scale, orientation, etc. of the original image to transform the image into a form suitable for object detection.

[0079] With the development of camera technology, most input images consist of millions to tens of millions of pixels, making direct processing of such large images undesirable. To effectively perform object detection, a preprocessing procedure is needed to appropriately modify the input image. This process is mathematically comprised of coordinate system transformations.

[0080] It is obvious that any four points in the input image can be mapped to the four vertices of the processed image, and any coordinate system transformation can be used to generate any processed image. However, during the coordinate system transformation, when using any nonlinear transformation function, it is necessary to be able to obtain the inverse transformation of the feature region of the input image from the bounding box obtained by the feature region detector. For example, if an affine transformation is used, which involves a linear transformation by mapping any four points in the input image to the four vertices of the processed image, its inverse transformation can be easily obtained, and therefore it is preferred.

[0081] As an example of methods for determining any four points within an input image, one could consider directly using the four vertices of the input image. Alternatively, one could add blank areas to the input image or crop a portion of it to make the horizontal and vertical lengths proportional. Various interpolation methods could also be applied to reduce the size of the input image.

[0082] First Feature Region Detection Steps

[0083] The purpose of this step is to first detect the areas where pets exist and the species of the animals within the preprocessed image, thereby defining the first feature regions that can be used in the second feature region detection step described later, and at the same time selecting the second feature region detector that is most suitable for each pet species, thereby improving the final feature point detection performance.

[0084] In this process, anyone with general knowledge in the relevant field can easily combine either object detection or classification methods. However, it is known that artificial neural network-based methods outperform previous methods, therefore, it is preferable to use feature detection techniques based on artificial neural networks whenever possible. For example, an algorithm for detecting objects of various sizes in a single image, namely the SSD (Single-Shot Multibox Detection) feature detector, can be used in artificial neural networks.

[0085] The input image, normalized by the preprocessor described above, is then layered into feature images from the first to the nth feature image using an artificial neural network. At this point, the methods for extracting feature images at each layer can be mechanically learned during the learning steps of the artificial neural network.

[0086] The extracted hierarchical feature images are combined with the prior box catalogs corresponding to each hierarchical level to generate bounding boxes, individual species, and confidence value catalogs. This computational process can also be mechanically learned during the learning steps of an artificial neural network. For example, the resulting values ​​are returned in the form shown in Table 1 below. At this point, the number of species that the neural network can identify is determined by the neural network design steps, which define "background" when there is no object by default.

[0087] Table 1

[0088]

[0089] These result boxes are merged through a Non-Maximum Suppression (NMS) step to return the final object detection result within the image. NMS, as a process of deriving the final feature region from multiple feature region candidates, can be based on, for example... Figure 6 Such a procedure considers probability values ​​as shown in Table 1 to generate the data.

[0090] The process is explained in detail below.

[0091] 1. Perform the following procedure separately for each species except the background.

[0092] A. In the bounding box catalog, exclude boxes whose probability of being the species is below a certain threshold. If no boxes remain, end with no results.

[0093] B. In the bounding box catalog, the box with the highest probability for that species will be designated as the first box (first boundary region) and excluded from the bounding box catalog.

[0094] C. For the remaining bounding box directories, execute the following procedures in order of probability.

[0095] i. Calculate the area ratio of the intersection and union of the first box.

[0096] ii. If the IoU is above a certain threshold, then this box overlaps with the first box. Merge with the first box.

[0097] D. Append the first box to the results box directory.

[0098] E. If there are still boxes remaining in the bounding box directory, then use the remaining boxes as objects and repeat from step C.

[0099] For two frames A and B, the ratio of the area of ​​their intersection to the area of ​​their union can be efficiently calculated as shown in Equation 1 below.

[0100]

Mathematical Formula 1

[0101]

[0102] That is, according to the present invention, the step of generating feature region candidates may include the step of selecting a first boundary region (first box) with the highest probability corresponding to a specific animal species in the feature image, and the step of calculating the area ratio (IoU) of the intersection and union with the first boundary region according to the probability value of other boundary regions besides the boundary region (first box) selected in the feature image, and including the boundary regions whose intersection and union areas are greater than the reference area ratio in the feature region candidates of the feature image.

[0103] The method for merging boxes that overlap with the first box during the process is described below. For example, it could be a method of merging by keeping the first box as is and removing the second box from the bounding box catalog (Hard NMS). Alternatively, it could be a method of merging by keeping the first box as is and reducing the weighted value of the probability that the second box is a specific species to a value between (0, 1), and only merging by removing it from the bounding box catalog if the resulting attenuation value is less than a certain threshold (Soft NMS).

[0104] As an embodiment proposed in this invention, a new method (ExpansionNMS) can be used to merge the first box (first feature region candidate) and the second box (first feature region candidate) according to probability values, as shown in the following mathematical formula 2.

[0105]

Mathematical Formula 2

[0106]

[0107]

[0108]

[0109] W n =P 1 ·W 1 +P 2 ·W 2

[0110] H n =P 1 ·H 1 +P 2 ·H 2

[0111] At this time, p 1 p 2 These are the probability values ​​of the first bounding box (candidate for the first feature region) and the second bounding box (candidate for the first feature region), respectively, C. (x,y) 1 C(x,y) 2 C (x,y) n These represent the (x, y) coordinates of the midpoints of the first frame, the second frame, and the merged frame, respectively. Similarly, W1, W2, and W... n These are the horizontal widths (H) of the first frame, the second frame, and the merged frame, respectively. 1 H 2 H n The vertical height is indicated. The probability value of the merged boxes can use the probability value of the first box. It is determined by considering the confidence value to be set for a specific species in each feature region candidate through the first feature region derived according to the extended NMS of the present invention.

[0112] That is, the center point C of the first feature region is set. (x,y) n As in mathematical formula 2, by targeting the center point C of the feature region... (x,y) 1 C (x,y) 2 Trust value p 1 p 2 It is determined by the weighted sum.

[0113] Alternatively, it could be the first feature region W. n As in Formula 2, the confidence values ​​p for the widths W1 and W2 of the candidate feature regions are used. 1 p 2 The height H of the first feature region is determined by the weighted sum of the values. n By targeting the height H of the candidate feature region 1 H 2 Trust value p 1 p 2 It is determined by the weighted sum.

[0114] The new bounding boxes generated according to the embodiments result in boxes with larger widths and spans when compared with existing Hard-NMS or Soft-NMS methods. According to this embodiment, preprocessing detection for multi-step detectors can include additional blanking structures, and these blankings are adaptively determined using the expanded NMS (ExpansionNMS) according to the invention.

[0115] Figure 8 An example is shown where a pet's characteristic regions are detected using the extended NMS according to the present invention, compared to existing NMS. Figure 8 (a) shows multiple candidate feature regions generated in the original image. Figure 8 (b) shows an example of the first feature region derived using existing NMS. Figure 8 (c) shows an example of a second feature region derived by applying the extended NMS according to the invention. Figure 8 As shown in (b), the existing NMS (Hard NMS, Soft NMS) selects the box with the highest confidence among multiple boxes (feature region candidates), thus having the possibility of obtaining the region detachment required for nasal texture in the subsequent second feature region detection process, such as the nose region.

[0116] Therefore, the present invention can apply a weighted average based on confidence values ​​to multiple boxes (candidate feature regions) as follows: Figure 8 As in (c), a box with a large width and height is designated as the first feature region (the pet's face region), and a second feature region (the nose region) for pet identification is detected within the first feature region. By further extending the first feature region as described in this invention, the error of the second feature region not being detected during subsequent execution can be reduced.

[0117] Finally, for such a defined one or more bounding boxes, the feature regions in the original image can be obtained by inverse transformation of any transformation process used in the preprocessing step. Depending on the composition, the feature regions in the original image can be adjusted to facilitate the execution of the second detection step, which will be described later.

[0118] First post-processing step

[0119] For each feature region of the input image obtained from the aforementioned first feature region setting step (S515), a first feature value can be generated by performing an additional post-processing step. For example, in order to obtain the brightness information (first feature value) of the first feature region of the input image, an operation such as the following mathematical formula 3 can be performed.

[0120]

Mathematical Expression 3

[0121] L x,y =0.299·R x,y +0.587·G x,y +0.114·B x,y

[0122]

[0123] V x,y =Max(R) x,y G x,y B x,y )

[0124]

[0125] At this point, L is the Luma value according to the BT.601 standard, and V is the lightness value defined in the HSV color space. M and N are the horizontal width and vertical height of the object's feature area.

[0126] By utilizing this additionally generated first feature value, it is possible to predict whether the first feature region obtained in the first feature region detection step (S515) is suitable for use in the application domain associated with this application. It is obvious that the additionally generated first feature value can be appropriately designed according to the application domain. When the conditions for the first feature value defined in the application domain are not met, the system can be selectively configured to omit the second feature region setting and object detection steps described later.

[0127] Second feature region detection step

[0128] The purpose of this step is to extract specific feature regions required in the application area from the area where the animal exists. For example, in an application area that can detect the positions of the eyes, nose, mouth, and ears in the animal's facial region, for instance, the first feature region detection step distinguishes between the animal's facial region and the animal's species information, and the second feature region detection step detects the positions of the eyes, nose, mouth, and ears based on the animal's species.

[0129] In this process, the second feature region detection step can consist of multiple independent feature region detectors for each animal species. For example, if the first feature region detection step can distinguish between dogs, cats, and hamsters, it is preferable to set up three second feature region detectors, each specifically designed for dogs, cats, and hamsters. By doing so, the learning complexity can be reduced by decreasing the number of features to be learned in a single feature region detector. Furthermore, it is obvious that less data is needed for neural network learning in terms of learning data collection.

[0130] Each second feature region detector is constructed independently, so that anyone with general knowledge can easily construct a single, independent detector. Each feature region detector is preferably constructed separately to suit the feature information to be detected in each species. Alternatively, to reduce the complexity of the system structure, some or all of the second feature region detectors can share the same structure, and a system suitable for each species can be constructed by changing the learning parameter values. Furthermore, as the second feature region detector, a feature region detector with the same construction as the first feature region detection step can be used; alternatively, a method can be considered to further reduce system complexity by only changing the learning parameter values ​​and the NMS method.

[0131] For one or more feature regions defined by the first feature region detection step and the first post-processing step, the species information detected in the first feature region detection step is used to determine which second feature region detector to use, and the determined second feature region detector is used to perform the second feature region detection step.

[0132] First, a preprocessing step is performed. At this stage, it is obvious that an inverse transformation process should be used during the coordinate transformation. In the second preprocessing step, the first feature region detected within the input image is transformed into the input image of the second feature region detector; therefore, defining the four points required for the transformation function as the four vertices of the first feature region is preferable.

[0133] The second feature region obtained by the second feature region detector is the value detected using the first feature region. Therefore, when calculating the second feature region within the overall input image, the first feature region should be taken into account.

[0134] For the second feature region obtained by the second feature region detector, a second feature value can be generated by performing an additional post-processing step similar to the first post-processing step. For example, to obtain image sharpness, a Sobel filter can be applied, or information such as the pose of the animal to be detected can be obtained by utilizing the detection and non-detection of feature regions and their relative positional relationships. In addition, image quality checks (e.g., focus blur, motion blur) can be performed as described later.

[0135] Using this additionally generated second feature value, it is possible to predict whether the feature region obtained in the second object detection step is suitable for use in the application domain associated with this application. It is obvious that the additionally generated second feature value can be appropriately designed according to the application domain. When the conditions for the second feature value defined in the application domain are not met, preferably, not only the second detection region but also the first detection region can be designed to conform to the data obtained in the application domain, such as excluding the first detection region from the detection results.

[0136] System Expansion

[0137] In this invention, an example is given of a system and configuration method that, by constituting a two-stage detection step, detects the location and species of an animal in a first specific location detection step and selects a detector to be used in a second specific location detection step based on the result.

[0138] This cascade configuration can be easily extended to a multi-layered cascade configuration. For example, it is possible to detect the entire animal in a first specific location detection step, detect the animal's facial and limb positions in a second specific location detection step, and detect the applied structures such as the positions of the eyes, nose, mouth, and ears on the face in a third specific location detection step.

[0139] By using this multi-layered cascaded structure, it is easy to design a system capable of simultaneously acquiring specific locations at various levels. When designing a multi-layered cascaded system, it is obvious that determining the number of layers requires considering factors such as the desired layer domain for the specific location, the overall system operating time, complexity, and the resources required to construct each individual feature region detector in order to design the optimal hierarchical structure.

[0140] Figure 7 This is a flowchart of a method for detecting an object corresponding to a pet's nose in a pet nose print management system according to the present invention.

[0141] The method for detecting an object for identifying a pet according to the present invention includes the steps of acquiring an original image of a pet (e.g., a puppy) (S710), determining a first feature region and the species of the pet by image processing of the original image (S720), and detecting an object for identifying the pet (e.g., a nose) within the first feature region based on the determined species of the pet (S730).

[0142] In step S710, while the application for pet object recognition is running, an original image including the pet is acquired using the activated camera. Here, illumination, focus, etc., can be adjusted to ensure successful pet photography. The acquired image can be provided as... Figure 5 as well as Figure 6 The input image is then used. The steps for determining the pet's species (S720) and detecting the pet as an object (S730) as described above can then be performed.

[0143] In step S720, a procedure for identifying the species of the pet is executed. According to the present invention, the step of determining the species of the pet (S720) may include a step of applying a first preprocessing to the original image, a step of identifying the species of the pet in the preprocessed image and setting the first feature region, and a step of extracting a first feature value by a first postprocessing to the first feature region.

[0144] The first preprocessing steps for the original image are as described above. Figure 5The steps described in S510 involve adjusting the size, scale, orientation, etc. of the original image to transform it into a form suitable for object detection.

[0145] The step of setting the first feature region is used to detect the area where pets exist in the image and the species of the pets. The purpose is to set the first feature region that can be used in the second feature region detection step, which will be described later, and at the same time select the second feature region detector that is most suitable for each pet species, thereby improving the final feature point detection performance.

[0146] According to the present invention, the step of setting the first feature region may include the steps of dividing the preprocessed image into multiple feature images using a learning neural network, applying a predefined bounding box to each of the multiple feature images, calculating a probability value within the bounding box according to the species of each pet, and when the calculated probability value is above a standard value for a specific animal species, constituting the first feature region as including the bounding box.

[0147] As mentioned earlier, the input image, normalized by the preprocessor, is layered through an artificial neural network to construct feature images from the first to the nth feature image. At this point, the methods for extracting feature images at each layer can be mechanically learned during the learning steps of the artificial neural network.

[0148] The extracted hierarchical feature images can be combined with the prior box catalogs corresponding to each hierarchical level to generate bounding boxes, individual categories, and confidence value catalogs. The results are output in the form shown in Table 1.

[0149] Subsequently, when the probability value of a specific animal species in a specific bounding box is above the standard value, the first feature region is defined as the region in which the bounding box can be included.

[0150] On the other hand, the process for determining the facial region (first feature region) of the pet as described above does not require high resolution and can therefore be performed at a relatively low resolution. That is, the step of applying the first preprocessing to the original image may include the step of transforming the original image to a first resolution image lower than the original resolution and the step of applying the first preprocessing to the image transformed to the first resolution.

[0151] On the other hand, if a first feature region is defined for identifying the pet's species, a first feature value is extracted through a first post-processing step targeting the first feature region. This is to first determine whether the dog's nose print image extracted from the acquired image is suitable as data for subsequent learning or recognition.

[0152] That is, when the first feature value is greater than the standard value, object detection for identifying pets is performed; when the first feature value is less than the standard value, no additional processing is performed, but processing for other images is performed instead. The first feature value may vary depending on the embodiment; for example, brightness information of the image to be processed can be used.

[0153] In step S730, the detection of an object used to identify the pet is performed. Various body parts can be used as objects for pet identification, such as eyes, nose, mouth, and ears, but the nose, which is used for nose prints, will be a representative example. This step is performed considering the species of the pet previously identified. When the pet is a puppy, object detection best suited for puppy identification can be performed. The best-suited object detection may vary depending on the animal species. Furthermore, when the captured image contains multiple pet species, object detection for identification can be performed for each individual animal.

[0154] The steps of detecting an object for identifying a pet may include applying a second preprocessing step to a first feature region for identifying the species of the pet, setting a second feature region for identifying the pet based on the species of the pet in the first feature region of the second preprocessing, and applying a second postprocessing step to the second feature region.

[0155] Similar to the first preprocessing, the second preprocessing for detecting objects used to identify pets involves processes such as adjusting image size. The second preprocessing for the first feature region can be performed at a second resolution higher than the first resolution where the first preprocessing was applied. This differs from the process of determining the animal species because the process of detecting objects used to identify pets (e.g., nose) and examining identification data (nose print image) requires a relatively high-quality image. Subsequently, the second feature region is defined for the preprocessed image as an object used to identify pets.

[0156] The step of setting the second feature region includes setting the second feature region (e.g., the nose region) based on the probability of setting an object (e.g., the nose) for identifying the pet in the first feature region (e.g., the face region) according to the pet's species. In the aforementioned S720 step, if the pet's species is determined, the most suitable single feature region detector and parameters can be selected according to that species, and the object for identifying the pet (e.g., the nose region) can be detected with lower computational complexity using the selected detector and parameters.

[0157] Post-processing can be performed to check whether an image of a pet, detected as a second feature region, is suitable for subsequent learning or recognition. As a result of the post-processing, a second feature value representing the image's suitability can be derived. When the second feature value is greater than a standard value, the image including the second feature region is transmitted to the server.

[0158] Figure 13 This is a block diagram of an electronic device 1300 according to the present invention. The electronic device 1300 according to the present invention may include a camera 1310, a processor 1320, a communication module 1330, a memory 1340, and a display 1350.

[0159] The camera 1310 may include a CCD (charge-coupled device) or CMOS (complementary metal-oxide semiconductor) that generates image signals from an optical module such as a lens and input light, and generates image data by capturing images, which is then provided to the processor 1320.

[0160] Processor 1320 controls the modules of electronic device 1300 and performs operations required for image processing. Processor 1320 may be composed of multiple microprocessors (processing circuits) depending on its function. As previously described, processor 1320 may detect an object (e.g., nose) for identification of a pet (e.g., a puppy) and perform a validity judgment on the image of that object.

[0161] The communication module 1330 can send or receive external entities and data via wired / wireless networks. In particular, the communication module 1330 learns or recognizes data exchanged for artificial intelligence-based processing through communication with a server.

[0162] Additionally, the electronic device 1300 may include a memory 1340 for storing necessary information for image data and image processing, and a display 1350 for outputting images to the user, and may include various modules depending on the application.

[0163] The electronic device 1300 according to the present invention includes a camera 1310 for generating an original image of a pet, a processor 1320 for determining a first feature region and the species of the pet by image processing of the original image, a processor 1320 for detecting an object for identifying the pet based on the determined pet species in the first feature region, and a communication module 1330 for transmitting an image of the object to a server if the object for identifying the pet is valid.

[0164] According to the present invention, the processor 1320 may apply a first preprocessing to the original image, determine the species of the pet in the preprocessed image and set a first feature region, and extract a first feature value by a first postprocessing to the first feature region.

[0165] According to the present invention, the processor 1320 uses a learning neural network to generate a plurality of feature images from the preprocessed image, and applies a predefined bounding box to each of the plurality of feature images, calculating a probability value within the bounding box according to the species of each pet. When the calculated probability value is above a standard value for a specific animal species, a first feature region including the bounding box can be constituted.

[0166] According to the present invention, when the first feature value is greater than the standard value, object detection for identifying pets is performed, and when the first feature value is less than the standard value, additional processing is omitted.

[0167] According to the present invention, the processor 1320 can transform the original image into an image with a first resolution smaller than the original resolution, and apply a first preprocessing to the image transformed to the first resolution.

[0168] According to the present invention, the processor 1320 may apply a second preprocessing to a first feature region for identifying the species of a pet, wherein a second feature region for pet identification is set based on the species of the pet in the first feature region of the second preprocessing, and a second postprocessing is applied to the second feature region to extract a second feature value.

[0169] According to the present invention, the second preprocessing for the first feature region can be performed at a second resolution that is higher than the first resolution of the first preprocessing applied to the settings for the first feature region.

[0170] According to the present invention, the processor 1320 can set a second feature region based on the probability that an object for identifying the pet is set in a first feature region according to the species of the pet.

[0171] According to the present invention, when the second feature value is greater than the standard value, the image including the second feature region can be transmitted to the server.

[0172] According to the present invention, processor 1320 can generate feature region candidates for determining the species of a pet in an image, and generate a first feature region with a determined location and size based on the confidence value of each feature region candidate.

[0173] Figure 9 This is a flowchart of a method for processing images of objects used to identify pets.

[0174] A method for processing an image of an object for identifying a pet according to the present invention includes the steps of acquiring an image of a pet (S910), generating candidate feature regions for determining the species of the pet in the image (S920), setting a first feature region for determining the location and size based on a confidence value of each of the candidate feature regions (S930), setting a second feature region for identifying the pet in the first feature region (S940), and acquiring an image of the object in the second feature region (S950).

[0175] According to the present invention, the step of generating feature region candidates may include the step of generating feature images hierarchically using an artificial neural network, the step of calculating the probability value of a specific species of pet in each boundary region for each feature image by applying a predefined boundary region, and the step of generating feature region candidates by taking the probability value into account.

[0176] The input image, normalized by the preprocessor, is used to generate feature images from the first feature image to the nth feature image in layers through an artificial neural network. The method of extracting feature images according to each layer can be mechanically learned in the learning steps of the artificial neural network.

[0177] The extracted layered feature images can be combined with a catalog of predefined boundary regions (bounding boxes) corresponding to each layer, and a catalog of probability values ​​to be set for specific animal species in the boundary regions, as shown in Table 1. Here, when it is impossible to determine whether it is a specific animal species, it can be defined as "background".

[0178] Subsequently, the extended NMS according to the invention is applied to generate candidates for first feature regions (feature region candidates) in each feature image for identifying a species such as a pet's face. Each feature region candidate can be derived using the previously derived probability values ​​according to a specific animal species.

[0179] According to the present invention, the step of generating candidate feature regions may include selecting a first boundary region in the feature image that has the highest probability of corresponding to a specific animal species, and calculating the overlap with the first boundary region in the feature image according to the probability values ​​of other boundary regions in the feature image, and including boundary regions whose overlap is greater than a reference overlap in the candidate feature regions of the feature image. In this case, for example, the area ratio of the intersection to the union of two boundary regions can be used for overlap evaluation.

[0180] That is, the following procedure can generate a file such as... Figure 8 Candidates for the characteristic region like (a).

[0181] 1. Perform the following procedure separately for each species except the background.

[0182] A. In the bounding box catalog, exclude boxes whose probability of being the species is below a certain threshold. If no boxes remain, end with no results.

[0183] B. In the bounding box catalog, the box with the highest probability for that species will be designated as the first box (first boundary region) and excluded from the bounding box catalog.

[0184] C. For the remaining bounding box directories, execute the following procedures in order of probability.

[0185] i. Calculate the overlap with the first box. For example, you can use the area ratio of the intersection to the union (Intersection over Union).

[0186] ii. If the overlap value is higher than a certain threshold, then its frame is a frame that overlaps with the first frame. It is then merged with the first frame.

[0187] D. Append the first box to the results box directory.

[0188] E. If there are still boxes remaining in the bounding box directory, then use the remaining boxes as objects and repeat from step C.

[0189] For two boxes A and B, for example, the ratio of the area of ​​their intersection to their union can be calculated as efficiently as in mathematical formula 1 described above.

[0190] Reference Figure 8 As mentioned above, a first feature region (e.g., the facial region of a pet) can be derived from each of the feature region candidates derived according to the present invention based on the confidence value of each feature region candidate.

[0191] According to the present invention, the center point C of the first feature region is set. (x,y) n As in Equation 2, the candidate center point C of the feature region can be selected. (x,y) 1 C (x,y) 2 Trust value p 1 p 2 The weighted sum is used to determine this.

[0192] According to the present invention, the width W of the first feature region n As in Equation 2, the confidence values ​​p for the widths W1 and W2 of the candidate feature regions can be used. 1 p 2 The height H of the first feature region is determined by the weighted sum. n The height H of the candidate feature region can be selected. 1 H 2 Trust value p1 p 2 The weighted sum is used to determine this.

[0193] This invention applies a weighted average based on confidence values ​​to multiple boxes (candidate feature regions), designating the box with the largest width and height as the first feature region (the pet's face region). A second feature region (the nose region) for pet identification is then detected within this first feature region. By further expanding the first feature region as described in this invention, the error of the second feature region not being detected during subsequent execution can be reduced.

[0194] Next, within the first feature region (e.g., the puppy's facial region), a detection for the second feature region (e.g., the nose region) used for pet identification is performed. This step takes into account the species of the pet previously identified. When the pet is a puppy, an object detection best suited for puppy identification can be performed. The best-suited object detection may vary depending on the animal species.

[0195] The step of setting the second feature region includes setting the second feature region (e.g., nose region) based on the probability of setting an object (e.g., nose) for identifying the pet in the first feature region (e.g., face region) according to the species of the pet.

[0196] Additionally, post-processing can be performed to check whether an image of a pet detected as a second feature region is suitable for subsequent learning or recognition. As a result of the post-processing, a second feature value representing the image's suitability can be derived. When the second feature value is greater than a standard value, the image including the second feature region is transmitted to the server; when the second feature value is less than the standard value, the image including the second feature region is discarded.

[0197] The electronic device 1300 according to the present invention includes a camera 1310 for generating an image including a pet and a processor 1330 for processing an image provided from the camera 1320 to generate an image of an object for identifying the pet. The processor 1330 may set feature region candidates for determining the species of the pet in the image, generate a first feature region for determining the location and size based on a confidence value of each of the feature region candidates, include setting a second feature region for identifying the object of the pet in the first feature region, and acquiring an image of the object in the second feature region.

[0198] According to the present invention, the processor 1320 can use an artificial neural network to hierarchically generate multiple feature images of the image, calculate a probability value for a specific species of pet to be set in each boundary region for each of the feature images, and generate the feature region candidate by taking the probability value into account.

[0199] According to the present invention, the processor 1320 selects a first boundary region in the feature image that corresponds to a specific animal species with the highest probability, and then, for other boundary regions in the feature image besides the selected first boundary region, calculates the overlap with the first boundary region in order of the probability values, and includes boundary regions with an overlap greater than a reference overlap as candidate feature regions in the feature image. At this time, for the purpose of overlap evaluation, for example, the area ratio of the intersection to the union of two boundary regions can be used.

[0200] According to the present invention, the center point of the first feature region can be determined by the weighted sum of the confidence values ​​of the candidate center points of the feature region.

[0201] According to the present invention, the width of the first feature region may be determined by a weighted sum of confidence values ​​for the widths of candidate feature regions, and the height of the first feature region may be determined by a weighted sum of confidence values ​​for the heights of candidate feature regions.

[0202] According to the present invention, the processor 1320 can control the camera 1310 to detect the changed position of the object in the next image, determine whether the image of the object whose position has changed in the next image is suitable for artificial intelligence-based learning or recognition, and perform the next shooting while setting the focus to the changed position.

[0203] According to the present invention, the processor 1320 may set a first feature region in an image for determining the species of a pet, and set a second feature region within the first feature region including an object for identifying the pet.

[0204] According to the present invention, the processor 1320 can determine whether the image of the object used to identify the pet is suitable for artificial intelligence-based learning or recognition.

[0205] According to the present invention, the processor 1320 can determine whether the quality of the image of the object meets the reference conditions. When the quality meets the reference conditions, the image of the object is transmitted to the server. When the quality does not meet the reference conditions, the image of the object is discarded and the camera 1310 is controlled to perform the shooting of the next image.

[0206] The image of the object derived through the process described above (e.g., a nose print image) is checked to determine its suitability for artificial intelligence-based learning or recognition. Image quality checks can be performed using various quality conditions, defined by the neural network designer. These conditions may include, for example, a photograph of a real dog, clear nose prints, absence of foreign objects, a frontal view, and a minimum proportion of surrounding blank space. Such conditions can be quantified and objectiveized. Storing poor-quality images in the neural network can degrade overall network performance; therefore, images below a certain quality standard are preferably filtered beforehand. These filtering processes can be performed in the first or second post-processing step described above.

[0207] As an example for inspecting the image quality of an object, a method for detecting quality degradation caused by large focus shift and quality degradation caused by camera or object shaking will be described.

[0208] The method for filtering images of objects for identifying pets according to the present invention includes the steps of acquiring an image of a pet (S1210), determining the species of the pet in the image and setting a first feature region (S1220), setting a second feature region for identifying the pet within the first feature region, taking into account the determined species of the pet (S1230), and checking the quality of the image of the object within the second feature region to determine whether the image of the object is suitable for artificial intelligence-based learning or recognition (S1240).

[0209] Furthermore, according to the present invention, post-processing (quality checking) is performed on the first feature region, and the second feature region detection and suitability judgment can only be performed when the image of the first feature region has appropriate quality. That is, the step of setting the first feature region includes checking the quality of the object's image in the first feature region to determine whether the object's image is suitable for artificial intelligence-based learning or recognition, and when it is determined that the object's image in the first feature region is suitable for artificial intelligence-based learning or recognition, the second feature region can be set. In the first feature region, if it is determined that the object's image is not suitable for artificial intelligence-based learning or recognition, the image of the current frame can be discarded and the image of the next frame can be captured.

[0210] The quality check (first post-processing) for the first feature region can also be omitted according to the embodiment. That is, the first post-processing process can be omitted and the second feature region detection can be performed directly.

[0211] The quality inspection of the object's image is performed by applying different weighting values ​​according to the location of the first feature region or the second feature region.

[0212] In the first post-processing step, the aforementioned brightness evaluation can be performed as a method for image quality inspection of the object. For example, the above mathematical formula 2 is performed on the first feature region to extract the brightness value specified by the BT.601 standard and the brightness (Value) information in the HSV color space on a pixel-by-pixel basis. When the average value is less than the first brightness reference value, the image is judged to be too dark; when the average value is greater than the second brightness reference value, the image is judged to be too bright. If the image is too dark or too bright, subsequent steps such as the second feature region detection can be omitted and the processing can be terminated. Alternatively, the judgment can be made by assigning weighted values ​​to regions determined to be important in the first feature region.

[0213] According to the present invention, the step of determining whether an image of an object is suitable for artificial intelligence-based learning or recognition may include determining the degree of defocus blur in the image of the object.

[0214] The method for detecting quality degradation due to focus shift (defocus blur) is as follows. Defocus blur refers to the phenomenon where the camera is out of focus, causing the target area (e.g., the NOT gate neighborhood) to become blurred. Examples of defocus blur include photos taken during automatic focus adjustment in a mobile phone camera.

[0215] To identify images experiencing defocus blur, high-frequency components (components with frequencies greater than a certain value) can be extracted. In an image, high-frequency components are primarily located at points where brightness and color change abruptly, i.e., object boundaries within the image, while low-frequency components are mainly located at points with similar brightness and color to their surroundings. Therefore, the better the focus and the sharper the image, the stronger the distribution of high-frequency components. This can be determined, for example, using the Laplacian operator. The Laplacian operator performs a second-order derivative on the input signal, retaining the high-frequency components and effectively removing the low-frequency components. Therefore, the Laplacian operator can be used to effectively find object boundaries within an image, and the sharpness of these boundaries can be numerically determined.

[0216] For example, by applying a 5x5 LoG (Laplacian of Gaussian) convolution operation, as shown in Equation 4 below, to the input photo, information about the location and sharpness of the boundary lines within the image can be obtained.

[0217]

Mathematical Expression 4

[0218]

[0219] Photos with less defocus blur and sharp edges will fall within the range of Laplacian operator values, from 0 to relatively large values. Conversely, photos with greater defocus blur and blurred edges will fall within the same range. Therefore, sharpness can be assessed by modeling the distribution of Laplacian operator values.

[0220] As an example of this method, applying the Laplacian operator allows us to assess image sharpness using the dispersion values. Alternatively, histogram analysis can yield a decimal distribution of the Laplacian values, and various statistical methods can be employed, such as calculating the distribution ratio of the highest and lowest intervals. This method can be selectively applied depending on the target application area.

[0221] That is, according to the present invention, the step of determining the degree of focus shift of the object may include the steps of applying a Laplacian operator that performs second-order differentiation on the image of the second feature region, extracting an image representing the distribution map of high-frequency components, and calculating a value representing the focus shift of the image of the second feature region from the distribution map of high-frequency components.

[0222] Within the nose region, the importance of sharpness varies depending on its location. Specifically, the central part of the nose is highly probable in the image, while the probability of sharpness increases towards the image edges, particularly in the outer corners of the nose or the hair areas around it. To reflect this spatial characteristic, one could divide the image into regions and assign different weights to each region to determine sharpness. For example, one could divide the image into nine parts, or draw an ellipse with the image center as a reference point to define a region of interest, and then multiply that region by a weighting value greater than 1 (w).

[0223] That is, according to the present invention, the weighting value applied to the central part of the second feature region can be set to be greater than the weighting value applied to the peripheral part of the second feature region. By applying a larger weighting value to the central part than to the peripheral part, image quality checks can be performed intensively on identifiable objects such as the nose print of a puppy.

[0224] The defocus blur score, determined using the Laplacian operator, indicates the presence of blurred boundaries in the image as it approaches 0, while a higher value indicates stronger boundary presence. Therefore, if the defocus blur score is greater than a threshold value, the image can be classified as sharp; otherwise, it is considered blurry. These thresholds can be determined through pre-collected experience or adaptively by accumulating and observing multiple input images in each camera session.

[0225] According to the present invention, the step of determining whether an object image is suitable for artificial intelligence-based learning or recognition may include the step of determining the degree of shaking of the object in the object image.

[0226] The following explains how to detect motion blur, a quality degradation caused by camera shake. Motion blur refers to the phenomenon where, during the camera's exposure time, the relative position of the subject and the camera shifts, causing the target area to shake as if the image were being captured. This type of blur can occur when a mobile phone camera is set to a long exposure time in low-light conditions, possibly due to the movement of a small dog during the exposure time or the user's hand tremor.

[0227] To analyze the features of these images, various edge detectors can be used. For example, the Canny edge detector is well-known for its effectiveness in detecting continuous, connected boundary lines.

[0228] exist Figure 10 An example of an image obtained by applying a Canny edge detector to an image that is swaying along its upper diagonal. For example... Figure 10 As shown, the results of applying the Canny edge detector to the image confirm that the boundary line of the nasal wrinkle region always occurs in the diagonal ( / ) direction.

[0229] By analyzing the directionality of the boundary line, it is possible to effectively determine whether it is swaying. Taking the directionality analysis method as an example, the boundary line detected by the Canny edge detector is always connected to the surrounding pixels. Therefore, the directionality can be analyzed by analyzing the connection relationship with the surrounding pixels. According to an embodiment of the present invention, the overall sway direction and degree can be calculated by analyzing the pattern distribution in a pixel block of a certain size where the boundary line is located in an image where the Canny edge detector is applied.

[0230] Figure 11 An example is shown of a pattern of pixel blocks used to determine whether a boundary line is wobbly in a result image of an image obtained by applying the Canny boundary line detector.

[0231] For example, such as Figure 11 As shown in (a), a detailed explanation will be given using the case where a 3x3 pixel boundary is detected as an example. For ease of explanation, as... Figure 11 As shown, each of the nine pixels is assigned a number based on its position.

[0232] For the central pixel 5, it can be assumed that it is always judged as a boundary line. If pixel 5 is not a boundary line, then the 3x3 pixel array is not a boundary line array, and therefore can be skipped or counted as a non-boundary line pixel.

[0233] When the central pixel 5 is a boundary line pixel, a total of 2 can be defined based on whether the other 8 pixels are boundary lines. 8 = 256 patterns. For example, in Figure 11 In (a), based on whether the boundary line of the {1st, 2nd, 3rd, 4th, 6th, 7th, 8th, 9th}th pixel is present, it is a (01000100) pattern, which can be converted to decimal and named the 68th pattern. These naming methods may be changed to facilitate implementation.

[0234] If the pattern is defined in this way, the start point, end point, and direction of the boundary line can be defined according to the pattern's configuration. For example, the 68th pattern can be defined as having a boundary line starting from the lower left corner (number 7) and ending at the top (number 2). Based on this, the corresponding pattern {diagonal upper right corner (↗) direction, steep angle} can be defined using patterns.

[0235] Analyze using the same method Figure 11 The pattern (b) is as follows. This pattern is (01010000) and can be named pattern number 80. The boundary line starts from the left (number 4) and ends at the top (number 2), so the corresponding pattern {diagonal upper right corner (↗) direction, middle angle} can be defined by the pattern.

[0236] Using this method, a lookup table of 256 patterns can be created. The possible combinations can then be defined in the following eight directions.

[0237] Vertical (↑)

[0238] Diagonal top right corner (↗) {steep, middle, shallow} angle

[0239] Horizontal (→)

[0240] Diagonal bottom right corner (↘) {shallow, medium, steep} angle

[0241] Building upon these methods, directional statistics of boundary line pixels can be compiled from the results images of the Canny edge detector. Based on this statistical information, it is possible to effectively determine whether motion blur has occurred in the relevant image. These judgment criteria can be derived from empirically designed classification methods or by utilizing machine learning methods, which are readily apparent from large amounts of data. Such methods can employ, for example, decision trees, random forests, or classifiers utilizing deep neural networks.

[0242] That is, according to the present invention, the step of determining the degree of shaking of an object may include, for example: Figure 10 As shown, a Canny edge detector is applied to the image of the second feature region to construct a boundary line image consisting of continuously connected boundary lines in the object image. Analysis is performed as follows: Figure 10 Such a boundary line image includes the distribution of the directional pattern of the blocks of the boundary line, and the step of calculating a value representing the degree of swaying of the object from the distribution of the directional pattern.

[0243] When compiling the aforementioned statistical information, it is obvious that the nose region contains more information than the surrounding regions. Therefore, methods such as collecting statistical information separately in certain regions within the image and assigning weighted values ​​can be used. As an example of this method, the method used in the defocus blur discrimination using the aforementioned Laplacian operator can be employed. That is, the step of calculating the degree of object sway from the distribution of the orientation pattern includes applying the weighted value of each block in the second feature region to calculate the distribution degree of the orientation pattern, and the weighted value of the block located at the center of the second feature region can be set to be greater than the weighted value of the block located at the periphery of the second feature region.

[0244] Figure 12 This is a flowchart of a method for filtering images of objects used to identify pets. The method for filtering images of objects used to identify pets according to the present invention includes the steps of acquiring an image including a pet (S1210), determining the species of the pet in the image and setting a first feature region (S1220), setting a second feature region including the object used to identify the pet within the first feature region, taking into account the determined pet species (S1230), and checking the quality of the object image in the second feature region to determine whether the object image is suitable for artificial intelligence-based learning or recognition (S1240). The quality check of the object image is performed by applying different weighting values ​​to the locations of the first or second feature regions.

[0245] On the other hand, after the step of setting the first feature region (S1220), a step of determining whether the object image is suitable for artificial intelligence-based learning or recognition can be performed by checking the quality of the object image in the first feature region (S1230). At this time, when it is determined that the object image in the first feature region is suitable for artificial intelligence-based learning or recognition, a second feature region can be set. According to the embodiment, the quality check (first post-processing) for the first feature region can also be omitted.

[0246] The step of determining whether an object image is suitable for AI-based learning or recognition by examining the quality of the object image in the first feature region may include determining whether the brightness in the first feature region falls within a reference range. This step may include extracting Luma information according to the BT.601 standard and brightness information in the HSV color space from the first feature region, and determining whether their average value is between a first threshold and a second threshold. When calculating the average value in this step, different weighting values ​​may be applied based on the location within the image.

[0247] According to the present invention, the step of determining whether an object image is suitable for artificial intelligence-based learning or recognition may include the step of determining the degree of object focus shift (defocus blur) in the object image.

[0248] According to the present invention, the step of determining the degree of focus shift of the object may include the steps of applying a Laplacian operator that performs second-order differentiation on the image of the second feature region to extract an image representing the distribution map of high-frequency components, and the step of calculating a value representing the focus shift of the image of the second feature region from the distribution map of high-frequency components.

[0249] According to the present invention, the weighting value applied to the center of the first feature region or the second feature region can be set to be greater than the weighting value applied to the periphery of the first feature region or the second feature region.

[0250] According to the present invention, the step of determining whether an object image is suitable for artificial intelligence-based learning or recognition may include the step of determining the degree of shaking of the object in the object image.

[0251] According to the present invention, the step of determining the degree of object sway may include the steps of applying a Canny edge detector to an image of a second feature region to form a boundary line image consisting of boundary lines that extend continuously in the object image, the steps of analyzing the directional pattern distribution of blocks including the boundary lines in the boundary line image, and the steps of calculating a value representing the degree of object sway from the distribution of the directional pattern.

[0252] According to the present invention, the step of calculating a value representing the degree of swaying of an object from the distribution of the orientation pattern includes the step of calculating the degree of distribution of the orientation pattern by applying a weighted value according to each block in the second feature region, and the weighted value of the block located at the center of the second feature region can be set to be greater than the weighted value of the block located at the periphery of the second feature region.

[0253] The electronic device 1300 according to the present invention includes a camera 1310 for generating images including pets and a processor 1320 for processing images provided from the camera 1310 to generate an image of an object for identifying the pet. The processor 1320 sets a first feature region for determining the species of the pet in the image, takes into account the determined pet species, sets a second feature region within the first feature region including an object for identifying the pet, and sets up a function to check the quality of the object image in the second feature region to determine whether the object image is suitable for artificial intelligence-based learning or recognition.

[0254] According to the present invention, the processor 1310 can check the quality of the object image in the first feature region to determine whether the object image is suitable for artificial intelligence-based learning or recognition. Here, the second feature region detection and quality check are only performed when the object image in the first feature region is suitable for artificial intelligence-based learning or recognition.

[0255] According to the present invention, the processor 1310 can determine whether the brightness in the first feature region falls within a reference range. According to an embodiment, the quality check (first post-processing) of the first feature region can be omitted.

[0256] Here, the quality inspection of the object image can be performed by applying different weighting values ​​based on the location of the first feature region or the second feature region.

[0257] According to the present invention, the processor 1310 can determine the degree of focus shift of the object from the object image.

[0258] According to the present invention, the processor 1310 can extract an image representing a high-frequency component distribution map from the image of the second feature region, and calculate the value of the focus offset of the image representing the second feature region from the distribution map of the high-frequency components.

[0259] According to the present invention, the weighting value applicable to the center of the first feature region or the second feature region can be set to be greater than the weighting value applicable to the periphery of the first feature region or the second feature region.

[0260] According to the present invention, the processor 1310 can determine the degree of shaking of the object from the image of the object.

[0261] According to the present invention, the processor 1310 can construct a boundary line image composed of the boundary lines of the image of the second feature region, analyze the distribution of the orientation pattern of the blocks including the boundary lines in the boundary line image, and calculate a value representing the degree of object swaying from the distribution of the orientation pattern.

[0262] According to the present invention, the processor 1310 may be configured to calculate the distribution degree of the orientation pattern by applying the weighted values ​​of the blocks in the second feature region, wherein the weighted value of the block located at the center of the second feature region is greater than the weighted value of the block located at the periphery of the second feature region.

[0263] This embodiment and the accompanying drawings are merely illustrative of a portion of the technical concept included in this invention. It is obvious that variations and specific embodiments that can be readily derived by those skilled in the art within the scope of the technical concept included in the specification and drawings of this invention are all included within the scope of the claims of this invention.

[0264] Therefore, the concept of the present invention should not be limited to the illustrated embodiments, not only to the appended claims, but also to all concepts that are equivalent or modified from the claims.

Claims

1. A method for detecting an object used to identify a pet, wherein, include: The steps to obtain an original image of the pet; The steps of determining a first feature region corresponding to the facial region of the pet and the species of the pet through image processing of the original image; as well as Based on the determined pet species, a second feature region is defined within the first feature region, thereby detecting an object used to identify the pet. The steps for detecting the object used to identify the pet include: The step of selecting a feature region detector corresponding to a specific species of the pet from multiple independent feature region detectors that are species-specific for each animal; as well as The step of defining the second feature region within the first feature region using the selected feature region detector. The feature region detector uses an algorithm that employs an artificial neural network to detect objects of various sizes in a single image.

2. The method according to claim 1, wherein, The steps for determining the species of the pet include: The first preprocessing step is applied to the original image; The steps of determining the species of the pet in the preprocessed image and setting the first feature region; The step of extracting the first feature value through a first post-processing of the first feature region.

3. The method according to claim 2, wherein, The steps for setting the first feature region include: The step of generating multiple feature images from the preprocessed image using a learning neural network; The step of applying a predefined bounding box to each of the plurality of feature images; The steps of calculating probability values ​​for each pet species within the bounding box; and When the calculated probability value is above the standard value for a specific animal species, the first feature region is defined as including the bounding box.

4. The method according to claim 2, wherein, When the first feature value is greater than the standard value, object detection for identifying the pet is performed. When the first feature value is less than the standard value, the additional processing is omitted.

5. The method according to claim 2, wherein, The first preprocessing step applied to the original image includes: The step of transforming the original image into a first resolution image with a resolution lower than that of the original image; and The first preprocessing is applied to the image transformed to the first resolution.

6. The method according to claim 1, wherein, The steps for detecting the object used to identify the pet include: The second preprocessing step is applied to the first characteristic region used to identify the species of the pet; The step of setting a second feature region for identifying the pet based on the pet's species in the first feature region of the second preprocessing; and The step of extracting the second feature value by applying a second post-processing to the second feature region.

7. The method according to claim 6, wherein, The second preprocessing for the first feature region is performed at a second resolution higher than the first resolution of the first preprocessing applied to the first feature region.

8. The method according to claim 6, wherein, The steps for setting the second feature region include: The step of setting the second feature region based on the probability that an object for identifying the pet exists in the first feature region according to the species of the pet.

9. The method according to claim 6, wherein, When the second feature value is greater than the standard value, the image including the second feature region is transmitted to the server.

10. The method according to claim 1, wherein, The steps for generating the first feature region include: The steps of generating candidate feature regions for determining the species of the pet in the image; and The step of generating a first feature region with a determined position and size based on the confidence value of each of the candidate feature regions.

11. An electronic device for detecting an object used to identify a pet, wherein, The electronic device includes: The camera generates an original image including the pet; The processor determines a first feature region corresponding to the facial region of the pet and the species of the pet through image processing of the original image, and sets a second feature region within the first feature region based on the determined pet species, thereby detecting an object for identifying the pet; and The communication module, when valid for identifying the object of the pet, transmits the image of the object to the server. The processor selects a feature region detector corresponding to the determined species of the pet from a plurality of independent feature region detectors that are species-specific for each animal, and uses the selected feature region detector to set the second feature region within the first feature region. The feature region detector uses an algorithm that employs an artificial neural network to detect objects of various sizes in a single image.

12. The electronic device according to claim 11, wherein, The processor applies a first preprocessing step to the original image. And in the preprocessed image, the species of the pet is determined and the first feature region is set. And the first feature value is extracted through a first post-processing of the first feature region.

13. The electronic device according to claim 12, wherein, The processor uses a learned neural network to generate multiple feature images from the preprocessed image. And a predefined bounding box is applied to each of the multiple feature images. Furthermore, probability values ​​are calculated for each pet species within the bounding box. Furthermore, when the calculated probability value is above the standard value for a specific animal species, the first feature region is defined as including the bounding box.

14. The electronic device according to claim 12, wherein, When the first feature value is greater than the standard value, object detection for identifying the pet is performed. When the first feature value is less than the standard value, the additional processing is omitted.

15. The electronic device according to claim 12, wherein, The processor transforms the original image into a first resolution image that is lower than the original resolution. The first preprocessing is then applied to the image transformed to the first resolution.

16. The electronic device according to claim 11, wherein, The processor applies a second preprocessing step targeting a first feature region used to identify the species of the pet. And in the first feature region of the second preprocessing, a second feature region for identifying the pet is set based on the pet's species. Furthermore, a second post-processing step is applied to the second feature region to extract the second feature value.

17. The electronic device according to claim 16, wherein, The second preprocessing for the first feature region is performed at a second resolution higher than the first resolution of the first preprocessing applied to the first feature region.

18. The electronic device according to claim 16, wherein, The processor sets the second feature region based on the probability that an object for identifying the pet exists in the first feature region according to the species of the pet.

19. The electronic device according to claim 16, wherein, When the second feature value is greater than the standard value, the image including the second feature region is transmitted to the server.

20. The electronic device according to claim 11, wherein, The processor generates candidate feature regions for identifying the species of the pet in the image. A first feature region with a determined position and size is generated based on the confidence value of each of the candidate feature regions.

Citation Information

Patent Citations

  • Computer program and terminal for providing information about individual animals on basis of animal face and noseprint images

    WO2020075888A1