Animal face region-of-interest extraction method and device and medium
By performing key point detection and affine transformation on pet facial images, the difficulty in ROI extraction caused by changes in pet facial posture is solved, and the accuracy of pet facial recognition and the precision of ROI areas are improved.
Patent Information
- Application Number
- CN202511172894.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
In pet facial recognition, the position and shape of key feature points change due to the unpredictable changes in posture and expression, making it difficult to extract complete and accurate regions of interest, affecting the recognition effect.
By detecting facial key points on animal facial images, generating a key point set, calculating the coordinates after rotation correction, generating an affine transformation matrix, and using the affine transformation matrix to transform the region of interest in the image, the spatially aligned region of interest is output.
It effectively solves the difficulty of ROI extraction caused by changes in pet facial posture, improves the accuracy of ROI area extraction, and filters out low-quality images through facial quality detection to provide high-quality data.
Smart Images

Figure CN120673046A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of image processing technology. More specifically, the present disclosure relates to a method, device, and medium for extracting a region of interest (ROI) from an animal face. Background Art
[0002] With the gradual relaxation of domestic pet ownership policies, China's pet industry is entering a golden period of rapid development. Accurate identification of pets and their specific identities are essential in many everyday scenarios, such as pet vaccinations, pet insurance processing, pet community management, and multi-pet feeding. This is where pet identification technology emerges. Similar in function to facial recognition, it aims to accurately determine a pet's identity, providing strong support for pet management and services, and helping the pet industry flourish in a more standardized, orderly, and refined direction.
[0003] Common pet identification solutions include: implanting a chip in the pet's body for positioning; utilizing the pet's facial features for identification; and utilizing the pet's nose prints for identification. For example, in the context of pet facial feature recognition, a region of interest (ROI) refers to key, unique and recognizable areas of a pet's face, such as the eyes, nose, mouth, and ears. By focusing on these areas, extracting and analyzing them, the pet's identity can be identified. However, unlike humans, pets cannot strike poses according to commands; their postures and expressions are constantly changing. When a pet's face is tilted, turned sideways, or exhibits a specific expression, the position and shape of key feature points change significantly, making it difficult for the system to extract a complete and accurate ROI, thus affecting recognition effectiveness.
[0004] In view of this, there is an urgent need to provide an animal facial region of interest extraction solution to improve the accuracy of pet facial recognition. Summary of the Invention
[0005] In order to at least solve one or more technical problems mentioned above, the present disclosure proposes a solution for extracting an animal facial region of interest in multiple aspects.
[0006] In a first aspect, the present disclosure provides a method for extracting regions of interest (ROIs) on animal faces, the method comprising: performing facial key point detection on an image to be processed containing an animal face to generate a set of key points; calculating the rotation-corrected coordinates of each point in the key point set based on the positional relationship of a preset subset of key points in the key point set under a standard frontal face posture; generating reference points for aligning the regions of interest based on the rotation-corrected coordinates to determine an affine transformation matrix; and performing an affine transformation on the regions of interest in the image to be processed using the affine transformation matrix to output spatially aligned regions of interest.
[0007] In a second aspect, the present disclosure provides an electronic device comprising: a processor; and a memory storing computer instructions implemented by a computer for extracting an area of interest on an animal's face. When the computer instructions are executed by the processor, the electronic device executes the method described in the first aspect.
[0008] In a third aspect, the present disclosure provides a computer-readable storage medium containing computer-implemented program instructions for extracting an animal face region of interest. When the program instructions are executed by a processor, the method described in the first aspect is implemented.
[0009] Through the animal facial region of interest extraction method, device and medium provided above, the disclosed embodiment generates a key point set by performing facial key point detection on the image to be processed containing the animal face; calculates the rotation-corrected coordinates of each point in the key point set based on the positional relationship of a preset key point subset in the key point set under a standard frontal face posture; generates reference points for aligning the region of interest based on the rotation-corrected coordinates to determine the affine transformation matrix; uses the affine transformation matrix to perform affine transformation on the region of interest in the image to be processed to output a spatially aligned region of interest, which effectively solves the problem of difficulty in extracting the pet face region of interest due to reasons such as the independent rotation of the pet's ears through affine transformation.
[0010] Furthermore, in some embodiments, by correcting the reference points of the affine transformation, the problem of large differences in the cropped facial regions when the pet's posture changes greatly is solved, thereby effectively improving the accuracy of ROI region extraction.
[0011] Furthermore, in some embodiments, low-quality images are filtered out through face quality detection, thereby providing high-quality image data for ROI region extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein: Figure 1 An exemplary flow chart of a method 100 for extracting an animal face region of interest provided by some embodiments of the present disclosure is shown; Figure 2 (a) to Figure 2 (d) show an exemplary process of extracting the region of interest of an animal face; Figure 3 An exemplary flow chart of a method 300 for extracting an animal face region of interest provided by some other embodiments of the present disclosure is shown; Figure 4FIG2 shows an exemplary flow chart of a method 400 for extracting an animal face region of interest provided by yet other embodiments of the present disclosure; Figure 5 A schematic diagram showing distribution features of animal faces provided by some embodiments of the present disclosure; FIG6 (a) to FIG6 (c) are schematic diagrams showing images identified as low quality according to some embodiments of the present disclosure; FIG7 (a) to FIG7 (b) show schematic diagrams of image quality evaluation indicators provided by some embodiments of the present disclosure; Figure 8 A schematic structural diagram of an electronic device provided by some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.
[0014] It should be understood that the terms “include” and “comprising” used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0015] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0016] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0017] When using pet facial feature information for recognition, pets cannot strike poses according to commands like humans do; their postures and expressions are highly variable. When a pet's face is tilted, turned sideways, or in a specific expression, the position and shape of key feature points can significantly change, making it difficult for the system to extract a complete and accurate ROI. Accurately extracting facial regions of interest (ROIs) is a critical step in the preprocessing stage of facial feature extraction. Ensuring consistent facial ROI extraction across various angles and scenarios can reduce the subsequent training challenge and further improve recognition accuracy. However, pet faces differ significantly from human faces. Take cats and humans as an example: human faces have diverse contours, are generally flat, and have evenly distributed facial features; in contrast, cats have shorter faces with a prominent snout and an overall triangular shape. Human ears are located on either side of the head and have a fixed shape; cat ears are located on either side of the head and can rotate independently. These differences complicate ROI extraction for cat faces.
[0018] In order to solve the problems in the above scenarios, this disclosure proposes a solution for extracting animal facial regions of interest. Through the key point position correction mechanism of the turning ratio, it effectively solves the problem of deviation in the nose position caused by the turning of the pet's head, and improves the accuracy of ROI extraction.
[0019] The following combination Figures 1 to 8 The disclosed solution is described in detail.
[0020] Figure 1 FIG. 1 shows an exemplary flow chart of a method 100 for extracting an animal face region of interest provided by some embodiments of the present disclosure. Figure 1 As shown, the method includes: Step S101 : performing facial key point detection on an image to be processed containing an animal face to generate a key point set.
[0021] Step S102 : calculating the rotation-corrected coordinates of each point in the key point set based on the positional relationship of a preset key point subset in the key point set under the standard frontal face posture.
[0022] Step S103 : generating reference points for alignment of the region of interest based on the rotation-corrected coordinates to determine an affine transformation matrix.
[0023] Step S104 : performing affine transformation on the ROI in the image to be processed using an affine transformation matrix to output a spatially aligned ROI.
[0024] Prior to the above steps, the method may further include obtaining an image to be processed that includes an animal face. The image to be processed may be captured by a camera, or may be obtained from the Internet, a standard image library, or other third-party image libraries. This disclosure does not limit the image acquisition method. For example, an image to be processed that includes an animal face may be obtained from an animal image library. The animal may refer to an animal whose facial three-dimensional structure has multiple prominent parts. For example, cats, canines, primates, equines, birds, etc. The prominent part of cats and canines is the nose.
[0025] In step S101, facial keypoint detection is performed on an image containing an animal face. The image containing an animal face can be input into a trained facial detection network, which outputs a set of keypoints for the animal face. Taking a cat face as an example, the keypoint set may include, but is not limited to, keypoints for the eye region, the nose region, the chin region, and the ear region. Optionally, as shown in Figure 2(a), the keypoint set may include keypoints for the eye region {0 for the left eye, 1 for the right eye}, keypoints for the nose region {2}, keypoints for the ear triangle contour {3 for the first ear root location of the left ear outer contour, 4 for the vertex of the left ear outer contour, 5 for the second ear root location of the left ear outer contour, 6 for the second ear root location of the right ear outer contour, 7 for the vertex of the right ear outer contour, and 8 for the first ear root location of the right ear outer contour}, and keypoints for the chin region {9}. These keypoints locate the outermost anchor points of the facial triangle (e.g., the eyes, ear roots, and chin contour points), have significant biological plausibility, and can adapt to the differences between different cat faces.
[0026] In step S102, the standard frontal face posture means that the geometric relationship between the key points on the animal's face satisfies that the line connecting the symmetrical key points is in a horizontal position, and the key point of the chin area is located directly below the midpoint of the line connecting the symmetrical key points.
[0027] The preset key point subset in the key point set refers to a subset of key points selected from the key point set. As shown in Figure 2(a), the preset key point subset in the key point set includes the relatively fixed symmetrical key points {5, 6} and the chin region key point {9}. In some embodiments, the preset key point subset in the key point set may further include the relatively fixed symmetrical key points {0, 1} and the chin region key point {9}, or the relatively fixed symmetrical key points {3, 8} and the chin location key point {9}.
[0028] By establishing a rotation correspondence between a preset key point subset and a standard frontal face posture, the inclination angle between the line connecting the symmetrical key points in the preset key point subset and the horizontal line is calculated to obtain the rotation angle.
[0029] In some embodiments, the preset key point subset includes two symmetrically positioned key points in the animal's face and a key point in the chin area. Based on the positional relationship of the preset key point subset in the key point set under the standard frontal face posture, the rotation-corrected coordinates of each point in the key point set in the rotation-corrected image are calculated, including: calculating the image rotation angle based on the positional relationship between the line connecting the two symmetrically positioned key points and the horizontal line in the image to be processed; performing the same rotation transformation on each point in the key point set based on the image rotation angle to generate the rotation-corrected coordinates.
[0030] In some embodiments, a rotation transformation may be performed on the image to be processed based on the image rotation angle to generate a rotation-corrected image, so that the face of the animal in the rotation-corrected image meets the requirement of a standard frontal face posture.
[0031] Assume that a preset keypoint subset includes the symmetrical ear base keypoints {5, 6} and the chin keypoint {9}. The midpoint of the line connecting the symmetrical keypoints {5, 6} is the rotation center. The line connecting the symmetrical keypoints {5, 6} is adjusted to be horizontal (i.e., the two points have the same vertical coordinate), so that the chin keypoint {9} is located below the line connecting the symmetrical keypoints {5, 6}. A rotation matrix A is generated based on the angle between the line connecting the symmetrical keypoints {5, 6} and the horizontal line, as well as the coordinates of the rotation center. Rotation matrix A is used to rotate the image to produce a rotated image, so that the animal's chin keypoint {9} is located below. If the chin is located above during the rotation, a further 180° rotation based on the angle can be used to move the chin keypoint {9} below.
[0032] Use the same rotation matrix A to transform the key point set to obtain the corrected key point coordinates. Key point set: eye area key points {0 is the left eye, 1 is the right eye}, nose position key points {2}, ear triangle contour area key points {3 is the first ear root position point of the left ear outer contour, 4 is the vertex of the left ear outer contour, 5 is the second ear root position point of the left ear outer contour, 6 is the second ear root position point of the right ear outer contour, 7 is the vertex of the right ear outer contour, 8 is the first ear root position point of the right ear outer contour}, chin area key points {9}. The corrected key point coordinates {landmarkRot[0], landmarkRot[1]...landmarkRot[9]} can be generated by the rotation matrix A. The corrected key point coordinates refer to the precise position coordinates of the key points in the original image (such as the eyes, nose, chin and other feature points of the pet's face) in the newly generated rotated corrected image after the same rotation transformation as the image.
[0033] In step S103, generating reference points for aligning the region of interest based on the rotationally corrected coordinates may include selecting a first key point subset and a second key point subset from the corrected coordinates, wherein the first key point subset includes key point coordinates corresponding to the first side eye, the first side ear base and the nose, and the second key point subset includes key point coordinates corresponding to the second side eye, the second side ear base and the nose, and the second side is opposite to the first side; determining the horizontal coordinate of the first reference point as the minimum value of the horizontal coordinates of all key points in the first key point subset; determining the vertical coordinate of the first reference point as the minimum value of the vertical coordinates of all key points in the first key point subset; determining the horizontal coordinate of the second reference point as the maximum value of the horizontal coordinates of all key points in the second key point subset; determining the vertical coordinate of the second reference point as the minimum value of the vertical coordinates of all key points in the second key point subset; and determining the third reference point as the coordinate corresponding to the key point of the chin area.
[0034] As shown in FIG2( b ), on the rotation-corrected image, the reference points for ROI alignment may include a first reference point R11 , a second reference point R22 , and a third reference point R33 , where R33 is a key point of the chin region.
[0035] The x-coordinate of the first reference point R11 is landmarkRot[0], landmarkRot[2], landmarkRot[3], and landmarkRot[5] (not shown in the figure), and the minimum value of the x-coordinate of these four points (min_x1); The vertical coordinate y of the first reference point R11 is landmarkRot[0], landmarkRot[2], landmarkRot[3], landmarkRot[5] (not shown in the figure), and the minimum value of the vertical coordinate y of these four points is (min_y1).
[0036] The minimum x value (min_x1) and the minimum y value (min_y1) are combined into a first reference point R11 (min_x1, min_y1).
[0037] The x-coordinate of the second reference point R22 is landmarkRot[1], landmarkRot[2], landmarkRot[6], and landmarkRot[8] (not shown in the figure), and the maximum value of the x-coordinate of these four points is (max_x2); The ordinate y of the second reference point R22 is landmarkRot[1], landmarkRot[2], landmarkRot[6], and landmarkRot[8] (not shown in the figure), and the minimum value of the ordinate y of these four points is (min_y2).
[0038] The third reference point R33 is landmarkRot[9], which is the coordinate of the chin position in the rotation-corrected image.
[0039] In step 104, an affine transformation is performed on the ROI in the image to be processed using an affine transformation matrix to output a spatially aligned ROI. An image alignment (affine transformation) process based on three sets of corresponding points can be used to align specific key points in the rotation-rectified image to the target position to generate a standardized aligned image (for example, Size). The value of w is related to the model input parameters. For example, the model input image requirement is , then w is 112. In some embodiments, the standardized alignment image can also be Size, where the values of w1 and w2 are not equal.
[0040] Here, it is assumed that the aligned images are unified as Size, w can be 112.
[0041] Determine three reference points in the rotation correction image, including {point11, point22, point33}; Three target reference points are determined in the aligned target image, including {point44, point55, point66}.
[0042] A mapping relationship is established between the three reference points determined in the rotation-corrected image and the three target reference points in the aligned target image. Point 11 is defined as the upper left anchor point point 44 (0, 0) of the aligned target image, point 22 is defined as the upper right anchor point point 55 (0, w) of the aligned target image, and point 33 is defined as the key point of the chin area as the lower right anchor point ( ,w). scaleWidth_new represents the recalculated steering offset ratio.
[0043] The affine transformation matrix B is calculated based on the coordinates of {point11, point22, point33} in the original image and {point44, point55, point66} in the aligned target image. The affine transformation matrix is applied to output the spatially aligned region of interest. The mapping relationship between points11 and point44 fixes the left boundary to ensure that the left side of the head is aligned with the left edge of the image. The mapping relationship between points22 and point55 fixes the lower boundary to ensure that the bottom of the chin is aligned with the lower edge of the image. The mapping relationship between points33 and point66 dynamically sets the right boundary to adapt the width to the rotation offset ratio. The lower right point dynamically simulates the region when the head is turned, achieving a normalized output to eliminate rotation and scale differences and maintain the alignment of key facial structures.
[0044] The above three-point alignment method can effectively handle the deformation caused by head rotation while maintaining the position consistency of key facial structures, providing standardized input for subsequent feature extraction and recognition.
[0045] Figure 3 FIG. 3 shows an exemplary flow chart of a method 300 for extracting an animal face region of interest provided by some embodiments of the present disclosure. It is understood that the method 300 is a method for extracting an animal face region of interest provided by some embodiments of the present disclosure. Figure 1 Therefore, the above combined with the further limitation and / or expansion of the method 100 Figure 1 The relevant detailed description of is also applicable to the following. The method further modifies the reference points used for alignment of the region of interest according to the profile posture of the animal's face.
[0046] like Figure 3 As shown, at step S304, it can include calculating the steering offset ratio based on the lateral offset between the first reference point, the second reference point and the third reference point; then, when it is determined that the animal's face is in a side face posture based on the steering offset ratio, the first reference point or the second reference point is corrected.
[0047] Calculating the steering offset ratio based on the lateral offsets between the first reference point, the second reference point and the third reference point may include calculating the first lateral offset between the horizontal coordinate of the first reference point and the horizontal coordinate of the second reference point; calculating the second lateral offset between the horizontal coordinate of the first reference point and the horizontal coordinate of the third reference point; and calculating the steering offset ratio based on the ratio of the first lateral offset to the second lateral offset.
[0048] Continuing with the reference points shown in FIG2 (b) as an example, calculate the first lateral offset X12 between the first reference point and the second reference point; calculate the second lateral offset X13 between the first reference point and the third reference point; and calculate the steering offset scaleWidth: .
[0049] In some embodiments, the profile face posture can be determined based on the distribution characteristics of the coordinates after rotation correction.
[0050] In other embodiments, an elevation scoring parameter for quantitatively evaluating the head-up state, a depression scoring parameter for quantitatively evaluating the head-down state, an auxiliary scoring parameter for assisting in the evaluation of the head-down state, and a yaw scoring parameter for quantitatively evaluating the left and right head turning angles can be defined based on the relationship between head posture changes and multiple distance measurements. The elevation scoring parameter, depression scoring parameter, auxiliary scoring parameter, and yaw scoring parameter are jointly judged using a preset threshold to determine the current posture of the animal's face. After first judging the posture of the animal's face, the degree of the animal's face turning is further judged based on the steering offset ratio, thereby triggering correction compensation for the rotation-corrected coordinates.
[0051] The head turn offset ratio is a normalized metric that quantifies the degree of horizontal (left-right) or vertical (head-up / head-down) head rotation of a pet. It is calculated by comparing the asymmetry of key facial points and converting the head turn ratio into a proportional value between 0% and 100%.
[0052] In some embodiments, when the key point set includes: key points of the eye region, key points of the ear root region, key points of the nose region, and key points of the chin region, when the animal's face is determined to be in a side profile posture based on the steering offset ratio, the first reference point or the second reference point is corrected, which may include determining the first reference point, the second reference point, and the third reference point for aligning the region of interest based on the coordinates after rotation correction; calculating the steering offset ratio based on the lateral offset between the first reference point, the second reference point, and the third reference point; when it is determined that the steering offset ratio is greater than a preset first threshold, reducing the horizontal coordinate of the first reference point based on the steering offset ratio to generate an adjusted first reference point. At this time, the reference points used for aligning the region of interest are updated to the adjusted first reference point, the second reference point, and the third reference point.
[0053] In some embodiments, when the steering offset ratio is determined to be less than a preset second threshold, the horizontal coordinate of the second reference point is increased based on the steering offset ratio to generate an adjusted second reference point. The reference points used for ROI alignment are then updated to include the first reference point, the adjusted second reference point, and the third reference point.
[0054] Then, after generating the adjusted first reference point or the adjusted second reference point, the steering offset ratio is recalculated using the horizontal coordinate of the adjusted first reference point; or the steering offset ratio is recalculated using the horizontal coordinate of the adjusted second reference point until convergence.
[0055] When the steering offset ratio indicates that the animal's face is in a profile state, the spatial position of the corrected key point coordinates is adaptively adjusted, including expanding the horizontal coordinate of the first reference point (i.e., adjusting the horizontal coordinate of the first reference point in a direction with smaller pixel coordinate values) when it is determined that the steering offset ratio is greater than a preset first threshold, and generating an adjusted first reference point; or, when it is determined that the steering offset ratio is less than a preset second threshold, expanding the second reference point (i.e., adjusting the horizontal coordinate of the first reference point in a direction with larger pixel coordinate values) and generating an adjusted second reference point.
[0056] Use the adjusted first reference point to update the first reference point obtained by preliminary calculation, or use the adjusted second reference point to update the second reference point obtained by preliminary calculation, and keep the third reference point unchanged, return to the step of calculating the steering offset ratio based on the positional relationship between the first reference point and the second reference point, and recalculate the new steering offset ratio until the new steering offset ratio falls into the preset interval or reaches the maximum number of iterations (i.e., convergence).
[0057] In extreme profile situations (e.g., when the scaleWidth value is extremely small or extremely large), the cat's head is facing left or right at a wide angle. The nose area will significantly protrude from the original ear base line (the line connecting R11 and R22). To more accurately depict the head contour, the ear base point position needs to be adaptively adjusted based on the turning direction.
[0058] For right-side face processing (when scaleWidth is large), when a right-side face posture is detected (scaleWidth>>preset threshold), the x-coordinate value of the left ear root point R11 is proportionally reduced to generate a new coordinate R11_new (the coordinate value is closer to the left). This adjustment moves the left ear root inward to better adapt to the contour changes caused by the protrusion of the nose. As shown in Figure 2 (c), the horizontal coordinate of the R11_new point shown in Figure 2 (c) is obtained by proportionally reducing the horizontal coordinate of the R11 point shown in Figure 2 (b). Assuming that the horizontal coordinate of the R11 point is x, according to the formula , thereby expanding R11 to ensure that the nose area can be cut into the target image.
[0059] Alternatively, for left-side facial processing (when scaleWidth is smaller), if a left-side facial pose is detected (scaleWidth << a preset threshold), the x-coordinate of the right ear root point R22 is proportionally increased to generate a new coordinate R22_new (not shown). This adjustment moves the right ear root outward, better matching the head profile when the nose is protruding.
[0060] This adaptive adjustment mechanism dynamically corrects the position of the ear base reference point so that the geometric model composed of key points can more accurately reflect the anatomical features of the protruding nose under extreme turning postures, providing a more accurate benchmark reference for subsequent posture analysis and contour modeling.
[0061] In some embodiments, when processing the right side of the face (when the scaleWidth value is larger), the third horizontal offset between the horizontal coordinate of the corrected new coordinate R11_new (the corrected first reference point) and the horizontal coordinate of the original second reference point point22Rot is calculated. .
[0062] Calculate the fourth horizontal offset between the horizontal coordinate of the corrected new coordinate point11Rot_new (the corrected first reference point) and the horizontal coordinate of the original third reference point point33Rot .
[0063] Then calculate the new steering offset ratio scaleWidth_new: .
[0064] When the new steering offset ratio falls into the preset range or reaches the maximum number of iterations, the subsequent method process is entered.
[0065] For left-side face processing (when scaleWidth is smaller), calculate the fifth horizontal offset between the horizontal coordinate of the original first reference point R11 and the horizontal coordinate of the new coordinate R22_new (the corrected second reference point) Calculate the second lateral offset X13 between the horizontal coordinate of the original first reference point R11 and the horizontal coordinate of the original third reference point R33. Then calculate the new steering offset ratio scaleWidth_new: .
[0066] When the new steering offset ratio falls into the preset range or reaches the maximum number of iterations, the subsequent method process is entered.
[0067] After the reference points are corrected, a corrected affine transformation matrix is determined based on the corrected reference points. Then, at step S305, an affine transformation is performed on the ROI in the image to be processed using the corrected affine transformation matrix to output a spatially aligned ROI.
[0068] After processing the right side of the face (when the scaleWidth value is large), mapping the corrected first reference point R11_new back to the position of the rotation correction image includes: -1 , and obtain point11 on the corresponding rotation-corrected image. Mapping the second reference point R22 back to the position of the rotation-corrected image includes: transforming R22 by the rotation inverse matrix A -1 , and obtain the corresponding point22 on the rotation-rectified image. The key point of the chin area of the cat face does not need to be transformed, and corresponds to the point33 landmark on the rotation-rectified image [9].
[0069] After processing the left side of the face (when the scaleWidth value is small), mapping the first reference point R11 back to the position of the rotation-corrected image includes: rotating R11 through the inverse rotation matrix A -1 , and obtain the point 11 on the corresponding rotation-corrected image. Mapping the corrected second reference point R22_new back to the position of the rotation-corrected image includes: transforming R22_new by the inverse rotation matrix A -1 , and obtain the corresponding point22 on the rotation-rectified image. The key point of the chin area of the cat face does not need to be transformed, and corresponds to the point33 landmark on the rotation-rectified image [9].
[0070] After the above processing, the spatially aligned region of interest is finally output, as shown in Figure 2 (d). When the animal face is a cat face, the ROI region extraction method provided by the embodiment of the present disclosure is used to overcome the problem that when the cat face is in a large-angle side face, the ear changes greatly, which has a significant impact on the ROI region extraction. In the embodiment of the present disclosure, during the ROI region extraction process, the information extraction of the ear region is reduced. After alignment, the points in the upper left corner mainly consider the positions of the key points {0, 2, 3, 5} of the cat face, and the points in the upper right corner after alignment mainly consider the positions of the key points {0, 2, 6, 8} of the cat face. The points in the lower right corner are dynamically aligned through the key points of the chin area, thereby eliminating rotation and scale differences and keeping the key facial structures aligned.
[0071] Furthermore, considering the relationship between different profiles of the face, which has a greater impact on the position of the nose, an adaptive compensation correction method is used to adjust the position of the reference point for calibrating and cropping the ear part in the rotation-corrected image, so that the nose part can be maintained in the output alignment image, thereby obtaining the most beneficial alignment effect with minimal deformation.
[0072] Figure 4 FIG4 shows an exemplary flow chart of a method 400 for extracting an animal face region of interest provided by another embodiment of the present disclosure. It can be understood that the method 400 is a method for extracting an animal face region of interest provided by another embodiment of the present disclosure. Figure 1 or Figure 3 Therefore, the above combined with the further limitation and / or expansion of the method 100 or method 300 Figure 1 and Figure 3 The relevant detailed description of also applies to the following. The method further performs image quality detection on the image to be processed based on the key point set; based on the image quality detection result, the image to be processed that does not meet the image quality detection requirements is filtered. Figure 4 As shown, the method includes: At step S402, an image quality check is performed on the image to be processed based on the key point set. If the image quality check result indicates no image quality issues, the method proceeds to the subsequent relevant steps. If the image quality check result indicates image quality issues, images to be processed that do not meet the image quality check requirements are filtered out.
[0073] In the above steps, performing image quality detection on the image to be processed based on the key point set may include at least one of the following: abnormal facial posture detection; detection of the proportion of effective pixels occupied by the face in the image; and detection of the relationship between the face and the edge of the image.
[0074] In some embodiments, abnormal facial posture detection may include: selecting a third key point subset from the key point set, the third key point subset including two symmetrical key points in the eye area, two symmetrical key points in the ear base area, a key point in the nose area, and a key point in the chin area; judging whether the animal's face is in an abnormal posture based on the positional relationship of the key points in the third key point subset, the abnormal posture including at least one of the following: facial abnormality caused by the head elevation angle; facial abnormality caused by the head depression angle; facial abnormality caused by the head yaw angle.
[0075] When the abnormal posture is that the head elevation angle causes facial abnormality and / or the head yaw angle causes facial abnormality, whether the animal's face is in an abnormal posture is judged based on the positional relationship of the key points in the third key point subset, including: calculating a first distance value based on two symmetrical key points in the eye area; calculating a second distance value based on the midpoint position of the line connecting the two symmetrical key points in the eye area and the key points in the nose area; judging whether the head elevation angle causes facial abnormality based on the relationship between the ratio of the first distance value and the second distance value and a preset elevation angle threshold; and / or judging whether the head yaw angle causes facial abnormality based on the relationship between the ratio of the first distance value and the second distance value and a preset yaw angle threshold.
[0076] When the abnormal posture is caused by the head pitch angle causing facial abnormality, judging whether the animal's face is in an abnormal posture based on the positional relationship of the key points in the third key point subset can include: calculating a first distance value based on two symmetrical key points in the eye area; calculating a third distance value based on two symmetrical key points in the ear root area; calculating a fourth distance value based on the key points in the nose area and the key points in the chin area; judging whether the head pitch angle is abnormal based on the relationship between the ratio of the first distance value and the second distance value and a preset elevation angle threshold; judging whether the head yaw angle is abnormal based on the relationship between the ratio of the first distance value and the second distance value and a preset yaw angle threshold; judging whether the head pitch angle is abnormal based on the ratio of the first distance value and the third distance value and the ratio of the first distance value to the fourth distance value.
[0077] The eye distance DistEyes between the two eyes is calculated based on the key point 0 at the left eye position and the key point 1 at the right eye position, such as Figure 5 As shown, line segment D1 represents the first distance value.
[0078] Get the coordinates of the midpoint of the line connecting the two eyes, and then calculate the nose offset distance DistMidEyesNose between the midpoint and the nose position key point 2, such as Figure 5 As shown, line segment D2 represents the second distance value.
[0079] An elevation score parameter, ElevaScore, is calculated based on the second distance value and the first distance value. For example, the elevation score parameter can be calculated based on the ratio of the second distance value to the first distance value, or based on the ratio of the first distance value to the second distance value. Assuming the elevation score parameter is calculated based on the ratio of the second distance value to the first distance value, when the elevation score parameter is greater than a preset elevation threshold, it indicates that the elevation angle is too high. An image with an excessively high elevation angle is shown in Figure 6(a).
[0080] In some embodiments, a yaw angle scoring parameter, YawScore, can also be calculated based on the second distance value and the first distance value. For example, the yaw angle scoring parameter can be calculated based on the ratio of the second distance value to the first distance value, or based on the ratio of the first distance value to the second distance value. Assuming the yaw angle scoring parameter is calculated based on the second distance value and the first distance value, if the yaw angle scoring parameter is less than a preset yaw threshold, the yaw angle is determined to be excessive. An image with an excessively large yaw angle is shown in Figure 6(b).
[0081] When an animal is in the head-up position or in a head-turning posture, the distance between the eyes remains essentially unchanged. However, in the head-up position, the distance between the nose and the midpoint of the line connecting the two eyes decreases. In a head-turning posture, the nose deviates laterally, causing the distance between the nose and the midpoint of the line connecting the two eyes to increase. Therefore, the present disclosure utilizes the different sensitivities of eye-nose distance to elevation and yaw angles, combined with the stability of eye distance, to achieve detection of both states using the same ratio.
[0082] In some embodiments, the first midpoint coordinate value of the key points inside the two ears is calculated based on the key points inside the two ears. The second midpoint coordinate value of the key points of the two eyes is calculated based on the key points of the two eyes, and the third distance value DistMidEyesMidEars between the two midpoints is calculated based on the first midpoint coordinate value and the second midpoint coordinate value. Figure 5 As shown, line segment D3 represents the third distance value.
[0083] The fourth distance value DistNoseChin is calculated based on the key points of the nose position and the key points of the chin position. Figure 5 As shown, line segment D4 represents the fourth distance value.
[0084] A depression angle scoring parameter, DepressScore, is calculated based on the first and third distance values. For example, the depression angle scoring parameter can be calculated based on the ratio of the third distance value to the first distance value, or based on the ratio of the first and third distance values. An auxiliary scoring parameter, AssistScore, is calculated based on the first and fourth distance values. For example, the auxiliary scoring parameter can be calculated based on the ratio of the fourth distance value to the first distance value, or based on the ratio of the first and fourth distance values.
[0085] Assume that the ratio of the first distance value to the third distance value is used to calculate the depression angle scoring parameter, and the ratio of the first distance value to the third distance value is used to calculate the auxiliary scoring parameter. When the depression angle scoring parameter is less than the preset depression angle threshold and the auxiliary scoring parameter is greater than the preset auxiliary scoring threshold, it indicates that the depression angle is too large. An image with an excessively large depression angle is shown in Figure 6(c).
[0086] When an animal's face is in a state of lowering its head, the stability of the first distance value during the lowering process is utilized. The third distance value between the midpoint of the eye and the midpoint of the ear will increase when the head is lowered, and the fourth distance value will decrease when the head is lowered. These are used to jointly detect the problem of the animal lowering its head at an excessive angle, effectively improving the accuracy of the detection results.
[0087] This disclosure proposes using a multi-feature collaborative approach to identify issues such as excessive elevation (looking up), low pitch (looking down), or excessive yaw (wide-angle profile) in images. These issues indicate serious posture deviations and should be discarded. The image collection that has undergone image detection can provide useful, high-quality data for ROI extraction, thereby improving the accuracy of ROI proposal.
[0088] In other embodiments, the detection of the effective pixel ratio of the face in the image may include: calculating the diameter length value of the minimum enclosing circle based on the positional relationship of all key points in the key point set, the minimum enclosing circle being the smallest circle that contains all key points in the key point set; if the ratio of the diameter length value to the diagonal length of the image to be processed is less than a preset first distance threshold, it indicates that the effective pixel ratio of the face in the image is too low; or, if the ratio of the diameter length value to the diagonal length of the image to be processed is greater than a preset second distance threshold, it indicates that the effective pixel ratio of the face in the image is too high.
[0089] As shown in Figure 7(a), the keypoint set includes: eye keypoints {0 for the left eye, 1 for the right eye}, nose keypoints {2}, ear triangle contour keypoints {3 for the first ear root of the left ear, 4 for the vertex of the left ear, 5 for the second ear root of the left ear, 6 for the second ear root of the right ear, 7 for the vertex of the right ear, and 8 for the first ear root of the right ear}, and chin keypoints {9}. The minimum enclosing circle is the smallest circle that encompasses all keypoints in a given keypoint set. For example, OpenCV can be used to calculate the minimum enclosing circle and obtain its diameter. The relationship between the diameter and a threshold is used to determine whether the face occupies a normal proportion of the image's effective pixels. If the diameter is below the threshold, the face's proportion is too low, making it impossible to detect specific details; if the diameter is above the threshold, the face's proportion is too high, potentially resulting in loss of details. The actual threshold can be adjusted based on the specific application and camera focal length.
[0090] The disclosed embodiment quantifies the proportion of the face in the image through the minimum enclosing circle diameter, effectively solving the quality control problems of faces that are too small (insufficient recognition features) and faces that are too large (loss of edge features).
[0091] In some embodiments, detecting the relationship between a face and an image edge can include calculating a minimum bounding rectangle based on all key points in a key point set and obtaining the coordinate values of its four vertices. The minimum bounding rectangle is the rectangle with the smallest area that encompasses all key points in the key point set. If the shortest distance from any vertex to the four edges of the image is less than a preset distance threshold, the face is determined to be too close to the edge of the image. As shown in Figure 7(b), the minimum bounding rectangle includes all key points in the key point set. For example, the minimum bounding rectangle can be calculated using OpenCV, and the coordinate values of the four vertices can be simultaneously obtained. Based on the coordinate values of the four vertices and the shortest distance to the four edges of the image, it is determined whether the facial area is too close to the image edge, which may result in the truncation of important features.
[0092] When an object rotates, the orientation of the minimum enclosing rectangle directly reflects the object's direction, while the minimum enclosing circle cannot provide directional information. Calculating the distance between the rectangle's vertices and the image edge allows for a more accurate determination of their relationship. When a cat's head tilts, the minimum enclosing rectangle rotates with it, accurately reflecting its actual space occupancy.
[0093] The disclosed embodiment uses a minimum bounding rectangle for rapid initial screening (small computational effort), and then accurately detects the minimum bounding rectangle of the image that passes the initial screening, thereby constructing a hierarchical quality assessment system that takes into account both efficiency and accuracy.
[0094] If in Figure 1 On this basis, after performing image quality detection, the rotation correction processing is calculated for the image to be processed that meets the image quality detection, and then the reference points for aligning the region of interest are generated to determine the affine transformation matrix, and then the affine transformation matrix is used to perform affine transformation on the region of interest in the image to be processed to output the spatially aligned region of interest.
[0095] If in Figure 3 Based on the image quality test, after performing image quality testing, a rotation correction process is calculated for the image to be processed that meets the image quality test, and then reference points for ROI alignment are generated. Then, the reference points corresponding to the ROI are corrected according to the profile posture of the animal's face. After the reference points are corrected, a corrected affine transformation matrix is determined based on the corrected reference points. Finally, in step S406, an affine transformation is performed on the ROI in the image to be processed using the corrected affine transformation matrix to output a spatially aligned ROI.
[0096] In some embodiments, the method may further include defining an elevation scoring parameter for quantitatively evaluating the head-raising state, a depression scoring parameter for quantitatively evaluating the head-lowering state, an auxiliary scoring parameter for assisting in the evaluation of the head-lowering state, and a yaw scoring parameter for quantitatively evaluating the left and right head turning angles based on the relationship between the head posture change and multiple distance measurement values; and jointly judging the elevation scoring parameter, depression scoring parameter, auxiliary scoring parameter, and yaw scoring parameter through a preset threshold to determine the current posture of the animal's face.
[0097] First, define complete thresholds for all parameters. These parameters include the elevation scoring parameter ElevaScore, the depression scoring parameter DepressScore, the auxiliary scoring parameter AssistScore, and the new yaw scoring parameter ScaleYawScore. For each parameter, set a threshold range based on specific detection requirements. For example, the elevation scoring parameter sets the elevation threshold range; the depression scoring parameter sets the depression threshold range; the auxiliary scoring parameter sets the auxiliary scoring threshold range; and the new yaw scoring parameter sets the yaw threshold range.
[0098] The new yaw angle score parameter is calculated based on certain keypoints in the keypoint set: the ear triangle contour keypoints {3 is the first ear root location point of the left ear's outer contour, 5 is the second ear root location point of the left ear's outer contour, 6 is the second ear root location point of the right ear's outer contour, and 8 is the first ear root location point of the right ear's outer contour}, and the chin region keypoint {9}. Reference point 1 is formed by selecting the minimum horizontal coordinate and the minimum vertical coordinate from the horizontal coordinates of keypoints {3, 5}. Reference point 2 is formed by selecting the maximum horizontal coordinate and the minimum vertical coordinate from the vertical coordinates of keypoints {6, 8}. Using chin region keypoint {9} as reference point 3, the lateral offset P12 between reference points 1 and 2 is calculated, as is the lateral offset P13 between reference points 1 and 3. The new yaw angle score parameter, ScaleYawScore, is calculated based on the ratio of offsets P12 and P13.
[0099] Assuming a single threshold judgment method is adopted, the five-posture classification rules are defined by combining the elevation scoring parameter ElevaScore, the depression scoring parameter DepressScore, the auxiliary scoring parameter AssistScore, and the new yaw angle scoring parameter ScaleYawScore.
[0100] For example, when the elevation angle scoring parameter ElevaScore, the depression angle scoring parameter DepressScore, the auxiliary scoring parameter AssistScore, and the new yaw angle scoring parameter ScaleYawScore all fall within their prescribed threshold ranges, the posture is defined as a frontal face.
[0101] When the depression angle scoring parameter DepressScore exceeds its specified threshold range, and the auxiliary scoring parameter AssistScore and the new yaw angle scoring parameter ScaleYawScore do not exceed their specified threshold ranges, this posture is defined as a head-down face.
[0102] When the elevation angle scoring parameter ElevaScore exceeds its specified threshold range, and the depression angle scoring parameter DepressScore and the new yaw angle scoring parameter ScaleYawScore do not exceed their specified threshold ranges, this posture is defined as an upward-looking face.
[0103] When the elevation angle scoring parameter ElevaScore and the depression angle scoring parameter DepressScore are both within their specified threshold ranges, and the new yaw angle scoring parameter ScaleYawScore is greater than the preset left face deflection threshold, this posture is defined as a left-turned face.
[0104] When the elevation angle scoring parameter ElevaScore and the depression angle scoring parameter DepressScore are both within their specified threshold ranges, and the new yaw angle scoring parameter ScaleYawScore is less than the preset right face deflection threshold, this posture is defined as a right-turned face.
[0105] Other postures include all other postures that do not meet any of the above conditions or whose parameters exceed the maximum threshold.
[0106] The disclosed embodiment can test the logic through a large amount of high-quality cat face data and adjust the threshold range to reduce the classification error rate. "Other types of postures" can also reasonably cover low-quality images (such as extreme postures). For example, if the elevation angle scoring parameter ElevaScore does not exceed its specified threshold range, the depression angle scoring parameter DepressScore does not exceed its specified threshold range, and the new yaw angle scoring parameter ScaleYawScore is greater than the preset left face deflection threshold, it is defined as a left-turned face. Alternatively, the posture range can be expanded by setting a floating range of the specified threshold range. For example, if the current elevation angle scoring parameter ElevaScore is within the floating range of the specified threshold range, it is considered to be a slight head raise, the depression angle scoring parameter DepressScore does not exceed its specified threshold range, and the new yaw angle scoring parameter ScaleYawScore is greater than the preset left face deflection threshold, then the facial posture is in a head-up and left-turned state. Since the head-up degree is within the floating range, this posture can still be classified as a left-turned face state. However, when the elevation angle scoring parameter ElevaScore exceeds the maximum threshold allowed by the elevation angle, only one indicator is needed to determine that the facial posture at this time is excessively raised (i.e., the elevation angle is too large), and the image is classified as excessively high elevation. The image can be discarded or classified as another category.
[0107] In summary, the disclosed embodiment utilizes facial key point information to realize the judgment of the three-dimensional posture of the animal face, and filters out low-quality images based on the detected posture angle, the effective pixel ratio of the animal face occupied by the image (i.e., the distance), and the relationship between the external rectangular frame and the image edge, thereby screening out high-quality image data to be processed for ROI area extraction, which is beneficial to improving the accuracy of ROI area extraction.
[0108] Reference below Figure 8 , Figure 8 Schematic diagram of the structure of the electronic device provided by some embodiments of the present disclosure is shown. Figure 8 As shown, the electronic device at least includes a memory 801 and a processor 802. For example, the electronic device may also include Figure 8 More or fewer components (e.g., network interfaces, display devices, etc.) may be shown.
[0109] In particular, according to the embodiments provided by the present disclosure, the above reference flow chart Figure 1 , Figure 3 or Figure 4The described process can be implemented as a computer software program. For example, the embodiment provided by the present disclosure includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit (CPU), the above-mentioned functions defined in the system of the present disclosure are executed. The electronic device can be a terminal such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, other mobile Internet devices (Mobile Internet Devices, MID), a PAD, a desktop computer, etc. Figure 8 There is no limitation on the structure of the electronic device.
[0110] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and computer program products described in accordance with the various embodiments provided by this disclosure. Among them, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, and may also be provided in a processor.
[0113] As another aspect, embodiments of the present disclosure further provide a computer-readable storage medium. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not incorporated into the electronic device. The computer-readable storage medium stores one or more programs, which, when used by one or more processors, execute the method for extracting regions of interest from animal faces described in the present disclosure.
[0114] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention herein is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0115] Although a plurality of embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may conceive of many modifications, changes, and alternatives without departing from the ideas and spirit of the present disclosure. It should be understood that in practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The appended claims are intended to define the scope of protection of the present disclosure and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for extracting regions of interest from animal faces, characterized in that: The method includes: Perform facial key point detection on the image to be processed containing the animal face and generate a key point set; Calculating the rotation-corrected coordinates of each point in the key point set based on the positional relationship of a preset key point subset in the key point set under a standard frontal face pose; generating reference points for alignment of the region of interest based on the rotation-corrected coordinates to determine an affine transformation matrix; An affine transformation is performed on the region of interest in the image to be processed using the affine transformation matrix to output a spatially aligned region of interest.
2. The method according to claim 1, characterized in that The key point set includes at least: key points of the eye area, key points of the ear root area, key points of the nose area and key points of the chin area.
3. The method according to claim 2, characterized in that The method further includes: The reference point is corrected according to the side face posture of the animal's face.
4. The method according to claim 3, characterized in that When the reference points include a first reference point, a second reference point, and a third reference point, correcting the reference points according to the side face posture of the animal's face includes: Calculating a steering offset ratio according to lateral offsets between the first reference point, the second reference point, and the third reference point; When it is determined based on the steering offset ratio that the animal's face is in a side profile posture, the first reference point or the second reference point is corrected.
5. The method according to claim 4, characterized in that Calculating a steering offset ratio according to lateral offsets between the first reference point, the second reference point, and the third reference point includes: Calculating a first lateral offset between the abscissa of the first reference point and the abscissa of the second reference point; Calculating a second lateral offset between the abscissa of the first reference point and the abscissa of the third reference point; A steering offset ratio is calculated according to a ratio of the first lateral offset to the second lateral offset.
6. The method according to claim 4, characterized in that When it is determined based on the steering offset ratio that the face of the animal is in a profile posture, correcting the first reference point or the second reference point includes: When it is determined that the steering offset ratio is greater than a preset first threshold, reducing the horizontal coordinate of the first reference point based on the steering offset ratio to generate an adjusted first reference point; or When it is determined that the steering offset ratio is less than a preset second threshold, the horizontal coordinate of the second reference point is increased based on the steering offset ratio to generate an adjusted second reference point.
7. The method according to claim 1, characterized in that The preset key point subset includes two symmetrical key points on the animal's face and a key point in the chin area. Based on the positional relationship of the preset key point subset in the key point set under a standard frontal face posture, the rotation-corrected coordinates of each point in the key point set are calculated, including: Calculating an image rotation angle according to a positional relationship between a line connecting the two symmetrically positioned key points and a horizontal line in the image to be processed; A rotation transformation is performed on each point in the key point set based on the image rotation angle to generate rotation-corrected coordinates.
8. The method according to claim 1, characterized in that Generating reference points for alignment of the region of interest based on the rotation-corrected coordinates includes: Selecting a first keypoint subset and a second keypoint subset from the rotation-corrected coordinates, wherein the first keypoint subset includes keypoint coordinates corresponding to a first-side eye, a first-side base of an ear, and a nose, and the second keypoint subset includes keypoint coordinates corresponding to a second-side eye, a second-side base of an ear, and a nose, the second side being opposite to the first side; Determine the horizontal coordinate of the first reference point as the minimum value of the horizontal coordinates of all key points in the first key point subset; Determine the ordinate of the first reference point as the minimum value of the ordinates of all key points in the first key point subset; Determine the horizontal coordinate of the second reference point as the maximum horizontal coordinate of all key points in the second key point subset; Determine the ordinate of the second reference point as the minimum value of the ordinates of all key points in the second key point subset; The third reference point is determined to be the coordinate corresponding to the key point of the chin area.
9. The method according to claim 1, characterized in that After performing facial key point detection on the image to be processed containing the animal face and generating a key point set, the method further includes: performing image quality detection on the image to be processed based on the key point set; Based on the image quality detection results, the images to be processed that do not meet the image quality detection requirements are filtered out.
10. The method according to claim 9, characterized in that Image quality detection is performed on the image to be processed based on the key point set, including at least one of the following: abnormal facial posture detection; detection of the proportion of effective pixels occupied by the face in the image; and detection of the relationship between the face and the edge of the image.
11. The method according to claim 10, characterized in that The abnormal facial posture detection includes: Selecting a third key point subset from the key point set, the third key point subset comprising two symmetrical key points in the eye region, two symmetrical key points in the ear root region, a key point in the nose region, and a key point in the chin region; Whether the animal's face is in an abnormal posture is determined based on the positional relationship of the key points in the third key point subset, and the abnormal posture includes at least one of the following: the head's elevation angle causes facial abnormality; the head's depression angle causes facial abnormality; the head's yaw angle causes facial abnormality.
12. The method according to claim 11, characterized in that When the abnormal posture is a facial abnormality caused by the head elevation angle and / or the head yaw angle, determining whether the animal's face is in an abnormal posture based on the positional relationship of the key points in the third key point subset includes: Calculate a first distance value based on two symmetrical key points in the eye area; Calculate a second distance value based on the midpoint of the line connecting the two symmetrical key points of the eye region and the key point of the nose region; Determining whether the head elevation angle causes facial abnormality based on a relationship between a ratio of the first distance value to the second distance value and a preset elevation angle threshold; and / or, Whether the head yaw angle causes facial abnormality is determined according to a relationship between a ratio of the first distance value to the second distance value and a preset yaw angle threshold.
13. The method according to claim 12, characterized in that When the abnormal posture is caused by a head depression angle causing facial abnormality, determining whether the animal's face is in an abnormal posture based on the positional relationship of the key points in the third key point subset further includes: Calculate the third distance value based on the two symmetrical key points in the ear root area; Calculate a fourth distance value based on the key points of the nose area and the key points of the chin area; Whether the head depression angle causes facial abnormality is determined based on the ratio of the first distance value to the third distance value and the ratio of the first distance value to the fourth distance value.
14. The method according to claim 10, characterized in that The detection of the effective pixel ratio of the face to the image includes: Calculating the diameter length of a minimum enclosing circle according to the positional relationship of all key points in the key point set, wherein the minimum enclosing circle is the smallest circle that contains all key points in the key point set; If the ratio of the diameter length value to the diagonal length of the image to be processed is less than a preset first distance threshold, it means that the proportion of effective pixels occupied by the face in the image is too low; or If the ratio of the diameter length value to the diagonal length of the image to be processed is greater than a preset second distance threshold, it means that the face occupies too high a proportion of the effective pixels of the image.
15. The method according to claim 10, characterized in that The detection of the relationship between the face and the image edge includes: Calculating a minimum bounding rectangle based on all key points in the key point set, and obtaining coordinate values of four vertices thereof, wherein the minimum bounding rectangle is a rectangle with the minimum area that includes all key points in the key point set; If the shortest distance from any vertex to the four edges of the image is less than a preset distance threshold, it is determined that the face is too close to the edge of the image.
16. An electronic device, characterized in that: The device includes: processor; and, A memory storing computer instructions implemented by a computer for extracting an area of interest on an animal's face, wherein when the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 15.
17. A computer-readable storage medium, characterized in that The invention comprises computer-implemented program instructions for extracting an area of interest on an animal's face, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 15 is implemented.
Citation Information
Patent Citations
Method of facial landmark detection
CN103443804A
A face multi-area fusion expression recognition method based on depth learning
CN109344693A
Face rotation model generation method and device, computer device and storage medium
CN110826395A
Facial component searching device and facial component searching method
JP2003281539A
Formulation of compost decomposition using microbial agent and Fertilizer manufactured by using thereof and livestock excrement
KR102235297B1