Animal face region of interest extraction method, device and medium
By performing key point detection and affine transformation on pet facial images, rotation-corrected coordinates are generated, and the reference point is adaptively corrected. This solves the difficulty of ROI extraction caused by changes in pet facial pose and improves the accuracy of pet facial recognition.
Patent Information
- Application Number
- CN202511172894.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-21
AI Technical Summary
In pet facial recognition, the unpredictable changes in posture and expression can alter the position and shape of key feature points, making it difficult to extract complete and accurate regions of interest, thus affecting recognition performance.
By detecting facial key points in animal facial images, a set of key points is generated. Based on the positional relationship of a preset subset of key points in a standard frontal face pose, the coordinates after rotation correction are calculated, an affine transformation matrix is generated, and an affine transformation is performed on the image to output a spatially aligned region of interest. Adaptive correction reference points are used to handle pose changes.
It effectively solves the problem of difficulty in extracting regions of interest (ROIs) from pet faces, improves the accuracy of ROI extraction, and ensures the consistency and accuracy of ROI extraction under different poses.
Smart Images

Figure CN120673046B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of image processing. More specifically, the present disclosure relates to an animal face region of interest extraction method, device and medium. BACKGROUND
[0002] With the gradual relaxation of domestic pet-keeping policies, the pet industry in China is entering a golden period of rapid development. In many life scenarios such as pet vaccination, pet insurance handling, pet community management, and multi-pet feeding, it is necessary to accurately identify the identity of the pet and assign specific identity information to it. Pet identification technology has emerged as the times require, and its function is similar to that of human face recognition technology, aiming to accurately determine the identity information of the pet, thereby providing strong support for the management and service of the pet, and helping the pet industry to develop vigorously in a more standardized, orderly and refined direction.
[0003] Common pet identification solutions include implanting a chip in the pet's body, using the chip for positioning, using pet facial feature information for identification, using pet noseprint for identification, etc. For example, in the context of pet facial feature identification, the region of interest (ROI) refers to the key areas of the pet's face that have unique characteristics and recognition, such as the eyes, nose, mouth, ears, etc. By focusing on the extraction and analysis of these areas, the identification of the pet's identity can be achieved. However, pets, unlike humans, cannot pose according to instructions, and their poses and expressions change unpredictably. When the pet's face is tilted, turned sideways, or in a special expression state, the position and shape of the key feature points will change significantly, making it difficult for the system to extract complete and accurate regions of interest, thereby affecting the identification effect.
[0004] In view of this, there is an urgent need to provide an animal face region of interest extraction solution to improve the accuracy of pet face recognition. SUMMARY
[0005] To at least solve one or more of the technical problems mentioned above, the present disclosure proposes an animal face region of interest extraction solution in various aspects.
[0006] In a first aspect, the present disclosure provides an animal face region of interest extraction method, comprising: performing face key point detection on a to-be-processed image containing an animal face to generate a key point set; calculating the coordinates of each point in the key point set after rotation correction based on the positional relationship of a preset key point subset in the key point set in a standard front face pose; generating a reference point for aligning the region of interest based on the coordinates after rotation correction to determine an affine transformation matrix; and performing affine transformation on the region of interest in the to-be-processed image using the affine transformation matrix to output a spatially aligned region of interest.
[0007] In a second aspect, the present disclosure provides an electronic device, comprising: a processor; and a memory storing computer instructions for extracting a region of interest of an animal face implemented by a computer, which, when executed by the processor, causes the electronic device to perform the method described in the first aspect.
[0008] In a third aspect, the present disclosure provides a computer-readable storage medium containing program instructions for extracting a region of interest of an animal face implemented by a computer, which, when executed by the processor, causes the implementation of the method described in the first aspect.
[0009] Through the animal face region of interest extraction method, device and medium provided as above, the embodiments of the present disclosure generate a key point set by performing face key point detection on a to-be-processed image containing an animal face; calculate the coordinates of each point in the key point set after rotation correction based on the positional relationship of a preset key point subset in the key point set in a standard front face posture; generate a reference point for region of interest alignment based on the coordinates after rotation correction to determine an affine transformation matrix; and perform affine transformation on the region of interest in the to-be-processed image using the affine transformation matrix to output a spatially aligned region of interest, which effectively solves the problem of difficulty in extracting a region of interest of a pet face caused by independent rotation of pet ears and the like by effectively solving the affine transformation.
[0010] Further, in some embodiments, the problem of large differences in the cropped face region when the posture of the pet changes greatly is solved by correcting the reference point of the affine transformation, effectively improving the accuracy of ROI region extraction.
[0011] Further, in some embodiments, low-quality images are filtered out through face quality detection, providing high-quality image data for ROI region extraction. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example, and wherein like reference numerals refer to like elements throughout. In the drawings:
[0013] Figure 1 An exemplary flowchart of an animal face region of interest extraction method 100 provided by some embodiments of the present disclosure is shown;
[0014] FIGS. 2(a)~2(d) show an exemplary process of animal face region of interest extraction;
[0015] Figure 3An exemplary flow chart of the animal face region of interest extraction method 300 provided by some embodiments of the present disclosure is shown.
[0016] Figure 4 An exemplary flow chart of the animal face region of interest extraction method 400 provided by some embodiments of the present disclosure is shown.
[0017] Figure 5 A schematic diagram of the distribution characteristics of the animal face provided by some embodiments of the present disclosure is shown.
[0018] FIGS. 6(a)-6(c) show schematic diagrams of images identified as low quality provided by some embodiments of the present disclosure.
[0019] FIGS. 7(a)-7(b) show schematic diagrams of image quality evaluation indicators provided by some embodiments of the present disclosure.
[0020] Figure 8 A schematic diagram of the structure of an electronic device provided by some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of, rather than all of, the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0022] It should be understood that the terms “comprise” and “include” used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0023] It should also be understood that the terms used in the specification of the present disclosure are merely for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, the singular forms “a,” “an,” and “the” are intended to include the plural forms, unless the context clearly indicates otherwise. It should be further understood that the term “and / or” used in the specification and claims of the present disclosure means one or more of the associated listed items as well as all possible combinations of the items and includes the combinations.
[0024] As used in the specification and claims, the term “if’ can be interpreted as meaning “when,” or “upon,” or “in response to a determination,” or “in response to a detection” depending on the context. Similarly, the phrase “if it is determined” or “if [the recited condition or event] is detected” can be interpreted as meaning “upon determining” or “in response to a determining” or “upon detecting [the recited condition or event]” or “in response to a detection of [the recited condition or event],” depending on the context.
[0025] In the recognition using pet facial feature information, pets cannot pose as instructed like humans, and their postures and expressions change unpredictably. When the pet face is tilted, turned sideways, or in a special expression state, the position and shape of the key feature points will change greatly, making it difficult for the system to extract a complete and accurate ROI. In the preprocessing stage of facial feature extraction, accurately extracting the face region of interest (ROI) is a key step. Only by ensuring the consistency of face ROI extraction under various angles and scenes can the training difficulty of subsequent recognition accuracy be reduced, and the recognition accuracy be further improved. However, there are great differences between pet faces and human faces. Taking cats and humans as examples: human faces have diverse facial contours, which are usually flat, and the features are evenly distributed; while cat faces have shorter facial contours, with protruding mouth and nose, and the overall shape is triangular. Human ears are located on both sides of the head and have a fixed shape; cat ears are located on the top of the head and can rotate independently. These differences increase the difficulty of cat face ROI extraction.
[0026] To solve the problems in the above scenarios, the present disclosure provides an animal face region of interest extraction scheme, which effectively solves the problem of deviation of the nose position caused by the pet head turning through a key point position correction mechanism of the turning ratio, and improves the accuracy of ROI extraction.
[0027] The following will be described in detail Figures 1-8 The scheme of the present disclosure will be described in detail.
[0028] Figure 1 An exemplary flowchart of an animal face region of interest extraction method 100 provided by some embodiments of the present disclosure is shown. As shown in Figure 1 The method comprises:
[0029] Step S101, detecting facial key points of a to-be-processed image containing an animal face to generate a key point set.
[0030] Step S102, calculating the coordinates of each point in the key point set after rotation correction based on the positional relationship of a preset key point subset in the key point set under a standard front face posture.
[0031] Step S103, generating a reference point for region of interest alignment based on the coordinates after rotation correction to determine an affine transformation matrix.
[0032] In step S104, affine transformation is performed on the region of interest in the to-be-processed image by using the affine transformation matrix to output a spatially aligned region of interest.
[0033] Before the above steps, the method can further include obtaining the to-be-processed image containing the animal face. The to-be-processed image can be captured by a camera, and can also be obtained from a network, a standard image library, or other third-party image library, and the image obtaining manner is not limited by the disclosure. For example, the to-be-processed image containing the animal face can be obtained from an animal image library. The animal can refer to an animal whose face has multiple parts with prominent features. For example, felines, canines, primates, equines, birds, etc. The prominent parts of felines and canines are the nose.
[0034] In step S101, the to-be-processed image containing the animal face is subjected to face key point detection. The to-be-processed image containing the animal face can be input into the trained face detection network to output a key point set of the animal face. Taking a cat face as an example, the key point set can include but is not limited to key points of the eye region, key points of the nose region, key points of the chin region, key points of the ear region, etc. Optionally, as shown in FIG. 2(a), the key point set can include key points of the eye region {0 for the left eye and 1 for the right eye}, key points of the nose region {2}, ear triangular contour part key points {3 for the first ear root position point of the left ear outer contour, 4 for the top point of the left ear outer contour, 5 for the second ear root position point of the left ear outer contour, 6 for the second ear root position point of the right ear outer contour, 7 for the top point of the right ear outer contour, and 8 for the first ear root position point of the right ear outer contour}, and key points of the chin region {9}. These key points locate the anchor points at the outermost periphery of the face triangular region (such as the eyes, ear roots, and chin contour points), have significant biological rationality, and can adapt to the differences of different cat faces.
[0035] In step S102, the standard front face posture refers to a geometric relationship between the animal face key points satisfying that the line between the symmetric key points is in a horizontal position and the key point of the chin region is located directly below the midpoint of the line between the symmetric key points.
[0036] The preset key point subset in the key point set refers to part of the key points selected from the key point set. As shown in FIG. 2(a), the preset key point subset in the key point set includes the symmetric key points {5, 6} and the key point {9} of the chin region which are relatively fixed in position. In some embodiments, the preset key point subset in the key point set can also select the symmetric key points {0, 1} and the key point {9} of the chin region which are relatively fixed in position, or select the symmetric key points {3, 8} and the key point {9} of the chin region which are relatively fixed in position.
[0037] The rotation angle is obtained by establishing a rotation correspondence between the preset key point subset and the standard front face posture, and calculating the inclination angle of the line connecting the symmetric key points in the preset key point subset and the horizontal line.
[0038] In some embodiments, the preset key point subset includes two position-symmetric key points in the animal face and a key point in the chin region, and the rotated coordinates of each point in the key point set in the rotation-corrected image are calculated based on the positional relationship of the preset key point subset in the key point set in the standard front face posture, including: calculating the image rotation angle according to the positional relationship between the line connecting the two position-symmetric key points and the horizontal line in the image to be processed; performing the same rotation transformation on each point in the key point set based on the image rotation angle to generate the rotated coordinates.
[0039] In some embodiments, the rotation transformation can also be performed on the image to be processed based on the image rotation angle to generate a rotation-corrected image, so that the animal face in the rotation-corrected image meets the requirements of the standard front face posture.
[0040] Suppose that the preset key point subset includes the key points {5, 6} in the symmetric ear root region and the key point {9} in the chin position, the midpoint position of the line connecting the symmetric key points {5, 6} is the rotation center, the line connecting the symmetric key points {5, 6} is adjusted to be horizontal (i.e., the two points have the same vertical coordinates), and the key point {9} in the chin position is located below the line connecting the symmetric key points {5, 6}. According to the angle between the line connecting the symmetric key points {5, 6} and the horizontal line, and the coordinate position of the rotation center, a rotation matrix A is generated. The image to be processed is rotated using the rotation matrix A to obtain a rotation-corrected image, so that the key point {9} in the chin position of the animal in the rotation-corrected image is below. If the chin position is above during the rotation, rotating 180° based on the angle can make the key point {9} in the chin position be below.
[0041] The same rotation matrix A is used to transform the key point set to obtain the corrected key point coordinates. The key point set includes eye part key points {0 for the left eye and 1 for the right eye}, nose position key points {2}, ear triangular contour part key points {3 is the first ear root position point of the left ear outer contour, 4 is the top point of the left ear outer contour, 5 is the second ear root position point of the left ear outer contour, 6 is the second ear root position point of the right ear outer contour, 7 is the top point of the right ear outer contour, and 8 is the first ear root position point of the right ear outer contour}, and chin region key points {9}. The corrected key point coordinates {landmarkRot[0], landmarkRot[1], …, landmarkRot[9]} can be generated by the rotation matrix A. The corrected key point coordinates refer to the accurate position coordinates of the key points (such as the eye, nose, chin, and other feature points of the pet face) in the original image after the same rotation transformation as the image.
[0042] In step S103, the reference points for the alignment of the region of interest are generated based on the rotation-corrected coordinates, which can include selecting a first key point subset and a second key point subset from the corrected coordinates, wherein the first key point subset includes key point coordinates corresponding to the first side eye, the first side ear root, and the nose, and the second key point subset includes key point coordinates corresponding to the second side eye, the second side ear root, and the nose, the second side being opposite to the first side; determining the horizontal coordinate of the first reference point as the minimum value of the horizontal coordinates of all key points in the first key point subset; determining the vertical coordinate of the first reference point as the minimum value of the vertical coordinates of all key points in the first key point subset; determining the horizontal coordinate of the second reference point as the maximum value of the horizontal coordinates of all key points in the second key point subset; determining the vertical coordinate of the second reference point as the minimum value of the vertical coordinates of all key points in the second key point subset; and determining the third reference point as the coordinate corresponding to the key point of the chin region.
[0043] As shown in FIG. 2(b), in the rotation-corrected image, the reference points for the alignment of the region of interest can include a first reference point R11, a second reference point R22, and a third reference point R33, which is the key point of the chin region.
[0044] The horizontal coordinate x of the first reference point R11 is landmarkRot[0], landmarkRot[2], landmarkRot[3], and landmarkRot[5] (not shown in the figure), and the minimum value (min_x1) of the horizontal coordinates x of these four points is
[0045] The longitudinal coordinate y of the first reference point R11 is landmarkRot[0], landmarkRot[2], landmarkRot[3], and landmarkRot[5] (not shown in the figure), and the minimum value (min_y1) of the longitudinal coordinates y of the four points.
[0046] The minimum value (min_x1) of x and the minimum value (min_y1) of y are combined into the first reference point R11 (min_x1, min_y1).
[0047] The longitudinal coordinate y of the second reference point R22 is landmarkRot[1], landmarkRot[2], landmarkRot[6], and landmarkRot[8] (not shown in the figure), and the minimum value (min_y2) of the longitudinal coordinates y of the four points.
[0048] The longitudinal coordinate y of the second reference point R22 is landmarkRot[1], landmarkRot[2], landmarkRot[6], and landmarkRot[8] (not shown in the figure), and the minimum value (min_y2) of the longitudinal coordinates y of the four points.
[0049] The third reference point R33 is landmarkRot[9], which is the coordinate of the chin position in the rotation-corrected image.
[0050] In step 104, an affine transformation is performed on the region of interest in the image to be processed using the affine transformation matrix to output a spatially aligned region of interest. An image alignment (affine transformation) process based on three sets of corresponding points can be used to align specific key points in the rotation-corrected image to target positions to generate a standardized aligned image (for example, the size can be The value of w is related to the input parameters of the model, for example, the model input image requires w to be 112. In some embodiments, the standardized aligned image can also be size, where the values of w1 and w2 are not equal.
[0051] Here, it is assumed that the aligned image is uniform size, and w can be 112.
[0052] The three reference points in the rotation-corrected image include {point11, point22, and point33}.
[0053] The three target reference points in the aligned target image include {point44, point55, and point66}.
[0054] A mapping relationship is established between the three reference points determined in the rotation-corrected image and the three target reference points in the aligned target image. Point 11 is defined as the upper left anchor point (point 44, 0, 0) of the aligned target image, point 22 is defined as the upper right anchor point (point 55, 0, w) of the aligned target image, and point 33, the key point of the chin region, is defined as the lower right anchor point of the aligned target image. ,w). scaleWidth_new represents the recalculated steering offset ratio.
[0055] The affine transformation matrix B is calculated based on the coordinates of {point11, point22, point33} in the original image and {point44, point55, point66} in the aligned target image. The affine transformation matrix is applied to output the spatially aligned region of interest. The left boundary is fixed using the mapping relationship between points 11 and 44 to ensure the left side of the head aligns with the left edge of the image. The lower boundary is fixed using the mapping relationship between points 22 and 55 to ensure the bottom of the chin aligns with the lower edge of the image. The right boundary is dynamically set using the mapping relationship between points 33 and 66 to adapt the width according to the rotation offset ratio. The lower right point dynamically simulates the region during head rotation, achieving a standardized output to eliminate rotation and scale differences and maintain the alignment of key facial structures.
[0056] The above three-point alignment method can effectively handle the deformation caused by head turning, while maintaining the positional consistency of key facial structures, providing standardized input for subsequent feature extraction and recognition.
[0057] Figure 3 An exemplary flowchart of an animal facial region of interest extraction method 300 provided in some embodiments of this disclosure is shown. It will be understood that method 300 is a... Figure 1 Further limitations and / or extensions of Method 100. Therefore, the foregoing is combined with Figure 1 The relevant detailed description also applies below. This method further refines the reference points used for region-of-interest alignment based on the animal's facial profile.
[0058] like Figure 3 As shown, at step S304, the calculation of the steering offset ratio may be included based on the lateral offset between the first reference point, the second reference point and the third reference point; then, when it is determined that the animal's face is in a side profile posture based on the steering offset ratio, the first reference point or the second reference point may be corrected.
[0059] According to the lateral offset between the first reference point, the second reference point and the third reference point, the calculation of the steering offset ratio can include calculating a first lateral offset between the abscissa of the first reference point and the abscissa of the second reference point; calculating a second lateral offset between the abscissa of the first reference point and the abscissa of the third reference point; and calculating the steering offset ratio according to the ratio of the first lateral offset to the second lateral offset.
[0060] Continuing with the example of the reference points shown in Figure 2(b), the first lateral offset X12 between the first reference point and the second reference point is calculated; the second lateral offset X13 between the first reference point and the third reference point is calculated; and the steering offset ratio scaleWidth is calculated:
[0061] .
[0062] In some embodiments, the side face posture can be determined according to the distribution characteristics of the rotation-corrected coordinates.
[0063] In other embodiments, the elevation angle score parameter for quantitatively evaluating the head-up state, the depression angle score parameter for quantitatively evaluating the head-down state, the auxiliary score parameter for assisting in evaluating the head-down state, and the yaw angle score parameter for quantitatively evaluating the left and right turning angles can be defined according to the relationship between the head posture change and the plurality of distance measurement values; and the current posture of the animal face can be determined by jointly judging the elevation angle score parameter, the depression angle score parameter, the auxiliary score parameter and the yaw angle score parameter through a pre-set threshold. After the posture of the animal face is determined, the degree of turning of the animal face is further determined according to the steering offset ratio, so as to trigger the correction and compensation of the rotation-corrected coordinates.
[0064] The steering offset ratio refers to a normalized index quantifying the degree of turning of the pet head in the horizontal direction (left and right turning) or the vertical direction (head-up / head-down). It is calculated by comparing the asymmetry of the facial key points, and converts the degree of head turning into a ratio value between 0% and 100%.
[0065] In some embodiments, when the key point set includes key points of an eye region, key points of an ear root region, key points of a nose region, and key points of a chin region, and the animal face is determined to be in a profile pose based on the turning offset ratio, the modification of the first reference point or the second reference point can include determining the first reference point, the second reference point, and the third reference point for the region of interest alignment according to the rotated corrected coordinates; calculating the turning offset ratio according to the lateral offset between the first reference point, the second reference point, and the third reference point; when the turning offset ratio is determined to be greater than a preset first threshold, reducing the horizontal coordinate of the first reference point based on the turning offset ratio to generate an adjusted first reference point. At this time, the reference points for the region of interest alignment are updated to the adjusted first reference point, the second reference point, and the third reference point.
[0066] In some embodiments, when the turning offset ratio is determined to be less than a preset second threshold, the horizontal coordinate of the second reference point is increased based on the turning offset ratio to generate an adjusted second reference point. At this time, the reference points for the region of interest alignment are updated to the first reference point, the adjusted second reference point, and the third reference point.
[0067] Then, after generating the adjusted first reference point or the adjusted second reference point, the turning offset ratio is recalculated using the horizontal coordinate of the adjusted first reference point, or the turning offset ratio is recalculated using the horizontal coordinate of the adjusted second reference point until convergence.
[0068] When the turning offset ratio indicates that the animal face is in a profile state, the adaptive adjustment of the spatial position of the corrected key point coordinates includes, when the turning offset ratio is determined to be greater than a preset first threshold, performing an outward expansion process on the horizontal coordinate of the first reference point (i.e., adjusting the horizontal coordinate of the first reference point to a direction with a smaller pixel coordinate value), to generate an adjusted first reference point; or, when the turning offset ratio is determined to be less than a preset second threshold, performing an outward expansion process on the second reference point (i.e., adjusting the horizontal coordinate of the first reference point to a direction with a larger pixel coordinate value), to generate an adjusted second reference point.
[0069] The first reference point is updated using the adjusted first reference point, or the second reference point is updated using the adjusted second reference point, and the third reference point remains unchanged, and the step of calculating the turning offset ratio based on the positional relationship between the first reference point and the second reference point is returned to, to recalculate a new turning offset ratio until the new turning offset ratio falls within a preset interval or reaches a maximum number of iterations (i.e., convergence).
[0070] In extreme profile cases (e.g. scaleWidth value is very small or very large), the cat's head is in a large angle left or right profile pose, in which the nose region will protrude significantly from the original ear root connecting line (the connecting line of R11 and R22). In order to more accurately describe the head contour, it is necessary to adaptively adjust the ear root point position according to the turning direction.
[0071] For right profile processing (scaleWidth value is large), when the right profile pose is detected (scaleWidth >> preset threshold), the x coordinate value of the left ear root point R11 is proportionally reduced to generate a new coordinate R11_new (the coordinate value is closer to the left side). This adjustment moves the left ear root inward, better adapting to the contour changes caused by the protruding nose. As shown in FIG. 2(c), the horizontal coordinate of the R11_new point shown in FIG. 2(c) is obtained by proportionally reducing the horizontal coordinate of the R11 point shown in FIG. 2(b). Assuming that the horizontal coordinate of the R11 point is x, according to the formula , the outward expansion of R11 is realized to ensure that the nose region can be cut into the target image.
[0072] Alternatively, for left profile processing (scaleWidth value is small), when the left profile pose is detected (scaleWidth << preset threshold), the x coordinate value of the right ear root point R22 is proportionally increased to generate a new coordinate R22_new (not shown in the figure). This adjustment moves the right ear root outward, better fitting the head contour characteristics when the nose protrudes.
[0073] This adaptive adjustment mechanism dynamically corrects the ear root reference point position, so that the geometric model composed of key points can more accurately reflect the anatomical features of the protruding nose in extreme turning poses, providing more accurate reference for subsequent pose analysis and contour modeling.
[0074] In some embodiments, in the right profile processing (scaleWidth value is large), a third horizontal offset between the horizontal coordinate of the corrected new coordinate R11_new (corrected first reference point) and the horizontal coordinate of the original second reference point point22Rot is calculated .
[0075] A fourth horizontal offset between the horizontal coordinate of the corrected new coordinate point11Rot_new (corrected first reference point) and the horizontal coordinate of the original third reference point point33Rot is calculated .
[0076] A new turning offset scaleWidth_new is calculated again:
[0077] .
[0078] When the new steering offset ratio falls into the preset interval or reaches the maximum number of iterations, the subsequent method flow is entered.
[0079] For left face processing (when scaleWidth value is small), calculate the fifth lateral offset between the horizontal coordinate of the original first reference point R11 and the horizontal coordinate of the new coordinate R22_new (the corrected second reference point) . Calculate the second lateral offset X13 between the horizontal coordinate of the original first reference point R11 and the horizontal coordinate of the original third reference point R33. Then calculate the new steering offset ratio scaleWidth_new:
[0080] .
[0081] When the new steering offset ratio falls into the preset interval or reaches the maximum number of iterations, the subsequent method flow is entered.
[0082] After completing the correction of the reference points, a corrected affine transformation matrix is determined based on the corrected reference points. Then, at step S305, an affine transformation is performed on the region of interest in the image to be processed using the corrected affine transformation matrix to output a spatially aligned region of interest.
[0083] For right face processing (when scaleWidth value is large), the corrected first reference point R11_new is mapped back to the position of the rotation-corrected image, including: passing R11_new through the inverse rotation matrix A -1 to obtain the corresponding point 11 point on the rotation-corrected image. The second reference point R22 is mapped back to the position of the rotation-corrected image, including: passing R22 through the inverse rotation matrix A -1 to obtain the corresponding point 22 point on the rotation-corrected image. The chin area key point of the cat face does not need to be transformed, and the corresponding point 33 point landmark[9] on the rotation-corrected image.
[0084] For left face processing (when scaleWidth value is small), the first reference point R11 is mapped back to the position of the rotation-corrected image, including: passing R11 through the inverse rotation matrix A -1 to obtain the corresponding point 11 point on the rotation-corrected image. The corrected second reference point R22_new is mapped back to the position of the rotation-corrected image, including: passing R22_new through the inverse rotation matrix A -1 to obtain the corresponding point 22 point on the rotation-corrected image. The chin area key point of the cat face does not need to be transformed, and the corresponding point 33 point landmark[9] on the rotation-corrected image.
[0085] After the above processing, the final output spatially aligned region of interest is shown in FIG. 2(d). When the animal face is a cat face, the ROI region extraction method provided by the embodiments of the present disclosure overcomes the problem that the large change in the ear greatly affects the ROI region extraction when the cat face is at a large angle side face. In the ROI region extraction process, the embodiments of the present disclosure reduce the information extraction of the ear region, and the point at the upper left corner after alignment mainly considers the positions of the cat face key points {0, 2, 3, 5}, the point at the upper right corner after alignment mainly considers the positions of the cat face key points {0, 2, 6, 8}, and the point at the lower right corner after alignment is dynamically aligned through the key points of the chin region, so as to eliminate the rotation and scale difference and keep the key structures of the face aligned.
[0086] Further, considering the relationship of different side faces, the nose position is greatly affected, and an adaptive compensation correction method is used to adjust the reference point position of the ear part calibrated and cropped in the rotation correction image, so that the nose part can be kept in the output aligned image, thereby obtaining the most beneficial alignment effect under small deformation.
[0087] Figure 4 An exemplary flowchart of an animal face region of interest extraction method 400 provided by another embodiment of the present disclosure is shown. It can be understood that the method 400 is a further limitation and / or expansion of the method 100 or the method 300 in the Figure 1 or Figure 3 Therefore, the foregoing related detailed description in conjunction with Figure 1 and Figure 3 is also applicable hereinafter. The method further performs image quality detection on the to-be-processed image based on the key point set; and filters the to-be-processed image that does not meet the image quality detection requirement based on the image quality detection result. As shown in Figure 4 , the method comprises:
[0088] At step S402, the image quality detection is performed on the to-be-processed image based on the key point set. If the image quality detection result indicates that there is no image quality problem, the method flow continues to the subsequent related steps. If the image quality detection result indicates that there is an image quality problem, the to-be-processed image that does not meet the image quality detection requirement is filtered.
[0089] In the above steps, the image quality detection performed on the to-be-processed image based on the key point set can include at least one of the following: face abnormal posture detection; face effective pixel proportion detection occupying the image; face and image edge relationship detection.
[0090] In some embodiments, the face abnormal posture detection can include: selecting a third key point subset from the key point set, the third key point subset including two symmetrical key points of an eye region, two symmetrical key points of an ear root region, a key point of a nose region, and a key point of a chin region; determining whether the animal face is in an abnormal posture according to the positional relationship of the key points in the third key point subset, the abnormal posture including at least one of the following: a head elevation angle causing the face to be abnormal; a head depression angle causing the face to be abnormal; and a head yaw angle causing the face to be abnormal.
[0091] When the abnormal posture is the head elevation angle causing the face to be abnormal and / or the head yaw angle causing the face to be abnormal, determining whether the animal face is in an abnormal posture according to the positional relationship of the key points in the third key point subset can include: calculating a first distance value according to the two symmetrical key points of the eye region; calculating a second distance value according to the midpoint position of the line connecting the two symmetrical key points of the eye region and the key point of the nose region; determining whether the head elevation angle causes the face to be abnormal according to the relationship between the ratio of the first distance value and the second distance value and a preset elevation angle threshold; and / or determining whether the head yaw angle causes the face to be abnormal according to the relationship between the ratio of the first distance value and the second distance value and a preset yaw angle threshold.
[0092] When the abnormal posture is the head depression angle causing the face to be abnormal, determining whether the animal face is in an abnormal posture according to the positional relationship of the key points in the third key point subset can include: calculating a first distance value according to the two symmetrical key points of the eye region; calculating a third distance value according to the two symmetrical key points of the ear root region; calculating a fourth distance value according to the key point of the nose region and the key point of the chin region; determining whether the head elevation angle is abnormal according to the relationship between the ratio of the first distance value and the second distance value and a preset elevation angle threshold; determining whether the head yaw angle is abnormal according to the relationship between the ratio of the first distance value and the second distance value and a preset yaw angle threshold; and determining whether the head depression angle is abnormal according to the ratio of the first distance value and the third distance value and the ratio of the first distance value and the fourth distance value.
[0093] The eye distance DistEyes between the left eye position key point 0 and the right eye position key point 1 is calculated as shown in FIG. 4, where the line segment D1 represents the first distance value. Figure 5
[0094] The midpoint coordinates of the line connecting the two eyes are obtained, and the nose offset distance DistMidEyesNose between the midpoint and the nose position key point 2 is calculated as shown in FIG. 5, where the line segment D2 represents the second distance value. Figure 5
[0095] An elevation score parameter ElevaScore is calculated according to the second distance value and the first distance value. For example, the elevation score parameter can be calculated by a ratio of the second distance value to the first distance value, or according to a ratio of the first distance value to the second distance value. Assuming that the elevation score parameter is calculated by a ratio of the second distance value to the first distance value, when the elevation score parameter is greater than a preset elevation threshold, it indicates that the elevation is too large. An image with too large elevation is shown in Fig. 6(a).
[0096] In some embodiments, a yaw score parameter YawScore can also be calculated according to the second distance value and the first distance value. For example, the yaw score parameter can be calculated by a ratio of the second distance value to the first distance value, or according to a ratio of the first distance value to the second distance value. Assuming that the yaw score parameter is calculated by a ratio of the second distance value to the first distance value, when the yaw score parameter is less than a preset yaw threshold, it indicates that the yaw is too large. An image with too large yaw is shown in Fig. 6(b).
[0097] When the animal is in the head-up state or the head-turning posture, the distance between the two eyes is basically unchanged. However, when in the head-up state, the distance between the midpoint of the line connecting the nose to the two eyes is reduced. When in the head-turning posture, the nose is laterally offset, which can cause the distance between the midpoint of the line connecting the nose to the two eyes to increase. Therefore, the present disclosure utilizes the different sensitivity of the eye-nose distance to the elevation and yaw angles, combined with the stability of the eye distance, to achieve the detection of the two states by the same ratio.
[0098] In some embodiments, a first midpoint coordinate value of the two-ear-inner key points is calculated according to the key points of the two-ear-inner. A second midpoint coordinate value of the two-eye key points is calculated according to the key points of the two eyes, and a third distance value DistMidEyesMidEars between the two midpoints is calculated according to the first midpoint coordinate value and the second midpoint coordinate value. As shown in Fig. 5, line segment D3 represents the third distance value. Figure 5
[0099] A fourth distance value DistNoseChin is calculated according to the key point of the nose position and the key point of the chin position. As shown in Fig. 6, line segment D4 represents the fourth distance value. Figure 5
[0100] The depression score parameter DepressScore is calculated according to the first distance value and the third distance value. For example, the depression score parameter can be calculated by the ratio of the third distance value to the first distance value, or according to the ratio of the first distance value to the third distance value. The assist score parameter AssistScore is calculated according to the first distance value and the fourth distance value. For example, the assist score parameter can be calculated by the ratio of the fourth distance value to the first distance value, or according to the ratio of the first distance value to the fourth distance value.
[0101] Assuming that the depression score parameter is calculated by the ratio of the first distance value to the third distance value, and the assist score parameter is calculated by the ratio of the first distance value to the third distance value, when the depression score parameter is less than a preset depression threshold value and the assist score parameter is greater than a preset assist score threshold value, it indicates that the depression angle is too large. The image with the depression angle being too large is shown in FIG. 6(c).
[0102] In the case that the animal's face is in a lowered head state, the third distance value between the midpoint of the eye and the midpoint of the ear increases when the head is lowered, and the fourth distance value decreases when the head is lowered, by using the stability of the first distance value in the process of lowering the head, to jointly detect the problem of the animal's head angle being too large, effectively improving the accuracy of the detection result.
[0103] The present disclosure proposes to determine the problems of the image's elevation angle being too large (head up), the image's depression angle being too low (head down), or the image's yaw angle being too large (large angle side face) by a multi-feature cooperative manner. These problems indicate that the image has a serious posture deviation, and the image needs to be discarded. The image set detected by the image detection can provide beneficial high-quality data for ROI region extraction, thereby improving the accuracy of the ROI region extraction.
[0104] In some embodiments, the face occupies the effective pixel ratio of the image detection can include: calculating a diameter length value of a minimum enclosing circle according to the positional relationship of all key points in the key point set, the minimum enclosing circle being the smallest circle containing all key points in the key point set; if the ratio of the diameter length value to the diagonal length of the image to be processed is less than a preset first distance threshold value, it indicates that the face occupies the effective pixel ratio of the image is too low; or, if the ratio of the diameter length value to the diagonal length of the image to be processed is greater than a preset second distance threshold value, it indicates that the face occupies the effective pixel ratio of the image is too high.
[0105] As shown in FIG. 7(a), the key point set includes: eye part key points {0 for left eye, 1 for right eye}, nose part key point {2}, ear triangular contour part key points {3 for the first ear root position point of the left ear outer contour, 4 for the top point of the left ear outer contour, 5 for the second ear root position point of the left ear outer contour, 6 for the second ear root position point of the right ear outer contour, 7 for the top point of the right ear outer contour, and 8 for the first ear root position point of the right ear outer contour}, and chin area key point {9}. The minimum enclosing circle refers to the smallest circle that can contain all the key points in the given key point set. For example, the minimum enclosing circle can be calculated using OpenCV to obtain the diameter of the minimum enclosing circle. Whether the effective pixel ratio of the face in the image is normal is determined by judging the relationship between the diameter and the threshold. If it is lower than the threshold, it means that the face occupies too low a proportion, and specific details cannot be detected. If it is higher than the threshold, it means that the face occupies too high a proportion, and details may be lost. The actual threshold can be adjusted according to the specific application and the camera focal length.
[0106] The embodiments of the present disclosure quantize the proportion of the face in the image by the diameter of the minimum enclosing circle, effectively solving the quality control problems of too small face (insufficient recognition features) and too large face (edge feature loss).
[0107] In some embodiments, the face-to-image edge relationship detection can include calculating a minimum circumscribed rectangle according to all the key points of the key point set, and obtaining the coordinate values of the four vertices of the minimum circumscribed rectangle, the minimum circumscribed rectangle being the smallest area rectangle containing all the key points in the key point set; if the shortest distance of any vertex to the four edges of the image is less than a preset distance threshold, it is determined that the face is too close to the image edge. As shown in FIG. 7(b), the minimum circumscribed rectangle includes all the key points in the key point set. For example, the minimum circumscribed rectangle can be calculated using OpenCV, and the coordinate values of the four vertices are obtained at the same time. Then, according to the coordinate values of the four vertices, the shortest distance to the four edges of the image is calculated to determine whether the face region is too close to the image edge, which may cause important features to be truncated.
[0108] When the object rotates, the direction change of the minimum circumscribed rectangle directly reflects the object orientation, while the minimum enclosing circle cannot provide direction information. The distance calculation from the rectangle vertex to the image edge can more accurately determine the relationship between the rectangle vertex and the image edge. When the cat's head is tilted, the minimum circumscribed rectangle rotates with the head, which can accurately reflect the actual space occupation.
[0109] The embodiments of the present disclosure use the minimum circumscribed rectangle for fast preliminary screening (small calculation amount), and then perform accurate detection of the minimum circumscribed rectangle on the images that pass the preliminary screening, to construct a hierarchical quality evaluation system, taking into account efficiency and accuracy.
[0110] If the face is too small, the face-to-image edge relationship detection can be performed on the face region in the image. Figure 1On the basis of the above, after performing image quality detection, a rotation correction process is calculated for the to-be-processed image that satisfies the image quality detection, and then a reference point for alignment of the region of interest is generated to determine an affine transformation matrix, and the affine transformation matrix is then used to perform affine transformation on the region of interest in the to-be-processed image to output the spatially aligned region of interest.
[0111] If in step S404 Figure 3 On the basis of the above, after performing image quality detection, a rotation correction process is calculated for the to-be-processed image that satisfies the image quality detection, and then a reference point for alignment of the region of interest is generated to determine an affine transformation matrix, and the affine transformation matrix is then used to perform affine transformation on the region of interest in the to-be-processed image to output the spatially aligned region of interest.
[0112] In some embodiments, the method can further include defining, according to the relationship between the head posture change and the plurality of distance measurement values, an elevation score parameter for quantitatively evaluating the head-up state, a depression score parameter for quantitatively evaluating the head-down state, an auxiliary score parameter for assisting in evaluating the head-down state, and a yaw score parameter for quantitatively evaluating the left and right turning angles; and jointly judging the elevation score parameter, the depression score parameter, the auxiliary score parameter, and the yaw score parameter by a preset threshold to determine the posture in which the animal face is currently located.
[0113] First, complete threshold values of all parameters are defined. The parameters include an elevation score parameter ElevaScore, a depression score parameter DepressScore, an auxiliary score parameter AssistScore, a new yaw angle score parameter ScaleYawScore, etc. For each parameter, a threshold range is set according to specific detection requirements. For example, the elevation score parameter sets an elevation threshold range; the depression score parameter sets a depression threshold range; the auxiliary score parameter sets an auxiliary score threshold range; and the new yaw angle score parameter sets a yaw threshold range.
[0114] The new yaw angle score parameter is calculated based on part of the key point set: ear triangular contour part key points {3 is the first ear root position point of the left ear outer contour, 5 is the second ear root position point of the left ear outer contour, 6 is the second ear root position point of the right ear outer contour, and 8 is the first ear root position point of the right ear outer contour} and chin region key point {9}. By selecting the minimum horizontal coordinate in the horizontal coordinates of the key points {3, 5} and the minimum vertical coordinate in the vertical coordinates, the minimum horizontal coordinate and the minimum vertical coordinate form a reference point 1. By selecting the maximum horizontal coordinate in the horizontal coordinates of the key points {6, 8} and the minimum vertical coordinate in the vertical coordinates, the maximum horizontal coordinate and the minimum vertical coordinate form a reference point 2. The chin region key point {9} is taken as a reference point 3. The horizontal offset P12 between the reference point 1 and the reference point 2 is calculated, the horizontal offset P13 between the reference point 1 and the reference point 3 is calculated, and the new yaw angle score parameter ScaleYawScore is calculated according to the ratio of the offset P12 and the offset P13.
[0115] Assuming that a single threshold judgment mode is adopted, the five posture classification rules are defined in combination of the elevation angle score parameter ElevaScore, the depression angle score parameter DepressScore, the auxiliary score parameter AssistScore and the new yaw angle score parameter ScaleYawScore.
[0116] For example, when the elevation angle score parameter ElevaScore, the depression angle score parameter DepressScore, the auxiliary score parameter AssistScore and the new yaw angle score parameter ScaleYawScore all do not exceed the specified threshold range, this posture is defined as a frontal face.
[0117] When the depression angle score parameter DepressScore exceeds the specified threshold range, and the auxiliary score parameter AssistScore and the new yaw angle score parameter ScaleYawScore all do not exceed the specified threshold range, this posture is defined as a low head face.
[0118] When the elevation angle score parameter ElevaScore exceeds the specified threshold range, and the depression angle score parameter DepressScore and the new yaw angle score parameter ScaleYawScore all do not exceed the specified threshold range, this posture is defined as a raised head face.
[0119] When the elevation angle score parameter ElevaScore and the depression angle score parameter DepressScore both do not exceed the specified threshold range, and the new yaw angle score parameter ScaleYawScore is greater than a preset left face deflection threshold, this posture is defined as a left turning face.
[0120] When the elevation score parameter ElevaScore and the depression score parameter DepressScore are both not exceeding their prescribed threshold range, and the new yaw score parameter ScaleYawScore is less than the preset right face turning threshold, the face pose is defined as right face turning.
[0121] Other class poses include all poses that do not meet any of the above conditions, or each parameter exceeds the maximum degree threshold.
[0122] The disclosure embodiments can test the logic with a large amount of high-quality cat face data, adjust the threshold range to reduce the error rate of classification, and reasonably cover low-quality pictures (such as extreme poses) in the other class poses. For example, when the elevation score parameter ElevaScore does not exceed its prescribed threshold range, the depression score parameter DepressScore does not exceed its prescribed threshold range, and the new yaw score parameter ScaleYawScore is greater than the preset left face turning threshold, it is defined as left face turning. Alternatively, the pose range can be expanded by setting a floating range of the prescribed threshold range, for example, the current elevation score parameter ElevaScore is within the floating range of the prescribed threshold range, which is considered as slight head lifting, the depression score parameter DepressScore does not exceed its prescribed threshold range, and the new yaw score parameter ScaleYawScore is greater than the preset left face turning threshold, at this time the face pose is in the state of head lifting and left turning. Since the degree of head lifting is within the floating range, the face pose can still be attributed to the state of left face turning. However, when the elevation score parameter ElevaScore exceeds the maximum threshold of the elevation angle, only one indicator is needed to determine that the face pose at this time is in the state of excessive head lifting (i.e., the elevation angle is too large), which is attributed to the state of excessive elevation angle, and the image can be discarded or attributed to the other class.
[0123] In summary, the disclosure embodiments utilize the face key point information to realize the judgment of the three-dimensional pose of the animal face, and filter out low-quality images according to the detected pose angle, the effective pixel ratio of the animal face in the image (i.e., the distance), and the relationship between the circumscribed rectangle frame and the image edge, so as to extract high-quality image data for the ROI region, which is conducive to improving the accuracy of ROI region extraction.
[0124] The following refers to Figure 8 , Figure 8 The structure schematic diagram of an electronic device provided by some embodiments of the disclosure is shown. As Figure 8 shown, the electronic device at least includes a memory 801 and a processor 802. For example, the electronic device can further include more or fewer components (such as a network interface, a display device, etc.) than Figure 8 shown.
[0125] In particular, embodiments provided in accordance with the present disclosure provide for the above-described processes to be performed in accordance with the flowcharts Figure 1 , Figure 3 or Figure 4 The processes described above can be implemented as computer software programs. For example, embodiments provided in accordance with the present disclosure include a computer program product comprising a computer program carried on a machine-readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication section, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-described functions defined in the system of the present disclosure are performed. The electronic device can be a terminal such as a smartphone (e.g., an Android phone, an iOS phone, etc.), a tablet computer, a palmtop computer, a Mobile Internet Device (MID), a PAD, a desktop computer, etc. Figure 8 The structure of the electronic device is not limited.
[0126] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium that can send, propagate or transmit a program for use by or in connection with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the methods and computer program products described in accordance with the various embodiments provided in this disclosure. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0128] The units or modules described in the embodiments provided in this disclosure can be implemented in software or hardware. The described units or modules can also be located in a processor.
[0129] In another aspect, this disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the animal facial region of interest extraction method described in this disclosure.
[0130] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0131] While several embodiments of the disclosure have been shown and described herein, it is to be understood that the embodiments are merely exemplary. Numerous changes, substitutions and equivalents can occur to those skilled in the art without departing from the spirit and scope of the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein can be employed in practicing the disclosure. It is intended that the following claims define the scope of the disclosure and that methods equivalent to those claims recited herein are within the scope and spirit of the disclosure. Therefore, the disclosure is not limited to the specific embodiments described herein, but only by the claims.
Claims
1. A method for extracting regions of interest (ROIs) on an animal's face, characterized in that, The method includes: Facial keypoint detection is performed on the image to be processed, which contains animal faces, to generate a keypoint set; Based on the positional relationship of a preset subset of key points in the key point set under the standard frontal face pose, calculate the rotationally corrected coordinates of each point in the key point set; Based on the rotation-corrected coordinates, reference points are generated for aligning the region of interest to determine the affine transformation matrix; The affine transformation matrix is used to perform an affine transformation on the region of interest in the image to be processed to output a spatially aligned region of interest; The process of generating reference points for region-of-interest alignment based on the rotated and corrected coordinates includes: A first keypoint subset and a second keypoint subset are selected from the rotationally corrected coordinates. The first keypoint subset includes keypoint coordinates corresponding to the first eye, the first ear root, and the nose. The second keypoint subset includes keypoint coordinates corresponding to the second eye, the second ear root, and the nose. The second side is opposite to the first side. The x-coordinate of the first reference point is determined to be the minimum x-coordinate of all key points in the first key point subset; The ordinate of the first reference point is determined to be the minimum value of the ordinates of all key points in the first key point subset; The x-coordinate of the second reference point is determined to be the maximum value of the x-coordinates of all key points in the second key point subset; The ordinate of the second reference point is determined to be the minimum value of the ordinates of all key points in the second key point subset; The third reference point is determined as the coordinates corresponding to the key points in the chin area.
2. The method according to claim 1, characterized in that, The set of key points includes at least the following: key points in the eye area, key points in the ear root area, key points in the nose area, and key points in the chin area.
3. The method according to claim 2, characterized in that, The method also includes: The reference point is adjusted based on the side profile of the animal's face.
4. The method according to claim 3, characterized in that, When the reference points include a first reference point, a second reference point, and a third reference point, the reference points are corrected according to the side profile of the animal's face, including: Calculate the steering offset ratio based on the lateral offset between the first reference point, the second reference point, and the third reference point; When determining that the animal's face is in a side profile posture based on the steering offset ratio, the first reference point or the second reference point is corrected.
5. The method according to claim 4, characterized in that, The steering offset ratio is calculated based on the lateral offset between the first reference point, the second reference point, and the third reference point, including: Calculate the first lateral offset between the x-coordinate of the first reference point and the x-coordinate of the second reference point; Calculate the second lateral offset between the x-coordinate of the first reference point and the x-coordinate of the third reference point; The steering offset ratio is calculated based on the ratio of the first lateral offset to the second lateral offset.
6. The method according to claim 4, characterized in that, When determining that the animal's face is in a side-facing posture based on the steering offset ratio, the first reference point or the second reference point is corrected, including: When the steering offset ratio is determined to be greater than a preset first threshold, the x-coordinate of the first reference point is reduced based on the steering offset ratio to generate an adjusted first reference point; or... When the steering offset ratio is determined to be less than a preset second threshold, the horizontal coordinate of the second reference point is increased based on the steering offset ratio to generate an adjusted second reference point.
7. The method according to claim 1, characterized in that, The preset key point subset includes two symmetrical key points on the animal's face and key points in the chin area. Based on the positional relationship of the preset key point subset in the key point set under a standard frontal facial pose, the rotation-corrected coordinates of each point in the key point set are calculated, including: Calculate the image rotation angle based on the positional relationship between the line connecting the two symmetrical key points and the horizontal line in the image to be processed. Based on the image rotation angle, a rotation transformation is performed on each point in the key point set to generate rotation-corrected coordinates.
8. The method according to claim 1, characterized in that, After performing facial landmark detection on the image to be processed, which contains animal faces, and generating a landmark set, the method further includes: Image quality detection is performed on the image to be processed based on the set of key points; Based on the image quality detection results, images that do not meet the image quality detection requirements are filtered out.
9. The method according to claim 8, characterized in that, Image quality detection is performed on the image to be processed based on the set of key points, including at least one of the following: abnormal facial pose detection; detection of the effective pixel ratio of the face in the image; and detection of the relationship between the face and the image edge.
10. The method according to claim 9, characterized in that, The abnormal facial posture detection includes: A third subset of key points is selected from the set of key points, which includes two symmetrical key points in the eye region, two symmetrical key points in the ear root region, key points in the nose region, and key points in the chin region. The positional relationship of key points in the third key point subset is used to determine whether the animal's face is in an abnormal posture. The abnormal posture includes at least one of the following: head tilt angle causing facial abnormality; head depression angle causing facial abnormality; head yaw angle causing facial abnormality.
11. The method according to claim 10, characterized in that, When the abnormal posture is caused by a head tilt angle resulting in facial abnormalities and / or a head yaw angle resulting in facial abnormalities, the determination of whether the animal's face is in an abnormal posture is based on the positional relationship of the key points in the third key point subset, including: The first distance value is calculated based on two symmetrical key points in the eye region; The second distance value is calculated based on the midpoint of the line connecting two symmetrical key points in the eye region and the key point in the nose region; Based on the relationship between the ratio of the first distance value and the second distance value and a preset elevation angle threshold, determine whether the head elevation angle causes facial abnormalities; and / or, Based on the relationship between the ratio of the first distance value and the second distance value and the preset yaw angle threshold, it is determined whether the head yaw angle causes facial abnormalities.
12. The method according to claim 11, characterized in that, When the abnormal posture is caused by a head tilt angle resulting in facial abnormalities, determining whether the animal's face is in an abnormal posture based on the positional relationship of key points in the third key point subset also includes: Calculate the third distance value based on two symmetrical key points in the ear root region; Calculate the fourth distance value based on the key points in the nose region and the key points in the chin region; The ratio of the first distance value to the third distance value and the ratio of the first distance value to the fourth distance value are used to jointly determine whether the head tilt angle causes facial abnormalities.
13. The method according to claim 9, characterized in that, The detection of the effective pixel percentage of the face in the image includes: The diameter of the minimum enclosing circle is calculated based on the positional relationship of all key points in the key point set. The minimum enclosing circle is the smallest circle that contains all key points in the key point set. If the ratio of the diameter length value to the diagonal length of the image to be processed is less than a preset first distance threshold, it indicates that the effective pixel proportion of the face in the image is too low; or, If the ratio of the diameter length value to the diagonal length of the image to be processed is greater than a preset second distance threshold, it indicates that the face occupies too high a percentage of the effective pixels in the image.
14. The method according to claim 9, characterized in that, The detection of the relationship between the face and the image edges includes: The minimum bounding rectangle is calculated based on all the key points in the key point set, and the coordinates of its four vertices are obtained. The minimum bounding rectangle is the rectangle with the smallest area that contains all the key points in the key point set. If the shortest distance from any vertex to the four edges of the image is less than a preset distance threshold, it is determined that the face is too close to the image edge.
15. An electronic device, characterized in that, The device includes: Processor; and, A memory storing computer instructions implemented by a computer for extracting regions of interest on an animal's face, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1-14.
16. A computer-readable storage medium, characterized in that, The method includes computer-implemented program instructions for extracting regions of interest on an animal's face, which, when executed by a processor, cause the method according to any one of claims 1-14 to be implemented.
Citation Information
Patent Citations
A face multi-area fusion expression recognition method based on depth learning
CN109344693A