Image processing method, device, and storage medium based on eye condition detection

The image processing method addresses suboptimal eye conditions in group photos by detecting and enhancing eye states, improving photo quality and user satisfaction through composite image processing.

JP7822369B2Active Publication Date: 2026-03-02DOUYIN VISION CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2023513999
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-31
Filing Date
2021-08-27
Publication Date
2026-03-02
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

Existing image capturing methods often result in suboptimal eye conditions in group photos, requiring repeated retakes and reducing user satisfaction due to inconsistent eye states among individuals.

Method used

An image processing method that detects and enhances the eye state of each face in a series of images, determining a target effect image where the eye state meets a predetermined condition and compositing it onto a reference image to improve the overall eye condition effect.

Benefits of technology

Enhances the quality of group photos by ensuring optimal eye conditions for all individuals, reducing the need for repeated shots and increasing user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007822369000002
    Figure 0007822369000002
  • Figure 0007822369000003
    Figure 0007822369000003
  • Figure 0007822369000004
    Figure 0007822369000004
Patent Text Reader

Abstract

The present disclosure discloses an image processing method, device, equipment, and storage medium using eye state detection. The method includes steps of detecting the eye state of a target face in a set of images to be processed, obtaining target area images whose eye state satisfies a predetermined condition, determining a target effect image corresponding to the target face from among the target area images, and finally combining the target effect image with a reference image in the set of images to be processed to obtain a target image corresponding to the set of images to be processed. In the present disclosure, by determining a target effect image for each face based on the eye state detection and then combining each target effect image with the reference image, the effect of the eye state of each face in the target image is improved, the quality of the target image is guaranteed, and user satisfaction with the target image is increased.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority from a Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on August 31, 2020, bearing application number 202010899317.9 and entitled "Image processing method, device and storage medium based on eye condition detection," the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the field of picture data processing, and more particularly to image processing methods, devices, equipment and storage media with eye state detection. [Background technology]

[0003] When taking a photo, there is a possibility that a problem may occur in the captured photo where the eye state of a person is not ideal (for example, "someone has their eyes closed"), which may require the user to retake the photo and may even require repeated retakes. In particular, when taking a group photo of multiple people, problems where the eye state is not ideal, such as "someone has their eyes closed" or "someone is not looking at the lens," are more likely to occur, which may result in repeated retakes, affecting the user's photography experience.

[0004] Currently, users generally manually select a photo in which most people have ideal eye conditions as the final group photo based on multiple photos taken repeatedly. However, since there are some problems with the eye conditions of people in the selected group photo being less than ideal, the optimal eye conditions of all people taken during the photo shoot cannot be expressed in the group photo, which reduces user satisfaction with the final group photo to some extent. Summary of the Invention [Problem to be solved by the invention]

[0005] In order to solve the above technical problems or at least part of them, the present disclosure provides an image processing method, device, equipment, and storage medium with eye condition detection, which can improve the eye condition effect of everyone in a group photo, ensure the quality of the group photo, and increase user satisfaction with the final group photo. [Means for solving the problem]

[0006] According to a first aspect, the present disclosure provides an image processing method with eye state detection, the method comprising: a step of detecting an eye state of a target face in an image set to be processed and acquiring a target area image in which the eye state satisfies a predetermined condition, wherein the image set to be processed includes images of a plurality of consecutive frames, and each of the images of the plurality of frames includes at least one face; determining a target effect image corresponding to the target face based on the target area image in which the eye state satisfies a predetermined condition; and combining a target effect image corresponding to the target face onto a reference image in the set of images to be processed to obtain a target image corresponding to the set of images to be processed.

[0007] In a preferred embodiment, the predetermined condition includes that the eye opening / closing degree value is greater than a predetermined opening / closing threshold value.

[0008] In a preferred embodiment, the step of detecting an eye state of a target face in the image set to be processed and acquiring a target area image in which the eye state satisfies a predetermined condition comprises: determining a face image of a target face from a set of images to be processed; The method includes a step of detecting an eye state for the face image of the target face, and acquiring, as a target area image, a face image of which the eye state satisfies a predetermined condition from among the face images of the target face.

[0009] In a preferred embodiment, the step of detecting an eye state for the face image of the target face includes: extracting an eye image of the target face from a face image of the target face; and performing eye state detection on the eye image of the target face.

[0010] In a preferred embodiment, the step of detecting an eye state of a target face in the image set to be processed and acquiring a target area image in which the eye state satisfies a predetermined condition comprises: determining an eye image of a target face from a set of images to be processed; The method includes a step of detecting an eye state for the eye images of the target face, and acquiring, as a target region image, an eye image whose eye state satisfies a predetermined condition from among the eye images of the target face.

[0011] In a preferred embodiment, the step of detecting an eye state for the eye image of the target face includes: determining position information of eye keypoints in the eye image of the target face; and determining an eye state corresponding to the eye image based on position information of the eye keypoints.

[0012] In a preferred embodiment, the step of determining position information of eye keypoints in the eye image of the target face comprises: The step includes inputting an eye image of the target face into a first model to obtain position information of eye keypoints in the eye image, wherein the position information of the eye keypoints is obtained by training the first model based on marked eye image samples.

[0013] In a preferred embodiment, the step of detecting an eye state for the eye image of the target face includes: determining eye state values ​​in the eye image of the target face, the eye state values ​​including an open state value and a closed state value; and determining an eye state corresponding to the eye image based on the eye state value.

[0014] In a preferred embodiment, the step of determining eye state values ​​in the eye image of the target face comprises: The step includes inputting an eye image of the target face into a second model to obtain eye state values ​​in the eye image, the eye state values ​​being obtained by training the second model based on marked eye image samples.

[0015] In a preferred embodiment, the eye state corresponding to the eye image is determined based on the ratio of the vertical eye opening width to the distance between the two horizontal eye corners, and the ratio of the vertical eye opening width to the distance between the two horizontal eye corners is determined based on the position information of the eye keypoints.

[0016] In a preferred embodiment, before the step of detecting an eye state of a target face in the image set to be processed, the method further comprises: The method includes a step of acquiring, in response to a trigger operation on the shutter button, a plurality of consecutive frames of preview images including a current image frame and having the current image frame as an end frame, as an image set to be processed.

[0017] In a preferred embodiment, before the step of synthesizing a target effect image corresponding to the target face onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed, the method further comprises: The method includes determining a current image frame in the set of images to be processed that corresponds to the pressing of the shutter button as a reference image.

[0018] In a preferred embodiment, the step of determining a target effect image corresponding to the target face based on a target area image in which the eye state satisfies a predetermined condition comprises: The method includes a step of determining, as a target effect image corresponding to the target face, a target area image in which the eye state of the target face satisfies a predetermined condition and in which the eye open / close degree value is greatest.

[0019] In a preferred embodiment, the step of determining a face image of a target face from the set of images to be processed comprises: performing face detection on a reference image in the image set to be processed to determine position information of each face in the reference image; The method includes a step of determining, based on the face position information, a face image corresponding to the position information of a target face among the faces on an image in the image set to be processed as the face image of the target face.

[0020] In a preferred embodiment, the step of determining a face image of a target face from the set of images to be processed comprises: performing face detection for each image in the set of images to be processed to obtain a face image; and determining a face image having a similarity greater than a predetermined similarity threshold as a face image of the target face.

[0021] According to a second aspect, the present disclosure provides an image processing device with eye state detection, the device comprising: a first detection module that detects an eye state of a target face in an image set to be processed and acquires a target area image in which the eye state satisfies a predetermined condition, the image set to be processed including a plurality of consecutive frames of images, each of the plurality of frames including at least one face; a first determination module that determines a target effect image corresponding to the target face based on a target area image in which the eye state satisfies a predetermined condition; a synthesis module for synthesizing a target effect image corresponding to the target face onto a reference image in the set of images to be processed to obtain a target image corresponding to the set of images to be processed.

[0022] According to a third aspect, the present disclosure provides a computer-readable storage medium having stored thereon instructions that, when executed by a terminal device, cause the terminal device to implement the above method.

[0023] According to a fourth aspect, the present disclosure provides an apparatus, comprising: a memory; a processor; and a computer program stored in the memory and executable by the processor, the computer program, when executed by the processor, realizing the above method. [Effects of the Invention]

[0024] Compared with the prior art, the technical solution provided by the present disclosure has the following advantages: The present disclosure provides an image processing method using eye state detection, which first detects the eye state of a target face in an image set to be processed, obtains a target area image in which the eye state satisfies a predetermined condition, then determines a target effect image corresponding to the target face from the target area image in which the eye state satisfies the predetermined condition, and finally, composites the target effect image onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed. In the present disclosure, by detecting the eye state and determining a target effect image for each face, and then composites the target effect image for each face onto the reference image, it is possible to improve the eye state effect for all people in the finally obtained target image, enhance the quality of the target image, and increase user satisfaction with the target image to a certain extent. [Brief explanation of the drawings]

[0025] The drawings herein are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification serve to explain the principles of the disclosure. In order to more clearly explain the technical solutions of the embodiments of the present disclosure or the prior art, the following briefly introduces drawings that need to be used in the description of the embodiments or the prior art, and it is obvious that those skilled in the art can obtain other drawings based on these drawings under the premise that they do not make efforts that amount to inventive step. [Figure 1] 1 is a flowchart of an image processing method with eye state detection according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of eye image extraction according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram of eye keypoints in an eye image according to an embodiment of the present disclosure. [Figure 4] 10 is a flowchart of another image processing method based on eye state detection according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a block diagram illustrating a configuration of an image processing device that detects an eye state according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a block diagram illustrating a configuration of an image processing device based on eye state detection according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0026] In order to make the above objectives, features and advantages of the present disclosure more clearly understood, the solutions of the present disclosure are further described below. It should be noted that, if not contradictory, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0027] In order to facilitate a thorough understanding of the present disclosure, numerous specific details are set forth in the following description; however, the present disclosure may be embodied in forms different from those described herein, and it is apparent that the embodiments in the specification are only some of the embodiments of the present disclosure, rather than all of the embodiments.

[0028] The eye conditions of people in an image (for example, a group photo) are a factor for evaluating image quality. For example, if the image is a group photo, in an actual shooting scene, in order to express the optimal eye conditions of all people in the group photo, multiple group photos are taken by repeating multiple re-shoots, and then an ideal group photo is manually selected from the multiple group photos.

[0029] The method of repeating the above-mentioned multiple re-shoots not only reduces people's photography experience, but also fails to ensure that the eye conditions of all people in the re-shoots are ideal, which affects users' satisfaction with the group photo.

[0030] Therefore, the present disclosure provides an image processing method using eye state detection, which first detects the eye state of a target face in an image set to be processed, obtains a target area image in which the eye state satisfies a predetermined condition, then determines a target effect image corresponding to the target face from the target area image in which the eye state satisfies the predetermined condition, and finally composites the target effect image onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed.

[0031] Based on the above-mentioned shooting scene, the image processing method by eye state detection provided by the embodiment of the present disclosure detects the eye states of people in the group photo after taking a group photo, determines a target effect image for each face in the group photo, and then composites the target effect image of each face onto the original group photo, thereby improving the eye state effect of everyone in the final group photo, improving the quality of the group photo, and increasing user satisfaction with the group photo.

[0032] Based on this, an embodiment of the present disclosure provides an image processing method based on eye state detection, and referring to FIG. 1, there is shown a flowchart of an image processing method based on eye state detection according to an embodiment of the present disclosure, which includes the following steps: Step S101: The eye state of a target face in an image set to be processed is detected, and a target region image in which the eye state satisfies a predetermined condition is obtained.

[0033] The specified condition includes that the eye opening / closing degree value is greater than a specified opening / closing threshold, and the image set to be processed includes multiple consecutive frames of images, each of which includes at least one face.

[0034] In a preferred embodiment, when a trigger operation on the shutter button (e.g., pressing the shutter button) is detected in a scene where a photograph is to be taken, a plurality of consecutive frames of preview images, including the current image frame and with the current image frame as the end frame, are obtained as images of the plurality of consecutive frames, thereby obtaining a set of images to be processed in an embodiment of the present disclosure.

[0035] In practical applications, in the camera's preview mode, the preview images in the camera's preview interface are stored in the form of a preview stream. In an embodiment of the present disclosure, when the camera's preview mode detects that the shutter button is pressed, not only the current image frame (i.e., the photo taken by the camera) but also the most recent N frame preview images are obtained from the preview pictures in the stored preview stream. The most recent N frame preview images, together with the current image frame, constitute an image set to be processed. Typically, the image set to be processed includes 8 or 16 frame images, and the embodiment of the present disclosure does not limit the number of images in the image set to be processed. In some other embodiments, the image set to be processed may include more frame images.

[0036] In another preferred embodiment, in a scene where a photograph is to be taken, if the current mode is a continuous shooting mode and a trigger operation of pressing the shutter button is detected, multiple frames of images are obtained by continuous shooting as consecutive multiple frame images, and further, an image set to be processed in an embodiment of the present disclosure is obtained.

[0037] In an embodiment of the present disclosure, after acquiring a set of images to be processed, the eye state of a target face in the set of images to be processed is detected, where the target face is a face corresponding to the same person among multiple faces in the set of images to be processed.

[0038] In a preferred embodiment, the step of detecting the eye state of a target face in an image set to be processed includes the steps of determining a facial image of the target face from the image set to be processed, then performing eye state detection on the facial image of the target face, and obtaining, as a target area image, a facial image of the target face whose eye state satisfies a predetermined condition.

[0039] The embodiments of the present disclosure provide at least two methods for determining a face image of a target face from a set of images to be processed, which are introduced below: In a preferred embodiment, face detection is performed on a reference image in a set of images to be processed to determine position information of each face on the reference image, and then, based on the position information of each face, a face image in an image in the set of images to be processed that corresponds to the position information of a target face in each face is determined as the face image of the target face.

[0040] In practical applications, during a single photograph, the current image frame corresponding to the shutter button press is generally an image in which most people's eyes are in good condition. Therefore, an embodiment of the present disclosure can determine the current image frame corresponding to the shutter button press in the image set to be processed as a reference image. In this way, the position information of each face is determined based on the reference image, and then a face image corresponding to the target face is further determined based on the position information of each face, thereby improving the accuracy of the face image corresponding to the target face.

[0041] In an embodiment of the present disclosure, after determining a reference image in a set of images to be processed, face detection is performed on the reference image based on a machine learning model, thereby determining position information of each face in the reference image. Since the position information of each face in multiple frames of images captured continuously during a single shooting is basically the same, face images corresponding to target faces in other images in the set of images to be processed can be further determined based on the position information of each face determined in the reference image. Here, in the set of images to be processed, face images at the same position in each image belong to the same person.

[0042] The facial image of the target face may be the smallest rectangular area that includes the target face, which is determined based on position information of the target face.

[0043] In another preferred embodiment, a facial image of a target face may be determined from a set of images to be processed by combining face detection and similarity calculation methods. Specifically, face detection is performed on each image in the set of images to be processed to obtain a facial image. Then, a facial image with a similarity greater than a predetermined similarity threshold is determined to be the facial image of the target face.

[0044] In practical applications, since the similarity of the facial image of the target face is high, in the embodiment of the present disclosure, after determining the facial image in each image in the image set to be processed, the facial image of the target face can be determined based on the similarity of the facial images. Specifically, the facial image whose similarity is greater than a predetermined similarity threshold is determined as the facial image of the target face.

[0045] In the embodiment of the present disclosure, in the process of performing eye state detection on a face image, first, an eye image is extracted from the face image, and then eye state detection is performed on the eye image to complete eye state detection of the corresponding face image. The embodiment of the present disclosure provides a method for performing eye state detection on an eye image, which will be described below.

[0046] In another preferred embodiment, in the process of detecting the eye state of a target face in an image set to be processed, after determining the eye image of the target face from the image set to be processed, eye state detection is performed on the eye image of the target face, and eye images of the target face whose eye state satisfies predetermined conditions can be obtained as target area images.

[0047] In practical applications, the machine learning model can perform eye detection on a reference image in a set of images to be processed, thereby determining the position information of the eyes on the reference image. Then, based on the position information of the eyes determined on the reference image, eye images corresponding to the eyes on each image in the set of images to be processed can be further determined. Here, eye images at the same position in each image in the set of images to be processed belong to the same person.

[0048] The eye image includes a minimum rectangular area of ​​the eye. Specifically, the eye image may be a minimum rectangular area including the left eye, a minimum rectangular area including the right eye, or a minimum rectangular area including both the left eye and the right eye.

[0049] In another preferred embodiment, eye detection and similarity calculation methods may be combined to determine eye images of a target face from a set of images to be processed. Specifically, eye detection is performed on each image in the set of images to be processed to obtain eye images. Then, eye images with similarity greater than a predetermined similarity threshold are determined to be eye images of the target face.

[0050] In an embodiment of the present disclosure, after detecting the eye states of target faces in a set of images to be processed, target region images are acquired in which the eye states of each face satisfy a predetermined condition. The target region images may be face images or eye images.

[0051] In the embodiments of the present disclosure, the eye state corresponding to the eye image may be determined based on the position information or eye state value of the eye keypoints in the eye image, or a combination of the position information and the eye state value of the eye keypoints. The embodiments of the present disclosure provide a specific implementation of determining the eye state corresponding to the eye image, which will be introduced hereinafter.

[0052] Step S102: Based on the target area image in which the eye state satisfies a predetermined condition, a target effect image corresponding to the target face is determined.

[0053] In practical applications, in ideal photographs, the eye states of all people are generally open, and the degree of eye opening should meet a certain standard. Therefore, in an embodiment of the present disclosure, a target area image whose eye state of a target face meets a predetermined condition is first determined, and then a target effect image corresponding to each face is further determined based on the determined target area image. The eye state meeting the predetermined condition means that the eye opening degree value is greater than a predetermined opening / closing threshold.

[0054] In a preferred embodiment, the target area image of the target face with the largest eye opening / closing degree value is determined as the target effect image corresponding to the target face, thereby improving the eye opening / closing degree of each face in the target image and further increasing the user's satisfaction with the target image.

[0055] In another preferred embodiment, one of the target area images of a target face is determined as a target effect image corresponding to the target face, thereby satisfying the user's basic requirements for the eye condition effect of the face in the target image.

[0056] In a preferred embodiment, if it is determined that the eye condition of a face (e.g., the first face) in the reference image satisfies a predetermined condition, synthesis is not performed on the first face, thereby improving the efficiency of image processing.

[0057] Step S103: A target effect image corresponding to the target face is composited onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed.

[0058] In an embodiment of the present disclosure, after determining a target effect image corresponding to each face, the target effect image is composited onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed.

[0059] Since the target image is obtained based on the target effect image, the target image can maximize the effect of the eye condition of all people on the image, thereby improving the quality of the target image and improving the user's satisfaction with the target image to a certain extent.

[0060] In a preferred embodiment, the target effect image corresponding to each face has position information, and based on the position information of the target effect image, the target effect image is composited at the corresponding position on the reference image.

[0061] In the implementation of the present disclosure, any one image in the set of images to be processed may be determined as the reference image. The embodiments of the present disclosure do not specifically limit the method for determining the reference image, and those skilled in the art may select the method according to their actual needs.

[0062] In an image processing method using eye state detection provided by an embodiment of the present disclosure, first, the eye state of a target face in an image set to be processed is detected, and a target area image whose eye state satisfies a predetermined condition is obtained. Then, a target effect image corresponding to the target face is determined from the target area image whose eye state satisfies the predetermined condition. Finally, the target effect image is superimposed on a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed. In the embodiment of the present disclosure, by detecting the eye state, determining the target effect image for each face, and then superimposing the target effect image for each face on the reference image, it is possible to improve the eye state effect of all people in the target image, enhance the quality of the target image, and improve user satisfaction with the target image to a certain extent.

[0063] In the image processing method by eye state detection provided by the embodiments of the present disclosure, the eye state corresponding to the eye image can be determined based on the position information of the eye key points.

[0064] In a preferred embodiment, corresponding eye images are extracted from eight frames of face images, as shown in Fig. 2. Then, for each eye image, position information of eye keypoints in the eye image is determined, and then an eye state corresponding to the eye image is determined based on the position information of the eye keypoints. In one implementation, the eye state corresponding to the eye image of the target face is set as the eye state of the face image corresponding to the target face, where the eye state can be represented by an eye open / close degree value.

[0065] As shown in Figure 3, the eye keypoints may be keypoint 1 on the left corner of the eye, keypoints 2 and 3 on the upper eyelid, keypoint 4 on the right corner of the eye, and keypoints 5 and 6 on the lower eyelid. After determining the above six eye keypoints, the eye open / close degree value is determined based on the position information of each eye keypoint.

[0066] In a preferred embodiment, the distance between key points 1 and 4 in Fig. 3 is taken as the distance between the two corners of the eyes in the horizontal direction, and the average value of the distance between key points 2 and 6 and the distance between key points 3 and 5 is taken as the vertical eye opening width. The ratio of the vertical eye opening width to the horizontal eye opening width is then determined as the eye opening degree value.

[0067] In a preferred embodiment, a machine learning model can be used to determine the position information of the eye keypoints, specifically, a first model is trained using eye image samples marked with the position information of the eye keypoints, and the eye images are input to the trained first model, which then outputs the position information of the eye keypoints in the eye images after processing by the first model.

[0068] In addition, in the image processing method for detecting an eye state provided by the embodiments of the present disclosure, an eye state corresponding to an eye image may be determined based on an eye state value. The eye state value includes an open state value and a closed state value. Specifically, the eye state value may be a value within a range of [0, 1], where a larger eye state value indicates a larger eye open / closed degree value, and conversely, a smaller eye state value indicates a smaller eye open / closed degree value. Specifically, the closed state value may be a value within a range of [0, 0.5], and the open state value may be a value within a range of [0.5, 1]. In some other embodiments, the closed state value may be a value within a range of [0, 0.5], and the open state value may be a value within a range of [0.5, 1].

[0069] In a preferred embodiment, a machine learning model can be used to determine the eye state value, specifically, the eye image samples marked with eye state values ​​are used to train a second model, and the eye images are input to the trained second model, which outputs the eye state value of the eye image after being processed by the second model.

[0070] In an embodiment of the present disclosure, the eye state corresponding to the eye image is determined according to the eye state value, and the target region image with the largest eye state value of the target face is determined as the target effect image.

[0071] In order to improve the accuracy of eye state detection, an embodiment of the present disclosure combines the position information and eye state value of the eye keypoints to determine the eye opening / closing degree value corresponding to the facial image, thereby improving the accuracy of the target effect image determined based on the eye opening / closing degree value and further improving the quality of the target image.

[0072] An embodiment of the present disclosure provides an image processing method based on eye state detection, and referring to FIG. 4, there is shown a flowchart of another image processing method based on eye state detection provided by an embodiment of the present disclosure, which includes the following steps: Step S401: Based on the image set to be processed, a face image belonging to a target face is determined.

[0073] The image set to be processed includes a plurality of consecutive frames of preview images, with the current image frame corresponding to the shutter button being the end frame.

[0074] Step S401 in the embodiment of the present disclosure can be understood by referring to the description of the above embodiment, and will not be described more than necessary here.

[0075] Step S402: Extract an eye image from the target face image.

[0076] Referring to FIG. 2, after a face image belonging to a target face is determined, an eye image belonging to the target face is extracted based on the face image of the determined target face.

[0077] In a preferred embodiment, eyes are detected in a face image, position information of the eyes in the face image is determined, and then a rectangular frame area including the eyes is determined based on the position information of the eyes, and the rectangular frame area is extracted from the face image. The image corresponding to the extracted rectangular area is defined as the eye image. The method of eye detection will not be described more than necessary.

[0078] In practical applications, since the eye states of both eyes on a facial image are considered to be basically the same, the extracted eye image in the embodiments of the present disclosure may include only one eye in the facial image, thereby improving the efficiency of image processing.

[0079] Step S403: Determine the eye state values ​​and position information of the eye key points in the eye image.

[0080] In an embodiment of the present disclosure, after extracting the eye image, the eye state values ​​and position information of the eye keypoints in the eye image are determined.

[0081] In a preferred embodiment, the eye state values ​​and the position information of the eye keypoints can be determined by a machine learning model, specifically, the eye image samples marked with the position information of the eye keypoints and the eye state values ​​are used to train a third model, and the eye images are input to the trained third model, and after being processed by the third model, the eye state values ​​and the position information of the eye keypoints of the eye images are output.

[0082] Step S404: Determine an eye open / close degree value corresponding to the face image based on the eye state value and the position information of the eye key points.

[0083] In an embodiment of the present disclosure, after determining the eye state value and the position information of the eye keypoints in the eye image, the ratio between the vertical eye opening width and the horizontal distance between the two eye corners in the face image is determined based on the position information of the eye keypoints, and then the ratio between the vertical eye opening width and the horizontal distance between the two eye corners and the eye state value are combined to determine the eye opening degree value corresponding to the face image.

[0084] In a preferred embodiment, referring to FIG. 3, the eye open / close degree value corresponding to the face image is calculated by equation (1), which is as follows:

number

[0085] Step S405: From the face images belonging to the target face, a face image whose eye opening / closing degree value is greater than a predetermined opening / closing threshold is determined.

[0086] In an embodiment of the present disclosure, after determining the face images belonging to the target face, for any one of the face images, first, face images in a closed state are removed based on the eye opening / closing degree value. Then, face images whose eye opening / closing degree value is equal to or less than a predetermined opening / closing threshold are removed based on the eye opening / closing degree value. The remaining face images may be sorted based on the eye opening / closing degree value to determine face images whose eye opening / closing degree value is greater than the predetermined opening / closing threshold.

[0087] In a preferred embodiment, if there is no face image of a certain face whose eye opening / closing degree value is greater than a predetermined opening / closing threshold, the effect of that face in the reference image may be retained without processing the face image corresponding to that face in the reference image.

[0088] Step S406: A target effect image corresponding to the target face is determined from the face images whose eye opening / closing degree value is greater than a predetermined opening / closing threshold value.

[0089] In an embodiment of the present disclosure, after determining the face images in which the eye occlusion degree value of each face is greater than a predetermined occlusion threshold, the method randomly selects one face image in which the eye occlusion degree value is greater than the predetermined occlusion threshold from the face images corresponding to the target face, and sets the randomly selected face image as a target effect image corresponding to the target face, thereby improving the eye condition effect of the target face in the target image. Target effect images corresponding to the other faces can be obtained by performing similar processing on face images corresponding to the other faces.

[0090] In a preferred embodiment, a larger eye opening / closing degree value indicates a larger eye opening degree and an optimal eye state can be expressed. Therefore, in an embodiment of the present disclosure, a face image with the largest eye opening / closing degree value is selected as a target effect image corresponding to the target face from face images with eye opening / closing degree values ​​greater than a predetermined opening / closing threshold, thereby maximizing the effect of the eye state of the target face in the target image.

[0091] Step S407: A target effect image corresponding to the target face is composited onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed.

[0092] In an embodiment of the present disclosure, after determining a target effect image corresponding to each face, each target effect image is composited onto a reference image to finally obtain a target image corresponding to the image set to be processed.

[0093] In a shooting scene involving multiple people, by taking a group photo only once based on the image processing method for detecting eye conditions provided by the present disclosure, the effect of the eye conditions of as many people as possible in the group photo can be improved, eliminating the need to retake photos multiple times, improving the user's experience with group photo shooting and providing users with a group photo that they are very satisfied with.

[0094] Based on the same inventive concept as the above method embodiment, the present disclosure further provides an image processing device based on eye state detection. Referring to FIG. 5 , the image processing device based on eye state detection is provided by the embodiment of the present disclosure, and the device comprises: a first detection module 501 that detects the eye state of a target face in an image set to be processed and acquires a target area image in which the eye state satisfies a predetermined condition, the image set to be processed including a plurality of consecutive frames of images, each of the plurality of frames including at least one face; a first determination module 502 for determining a target effect image corresponding to the target face based on the target area image in which the eye state satisfies a predetermined condition; and a synthesis module 503 that synthesizes a target effect image corresponding to the target face onto a predetermined reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed.

[0095] In a preferred embodiment, the predetermined condition includes that the eye opening / closing degree value is greater than a predetermined opening / closing threshold value.

[0096] In a preferred embodiment, the first detection module comprises: a first determination submodule for determining a face image of a target face from a set of images to be processed; and a first detection sub-module that performs eye state detection on the face image of the target face and acquires, as a target area image, a face image of the target face whose eye state satisfies a predetermined condition.

[0097] In a preferred embodiment, the first detection sub-module: an extraction sub-module for extracting an eye image of the target face from the face image of the target face; and a second detection sub-module that performs eye state detection on the eye image of the target face.

[0098] In a preferred embodiment, the first detection module comprises: a second determination sub-module for determining an eye image of a target face from the set of images to be processed; and a third detection sub-module that performs eye state detection on the eye images of the target face and acquires, as target area images, eye images of the target face whose eye states satisfy predetermined conditions.

[0099] In a preferred embodiment, the second detection module or the third detection sub-module: a third determination sub-module for determining position information of eye keypoints in the eye image of the target face; and a fourth determination sub-module for determining an eye state corresponding to the eye image based on position information of the eye keypoints.

[0100] In a preferred embodiment, the third determination sub-module specifically: An eye image of the target face is input into a first model to obtain position information of eye keypoints in the eye image, and the first model is obtained by training based on eye image samples in which position information of eye keypoints is marked.

[0101] In a preferred embodiment, the second detection module or the third detection sub-module: a fifth determination sub-module for determining eye state values ​​in the eye image of the target face, the eye state values ​​including an open state value and a closed state value; and a sixth determination sub-module that determines an eye state corresponding to the eye image based on the eye state value.

[0102] In a preferred embodiment, the fifth determination sub-module specifically: An eye image of the target face is input to a second model to obtain eye state values ​​in the eye image, and the second model is obtained by training based on eye image samples with marked eye state values.

[0103] In a preferred embodiment, the eye state corresponding to the eye image is determined based on the ratio of the vertical eye opening width to the distance between the two horizontal eye corners, and the ratio of the vertical eye opening width to the distance between the two horizontal eye corners is determined based on the position information of the eye keypoints.

[0104] In a preferred embodiment, the device comprises: The camera further includes an acquisition module that acquires, in response to a trigger operation on the shutter button, a plurality of consecutive frames of preview images including a current image frame and having the current image frame as an end frame, as an image set to be processed.

[0105] In a preferred embodiment, the device comprises: The camera further includes a second determination module for determining a current image frame in the image set to be processed, which corresponds to the pressing of the shutter button, as a reference image.

[0106] In a preferred embodiment, the first decision module specifically: Among the target area images in which the eye state of the target face satisfies a predetermined condition, the target area image in which the eye open / close degree value is the largest is determined as the target effect image corresponding to the target face.

[0107] In a preferred embodiment, the first determination sub-module: a seventh determination sub-module that performs face detection on a reference image in the image set to be processed and determines position information of each face in the reference image; and an eighth determination submodule that determines, based on the position information of each face, a face image corresponding to the position information of a target face among the faces on an image in the image set to be processed as the face image of the face.

[0108] In a preferred embodiment, the first determination sub-module: a fourth detection sub-module that performs face detection on each image in the image set to be processed to obtain a face image; and a ninth determination sub-module that determines a face image whose similarity is greater than a predetermined similarity threshold as a face image of the target face.

[0109] In an image processing device using eye state detection provided by an embodiment of the present disclosure, the eye state of a target face in an image set to be processed is detected, a target area image in which the eye state of the target face satisfies a predetermined condition is obtained, a target effect image corresponding to the face is determined from the target area image in which the eye state satisfies the predetermined condition, and finally, the target effect image is composited onto a reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed. In the embodiment of the present disclosure, by detecting the eye state and determining the target effect image for each face, and then composited onto the reference image, the eye state effect of all people in the finally obtained target image can be improved, the quality of the target image can be improved, and user satisfaction with the target image can be increased to a certain extent.

[0110] In addition, the embodiment of the present disclosure further provides an image processing device based on eye state detection, as shown in FIG. 6 : The image processing device for eye condition detection may include a processor 601, a memory 602, an input device 603, and an output device 604. The number of processors 601 in the image processing device for eye condition detection may be one or more, and one processor is shown as an example in Fig. 6. In some embodiments of the present disclosure, the processor 601, the memory 602, the input device 603, and the output device 604 may be connected by a bus or other method, and Fig. 6 shows an example of connection by a bus.

[0111] The memory 602 can store software programs and modules, and the processor 601 can execute the software programs and modules stored in the memory 602 to perform various functional applications and data processing of the image processing device based on eye condition detection. The memory 602 mainly includes a program storage area and a data storage area, and the program storage area stores an operating system, an application program required for at least one function, etc. The memory 602 may also include high-speed random access memory or non-volatile memory, such as at least one magnetic disk memory, flash memory, or other volatile solid-state memory. The input device 603 can be used to receive input numeric or character information and to generate signal input related to user settings and function control of the image processing device based on eye condition detection.

[0112] Specifically, in this embodiment, the processor 601 loads executable files corresponding to the processes of one or more application programs into the memory 602 in accordance with the following instructions, and the processor 601 executes the application programs stored in the memory 602, thereby realizing various functions of the image processing device based on the above-mentioned eye condition detection.

[0113] An embodiment of the present disclosure further provides a computer-readable storage medium, the computer-readable storage medium storing instructions, which, when executed by a terminal device, cause the terminal device to implement the above method.

[0114] It should be noted that, in this specification, relational terms such as "first" and "second" are used to distinguish one entity or operation from another and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise" and "comprises," or any other variations thereof, are intended to include non-exclusive inclusions, whereby a process, method, article, or device that includes a set of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent in such process, method, article, or device. Unless further limited, an element defined by the phrase "comprises" does not exclude that the process, method, article, or device that includes that element also includes other identical elements.

[0115] The foregoing are only specific embodiments of the present disclosure, which will enable those skilled in the art to understand or realize the present disclosure. Various modifications to these examples will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other examples without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the examples described herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image processing method based on eye state detection, comprising: determining a face image of each face determined by face detection on another image in the set of images to be processed based on position information of the face on a reference image in the set of images to be processed; a step of detecting an eye state of the face image of each of the faces and acquiring a target area image in which the eye state satisfies a predetermined condition, wherein the image set to be processed includes a plurality of consecutive frames of preview images including a current image frame and having the current image frame as an end frame, which are acquired in response to a trigger operation on a shutter button, and each of the plurality of frames of preview images includes at least one face; determining, as a target effect image corresponding to each face, a target area image with a maximum eye open / close degree value among the target area images whose eye states satisfy a predetermined condition; and combining a target effect image corresponding to each of the faces onto the reference image in the set of images to be processed to obtain a target image corresponding to the set of images to be processed; the reference image is any one of the images in the image set to be processed; A method characterized by:

2. The step of detecting an eye state of each face image includes: extracting an eye image from the face image; determining position information of eye keypoints in the eye image, the eye keypoints including a left eye corner keypoint, an upper eyelid first keypoint, an upper eyelid second keypoint, a right eye corner keypoint, a lower eyelid first keypoint, and a lower eyelid second keypoint, the upper eyelid first keypoint and the lower eyelid first keypoint being vertically aligned, and the upper eyelid second keypoint and the lower eyelid second keypoint being vertically aligned; the distance from the left eye corner key point to the right eye corner key point is determined as the horizontal eye corner distance, the average value of the distance from the upper eyelid first key point to the lower eyelid first key point and the distance from the upper eyelid second key point to the lower eyelid second key point is determined as the vertical eye opening width, and the ratio of the vertical eye opening width to the horizontal eye corner distance is determined as an eye opening degree value, and the eye state is determined based on the eye opening degree value.

3. The method of claim 2 , wherein the predetermined condition includes the eye closure value being greater than a predetermined closure threshold.

4. The step of determining position information of eye keypoints in the eye image comprises:

3. The method of claim 2, further comprising: inputting the eye image into a first model to obtain the position information of the eye keypoints in the eye image, the first model being obtained by training based on marked eye image samples.

5. The method of claim 1, further comprising: determining eye state values ​​in the eye image, the eye state values ​​including an open state value and a closed state value; The method of claim 2 , further comprising determining an eye state corresponding to the eye image based on the eye state value.

6. The step of determining an eye state value in the eye image comprises:

6. The method of claim 5, further comprising inputting the eye image into a second model to obtain eye state values ​​in the eye image, the second model being obtained by training the eye state values ​​based on marked eye image samples.

7. before the step of compositing a target effect image corresponding to each face onto the reference image in the image set to be processed to obtain a target image corresponding to the image set to be processed, 2. The method of claim 1, further comprising determining a current image frame in the set of images to be processed that corresponds to a shutter button press as a reference image.

8. The step of determining a face image of each face determined by face detection on another image in the image set to be processed based on position information on a reference image in the image set to be processed, comprises: performing face detection for each image in the set of images to be processed to obtain a face image; and determining a face image having a similarity greater than a predetermined similarity threshold as the face image for each face.

9. An image processing device using eye state detection, a first detection module that determines a face image of each face determined by face detection on another image in the image set to be processed based on position information on a reference image in the image set to be processed, detects an eye state of the face image of each face, and acquires a target area image in which the eye state satisfies a predetermined condition, wherein the image set to be processed includes a plurality of consecutive frames of preview images including a current image frame and having the current image frame as an end frame, which are acquired in response to a trigger operation on a shutter button, and each of the plurality of frame preview images includes at least one face; a first determination module that determines, among the target area images whose eye states satisfy a predetermined condition, a target area image whose eye open / close degree value is the largest as a target effect image corresponding to each of the faces; a synthesis module for synthesizing a target effect image corresponding to each of the faces onto the reference image in the set of images to be processed to obtain a target image corresponding to the set of images to be processed; the reference image is any one of the images in the image set to be processed; An apparatus characterized in that

10. The first detection module further comprises: extracting an eye image from the face image; determining position information of eye keypoints in the eye image, the eye keypoints including a left eye corner keypoint, an upper eyelid first keypoint, an upper eyelid second keypoint, a right eye corner keypoint, a lower eyelid first keypoint, and a lower eyelid second keypoint, wherein the upper eyelid first keypoint and the lower eyelid first keypoint are vertically aligned, and the upper eyelid second keypoint and the lower eyelid second keypoint are vertically aligned; The device according to claim 9, wherein the distance from the left eye corner key point to the right eye corner key point is set as the horizontal eye corner distance, the average value of the distance from the upper eyelid first key point to the lower eyelid first key point and the distance from the upper eyelid second key point to the lower eyelid second key point is set as the vertical eye opening width, and the ratio of the vertical eye opening width to the horizontal eye corner distance is determined as an eye opening degree value, and the eye state is determined based on the eye opening degree value.

11. A computer-readable storage medium having stored thereon instructions which, when executed by a terminal device, cause the terminal device to implement the method of any one of claims 1 to 8.

12. 9. An apparatus comprising: a memory; a processor; and a computer program stored in the memory and executable by the processor, the computer program implementing the method of any one of claims 1 to 8 when executed by the processor.

Citation Information

Patent Citations

  • Fatigue driving detection and early warning system based on machine vision

    CN110246305A

  • Electronic still camera and photographing method

    JP2000083210A

  • Image processing apparatus

    JP2002199202A

  • Method, device and program for discriminating state of eye

    JP2007323104A

  • Imaging apparatus, imaging method and imaging program

    JP2009027462A