Method, apparatus, and computer program product for closed eye defect repair
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本申请实施例提供了一种闭眼缺陷修复方法、装置、电子设备和计算机程序产品,以至少解决相关技术中通过替换图像中的闭眼区域得到的修复后图像的修复效果生硬、不自然且动态真实性较低的问题
Smart Images

Figure CN122554723A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to a method, device, electronic device and computer program product for repairing closed-eye defects. Background Technology
[0002] With the widespread use of electronic devices, taking photos using electronic devices has become a daily routine. However, when taking portrait photos using electronic devices, if the subject blinks at the moment the shutter is pressed, the resulting photo may have a quality issue where the subject's eyes are closed.
[0003] To correct closed-eye defects in images, techniques typically select an image with open eyes (i.e., an image of the eyes in an open state) from a pre-built image library to replace the closed-eye area in the original image. However, this open-eye image may not accurately reflect the actual open-eye state of the subject. Therefore, the corrected image obtained using this method often exhibits a stiff and unnatural appearance in the corrected area, and the corrected image often shows low continuity with preceding and following frames in the shooting preview interface, indicating low dynamic realism.
[0004] There is no effective solution to the problem that the restoration effect of the image obtained by replacing the closed eye area in the image is stiff, unnatural and has low dynamic realism. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and computer program product for repairing closed-eye defects, in order to at least solve the problems in related technologies where the repaired image obtained by replacing the closed-eye area in an image has a stiff, unnatural, and low dynamic realism.
[0006] In a first aspect, embodiments of this application provide a method for repairing eye-closed defects, applied to an electronic device. The method includes: taking a preview image of a target object and storing it in a preview stream buffer; if the current frame image includes an eye-closed defect of the target object, obtaining multiple consecutive images, including the current frame image, from the preview stream buffer; based on the multiple consecutive images, predicting the open-eye state of the target object at the current moment corresponding to the current frame image, and obtaining an open-eye state parameter; and repairing the eye-closed defect of the current frame image based on the open-eye state parameter, to obtain a repaired image.
[0007] In some embodiments, predicting the open-eye state of the target object at the current moment corresponding to the current frame image based on the multi-frame continuous images to obtain the open-eye state parameters includes: performing spatiotemporal feature encoding on the eye region of the target object in the multi-frame continuous images to obtain the open-eye state fusion feature of the target object; and predicting the open-eye state of the target object at the current moment based on the open-eye state fusion feature to obtain the open-eye state parameters.
[0008] In some embodiments, the step of performing spatiotemporal feature encoding on the eye region of the target object in the multiple consecutive images to obtain the eye-opening state fusion feature of the target object includes: for each of the consecutive images other than the current frame image, transforming and matching the eye region of the target object in the consecutive images from the current coordinate system corresponding to the consecutive images to the target coordinate system corresponding to the current frame image; and performing spatiotemporal feature encoding on multiple eye regions in the target coordinate system to obtain the eye-opening state fusion feature of the target object.
[0009] In some embodiments, the step of performing spatiotemporal feature encoding on the eye region of the target object in the multi-frame continuous images to obtain the fusion feature of the target object's open-eye state includes: extracting spatial features from the eye region of the target object in each frame of the continuous images to obtain a spatial feature map corresponding to each frame of the continuous images; concatenating the spatial feature maps corresponding to each frame of the continuous images into a tensor, and extracting temporal features from the tensor to obtain a temporal feature map; and performing global average pooling on the temporal feature map according to the time dimension to obtain the fusion feature of the target object's open-eye state.
[0010] In some embodiments, the eye-opening state parameters include static eye parameters and eye-opening motion parameters; the step of repairing the closed-eye defect in the current frame image based on the eye-opening state parameters to obtain a repaired image includes: acquiring a reference eye image of the target object; inputting the reference eye image and the eye-opening state parameters into a trained image generation model, so that the image generation model uses the reference eye image and the eye-opening state parameters as conditions to generate a repaired eye region; adding a motion blur effect to the repaired eye region based on the eye-opening motion parameters to obtain an enhanced eye region; and fusing the enhanced eye region with the current frame image to obtain the repaired image.
[0011] In some embodiments, the image generation model includes a first encoding module, a second encoding module, and a generation module; the step of inputting the reference eye image and the eye-opening state parameters into the trained image generation model, so that the image generation model uses the reference eye image and the eye-opening state parameters as conditions to generate a repaired eye region, includes: inputting the reference eye image into the first encoding module to encode the reference eye image into a first style vector based on the first encoding module; inputting the eye-opening state parameters into the second encoding module to encode the eye-opening state parameters into a second style vector based on the second encoding module; adding the first style vector and the second style vector and then inputting the result into the generation module, so that the generation module generates the repaired eye region based on the first style vector and the second style vector.
[0012] In some embodiments, the method further includes: generating and outputting a captured image based on the repaired image in response to a target operation.
[0013] Secondly, embodiments of this application provide an eye-closing defect repair device applied to an electronic device. The device includes: a preview shooting module for previewing and shooting a target object and storing it in a preview stream buffer; an acquisition module for acquiring multiple consecutive images, including the current frame image, from the preview stream buffer when the current frame image includes the eye-closing defect of the target object; a prediction module for predicting the eye-opening state of the target object at the current moment corresponding to the current frame image based on the multiple consecutive images, and obtaining an eye-opening state parameter; and a repair module for repairing the eye-closing defect of the current frame image based on the eye-opening state parameter, and obtaining a repaired image.
[0014] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the eye-closing defect repair method described in any one of the first aspects above.
[0015] Fourthly, embodiments of this application provide a computer program product, including a computer program, which, when run, causes the eye-closing defect repair method described in any one of the first aspects to be executed.
[0016] Compared to related technologies, the eye-closure defect repair method, apparatus, electronic device, and computer program product provided in this application can launch a camera application and display a shooting preview interface. If the current frame image in the preview stream buffer includes the eye-closure defect of the target object, multiple consecutive images, including the current frame image, can be obtained from the preview stream buffer. This preview stream buffer is used to store preview images captured by the camera on the target object. Then, based on the multiple consecutive images, the open-eye state of the target object at the current moment corresponding to the current frame image can be predicted to obtain open-eye state parameters. Subsequently, based on the open-eye state parameters, the eye-closure defect can be repaired on the current frame image to obtain a repaired image. Finally, the repaired image can be displayed on the shooting preview interface, and the repaired image is used to generate the shooting image. In this way, during the camera preview stage, when a closed-eye defect is detected in the current frame image, recent images related to the target object (i.e., multiple consecutive frames including the current frame image) can be acquired. Based on these multiple consecutive frames, the open-eye state of the target object at the current moment can be predicted, rather than simply replacing the closed-eye area in the current frame image with a static open-eye image. Therefore, the predicted open-eye state parameters can match the target object's recent eye-opening movement habits. This results in a natural and vivid restoration effect for the restored image obtained based on these parameters, while maintaining high coherence with preceding and following frames. This application solves the problem in related technologies where the restoration effect obtained by replacing the closed-eye area in an image is stiff, unnatural, and lacks dynamic realism, thus achieving a technical improvement in the restoration effect of image closed-eye defect repair.
[0017] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a method for repairing eye closure defects according to an embodiment of this application; Figure 2 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application; Figure 3 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 1 ; Figure 4 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 2 ; Figure 5 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 3 ; Figure 6 This is a schematic diagram of the structure of an eye-closing defect repair device according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0021] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0022] With the widespread use of electronic devices, taking photos using electronic devices has become a daily routine. However, when taking portrait photos using electronic devices, if the subject blinks at the moment the shutter is pressed, the resulting photo may have a quality issue where the subject's eyes are closed.
[0023] To correct closed-eye defects in images, techniques typically select an image with open eyes (i.e., an image of the eyes in an open state) from a pre-built image library to replace the closed-eye area in the original image. However, this open-eye image may not accurately reflect the actual open-eye state of the subject. Therefore, the corrected image obtained using this method often exhibits a stiff and unnatural appearance in the corrected area, and the corrected image often shows low continuity with preceding and following frames in the shooting preview interface, indicating low dynamic realism.
[0024] There is no effective solution to the problem that the restoration effect of the image obtained by replacing the closed eye area in the image is stiff, unnatural and has low dynamic realism.
[0025] In view of this, embodiments of this application provide a method for repairing eye closure defects. The method can launch a camera application and display a shooting preview interface. If the current frame image in the preview stream buffer includes the eye closure defect of the target object, multiple consecutive images, including the current frame image, can be obtained from the preview stream buffer. This preview stream buffer stores preview images captured by the camera on the target object. Then, based on the multiple consecutive images, the open-eye state of the target object at the current moment corresponding to the current frame image can be predicted to obtain open-eye state parameters. Subsequently, based on the open-eye state parameters, the eye closure defect can be repaired on the current frame image to obtain a repaired image. Finally, the repaired image can be displayed on the shooting preview interface, and the repaired image is used to generate the shooting image. In this way, during the camera preview stage, when a closed-eye defect is detected in the current frame image, recent images related to the target object (i.e., multiple consecutive frames including the current frame image) can be acquired. Based on these multiple consecutive frames, the open-eye state of the target object at the current moment can be predicted, rather than simply replacing the closed-eye area in the current frame image with a static open-eye image. Therefore, the predicted open-eye state parameters can match the target object's recent eye-opening movement habits. This results in a natural and vivid restoration effect for the restored image obtained based on these parameters, while maintaining high coherence with preceding and following frames. This application solves the problem in related technologies where the restoration effect obtained by replacing the closed-eye area in an image is stiff, unnatural, and lacks dynamic realism, thus achieving a technical improvement in the restoration effect of image closed-eye defect repair.
[0026] The following will combine Figure 1 This application describes a method for repairing eye closure defects according to one embodiment.
[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, personal computer, mobile phone, etc., or a wearable device that can achieve the above functions.
[0028] Please see Figure 1 , Figure 1 This is a flowchart of a method for repairing eye closure defects according to an embodiment of this application, as follows: Figure 1 As shown, the method includes: Step S101: Preview the target object and store it in the preview stream buffer.
[0029] In this embodiment, the electronic device may be equipped with a camera application, and the electronic device may launch the camera application in response to the user's operation (e.g., touching the icon of the camera application, or issuing a voice command to launch the camera application).
[0030] Once the camera application is launched, a shooting preview interface can be displayed in the display area of the electronic device. This shooting preview interface is used to display the image captured by the camera of the electronic device (i.e., the preview stream image, which includes multiple preview images) during the preview stage before taking a photo.
[0031] The aforementioned shooting preview interface may also include multiple controls, such as a shooting control. After the user clicks the shooting control, the electronic device can generate and output the captured image based on the preview image at the moment the shooting control is clicked.
[0032] Step S102: If the current frame image includes the closed-eye defect of the target object, acquire multiple consecutive images, including the current frame image, from the preview stream buffer. The preview stream buffer is used to store preview images acquired by the camera of the target object.
[0033] In this embodiment, a trained lightweight classification model or directly based on the geometric features of the target object's face in the current frame image can be used to determine whether the current frame image includes the target object's closed-eye defect.
[0034] As an example, it is possible to determine whether a target object has closed eyes based on the geometric features of the eye region in the current frame image (e.g., the aspect ratio of the eyes); or, the current frame image can be analyzed in real time based on a binary classifier used to determine whether a person has closed eyes, which can output a probability of closed eyes (0 to 1), and when the probability of closed eyes is greater than a certain threshold (e.g., 0.7), the current frame image includes the target object's closed eye defect.
[0035] When using a trained lightweight classification model to determine whether the current frame image contains the closed-eye defect of the target object, the geometric features of the target object's eye region in the current frame image can also be combined for auxiliary judgment to improve the robustness of this detection step.
[0036] In this embodiment, the preview stream buffer is used to store the preview image captured by the camera on the target object. That is, the preview image stored in the preview stream buffer is used to display on the shooting preview interface so that the user can preview the current shooting scene.
[0037] In the case where the current frame image in the preview stream buffer includes the closed-eye defect of the target object, a short-sequence preview stream image including the current frame image can be extracted from the preview stream buffer, that is, a multi-frame continuous image including the current frame image. For example, the multi-frame continuous image includes the current frame image t and historical frame images t-1 to t-4.
[0038] Step S103: Based on multiple consecutive images, predict the open-eye state of the target object at the current moment corresponding to the current frame image to obtain the open-eye state parameters.
[0039] In this embodiment, the static appearance features of the target object's eyes (e.g., shape and iris color) can be encoded based on multiple consecutive images, as well as the dynamic features of the target object's eye-opening motion (e.g., eyelid motion information in multiple consecutive images). Based on the aforementioned static appearance features and dynamic features of the target object's eyes, the eye-opening state of the target object at the current moment corresponding to the current frame image can be predicted to obtain eye-opening state parameters.
[0040] Compared to simply replacing the closed-eye region in the current frame image with a static open-eye image, the eye-opening state parameters predicted by the eye-opening defect repair method provided in this application can match the recent eye-opening movement habits of the target object. This makes the repair effect of the subsequently repaired image obtained based on the eye-opening state parameters natural and vivid, while maintaining high continuity with the previous and next frame images.
[0041] Figure 2 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application, such as... Figure 2 As shown, in one embodiment, step S103 includes: Step S201: For each consecutive image except the current frame, the eye region of the target object in the consecutive images is transformed from the current coordinate system corresponding to the consecutive images and matched to the target coordinate system corresponding to the current frame image.
[0042] In this embodiment, optical flow algorithms or feature point-based matching algorithms can be used to perform precise temporal alignment of the facial and eye regions of the target object in multiple consecutive images to eliminate displacement caused by minute head movements during the shooting process of the target object. This can provide more accurate input information for subsequent steps.
[0043] As an example, the Farneback dense optical flow algorithm can be used to calculate the optical flow field between two adjacent frames. Using the current frame image t as a reference, the eye region of the target object in the historical frame images t-1 to t-4 can be aligned to the target coordinate system corresponding to the current frame image t by reverse warping.
[0044] Step S202: Spatiotemporal feature encoding is performed on multiple eye regions in the target coordinate system to obtain the fusion features of the target object's open-eye state.
[0045] In this embodiment, spatiotemporal feature encoding can be performed on the eye region of the target object in multiple consecutive images to obtain the eye-opening state fusion feature of the target object. Specifically, spatial and temporal features of the eye region of the target object in consecutive images in the target coordinate system can be extracted, and the above spatial and temporal features can be fused to obtain the eye-opening state fusion feature of the target object. The eye-opening state fusion feature can include not only the static appearance features of the target object's eyes (e.g., shape and iris color), but also the dynamic features of the target object's eye-opening motion.
[0046] Figure 3 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 1 ,like Figure 3 As shown, in one embodiment, step S202 includes: Step S301: Extract spatial features from the eye region of the target object in each frame of continuous images to obtain the spatial feature map corresponding to each frame of continuous images.
[0047] Step S302: The spatial feature maps corresponding to each consecutive image frame are concatenated into a tensor, and temporal features are extracted from the tensor to obtain a temporal feature map.
[0048] Step S303: Perform global average pooling on the temporal feature map according to the time dimension to obtain the fused features of the target object's open-eye state.
[0049] In this embodiment, a lightweight spatiotemporal convolutional network can be used to encode the spatiotemporal features of the eye region of a target object in a continuous image.
[0050] Spatiotemporal convolutional networks can include two-dimensional convolutional neural network modules, temporal convolutional layer modules, and global average pooling layer modules.
[0051] First, the eye region of the target object in each consecutive frame of images (e.g., an image of size 64×64 (pixels)) can be input into a two-dimensional convolutional neural network module (e.g., including two convolutional layers) to extract spatial features of the eye region based on the two-dimensional convolutional neural network module, thereby obtaining the spatial feature map (C (number of channels) × H (height) × W (width)) corresponding to each consecutive frame of images.
[0052] Then, the spatial feature maps corresponding to each consecutive image frame can be stacked along the temporal dimension to obtain a tensor (5×C×H×W). This tensor can then be input into a temporal convolutional layer module (e.g., including a one-dimensional temporal convolutional layer with a kernel size of 3) to convolve the tensor along the temporal dimension, capturing dynamic changes within the tensor and obtaining a temporal feature map. Alternatively, this tensor can be input into a recurrent neural network module to convolve the tensor along the temporal dimension, capturing dynamic changes within the tensor and obtaining a temporal feature map.
[0053] Finally, the temporal feature map can be input into the global average pooling layer module, and the temporal feature map can be processed by global average pooling along the time dimension based on the global average pooling layer module to obtain the fused features of the target object's open-eye state.
[0054] Step S203: Based on the eye-opening state fusion features, predict the eye-opening state of the target object at the current moment to obtain the eye-opening state parameters.
[0055] In this embodiment, the eye-opening state fusion features can be input into a multilayer perceptron to predict the eye-opening state of the target object at the current moment based on the aforementioned eye-opening state fusion features.
[0056] Videos of blinking of the target object in its natural state can be collected, and the eye-opening state parameters of each blink image in the video can be manually labeled or estimated using a high-precision model. Then, a mean squared error loss function can be used to supervise the training of a multilayer perceptron using the labeled blink images, so that the trained multilayer perceptron can map the eye-opening state fusion features to the eye-opening state parameters of the target object at the current moment.
[0057] In one embodiment, the eye-opening state parameters include static eye parameters and eye-opening motion parameters. The static eye parameters may include the target object's eyelid opening degree (0 indicates fully closed, 1 indicates fully open), the eyeball's positional offset within the eye socket, and eyelid curvature parameters, etc. The eye-opening motion parameters may include the rate of change of the target object's eyelid opening degree, etc.
[0058] The aforementioned static eye parameters and eye-opening motion parameters define the eye geometry and eye-opening motion information of the target object at the previous time step.
[0059] Step S104: Based on the open-eye state parameters, repair the closed-eye defect of the current frame image to obtain the repaired image.
[0060] Figure 4 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 2 ,like Figure 4 As shown, in one embodiment, step S104 includes: Step S401: Obtain the reference eye image of the target object.
[0061] In this embodiment, a facial feature image that is closest to the head pose and shooting environment of the target object contained in the current frame image can be retrieved from a pre-built personal facial feature library, and this facial feature image can be used as a reference eye image; or, the user can specify the reference eye image himself.
[0062] Step S402: Input the reference eye image and the open-eye state parameters into the trained image generation model so that the image generation model uses the reference eye image and the open-eye state parameters as conditions to generate the repaired eye region.
[0063] In this embodiment, the image generation model may include a pre-trained generative adversarial network module (e.g., the StyleGAN2 model). The image generation model may include a first encoding module, a second encoding module, and a generation module.
[0064] Style vectors can be used to control the image generation process of the StyleGAN2 model. For example, the style input space W+ can be used to control the image generation process of the StyleGAN2 model. This style input space W+ can include multiple style vectors. Based on the idea of semantically guided image inpainting, it is assumed that there are certain directions in the W+ space that are related to the current eye state of the target object (eyelid opening and closing, eyeball position offset, etc.). By adding offsets in these directions, the eye attributes of the generated result can be linearly manipulated without changing the identity characteristics of the target object.
[0065] Figure 5 This is a flowchart of a method for repairing eye closure defects according to another embodiment of this application. Figure 3 ,like Figure 5 As shown, in one embodiment, step S402 includes: Step S501: Input the reference eye image into the first encoding module to encode the reference eye image into a first style vector based on the first encoding module.
[0066] In this embodiment, the first encoding module may include a convolutional neural network and a mapping network. A reference eye image can be input into the first encoding module, and the convolutional neural network in the first encoding module can extract the identity feature vector of the target object from the reference eye image.
[0067] Then, the aforementioned identity feature vector can be input into a mapping network composed of several fully connected layers. This mapping network maps the identity feature vector into a first style vector in the style input space W+. This first style vector encodes, for example, fixed information such as the target object's facial identity, skin texture, and eye structure (e.g., single or double eyelids, eye shape, etc.), and serves as the basis for generating the repaired eye region.
[0068] Step S502: Input the eye-opening state parameters into the second encoding module to encode the eye-opening state parameters into a second style vector based on the second encoding module.
[0069] In this embodiment, the second encoding module may include a mapping network that can input the open-eye state parameters into the second encoding module. The mapping network, which consists of several fully connected layers in the second encoding module, maps the open-eye state parameters into a style offset, i.e., a second style vector. The second style vector encodes the style adjustment required from the reference eye state in the reference eye image to the target open-eye state (i.e., the state defined by the open-eye state parameters).
[0070] To ensure the stability and decoupling of the control, an amplitude constraint can be applied to the second style vector (e.g., L2 normalization followed by multiplication by a learnable scaling factor) to prevent the generated image from being distorted or identity drifted due to excessive style shift.
[0071] Step S503: The first style vector and the second style vector are added together and then input into the generation module so that the generation module generates the repaired eye area based on the first style vector and the second style vector.
[0072] In this embodiment, the first style vector and the second style vector can be added together to obtain the final conditional style vector. This conditional style vector can be input into the generation module, which then generates a state-controlled repaired eye region based on it. The repaired eye region inherits the identity features of the target object contained in the baseline eye image, and its eye geometry information (e.g., eyelid opening / closing, eyeball position offset, etc.) is controlled by eye-opening state parameters.
[0073] In this way, by decoupling the identity features of the target object from its open-eye state and injecting them into the style input space W+, continuous and precise control over the eye geometry information in the repaired eye region can be achieved, while ensuring that the generated result is highly consistent with the identity features of the target object.
[0074] Step S403: Based on the eye-opening motion parameters, add motion blur effect to the repaired eye area to obtain an enhanced eye area.
[0075] In this embodiment, the eyelid movement speed of the target object can be calculated based on eye-opening motion parameters, such as the rate of change of the target object's eyelid opening and closing degree. Then, based on this eyelid movement speed, a motion blur effect can be added to the repaired eye area to simulate eyelid movement during camera exposure.
[0076] As an example, when the eyelid movement speed exceeds a certain threshold, a Gaussian blur with a radius of 1 to 2 pixels can be applied to the eyelid edge region in the repaired eye area. The area of this Gaussian blur region can be positively correlated with the eyelid movement speed.
[0077] In this way, by making full use of the rich temporal information contained in the preview stream image, such as the micro-dynamics of eyelid movement speed and the inertia when the eyes are open, an enhanced eye region that matches the identity characteristics and blinking habits of the target object can be generated. This makes the restoration effect of the restored image obtained by fusing based on the enhanced eye region natural and vivid, while maintaining high coherence with the preceding and following frames.
[0078] Step S404: Fuse the enhanced eye region with the current frame image to obtain the repaired image.
[0079] In this embodiment, an improved Poisson fusion algorithm or a deep learning-based image fusion model (e.g., a generative adversarial network-based image fusion model) can be used to fuse the enhanced eye region with the current frame image. Taking the generative adversarial network-based image fusion model as an example, the discriminator module of this image fusion model can simultaneously receive the enhanced eye region and the surrounding skin region, enabling a smooth transition between the edge gradient of the enhanced eye region and the background of the current frame image, achieving seamless fusion.
[0080] In one embodiment, after step S104 above, the method further includes: displaying the repaired image on the shooting preview interface, wherein the repaired image is used to generate the shooting image.
[0081] In this embodiment, the repaired image can be directly displayed on the shooting preview interface to ensure that the image observed by the user operating the electronic device is always a preview image with the repaired closed eye defect, without having to wait for the perfect moment or repeatedly retake the shot. A high-quality shooting image can be obtained directly by pressing the shutter, which can improve the efficiency of portrait shooting.
[0082] In one embodiment, the method further includes: in response to a target operation, generating and outputting a captured image based on the repaired image.
[0083] In this embodiment, the user operating the electronic device can press the shutter button or touch the shooting control on the shooting preview interface the moment the repaired image is displayed on the shooting preview interface. The electronic device can respond to the above-mentioned target operation, generate and output the captured image based on the repaired image.
[0084] It should be noted that the information collection process involved in this application (e.g., image collection process, etc.) is carried out with the user's knowledge and permission, that is, the information collection process complies with the requirements of relevant standards and does not constitute an act that harms the public interest.
[0085] Through the above steps S101 to S104, firstly, the camera application can be launched and the shooting preview interface can be displayed. If the current frame image in the preview stream buffer includes the closed eye defect of the target object, multiple consecutive images including the current frame image can be obtained from the preview stream buffer. The preview stream buffer is used to store the preview images obtained by the camera shooting the target object. Then, based on the multiple consecutive images, the open eye state of the target object at the current moment corresponding to the current frame image can be predicted to obtain the open eye state parameters. Subsequently, based on the open eye state parameters, the closed eye defect of the current frame image can be repaired to obtain the repaired image. Finally, the repaired image can be displayed on the shooting preview interface, and the repaired image is used to generate the shooting image. In this way, during the camera preview stage, when a closed-eye defect is detected in the current frame image, recent images related to the target object (i.e., multiple consecutive frames including the current frame image) can be acquired. Based on these multiple consecutive frames, the open-eye state of the target object at the current moment can be predicted, rather than simply replacing the closed-eye area in the current frame image with a static open-eye image. Therefore, the predicted open-eye state parameters can match the target object's recent eye-opening movement habits. This results in a natural and vivid restoration effect for the restored image obtained based on these parameters, while maintaining high coherence with preceding and following frames. This application solves the problem in related technologies where the restoration effect obtained by replacing the closed-eye area in an image is stiff, unnatural, and lacks dynamic realism, thus achieving a technical improvement in the restoration effect of image closed-eye defect repair.
[0086] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution of the target. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0087] Corresponding to the eye closure defect repair method described in the above embodiments, Figure 6 A schematic diagram of a device for repairing eye closure defects according to an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.
[0088] Please see Figure 6 The eye-closed defect repair device 6 includes: a preview shooting module 60 for previewing the target object and storing it in a preview stream buffer; an acquisition module 61 for acquiring multiple consecutive images, including the current frame image, from the preview stream buffer when the current frame image includes the eye-closed defect of the target object; a prediction module 62 for predicting the open-eye state of the target object at the current moment corresponding to the current frame image based on the multiple consecutive images, and obtaining open-eye state parameters; and a repair module 63 for repairing the eye-closed defect of the current frame image based on the open-eye state parameters, and obtaining a repaired image.
[0089] In one embodiment, the prediction module 62 is further configured to perform spatiotemporal feature encoding on the eye region of the target object in multiple consecutive images to obtain the eye-opening state fusion feature of the target object; based on the eye-opening state fusion feature, to predict the eye-opening state of the target object at the current moment to obtain the eye-opening state parameter.
[0090] In one embodiment, the prediction module 62 is further configured to, for each consecutive image except the current frame image, transform and match the eye region of the target object in the consecutive images from the current coordinate system corresponding to the consecutive images to the target coordinate system corresponding to the current frame image; and perform spatiotemporal feature encoding on multiple eye regions in the target coordinate system to obtain the fusion feature of the target object's open-eye state.
[0091] In one embodiment, the prediction module 62 is further configured to extract spatial features from the eye region of the target object in each consecutive frame of images to obtain spatial feature maps corresponding to each consecutive frame of images; to concatenate the spatial feature maps corresponding to each consecutive frame of images into a tensor, and to extract temporal features from the tensor to obtain a temporal feature map; and to perform global average pooling on the temporal feature map according to the time dimension to obtain the fusion features of the target object's open-eye state.
[0092] In one embodiment, the eye-opening state parameters include static eye parameters and eye-opening motion parameters; the repair module 63 is further configured to acquire a reference eye image of the target object; input the reference eye image and the eye-opening state parameters into a trained image generation model, so that the image generation model uses the reference eye image and the eye-opening state parameters as conditions to generate a repaired eye region; based on the eye-opening motion parameters, add motion blur effect to the repaired eye region to obtain an enhanced eye region; fuse the enhanced eye region with the current frame image to obtain the repaired image.
[0093] In one embodiment, the image generation model includes a first encoding module, a second encoding module, and a generation module; the repair module 63 is further configured to input a reference eye image into the first encoding module to encode the reference eye image into a first style vector based on the first encoding module; input eye-opening state parameters into the second encoding module to encode the eye-opening state parameters into a second style vector based on the second encoding module; and add the first style vector and the second style vector together and input them into the generation module so that the generation module generates the repaired eye region based on the first style vector and the second style vector.
[0094] In one embodiment, the closed-eye defect repair device 6 further includes an imaging module for generating and outputting an image based on the repaired image in response to a target operation.
[0095] It should be noted that the aforementioned eye-closing defect repair device 6 can be or be applied to a computing service device with data processing, network communication and program running functions, such as a tablet computer, personal computer, mobile phone, server, etc., or a wearable device that can achieve the above functions.
[0096] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Figure 7 As shown, the electronic device 7 includes: at least one processor 70 ( Figure 7(Only one is shown in the image) a processor, a memory 71, and a computer program 72 stored in the memory 71 and executable on at least one processor 70, wherein the processor 70 executes the computer program 72 to implement the steps in any of the above embodiments of the method for repairing eye closure defects.
[0099] Electronic device 7 can be a computing service device with data processing, network communication and program execution functions, such as a tablet computer, personal computer, mobile phone, server, etc., or a wearable device that can achieve the above functions.
[0100] The electronic device 7 may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that... Figure 7 The example of electronic device 7 is merely an illustration and does not constitute a limitation on electronic device 7. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, environmental sensing units (e.g., inertial measurement units, cameras, audio sensors, etc.).
[0101] The processor 70 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0102] In some embodiments, memory 71 may be an internal storage unit of electronic device 7, such as a hard disk or memory of electronic device 7. In other embodiments, memory 71 may be an external storage device of electronic device 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on electronic device 7. In other embodiments, memory 71 may include both internal and external storage units of electronic device 7. Memory 71 is used to store operating system, application programs, bootloader, data, and other programs, such as the program code of computer program 72. Memory 71 may also be used to temporarily store data that has been output or will be output.
[0103] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above embodiments of the methods for repairing eye closure defects.
[0104] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to execute the steps described in the above-described methods for repairing eye closure defects.
[0105] This application implements all or part of the processes in the methods of the above embodiments, which can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a blindness repair device or electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.
[0106] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0107] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0108] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method of closed eye defect repair, the method comprising: Applied to electronic devices, the method includes: Take a preview shot of the target object and store it in the preview stream buffer; If the current frame image includes the closed-eye defect of the target object, multiple consecutive frames of images, including the current frame image, are obtained from the preview stream buffer; Based on the multiple consecutive images, the eye-opening state of the target object at the current moment corresponding to the current frame image is predicted to obtain the eye-opening state parameters. Based on the eye-opening state parameters, the closed-eye defect of the current frame image is repaired to obtain the repaired image.
2. The method of claim 1, wherein, Based on the multiple consecutive images, the eye-opening state of the target object at the current moment corresponding to the current frame image is predicted, and the eye-opening state parameters are obtained, including: Spatiotemporal feature encoding is performed on the eye region of the target object in the multi-frame continuous images to obtain the eye-opening state fusion feature of the target object; Based on the eye-opening state fusion features, the eye-opening state of the target object at the current time is predicted to obtain the eye-opening state parameters.
3. The method of claim 2, wherein, The step of performing spatiotemporal feature encoding on the eye region of the target object in the multi-frame continuous images to obtain the fused features of the target object's open-eye state includes: For each consecutive image frame other than the current frame image, the eye region of the target object in the consecutive images is transformed from the current coordinate system corresponding to the consecutive images and matched to the target coordinate system corresponding to the current frame image; Spatiotemporal feature encoding is performed on multiple eye regions in the target coordinate system to obtain the fusion feature of the open-eye state of the target object.
4. The method of claim 2, wherein, The step of performing spatiotemporal feature encoding on the eye region of the target object in the multi-frame continuous images to obtain the fused features of the target object's open-eye state includes: Spatial features are extracted from the eye region of the target object in each frame of the continuous image to obtain the spatial feature map corresponding to each frame of the continuous image; The spatial feature maps corresponding to each frame of the continuous images are concatenated into a tensor, and temporal features are extracted from the tensor to obtain a temporal feature map; The temporal feature map is subjected to global average pooling according to the time dimension to obtain the eye-opening state fusion feature of the target object.
5. The method according to any one of claims 1 to 4, characterized in that, The eye-opening state parameters include static eye parameters and eye-opening movement parameters; The process of repairing the closed-eye defect in the current frame image based on the open-eye state parameters, resulting in a repaired image, includes: Obtain a baseline eye image of the target object; The reference eye image and the eye-opening state parameters are input into the trained image generation model so that the image generation model uses the reference eye image and the eye-opening state parameters as conditions to generate the repaired eye region. Based on the eye-opening motion parameters, a motion blur effect is added to the repaired eye area to obtain an enhanced eye area; The enhanced eye region is fused with the current frame image to obtain the repaired image.
6. The method of claim 5, wherein, The image generation model includes a first encoding module, a second encoding module, and a generation module; The step of inputting the reference eye image and the eye-opening state parameters into a trained image generation model, so that the image generation model uses the reference eye image and the eye-opening state parameters as conditions to generate the repaired eye region, includes: The reference eye image is input into the first encoding module to encode the reference eye image into a first style vector based on the first encoding module; The eye-opening state parameters are input into the second encoding module to encode the eye-opening state parameters into a second style vector based on the second encoding module; The first style vector and the second style vector are added together and then input into the generation module so that the generation module generates the repaired eye area based on the first style vector and the second style vector.
7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: In response to the target operation, a captured image is generated and output based on the repaired image.
8. A closed eye defect repair device, characterized by, Applied to electronic devices, the device includes: The preview shooting module is used to preview the target object and store it in the preview stream buffer; The acquisition module is used to acquire multiple consecutive images, including the current frame image, from the preview stream buffer when the current frame image includes the closed eye defect of the target object; The prediction module is used to predict the open-eye state of the target object at the current moment corresponding to the current frame image based on the multi-frame continuous images, and obtain the open-eye state parameters. The repair module is used to repair the closed-eye defect of the current frame image based on the open-eye state parameters to obtain the repaired image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for repairing eye closure defects as described in any one of claims 1 to 7.
10. A computer program product, characterised in that, Includes a computer program, which, when run, causes the method for repairing eye closure defects as described in any one of claims 1 to 7 to be performed.