Image processing method and device, model training method and device, electronic equipment and medium

By automatically detecting and adjusting the eye status in the image processing interface and combining multiple visual features for repair, the problem of abnormal eye images in the existing technology is solved, and a natural and realistic eye repair effect and a user-friendly operation experience are achieved.

CN120725923AActive Publication Date: 2025-09-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510733770.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-30
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing technologies in image processing find it difficult to effectively repair eye image abnormalities caused by blinking, light interference, etc., especially problems such as closed eyes, inconsistent gaze direction, and dull eyes, and the user operation chain is lengthy.

Method used

By automatically detecting the eye status in the image display interface, displaying adjustment controls, extracting visual features and making eye adjustments based on these features, and combining information such as eyebrow texture, facial light and shadow distribution, and eye closure degree, automated and personalized eye repair is achieved.

Benefits of technology

It improves the user experience, achieves the naturalness and authenticity of eye repair, lowers the operation threshold, ensures the natural transition of light and shadow in the eyes, and avoids blurred edges of clothing and destruction of the outline of head clothing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725923A_ABST
    Figure CN120725923A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, a model training method and device, electronic equipment and a medium, and relates to the field of image processing, in particular to the technical fields of computer vision, deep learning and the like. According to the specific implementation scheme, eye state detection is carried out on a target image displayed in an image display interface; displaying an eye adjustment control on an image display interface in response to a target object of which the eye state does not meet a set condition in the target image; in response to the detected interaction operation for the eye adjustment control, visual features of the target object are extracted, and the visual features at least comprise facial local features; and based on the visual features, performing eye adjustment on the target object to obtain an adjusted target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, specifically to technical fields such as computer vision and deep learning, and in particular to an image processing method, model training method, device, electronic device and medium. Background Art

[0002] During image capture, unexpected factors like blinking and light interference often lead to eye anomalies in photos, such as closed eyes, inconsistent gaze, and dull gaze. While advances in image processing technology have enabled general facial restoration, such as blemish removal and expression adjustment, the effectiveness of eye restoration still needs to be improved. Summary of the Invention

[0003] The present disclosure provides an image processing method, a model training method, an apparatus, an electronic device, and a medium.

[0004] According to one aspect of the present disclosure, there is provided an image processing method, comprising:

[0005] Performing eye state detection on the target image displayed in the image display interface;

[0006] In response to the target image including a target object whose eye state does not meet a set condition, displaying an eye adjustment control on the image display interface;

[0007] In response to detecting an interactive operation on the eye adjustment control, extracting visual features of the target object, the visual features including at least local facial features;

[0008] Based on the visual features, the eyes of the target object are adjusted to obtain an adjusted target image.

[0009] According to another aspect of the present disclosure, a model training method is provided, comprising:

[0010] Obtaining a model to be trained and sample images, wherein the sample images include sample objects whose eye conditions do not meet set conditions, and the sample images correspond to eye annotation data;

[0011] Inputting the sample image into the model to be trained to extract visual features of the sample object, and performing eye adjustments on the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features;

[0012] The model to be trained is trained based on the eye data of the sample object in the adjusted sample image and the eye annotation data.

[0013] According to another aspect of the present disclosure, there is provided an image processing apparatus, comprising:

[0014] A detection module is used to detect the eye state of the target image displayed in the image display interface;

[0015] a control display module, configured to display eye adjustment controls on the image display interface in response to the target image including a target object whose eye state does not meet a set condition;

[0016] a feature extraction module, configured to extract visual features of the target object in response to detecting an interactive operation on the eye adjustment control, the visual features including at least local facial features;

[0017] An adjustment module is used to perform eye adjustment on the target object based on the visual features to obtain an adjusted target image.

[0018] According to another aspect of the present disclosure, there is provided a model training device, comprising:

[0019] an acquisition module, configured to acquire a model to be trained and sample images, wherein the sample images include sample objects whose eye conditions do not meet set conditions, and the sample images correspond to eye annotation data;

[0020] An input module, configured to input the sample image into the model to be trained to extract visual features of the sample object, and perform eye adjustments on the sample object based on the visual features, thereby obtaining an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features;

[0021] A training module is used to train the model to be trained based on the eye data of the sample object in the adjusted sample image and the eye annotation data.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0023] at least one processor; and

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method proposed in the above-mentioned first aspect of the present disclosure or execute the model training method proposed in the above-mentioned other aspect of the present disclosure.

[0026] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the image processing method proposed in the above-mentioned first aspect of the present disclosure or the model training method proposed in the above-mentioned other aspect of the present disclosure.

[0027] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the image processing method proposed in the above-mentioned first aspect of the present disclosure or implements the model training method proposed in the above-mentioned other aspect of the present disclosure.

[0028] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0030] Figure 1 is a flowchart of an image processing method provided according to an embodiment of the present disclosure;

[0031] Figure 2 is a flowchart of an image processing method provided according to another embodiment of the present disclosure;

[0032] Figure 3 is a flowchart of an image processing method provided according to another embodiment of the present disclosure;

[0033] Figure 4 is a flowchart of a model training method provided according to another embodiment of the present disclosure;

[0034] Figure 5 is a flowchart of a model training method provided according to another embodiment of the present disclosure;

[0035] Figure 6 A schematic structural diagram of an image processing device according to an embodiment of the present disclosure;

[0036] Figure 7 A schematic structural diagram of a model training device according to an embodiment of the present disclosure;

[0037] Figure 8 It is a block diagram of an electronic device used to implement the image processing method or model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0039] In the related art, the repair function is triggered actively by the user, such as manually selecting the "eye-closed repair" tool to perform eye-closed repair, and the user operation chain is lengthy.

[0040] Therefore, to address at least one of the above-mentioned problems, the present disclosure proposes an image processing method, model training method, device, electronic device, and medium. The present disclosure can be applied to network storage, image processing tools, and the like.

[0041] The following describes the image processing method, model training method, device, electronic device and medium of the embodiments of the present disclosure with reference to the accompanying drawings.

[0042] Figure 1 It is a flowchart of an image processing method provided according to an embodiment of the present disclosure.

[0043] The embodiment of the present disclosure is described by taking the image processing method configured in an image processing device as an example. The image processing device can be applied to any electronic device to enable the electronic device to perform a text processing function.

[0044] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a mobile phone, tablet computer, personal digital assistant, wearable device, etc., which are hardware devices with various operating systems, touch screens and / or display screens.

[0045] like Figure 1 As shown, the image processing method may include the following steps S101 to S104:

[0046] Step S101 : performing eye state detection on a target image displayed in an image display interface.

[0047] Among them, the image display interface includes an image preview interface, an image editing interface and other interfaces for image display. The image display interface may include interactive controls such as image cropping controls, image deletion controls and image sharing controls; the target image refers to at least one image displayed in the image display interface. For example, the objects in the target image include people and / or animals; the eye state detection includes at least one of eye closure detection, inconsistent line of sight direction detection and dull eyes detection.

[0048] As an example, eye state detection can be performed on a target image using a trained eye state detection model.

[0049] When the eye state detection is eye closure detection, as an example, the EAR (Eye Aspect Ratio) corresponding to each object in the target image can be obtained; and the eye closure detection is performed on the target image based on the EAR.

[0050] It should be noted that face detection may be performed on the target image first to obtain face detection frames, and then eye state detection may be performed based on each face detection frame.

[0051] Step S102 : in response to the target image including a target object whose eye status does not meet the set condition, displaying an eye adjustment control on the image display interface.

[0052] Among them, the target object may refer to at least one object in the target image with closed eyes, inconsistent line of sight or dull eyes; the eye adjustment control is used to trigger the eye adjustment function, and the properties of the eye adjustment control such as size, color and display position can be pre-set. For example, the eye adjustment control may refer to an eye adjustment button.

[0053] Therefore, when it is detected that the target image includes a target object whose eye status does not meet the set conditions, the eye adjustment controls are displayed on the image display interface, which can transform the "user looking for function" mode into the "function reaching user" mode, lower the operation threshold and improve the user experience.

[0054] Step S103 : In response to detecting an interactive operation on the eye adjustment control, extracting visual features of the target object, where the visual features include at least local facial features.

[0055] Among them, interactive operations may include trigger operations such as click operations, double-click operations, and touch operations; visual features may include local facial features, hair color, skin color, etc.; local facial features may include eye features, nose features, and facial feature position features, etc.

[0056] As an example, visual features of a target object may be extracted by an encoder.

[0057] Step S104: performing eye adjustment on the target object based on the visual features to obtain an adjusted target image.

[0058] As an example, based on the correspondence between visual features and open-eye images, an open-eye image library can be searched for an open-eye image that matches the visual features of the target object; based on the queried open-eye image, the eyes of the target object are adjusted to obtain an adjusted target image.

[0059] As another example, the visual features may be input into a trained decoder to obtain an adjusted target image output by the decoder; wherein the decoder is used to indicate a mapping relationship between the visual features and the adjusted target image.

[0060] As a result, the decoder can accurately map the adjusted target image based on visual features and efficiently complete eye adjustment.

[0061] The image processing method of the embodiment of the present disclosure performs eye state detection on the target image displayed in the image display interface, and displays eye adjustment controls on the image display interface in response to the target image including a target object whose eye state does not meet the set conditions. This can transform the "user looking for function" mode into the "function reaching user" mode, thereby lowering the operation threshold; based on the visual characteristics of the target object, the eye of the target object is adjusted, which can realize automatic and personalized repair and adjustment of the eyes, thereby improving the user experience.

[0062] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0063] Visual features include eyebrow features, facial light and shadow distribution features, and eye closure degree features. In order to clearly illustrate how the aforementioned features are extracted in any embodiment of the present disclosure, the present disclosure also proposes an image processing method.

[0064] Figure 2 It is a flowchart of an image processing method provided according to another embodiment of the present disclosure.

[0065] like Figure 2 As shown, the image processing method may include the following steps S201 to S206:

[0066] Step S201 : performing eye state detection on a target image displayed in an image display interface.

[0067] Step S202 : in response to the target image including a target object whose eye status does not meet the set condition, displaying an eye adjustment control on the image display interface.

[0068] For explanations of step S201 and step S202, reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details are given here.

[0069] Step S203, in response to detecting an interactive operation on the eye adjustment control, obtaining a first area image corresponding to the eyebrow area of ​​the target object based on the target image; obtaining eyebrow texture features of the target object based on the first area image; and obtaining eyebrow features based on the eyebrow texture features.

[0070] Among them, the first area image can be obtained based on the eyebrow detection frame; the eyebrow texture feature is used to indicate the direction, thickness, density, etc. of the eyebrow texture.

[0071] In any embodiment of the present disclosure, grayscale conversion is performed on the first region image to obtain a corresponding first grayscale image; based on the first grayscale image, a grayscale co-occurrence matrix is ​​obtained; and based on the grayscale co-occurrence matrix, eyebrow texture features are determined.

[0072] Among them, statistics such as contrast, energy, and entropy can be calculated based on the gray-level co-occurrence matrix, and the eyebrow texture features can be determined based on the statistics.

[0073] The grayscale co-occurrence matrix can reflect the spatial distribution of eyebrow grayscale and accurately reflect its texture characteristics. Therefore, obtaining eyebrow texture characteristics based on the grayscale co-occurrence matrix helps to fit the original eyebrow texture when making eye adjustments, so that the adjusted and repaired eyes and eyebrows are naturally connected, avoiding a sense of stiffness and improving the authenticity and naturalness of eye adjustment and repair.

[0074] In any one embodiment of the present disclosure, the first area image is grayscale converted to obtain a corresponding first grayscale image; based on the set local binary pattern LBP operator, the first LBP value corresponding to any pixel in the first grayscale image is obtained; based on the first LBP value corresponding to any pixel, the eyebrow texture feature is obtained.

[0075] The set LBP operator may include an n*n LBP operator and a circular LBP operator, where n is a positive integer.

[0076] As an example, an LBP feature map can be obtained based on the first LBP value corresponding to any pixel in the first region image; and the LBP feature map is used as the eyebrow texture feature.

[0077] As another example, the number of occurrences of the first LBP values ​​may be counted; and the eyebrow texture feature may be obtained based on the number of occurrences of each first LBP value.

[0078] The LBP operator takes a neighborhood centered on a pixel, compares the grayscale values ​​of the neighborhood pixels with those of the central pixel, and generates a binary code. It can perceive subtle differences in grayscale between pixels, and thus accurately capture the grayscale and texture changes of the eyebrows. Therefore, the eyebrow texture features are determined based on the first LBP value, which helps to achieve a natural connection between the eye and eyebrow textures, making the adjusted target image more realistic and natural.

[0079] As a possible implementation method, eyebrow texture features can be used as eyebrow features.

[0080] As another possible implementation manner, image features are extracted from the first region image to obtain image features; and eyebrow features are obtained based on the image features and eyebrow texture features.

[0081] Among them, image features can be extracted from the first region image through an image feature extraction network; the image features and eyebrow texture features can be spliced ​​to obtain eyebrow features.

[0082] Image features can reflect global information such as the overall shape and color distribution of the eyebrows, while eyebrow texture features can reflect detailed information such as the local texture direction and thickness of the eyebrows. Therefore, combining image features and eyebrow texture features for eye adjustment can make the adjusted eyes and eyebrows more compatible in shape and texture, and the repair effect more realistic.

[0083] Step S204: Based on the target image, a second region image corresponding to the facial region of the target object is obtained; grayscale conversion is performed on the second region image to obtain a second grayscale image; based on the second grayscale image, a grayscale histogram is obtained to obtain facial light and shadow distribution characteristics.

[0084] The facial region of the target object may refer to the entire facial region of the target object, or may refer to a partial facial region including the eye region of the target object.

[0085] Among them, the grayscale histogram can reflect the distribution of pixels with different grayscale values ​​in the image, and thus indirectly reflect the overall distribution of light and shadow.

[0086] There may be shadows around the eyes, such as the shadow cast by a hat brim on the face. The facial light and shadow distribution characteristics can reflect the light and shadow distribution in the area around the eyes. Therefore, adjusting the eyes based on the facial light and shadow distribution characteristics can avoid inconsistencies between the eye light and shadow and the light and shadow in the area around the eyes, that is, abnormal eye brightness after adjustment, and ensure a natural transition between the eye light and shadow.

[0087] Step S205 : acquiring eye key points of the target object based on the target image; determining the eyelid closure rate based on the eye key points, and determining the eye closure degree feature based on the eyelid closure rate.

[0088] Among them, the eye key points may include the inner canthus key points, the outer canthus key points, and the key points of the upper and lower eyelids.

[0089] As an example, the eyelid closure rates of the two eyes of the target object may be obtained; and the eyelid closure rates of the two eyes of the target object may be used as eye closure degree features.

[0090] Adjusting the eyes based on the degree of eye closure can accurately locate the state of the eyes, and then make targeted adjustments to the eyelid shape, skin texture, etc., making the eye adjustment effect more realistic and natural.

[0091] It should be noted that the order of steps 203 to 205 is not used to limit the order of obtaining each feature in the visual feature, that is, the present disclosure does not limit the order of obtaining eyebrow features, facial light and shadow distribution features, and eye closure degree features.

[0092] In any of the embodiments of the present disclosure, in response to detecting an interactive operation on an eye adjustment control, an eye adjustment interface is displayed, and the eye adjustment progress is displayed on the eye adjustment interface.

[0093] The eye adjustment interface does not include the interactive controls displayed in the image display interface, that is, in response to detecting an interactive operation on the eye adjustment control, the user cannot interact with the target image.

[0094] For example, the eye adjustment progress can be displayed on the eye adjustment interface through a toast prompt.

[0095] It should be noted that, in response to detecting an interactive operation on the eye adjustment control, the eye adjustment interface can be displayed and the visual features can be extracted simultaneously.

[0096] Therefore, in response to detecting an interactive operation on the eye adjustment control, displaying the eye adjustment interface can quickly start the eye adjustment process; displaying the adjustment progress can allow the user to understand the progress of the eye adjustment more intuitively.

[0097] In any one embodiment of the present disclosure, in response to detecting a closing operation on the eye adjustment interface, the eye adjustment interface is exited, and the target image before eye adjustment is displayed on the image display interface; or, in response to detecting a saving operation on the adjusted target image displayed in the eye adjustment interface, the adjusted target image is stored.

[0098] For example, the eye adjustment interface includes a page control control, a save control, and an exit save control. The page control control is in an open state during the display of the eye adjustment interface.

[0099] The closing operation on the eye adjustment interface may refer to a closing operation on a page control control, or an interactive operation on an exit and save control.

[0100] As a result, the user can quickly return from the eye adjustment interface to the target image displayed on the image display interface, or quickly and conveniently save the adjusted target image in the eye adjustment interface, meeting the different needs of users and thus improving the user experience.

[0101] Step S206 : performing eye adjustment on the target object based on the visual features to obtain an adjusted target image.

[0102] For explanation of step S206, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.

[0103] The image processing method of the disclosed embodiment combines eyebrow texture features to obtain eyebrow features, which helps to achieve a natural connection between the eye and eyebrow texture, making the adjusted target image more realistic and natural; combining facial light and shadow distribution features to adjust the eyes can avoid the inconsistency between the eye light and shadow and the light and shadow of the area around the eyes, ensuring a natural transition of the eye light and shadow; combining eye closure degree features to adjust the eyes can accurately locate the eye state, and then make targeted adjustments to the eyelid shape, skin texture, etc., making the eye adjustment effect more realistic and natural.

[0104] Visual features also include clothing collar texture features and head clothing geometric features. In order to clearly illustrate how the aforementioned features are obtained in any embodiment of the present disclosure, the present disclosure also proposes an image processing method.

[0105] Figure 3 It is a flowchart of an image processing method provided according to another embodiment of the present disclosure.

[0106] like Figure 3 As shown, the image processing method may include the following steps S301 to S305:

[0107] Step S301 : performing eye state detection on a target image displayed in an image display interface.

[0108] Step S302 : in response to the target image including a target object whose eye status does not meet the set condition, displaying an eye adjustment control on the image display interface.

[0109] For explanations of step S301 and step S302, reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.

[0110] Step S303, in response to detecting an interactive operation on the eye adjustment control, based on the target image, obtaining a third area image corresponding to the collar area of ​​the target object's clothing; obtaining a second LBP value corresponding to any pixel in the third area image; and determining the texture feature of the collar of the clothing based on the number of occurrences of any second LBP value.

[0111] The third region image may include a collar region of clothing, and may also include a local facial region of the target object, such as a chin region. For example, the collar region of clothing may include a collar region of a bachelor's gown.

[0112] The LBP operator for obtaining the second LBP value and the LBP operator for obtaining the first LBP value may be the same or different.

[0113] Adjusting the eyes based on the texture features of the neckline area of ​​the clothing can not only achieve matching of the adjusted eyes with the overall texture style of the clothing, but also preserve the edges of the clothing, avoiding blurred edges of the clothing after eye adjustment.

[0114] In any embodiment of the present disclosure, the clothing collar texture feature includes at least one of the following: an average value of the number of occurrences of the second LBP value; a variance of the number of occurrences of the second LBP value; and an energy of the number of occurrences of the second LBP value.

[0115] The average value reflects the overall distribution of the texture, the variance reflects the degree of distribution dispersion, and the energy shows the complexity of the texture. Combining these values ​​can fully reflect the texture characteristics of the clothing collar.

[0116] Step S304 , performing edge detection on the headwear of the target object, and obtaining contour information of the headwear based on the edge detection result; and obtaining geometric features of the headwear based on the contour information.

[0117] Among them, headwear includes clothing / decorations worn on the head such as headscarves and hats.

[0118] The geometric characteristics of headwear can reflect the outline and shape of the headwear, and provide reference information for eye adjustment. Therefore, adjusting the eyes based on the geometric characteristics of the headwear can make the adjusted eyes and headwear more visually coordinated and natural.

[0119] In any embodiment of the present disclosure, the headwear includes a hat, and the contour information includes at least one of the following: the brim length of the brim; the curvature of the brim; the brim length of the crown; and the curvature of the crown. Combining the aforementioned data can more accurately characterize the contour of the hat.

[0120] Illustratively, the hat may include a graduate cap, a peaked cap, and the like.

[0121] Based on the geometric features of the hat, the relationship between the headwear and the face can be accurately located. Therefore, when adjusting the eyes based on the geometric features of the hat, on the one hand, it can avoid destroying the outline of the headwear, and on the other hand, it can make the adjusted eyes and headwear more harmonious and natural as a whole.

[0122] It should be noted that the order of step 303 to step 304 is not used to limit the order of obtaining each feature in the visual feature, that is, the present disclosure does not limit the order of obtaining clothing collar texture features and head clothing geometric features.

[0123] Step S305 : performing eye adjustment on the target object based on the visual features to obtain an adjusted target image.

[0124] For explanation of step S305, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.

[0125] The image processing method of the disclosed embodiment combines the texture features of the neckline area of ​​clothing to perform eye adjustment, which not only achieves the matching of the adjusted eyes with the overall texture style of the clothing, but also preserves the edges of the clothing, avoiding the blurring of the edges of the clothing after the eye adjustment; the geometric features of the head clothing can reflect the outline and shape of the head clothing, providing reference information for the eye adjustment. Combining the geometric features of the head clothing to perform eye adjustment can make the adjusted eyes and head clothing more visually coordinated and natural, thereby improving the quality of the adjusted image.

[0126] This disclosure also proposes a model training method, Figure 4 It is a flowchart of a model training method provided according to another embodiment of the present disclosure.

[0127] The embodiment of the present disclosure takes the model training method as an example in which it is configured in a model training device. The model training device can be applied to any electronic device so that the electronic device can perform the model training function.

[0128] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a mobile phone, tablet computer, personal digital assistant, wearable device, etc., which are hardware devices with various operating systems, touch screens and / or display screens.

[0129] like Figure 4 As shown, the model training method may include the following steps S401 to S403:

[0130] Step S401 : obtaining a model to be trained and sample images, wherein the sample images include sample objects whose eye states do not meet set conditions, and the sample images correspond to eye annotation data.

[0131] Among them, the model to be trained may include an encoder to be trained and a decoder to be trained; the number of sample images is at least one, and the sample images can be images of various scenes and various lighting conditions; the number of sample objects whose eye conditions do not meet the set conditions included in the sample images is at least one; and the eye annotation data may include a reference eye image whose eye conditions corresponding to the sample objects meet the set conditions.

[0132] For example, the scenes may include graduation season scenes, travel scenes, daily life scenes, artistic photo scenes, etc.; the lighting conditions may include outdoor strong light, indoor lighting, etc.

[0133] Step S402: Input the sample image into the model to be trained to extract the visual features of the sample object, and perform eye adjustments on the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features.

[0134] The encoder to be trained is used to extract visual features of the sample object; the decoder to be trained is used to perform eye adjustment on the sample object based on the visual features and output the adjusted sample image.

[0135] The encoder to be trained includes a first branch and a second branch. When the visual features include eyebrow features, facial light and shadow distribution features, eye closure degree features, clothing collar texture features, and head clothing geometric features, the first branch can be used to extract eyebrow features and eye closure degree features, and the second branch can be used to extract facial light and shadow distribution features, clothing collar texture features, and head clothing geometric features.

[0136] Step S403 : training the to-be-trained model based on the eye data and eye annotation data of the sample object in the adjusted sample image.

[0137] The eye data may include the eye image of the sample object in the adjusted sample image.

[0138] As an example, a first loss may be calculated based on the difference between the eye data of the sample object and the eye annotation data in the adjusted sample image; and the model to be trained may be trained based on the first loss.

[0139] Exemplarily, the first loss may be calculated based on an image difference between an eye image of the sample object in the adjusted sample image and a reference eye image.

[0140] The model training method of the embodiment of the present disclosure inputs a sample image into the model to be trained to extract the visual features of the sample object, and adjusts the eyes of the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features; the model to be trained is trained based on the eye data and eye annotation data of the sample object in the adjusted sample image; the model to be trained can learn the extraction of visual features and the relationship between visual features and eye adjustment, thereby realizing automated and personalized repair and adjustment of the eyes and improving user experience.

[0141] The sample image also corresponds to clothing annotation data and eye occlusion annotation data. In order to clearly explain how any embodiment of the present disclosure trains the training model based on the eye data and eye annotation data of the sample object in the adjusted sample image, the present disclosure also proposes a model training method.

[0142] Figure 5 It is a flowchart of a model training method provided according to another embodiment of the present disclosure.

[0143] like Figure 5 As shown, the model training method may include the following steps S501 to S504:

[0144] Step S501 : determining a first loss based on a difference between eye data and eye annotation data of a sample object in an adjusted sample image.

[0145] In any embodiment of the present disclosure, the eye data includes eye key point data, and correspondingly, the eye annotation data includes eye key point annotation data.

[0146] Exemplarily, the eye key point data includes the key points of the inner canthus, the outer canthus, and the key points of the upper and lower eyelids.

[0147] Eye key point data can accurately locate the eye structure and morphology. Therefore, model training based on eye key point data and eye key point annotation data can enable the model to deeply learn the laws of eye features, making it more closely fit the real eye shape during repair and adjustment, avoiding repair distortion, making the repaired eyes natural and vivid, and improving the repair effect and overall image quality.

[0148] It should be noted that the eye data may also include data used to assist in eye adjustment, such as eyebrow key point data.

[0149] Step S502 : determining a second loss based on a difference between the clothing data of the sample object and the clothing annotation data in the adjusted sample image.

[0150] In any one of the embodiments of the present disclosure, the wearing clothing data includes at least one of clothing neckline contour data, hat brim edge data, and shoulder line position data; correspondingly, the wearing clothing annotation data includes at least one of clothing neckline contour annotation data, hat brim edge annotation data, and shoulder line position annotation data.

[0151] As an example, edge detection may be performed on the adjusted sample image, and based on the edge detection result, at least one of clothing collar contour data, hat brim edge data, and shoulder line position data may be obtained.

[0152] Combining clothing collar outline data, hat brim edge data, shoulder line position data and corresponding annotation data for model training can not only enable the model to consider more relevant information when making eye adjustments, but also avoid problems such as dislocation and blurring of clothing edges (such as clothing collar outline, hat brim edge, shoulder line).

[0153] Step S503 : determining a third loss based on the difference between the eye occlusion data and the eye occlusion annotation data of the sample object in the adjusted sample image.

[0154] The eye occlusion data may be used to indicate whether there is an occluder on the eye, and may also be used to indicate information such as the position, shape, and outline of the occluder.

[0155] In any embodiment of the present disclosure, the eye occlusion data includes hair occlusion data and / or glasses occlusion data; correspondingly, the eye occlusion annotation data includes hair occlusion annotation data and / or glasses occlusion annotation data.

[0156] Exemplarily, the hair occlusion data includes data such as the outline and color of the hair that occludes the eyes; the glasses occlusion data includes data such as the outline, color, and position of the glasses.

[0157] Combining hair occlusion data, glasses occlusion data and corresponding annotation data for model training allows the model to fully learn the eye features under occlusion, avoiding problems such as hairstyle distortion and glasses disappearance after eye adjustment.

[0158] Step S504: training the model to be trained based on at least one of the first loss, the second loss, and the third loss.

[0159] As an example, the model to be trained can be trained based on any one or any two of the first loss, the second loss, and the third loss.

[0160] As another example, the first loss, the second loss, and the third loss may be weightedly summed to obtain a target loss; and the model to be trained may be trained based on the target loss.

[0161] The model training method of the embodiment of the present disclosure determines a first loss based on the difference between the eye data and the eye annotation data of the sample object in the adjusted sample image, which can accurately measure the eye repair adjustment error, guide the model to optimize the eye repair effect, and improve the repair accuracy; determines a second loss based on the difference between the wearing clothing data and the wearing clothing annotation data of the sample object in the adjusted sample image, which can enable the model to learn clothing-related features, assist in eye repair, and enhance the coordination between eye repair and overall performance; determines a third loss based on the difference between the eye occlusion data and the eye occlusion annotation data of the sample object in the adjusted sample image, which can enable the model to adapt to eye occlusion conditions, accurately repair the eyes when occlusion exists, and improve the repair robustness; trains the model to be trained based on at least one of the first loss, the second loss, and the third loss, which can comprehensively optimize the model based on multiple factors and improve the eye repair quality of the model.

[0162] To clearly illustrate how this disclosure trains the model and how to use the model for eye adjustment, the following uses closed-eye restoration for graduation season images as an example:

[0163] Model training process:

[0164] (1) Obtain a training dataset containing graduation season portraits. The sample images in the training dataset cover different lighting conditions (strong outdoor light, indoor ceremonial light), different bachelor's degree clothing styles (including different colors such as black / red / blue, high collar / low collar, different fabrics such as drapey fabrics, etc.), graduation cap wearing angles (straight, slightly tilted), hairstyles (short hair for boys, loose hair for girls, bun hair), etc. Among them, the sample images correspond to eye key point annotation data, clothing neckline outline annotation data, hat brim edge annotation data, shoulder line position annotation data, hair occlusion annotation data, and glasses occlusion annotation data.

[0165] (2) Inputting sample images in the training data set into the model to be trained; wherein the model to be trained includes an encoder and a decoder, and the encoder includes a first branch and a second branch.

[0166] (3) The first branch is used to extract the eyebrow features and eye closure degree features of the sample object in the sample image, and the second branch is used to extract the facial light and shadow distribution features, clothing collar texture features, and head clothing geometric features of the sample object in the sample image.

[0167] (4) Fusing the features extracted by the first branch and the second branch to obtain a fused feature; inputting the fused feature into the decoder for eye adjustment to obtain an adjusted sample image output by the decoder.

[0168] (5) Obtain the eye key point data, clothing collar outline data, hat brim edge data, shoulder line position data, hair occlusion data, and glasses occlusion data in the adjusted sample image, and train the training model based on the difference between the aforementioned data and the corresponding labeled data to obtain a trained target model.

[0169] Model application process:

[0170] When an interactive operation on the eye adjustment control displayed in the image display interface is detected, the target image displayed in the image display interface is input into the target model, so as to extract the eyebrow features and eye closure degree features of the target object in the target image through the first branch of the encoder in the target model; the facial light and shadow distribution features, clothing collar texture features and head clothing geometric features of the target object are extracted through the second branch of the encoder; the fusion features of the features extracted by the first branch and the second branch are decoded by the decoder to adjust the eyes of the target object, and finally the adjusted target image is obtained.

[0171] This disclosure optimizes the hairstyle, clothing (such as bachelor's gown collars, graduation caps) and other features of graduation season portraits. It can ensure that the light and shadow of the eyes after closed-eye restoration are consistent with the shadow of the graduation cap, and the eye skin tone matches the bachelor's gown collar area, avoiding problems such as blurred clothing edges, abnormal eye brightness caused by conflicts between light and shadow and hat structure, and distortion of hairstyle texture.

[0172] With the above Figures 1 to 3 Corresponding to the image processing method provided in the embodiment, the present disclosure also provides an image processing device. Since the image processing device provided in the embodiment of the present disclosure is consistent with the above-mentioned Figures 1 to 3 The image processing method provided in the embodiment corresponds to the above, so the implementation of the image processing method is also applicable to the image processing device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.

[0173] Figure 6 3 is a structural diagram of an image processing device provided according to an embodiment of the present disclosure.

[0174] like Figure 6 As shown, the image processing device 600 may include: a detection module 610, a control display module 620, a feature extraction module 630 and an adjustment module 640.

[0175] The detection module 610 is used to detect the eye state of the target image displayed in the image display interface;

[0176] A control display module 620 is configured to display eye adjustment controls on the image display interface in response to a target image including a target object whose eye status does not meet a set condition;

[0177] a feature extraction module 630 for extracting visual features of a target object in response to detecting an interactive operation on an eye adjustment control, the visual features including at least local facial features;

[0178] The adjustment module 640 is configured to perform eye adjustment on the target object based on the visual features to obtain an adjusted target image.

[0179] In a possible implementation of the embodiment of the present disclosure, the visual feature includes an eyebrow feature, and the feature extraction module 630 is specifically configured to:

[0180] Based on the target image, obtaining a first region image corresponding to the eyebrow region of the target object;

[0181] Based on the first region image, acquiring eyebrow texture features of the target object;

[0182] Based on the eyebrow texture features, the eyebrow features are obtained.

[0183] In a possible implementation of the embodiment of the present disclosure, the feature extraction module 630 is specifically configured to:

[0184] Performing image feature extraction on the first region image to obtain image features;

[0185] Based on the image features and eyebrow texture features, the eyebrow features are obtained.

[0186] In a possible implementation of the embodiment of the present disclosure, the feature extraction module 630 is specifically configured to:

[0187] Performing grayscale conversion on the first region image to obtain a corresponding first grayscale image;

[0188] Based on the first grayscale image, obtaining a gray level co-occurrence matrix;

[0189] Based on the gray-level co-occurrence matrix, the eyebrow texture features are determined.

[0190] In a possible implementation of the embodiment of the present disclosure, the feature extraction module 630 is specifically configured to:

[0191] Performing grayscale conversion on the first region image to obtain a corresponding first grayscale image;

[0192] Based on the set local binary pattern (LBP) operator, obtaining a first LBP value corresponding to any pixel in the first grayscale image;

[0193] Based on the first LBP value corresponding to any pixel, the eyebrow texture feature is obtained.

[0194] In a possible implementation of the embodiment of the present disclosure, the visual features include facial light and shadow distribution features, and the feature extraction module 630 is specifically configured to:

[0195] Based on the target image, obtaining a second region image corresponding to the facial region of the target object;

[0196] Performing grayscale conversion on the second region image to obtain a second grayscale image;

[0197] Based on the second grayscale image, a grayscale histogram is obtained to obtain facial light and shadow distribution characteristics.

[0198] In a possible implementation of the embodiment of the present disclosure, the visual feature includes an eye closure degree feature, and the feature extraction module 630 is specifically configured to:

[0199] Based on the target image, obtain the eye key points of the target object;

[0200] Based on the eye key points, the eyelid closure rate is determined to determine the eye closure degree feature based on the eyelid closure rate.

[0201] In a possible implementation of the embodiment of the present disclosure, the visual feature includes a texture feature of a clothing collar, and the feature extraction module 630 is specifically configured to:

[0202] Based on the target image, obtaining a third region image corresponding to a collar region of clothing of the target object;

[0203] Obtaining a second LBP value corresponding to any pixel in the third region image;

[0204] The clothing collar texture feature is determined based on the number of occurrences of any second LBP value.

[0205] In a possible implementation of the embodiment of the present disclosure, the clothing collar texture feature includes at least one of the following:

[0206] The average number of occurrences of the second LBP value;

[0207] The variance of the number of occurrences of the second LBP value;

[0208] The energy of the number of occurrences of the second LBP value.

[0209] In a possible implementation of the embodiment of the present disclosure, the visual features include geometric features of head clothing, and the feature extraction module 630 is specifically configured to:

[0210] Performing edge detection on the headwear of the target object, and obtaining contour information of the headwear based on the edge detection result;

[0211] Based on the contour information, the geometric features of the head clothing are obtained.

[0212] In a possible implementation of the embodiment of the present disclosure, the headwear includes a hat, and the outline information includes at least one of the following:

[0213] The length of the brim of the hat;

[0214] the curvature of the brim;

[0215] The length of the brim of the hat;

[0216] The curvature of the crown.

[0217] In a possible implementation of the embodiment of the present disclosure, the adjustment module 640 is specifically configured to:

[0218] The visual features are input into a decoder to obtain an adjusted target image output by the decoder; wherein the decoder is used to indicate a mapping relationship between the visual features and the adjusted target image.

[0219] In a possible implementation of the embodiment of the present disclosure, the device further includes an interface display module, which is configured to:

[0220] In response to detecting an interactive operation on the eye adjustment control, an eye adjustment interface is displayed, and the eye adjustment progress is displayed on the eye adjustment interface.

[0221] In a possible implementation of the embodiment of the present disclosure, the apparatus further includes an operating module, configured to:

[0222] In response to detecting a closing operation on the eye adjustment interface, exiting the eye adjustment interface, and displaying the target image before the eye adjustment is performed on the image display interface;

[0223] Alternatively, in response to detecting a save operation on the adjusted target image displayed in the eye adjustment interface, the adjusted target image is stored.

[0224] The image processing device of the disclosed embodiment performs eye state detection on a target image displayed in an image display interface, and displays eye adjustment controls in the image display interface in response to a target object included in the target image whose eye state does not meet set conditions. This can transform the "user looking for function" mode into the "function reaching user" mode, thereby lowering the operation threshold; based on the visual features of the target object, the eye of the target object is adjusted, which can realize automatic and personalized repair and adjustment of the eyes, thereby improving the user experience.

[0225] With the above Figures 4 and 5 Corresponding to the model training method provided in the embodiment, the present disclosure also provides a model training device. Figures 1 to 5The model training method provided in the embodiment corresponds to the embodiment, so the implementation method of the model training method is also applicable to the model training device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.

[0226] Figure 7 Schematic diagram of the structure of a model training device provided according to an embodiment of the present disclosure.

[0227] like Figure 7 As shown, the model training device 700 may include: an acquisition module 710, an input module 720 and a training module 730.

[0228] The acquisition module 710 is used to acquire a model to be trained and sample images, wherein the sample images include sample objects whose eye conditions do not meet the set conditions, and the sample images correspond to eye annotation data;

[0229] An input module 720 is configured to input a sample image into the model to be trained to extract visual features of the sample object, and perform eye adjustments on the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features;

[0230] The training module 730 is configured to train the to-be-trained model based on the eye data and eye annotation data of the sample object in the adjusted sample image.

[0231] In a possible implementation of the embodiment of the present disclosure, the sample image further corresponds to clothing annotation data and eye occlusion annotation data, and the training module 730 is specifically configured to:

[0232] determining a first loss based on a difference between the eye data of the sample object and the eye annotation data in the adjusted sample image;

[0233] determining a second loss based on a difference between the clothing data of the sample object and the labeled clothing data in the adjusted sample image;

[0234] determining a third loss based on a difference between the eye occlusion data of the sample object in the adjusted sample image and the eye occlusion annotation data;

[0235] The model to be trained is trained based on at least one of the first loss, the second loss, and the third loss.

[0236] In a possible implementation of the embodiment of the present disclosure, the eye data includes eye key point data;

[0237] The clothing data includes at least one of clothing collar outline data, hat brim edge data, and shoulder line position data;

[0238] The eye occlusion data includes hair occlusion data and / or glasses occlusion data.

[0239] The model training device of the embodiment of the present disclosure inputs a sample image into the model to be trained to extract the visual features of the sample object, and adjusts the eyes of the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features; the model to be trained is trained based on the eye data and eye annotation data of the sample object in the adjusted sample image; the model to be trained can learn the extraction of visual features and the relationship between visual features and eye adjustment, thereby realizing automated and personalized repair and adjustment of the eyes and improving user experience.

[0240] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0241] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0242] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 into a RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0243] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0244] The computing unit 801 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the image processing method and / or the model training method. For example, in some embodiments, the image processing method and / or the model training method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image processing method and / or the model training method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the image processing method and / or the model training method in any other appropriate manner (e.g., by means of firmware).

[0245] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0246] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0247] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0248] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0249] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0250] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0251] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0252] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0253] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An image processing method, comprising: Performing eye state detection on the target image displayed in the image display interface; In response to the target image including a target object whose eye state does not meet a set condition, displaying an eye adjustment control on the image display interface; In response to detecting an interactive operation on the eye adjustment control, extracting visual features of the target object, the visual features including at least local facial features; Based on the visual features, the eyes of the target object are adjusted to obtain an adjusted target image.

2. The method according to claim 1, wherein The visual feature includes an eyebrow feature, and extracting the visual feature of the target object includes: Based on the target image, obtaining a first region image corresponding to the eyebrow region of the target object; acquiring eyebrow texture features of the target object based on the first region image; The eyebrow feature is obtained based on the eyebrow texture feature.

3. The method according to claim 2, wherein: The step of obtaining the eyebrow feature based on the eyebrow texture feature includes: performing image feature extraction on the first region image to obtain image features; The eyebrow feature is obtained based on the image feature and the eyebrow texture feature.

4. The method according to claim 2, wherein: The acquiring the eyebrow texture feature of the target object based on the first region image includes: Performing grayscale conversion on the first region image to obtain a corresponding first grayscale image; Based on the first grayscale image, obtaining a gray level co-occurrence matrix; The eyebrow texture feature is determined based on the gray level co-occurrence matrix.

5. The method according to claim 2, wherein: The acquiring the eyebrow texture feature of the target object based on the first region image includes: Performing grayscale conversion on the first region image to obtain a corresponding first grayscale image; Obtaining a first LBP value corresponding to any pixel in the first grayscale image based on a set local binary pattern (LBP) operator; The eyebrow texture feature is obtained based on the first LBP value corresponding to any pixel.

6. The method according to claim 1, wherein The visual features include facial light and shadow distribution features, and extracting the visual features of the target object includes: Based on the target image, acquiring a second region image corresponding to the facial region of the target object; performing grayscale conversion on the second region image to obtain a second grayscale image; Based on the second grayscale image, a grayscale histogram is obtained to obtain the facial light and shadow distribution characteristics.

7. The method according to claim 1, wherein The visual feature includes an eye closure degree feature, and extracting the visual feature of the target object includes: Based on the target image, acquiring eye key points of the target object; Based on the eye key points, an eyelid closure rate is determined to determine the eye closure degree feature based on the eyelid closure rate.

8. The method according to claim 1, wherein The visual features include clothing collar texture features, and extracting the visual features of the target object includes: Based on the target image, acquiring a third region image corresponding to a collar region of clothing of the target object; Obtaining a second LBP value corresponding to any pixel in the third region image; The clothing collar texture feature is determined based on the number of occurrences of any one of the second LBP values.

9. The method according to claim 8, wherein The clothing collar texture features include at least one of the following: an average of the number of occurrences of the second LBP value; the variance of the number of occurrences of the second LBP value; The energy of the number of occurrences of the second LBP value.

10. The method according to claim 1, wherein The visual features include geometric features of head clothing, and extracting the visual features of the target object includes: Performing edge detection on the headwear of the target object, and acquiring contour information of the headwear based on the edge detection result; Based on the contour information, the geometric features of the headwear are obtained.

11. The method according to claim 10, wherein: The headwear includes a hat, and the outline information includes at least one of the following: The length of the brim of the hat; the curvature of the brim; The length of the brim of the hat; The curvature of the crown.

12. The method according to any one of claims 1 to 11, wherein: The step of performing eye adjustment on the target object based on the visual features to obtain an adjusted target image includes: The visual features are input into a decoder to obtain the adjusted target image output by the decoder; wherein the decoder is used to indicate a mapping relationship between the visual features and the adjusted target image.

13. The method according to any one of claims 1 to 11, wherein: The method further comprises: In response to detecting an interactive operation on the eye adjustment control, an eye adjustment interface is displayed, and the eye adjustment progress is displayed on the eye adjustment interface.

14. The method according to claim 13, wherein The method further comprises: In response to detecting a closing operation on the eye adjustment interface, exiting the eye adjustment interface, and displaying the target image before the eye adjustment on the image display interface; Alternatively, in response to detecting a save operation on the adjusted target image displayed in the eye adjustment interface, the adjusted target image is stored.

15. A model training method comprising: Obtaining a model to be trained and sample images, wherein the sample images include sample objects whose eye conditions do not meet set conditions, and the sample images correspond to eye annotation data; Inputting the sample image into the model to be trained to extract visual features of the sample object, and performing eye adjustments on the sample object based on the visual features to obtain an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features; The model to be trained is trained based on the eye data of the sample object in the adjusted sample image and the eye annotation data.

16. The method according to claim 15, wherein The sample image also corresponds to clothing annotation data and eye occlusion annotation data, and the training of the to-be-trained model based on the eye data of the sample object in the adjusted sample image and the eye annotation data includes: determining a first loss based on a difference between the eye data of the sample object in the adjusted sample image and the eye annotation data; determining a second loss based on a difference between the clothing data of the sample object in the adjusted sample image and the clothing annotation data; determining a third loss based on a difference between the eye occlusion data of the sample object in the adjusted sample image and the eye occlusion annotation data; The model to be trained is trained based on at least one of the first loss, the second loss, and the third loss.

17. The method according to claim 16, wherein The method comprises at least one of the following: The eye data includes eye key point data; The clothing data includes at least one of clothing neckline outline data, hat brim edge data, and shoulder line position data; The eye occlusion data includes hair occlusion data and / or glasses occlusion data.

18. An image processing apparatus, comprising: A detection module is used to detect the eye state of the target image displayed in the image display interface; a control display module, configured to display eye adjustment controls on the image display interface in response to the target image including a target object whose eye state does not meet a set condition; a feature extraction module, configured to extract visual features of the target object in response to detecting an interactive operation on the eye adjustment control, the visual features including at least local facial features; An adjustment module is used to perform eye adjustment on the target object based on the visual features to obtain an adjusted target image.

19. A model training device comprising: an acquisition module, configured to acquire a model to be trained and sample images, wherein the sample images include sample objects whose eye conditions do not meet set conditions, and the sample images correspond to eye annotation data; An input module, configured to input the sample image into the model to be trained to extract visual features of the sample object, and perform eye adjustments on the sample object based on the visual features, thereby obtaining an adjusted sample image output by the model to be trained; wherein the visual features include at least local facial features; A training module is used to train the model to be trained based on the eye data of the sample object in the adjusted sample image and the eye annotation data.

20. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14 or 15 to 17.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method of any one of claims 1-14 or 15-17.

22. A computer program product comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 14 or 15 to 17.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and readable storage medium

    CN110544317A

  • Image processing method and device, equipment and storage medium

    CN113139486A

  • Face image processing method and device, equipment and computer readable storage medium

    CN113642364A

  • Eye drive control method and device, storage medium and electronic equipment

    CN113946221A

  • Shading plate, control system and control method for relieving eye fatigue of driver

    CN118665125A