Image screening method and device, electronic equipment and storage medium
By automatically obtaining and filtering the images in the character album of the target characters, the problem of cumbersome and error-prone virtual image generation process in the prior art is solved, and efficient and accurate virtual image generation is achieved.
Patent Information
- Application Number
- CN202311595989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, when generating a user's virtual image, the user needs to manually select the training image, which is cumbersome and prone to errors, affecting the generation effect of the virtual image.
An image filtering method is provided, in response to the virtual image generation request of the target character, obtain the character album of the target character, and perform image filtering processing on the images in the album, and automatically obtain multiple character images of the target character without manual selection by the user.
It improves image acquisition efficiency, reduces error occurrence, improves the generation effect of virtual images, and improves the user experience.
Smart Images

Figure CN120045734A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image screening method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology and the continuous improvement of users' entertainment needs, a function of generating virtual images based on character images has emerged and is widely used in various scenarios.
[0003] When generating a user's virtual image, multiple images of the user need to be used for training to obtain a virtual image that is unique to the user. In the related art, the user needs to manually select training images, which is cumbersome on the one hand and prone to errors on the other hand, affecting the generation effect of the virtual image. Summary of the invention
[0004] In order to overcome the problems existing in the related art, the present disclosure provides an image screening method, device, electronic device and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image screening method, comprising:
[0006] In response to a request for generating a virtual image of a target person, obtaining a photo album of the target person; the photo album includes at least one image of the target person;
[0007] Perform image screening processing on at least one character image in the character album to obtain a target character image; the target character image is used to generate a virtual image of the target character.
[0008] Optionally, the method further comprises:
[0009] Determining whether the number of the target person images obtained by screening is less than a first number threshold;
[0010] If the number of the target person images obtained by the screening is less than the first number threshold, obtaining at least one person image manually uploaded by a user;
[0011] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0012] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0013] Optionally, the method further comprises:
[0014] In response to the virtual image generation request, detecting whether a photo album of the target person is stored in the electronic device;
[0015] If the electronic device does not store a photo album of the target person, obtaining at least one person image manually uploaded by a user;
[0016] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0017] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0018] Optionally, the image screening process includes at least one of the following:
[0019] Performing angle detection on the at least one person image, and determining a person image in which the face of the target person is in a frontal posture from the at least one person image;
[0020] Performing image quality detection on the at least one person image, and determining a person image whose image quality meets a preset condition from the at least one person image;
[0021] Performing occlusion detection on the at least one person image, and determining a person image in which the face of the target person is not occluded from the at least one person image;
[0022] Performing human body detection on the at least one human image, and determining a human body image that only includes a human body from the at least one human image.
[0023] Optionally, performing angle detection on the at least one person image to determine a person image in which the face of the target person is in a frontal posture from the at least one person image includes:
[0024] Positioning key points of the at least one person image to obtain coordinate information of multiple contour key points of facial features in a face region of the person image;
[0025] Determining three-dimensional posture angle information of a face region in the character image based on a coordinate difference between the coordinate information of the plurality of contour key points and the coordinate information of corresponding reference points;
[0026] According to the three-dimensional posture angle information, a person image with a face in a frontal posture is determined.
[0027] Optionally, the performing image quality detection on the at least one person image to determine a person image whose image quality meets a preset condition from the at least one person image comprises at least one of the following:
[0028] Performing blur detection on the at least one person image to obtain blur parameters of the person image; and determining, based on the blur parameters, a person image whose blur parameters satisfy a first condition from the at least one person image;
[0029] Traversing the character image by using a sliding window to obtain a plurality of window images corresponding to the character image; determining a first ratio between the number of window images whose average brightness is greater than a first brightness threshold and the number of all window images in the plurality of window images; determining a character image whose first ratio satisfies a second condition based on the first ratios of the plurality of character images;
[0030] Use a sliding window to traverse the character image to obtain multiple window images corresponding to the character image; determine a second ratio between the number of window images whose brightness mean is less than a second brightness threshold and the number of all window images in the multiple window images; and determine a character image whose second ratio meets a third condition based on the second ratios of the multiple character images.
[0031] Optionally, performing occlusion detection on the at least one person image to determine a person image in which the face of the target person is not occluded from the at least one person image includes:
[0032] Using the Ghost module in the occlusion detection model, convolution processing is performed on the person image to obtain a plurality of first feature maps corresponding to the person image; and based on the plurality of first feature maps, a plurality of second feature maps corresponding to each of the first feature maps are generated; the second feature maps are similar to the corresponding first feature maps;
[0033] Based on the first feature map and the second feature map corresponding to the person image, the person image is classified to obtain a classification result; the classification result indicates the probability that the face in the person image is occluded; the occlusion detection model is trained based on the modified cross entropy loss function;
[0034] Based on the classification result, a person image in which the face of the target person is not blocked is determined from the at least one person image.
[0035] Optionally, before performing the image screening process, the method further includes:
[0036] Performing face detection on at least one person image to determine the number of face regions in the person image;
[0037] When the character image only contains one face area, determine the third ratio between the area of the face area and the image area of the character image; determine the character image to be subjected to image screening processing based on the third ratio corresponding to the at least one character image; wherein the third ratio of the character image to be subjected to image screening processing is greater than a third threshold value.
[0038] According to a second aspect of an embodiment of the present disclosure, there is provided an image screening device, comprising:
[0039] A first acquisition module, configured to acquire a photo album of the target person in response to a request for generating a virtual image of the target person; the photo album includes at least one image of the target person;
[0040] The processing module is used to perform image screening processing on at least one character image in the character album to obtain a target character image; the target character image is used to generate a virtual image of the target character.
[0041] Optionally, the device further includes: a second acquisition module, configured to:
[0042] Determining whether the number of the target person images obtained by screening is less than a first number threshold;
[0043] If the number of the target person images obtained by the screening is less than the first number threshold, obtaining at least one person image manually uploaded by a user;
[0044] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0045] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0046] Optionally, the device further includes: a second acquisition module, configured to:
[0047] In response to the virtual image generation request, detecting whether a photo album of the target person is stored in the electronic device;
[0048] If the electronic device does not store a photo album of the target person, obtaining at least one person image manually uploaded by a user;
[0049] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0050] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0051] Optionally, the processing module is used to perform at least one of the following:
[0052] Performing angle detection on the at least one person image, and determining a person image in which the face of the target person is in a frontal posture from the at least one person image;
[0053] Performing image quality detection on the at least one person image, and determining a person image whose image quality meets a preset condition from the at least one person image;
[0054] Performing occlusion detection on the at least one person image, and determining a person image in which the face of the target person is not occluded from the at least one person image;
[0055] Performing human body detection on the at least one human image, and determining a human body image that only includes a human body from the at least one human image.
[0056] Optionally, the processing module is used to:
[0057] Positioning key points of the at least one person image to obtain coordinate information of multiple contour key points of facial features in a face region of the person image;
[0058] Determining three-dimensional posture angle information of a face region in the character image based on a coordinate difference between the coordinate information of the plurality of contour key points and the coordinate information of corresponding reference points;
[0059] According to the three-dimensional posture angle information, a person image with a face in a frontal posture is determined.
[0060] Optionally, the processing module is used to perform at least one of the following:
[0061] Performing blur detection on the at least one person image to obtain blur parameters of the person image; and determining, based on the blur parameters, a person image whose blur parameters satisfy a first condition from the at least one person image;
[0062] Traversing the character image by using a sliding window to obtain a plurality of window images corresponding to the character image; determining a first ratio between the number of window images whose average brightness is greater than a first brightness threshold and the number of all window images in the plurality of window images; determining a character image whose first ratio satisfies a second condition based on the first ratios of the plurality of character images;
[0063] Use a sliding window to traverse the character image to obtain multiple window images corresponding to the character image; determine a second ratio between the number of window images whose brightness mean is less than a second brightness threshold and the number of all window images in the multiple window images; and determine a character image whose second ratio meets a third condition based on the second ratios of the multiple character images.
[0064] Optionally, the processing module is used to:
[0065] Using the Ghost module in the occlusion detection model, convolution processing is performed on the person image to obtain a plurality of first feature maps corresponding to the person image; and based on the plurality of first feature maps, a plurality of second feature maps corresponding to each of the first feature maps are generated; the second feature maps are similar to the corresponding first feature maps;
[0066] Based on the first feature map and the second feature map corresponding to the person image, the person image is classified to obtain a classification result; the classification result indicates the probability that the face in the person image is occluded; the occlusion detection model is trained based on the modified cross entropy loss function;
[0067] Based on the classification result, a person image in which the face of the target person is not blocked is determined from the at least one person image.
[0068] Optionally, the processing module is further used to:
[0069] Performing face detection on at least one person image to determine the number of face regions in the person image;
[0070] When the character image only contains one face area, determine the third ratio between the area of the face area and the image area of the character image; determine the character image to be subjected to image screening processing based on the third ratio corresponding to the at least one character image; wherein the third ratio of the character image to be subjected to image screening processing is greater than a third threshold value.
[0071] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0072] a memory for storing processor-executable instructions;
[0073] A processor connected to the memory;
[0074] Wherein, the processor is configured to execute the image screening method as described in any embodiment of the first aspect of the present disclosure.
[0075] According to the fourth aspect of the embodiments of the present disclosure, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can perform the image screening method as described in any embodiment of the first aspect of the present disclosure.
[0076] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:
[0077] In response to the received virtual image generation request of the target person, the disclosed embodiment directly obtains the character album of the target person from the electronic device to obtain multiple character images of the target person, without the need for manual selection by the user, which can effectively improve the efficiency of image acquisition and is not prone to errors. In addition, the character images in the character album of the target person are subjected to image screening processing to obtain the target person image that meets the requirements, so that the target person image can be used in the subsequent process to generate a virtual image that is more in line with the actual image of the target person, effectively improving the generation effect of the virtual image of the target person; no user operation is required during the entire processing process, thereby improving the user experience.
[0078] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0080] Figure 1 is a schematic diagram of a process of an image screening method according to an exemplary embodiment Figure 1 ;
[0081] Figure 2 is a schematic diagram showing the positions of key points of the contour in a face image according to an exemplary embodiment;
[0082] Figure 3 is a schematic diagram of a process of an image screening method according to an exemplary embodiment Figure 2 ;
[0083] Figure 4 is a flow chart of a method for automatically screening human images according to an exemplary embodiment;
[0084] Figure 5 is a schematic structural diagram of an image screening device according to an exemplary embodiment;
[0085] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0086] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.
[0087] The present disclosure provides an image screening method, such as Figure 1 As shown, Figure 1 is a schematic diagram of a process of an image screening method according to an exemplary embodiment Figure 1 The image screening method comprises:
[0088] Step S101, in response to a request for generating a virtual image of a target person, obtaining a photo album of the target person; the photo album includes at least one image of the target person;
[0089] Step S102, performing image screening processing on at least one person image in the person album to obtain a target person image; the target person image is used to generate a virtual image of the target person.
[0090] The image screening method shown in the embodiment of the present disclosure can be applied to any electronic device, which can be: a smart phone, a tablet computer or a wearable electronic device, etc. Alternatively, the image screening method can also be applied to a server, and the server can respond to a virtual image generation request sent by the electronic device, and screen the character images by obtaining the character images collected or stored by the electronic device.
[0091] In step S101, in response to a received virtual image generation request, a target person corresponding to the virtual image generation request may be determined, and a person album of the target person may be obtained from the electronic device.
[0092] Here, the avatar generation request may be generated based on an avatar generation operation triggered by a user.
[0093] The user-triggered virtual image generation operation may be a user clicking on a virtual image generation request area in an application interface currently displayed on the electronic device, or may be triggering virtual image generation through voice recognition, or may be other operations that can trigger virtual image generation, which is not limited in the embodiments of the present disclosure.
[0094] The virtual image generation request may carry the character information of the target person, so as to determine the target person according to the character information carried in the virtual information generation request; and obtain a character album matching the character information of the target person from the electronic device.
[0095] Here, the character information of the target person may be information used to identify the target person; in some embodiments, the character information of the target person may be a facial image of the target person, or facial feature information of the target person.
[0096] It is understandable that if the character information of the target person carried in the virtual image generation request is a facial image of the target person, facial feature extraction can be performed based on the facial image to obtain the facial feature information of the target person.
[0097] A person album may include one or more person images; and the multiple person images in a person album are all images of the same person.
[0098] It should be noted that some electronic devices can perform face detection on multiple stored images to filter out images containing faces; and perform face recognition on images containing faces to classify the images containing faces based on the face recognition results to obtain multiple character albums; images in different character albums are face images of different characters; and images in the same character album are face images of the same character.
[0099] In the embodiment of the present disclosure, the character information of the target person may be matched with the character information corresponding to multiple character albums stored in the electronic device to determine the character album corresponding to the target person.
[0100] It is worth noting that in the process of generating the virtual image of the target person, multiple character images of the target person are needed as training images to train the virtual image of the target person. In the related art, the user is usually required to manually select multiple character images of the target person. In the case where a large number of images are stored in the electronic device, the user is required to check them one by one to determine the multiple character images. The whole process is time-consuming and labor-intensive. It is also easy to accidentally click on images of other non-target persons during the image viewing process, affecting the training results. The disclosed embodiment can obtain multiple character images of the target person by obtaining the character album corresponding to the target person, without the need for manual selection by the user, and the image acquisition process is very efficient and not prone to errors.
[0101] In step S102, after obtaining the photo album of the target person, image screening processing may be performed on at least one person image in the photo album to obtain the target person image; so that the target person image can be used as a training image to generate a virtual image of the target person.
[0102] It should be noted that the target person image can be a person image in the target person's photo album that meets the image screening requirements. The target person image is used as a training image and input into the virtual image generation model to generate a virtual image of the target person.
[0103] It is worth noting that the virtual image generation model has certain requirements for the input training images. If the image quality of the input training images is not high, it will affect the virtual image generation effect of the virtual image generation model.
[0104] Here, the image screening process may be to screen the image size of the person image to screen out the target person image that meets the size requirement; and / or, it may be to screen the image brightness of the person image to screen out the target person image that meets the brightness requirement; and / or, it may be to screen the clarity of the face area in the person image to screen out the target person image whose face area meets the clarity requirement, etc. The embodiments of the present disclosure are not limited to this.
[0105] The disclosed embodiment performs image screening processing on the character images in the character album of the target character to obtain the target character images that meet the requirements, so as to use the target character images to generate a virtual image that is more consistent with the actual image of the target character.
[0106] In response to the received virtual image generation request of the target person, the disclosed embodiment directly obtains the character album of the target person from the electronic device to obtain multiple character images of the target person, without the need for manual selection by the user, which can effectively improve the efficiency of image acquisition and is not prone to errors. In addition, the character images in the character album of the target person are subjected to image screening processing to obtain the target person image that meets the requirements, so that the target person image can be used in the subsequent process to generate a virtual image that is more in line with the actual image of the target person, effectively improving the generation effect of the virtual image of the target person; no user operation is required during the entire processing process, thereby improving the user experience.
[0107] Optionally, the method further comprises:
[0108] Determining whether the number of the target person images obtained by screening is less than a first number threshold;
[0109] If the number of the target person images obtained by the screening is less than the first number threshold, obtaining at least one person image manually uploaded by a user;
[0110] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0111] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0112] In the disclosed embodiment, the first quantity threshold may be the minimum number of training images required to be input into the avatar generation model. The first quantity threshold may be set according to a specific avatar generation model. For example, the first quantity threshold may be 20, which means that at least 20 images of the target person need to be input into the avatar generation model to generate the avatar of the target person.
[0113] After performing image screening based on the target person's photo album, it can be determined whether the number of the target person images screened is less than a first quantity threshold; that is, it can be determined whether the number of the target person images screened reaches the minimum number of training images required to be input into the virtual image generation model.
[0114] If the number of the target person images obtained through screening is greater than or equal to the first number threshold, a virtual image of the target object may be generated based on the target person images obtained through screening.
[0115] If the number of the target person images obtained through screening is less than the first number threshold, at least one person image manually uploaded by the user is obtained.
[0116] It is worth noting that when the number of images of the target person automatically screened based on the target person's photo album is insufficient, the user may manually upload images.
[0117] In some embodiments, if the number of the screened target person images is less than a first number threshold, a prompt message is output, where the prompt message is used to prompt the user to upload the training image.
[0118] After obtaining the person images manually uploaded by the user, person images containing the target person may be screened out from the person images manually uploaded by the user.
[0119] It should be noted that face recognition can be performed on the character images manually uploaded by the user; based on the face recognition results, character images containing the target person are screened out from the character images manually uploaded by the user; so as to reduce the situation where the character images manually uploaded by the user for virtual image training and the character images screened from the character album do not belong to the same person.
[0120] After obtaining character images including the target character from character images manually uploaded by the user, image screening processing may be performed on the obtained character images including the target character.
[0121] It is worth noting that the screening threshold of the image screening process performed on the person images manually uploaded by the user is smaller than the screening threshold of the image screening process performed on the person images in the person album.
[0122] It should be noted that since the person images manually uploaded by users are person images that have been filtered by users, and the person images manually uploaded by users may be images that are automatically filtered out during the automatic screening process, if you continue to filter according to the screening threshold (or screening requirements) of automatic filtering, you may not be able to obtain person images that meet the requirements from the person images manually uploaded by users.
[0123] In some embodiments, two sets of screening thresholds may be pre-stored. The first set of screening thresholds may be used when performing image screening on person images in a person album to obtain target person images; the second set of screening thresholds may be used when performing image screening on person images manually uploaded by users.
[0124] It can be understood that the target person image obtained after the image screening process and the target person image screened based on the person's album are used together as training images and input into the virtual image generation model to generate a virtual image of the target person.
[0125] It is worth noting that when obtaining training images of the virtual image of the target person, the embodiment of the present disclosure may give priority to using the person images automatically screened based on the person album; the person images manually uploaded by the user are obtained only when the number of person images automatically screened based on the person album is less than a first quantity threshold; and in order to reduce the situation where the person images manually uploaded by the user and the automatically screened person images do not belong to the same person, person images containing the target person may be screened out from the person images manually uploaded by the user; so as to reduce the mismatch between the generated virtual image and the target person caused by the inconsistency of the characters in the two sets of person images due to user misoperation.
[0126] By setting the screening threshold of the image screening process corresponding to the character images manually uploaded by the user to be smaller than the screening threshold of the image screening process corresponding to the character images in the character album, a sufficient number of character images that meet the screening threshold can be screened out from the character images containing the target character manually uploaded by the user, so that in the subsequent process, the target character images manually uploaded by the user and the target character images automatically screened based on the character album can be used to generate a virtual image that is more consistent with the actual image of the target character, thereby effectively improving the generation effect of the virtual image of the target character.
[0127] Optionally, the method further comprises:
[0128] In response to the virtual image generation request, detecting whether a photo album of the target person is stored in the electronic device;
[0129] If the electronic device does not store a photo album of the target person, obtaining at least one person image manually uploaded by a user;
[0130] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0131] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0132] In an embodiment of the present disclosure, in response to a received virtual image generation request, it may be first detected whether the electronic device stores a photo album of the target person.
[0133] It should be noted that detecting whether the electronic device stores a photo album of the target person may also be detecting whether the electronic device provides a function of generating a photo album of the target person.
[0134] If the electronic device has a function of generating a person album, obtain the person information of the target person carried in the virtual image generation request; and based on the person information of the target person, detect whether the electronic device stores a person album of the target person. It is worth noting that if the electronic device has a function of generating a person album, but the electronic device does not store a person album of the target person, it means that the electronic device may not store a person image of the target person.
[0135] If the electronic device does not have a function of generating a person album, it means that the image of the target person may be stored in the electronic device, but there is no need to automatically select the image of the target person from the numerous stored images through the electronic device.
[0136] If it is detected that the electronic device does not store the target person's photo album, at least one person image manually uploaded by the user may be obtained; and the person image containing the target person may be screened out from the person images manually uploaded by the user.
[0137] It should be noted that face recognition can be performed on the character images manually uploaded by the user; and based on the face recognition results, character images containing the target person are screened out from the character images manually uploaded by the user; so as to reduce the situation where the character belonging to the character image manually uploaded by the user for virtual image training is different from the target person indicated by the virtual image generation request.
[0138] After obtaining character images including the target character from character images manually uploaded by the user, image screening processing may be performed on the obtained character images including the target character.
[0139] It is worth noting that the screening threshold of the image screening process performed on the person images manually uploaded by the user is smaller than the screening threshold of the image screening process performed on the person images in the person album.
[0140] It should be noted that since the person images manually uploaded by the user are person images that have been filtered by the user, in order to speed up the screening efficiency, the screening threshold of the image screening process corresponding to the person images manually uploaded by the user can be set smaller than the screening threshold of the image screening process corresponding to the person images in the person album, so as to screen out a sufficient number of person images that meet the screening threshold from the person images manually uploaded by the user that contain the target person.
[0141] In some embodiments, two sets of screening thresholds may be pre-stored. The first set of screening thresholds may be used when performing image screening on person images in a person album to obtain target person images; the second set of screening thresholds may be used when performing image screening on person images manually uploaded by users.
[0142] It is understandable that the target person image obtained after the image screening process can be used as a training image and input into the virtual image generation model to generate a virtual image of the target person.
[0143] It is worth noting that when obtaining the training image of the virtual image of the target person, the embodiment of the present disclosure may first detect whether the target person's character album is stored in the electronic device, so as to give priority to using the character images in the target person's character album; when the electronic device does not store the target person's character album, obtain the character image manually uploaded by the user; and in order to reduce the situation where the character corresponding to the character image manually uploaded by the user is different from the target person indicated by the virtual image generation request, the character image containing the target person can be filtered out from the character images manually uploaded by the user.
[0144] By setting the screening threshold of the image screening process corresponding to the character images manually uploaded by the user to be smaller than the screening threshold of the image screening process corresponding to the character images in the character album, a sufficient number of character images that meet the screening threshold can be screened out from the character images containing the target character manually uploaded by the user, so that in the subsequent process, the target character images manually uploaded by the user can be used to generate a virtual image that is more consistent with the actual image of the target character, thereby effectively improving the generation effect of the virtual image of the target character.
[0145] Optionally, the image screening process includes at least one of the following:
[0146] Performing angle detection on the at least one person image, and determining a person image in which the face of the target person is in a frontal posture from the at least one person image;
[0147] Performing image quality detection on the at least one person image, and determining a person image whose image quality meets a preset condition from the at least one person image;
[0148] Performing occlusion detection on the at least one person image, and determining a person image in which the face of the target person is not occluded from the at least one person image;
[0149] Performing human body detection on the at least one human image, and determining a human body image that only includes a human body from the at least one human image.
[0150] It should be noted that since the virtual image generation model has certain requirements for the input training images, for example, if the image quality of the training images cannot meet the requirements, the virtual image generation model may be unable to generate a virtual image of the target person; or the generated virtual image may not match the target person.
[0151] In order to reduce the impact of low image quality of training images on virtual image generation, the disclosed embodiment may perform image screening processing on character images contained in character albums and / or character images manually uploaded by users, including at least one of the following: angle detection; image quality detection; occlusion detection and human body detection.
[0152] By performing angle detection on the face area in the character image, the face posture in the character image is determined. It is worth noting that in order to ensure that the virtual image generation model can obtain more and more accurate facial feature information from the face area of the training image (i.e., the target character image), the embodiment of the present disclosure can detect the angle information corresponding to the face area in the character image when screening the training image (i.e., the target character image); based on the angle information corresponding to the face area, the character image of the target character with the face in a frontal posture is screened.
[0153] The image quality detection of the person image is performed to determine the image quality detection result of the person image; and according to the image quality detection result, a target person image is screened out from at least one person image.
[0154] In some embodiments, the image quality detection result may be used to indicate the brightness, contrast and / or resolution of the character image.
[0155] It is worth noting that the image quality mainly depends on information such as the image brightness, contrast and resolution; the image quality of the person image can be determined by performing image quality detection on the person image to obtain image quality detection results indicating the brightness, contrast and / or resolution of the person image.
[0156] By performing occlusion detection on the person image, the occlusion status of the face area in the person image is determined, and according to the occlusion detection result, a person image in which the face of the target person is not occluded is determined from at least one person image.
[0157] It is understandable that if the key points in the face area of the character image (such as the facial features area in the face) are largely blocked, the virtual image generation model may not be able to extract the facial feature information of the target person from the character image, resulting in the generated virtual image not matching the target person. Therefore, in order to ensure the generation effect of the virtual image of the target person, the character image can be subjected to occlusion detection to determine a character image in which the face of the target person is not blocked from at least one character image as a training image.
[0158] Human body detection is performed on human images to determine the number of human body areas contained in the human images. Based on the number of human body areas contained in the human images, a human image containing only one human body is determined from at least one human image, thereby reducing the impact of human body feature information of other non-target persons in the human image on the generation effect of the target person's virtual image.
[0159] It is worth noting that the virtual image generation model can extract features from the human body region of the character image to obtain human body feature information, so as to determine the body proportion parameters of the virtual image based on the human body feature information. If the character image contains multiple human body regions, the body proportions of the virtual image finally generated by the virtual image generation model will be different from the body proportions of the target person, and the virtual image will not fit the actual image of the target person.
[0160] Here, the specific human body detection method can be selected according to actual needs, and the embodiments of the present disclosure do not specifically limit this. For example, human body detection can be performed on a person image using a nanodet model.
[0161] The disclosed embodiment performs angle detection, image quality detection, occlusion detection and / or body detection on character images included in character albums and / or character images manually uploaded by users, so as to screen out target character images that meet the training image requirements from the above character images; so as to use the target character images for training to generate a virtual image that is more consistent with the actual image of the target character.
[0162] Optionally, performing angle detection on the at least one person image to determine a person image in which the face of the target person is in a frontal posture from the at least one person image includes:
[0163] Positioning key points of the at least one person image to obtain coordinate information of multiple contour key points of facial features in a face region of the person image;
[0164] Determining three-dimensional posture angle information of a face region in the character image based on a coordinate difference between the coordinate information of the plurality of contour key points and the coordinate information of corresponding reference points;
[0165] According to the three-dimensional posture angle information, a person image with a face in a frontal posture is determined.
[0166] In the disclosed embodiment, key points of the face region in the character image can be located to determine multiple contour key points corresponding to the facial features in the face region, and obtain coordinate information of the multiple contour key points.
[0167] Here, the contour key points are used to determine the contour shape of the facial features. Figure 2 As shown, Figure 2 The figure is a schematic diagram showing the positions of key points of the inner contour of a face image according to an exemplary embodiment.
[0168] It should be noted that the number of contour key points obtained varies depending on the key point positioning algorithm used. The more contour key points there are, the closer the contour shape of the facial features determined based on the contour key points is to the contour shape of the facial features of the target person, but the processing time is also longer.
[0169] In some embodiments, key point positioning is performed on the at least one character image to obtain coordinate information of 106 contour key points of facial features in the face region of the character image.
[0170] The three-dimensional posture angle information of the face area can be determined by comparing the coordinate information of multiple contour key points with the coordinate information of reference points corresponding to the multiple contour key points.
[0171] In some embodiments, the three-dimensional posture angle information may be Euler angles representing the posture of the face: pitch angle, yaw angle, and roll angle.
[0172] It is worth noting that the pitch angle is used to indicate the angle of rotation around the X-axis; here, the pitch angle can be used to describe the angle of the face raising or lowering the head; the yaw angle is used to indicate the angle of rotation around the Y-axis; here, the yaw angle can be used to describe the angle of the face turning left and right; the roll angle is used to indicate the angle of rotation around the Z-axis; here, the roll angle can be used to describe the angle of the face turning left and right.
[0173] In some embodiments, the roll angle corresponding to the face area can be determined based on the coordinate information of the contour key points corresponding to the eye area in the character image; the character image is rotated and adjusted according to the roll angle; and based on the coordinate information of multiple contour key points of the rotated character image and the coordinate information of the reference points corresponding to the multiple contour key points, the yaw angle and pitch angle of the face area are determined.
[0174] According to the comparison result between the three-dimensional posture angle information and the preset angle range, a character image with a face in a frontal posture can be determined from at least one character image.
[0175] In some embodiments, the preset angle range may include: a first angle range corresponding to the yaw angle; and a second angle range corresponding to the pitch angle.
[0176] It should be noted that the size of the roll angle of the face area will not affect the frontal posture of the face; therefore, only the yaw angle of the face area can be compared with the first angle range, and the pitch angle of the face area can be compared with the second angle range to determine the image of the person with the face in the frontal posture.
[0177] If the yaw angle of the face area is within the first angle range, and the pitch angle of the face area is also within the second angle range, it means that the character image is a character image with the face in a frontal posture; if the yaw angle of the face area is not within the first angle range, and / or the pitch angle of the face area is not within the second angle range, it means that the character image is not a character image with the face in a frontal posture.
[0178] It is worth noting that the first angle range corresponding to the character image manually uploaded by the user at least includes the first angle range corresponding to the character image in the character album; and / or, the second angle range corresponding to the character image manually uploaded by the user at least includes the second angle range corresponding to the character image in the character album.
[0179] The disclosed embodiment locates key points of a face region in a character image to obtain coordinate information of multiple contour key points of facial features in the face region; and determines three-dimensional posture angle information of the face region based on the coordinate difference between the coordinate information of the multiple contour key points of the facial features and the coordinate information of reference points corresponding to the multiple contour key points; in this way, the facial posture in the character image can be determined more accurately, so as to screen out character images with faces in a frontal posture based on the three-dimensional posture angle information of the face region.
[0180] Optionally, the performing image quality detection on the at least one person image to determine a person image whose image quality meets a preset condition from the at least one person image comprises at least one of the following:
[0181] Performing blur detection on the at least one person image to obtain blur parameters of the person image; and determining, based on the blur parameters, a person image whose blur parameters satisfy a first condition from the at least one person image;
[0182] Traversing the character image by using a sliding window to obtain a plurality of window images corresponding to the character image; determining a first ratio between the number of window images whose average brightness is greater than a first brightness threshold and the number of all window images in the plurality of window images; determining a character image whose first ratio satisfies a second condition based on the first ratios of the plurality of character images;
[0183] Use a sliding window to traverse the character image to obtain multiple window images corresponding to the character image; determine a second ratio between the number of window images whose brightness mean is less than a second brightness threshold and the number of all window images in the multiple window images; and determine a character image whose second ratio meets a third condition based on the second ratios of the multiple character images.
[0184] It should be noted that during the image acquisition process, the captured image may be overexposed, underexposed or blurred due to the influence of factors such as the shooting environment, shooting jitter, shooting distance or focus distance mismatch, which seriously affects the image quality. Based on this, the disclosed embodiment can determine the image quality of the portrait image from the three aspects of overexposure, underexposure or blurred image.
[0185] Perform blur detection on the person image to obtain the blur parameter of the person image. Here, the blur parameter is used to describe the blur degree of the person image; if the parameter value of the blur parameter is smaller, the blur degree of the person image is smaller, and the clarity of the person image is greater; conversely, if the parameter value of the blur parameter is larger, the blur degree of the person image is larger, and the clarity of the person image is smaller.
[0186] It should be noted that image clarity and image blur are two concepts that describe the clarity of an image. The clearer the image, the higher the quality, the greater the clarity and the smaller the blur. Similarly, the blurrier the image, the lower the quality, the smaller the clarity and the greater the blur.
[0187] The first condition may be that the blur parameter is less than or equal to a preset blur threshold. It is worth noting that the blur threshold corresponding to the character image manually uploaded by the user is greater than the blur threshold corresponding to the character image in the character album.
[0188] The size of the sliding window can be set according to actual needs, and the size of the sliding window is smaller than the size of the character image.
[0189] A sliding window can be used to traverse the character image to obtain multiple window images corresponding to the character image; the brightness mean of each window image in the multiple window images is determined; and based on the comparison result between the brightness mean of the window image and a preset first brightness threshold, it is determined whether the window image is a highlighted window image.
[0190] The number of window images of highlighted window images can be determined from multiple window images corresponding to the character image; a first ratio between the number of window images and the total number of window images corresponding to the character image can be determined; and based on the first ratios of the multiple character images, a character image whose first ratio satisfies a second condition can be determined.
[0191] The second condition may be that the first ratio is greater than or equal to a preset first threshold. It is worth noting that the first threshold corresponding to the character image manually uploaded by the user is greater than the first threshold corresponding to the character image in the character album.
[0192] It can be understood that if there are more highlighted window images among the multiple window images corresponding to the person image, it means that the person image may be an overexposed image.
[0193] A sliding window can be used to traverse the character image to obtain multiple window images corresponding to the character image; the brightness mean of each window image in the multiple window images is determined; and based on the comparison result between the brightness mean of the window image and a preset second brightness threshold, it is determined whether the window image is a low-brightness window image.
[0194] The number of window images of low-brightness window images can be determined from multiple window images corresponding to the character image; a second ratio between the number of window images and the total number of window images corresponding to the character image can be determined; and based on the second ratios of the multiple character images, a character image whose second ratio satisfies a third condition can be determined.
[0195] The third condition may be that the second ratio is greater than or equal to a preset second threshold. It is worth noting that the second threshold corresponding to the character image manually uploaded by the user is greater than the second threshold corresponding to the character image in the character album.
[0196] It can be understood that if there are more low-brightness window images among the multiple window images corresponding to the person image, it means that the person image may be an underexposed image.
[0197] In this way, by detecting the blur parameters of the character image and / or the first ratio and the second ratio reflecting the image brightness distribution of the character image, the blur and exposure of the character image can be determined, and then the image quality of the character image can be determined according to the blur and / or exposure of the character image, so as to screen out character images with better image quality as training images, thereby improving the generation effect of the virtual image.
[0198] Optionally, performing occlusion detection on the at least one person image to determine a person image in which the face of the target person is not occluded from the at least one person image includes:
[0199] Using the Ghost module in the occlusion detection model, convolution processing is performed on the person image to obtain a plurality of first feature maps corresponding to the person image; and based on the plurality of first feature maps, a plurality of second feature maps corresponding to each of the first feature maps are generated; the second feature maps are similar to the corresponding first feature maps;
[0200] Based on the first feature map and the second feature map corresponding to the person image, the person image is classified to obtain a classification result; the classification result indicates the probability that the face in the person image is occluded; the occlusion detection model is trained based on the modified cross entropy loss function;
[0201] Based on the classification result, a person image in which the face of the target person is not blocked is determined from the at least one person image.
[0202] In the disclosed embodiment, a pre-trained occlusion detection model may be used to perform occlusion detection on a person image to determine whether the face of a target person in the person image is occluded.
[0203] The occlusion detection model may include: a Ghost module and a classification module; wherein the Ghost module is a module that uses a series of linear operations to generate a feature map.
[0204] The Ghost module can be used to perform convolution processing on the character image to obtain multiple first feature maps corresponding to the character image; and multiple second feature maps similar to the first feature map can be generated by performing linear operations on the first feature map.
[0205] It should be noted that the feature extraction process of the Ghost module may include two operations; the first operation is to compress the number of channels of the input image using 1×1 convolution to obtain a compressed feature map. The second operation is to use depthwise separable convolution to obtain a similar feature map of the compressed feature map. In this way, the Ghost module can generate more feature maps using fewer parameters, that is, while ensuring the accuracy of the network, it can reduce network parameters and calculation amount, thereby improving the calculation speed and reducing latency.
[0206] A classification module may be used to classify the character image based on the first feature map and the second feature map of the character image to obtain a classification result.
[0207] The classification result output by the classification module is a value between 0 and 1, which is used to indicate the probability that the face in the person image is occluded.
[0208] After the classification result of the person image is obtained, the classification result may be compared with a preset classification threshold to determine a person image in which the face of the target person is not blocked.
[0209] Here, the occlusion detection model can be trained based on the modified cross entropy loss Focal loss function. It should be noted that the Focal loss function is a function modified based on the cross entropy loss function. Two hyperparameters are added on the basis of the cross entropy function to control the weight of the positive and negative samples and the weight of the difficult and easy classification samples, so as to achieve the effect of balancing the training samples by controlling the loss.
[0210] The focal loss function can make the occlusion boundary in the classification more obvious and is more conducive to threshold control.
[0211] In the disclosed embodiment, the Ghost module is used to extract features from the character image, and more feature maps are generated with fewer parameters, which greatly reduces the processing memory and processing time; and the Focal loss function is introduced to solve the problem of imbalance between difficult and easy samples in the training process of the occlusion detection model, thereby improving the detection accuracy of the occlusion detection model.
[0212] Optionally, before performing the image screening process, the method further includes:
[0213] Performing face detection on at least one person image to determine the number of face regions in the person image;
[0214] When the character image only contains one face area, determine the third ratio between the area of the face area and the image area of the character image; determine the character image to be subjected to image screening processing based on the third ratio corresponding to the at least one character image; wherein the third ratio of the character image to be subjected to image screening processing is greater than a third threshold value.
[0215] In the disclosed embodiment, before performing image screening processing on the person images contained in the person album and / or the person images manually uploaded by the user, face detection may be performed on the person images to determine the number of face areas in the person images.
[0216] It can be understood that by determining the number of facial areas in a person image, it is possible to determine whether the person image contains a facial area, and if the person image contains a facial area, it is possible to determine whether the person image is a single photo of the target person or a photo of the target person with other non-target persons.
[0217] It is worth noting that the virtual image generation model will extract the facial feature information of the target person based on the facial area in the input training image; if the person image input to the virtual image generation model also includes other non-target persons, the virtual image generation model will also extract the facial feature information of the non-target persons, which may affect the subsequent virtual image generation effect.
[0218] When the number of facial regions in the person image is one, that is, the person image is a single-person photo, determine the area of the facial region in the person image; and based on the area of the facial region and the image area of the person image, determine the third ratio between the area of the facial region and the image area of the person image.
[0219] In some embodiments, the mask area corresponding to the portrait mask can be determined by acquiring the portrait mask corresponding to the portrait image; and the third ratio between the area of the face region and the image area of the character image can be determined based on the ratio between the mask area corresponding to the portrait mask and the image area of the character image.
[0220] Here, the portrait mask is used to indicate the portrait area in the portrait image. It should be noted that if the portrait image is multiplied by the portrait mask, the portrait area in the portrait image is retained after the multiplication, while the background area other than the portrait area is eliminated, thereby achieving the effect of extracting the portrait area from the portrait image.
[0221] In some embodiments, the area of the face region can be represented by the number of pixels included in the face region, and the area of the person image can be represented by the number of pixels included in the person image. Accordingly, the third ratio between the area of the face region and the image area of the person image can be represented by the ratio between the number of pixels included in the face region and the number of pixels included in the person image.
[0222] After obtaining the third ratio corresponding to the person image, the third ratio of the person image may be compared with a preset third threshold value, and according to the comparison result, the person image to be subjected to image screening processing may be determined from the person images.
[0223] If the third ratio of the character image is greater than the third threshold, it can be determined that the character image is a character image to be subjected to image screening processing; if the third ratio of the character image is less than or equal to the third threshold, it can be determined that the character image is not a character image to be subjected to image screening processing.
[0224] It is worth noting that the third threshold corresponding to the person image manually uploaded by the user is greater than the third threshold corresponding to the person image in the person album.
[0225] Before performing image screening processing on the person images contained in the person album and / or the person images manually uploaded by the user, the disclosed embodiment performs an initial screening on the person images based on the number of facial regions in the person images and the area ratio of the facial regions, so as to screen out the person images that can be further subjected to subsequent image screening processing. This can effectively reduce the number of person images that need to be processed in the subsequent image screening processing and improve the processing efficiency of the image screening processing.
[0226] The present disclosure also provides an image screening method, such as Figure 3 As shown, Figure 3 is a schematic diagram of a process of an image screening method according to an exemplary embodiment Figure 2 The method comprises:
[0227] Step S201, in response to a request for generating a virtual image of a target person, detecting whether a person's photo album is stored in the electronic device;
[0228] It should be noted that after receiving a request to generate a virtual information image, it can be detected whether the electronic device has a photo album of the person stored. If the photo album of the person is stored, an automatic screening algorithm can be used to screen images.
[0229] Step S202, if a person photo album is stored in the electronic device, obtain the person photo album of the target person;
[0230] Here, the character album includes at least one character image including the target character.
[0231] If the electronic device stores photo albums of people, the electronic device can cyclically display any person image in each photo album; the user can manually select a target person from the displayed person images; and the photo albums of people in the electronic device are traversed using an automatic screening algorithm to screen out the photo albums of the target person.
[0232] Step S203, performing face detection on at least one person image in the person album of the target person, and determining the number of face regions in the person image;
[0233] The SCRFD face detection model can be used to perform face detection on human images and determine the number of face regions in the human images, so as to filter human images that do not contain face regions and human images that contain multiple face regions.
[0234] For a person image with a single face region that has been detected, the coordinates of the upper left corner and the lower right corner of the face region, as well as the coordinates of the five key points of the face region, can be obtained.
[0235] Step S204, when the person image contains only one face region, determining a third ratio between the area of the face region and the image area of the person image; and determining the person image to be subjected to image screening processing according to the third ratio corresponding to the at least one person image;
[0236] The face ratio of the person images of the single face area that has passed the detection is filtered. The coordinates of the upper left corner and the lower right corner of the face area are used to calculate the percentage with the image size of the person image, and the person images with a face ratio less than α% are filtered.
[0237] Step S205, performing image screening processing on at least one person image in the person album to obtain a target person image;
[0238] Here, the target person image is used to generate a virtual image of the target person.
[0239] It is worth noting that performing image screening processing on at least one person image in the person album may include:
[0240] Performing image quality detection on the at least one person image;
[0241] Performing occlusion detection on the at least one person image;
[0242] Performing angle detection on the at least one person image;
[0243] Perform human body detection on the at least one person image.
[0244] In some embodiments, performing image quality detection on the at least one person image may include:
[0245] Using the Laplace blur detection algorithm, the character images with blur values less than β are filtered;
[0246] Use an 8*8 sliding window to traverse the V channel image corresponding to the person image, set the person image whose sliding window brightness mean is greater than γ and occupies more than ζ of the person image as an overexposed image, and filter it;
[0247] Use an 8*8 sliding window to traverse the V channel image corresponding to the person image, set the person image whose sliding window brightness mean is greater than δ and occupies more than ζ of the person image as an underexposed image, and filter it;
[0248] In some embodiments, performing occlusion detection on the at least one person image includes:
[0249] The occlusion detection model is used to filter images of people with occluded faces.
[0250] It is worth noting that the input of the occlusion detection model is the person image after face detection, which is enlarged by 1.414 times. If the face area in the person image is rotated at an angle, the person image is corrected and the corrected person image is input into the occlusion detection model; the output of the occlusion detection model is a value between 0 and 1.
[0251] The occlusion detection model also refers to the basic module of GhostNet. The design model is trained by adding Focal loss and optimizing the number of channels.
[0252] In some embodiments, performing angle detection on the at least one person image includes:
[0253] Get the coordinates of 106 key points in the face area of the character image;
[0254] Determine the roll angle based on the coordinates of the key points corresponding to the eye area;
[0255] The yaw angle and pitch angle are determined by comparing the coordinates of the 106 key points with the template coordinates.
[0256] Filter the character images with yaw angle > η or pitch angle > θ.
[0257] Here, the roll angle is determined by the angle between the eye coordinates and the horizontal line. The yaw angle and pitch angle are obtained by comparing the 106 key points with the template coordinates and using the least squares method to solve the rotation matrix.
[0258] In some embodiments, performing human body detection on the at least one person image includes:
[0259] Use the nanodet model for human detection and filter images containing multiple human bodies.
[0260] For example, Figure 4 As shown, Figure 4 The figure is a flowchart of a method for automatically screening human images according to an exemplary embodiment.
[0261] Step S206, determining whether the number of the target person images obtained through screening is less than a first number threshold; if the number of the target person images obtained through screening is less than the first number threshold, obtaining at least one person image manually uploaded by the user;
[0262] It is worth noting that for users who do not have enough person images in the person album, a manual screening algorithm can be provided. It is worth noting that the manual screening adds face recognition to ensure that the target person corresponding to the person album is the same person; and the threshold requirement for the partial image screening process in the manual screening algorithm is lower.
[0263] Step S207, filtering out a person image containing the target person from at least one person image manually uploaded by the user; performing image screening processing on the filtered person image containing the target person;
[0264] In the case where the electronic device has a photo album and the user manually uploads a person's image, in order to avoid inconsistency between the uploaded person's image and the person's image filtered out from the photo album, a face detection algorithm can be introduced to identify the person's images manually uploaded by the user and filter out the person's images of non-target persons.
[0265] Step S208, if the electronic device does not store a person album, obtaining at least one person image manually uploaded by the user;
[0266] Step S209, filtering out a person image containing the target person from at least one person image manually uploaded by the user; performing image screening processing on the filtered person image containing the target person;
[0267] In the case that the electronic device does not have a person photo album, the person images manually uploaded by the user are directly filtered.
[0268] The present disclosure provides an image screening device. Figure 5 is a schematic diagram of the structure of an image screening device according to an exemplary embodiment. Figure 5 As shown, the image screening device 200 comprises:
[0269] The first acquisition module 201 is used to obtain a photo album of the target person in response to a request for generating a virtual image of the target person; the photo album includes at least one image of the target person;
[0270] The processing module 202 is used to perform image screening processing on at least one person image in the person album to obtain a target person image; the target person image is used to generate a virtual image of the target person.
[0271] Optionally, the device further includes: a second acquisition module 203, configured to:
[0272] Determining whether the number of the target person images obtained by screening is less than a first number threshold;
[0273] If the number of the target person images obtained by the screening is less than the first number threshold, obtaining at least one person image manually uploaded by a user;
[0274] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0275] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0276] Optionally, the device further includes: a second acquisition module 203, configured to:
[0277] In response to the virtual image generation request, detecting whether a photo album of the target person is stored in the electronic device;
[0278] If the electronic device does not store a photo album of the target person, obtaining at least one person image manually uploaded by a user;
[0279] Filtering a person image including the target person from at least one person image manually uploaded by a user;
[0280] The filtered person images containing the target person are subjected to image screening processing; wherein a screening threshold of the image screening processing performed on the person images manually uploaded by the user is smaller than a screening threshold of the image screening processing performed on the person images in the person album.
[0281] Optionally, the processing module 202 is configured to perform at least one of the following:
[0282] Performing angle detection on the at least one person image, and determining a person image in which the face of the target person is in a frontal posture from the at least one person image;
[0283] Performing image quality detection on the at least one person image, and determining a person image whose image quality meets a preset condition from the at least one person image;
[0284] Performing occlusion detection on the at least one person image, and determining a person image in which the face of the target person is not occluded from the at least one person image;
[0285] Performing human body detection on the at least one human image, and determining a human body image that only includes a human body from the at least one human image.
[0286] Optionally, the processing module 202 is used to:
[0287] Positioning key points of the at least one person image to obtain coordinate information of multiple contour key points of facial features in a face region of the person image;
[0288] Determining three-dimensional posture angle information of a face region in the character image based on a coordinate difference between the coordinate information of the plurality of contour key points and the coordinate information of corresponding reference points;
[0289] According to the three-dimensional posture angle information, a person image with a face in a frontal posture is determined.
[0290] Optionally, the processing module 202 is configured to perform at least one of the following:
[0291] Performing blur detection on the at least one person image to obtain blur parameters of the person image; and determining, based on the blur parameters, a person image whose blur parameters satisfy a first condition from the at least one person image;
[0292] Traversing the character image by using a sliding window to obtain a plurality of window images corresponding to the character image; determining a first ratio between the number of window images whose average brightness is greater than a first brightness threshold and the number of all window images in the plurality of window images; determining a character image whose first ratio satisfies a second condition based on the first ratios of the plurality of character images;
[0293] Use a sliding window to traverse the character image to obtain multiple window images corresponding to the character image; determine a second ratio between the number of window images whose brightness mean is less than a second brightness threshold and the number of all window images in the multiple window images; and determine a character image whose second ratio meets a third condition based on the second ratios of the multiple character images.
[0294] Optionally, the processing module 202 is used to:
[0295] Using the Ghost module in the occlusion detection model, convolution processing is performed on the person image to obtain a plurality of first feature maps corresponding to the person image; and based on the plurality of first feature maps, a plurality of second feature maps corresponding to each of the first feature maps are generated; the second feature maps are similar to the corresponding first feature maps;
[0296] Based on the first feature map and the second feature map corresponding to the person image, the person image is classified to obtain a classification result; the classification result indicates the probability that the face in the person image is occluded; the occlusion detection model is trained based on the modified cross entropy loss function;
[0297] Based on the classification result, a person image in which the face of the target person is not blocked is determined from the at least one person image.
[0298] Optionally, the processing module 202 is further configured to:
[0299] Performing face detection on at least one person image to determine the number of face regions in the person image;
[0300] When the character image only contains one face area, determine the third ratio between the area of the face area and the image area of the character image; determine the character image to be subjected to image screening processing based on the third ratio corresponding to the at least one character image; wherein the third ratio of the character image to be subjected to image screening processing is greater than a third threshold value.
[0301] Figure 6 1 is a block diagram of an electronic device according to an exemplary embodiment. For example, the electronic device may be a smart phone, a tablet computer, etc.
[0302] Reference Figure 6 The electronic device 80 may include one or more of the following components: a processing component 83, a memory 84, a power component 85, a multimedia component 86, an audio component 87, an input / output (I / O) interface 88, a sensor component 89, and a communication component 810.
[0303] The processing component 83 generally controls the overall operation of the electronic device 80, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 83 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 83 may include one or more modules to facilitate the interaction between the processing component 83 and other components. For example, the processing component 83 may include a multimedia module to facilitate the interaction between the multimedia component 86 and the processing component 83.
[0304] The memory 84 is configured to store various types of data to support operations on the electronic device 80. Examples of such data include instructions for any application or method operating on the electronic device 80, contact data, phone book data, messages, pictures, videos, etc. The memory 84 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0305] The power supply assembly 85 provides power to the various components of the electronic device 80. The power supply assembly 85 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 80.
[0306] The multimedia component 86 includes a screen that provides an output interface between the electronic device 80 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 86 includes a front camera and / or a rear camera. When the electronic device 80 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0307] The audio component 87 is configured to output and / or input audio signals. For example, the audio component 87 includes a microphone (MIC), and when the electronic device 80 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 84 or sent via the communication component 810. In some embodiments, the audio component 87 also includes a speaker for outputting audio signals.
[0308] I / O interface 88 provides an interface between processing component 83 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button and lock button.
[0309] The sensor assembly 89 includes one or more sensors for providing various aspects of status assessment for the electronic device 80. For example, the sensor assembly 89 can detect the open / closed state of the electronic device 80, the relative positioning of components, such as the display and keypad of the electronic device 80, and the sensor assembly 89 can also detect the position change of the electronic device 80 or a component of the electronic device 80, the presence or absence of user contact with the electronic device 80, the orientation or acceleration / deceleration of the electronic device 80, and the temperature change of the electronic device 80. The sensor assembly 89 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 89 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 89 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0310] The communication component 810 is configured to facilitate wired or wireless communication between the electronic device 80 and other devices. The electronic device 80 can access a wireless network based on a communication standard, such as Wi-Fi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 810 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 810 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0311] In an exemplary embodiment, the electronic device 80 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above methods.
[0312] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 84 including instructions, and the instructions can be executed by the processor 820 of the electronic device 80 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0313] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0314] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image screening method, characterized in that, the method includes: responding to a virtual image generation request of a target person, and obtaining a person album of the target person; the person album includes at least one person image containing the target person; performing image screening processing on at least one person image in the person album to obtain a target person image; the target person image is used to generate a virtual image of the target person.
2. The method according to claim 1, characterized in that, the method further includes: determining whether the number of the obtained target person images is less than a first quantity threshold; if the number of the obtained target person images is less than the first quantity threshold, obtaining at least one person image manually uploaded by the user; screening out the person images containing the target person from the at least one person image manually uploaded by the user; performing image screening processing on the screened person images containing the target person; wherein, the screening threshold for the image screening processing performed on the person images manually uploaded by the user is less than the screening threshold for the image screening processing performed on the person images in the person album.
3. The method according to claim 1, characterized in that, the method further includes: responding to the virtual image generation request, and detecting whether a person album of the target person is stored in the electronic device; if the person album of the target person is not stored in the electronic device, obtaining at least one person image manually uploaded by the user; screening out the person images containing the target person from the at least one person image manually uploaded by the user; performing image screening processing on the screened person images containing the target person; wherein, the screening threshold for the image screening processing performed on the person images manually uploaded by the user is less than the screening threshold for the image screening processing performed on the person images in the person album.
4. The method according to any one of claims 1 to 3, characterized in that, the performing of the image screening processing includes at least one of the following: performing angle detection on the at least one person image, and determining, from the at least one person image, a person image in which the face of the target person is in a frontal pose; performing image quality detection on the at least one person image, and determining, from the at least one person image, a person image whose image quality meets a preset condition; performing occlusion detection on the at least one person image, and determining, from the at least one person image, a person image in which the face of the target person is not occluded; performing human body detection on the at least one person image, and determining, from the at least one person image, a person image containing only one human body.
5. The method according to claim 4, characterized in that, the performing of the angle detection on the at least one person image and determining, from the at least one person image, a person image in which the face of the target person is in a frontal pose includes: performing key point positioning on the at least one person image to obtain coordinate information of multiple contour key points of facial features in the face area of the person image; Determine the three-dimensional pose angle information of the face region within the human image based on the coordinate differences between the coordinate information of the multiple contour key points and the coordinate information of the corresponding reference points; Determine a human image with a frontal face pose according to the three-dimensional pose angle information.
6. The method according to claim 4, wherein, the performing image quality detection on the at least one human image and determining a human image with image quality meeting a preset condition from the at least one human image includes at least one of the following: Performing blur detection on the at least one human image to obtain a blur parameter of the human image; determining a human image with a blur parameter meeting a first condition from the at least one human image according to the blur parameter; Traversing the human image by using a sliding window to obtain a plurality of window images corresponding to the human image; determining a first ratio between the number of window images with a brightness average value greater than a first brightness threshold among the plurality of window images and the total number of window images; determining a human image with the first ratio meeting a second condition according to the first ratios of the plurality of human images; Traversing the human image by using a sliding window to obtain a plurality of window images corresponding to the human image; determining a second ratio between the number of window images with a brightness average value less than a second brightness threshold among the plurality of window images and the total number of window images; determining a human image with the second ratio meeting a third condition according to the second ratios of the plurality of human images.
7. The method according to claim 4, wherein, the performing occlusion detection on the at least one human image and determining a human image with the face of the target person not occluded from the at least one human image includes: Performing convolution processing on the human image by using a Ghost module in an occlusion detection model to obtain a plurality of first feature maps corresponding to the human image; and generating a plurality of second feature maps corresponding to each of the first feature maps based on the plurality of first feature maps; the second feature maps are similar to the corresponding first feature maps; Classifying the human image based on the first feature maps and the second feature maps corresponding to the human image to obtain a classification result; the classification result indicates the probability that the face in the human image is occluded; the occlusion detection model is trained based on a modified cross-entropy loss function; Determining a human image with the face of the target person not occluded from the at least one human image based on the classification result.
8. The method according to claim 4, wherein, before performing the image screening process, the method further includes: Performing face detection on at least one human image to determine the number of face regions within the human image; When the human image only includes one face region, determining a third ratio between the area of the face region and the image area of the human image; determining a human image to be subjected to image screening processing according to the third ratios corresponding to the at least one human image; wherein, the third ratio of the human image to be subjected to image screening processing is greater than a third threshold.
9. An image screening device, characterized in that, the device includes: a first acquisition module, configured to acquire a person album of the target person in response to a virtual image generation request of the target person; the person album includes at least one person image containing the target person; a processing module, configured to perform image screening processing on at least one person image in the person album to obtain a target person image; the target person image is used to generate a virtual image of the target person.
10. An electronic device, characterized in that, it includes: a processor; a memory for storing executable instructions; wherein, the processor is configured to: when executing the executable instructions stored in the memory, implement the image screening method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enable the electronic device to execute the image screening method according to any one of claims 1 to 8.