Identification photo generation method and device, readable storage medium and program product
By predicting the user's part characteristics and deforming the reference diagram in shape, a more natural and personalized ID photo is generated, the problem of lack of personalization and poor visual effects in the existing technology ID photo generation method is solved.
Patent Information
- Application Number
- CN202510085300.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The existing method of generating ID photos leads to the same clothing for everyone through sticking, lacking personalization, and there are often color differences, skin color, and light and shadow processing between the neck, face and collar. The sense of incongruity is obvious, resulting in poor visual effects of the generated ID photos.
By obtaining the reference image and shooting images, predicting the user's part based on the prediction model, obtaining the predicted neck length and shoulder width, and deforming the target dressing effect in the reference image based on these prediction results, and generating a preliminary document photo containing the face.
The visual effect of generating ID photos is improved, making the dressing effect more natural and accurate, enhancing personalized characteristics, and reducing color aberration and sense of incongruity.
Smart Images

Figure CN119991878A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer device, computer-readable storage medium and computer program product for generating a certificate photo. Background Art
[0002] With the development of computer technology and Internet technology, users have an increasing demand for ID photos, which are widely used in job hunting, exam registration, and enrollment. However, there are many inconveniences in traditional photo studios. For example, users need to arrange time to go there, which is time-consuming and laborious, and the quality of the photos varies. Generating ID photos through the Internet is more convenient. Users can upload photos anytime and anywhere to generate ID photos, without having to go out and queue up, and get the results instantly. In addition, artificial intelligence (AI) technology can also provide personalized services such as beautifying ID photos and background adjustment, which not only improves efficiency, but also reduces shooting costs, providing users with more flexible choices.
[0003] However, the current ID photo generation method mainly uses AI ID photo technology to provide diversified clothing according to user needs, such as formal wear, casual wear or ethnic costumes, to meet the requirements of users in different occasions. This is difficult to achieve in traditional shooting, and users do not need to prepare multiple sets of clothing. However, the above-mentioned AI dressing method mainly uses the method of mapping to directly splice the faces of different users onto a fixed template, resulting in the same clothing for everyone, lack of personalization, and often color difference, skin color, and light and shadow processing between the neck, face and collar. There are obvious problems such as inconsistency, which leads to poor visual effects of the generated ID photos. Therefore, how to effectively improve the visual effects of generated ID photos has become an urgent problem to be solved. Summary of the invention
[0004] Based on this, the present application provides a method, apparatus, computer device, computer-readable storage medium and computer program product for generating ID photos, which can effectively improve the visual effect of generating ID photos.
[0005] On the one hand, the present application provides a method for generating an ID photo, comprising: obtaining a reference image and a captured image, wherein the captured image contains a user's face; predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; deforming the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
[0006] In one embodiment, after obtaining the reference image and the captured image, the method further includes: performing face detection on the captured image to obtain a face detection frame; cropping the face detection frame to obtain a face image; predicting the part of the user based on the face image and the reference image to obtain a predicted neck length and a predicted shoulder width, including: inputting the face image and the reference image into a prediction model, predicting the part of the user through the prediction model, and obtaining the predicted neck length and predicted shoulder width of the face image; or inputting the face image and the reference image into a prediction model, predicting the part of the user through the prediction model, and obtaining the predicted neck deformation coefficient and predicted shoulder deformation coefficient of the reference image.
[0007] In one of the embodiments, generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image includes: splicing the deformed reference image and the face image to obtain a textured image; extracting reference image features of the reference image; identifying image content in the reference image and converting the image content into a reference text description; generating a preliminary ID photo containing the face based on the textured image, the reference image features and the reference text description.
[0008] In one of the embodiments, after the deformed reference image and the face image are spliced to obtain a texture image, the method further includes: generating a line texture image based on the texture image; generating a preliminary ID photo containing the face based on the texture image, the reference image features and the reference text description, including: processing the line texture image, the reference image features and the reference text description through an image generation model to obtain a preliminary ID photo containing the face.
[0009] In one of the embodiments, after generating a preliminary ID photo including the face based on the reference image, the deformed reference image and the captured image, the method further includes: performing edge optimization on the preliminary ID photo to obtain a target ID photo including the face.
[0010] In one of the embodiments, the method further includes: performing face detection on the preliminary ID photo to obtain a face detection frame; cropping the face detection frame to obtain a partial face image; performing edge optimization on the preliminary ID photo to obtain a target ID photo containing the face, including: performing edge optimization on the partial face image through an image generation model to obtain a target ID photo containing the face.
[0011] In one of the embodiments, the prediction of the user's body part is achieved through a trained prediction model, and the training steps of the prediction model include: obtaining a sample face image set and a label of each sample face image in the sample face image set; the label identifies the neck length and shoulder width in each sample face image; using the sample face image set and the label of each sample face image as training data, training the initial prediction model to obtain the trained prediction model.
[0012] On the one hand, the present application also provides a device for generating an ID photo, comprising: an acquisition module, used to acquire a reference image and a captured image, wherein the captured image contains a user's face; a prediction module, used to predict the user's body parts based on the captured image and the reference image, and obtain a predicted neck length and a predicted shoulder width; a deformation module, used to deform the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width, and obtain a deformed reference image; a generation module, used to generate a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
[0013] On the one hand, the present application also provides a computer device, including a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program: obtaining a reference image and a captured image, the captured image containing a user's face; predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; deforming the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
[0014] On the one hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps: obtaining a reference image and a captured image, wherein the captured image contains a user's face; predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; deforming the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
[0015] On the one hand, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the following steps: obtaining a reference image and a captured image, wherein the captured image contains a user's face; predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; deforming the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
[0016] The above-mentioned ID photo generation method, device, computer equipment, computer-readable storage medium and computer program product obtain a reference image and a captured image, wherein the captured image contains the user's face, and predicts the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; further, the target dressing effect in the reference image is deformed based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image, and a preliminary ID photo containing a face is generated based on the reference image, the deformed reference image and the captured image. Since the user's body parts can be predicted based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width, the target dressing effect in the reference image can be deformed based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image that better matches the user's face contained in the captured image, thereby making the dressing effect of the preliminary ID photo containing the user's face generated based on the reference image, the deformed reference image and the captured image more natural and accurate, effectively improving the visual effect of the generated ID photo. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A diagram showing an application environment of a method for generating a certificate photo in an embodiment;
[0019] Figure 2 Schematic diagram of a process of generating a certificate photo in one embodiment;
[0020] Figure 3 A schematic diagram of the overall process of a method for generating a certificate photo provided in an embodiment;
[0021] Figure 4A schematic diagram of a process for generating an initial preliminary ID photo in one embodiment;
[0022] Figure 5 A schematic diagram of a process of generating a certificate photo using a pre-trained large model in one embodiment;
[0023] Figure 6 It is a structural block diagram of a device for generating a certificate photo in one embodiment;
[0024] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and beneficial effects of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0026] The ID photo generation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The server 104 can be a background server of the image application. After the terminal 102 obtains the reference image and the captured image, the captured image contains the user's face. The terminal can send the obtained reference image and the captured image containing the user's face to the background server of the image application, that is, the server 104, so that the server 104 predicts the user's part based on the captured image and the reference image, obtains the predicted neck length and the predicted shoulder width, and deforms the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain the deformed reference image; further, the server 104 can generate a preliminary ID photo containing a face based on the reference image, the deformed reference image and the captured image, and return the preliminary ID photo containing a face to the terminal 102, so that the terminal 102 displays the preliminary ID photo containing a face.
[0027] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0028] In an exemplary embodiment, Figure 2 As shown, a method for generating a certificate photo is provided, and the method is applied to Figure 1 The terminal in the example is used to illustrate, including the following steps 202 to 208. Among them:
[0029] Step 202: Acquire a reference image and a captured image, where the captured image contains the user's face.
[0030] The reference image refers to an image containing the target dressing effect. That is, the reference image in this application can be a pre-configured template reference image, or it can be an image containing the target dressing effect selected in real time. The target dressing effect can include a variety of clothing dressing effects, such as formal wear, casual wear or ethnic clothing to meet the requirements of different occasions, and can also include a variety of accessories dressing effects, such as hair accessories, earrings, necklaces and other accessories, and can also include a variety of makeup dressing effects, etc.
[0031] It can be understood that the target dressing effects in the present application include but are not limited to: diversified clothing dressing effects, accessories dressing effects, makeup dressing effects, etc., and may also include other customized dressing effects.
[0032] The captured image refers to an image containing the face of the target user. It is understood that the captured image in this application can be a real-time image taken by the user, or a historical image stored in the database. The captured image in this application needs to contain the face of the target user. It can be an image containing only the face area, or an image containing the half body or the whole body of the target user. For example, Figure 3 As shown in FIG. 1 , it is a schematic diagram of the overall process of the method for generating a certificate photo provided by the present application. The captured image may be input by the operator (the user currently using the terminal device). Figure 3 The original image shown in (i.e., the captured image containing the user's face).
[0033] Step 204 , predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width.
[0034] The predicted neck length refers to the predicted neck length in the captured image, or refers to the predicted neck deformation coefficient of the reference image relative to the captured image.
[0035] It can be understood that the predicted neck length in the captured image can also be converted into a predicted neck deformation coefficient of the reference image relative to the captured image.
[0036] The predicted shoulder width refers to the predicted shoulder width in the captured image, or refers to the predicted shoulder deformation coefficient of the reference image relative to the captured image.
[0037] Optionally / exemplarily, the devices used by different users (operating objects) can interact with the image application (or image generation system). When the user (operating object) wants to generate a personalized ID photo, the user can open the image application (Application, APP) on the terminal through a trigger operation, and enter the image generation page of the image application through a selection operation, that is, the user can log in to the image application (such as the image generation application) through a trigger operation. Further, the user can trigger an information input operation or an information selection operation in the generation page of the image application displayed on the terminal, so that the terminal responds to the above-mentioned information input operation or information selection operation triggered by the user, obtains the reference image and the captured image input by the user, and the captured image contains the face of the target user. Further, the terminal can predict the user's body parts based on the captured image and the reference image to obtain the predicted neck length and the predicted shoulder width. For example, the terminal can first perform face detection on the captured image to obtain a face detection frame, and crop the face detection frame to obtain a face image; further, the terminal can call a pre-trained prediction model, and input the cropped face image and the reference image in the captured image as input parameters into the prediction model, and after being processed by the prediction model, the predicted neck length and predicted shoulder width corresponding to the face image are output; or, the predicted neck deformation coefficient and predicted shoulder deformation coefficient of the reference image (relative to the face image) are output.
[0038] It can be understood that the method provided by the present application can be implemented through the interaction between the terminal and the background server of the image application, or it can be implemented through the interaction between the front end and the back end of the terminal, that is, the front end of the terminal is used to display the generated preliminary ID photo containing the face of the target user, and the back end of the terminal is equivalent to the background server, which is used to perform logical processing such as prediction and conversion on the reference image and captured image input by the operation object.
[0039] For example, Figure 3As shown in , when a user (operating object) wants to generate a personalized ID photo, the user can open the image application (Application, APP) on the terminal through a trigger operation, and enter the image generation page of the image application through a selection operation. Furthermore, the user can trigger an information input operation or an information selection operation in the generation page of the image application displayed on the terminal, so that the terminal responds to the above-mentioned information input operation or information selection operation triggered by the user, and obtains the information input by the user. Figure 3 The reference image and the original image shown in are the captured images, and the original image contains the face of the target user Xiao A.
[0040] Furthermore, the terminal performs face detection on the captured image to obtain a face detection frame, and crops the face detection frame to obtain a face image; the terminal can call a pre-trained prediction model, and input the cropped face image and the reference image in the captured image as input parameters into the prediction model, and after being processed by the prediction model, the predicted neck deformation coefficient and the predicted shoulder deformation coefficient of the reference image (relative to the face image) are output.
[0041] Step 206 , performing shape deformation on the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image.
[0042] Among them, the deformed reference image refers to the image obtained by shape deformation of the dressed image in the reference image. For example, the deformed reference image in the present application can be the image obtained by shape deformation of the dressed image of the person in the reference image. The reason for the shape deformation processing is to adapt to the body shape features of different users (faces) contained in the input captured image.
[0043] Step 208, generating a preliminary ID photo containing a face based on the reference image, the deformed reference image and the captured image.
[0044] The preliminary ID photo refers to a preliminary image containing the target user's face. The preliminary ID photo in this application can be a preliminary dressing image output after being processed by the image generation model. For example, the preliminary ID photo in this application can be Figure 3 The first step shown in generates the image corresponding to the result.
[0045] A preliminary ID photo containing a face means that the ID photo contains the face of the target user contained in the captured image. Figure 3 The image corresponding to the first step generation result shown in includes the face of the target user Xiao A and the clothes dressed as a student.
[0046] Specifically, the terminal predicts the user's body parts based on the captured image and the reference image, and after obtaining the predicted neck length and predicted shoulder width, the terminal can deform the target dressing effect in the reference image based on the predicted neck length and predicted shoulder width to obtain a deformed reference image, and generate a preliminary ID photo containing the user's face based on the reference image, the deformed reference image, and the captured image. For example, the terminal can splice the deformed reference image and the face image to obtain a texture image, extract the reference image features of the reference image, identify the image content in the reference image, and convert the image content into a reference text description. Finally, the terminal can process the texture image, the reference image features, and the reference text description through a pre-trained image generation model, and output a preliminary ID photo containing the user's face.
[0047] For example, Figure 3 As shown in , the terminal inputs the cropped face image and the reference image in the captured image as input parameters into the prediction model. After being processed by the prediction model, the predicted neck deformation coefficient and the predicted shoulder deformation coefficient of the reference image (relative to the face image) are output. The terminal can then perform the prediction based on the predicted neck deformation coefficient and the predicted shoulder deformation coefficient. Figure 3 The target dress effect (student dress) in the reference image shown in the figure is deformed to obtain the deformed reference image, and the deformed reference image and the face image are spliced to obtain the following: Figure 3 At the same time, the terminal can Figure 3 The CLIP image feature extraction model shown in FIG. 1 extracts the reference image features of the reference image and uses the Figure 3 The VLM visual-textual multimodal model shown in the figure recognizes the image content in the reference image and converts the image content into a reference text description. Finally, the terminal can use Figure 3 The pre-trained image generation model shown in , namely the MV large model, processes the texture image, reference image features and reference text description, and outputs the following Figure 3 The preliminary ID photo containing the user's face shown in is the first step to generate the result image.
[0048] In this embodiment, by obtaining a reference image and a captured image, the captured image includes the user's face, and based on the captured image and the reference image, the user's body parts are predicted to obtain a predicted neck length and a predicted shoulder width; further, based on the predicted neck length and the predicted shoulder width, the target dressing effect in the reference image is deformed to obtain a deformed reference image, and based on the reference image, the deformed reference image and the captured image, a preliminary ID photo including the face is generated. Since the user's body parts can be predicted based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width, the target dressing effect in the reference image is deformed based on the predicted neck length and the predicted shoulder width, so that a deformed reference image that better matches the user's face included in the captured image can be obtained, thereby making the dressing effect of the preliminary ID photo including the user's face generated based on the reference image, the deformed reference image and the captured image more natural and accurate, and effectively improving the visual effect of the generated ID photo.
[0049] In an exemplary embodiment, after acquiring the reference image and capturing the image, the method further includes:
[0050] Perform face detection on the captured image to obtain a face detection frame;
[0051] Crop the face detection frame to obtain the face image;
[0052] The predicting of the user's body part based on the face image and the reference image to obtain a predicted neck length and a predicted shoulder width includes:
[0053] Input the face image and the reference image into the prediction model, and use the prediction model to predict the user's body parts to obtain the predicted neck length and predicted shoulder width of the face image; or,
[0054] The face image and the reference image are input into a prediction model, and the user's body parts are predicted by the prediction model to obtain a predicted neck deformation coefficient and a predicted shoulder deformation coefficient of the reference image.
[0055] Specifically, the generation of a certificate photo containing a personalized service costume is used as an example for explanation. Figure 4 As shown in FIG. 1 , it is a schematic diagram of the process of generating an initial preliminary ID photo. Assume that the terminal obtains the user input such as Figure 4 After the reference image and the captured image shown in FIG. Figure 4The captured image shown in the figure is used for face detection to obtain a face detection frame, and the face detection frame is cropped to obtain a face image; further, the terminal can input the face image and the reference image into the prediction model, and predict the user's body part through the prediction model to obtain the predicted neck length and predicted shoulder width of the face image; or, the terminal can input the face image and the reference image into the prediction model, and predict the user's body part through the prediction model to obtain the predicted neck deformation coefficient and predicted shoulder deformation coefficient of the reference image. As a result, not only high-quality ID photos can be generated, but also the naturalness and accuracy of the dressing effect can be greatly improved.
[0056] In an exemplary embodiment, the step of generating a preliminary ID photo containing a face based on the reference image, the deformed reference image, and the captured image includes:
[0057] The deformed reference image and the face image are spliced to obtain a texture image;
[0058] extracting reference image features of the reference image;
[0059] Identify the image content in the reference image and convert the image content into a reference text description;
[0060] Generate a preliminary ID photo containing a face based on the texture image, reference image features and reference text description.
[0061] Specifically, the example of generating a certificate photo containing a personalized service costume is used for explanation. Figure 4 The face image and reference image shown in are input into the prediction model, and the user's body parts are predicted by the prediction model to obtain the predicted neck deformation coefficient and the predicted shoulder deformation coefficient of the reference image (relative to the face image). Then, the terminal can deform the shape of the student dress effect in the reference image based on the predicted neck deformation coefficient and the predicted shoulder deformation coefficient to obtain the deformed reference image, and splice the deformed reference image with the face image to obtain the following: Figure 4 At the same time, the terminal can Figure 4 The CLIP image feature extraction model shown in FIG. 1 extracts the reference image features of the reference image and uses the Figure 4 The VLM visual-textual multimodal model shown in the figure recognizes the image content in the reference image and converts the image content into a reference text description. Finally, the terminal can use Figure 4 The pre-trained image generation model shown in , namely the MV large model, processes the texture image, reference image features and reference text description, and outputs the following Figure 4 The preliminary ID photo containing the user's face is shown in . This makes it possible not only to generate high-quality ID photos, but also to greatly improve the naturalness and accuracy of the dressing effect.
[0062] In one exemplary embodiment, after splicing the deformed reference image and the face image to obtain a texture image, the method further includes:
[0063] Based on the map image, generate a line map image;
[0064] The step of generating a preliminary ID photo containing the face of the person based on the texture image, the reference image features and the reference text description includes:
[0065] The line map image, reference image features and reference text description are processed through an image generation model to obtain a preliminary ID photo containing a face.
[0066] Specifically, the example of generating a ID photo containing a personalized service costume is used for explanation. Assume that the terminal deforms the student costume effect in the reference image based on the predicted neck deformation coefficient and the predicted shoulder deformation coefficient to obtain a deformed reference image, and then splices the deformed reference image with the face image to obtain the following image: Figure 4 After the map image shown in , the terminal can generate a line map image based on the map image. At the same time, the terminal can generate a line map image based on the map image. Figure 4 The CLIP image feature extraction model shown in FIG. 1 extracts the reference image features of the reference image and uses the Figure 4 The VLM visual-textual multimodal model shown in the figure recognizes the image content in the reference image and converts the image content into a reference text description. Finally, the terminal can use Figure 4 The pre-trained image generation model shown in , namely the MV large model, processes the line map image, reference image features and reference text description, and the output is as follows Figure 4 The preliminary ID photo containing the user's face is shown in . This makes it possible not only to generate high-quality ID photos, but also to greatly improve the naturalness and accuracy of the dressing effect.
[0067] In an exemplary embodiment, after generating a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image, the method further includes:
[0068] The preliminary ID photo is edge optimized to obtain the target ID photo containing the face.
[0069] Specifically, the generation of a certificate photo containing a personalized service costume is used as an example for explanation. Figure 5 As shown in the figure, it is a flowchart of generating a ID photo through a pre-trained large model. Figure 5 The pre-trained image generation model shown in , namely the MV large model, processes the line map image, reference image features and reference text description, and the output is as follows Figure 3 After the preliminary ID photo containing the user's face shown in the first step generates the result, Figure 3 As shown in , the terminal can use the result of the first step, i.e., the preliminary ID photo, as an input parameter and re-enter the Figure 5 The pre-trained image generation model, i.e., the MV large model, is used in the MV large model, so that the image generation model, i.e., the MV large model, performs edge optimization on the preliminary ID photo, such as optimizing the face edge and hair edge in the preliminary ID photo, and the output is as follows Figure 3 The visual effect shown in the figure is more harmonious and realistic, and the target ID photo containing the face of the target user Xiao A, that is, the final result image, can be generated not only with high quality, but also with greatly improved naturalness and accuracy of the dressing effect.
[0070] In one exemplary embodiment, the method further comprises:
[0071] Perform face detection on the preliminary ID photo to obtain a face detection frame;
[0072] Crop the face detection frame to obtain a partial face image;
[0073] The step of performing edge optimization on the preliminary ID photo to obtain a target ID photo containing the face of the person includes:
[0074] The edge of the local face image is optimized through the image generation model to obtain the target ID photo containing the face.
[0075] Specifically, the generation of a certificate photo containing a personalized service costume is used as an example for explanation. Figure 5 The pre-trained image generation model shown in , namely the MV large model, processes the line map image, reference image features and reference text description, and the output is as follows Figure 3 After the preliminary ID photo containing the user's face shown in the first step generates the result, Figure 3 As shown in , the terminal can perform face detection on the preliminary ID photo to obtain a face detection frame, and crop the face detection frame to obtain a partial face image; further, the terminal can use the first step generation result, that is, the partial face image captured in the preliminary ID photo, as an input parameter, and re-enter as shown in Figure 5 The pre-trained image generation model, i.e., the MV large model, is used in the MV large model, so that the image generation model, i.e., the MV large model, performs edge optimization on the local face image, such as optimizing the face edge and hair edge in the local face image, and the output is as follows Figure 3 The visual effect shown in the figure is more harmonious and realistic, and the target ID photo containing the face of the target user Xiao A, that is, the final result image, can be generated not only with high quality, but also with greatly improved naturalness and accuracy of the dressing effect.
[0076] In an exemplary embodiment, the prediction of the user's position is achieved by a trained prediction model, and the training steps of the prediction model include:
[0077] Obtain a sample face image set and a label of each sample face image in the sample face image set; the label identifies the neck length and shoulder width in each sample face image;
[0078] The sample face image set and the label of each sample face image are used as training data to train the initial prediction model to obtain a trained prediction model.
[0079] Specifically, the terminal can obtain a sample face image set and a label of each sample face image in the sample face image set, and use the sample face image set and the label of each sample face image as training data to train the initial prediction model to obtain a trained prediction model. Among them, the label is used to identify the neck length and shoulder width in each sample face image. That is, the training data set of the prediction model in this application includes a large number of face images with neck length and shoulder width annotated. These images cover individuals of various ages, genders and body shapes, ensuring the diversity and representativeness of the data set. The data preprocessing steps include image standardization, cropping and data enhancement to improve the generalization ability of the model. The model uses mean square error (MSE) as a loss function to optimize the prediction accuracy of neck length and shoulder width. As a result, the neck length and shoulder width of different users can be predicted from different face images input by the user.
[0080] The present application also provides an application scenario, which applies the above-mentioned ID photo generation method. The method provided in the embodiment of the present application can be applied to various personalized ID photo generation scenarios. The following takes the scenario of user interaction with the image generation system as an example to illustrate the ID photo generation method provided in the embodiment of the present application.
[0081] That is, this application proposes a method for generating ID photo replacement based on reference images. First, the neck length and shoulder width prediction model is used to estimate the neck length and shoulder width of different users. Then, the reference image is deformed to adapt to the body shape of different users, and then the user image is detected and cropped for mapping. The generation process is divided into two stages:
[0082] 1. Stage 1: Generate preliminary dressing images using the MV large model combined with image embedding and line drawings of the reference image.
[0083] 2. The second stage: the generated image is cropped for the face, and the cropped image is input into the MV model again for further generation. In this stage, the focus is on optimizing the edges of the face and hair to improve the harmony and authenticity of the overall effect.
[0084] Through the above steps, this solution can generate high-quality ID photo replacement images to achieve a more natural and realistic effect. Figure 3 As shown in, including:
[0085] 1. Neck length and shoulder width prediction model
[0086] Accurate prediction of human features such as neck length and shoulder width is crucial. Accurate estimation of these features is important for generating ID photos. This method proposes a deep learning-based model that can predict neck length and shoulder width from user-input face images. The training dataset of the model includes a large number of face images annotated with neck length and shoulder width. These images cover individuals of various ages, genders, and body shapes, ensuring the diversity and representativeness of the dataset. The data preprocessing steps include image standardization, cropping, and data augmentation to improve the generalization ability of the model. The model uses mean squared error (MSE) as the loss function to optimize the prediction accuracy of neck length and shoulder width.
[0087] 2. First stage generation
[0088] The first stage redraws the image after mapping (implemented by the MV large model). The first stage includes: obtaining the description text of the ID photo, generating the ID photo latent using the MV large model based on the ID photo description text (the description text converted from the image content in the map by using the model to recognize the image), the map image latent (i.e., the latent space image features), the reference image embedding, and the map image line drawing, and decoding the features of the latent to obtain the initial ID photo.
[0089] 3. Second stage generation
[0090] In order to make the ID photo more harmonious, the edges of the face are redrawn again.
[0091] The beneficial effects of the technical solution of this application include:
[0092] The technical solution provided by this application has significant advantages in the generation of ID photo costume changes. First, the user's body shape is accurately estimated through the neck length and shoulder width prediction model, making the costume change effect more natural and realistic. Secondly, the shape deformation of the reference image and the precise mapping of the user image ensure the efficiency and accuracy of the costume change process. The generation process is divided into two stages. The initial generation uses the MV large model combined with image embedding and line drawings to ensure the richness of details in the costume change image; the subsequent optimization stage focuses on the processing of the face and hair edges, which improves the harmony and authenticity of the image. This method can not only generate high-quality ID photos, but also greatly improve the naturalness and accuracy of the costume change effect.
[0093] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0094] Based on the same inventive concept, the embodiment of the present application also provides a device for generating a certificate photo for implementing the method for generating a certificate photo involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the method above, so the specific limitations in one or more embodiments of the device for generating a certificate photo provided below can refer to the limitations of the method for generating a certificate photo above, and will not be repeated here.
[0095] In an exemplary embodiment, Figure 6 As shown, a device for generating a certificate photo is provided, comprising: an acquisition module 602, a prediction module 604, a deformation module 606 and a generation module 608, wherein:
[0096] The acquisition module 602 is used to acquire a reference image and a captured image, wherein the captured image contains the user's face.
[0097] The prediction module 604 is used to predict the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width.
[0098] The deformation module 606 is used to perform shape deformation on the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image.
[0099] The generation module 608 is used to generate a preliminary ID photo containing the face of the person based on the reference image, the deformed reference image and the captured image.
[0100] In one embodiment, the device also includes: a detection module, which is used to perform face detection on the captured image to obtain a face detection frame; a cropping module, which is used to crop the face detection frame to obtain a face image; the prediction module is also used to input the face image and the reference image into a prediction model, predict the user's part through the prediction model, and obtain the predicted neck length and predicted shoulder width of the face image; or, input the face image and the reference image into a prediction model, predict the user's part through the prediction model, and obtain the predicted neck deformation coefficient and predicted shoulder deformation coefficient of the reference image.
[0101] In one embodiment, the device also includes: a splicing module, which is used to splice the deformed reference image and the face image to obtain a map image; an extraction module, which is used to extract reference image features of the reference image; a conversion module, which is used to identify the image content in the reference image and convert the image content into a reference text description; the generation module is also used to generate a preliminary ID photo containing the face based on the map image, the reference image features and the reference text description.
[0102] In one embodiment, the generation module is also used to generate a line map image based on the map image; the device also includes: a processing module, used to process the line map image, the reference image features and the reference text description through an image generation model to obtain a preliminary ID photo containing the face.
[0103] In one embodiment, the device further includes: an optimization module, configured to perform edge optimization on the preliminary ID photo to obtain a target ID photo containing the face.
[0104] In one embodiment, the device also includes: a detection module, which is used to perform face detection on the preliminary ID photo to obtain a face detection frame; a cropping module, which is used to crop the face detection frame to obtain a partial face image; and the optimization module is also used to perform edge optimization on the partial face image through an image generation model to obtain a target ID photo containing the face.
[0105] In one embodiment, the acquisition module is also used to obtain a sample face image set and a label of each sample face image in the sample face image set; the label identifies the neck length and shoulder width in each sample face image; the device also includes: a training module, used to use the sample face image set and the label of each sample face image as training data to train the initial prediction model to obtain the trained prediction model.
[0106] Each module in the above-mentioned ID photo generation device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0107] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through Wi-Fi, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method for generating a certificate photo is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0108] Those skilled in the art will understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0109] In an exemplary embodiment, in one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0110] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0111] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0113] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0114] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0115] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for generating a certificate photo, characterized in that: The method comprises: Acquire a reference image and a captured image, wherein the captured image contains a user's face; Predicting the user's body parts based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; Performing shape deformation on the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; Based on the reference image, the deformed reference image and the captured image, a preliminary ID photo containing the face of the person is generated.
2. The method according to claim 1, characterized in that After obtaining the reference image and capturing the image, the method further includes: Performing face detection on the captured image to obtain a face detection frame; Cropping the face detection frame to obtain a face image; The predicting of the user's body part based on the face image and the reference image to obtain a predicted neck length and a predicted shoulder width includes: Inputting the face image and the reference image into a prediction model, predicting the user's body parts through the prediction model, and obtaining a predicted neck length and a predicted shoulder width of the face image; or The face image and the reference image are input into a prediction model, and the user's body parts are predicted by the prediction model to obtain a predicted neck deformation coefficient and a predicted shoulder deformation coefficient of the reference image.
3. The method according to claim 2, characterized in that The step of generating a preliminary ID photo containing the face of the person based on the reference image, the deformed reference image and the captured image comprises: Splicing the deformed reference image and the face image to obtain a texture image; extracting reference image features of the reference image; Identifying image content in the reference image and converting the image content into a reference text description; Based on the texture image, the reference image features and the reference text description, a preliminary ID photo containing the face of the person is generated.
4. The method according to claim 3, characterized in that After the deformed reference image and the face image are spliced to obtain a texture image, the method further includes: Based on the map image, generate a line map image; The step of generating a preliminary ID photo containing the face of the person based on the texture image, the reference image features and the reference text description includes: The line map image, the reference image features and the reference text description are processed through an image generation model to obtain a preliminary ID photo containing the face of the person.
5. The method according to any one of claims 1 to 4, characterized in that: After generating a preliminary ID photo containing the face of the person based on the reference image, the deformed reference image and the captured image, the method further includes: The preliminary ID photo is edge-optimized to obtain a target ID photo that includes the face of the person.
6. The method according to claim 5, characterized in that The method further comprises: Performing face detection on the preliminary ID photo to obtain a face detection frame; Cropping the face detection frame to obtain a partial face image; The step of performing edge optimization on the preliminary ID photo to obtain a target ID photo containing the face of the person includes: The edge of the partial face image is optimized through an image generation model to obtain a target ID photo containing the face.
7. The method according to claim 1, characterized in that The prediction of the user's position is achieved by a trained prediction model, and the training steps of the prediction model include: Obtaining a sample face image set and a label of each sample face image in the sample face image set; the label identifies the neck length and shoulder width in each sample face image; The sample face image set and the labels of each sample face image are used as training data to train the initial prediction model to obtain the trained prediction model.
8. A device for generating a certificate photo, characterized in that: The device comprises: An acquisition module, used to acquire a reference image and a captured image, wherein the captured image contains a user's face; A prediction module, configured to predict the user's body part based on the captured image and the reference image to obtain a predicted neck length and a predicted shoulder width; A deformation module, configured to perform shape deformation on the target dressing effect in the reference image based on the predicted neck length and the predicted shoulder width to obtain a deformed reference image; A generation module is used to generate a preliminary ID photo containing the face based on the reference image, the deformed reference image and the captured image.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.