Image processing method, and training method and device of image processing model
By obtaining target person images and clothing images for masking and image processing model training, the problems of low dressing efficiency and poor effect are solved, and a dressing effect that efficiently retains clothing details is achieved.
Patent Information
- Application Number
- CN202410275109.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the efficiency of portrait dressing is low and the dressing effect is poor, and it is difficult to effectively retain the detailed features of the character except for the parts to be dressed.
By obtaining the image of the target person, the key points of the human body and the clothing image, the masking and image processing model training is performed to generate the image of the target person wearing the target clothing, and the feature extraction network and denoising network are used to retain the detailed features.
Improves the efficiency of changing clothes, retains the detailed features of the target clothes, and improves the effect of changing clothes.
Smart Images

Figure CN120635249A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image processing method, a training method and a device for an image processing model. Background Art
[0002] Portrait dressing technology involves fusing a human image with a clothing template image to create an image of the person wearing the template clothing. This allows users to visualize the effect of the template clothing without having to wear the actual clothing. Portrait dressing technology is widely used in online shopping, clothing displays, fashion design, and offline shopping. Improving the efficiency of dressing and the quality of the resulting images are currently unresolved challenges. Summary of the Invention
[0003] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.
[0004] The first embodiment of the present disclosure provides an image processing method, including:
[0005] Acquire a first image containing a target person, target body key points corresponding to the target person in the first image, and a second image containing target clothing;
[0006] Masking the part of the target person corresponding to the part to be changed in the first image to obtain a third image including a target mask corresponding to the part to be changed;
[0007] placing the second image within the target mask of the third image to obtain a fourth image;
[0008] The second image, the fourth image and the target human key points are input into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
[0009] A second embodiment of the present disclosure provides a method for training an image processing model, including:
[0010] Acquire a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image;
[0011] Segmenting the first sample image to obtain a second sample image corresponding to the clothing of the part of the sample person to be changed;
[0012] Masking the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image including a sample mask corresponding to the part to be changed;
[0013] placing the second sample image within the sample mask in the third sample image to obtain a fourth sample image;
[0014] Inputting the second sample image, the fourth sample image and the sample human body key points into an initial image processing model to obtain a predicted image;
[0015] The initial image processing model is modified according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
[0016] A third embodiment of the present disclosure provides an image processing device, including:
[0017] A first acquisition module is configured to acquire a first image containing a target person, target body key points corresponding to the target person's body parts to be changed in the first image, and a second image containing target clothing;
[0018] a masking module, configured to mask the part of the target person to be changed corresponding to the part to be changed in the first image, so as to obtain a third image including a target mask corresponding to the part to be changed;
[0019] a second acquisition module, configured to place the second image within a target mask of the third image to acquire a fourth image;
[0020] The third acquisition module is used to input the second image, the fourth image and the target human key points into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
[0021] A fourth embodiment of the present disclosure provides a training device for an image processing model, comprising:
[0022] A first acquisition module is configured to acquire a first sample image containing a sample person and sample body key points corresponding to the part of the sample person to be dressed up in the first sample image;
[0023] a segmentation module, configured to segment the first sample image to obtain a second sample image corresponding to the clothing of the sample person;
[0024] a masking module, configured to mask the part to be changed corresponding to the sample person in the first sample image, so as to obtain a third sample image including a sample mask corresponding to the part to be changed;
[0025] a second acquisition module, configured to place the second sample image within a sample mask in the third sample image to acquire a fourth sample image;
[0026] a third acquisition module, configured to input the second sample image, the fourth sample image, and the sample human body key points into an initial image processing model to acquire a predicted image;
[0027] A correction module is used to correct the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
[0028] The fifth aspect embodiment of the present disclosure proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the image processing method proposed in the first aspect embodiment of the present disclosure, or implements the image processing model training method proposed in the second aspect embodiment of the present disclosure.
[0029] The sixth aspect embodiment of the present disclosure proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the image processing method proposed in the first aspect embodiment of the present disclosure, or implements the image processing model training method proposed in the second aspect embodiment of the present disclosure.
[0030] The image processing method, image processing model training method, and device provided by the present disclosure have the following beneficial effects:
[0031] In the disclosed embodiment, a first image containing a target person, target human body key points corresponding to the target person in the first image, and a second image containing target clothing are obtained, the part to be changed corresponding to the target person in the first image is masked to obtain a third image containing the target mask corresponding to the part to be changed, and then the second image is placed in the target mask of the third image to obtain a fourth image. Finally, the second image, the fourth image, and the target human body key points are input into the trained and generated image processing model to obtain a target image, wherein the target person in the target image is wearing target clothing. In this way, the part to be changed in the first image containing the target person can be masked, and the target clothing can be placed in the masked area of the first image to obtain the fourth image. Not only does it not need to separate the person from the image, thereby improving the efficiency of changing clothes, but also in the process of processing the fourth image by the image processing model, in order to retain the detailed features of other areas except the part to be changed, the detailed features of the target clothing can be better retained, thereby improving the changing effect of the target image.
[0032] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0034] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure;
[0035] Figure 2 A flowchart of an image processing method provided by an embodiment of the present disclosure;
[0036] Figure 3 A flowchart of a method for training an image processing model provided by one embodiment of the present disclosure;
[0037] Figure 4 A flowchart of a method for training an image processing model provided by another embodiment of the present disclosure;
[0038] Figure 5 A flowchart of a method for training an image processing model provided by another embodiment of the present disclosure;
[0039] Figure 6 A schematic structural diagram of an image processing device provided by an embodiment of the present disclosure;
[0040] Figure 7 A schematic diagram of the structure of a training device for an image processing model provided in one embodiment of the present disclosure;
[0041] Figure 8 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0042] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0043] The following describes the image processing method, image processing model training method and device of the embodiments of the present disclosure with reference to the accompanying drawings.
[0044] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the image processing method may include the following steps:
[0045] Step 101: Acquire a first image containing a target person, target body key points corresponding to the target person in the first image, and a second image containing target clothing.
[0046] The first image may include the upper body, lower body, or full body of the target person, but this disclosure does not limit this.
[0047] The part to be replaced can be the upper body, lower body, or whole body of the target person.
[0048] In some embodiments, the human body key points of the target person in the first image may be extracted to obtain the target human body key points. Alternatively, the human body key points may be extracted only for the part of the target person to be dressed up.
[0049] Among them, the key points of the human body include the neck, arms, shoulders, waist, waist, legs, ankles, etc.
[0050] The target clothing is the clothing to be changed into for the target person, that is, the clothing that the target person wears after changing into the clothing. The target clothing is different from the clothing worn by the target person in the first image.
[0051] Step 102 : Mask the part of the target person corresponding to the part to be replaced in the first image to obtain a third image including a target mask corresponding to the part to be replaced.
[0052] The target mask may be a polygon, a rectangle, etc., and the present disclosure does not specifically limit the shape of the mask.
[0053] In the disclosed embodiment, the color of the target mask is not specifically limited, but the color of the target mask needs to cover the clothes of the part to be changed. For example, the color of the target mask can be gray.
[0054] In some embodiments, a target bounding rectangle corresponding to a target human body key point of the part to be replaced in the first image may be determined first; then, pixels within the target bounding rectangle are set to a third pixel value to obtain a third image containing a target mask.
[0055] The area where the target circumscribed rectangle is located is the area where the target mask is located.
[0056] For example, if you need to obtain a rectangular box for the target person's upper body, you can calculate the circumscribed rectangle based on the key points of the neck, arms, shoulders, and waist. If you need to obtain a rectangular box for the target person's lower body, you can calculate the circumscribed rectangle based on the key points of the waist, legs, and ankles.
[0057] The third pixel value may be 128 or the like.
[0058] In some embodiments, the target mask may also be manually defined, which is not limited in this disclosure.
[0059] Step 103: Place the second image within the target mask of the third image to obtain a fourth image.
[0060] In some embodiments, the second image needs to be on top of the target mask. There is no limitation on the size of the second sample image within the target mask.
[0061] It should be noted that because the sample image contains features other than the part to be changed, the image processing model preserves these detailed features during processing of the fourth image. Therefore, by placing the second image within the sample mask in the third image, the detailed features of the clothing, such as the number and shape of petals, can be preserved. The second image can then be used to extract macroscopic features of the clothing, such as its overall structure, shape, and layout.
[0062] In step 104 , the second image, the fourth image, and the key points of the target person are input into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
[0063] In some embodiments, the image processing model can be trained in the following manner: first obtain a first sample image containing a sample person and sample human body key points corresponding to the sample person in the first sample image, then segment the first sample image to obtain a second sample image corresponding to the clothing of the sample person's part to be changed, cover the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image containing the sample mask corresponding to the part to be changed, place the second sample image in the sample mask in the third sample image to obtain a fourth sample image, input the second sample image, the fourth sample image and the sample human body key points into the initial image processing model to obtain a predicted image, and finally, based on the difference between the predicted image and the first sample image, correct the initial image processing model to obtain a trained image processing model.
[0064] In some embodiments, the image processing model can also be trained in the following manner: first obtain a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image; mask the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image containing a sample mask corresponding to the part to be changed; obtain a clothing image corresponding to the sample clothing to be changed into by the sample person, and a label image of the sample person wearing the sample clothing, place the clothing image in the sample mask in the third sample image to obtain a fourth sample image, input the clothing image, the fourth sample image and the sample body key points into the initial image processing model to obtain a predicted image, and finally, according to the difference between the predicted image and the label image, correct the initial image processing model to obtain a trained image processing model.
[0065] The initial image processing model may be an untrained model or a model that requires fine-tuning. For example, the initial image processing model may be an untrained stable diffusion model or other diffusion model, or a trained stable diffusion model. This disclosure does not limit this.
[0066] In the disclosed embodiment, a first image containing a target person, target human body key points corresponding to the target person in the first image, and a second image containing target clothing are obtained, the part to be changed corresponding to the target person in the first image is masked to obtain a third image containing the target mask corresponding to the part to be changed, and then the second image is placed in the target mask of the third image to obtain a fourth image. Finally, the second image, the fourth image, and the target human body key points are input into the trained and generated image processing model to obtain a target image, wherein the target person in the target image is wearing target clothing. In this way, the part to be changed in the first image containing the target person can be masked, and the target clothing can be placed in the masked area of the first image to obtain the fourth image. Not only does it not need to separate the person from the image, thereby improving the efficiency of changing clothes, but also in the process of processing the fourth image by the image processing model, in order to retain the detailed features of other areas except the part to be changed, the detailed features of the target clothing can be better retained, thereby improving the changing effect of the target image.
[0067] Figure 2 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the image processing method may include the following steps:
[0068] Step 201 : Acquire a first image containing a target person, target body key points corresponding to the target person in the first image, and a second image containing target clothing.
[0069] Step 202 : Mask the part of the target person corresponding to the part to be replaced in the first image to obtain a third image including a target mask corresponding to the part to be replaced.
[0070] Step 203: Place the second image within the target mask of the third image to obtain a fourth image.
[0071] The specific implementation of step 201 and step 203 can refer to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.
[0072] In step 204 , the second image, the fourth image, and the target human key points are respectively input into a feature extraction network in the image processing model to obtain a first feature corresponding to the second image, a second feature corresponding to the fourth image, and a third feature corresponding to the target human key points.
[0073] In some embodiments, the image processing model may include a feature extraction network, a denoising network, a decoder, etc.
[0074] The feature extraction network is used to extract features from the second image, the fourth image, and the target human body key points. In some embodiments, the feature extraction network can be a multimodal feature extraction network, so that feature extraction can be performed on the second image, the fourth image, and the target human body key points respectively.
[0075] In some embodiments, the feature extraction network may also include feature extraction sub-networks corresponding to the second image, the fourth image, and the target body key points. For example, the first feature extraction sub-network is used to extract features from the second image, the second feature extraction sub-network is used to extract features from the fourth image, and the third feature extraction sub-network is used to extract features from the target body key points.
[0076] In some embodiments, the feature extraction subnetwork corresponding to the second image may include an image feature extraction network and a feature mapping network, the feature extraction subnetwork corresponding to the fourth image may be an encoder, and the feature extraction subnetwork corresponding to the target human body key points may be a control network.
[0077] In some embodiments, if the image processing network is obtained by fine-tuning an already trained Stable Diffusion model, to improve the fine-tuning efficiency of the trained Stable Diffusion model, only the upsampling network of the denoising network in the Stable Diffusion model can be trained, without changing the parameters of the downsampling network. However, since one of the input data of the trained downsampling network is the language features corresponding to the text description of the clothing to be changed, when the text description of the clothing is changed to a picture of the clothing (i.e., the second image), the features of the second image need to be converted to match the input data type of the downsampling network.
[0078] Therefore, in some embodiments, the second image can be input into the image feature extraction network in the feature extraction network to obtain the initial features corresponding to the second image; the initial features corresponding to the second image can be input into the feature mapping network in the feature extraction network to obtain the first features, wherein the dimension of the first feature is the preset dimension corresponding to the downsampling network.
[0079] The image feature extraction network may be an encoder of a trained contrastive language-image pre-training (CLIP) network.
[0080] The feature mapping network may be a trained feature conversion network, configured to convert the dimensions of the initial features corresponding to the second image into a preset dimension corresponding to the downsampling network. The preset dimension corresponding to the downsampling network may be a language feature dimension.
[0081] In some embodiments, the fourth image can be input into an encoder in a feature extraction network to obtain the second features.
[0082] The encoder may be a variational autoencoder (VAE).
[0083] In some embodiments, the target human body key points can be input into the control network in the feature extraction network to obtain the third feature.
[0084] The third feature is used to control the posture of the target person in the generated target image so that the posture of the target person in the generated target image is consistent with the posture of the target person in the first image.
[0085] Step 205 : Generate a target mask image based on the dimension of the second feature and the fourth image.
[0086] In some embodiments, the pixel points within the target mask in the fourth image can be set to the first pixel value, and the other pixel points outside the target mask can be set to the second pixel value to obtain a fifth image, and then the fifth image can be scaled based on the dimension of the second feature to obtain the target mask image.
[0087] The first pixel value may be 1, ie, white, and the second pixel value may be 0, ie, black.
[0088] The dimension of the sample mask image is the same as the dimension of the second prediction feature.
[0089] In step 206, the first feature, the second feature, the third feature, the target mask image, and the random noise image are input into a denoising network in the image processing model to obtain the target feature.
[0090] In some embodiments, the size of the random noise image is the same as the size of the second feature.
[0091] In some embodiments, the first feature, the second feature, the target mask image and the random noise image are input into the downsampling network in the denoising network, and the first feature and the third feature are input into the upsampling network in the denoising network to obtain the denoised image output by the upsampling network. The random noise image is then updated based on the denoised image, and the step of obtaining the denoised image is returned to execute until the preset number of denoising times is reached, and the denoised image finally output by the upsampling network is determined as the target feature.
[0092] The preset number of noise reduction operations may be a preset number of noise reduction iterations, for example, the preset number of noise reduction operations may be 10 times, 100 times, etc. This disclosure does not limit this.
[0093] Specifically, after the random noise image is updated based on the denoised image, the first feature, the second feature, the target mask image and the updated noise image can be input into the downsampling network in the denoising network, and the first feature and the third feature can be input into the upsampling network in the denoising network to obtain the denoised image output again by the upsampling network, until the number of cycles reaches the preset denoising number.
[0094] Step 207: Input the target features into a decoder in the image processing model to obtain a target image.
[0095] In some embodiments, the decoder may be a VAE decoder corresponding to the VAE encoder. The decoder is used to decode the encoded features to obtain a clear target image.
[0096] In an embodiment of the present disclosure, a first image including a target person, target human key points corresponding to the target person in the first image, and a second image including target clothing can be first obtained, and the part to be changed corresponding to the target person in the first image can be masked to obtain a third image including a target mask corresponding to the part to be changed. The second image is then placed in the target mask of the third image to obtain a fourth image. The second image, the fourth image and the target human key points are respectively input into a feature extraction network in an image processing model to obtain a first feature corresponding to the second image, a second feature corresponding to the fourth image, and a third feature corresponding to the target human key points. Based on the dimension of the second feature and the fourth image, a target mask image is generated. The first feature, the second feature, the third feature, the target mask image and the random noise image are input into a denoising network in the image processing model to obtain target features. Finally, the target features are input into a decoder in the image processing model to obtain a target image. Therefore, based on the target mask image with the same dimension corresponding to the second feature, the denoising network only processes the part to be changed, and does not process other areas outside the part to be changed, thereby ensuring the consistency of other parts with the first sample image, which not only improves the dressing efficiency, but also further improves the dressing effect of the target image.
[0097] Figure 3 A flowchart of a method for training an image processing model provided in an embodiment of the present disclosure.
[0098] The embodiment of the present disclosure uses the example of the training method of the image processing model being configured in the training device of the image processing model. The training device of the image processing model can be applied to any electronic device so that the electronic device can perform the training function of the image processing model.
[0099] like Figure 3 As shown, the image processing method may include the following steps:
[0100] Step 301: Obtain a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image.
[0101] The first sample image may include the upper body, lower body, or full body of the sample person, but this disclosure does not limit this.
[0102] The part to be changed can be the upper body, lower body, or whole body of the sample character.
[0103] In some embodiments, the human body key points of the sample person in the first sample image may be extracted to obtain the sample human body key points. Alternatively, the human body key points may be extracted only on the part of the sample person to be dressed up.
[0104] Among them, the key points of the human body include the neck, arms, shoulders, waist, waist, legs, ankles, etc.
[0105] Step 302 : segment the first sample image to obtain a second sample image corresponding to the clothing of the part of the sample person to be changed.
[0106] In some embodiments, a clothing segmentation algorithm may be used to segment a second sample image corresponding to the clothing from the first sample image.
[0107] In some embodiments, elastic deformation may be performed on the clothes in the second sample image, that is, the posture of the clothes may be changed, thereby achieving data augmentation for the second sample image.
[0108] In the disclosed embodiment, the first sample image can be directly segmented to obtain the second sample image corresponding to the clothes, thereby eliminating the need to obtain images of the sample person after changing clothes as label data. The first sample image can be directly used as a label to train the model, thereby not only reducing the amount of model training data but also reducing the difficulty of obtaining training data.
[0109] Step 303 : Mask the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image including a sample mask corresponding to the part to be changed.
[0110] The sample mask may be a polygon, a rectangle, etc., and the present disclosure does not specifically limit the shape of the sample mask.
[0111] In the disclosed embodiment, the color of the sample mask is not specifically limited, but it is required to cover the clothes of the part to be changed. For example, the color of the sample mask can be gray.
[0112] In some embodiments, a bounding rectangle corresponding to key points of a sample human body of a part to be replaced in the first sample image is determined, and then pixels within the bounding rectangle are set to a third pixel value to obtain a third sample image including a sample mask.
[0113] For example, if you need to obtain a rectangular box for the upper body of a sample person, you can calculate the circumscribed rectangle based on the key points of the neck, arms, shoulders, and waist. If you need to obtain a rectangular box for the lower body of a sample person, you can calculate the circumscribed rectangle based on the key points of the waist, legs, and ankles.
[0114] The third pixel value may be 128 or the like.
[0115] In some embodiments, the sample mask may also be manually defined, which is not limited in this disclosure.
[0116] In some embodiments, data augmentation may be performed by randomly enlarging the sample mask.
[0117] Step 304: Place the second sample image within the sample mask in the third sample image to obtain a fourth sample image.
[0118] In some embodiments, the second sample image needs to be on an upper layer of the sample mask. There is no limitation on the size of the second sample image within the sample mask.
[0119] It should be noted that because the fourth sample image contains features of the sample person other than the part to be changed, the image processing model retains specific detailed features during processing of the fourth sample image. Therefore, placing the second sample image within the sample mask in the third sample image can preserve detailed features of the clothing, such as the number and shape of petals on the clothing. The second sample image can then be used to extract macro features of the clothing, such as its overall structure, shape, and layout.
[0120] Step 305 : Input the second sample image, the fourth sample image, and the sample human body key points into the initial image processing model to obtain a predicted image.
[0121] The initial image processing model may be an untrained model or a model that requires fine-tuning. For example, the initial image processing model may be an untrained stable diffusion model or other diffusion model, or a trained stable diffusion model. This disclosure does not limit this.
[0122] Step 306: Modify the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
[0123] In some embodiments, a mean square error loss function may be used to calculate the difference between the predicted image and the first sample image, thereby modifying the initial image processing model.
[0124] In some embodiments, the initial image processing model can be iteratively trained based on multiple sets of training data until a training termination condition is met, thereby obtaining a trained image processing model. The training termination condition can be a predetermined number of training cycles, a model loss less than a predetermined threshold, or other conditions. The training termination condition can be set based on actual needs and is not limited in this application.
[0125] In the embodiment of the present disclosure, a first sample image including a sample person and sample human body key points corresponding to the sample person in the first sample image are obtained, and then the first sample image is segmented to obtain a second sample image corresponding to the clothing of the sample person's part to be changed, the part to be changed corresponding to the sample person in the first sample image is masked to obtain a third sample image including the sample mask corresponding to the part to be changed, and the second sample image is placed in the sample mask in the third sample image to obtain a fourth sample image, and then the second sample image, the fourth sample image and the sample human body key points are input into the initial image processing model to obtain a predicted image, and finally, according to the difference between the predicted image and the first sample image, the initial image processing model is corrected to obtain a trained image processing model. Therefore, the first sample image is directly segmented to obtain the second sample image corresponding to the clothes, so there is no need to obtain the image of the sample person after changing clothes as label data. The first sample image can be directly used as the label to train the model, which reduces the difficulty of obtaining training data and the amount of data for model training. Then, based on the second sample image, the fourth sample image, the key points of the sample body and the first sample image, the initial image processing model is trained, which improves the accuracy and reliability of the image processing model and provides conditions for improving the image changing effect.
[0126] Figure 4 A flowchart of a training method for an image processing model provided by an embodiment of the present disclosure is shown as follows: Figure 4 As shown, the training method of the image processing model may include the following steps:
[0127] Step 401: Obtain a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image.
[0128] Step 402 : segment the first sample image to obtain a second sample image corresponding to the clothing of the sample person.
[0129] Step 403 : Mask the part of the sample person in the first sample image that is to be replaced, to obtain a third sample image including a sample mask corresponding to the part to be replaced.
[0130] Step 404: Place the second sample image within the sample mask in the third sample image to obtain a fourth sample image.
[0131] The specific implementation of steps 401 to 404 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.
[0132] In step 405, the second sample image, the fourth sample image, and the sample human body key points are respectively input into the initial feature extraction network in the initial image processing model to obtain the first prediction feature corresponding to the second sample image, the second prediction feature corresponding to the fourth sample image, and the third prediction feature corresponding to the sample human body key points.
[0133] In some embodiments, the initial feature extraction network may be a multimodal feature extraction network, so that features can be extracted from the second sample image, the fourth sample image, and the sample human body key points respectively.
[0134] Alternatively, the initial feature extraction network may also include feature extraction subnetworks corresponding to the second sample image, the fourth sample image, and the sample body key points, respectively. For example, the first feature extraction subnetwork may be used to extract features from the second sample image, the second feature extraction subnetwork may be used to extract features from the fourth sample image, and the third feature extraction subnetwork may be used to extract features from the sample body key points.
[0135] In some embodiments, if the initial image processing model is a trained stable diffusion model (StableDiffusion), the initial image processing model includes a denoising network (UNET), and the denoising network includes a downsampling network and an upsampling network. In order to improve the fine-tuning efficiency of the trained stable diffusion model, only the upsampling network can be trained. However, since one of the input data of the trained downsampling network is the language feature corresponding to the text description information of the target clothes, when the text description information of the clothes becomes a picture of the clothes (i.e., the second sample image), the features of the second sample image need to be converted to conform to the input data type of the downsampling network. Therefore, in the embodiment of the present disclosure, the second sample image needs to be input into the image feature extraction network in the initial feature extraction network to obtain the initial features corresponding to the second sample image, and then the initial features corresponding to the second sample image are input into the initial feature mapping network in the initial feature extraction network to obtain the first prediction feature, wherein the dimension of the first prediction feature is the preset dimension corresponding to the downsampling network.
[0136] The image feature extraction network may be an encoder of a trained contrastive language-image pre-training (CLIP) network.
[0137] The initial feature mapping network may be an untrained feature conversion network, configured to convert the dimensions of the initial features into a preset dimension corresponding to the downsampling network, wherein the preset dimension corresponding to the downsampling network may be a language feature dimension.
[0138] In some embodiments, the initial feature mapping network may be composed of one or several fully connected layers.
[0139] Among them, the image feature extraction network and the initial feature mapping network constitute the above-mentioned first feature extraction subnetwork.
[0140] In some embodiments, if the initial image processing model is a stable diffusion model that has been trained, the initial image processing model also includes a variational autoencoder (VAE) and a VAE decoder. Since the VAE encoder and the VAE decoder can encode and decode the input image more accurately, the autoencoder (VAE) encoder and the VAE decoder may not be fine-tuned. Therefore, in the embodiment of the present disclosure, the fourth sample image may also be input into the encoder in the initial feature extraction network to obtain the second prediction feature. Among them, the encoder in the initial feature extraction network may be a VAE encoder that has been trained.
[0141] Among them, the encoder in the initial feature extraction network can be understood as the above-mentioned second feature extraction subnetwork.
[0142] In some embodiments, the sample human body key points may be input into an initial control network (control net) in the initial feature extraction network to obtain a third prediction feature.
[0143] The third prediction feature is used to control the posture of the sample person in the generated prediction image, so that the posture of the sample person in the generated prediction image is consistent with the posture of the sample person in the first sample image.
[0144] The initial control network can be understood as the third feature extraction subnetwork mentioned above.
[0145] Step 406: Generate a sample mask image based on the dimension corresponding to the second prediction feature and the fourth sample image.
[0146] In some embodiments, the pixel points within the sample mask in the fourth sample image can be set to the first pixel value, and the other pixel points outside the sample mask can be set to the second pixel value to obtain a fifth sample image, and then the fifth sample image can be scaled based on the dimension of the second prediction feature to obtain a sample mask image.
[0147] The first pixel value may be 1, ie, white, and the second pixel value may be 0, ie, black.
[0148] The dimension of the sample mask image is the same as the dimension of the second prediction feature.
[0149] In step 407 , the first prediction feature, the second prediction feature, the third prediction feature, the sample mask image, and the random noise image are input into an initial denoising network in the initial image processing model to obtain a fourth prediction feature.
[0150] In some embodiments, the random noise image has the same size as the second prediction feature.
[0151] In some embodiments, the first prediction feature, the second prediction feature, the sample mask image and the random noise image can be input into the downsampling network in the initial denoising network, and the first prediction feature and the third prediction feature can be input into the initial upsampling network in the initial denoising network to obtain the predicted denoised image output by the initial upsampling network. The random noise image is then updated based on the predicted denoised image, and the step of obtaining the predicted denoised image is returned to execute until the preset denoising times is reached, and the predicted denoised image finally output by the initial upsampling network is determined as the fourth prediction feature.
[0152] The downsampling network can be understood as the downsampling network in the previously trained Stable Diffusion model and does not require further fine-tuning. The initial upsampling network can be the downsampling network in the previously trained Stable Diffusion model and does not require further fine-tuning.
[0153] The preset number of noise reduction operations may be a preset number of noise reduction iterations, for example, the preset number of noise reduction operations may be 10 times, 100 times, etc. This disclosure does not limit this.
[0154] Specifically, after the random noise image is updated based on the predicted denoising image, the first predicted feature, the second predicted feature, the sample mask image and the updated noise image can be input into the downsampling network in the initial denoising network, and the first predicted feature and the third predicted feature can be input into the initial upsampling network in the initial denoising network to obtain the predicted denoising image output again by the initial denoising network until the number of cycles reaches the preset denoising number.
[0155] Step 408: Input the fourth prediction feature into the decoder in the initial image processing model to obtain a predicted image.
[0156] The decoder may be a VAE decoder.
[0157] Step 409 : Modify the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
[0158] In some embodiments, each network structure in the initial image processing model may be modified based on the difference between the predicted image and the first sample image.
[0159] In some embodiments, if the initial image processing model is a stable diffusion model that has been trained, in order to improve training efficiency, the initial control network, the initial feature mapping network, and the initial downsampling network can be corrected only based on the difference between the predicted image and the first sample image, thereby improving training efficiency.
[0160] In some embodiments, if the initial image processing model is a stable diffusion model that has been trained, and only the initial control network, initial feature mapping network, and initial downsampling network are corrected, the first sample image can also be input into the encoder to obtain the fifth prediction feature corresponding to the first sample image. Then, based on the difference between the fourth prediction feature and the fifth prediction feature, the initial control network, initial feature mapping network, and initial downsampling network are corrected to obtain the trained image processing model.
[0161] Since the decoder does not require training and the amount of undecoded data is small, if the initial control network, the initial feature mapping network, and the initial downsampling network are corrected based on the difference between the fourth prediction feature and the fifth prediction feature, it can not only reduce the amount of calculation and improve the calculation efficiency, but also reduce the requirements for training equipment.
[0162] In an embodiment of the present disclosure, a first sample image including a sample person and sample human body key points corresponding to the sample person in the first sample image are first obtained, and then the first sample image is segmented to obtain a second sample image corresponding to the clothing of the part to be changed of the sample person, and the part to be changed corresponding to the sample person in the first sample image is masked to obtain a third sample image including a sample mask corresponding to the part to be changed, and the second sample image is placed in the sample mask in the third sample image to obtain a fourth sample image, and then the second sample image, the fourth sample image and the sample human body key points are respectively input into an initial feature extraction network in an initial image processing model to obtain a first prediction feature corresponding to the second sample image, a second prediction feature corresponding to the fourth sample image, and a third prediction feature corresponding to the sample human body key points, and a sample mask image is generated based on the dimension corresponding to the second prediction feature and the fourth sample image, and then the first prediction feature, the second prediction feature, the third prediction feature, the sample mask image and the random noise image are input into an initial denoising network in the initial image processing model to obtain a fourth prediction feature, and finally the fourth prediction feature is input into a decoder in the initial image processing model to obtain a predicted image. Therefore, based on the sample mask image with the same dimension corresponding to the second prediction feature, the initial denoising network can only process the parts to be changed, and not process other parts except the parts to be changed, thereby ensuring the consistency of other parts with the first sample image and improving the effect of the image processing model on changing the clothes of the characters in the image.
[0163] Figure 5 A flowchart of a training method for an image processing model provided by an embodiment of the present disclosure is shown as follows: Figure 5 As shown, this embodiment takes the fine-tuning of the trained stable diffusion model as an example to illustrate the training method of the image processing model, which specifically includes the following steps:
[0164] Step 501: Acquire a first sample image.
[0165] Step 502 : extracting key points of a sample person in the first sample image to obtain key points of the sample person.
[0166] Step 503: Determine the circumscribed rectangle corresponding to the key points of the sample human body of the part to be changed in the first sample image.
[0167] Step 504: Set the pixel points within the circumscribed rectangle to a third pixel value to obtain a third sample image.
[0168] Step 505 : segment the first sample image to obtain a second sample image corresponding to the clothing of the part of the sample person to be changed.
[0169] Step 506: Place the second sample image within the sample mask in the third sample image to obtain a fourth sample image.
[0170] Step 507: Input the second sample image into an image feature extraction network (CLIP IMAGE ENCODER) to obtain initial features corresponding to the second sample image.
[0171] Step 508: Input the initial features corresponding to the second sample image into an initial feature mapping network (IMAGEFEATURE PROJECTION) to obtain first prediction features.
[0172] Step 509: Input the fourth sample image into the VAE encoder (ENCODER) to obtain a second prediction feature.
[0173] In step 510 , the sample human body key points are input into an initial control network (CONTROL NET) to obtain a third prediction feature.
[0174] Step 511 : Set the pixel points within the sample mask in the fourth sample image to the first pixel value, and set the other pixel points outside the sample mask to the second pixel value, so as to obtain a fifth sample image.
[0175] Step 512: scale the fifth sample image based on the dimension of the second prediction feature to obtain a sample mask image.
[0176] Step 513: Generate a random noise image with the same dimension as the second prediction feature.
[0177] In step 514, the first predicted feature, the second predicted feature, the sample mask image, and the random noise image are input into the downsampling network in the initial denoising network, and the first predicted feature and the third predicted feature are input into the initial upsampling network in the initial denoising network (UNET) to obtain the predicted denoised image output by the initial upsampling network.
[0178] Step 515: Update the random noise image based on the predicted denoised image, and return to the step of obtaining the predicted denoised image until the preset number of denoising cycles is reached, and determine the predicted denoised image finally output by the initial upsampling network as the target prediction feature.
[0179] Step 516: Input the first sample image into the encoder to obtain a fifth prediction feature corresponding to the first sample image.
[0180] Step 517: According to the difference between the fourth prediction feature and the fifth prediction feature, the initial control network, the initial feature mapping network, and the initial downsampling network are modified to obtain a trained image processing model.
[0181] Figure 6 A schematic diagram of the structure of an image processing device provided in an embodiment of the present disclosure.
[0182] like Figure 6 As shown, the image processing device 600 may include:
[0183] The first acquisition module 601 is configured to acquire a first image containing a target person, target body key points corresponding to the target person's body parts to be changed in the first image, and a second image containing target clothing;
[0184] A masking module 602 is configured to mask the part of the target person to be replaced corresponding to the part to be replaced in the first image, so as to obtain a third image including a target mask corresponding to the part to be replaced;
[0185] A second acquisition module 603 is configured to place the second image within a target mask of the third image to acquire a fourth image;
[0186] The third acquisition module 604 is used to input the second image, the fourth image and the key points of the target human body into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
[0187] In the disclosed embodiment, a first image including a target person, target human body key points corresponding to the target person in the first image, and a second image including target clothing are first obtained, the part to be changed corresponding to the target person in the first image is masked to obtain a third image including the target mask corresponding to the part to be changed, and then the second image is placed in the target mask of the third image to obtain a fourth image, and finally the second image, the fourth image and the target human body key points are input into the trained and generated image processing model to obtain a target image, wherein the target person in the target image is wearing the target clothing. In this way, the part to be changed in the first image including the target person can be masked, and the target clothing can be placed in the masked area of the first image to obtain the fourth image. Not only does it not need to separate the person from the image, thereby improving the efficiency of changing clothes, but also in the process of processing the fourth image by the image processing model, in order to retain the detailed features of other areas except the part to be changed, the detailed features of the target clothing can be better retained, thereby improving the changing effect of the target image.
[0188] In order to implement the above embodiments, the present disclosure also provides an image processing device.
[0189] Figure 7 A schematic diagram of the structure of a training device for an image processing model provided in an embodiment of the present disclosure.
[0190] like Figure 7 As shown, the training device 700 of the image processing model may include:
[0191] A first acquisition module 701 is configured to acquire a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image;
[0192] A segmentation module 702 is configured to segment the first sample image to obtain a second sample image corresponding to the clothing of the sample person;
[0193] The masking module 703 is configured to mask the part of the sample person to be replaced corresponding to the part to be replaced in the first sample image, so as to obtain a third sample image including a sample mask corresponding to the part to be replaced;
[0194] A second acquisition module 704 is configured to place the second sample image within the sample mask in the third sample image to acquire a fourth sample image;
[0195] The third acquisition module 705 is used to input the second sample image, the fourth sample image and the sample human body key points into the initial image processing model to obtain a predicted image;
[0196] The correction module 706 is configured to correct the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
[0197] The functions and specific implementation principles of the above modules in the embodiments of the present disclosure can be referred to the above method embodiments and will not be repeated here.
[0198] The image processing device of the embodiment of the present disclosure first obtains a first sample image containing a sample person and sample human body key points corresponding to the sample person in the first sample image, then segments the first sample image to obtain a second sample image corresponding to the clothing of the sample person's part to be changed, masks the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image containing the sample mask corresponding to the part to be changed, places the second sample image in the sample mask in the third sample image to obtain a fourth sample image, and then inputs the second sample image, the fourth sample image and the sample human body key points into an initial image processing model to obtain a predicted image, and finally corrects the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model. Therefore, the first sample image is directly segmented to obtain the second sample image corresponding to the clothes, so there is no need to obtain the image of the sample person after changing clothes as label data. The first sample image can be directly used as the label to train the model, which reduces the difficulty of obtaining training data and the amount of data for model training. Then, based on the second sample image, the fourth sample image, the key points of the sample body and the first sample image, the initial image processing model is trained, which improves the accuracy and reliability of the image processing model and provides conditions for improving the image changing effect.
[0199] In order to implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the image processing model training method or image processing method proposed in the above embodiments of the present disclosure.
[0200] In order to implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the training method of the image processing model or the image processing method proposed in the above embodiments of the present disclosure.
[0201] Figure 8 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 8 The electronic device 12 shown is only an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0202] like Figure 8 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).
[0203] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of such architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.
[0204] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0205] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 8 Not shown, often called a "hard drive"). Although Figure 8 Although not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), a Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present disclosure.
[0206] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.
[0207] The electronic device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). This communication can occur via an input / output (I / O) interface 22. Furthermore, the electronic device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via the bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0208] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the above embodiments.
[0209] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.
[0210] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0211] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0212] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0213] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0214] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0215] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0216] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first image containing a target person, target body key points corresponding to the target person in the first image, and a second image containing target clothing; Masking the part of the target person corresponding to the part to be changed in the first image to obtain a third image including a target mask corresponding to the part to be changed; placing the second image within the target mask of the third image to obtain a fourth image; The second image, the fourth image and the target human key points are input into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
2. The method according to claim 1, characterized in that The step of inputting the second image, the fourth image, and the target human key points into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing, comprises: Inputting the second image, the fourth image, and the target human key point into a feature extraction network in the image processing model respectively to obtain a first feature corresponding to the second image, a second feature corresponding to the fourth image, and a third feature corresponding to the target human key point; generating a target mask image based on the dimension of the second feature and the fourth image; Inputting the first feature, the second feature, the third feature, the target mask image, and the random noise image into a denoising network in the image processing model to obtain a target feature; The target features are input into a decoder in the image processing model to obtain the target image.
3. The method according to claim 2, characterized in that Generating a target mask image based on the dimension of the second feature and the fourth image includes: Setting the pixel points within the target mask in the fourth image to the first pixel value and setting the other pixel points outside the target mask to the second pixel value to obtain a fifth image; The fifth image is scaled based on the dimension of the second feature to obtain the target mask image.
4. The method according to claim 2, characterized in that The step of inputting the first feature, the second feature, the third feature, the target mask image, and the random noise image into a denoising network in the image processing model to obtain the target feature includes: Inputting the first feature, the second feature, the target mask image, and the random noise image into a downsampling network in the denoising network, and inputting the first feature and the third feature into an upsampling network in the denoising network to obtain a denoised image output by the upsampling network; The random noise image is updated based on the denoised image, and the step of obtaining the denoised image is returned to be executed until a preset number of denoising times is reached, and the denoised image finally output by the upsampling network is determined as the target feature.
5. The method according to claim 4, characterized in that Inputting the second image, the fourth image, and the target human key point into a feature extraction network in the image processing model to obtain a first feature corresponding to the second image, a second feature corresponding to the fourth image, and a third feature corresponding to the target human key point, respectively, includes: Inputting the second image into an image feature extraction network in the feature extraction network to obtain initial features corresponding to the second image; Inputting the initial features corresponding to the second image into a feature mapping network in the feature extraction network to obtain the first features, wherein the dimension of the first features is a preset dimension corresponding to the downsampling network; Inputting the fourth image into an encoder in the feature extraction network to obtain the second feature; The target human body key points are input into a control network in the feature extraction network to obtain the third feature.
6. The method according to claim 1, characterized in that The step of masking the part of the target person to be replaced corresponding to the part to be replaced in the first image to obtain a third image including a target mask corresponding to the part to be replaced includes: Determine a target circumscribed rectangle corresponding to a target human body key point of the part to be changed in the first image; The pixel points within the target circumscribed rectangle are set to third pixel values to obtain the third image including the target mask.
7. A training method for an image processing model, characterized in that: The method comprises: Acquire a first sample image containing a sample person and sample body key points corresponding to the sample person in the first sample image; Segmenting the first sample image to obtain a second sample image corresponding to the clothing of the part of the sample person to be changed; Masking the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image including a sample mask corresponding to the part to be changed; placing the second sample image within the sample mask in the third sample image to obtain a fourth sample image; Inputting the second sample image, the fourth sample image and the sample human body key points into an initial image processing model to obtain a predicted image; The initial image processing model is modified according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
8. The method according to claim 7, characterized in that The step of inputting the second sample image, the fourth sample image, and the sample human body key points into an initial image processing model to obtain a predicted image includes: Inputting the second sample image, the fourth sample image, and the sample human body key points into the initial feature extraction network in the initial image processing model respectively to obtain a first prediction feature corresponding to the second sample image, a second prediction feature corresponding to the fourth sample image, and a third prediction feature corresponding to the sample human body key points; generating a sample mask image based on the dimension corresponding to the second prediction feature and the fourth sample image; Inputting the first prediction feature, the second prediction feature, the third prediction feature, the sample mask image, and the random noise image into an initial denoising network in the initial image processing model to obtain a fourth prediction feature; The fourth prediction feature is input into a decoder in the initial image processing model to obtain the predicted image.
9. The method according to claim 8, characterized in that The generating a sample mask image based on the dimension corresponding to the second prediction feature and the fourth sample image includes: Setting the pixel points within the sample mask in the fourth sample image to the first pixel value and setting the other pixel points outside the sample mask to the second pixel value to obtain a fifth sample image; The fifth sample image is scaled based on the dimension of the second prediction feature to obtain the sample mask image.
10. The method according to claim 8, characterized in that The step of inputting the first prediction feature, the second prediction feature, the third prediction feature, the sample mask image, and the random noise image into an initial denoising network in the initial image processing model to obtain a fourth prediction feature includes: Inputting the first prediction feature, the second prediction feature, the sample mask image, and the random noise image into a downsampling network in the initial denoising network, and inputting the first prediction feature and the third prediction feature into an initial upsampling network in the initial denoising network, so as to obtain a predicted denoised image output by the initial upsampling network; The random noise image is updated based on the predicted denoised image, and the step of obtaining the predicted denoised image is returned to execute until a preset number of denoising times is reached, and the predicted denoised image finally output by the initial upsampling network is determined as the fourth prediction feature.
11. The method according to claim 10, characterized in that The step of inputting the second sample image, the fourth sample image, and the sample human body key points into an initial feature extraction network in the initial image processing model to obtain a first prediction feature corresponding to the second sample image, a second prediction feature corresponding to the fourth sample image, and a third prediction feature corresponding to the sample human body key points includes: Inputting the second sample image into the image feature extraction network in the initial feature extraction network to obtain initial features corresponding to the second sample image; Inputting the initial features corresponding to the second sample image into the initial feature mapping network in the initial feature extraction network to obtain the first prediction features, wherein the dimension of the first prediction features is a preset dimension corresponding to the downsampling network; Inputting the fourth sample image into an encoder in the initial feature extraction network to obtain the second prediction feature; The sample human body key points are input into the initial control network in the initial feature extraction network to obtain the third prediction feature.
12. The method according to claim 11, characterized in that The modifying the initial image processing model to obtain a trained image processing model includes: Inputting the first sample image into the encoder to obtain a fifth prediction feature corresponding to the first sample image; According to the difference between the fourth prediction feature and the fifth prediction feature, the initial control network, the initial feature mapping network, and the initial downsampling network are modified to obtain the trained image processing model.
13. The method according to claim 7, characterized in that The step of masking the part to be changed corresponding to the sample person in the first sample image to obtain a third sample image containing a sample mask corresponding to the part to be changed includes: Determine the circumscribed rectangle corresponding to the key points of the sample human body of the part to be changed in the first sample image; The pixel points within the circumscribed rectangle are set to third pixel values to obtain the third sample image including the sample mask.
14. An image processing device, characterized in that: The device comprises: A first acquisition module is configured to acquire a first image containing a target person, target body key points corresponding to the target person's body parts to be changed in the first image, and a second image containing target clothing; a masking module, configured to mask the part of the target person to be changed corresponding to the part to be changed in the first image, so as to obtain a third image including a target mask corresponding to the part to be changed; a second acquisition module, configured to place the second image within a target mask of the third image to acquire a fourth image; The third acquisition module is used to input the second image, the fourth image and the target human key points into the trained image processing model to obtain a target image, wherein the target person in the target image wears the target clothing.
15. A training device for an image processing model, characterized in that: The device comprises: A first acquisition module is configured to acquire a first sample image containing a sample person and sample body key points corresponding to the part of the sample person to be dressed up in the first sample image; a segmentation module, configured to segment the first sample image to obtain a second sample image corresponding to the clothing of the sample person; a masking module, configured to mask the part to be changed corresponding to the sample person in the first sample image, so as to obtain a third sample image including a sample mask corresponding to the part to be changed; a second acquisition module, configured to place the second sample image within a sample mask in the third sample image to acquire a fourth sample image; a third acquisition module, configured to input the second sample image, the fourth sample image, and the sample human body key points into an initial image processing model to acquire a predicted image; A correction module is used to correct the initial image processing model according to the difference between the predicted image and the first sample image to obtain a trained image processing model.
16. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the image processing method as described in any one of claims 1 to 6 is implemented, or the training method of the image processing model as described in any one of claims 7 to 13 is implemented.
17. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the image processing method described in any one of claims 1 to 6, or implements the training method of the image processing model described in any one of claims 7 to 13.